lagrangeinterpolator, lagrangeinterpolator@awful.systems
Instance: awful.systems
Joined: 1 year ago
Posts: 0
Comments: 44
Posts and Comments by lagrangeinterpolator, lagrangeinterpolator@awful.systems
Posts by lagrangeinterpolator, lagrangeinterpolator@awful.systems
Comments by lagrangeinterpolator, lagrangeinterpolator@awful.systems
The OpenAI hacking incident has put everyone on edge. Anthropic, not to be outdone, decides to publicly announce that they have committed not one, not two, but three crimes! Come on! I was told this wasn’t a marketing campaign!
Before the days of AI 2027, he became famous in 2024 for posting one of the original pieces of writing in the line-go-up genre: Situational Awareness. It seems that like any good grifter, he used this opportunity to make money (in this case by starting a hedge fund).
There is something about this picture, a warehouse full of books to be destroyed by Anthropic for training data, that just makes me sick. Knowing that they are destroying books in the abstract is bad enough, but seeing the sheer scale of it makes me viscerally angry.

Karen Hao called her book Empire of AI for a very good reason. The empire has all these noble narratives of bringing civilization and progress to the colonies, but in reality the empire only sees resources to be extracted, repackaged, and sold back to their subjects at a fee. Kudos to those who saw this immediately and fought against it. Shame on those who believed in the narrative of progress, to the point where many of them are collaborators with the empire.
Whenever a smarmy fellow on Linkedin posts about how AI is making so much progress in coding or math, I want them to see this picture. I want them to see the videos of families in tears after a data center takes away their home with eminent domain, and other families who have to deal with the resulting pollution. Go beyond the bloodless math of tokens and instead see for yourself what this all truly costs. I hope it was worth it.
It gets worse. I found another article about the same podcast: Sam Altman says AI won’t shorten the workweek because humans secretly enjoy being busy.
“Technology, for a long time, has been promising people that they’re going to work less and they’re going to have all this leisure,” Altman said. “But somehow we never get the promise of the four-hour workweek at mass scale in society,” he added. “And I don’t expect AI to change that.”
“It’s like a relative game. People are very focused on how they’re doing relative to other people,” the CEO said.
“I think we’re all going to be much busier than we thought we were supposed to be in a post-superintelligence world. We’re still going to complain about it, but secretly we’re going to be happy,” Altman concluded.
The chuds are all at r/accelerate now. “We’re the unapologetically Pro-AI alternative to r/singularity, r/futurology, r/artificial & r/technology which are becoming filled with Luddites, Techno-Decelerationists, Techno-Reactionaries, & Anti-AIs.”
Clamuel did another podcast. I don’t want to waste my time watching it, but I want to sneer at some choice quotes from news articles breathlessly reporting this stuff as gospel 1, 2
“We are now, like, in the singularity”
Is the singularity in the room with us right now? Even the people on X the Everything (Including CSAM) App aren’t so convinced.
Altman said on the podcast that just a decade ago, the so-called singularity still felt like a distant and improbable dream — something he and his colleagues would discuss casually over lunch.
“Now we’re actually in the moment that we used to talk about at the lunch table in a very not-serious way,” he said. “I’ve been waiting for this my whole life, and I think it’s going to be incredible, hugely positive, awesome for the world.”
Do you feel it? Do you feel incredible? Do you feel hugely positive? Do you feel awesome, now that the singularity is here? Oh wait, he’s speaking in the future tense again, never mind. Force of habit.
It’s also funny how he talks about how positive things are gonna be immediately after the super scary hack done on Hugging Face (really due to incompetence on both sides). Oh no, the AI is gonna kill us all! It’s going to be incredible, hugely positive, awesome for the world.
Later in the podcast, Altman criticizes AI leaders who have repeatedly warned that AI is dangerous. He didn’t name Anthropic, his primary competition, but its CEO, Dario Amodei, is well-known for making dire predictions about the future in his calls for greater attention to safety.
“I also think some of the alternative visions painted by other companies are quite terrifying,” Altman said. “I’m going to make sure that gets pushed against and is not what happens.”
At least Clammy realizes that all the doom trolling about how AI will replace and destroy all humans turns out, surprisingly, to lead to overwhelming negative public opinion. But this statement truly reassures me that Scammy has good intentions about AI safety. He will make sure this is pushed against and is not what happens. He wouldn’t want another super scary, dangerous hack that just so conveniently turns out to be excellent marketing.
OpenAI CEO Sam Altman says one of his biggest concerns about AI isn’t whether it becomes too powerful in and of itself, but that a single company could control that power.
Yeah, it would be really bad for Anthropic to control all the power of AI. It would be a disaster to hand everything over to Anthropic.
“Every time that humanity has traded off its liberty for safety, it’s been a long-term net loss,” Altman said. “We are going to put this in the hands of people. We’re going to empower them. We are going to let society express its ideas and use this technology in the way they want.”
Society is expressing its ideas about AI, all right. But I feel like they aren’t what he has in mind.
Altman, meanwhile, said preventing such power from becoming overly concentrated has long been central to his vision for AI, even as his company continues to keep its own models closed.
“For so long now, I have felt focused on this singular goal of abundant intelligence and a belief that incredible human prosperity will come from that, as long as we don’t have a weird power concentration and kind of a new authoritarianism,” he said.
Well, we seem to have a weird power concentration and “kind of” a new authoritarianism, but I don’t see any abundant intelligence or incredible human prosperity.
Lots more sneers on the Better Offline subreddit.
It is surprising how many exceptionally strong mathematicians have started working for OpenAI and Anthropic. These people would have easily become professors at top universities if they stayed in academia. I think many mathematicians, especially the competitive ones at the top, have a “progress at any cost” attitude (and I’m sure the paychecks helped). As for the results, you still need good mathematicians to sift through all the output to identify that the proofs are valid.
I would honestly be positive about universities developing their own specialized math AI (in an ethical manner) to help mathematicians get these kinds of results, but right now, AI is inseparable from these evil companies. Thankfully, I believe this is a likely outcome in the future because the AI companies will one day implode.
From what I’ve seen, most prompts are in plain English. I suppose the part where the AI parses the statement correctly is much easier than the part where it boils a couple lakes in the process of bashing its head against the wall trying millions of different combinations of random shit from the literature to slap together a proof. For one of the big results (cycle double cover), the prompt specified that the AI could use 64 subagents and was required to not give up for at least 8 hours. The tokenmaxxers would be proud, we didn’t need that forest anyway. Thank god math doesn’t have a CTO to look at the expense reports.
Among the three big results I’ve looked at (unit distance problem, cycle double cover, Jacobian), two were counterexamples and one of them had a short 3 page proof using ideas from the 1970s. The Jacobian conjecture is an extreme case because a single counterexample is enough (for unit distance, you technically need a family of counterexamples), and it is easy to check with very basic computations. It is telling that all of these announcements came from OpenAI or Anthropic employees, who presumably have unlimited access to their AI. Nobody really knows how many resources they spent on this, or what else they tried. Nobody really seems to care about this question, either.
I think there is a phenomenon where supposedly hard questions are much easier than expected, because by chance nobody found the right approach for a while, and eventually it becomes famous as a “hard problem” which makes nobody want to attempt it.
What I’m more worried about is many people starting to use AI to try and prove small lemmas for them in their projects. Of course, a $200/mo subscription is absolutely necessary to them. This honestly feels like a repeat of Claude Code back in February. The software engineers eventually realized that AI is absurdly expensive after the AI companies realized that spending $14000/mo to service a $200/mo subscription is a bad idea. If the AI vendors couldn’t squeeze money out of rich software companies, what exactly are they gonna get out of poor mathematicians and universities? Also, there is the cognitive decline caused by overuse of LLMs that has yet to set in.
I think a serious possibility is that AI generated papers flood the zone with uninteresting incremental results that are eventually meaningless and full of mistakes. Right now, math is full of smart, dedicated people, so at least major results are reviewed carefully. But as AI alarmism drives away many honest people from the field, the remaining mathematicians will be burdened with far more work to review, and their cognitive faculties will be eroded by LLM use. Despite 4 years of development, $3 trillion of debt, mountains of stolen data, all the agents and harnesses and loops and other expensive tricks, as well as the advantages of Lean in math research, LLMs still hallucinate.
I believe this is happening with software, but at least there are objective consequences for screwing up there (guy gets his home directory deleted, email is sent on a guy’s behalf without permission, small business gets every customer subscription cancelled). But nothing bad happens if there is a mathematical mistake in a paper and nobody catches it. One could say to just provide a Lean proof, but there is still the issue of making sure the Lean code actually matches the content of the paper. Exactly what force will correct things?
Still, I don’t think this is the most likely possibility. The AI companies are extremely unsustainable financially, and it’s not like they’re very popular. Once they collapse, I believe there will be a re-evaluation of how LLMs should be used in research. If they are used (let alone trained), someone is going to have to pay the bills.
In the end, we have to ask ourselves the question of why one does math. To me, math is not really a field where you memorize trivia. The real value comes from being able to think abstractly and rigorously from first principles, and from understanding why something is true rather than just knowing it is true. It is another aspect of your ability to reason as a free human. A few dedicated people go into math research, but your skills can easily go to many places. If you’re starting undergrad, you have plenty of time to see how this all pans out before making a decision.
long rant about math
The recent big AI results in math have left me in quite a bad mood. I believe the main ingredient is Lean, which is a formal language resembling a programming language. Math proofs written in Lean can be verified deterministically with a computer, which really helps mitigate the hallucination problems of LLMs. Back in the days of pure scaling LLMs and Sam Altman talking about Dyson spheres, I was skeptical that LLMs would do math, but I did think that perhaps in the future, techniques using these formal languages could contribute to math. Well, it seems like OpenAI and Anthropic had the same idea and I underestimated their limitless checkbooks. Many of the biggest results were announced by mathematicians directly working for them (and presumably being paid a handsome amount).
For what it’s worth, after the last of these big announcements, I decided to try one of these AIs on one of my small problems that I couldn’t figure out. The AI did give a solution. That is, until I checked it thoroughly and realized that the it had a subtle but severe mistake that made it useless. I reprompted it, it failed again, and I ran out of tokens. I’m sure someone will tell me to shell out $200/mo for a pro subscription.
In the math and computer science research community, this is all anyone can really talk about right now. Honestly, after watching this whole AI bubble starting from the very beginning, I think the AI companies want to use marketing to stoke fear that all mathematicians will be replaced. But now, I am just too tired to argue. The amount of alarm and the extraordinary social pressure to use LLMs has soured me to this whole research thing. If becoming a researcher will one day require supporting these evil AI companies, I would rather just not. My dream job now is Factorio developer.
A lot of annoying people in technical areas view the world in terms of an intelligence hierarchy: the smartest people do math and physics, the slightly less smart people do coding, and the dumb people do everything else. So if AI can do math then it can do anything else. But, as an example, it is abundantly obvious now that AI is not replacing filmmaking. The techbros might be moved by arguments about how hilariously expensive video generation is, and how all these videos are 2 second clips stitched together so you won’t feel the uncanny valley. But the real reason is that nobody wants to watch slop made with no intention or feeling. Also, nobody wants to support the AI companies, which could not act more evil even if they tried.
The mania in math right now quite resembles the mania in software engineering back in December-February, when Claude Code definitely solved all coding. I don’t think the boosters expected that by April, everyone would be complaining about how expensive it all was while seeing an endless parade of vibe coding disasters (and no increase in productivity). Even if math research works out perfectly well (which is a still big if), it’s not going to pay the bills. They would need to find a use case in the real world, where hallucinations can cause serious damage and cannot be formally prevented. And they have certainly tried. Math will not change the fact that all of this will collapse.
Here’s A Taxonomy of Omnicidal Futures Involving Artificial Intelligence by Critch and Tsimerman, which approvingly cites our favorite AI 2027. Nobel disease is still a thing (I think the rationalist beliefs are quite pervasive among mathematicians).
I find it really funny how after he gets booed he says, “If you don’t care about science, that’s okay, because AI is going to touch everything else as well. Whatever path you choose, AI will become part of how work is done.” Yeah, if you’re worried that AI is only going to fuck up science, don’t worry, it’s going to fuck up everything else as well. Was he trying to stick to a (terrible) script, or is he genuinely this incapable of reading a room?
“When someone offers you a seat on the rocket ship, you do not ask which seat. You just get on.” No, my mom taught me about stranger danger. I know what to do when a sketchy old man named Eric Schmidt pulls up with a rocket ship that says FREE ICE CREAM.
“The rocket ship is here. Let me give you some advice. First, find a way to say yes. Listen.” Thanks for revealing how AI adoption is really about coercion. It doesn’t matter what you think, AI is inevitable and you ignorant Luddites are gonna have to find a way to like it.
Truly a masterclass in public speaking by Eric Schmidt. When the audience reacts negatively to what you said, just double down and shove it down their throats. You’re a billionaire, so you know better than them.
How many people, if they were given $1.3 million just once in their lifetime, would figure out far better uses for that money than this guy?
The last several years have been the monkey’s paw moment for rationalists, where they keep getting what they want and realizing it’s actually bad. As for why they keep getting what they want, just look at who’s funding them.
(Also featuring a “Chinese curse” that isn’t actually a phrase in Chinese. At least it’s not “may you live in interesting times”.)
I attended a town hall hosted by the department at my university supposedly for general discussion about department affairs. Considering the university had recently made moves such as adding “AI” into the very name of the department, I had suspicions that much of the discussion would be about AI. (I realize I’m doxxing myself but whatever.) I mostly came for the free food, but I was also interested in seeing what people thought about AI.
The event started with a talk by a prominent professor with major administrative power in the department, and indeed the talk was mostly about AI. His views were that he personally didn’t like AI, but he believed that it had changed the world (particularly in programming), and that it was going to stay. One of his justifications for pivoting the department to AI was ensuring universities had some say in AI and not letting all the control go to unaccountable corporations.
The reaction from the audience was a pleasant surprise to me. He asked everyone how much they were excited about AI (hardly anyone) and how much they were worried (most of the audience). By far the most amusing moment was when someone asked, “What if the assumption that AI is inevitable is wrong? What if AI does not live up to its promises?” (Sadly, I don’t remember the exact words that the person said.) The professor’s response was that by this point, there are so many trustworthy, smart, prominent people who definitely wouldn’t fall for scams, and they have adopted AI. He trusts those people, so he trusts that AI is genuine. I don’t know if the audience member accepted this explanation, but I hope not. Our modus operandi is FOMO.
The pizza was only ok, not really worth a 90 minute event.
This really goes to show how much they need to rely on the LLMentalist effect, despite the AI boosters insisting that the AI is totally different now, everything changed in the last few months. They do not care about creating a useful, reliable tool. That concept doesn’t even occur to them, since why do that when AI is magic?
In any case, they are incapable of creating a useful, reliable tool. Deep down, the only thing the AI companies have at their disposal is the ELIZA effect. OpenAI has every incentive not to truly eliminate AI psychosis, because they need engagement. They only want to mitigate the extreme cases where people go insane and cause bad PR for them. But mild AI psychosis is totally fine, it’s great when people are addicted to your product and make the numbers go up!
Somehow this is no worse than his usual fare, such as a thumbnail that is just a bunch of colored lines resembling a line chart but without representing any actual data, with some random marked points labeled “Dark Farms” and “Human Zoo”.
Unfortunately, our problem right now is not Donna the below-average Democrat but Donald the fascist. And when it comes to fascists I do not ask if they are above or below average.
The fire code thing really is an excellent example of LessWrong Brain. Fire truck drivers insist on needlessly large trucks (no citation) which makes roads 30% wider than they would otherwise be (no citation) which has “probably” “non-trivially” contributed to larger cars (no citation) leading to enough additional road fatalities to cancel out the lives saved by stricter fire codes (no citation).
The LessWrong Brain argument starts with a deliberately contrarian conclusion and proves it with a Rube Goldberg chain of logical syllogisms. Of course, citations are strictly optional, and they are free to misinterpret them as they see fit. The only real standard of each claim is “looks good to me”, but you are supposed to be impressed that they managed to string a dozen of them together to reveal some shocking, deep truth of the world that nobody else knows about. The AI 2027 nonsense is an infamous example of this.
He uses the word “fermi” which is cult jargon based on Fermi estimation, a.k.a. guessing shit with back-of-the-envelope calculations. Not exactly what you want if you want to convince people to reform fire codes, especially if you have zero citations for anything.
I guess people just aren’t rational enough, and the only reason the fire codes are so irrational is because people are emotional about fire codes. Firefighters are apparently revered as heroes, when it is the LWers who should be the heroes. After all, firefighters merely save people from fires, while LWers buy multimillion dollar mansions to talk about saving quadrillions of hypothetical people from hypothetical basilisks!
It’s fine, spyware is only a risk when it’s bad people’s spyware. It’s totally fine when it’s Anthropic™-approved spyware!
As for Mythos catching things, maybe they should have used Mythos on their very own Claude Code considering that it has hilariously obvious security exploits, such as this one which inserts an arbitrary string into a shell command. Actually, never mind I don’t see anything wrong here, maybe we should burn another $20k in electricity running Mythos on it again to find out.
RetroFed
In basically every case in history where people decided to kill a bad king, there was a period of chaos and violence that followed it. The killing of Charles I happened during the English Civil War, and the killing of Louis XVI happened during the French Revolution. This has happened many times in Chinese history, with the fall of an imperial dynasty leading to several decades of civil war (most recently in the early 1900s). But I guess if you have a big clever brain with big clever thoughts, you don’t need to look at history.
If the only way to get rid of a bad king is to kill him, he will do anything he can to defend his power, including using as much violence as necessary. (People generally do not like being killed.) Even if you successfully get rid of him, good luck establishing a proper government afterwards with all the violence you’ve caused. And who knows if the new king is gonna be better or worse? A better system would instead have a mechanism that replaces officials on a regular basis, say every few years, and ensure that these replacements are peaceful. Oh wait, that’s liberal democracy. If we do something boring like support democracy, how will people ever think of us as special, clever thinkers with bold, contrarian thoughts?
Bro, your system involves giving all the power to one person. You cannot then say they have no responsibility or that they’re “inoffensive” when they abuse it.
I’ve seen this story play out in software engineering: people were very impressed when the AI does unexpectedly well in one out of 50 attempts on an easy task, and so people decided to trust it for everything and turn their codebases into disasters. There was no great wave of new high-quality software. Instead, the only real result was that existing software has become far more buggy and insecure.
Now we have people using AI in science and math because it was impressive in random demonstrations of solving math problems. I now have friends asking me why I’m not using AI, and also saying that AI will be better than all mathematicians in 30 years or whatever. Do you really think I refuse to use AI out of ignorance? No, I know too much about it! I have seen the same story play out in software engineering, and what makes this any different?
I think the main difference here is that breaking RSA now just requires scaling up existing approaches, while breaking LWE or anything like that would need a major conceptual breakthrough. The former possibility is much more likely, and in any case, cryptographers are the most paranoid people on the planet for a reason.
Unfortunately, one can never be sure about much in cryptography until P vs NP is solved (and then some).
(Of course, just because some people say that scaling up is enough doesn’t mean it’s actually true. For breaking RSA, we know have Shor’s algorithm, while the only evidence AI bros have from superintelligence coming from scaling is “trust me bro”.)
This is what happens when your worldview is based on anime.
(A lot of anime has heavy themes, but most people understand that it’s not real life, just like all such art. Unlike Yud, most people’s worldviews on coding and math are based on actual coding and math.)
We can see that one 9 of availability is 90% = 0.9, two 9s is 99% = 0.99, three 9s is 99.9% = 0.999, etc. In general, for positive integers n, n 9s of availability is 1 - (1/10)^n, and we can extrapolate that to non-integer values of n. The value γ needed for 87.5% availability is the solution to 1 - (1/10)^γ = 7/8, or γ = log_10(8) = 0.903089987. γ is transcendental by Gelfond-Schneider (see this for a reference proof).
Right now, Sora is at zero 9s of availability.
By far the dumbest “feature” in the codebase is this thing called “Buddy” (described in a few places such as here). Honestly, I don’t really know what it’s for or what the point is.
Great, so they were planning on a gacha system where you can get an ASCII virtual pet that, uhh, occasionally makes comments? Truly a serious feature for a serious tool for the serious discipline of software engineering. Imagine if IntelliJ decided to pull this bullshit.
The Onion could not have come up with a better way to illustrate this very point.
Good luck telling the promptfondlers that LLMs are only useful for entertainment and not for any useful work.
I’m sure these English instructions work because they feel like they work. Look, these LLMs feel really great for coding. If they don’t work, that’s because you didn’t pay $200/month for the pro version and you didn’t put enough boldface and all-caps words in the prompt. Also, I really feel like these homeopathic sugar pills cured my cold. I got better after I started taking them!
No joke, I watched a talk once where some people used an LLM to model how certain users would behave in their scenario given their socioeconomic backgrounds. But they had a slight problem, which was that LLMs are nondeterministic and would of course often give different answers when prompted twice. Their solution was to literally use an automated tool that would try a bunch of different prompts until they happened to get one that would give consistent answers (at least on their dataset). I would call this the xkcd green jelly bean effect, but I guess if you call it “finetuning” then suddenly it sounds very proper and serious. (The cherry on top was that they never actually evaluated the output of the LLM, e.g. by seeing how consistent it was with actual user responses. They just had an LLM generate fiction and called it a day.)
AI seems good at purple prose and metaphors that don’t exactly make sense. No, I do not give a fuck about the “triangle of calm” when it comes to, of all things, the narrator taking off her shoes. No, I am not interested in how long the narrator sets the timer on the microwave when she makes literally the blandest meal of all time.
Now I’m sure the techbros truly think this is good “literary” writing. After all, they only care that the writing sounds flowery, because they seem to be very good at missing the actual meaning of everything. I remember Saltman saying that the movie Oppenheimer needed to be more optimistic to inspire more kids to become physicists (while also saying that The Social Network did that for startup founders).
The article’s entire premise is Musk saying some random shit. Remember how Musk said that he would land a man on Mars in 10 years 13 years ago? Honestly, I am incensed that people like Musk and Trump can just say shit and many people will just accept it. I can no longer tolerate it.
He says this after mentioning UBI. He really doesn’t want to confront the unfortunate fact that UBI is entirely a political issue. Whatever magical beliefs one may have about how AI can create wealth, the question of how to distribute it is a social arrangement. What exactly stops the wealthy from consolidating all that wealth for themselves? The goodness of their hearts? Or is it political pushback (and violence in the bad old days), as demonstrated in every single example we have in history?
I’d say the problem is even worse now. In previous eras, some wealthy people funded libraries and parks. Nowadays we see them donate to weirdo rationalist nonsense that is completely disconnected from reality.
This is followed by four whole paragraphs about how the office sucks and wouldn’t it be wonderful if AI got rid of all that. Guess what, we have remote work already! Remember how, during COVID, many software engineering jobs went fully remote, and it turned out that the work was perfectly doable and the workers’ lives improved? But then there were so many puff pieces by managers about the wonderful environment of the office, and back to the office they went. Don’t worry, when the magical AI is here, they’ll change their minds.
Yes, there are “mindless, stupid, inane things” like chores that are unavoidable. There are also other mindless, stupid, inane things that are entirely avoidable but exist anyway because some people base their entire lives around number go up.
I’d say that the great problems that last for decades do not fall purely to random bullshit and require serious advances in new concepts and understanding. But even then, the romanticized warrior culture view is inaccurate. It’s not like some big brain genius says “I’m gonna solve this problem” and comes up with big brain ideas that solve it. Instead, a big problem is solved after people make tons of incremental progress by trying random bullshit and then someone realizes that the tools are now good enough to solve the big problem. A better analogy than the Good Will Hunting genius is picking a fruit: you wait until it is ripe.
But math/CS research is not just about random bullshit go. The truly valuable part is theory and understanding, which comes from critically evaluating the results of whatever random bullshit one tries. Why did idea X work well with Y but not so well with Z, and where else could it work? So random bullshit go is a necessary part of the process, but I’d say research has value (and prestige) because of the theory that comes from people thinking about it critically. Needless to say, LLMs are useless at this. (In the Knuth example, the AI didn’t even prove that its construction worked.)
I think intelligence is overrated for research, and the most important quality for research is giving a shit. Solving big problems is mostly a question of having the right perspective and tools, and raw intelligence is not very useful without them. To do that, one needs to take time to develop opinions and feelings about the strengths and weaknesses of various tools.
Of course, every rule has exceptions, and there have been long standing problems that have been solved only when someone had the chutzpah to apply far more random bullshit than anyone had dared to try before.
The 31st try resulted in them only solving the problem for odd m, but the even m case was still open. So of course this happened:
Knuth did add a postscript on other friends maybe kinda vibing a possible solution for even m:
I find it slightly funny how Stappers suggested to the AI to use specific external tools that are actually reliable (like ORTools). This also makes me question how much the of the AI’s “insight” was a result of handholding and the rubber duck effect.
For context: 1. This is planned as a hard exercise for a textbook. 2. There are likely so many solutions that finding a general program that works (at least for enough values that you care to check) is like hitting the side of a barn with an arrow. Random bullshit go is an excellent strategy here. 3. The AIs did not provide proofs that their solutions worked. This is kind of a problem if you want to demonstrate that AI has understanding.
I can see them by clicking on the “Reasoned for 6m 34s” box on my desktop browser (I know, great UI design) without needing to log in. This did not work on my phone. In any case, I can copy it here:
Warning: very long slop
Exploring Hadamard matrix construction for order 428
For a Hadamard matrix of order 428, we need to confirm that it’s a multiple of 4, which it is (428 = 4 * 107). However, the challenge is that order 107 doesn’t work since it’s not a multiple of 4. A possible approach could involve using constructions like Williamson or Goethals-Seidel. Alternatively, I could try a Kronecker product of smaller matrices.
Examining constructions for Hadamard matrix of order 428
Since 107 ≡ 3 mod 4, Paley type I gives a Hadamard of order 108, but I need 428, not 432. That’s a little tricky. I wonder if there’s a known construction for order 4*107 using “Hadamard matrix of order 4q” or conference matrices. Could I adapt the generalized Paley approach for this?
Exploring conference matrix construction for Hadamard order 428
For q = 107, Paley construction doesn’t work since q ≡ 3 mod 4. But if I can create a conference matrix of order 214 (since 214 ≡ 2 mod 4), then a Hadamard matrix of order 428 would follow. Conference matrices require specific conditions, but this seems promising. Let’s focus on 214.
Searching for explicit Hadamard matrix of order 428
To generate a Hadamard matrix of order 428, it might be feasible since many such matrices exist up to large orders, but I need one explicitly with ±1 entries. A direct construction could involve methods like Goethals-Seidel, which uses circulant matrices to satisfy specific conditions. I’ll need to implement this method.
Exploring construction methods and searching for Hadamard matrix 428
For n=107, there might be a way to use Legendre sequences for Goethals-Seidel difference sets. I know there are constructions for order 4q where q is an odd prime power, but using approaches like Scarpis or Turyn may help too. I’ll search for an explicit Hadamard matrix of order 428 for a possible example or construction.
Exploring accessible sources and construction options for Hadamard matrix
It seems like there may be server timeouts, so I’ll need to try other sources like GitHub, arXiv, or design theory papers. We could construct the Hadamard matrix ourselves using known algorithms, but it’s heavy. If accessible, sequences from a Kharaghani paper could help, and SageMath might have relevant functions too.
Identifying base sequences to build Hadamard matrix
To create a Hadamard matrix of order 428, I think I can use base sequences of lengths 71 and 36, which combine to 107. I’ll look for Turyn-type sequences or T-sequences in accessible PDFs, potentially from open sources like arXiv. A useful file might be “H_428,” so I’ll search for it.
Checking GitHub for Hadamard matrix at 428 order
I found a GitHub project, “anymatrix,” which claims to have Hadamard matrices up to order 428. It might include the 428 matrix in MATLAB format. I’ll check the repository to see if it has the actual file with the matrix data. Let’s open it and explore!
Opening GitHub repository for 428 Hadamard matrix
The page at math.ntua.gr timed out before, but I might try downloading the file directly instead of opening it. Alternatively, I could explore the GitHub “anymatrix” repository, which likely has Hadamard matrices up to order 428. I’ll open the repository’s root to check its structure and find the matrix file.
Opening Hadamard folder
I’ll go ahead and open the Hadamard folder at link ID 81. Time to check what’s inside and see if the matrix or relevant files are there!
Trying GitHub raw content for Hadamard matrix
It looks like extracting the files might have failed because of JS requirements, so I’m considering an alternative: accessing raw content directly from GitHub. I’ll search the repository for relevant files containing the 428 matrix — once I find out the exact file names, I can access them with no issue!
Trying direct access to Hadamard file
It seems the math.ntua listing might still be promising, with both text and zip formats for the Hadamard matrix. I’ll try opening the direct file URL instead of the page, hoping it’ll avoid the timeout issues. It’s worth retrying, as the timeout could have been temporary. Let’s try again!
Trying alternative ways to access Hadamard matrix
Since opening the math.ntua.gr server is
Checking old Hadamard links
The page seems to be from 2001, and it has a link to Neil Sloane’s library list, which might include the H428 matrix. However, the link for item 8 points to “www.research.att.com ,” which could have moved. Let’s click and see if it’s still accessible.
Trying alternative download for Hadamard matrix
The ATT site removed the old page, and Sloane’s library now points to a page that times out. Another option might be using Anymatrix, but GitHub listings aren’t working. I’ll try accessing the IPM website directly to download the matrix — maybe the file’s URL works better with HTTP instead of HTTPS. Let’s give it a try.
Searching for accessible Hadamard matrix
The page redirects to HTTPS and times out, which might indicate blocking. Maybe the matrix is hosted elsewhere, like in a plain text or PDF format. I’ll search for terms like “h428.txt” or “hadamard428.txt” on cs.uleth.ca — there’s a mention of a PDF that could include it. Let’s give that a try.
Opening ADTHM page for possible matrix
I’ve found a listing on cs.uleth, and there’s a link on the ADTHM page that might lead to the matrix in text format. It could be what I’m looking for, so I’ll go ahead and open the ADTHM page to check it out. Let’s see if it includes the matrix!
Accessing matrix text through link
Awesome! I found a link labeled “Entries as text” (ID: 26). This could be the matrix in the text format I’m looking for. I’ll go ahead and click the link to download it. Let’s see if this works!
Parsing the XPM image to a matrix
The H428.html contains XPM image data, not matrix entries. It seems to represent values with two colors, p and m, which likely stand for +1 and -1. I can map p to 1 and m to -1 (just to simplify things) and validate that the Hadamard property holds by checking H*H^T = 428I. I’ll fetch and parse the full content to confirm.
Downloading H428.html
Alright, I’ll use the container.download method to fetch the H428.html file. I previously opened the URL, so that should meet the requirement. Now, I’ll download the file directly to /mnt/data/H428.html and get the content I need. This should work smoothly and efficiently!
Understanding limitations for container.download
It seems the container.download method restricts us to file types other than text/html — likely just images or data. Since I can’t use click or fetch via Python, I’ll need to rely on web.run for the content. However, the 200-word limit means we can’t fully retrieve the 428 lines. I may need to process it in parts.
Handling 200-line limits in tool response
Since the tool response gets cut off around 200 words, I can’t present more than that in the final output. However, I can strategically open various line ranges and parse those internal portions to reconstruct the full matrix. This should allow me to gather all the necessary data without breaking the word limits.
Handling file retrieval and parsing
To retrieve the full content, I can call the tool multiple times from different lines, reconstructing the file. However, retrieving 428 lines in total means a heavy operation, spread across several calls. Given the content-type restrictions, I can process the data using Python, but I need to work with a local file.
Figuring out how to retrieve and process data
To get more lines, I can use a tool to fetch a longer response, but there’s still a limit to the tool’s output. Since the H428.html file is too large, I could check if there’s a plain text version like “H428.txt” in the directory that can be downloaded. I’ll also look for any relevant comments in the HTML.
Looking for matrix in different format
The file starts with “/* XPM */,” indicating it’s actually plain text despite being labeled as HTML. I could try adding a download parameter or using a raw link, but the content type restriction remains. Next, I’ll check if there’s a plain text version of the matrix, like “H428.txt” in the directory.
Considering ways to fetch and process the file
There’s no plain text, but the XPM format might be parseable. I could gradually access the file using the web tool, then reconstruct it with Python. One option is fetching parts with web tool, accumulating in Python, and then processing. I could work around content type restrictions with query parameters for text/plain.
Exploring options to access content
The H428.html doesn’t seem like real HTML. Maybe I can access it through a ?download=1 query or H428.xpm? I might also try the enlarge image link (id 25). If it’s a GIF/PNG, I could download it and extract pixel values, though this might involve some image processing. I’ll check it out.
Downloading and inspecting image
The image is a PNG, so I can download it using the container. I’ll pull down H428_big.png and check the image’s pixel structure. It’s likely 856x856, possibly scaled at 2x, though the matrix should be 428x428. I’ll inspect the image further after downloading it.
Opening image for inspection
The file is small, just 5kb, which likely means it’s compressed and may be a 428x428 image with a palette. I’ll go ahead and open the image to inspect it using the container, and then I can analyze the pixels to understand how it’s structured further.
Processing image for matrix conversion
We have the image, so it’s time to use Python to read it and convert the pixel colors into matrix entries. First, I need to check the image’s pixel dimensions—more likely it will be 428x428 or a multiple of that. I’ll use Python with PIL to open the image and confirm
The AI people are still infatuated with math. The Epoch AI staff, after being thoroughly embarrassed last year by the FrontierMath scandal, have now decided to make a new FrontierMath Open Problems benchmark, this time with problems that people might give a shit about!
I decided to look at one of the easiest "moderately interesting" problems and noticed that GPT-5.2 Pro managed to solve a warm up version of the problem, i.e. a version that had been previously solved. Wow, these reasoning models sure are capable of math! So I was curious and looked at the reasoning trace and it turns out that … the model just found an obscure website with the right answer and downloaded it. Well, I guess you could say it has some impressive reasoning as it figures out how to download and parse the data, maybe.
Hey, you’re selling them short: there are also ReLU and softmax activation functions thrown around here and there. Clankers aren’t just linear transformations! /j
I am a computer science PhD so I can give some opinion on exactly what is being solved.
First of all, the problem is very contrived. I cannot think of what the motivation or significance of this problem is, and Knuth literally says that it is a planned homework exercise. It’s not a problem that many people have thought about before.
Second, I think this problem is easy (by research standards). The problem is of the form: “Within this object X of size m, find any example of Y.” The problem is very limited (the only thing that varies is how large m is), and you only need to find one example of Y for each m, even if there are many such examples. In fact, Filip found that for small values of m, there were tons of examples for Y. In this scenario, my strategy would be “random bullshit go": there are likely so many ways to solve the problem that a good idea is literally just trying stuff and seeing what sticks. Knuth did say the problem was open for several weeks, but: 1. Several weeks is a very short time in research. 2. Only he and a couple friends knew about the problem. It was not some major problem many people were thinking about. 3. It’s very unlikely that Knuth was continuously thinking about the problem during those weeks. He most likely had other things to do. 4. Even if he was thinking about it the whole time, he could have gotten stuck in a rut. It happens to everyone, no matter how much red site/orange site users worship him for being ultra-smart.
I guess “random bullshit go” is served well by a random bullshit machine, but you still need an expert who actually understands the problem to read the tea leaves and evaluate if you got something useful. Knuth’s narrative is not very transparent about how much Filip handheld for the AI as well.
I think the main danger of this (putting aside the severe societal costs of AI) is not that doing this is faster or slower than just thinking through the problem yourself. It’s that relying on AI atrophies your ability to think, and eventually even your ability to guard against the AI bullshitting you. The only way to retain a deep understanding is to constantly be in the weeds thinking things through. We’ve seen this story play out in software before.
I was pissed when my (non-academic) friends saw this and immediately started talking about how mathematicians and computer scientists need to use AI from now on.
scott jumpscare
Baldur Bjarnason’s essay remains evergreen.
Even when looking at Knuth’s account of what happened, you can already tell that the AI is receiving far more credit than what it actually did. There is something about a nondeterministic slot machine that makes it feel far more miraculous when it succeeds, while reliable tools that always do their job are boring and stupid. The downsides of the slot machine never register in comparison to the rewards. Does it feel so miraculous when I get an idea after experimenting in Mathematica?
I feel like math research is particularly susceptible to this, because it is the default that almost all of one’s attempts do not succeed. So what if most of the AI’s attempts do not succeed? But if it is to be evaluated as a tool, we have to check if the benefits outweigh the costs. Did it give me more productive ideas, or did it actually waste more of my time leading me down blind alleys? More importantly, is the cognitive decline caused by relying on slot machines going to destroy my progress in the long term? I don’t think anyone is going to do proper experiments for this in math research, but we have already seen this story play out in software. So many people were impressed by superficial performances, and now we are seeing the dumpster fire of bloat, bugs, and security holes. No, I don’t think I want that.
And then there is the narrative of not evaluating AI as an objective tool based on what it can actually do, but instead as a tidal wave of Unending Progress that will one day sweep away those elitists with actual skills. Random lemmas today mean the Millennium Prize problems tomorrow! This is where the AI hype comes from, and why people avoid, say, comparing AI with Mathematica. To them I say good luck. We have dumped hundreds of billions of dollars into this, and there are only so many more hundreds of billions of dollars left. Were these small positive results (and significant negatives) worth hundreds of billions of dollars, or perhaps were there better things that these resources could have been used for?
Don’t worry, there’s always Effective Altruism if you ever feel guilty about causing the suffering of regular people. Just say you’re going to donate your money at some point eventually in the future. There you go, 40 trillion hypothetical lives saved!