scruiser, scruiser@awful.systems

Instance: awful.systems
Joined: 3 years ago
Posts: 1
Comments: 149

RSS feed

Posts and Comments by scruiser, scruiser@awful.systems

It is this continuing slippage of standards that makes me appreciate a hard line against any and all genAI that place like awful.systems have. You concede one small usage and the boosters will keep pushing for more.


chinese confabulation engine

The best case scenario (for them) is that is a Chinese confabulation engine. But that requires that the logs mention ChatGPT as an in-joke. And code or logs of interfacing with a local instance of the Chinese confabulation are for some reason not available for totally innocent reasons. The more likely scenario is that it was ChatGPT.


libertarian socialist ideals

wtf are these ideals supposed to be?

If its “libertarian, but doesn’t want people dying in the streets or dying of preventable diseases”, I will give them the tiniest modicum of credit for being better than standard libertarians, but being libertarian at all still leaves them deep in the negative on credibility in my books.

pro-non-corporate GenAI technology

By which I assume they mean corporate models that got released as open weight models (and at-best/at-most got a touch of fine-tuning from a community effort), but still ultimately originated from mass plagiarism and are still not useful beyond generating slop…


I don’t have anything especially useful or insightful to add, but I just want to thank our mods for acting decisively to defederate us.

Edit: seeing the discussion over on https://awful.systems/post/8293548 I think our mods were actually extraordinarily patient in giving db0 a pretty solid chance to own up to their mistakes and commit to no LLM usage in the future and they instead doubled down hard.


Good on you for pushing that point. Trying to pass off the logging details identifying their llm as ChatGPT as a “joke” was incredibly stupid, but more people might have bought that lie if you didn’t press that followup question.


Even Scott’s fantasy dream scenario for what prediction markets could be like and what questions they could answer feels… …deliberately naive? …like libertarian brainrot? …disconnected from reality?

Ask yourself: what are the big future-prediction questions that important disagreements pivot around? When I try this exercise, I get things like:

Will the AI bubble pop? Will scaling get us all the way to AGI? Will AI be misaligned?

Huge amounts of money are being dumped into a bubble based on hype, so to hope a predict market would or could make better predictions than the existing business-idiot VCs funding this bubble feels hopelessly naive in a libertarian kind of way. There is already a method of aggregating the wisdom of the crowd and it is failing to incredibly lazy hype and PR.

Will Trump turn America into a dictatorship? Make it great again? Somewhere in between?

Again, there is already a mechanism for aggregating wisdom of the crowds, its called an election, and its also failed to get a answer predicated on reality or truth, so again, it seems incredibly naive to expect prediction markets to do better!

Will YIMBY policies lower rents? How much?

I mean, the councils and communities making these decision already ignore or overlook longer-term broader predictions of economic impact in favor of immediate home-owner value, I don’t see why Scott would expect prediction markets to help decision making go better here.

Overall, it feels like Scott is overlooking the way decision making often already ignores science and experts. Society doesn’t have a problem making decent predictions compared to the problems it has communicating expert opinions to the public effectively and crafting policy aligned with the public interest.


The prediction markets seem to have all the basic problems that sneerclubbers: problems with resolution mechanisms, all sorts of insider trading and gaming the market, people using it for gambling…

Various prediction markets have made various half-assed attempts at solutions, but so far nothing seems to actually work well enough to make prediction markets nearly as useful as rationalists expected.


Some of the change probably involves the discovery of a natural bat coronavirus with a furin cleavage site last October, but I’m surprised by the extent of the decline.

That actually seems like the prediction market sort of did its job in this case? I mean, 27% yes is still too high, but actually changing in response to real evidence is much better than my low low expectations for prediction markets. It seems like he should take his own advice and actually take the prediction market seriously in this case.


Yeah that was a good article. I think that is one of the fundamental issues with rationalists, they are basically a group formed around neat sci-fi ideas and not actually getting anything done, and their strong libertarian biases prevent them from actually pursuing the strategies that would be most effective for many of their nominal goals.


Their proposed sort of solution (controlled miscalibration) even amounts to forcing the model to generalize less by memorizing more, which used to be the opposite of why you would choose to use this type of topography.

Yeah, it does seem to be running into the basic issue that what boosters want LLMs to be (all knowing oracle) is in sharp contrast to what LLMs actually are (churn out statistically plausible content).


You’ve described the problem with generalization yes. Well, you could maybe sort of train it not to generate “all men are cats”, but then that might also prevent it from making the more correct generalization “all cats are mortal” or even completely valid generalizations like combing “all men are mortal” and “Socrates is man” to get “Socrates is mortal”.

The problem with monofacts is a bit more subtle. Let’s say the fact that “John Smith was born in Seattle in 1982, earned his PhD from Stanford in 2008, and now leads AI research at Tech Corp,” appears only once in the training data set. Some of the other words the model will have seen multiple times and be able to generate tokens in the right way for. Like Seattle as a location in the US, Stanford as a college, 2008 as a date, etc. But the combination describing a fact about John Smith appearing uniquely trains the model to try to generate facts that are unique combinations of data. So the model might try to make up a fact like “Jane Doe was born in Omaha in 1984, earned her master from Caltech in 2006, and is now CEO of Tech Corp” because it fits the pattern of a unique fact that was in its training data set.


For the chain of thought instruction following model gpt-oss-20b, I’ve noticed its reasoning content often includes it talking about stuff it is supposed to avoid in the final output and it double checking that it doesn’t have that forbidden output. So it would waste tokens talking about pink elephants in its reasoning content, but then do okayish at avoiding pink elephants in its final output.


Theoretically if the people responsible for that training and reinforcement did their jobs well then those patterns should only include true statements but if it was that easy then you wouldn’t have [insert the entire intellectual history of the human species].

I’m chiming in to agree with Architeuthis and mention a citation explaining more. LLMs have a hard minimum rate of hallucinations based on the rate of “monofacts” in their training data (https://arxiv.org/html/2502.08666v1). Basically, facts that appear independently and only once in the training data cause the LLM to “learn” that you can have a certain rate of disconnected “facts” that appear nowhere else, and cause it to in turn generate output similar to that, which in practice is basically random and thus basically guaranteed to be false.

And as Architeuthis says, the ability of LLMs to “generalize” basically means they compose true information together in ways that is sometimes false. So to the extent you want your LLM to ever “generalize”, you also get an unavoidable minimum of hallucinations that way.

So yeah, even given an even more absurdly big training data source that was also magically perfectly curated you wouldn’t be able to iron out the intrinsic flaws of LLMs.


-3 upvotes and 0 karma, but the article is absolutely right (they hate this post because it tells the truth). If Eliezer wants to influence public discourse and policy on an international level, he absolutely does need a respectable image (with maybe a touch of eccentricity in an allowable way). But apparently (what he thinks is) the literal end of the world isn’t enough to make him actually try for normie public image. Or maybe he has some galaxy brain plan about how looking like a weirdo actually helps his cause? If he does, I strongly suspect it is a rationalization.


Wonder of the goblin stuff is the start of some model collapse.

That is exactly it. Their official explanation avoids the phrase model collapse, but that is exactly what they describe: using the output of one model as training data for another amplified the occurrence of the word goblin (and other creatures), which apparently initially occurred because of their system prompt which was aimed at maximizing the Eliza effect (again they avoid an honest framing, but that is totally what they are doing and it is pretty gross considering all the cases of AI psychosis that have been occuring) by telling the model “You are an unapologetically nerdy, playful and wise AI mentor to a human. "


Widespread financial fraud which was legitimized and in some cases directly backed by EAs! Surely there are no parallels!


Zitron’s analogy is excellent because the bubble is multifactorial and the analogies that we can make are factor-to-factor. Here’s some things that caused the dot-com bubble; people were overly optimistic about:

Ed has also been clear there are a few factors that make this bubble worse (for the economy and the general public) than the dotcom bubble. For one, Ed is strongly convinced that GPU lifecycles are much shorter and worse than fiber optic life cycles. You build fiber optic infrastructure and it will last for decades. Meanwhile, GPUs used constantly at max load have life cycles of 3-5 years. The end result of the internet is also much more useful and less of a double-edged sword than the slop generators which churn out propaganda and spam.


I am a pretty big fan of Ed’s work, so I’m going to hold my nose and read Kelsey’s work thoroughly enough to do a line by line debunking:

Over the last two years, he has called the top repeatedly:

Well yes, but he has also explicitly said that the bubble peaking and popping would be a multiyear process. I’ve only kept up with his every article for the past year, but in the past year, his median guess for the bubble pop becoming undeniable was 2027. I guess making timelines with big events in 2027 and hedging on the median number is only for the rationalists? Also, we are already starting to see the narrative fray as Anthropic and OpenAI experiment with price hikes and struggle with getting ready for IPO, which would count as meeting his predictions for the start of the bubble pop.

In 2026, the focus is much more on alleging widespread, Enron- or FTX-tier outright fraud.

This is basically an admission that he can’t make the case in terms of the economics anymore.

??? Ed has been making the case for circular financing and investors being deceived because he thinks there are circular financing deals and investors being deceived. Ed has slightly softened on his position on exactly how useless or not LLMs are, but he is still holding to his economic case that the amount they cost isn’t worth the value they provide, extremely blatantly so once consumers start paying the real cost and not the VC-subsidized cost.

By almost every metric, AI progress from 2024 to 2026 has been much faster than AI progress from 2022 to 2024.

And she is quoting a rat-adjacent think-tank for proof that AI improvement has been exponential. Even among the rationalist, the case has been made that the benchmarks are not reflective of real world usage/value and that costs are growing with “capabilities”.

It can no longer argue that costs aren’t falling; they are.

Even accepting the premise that real costs have fallen, Kelsey fails to address Ed’s case that the costs LLM companies charge is massively subsidized. If real costs are 10x the current subsidized costs (which have already been pushed up as far they can be without losing customers), and model inference prices miraculously drop 5x (which Kelsey would treat as a given, but I think is pretty unlikely barring some radical paradigm shifts), that is still a 2x gap.

It is a straightforward crime to claim $2 billion in monthly revenue if you mean that you are giving away services that would have a $2 billion market value.

Yes, exactly. Technically OpenAI and Anthropic play games with ARR and “gross” revenue (i.e. magically excluding the cost of training the model in the first place), but in a just nation it would straightforwardly be a crime. Why does she find this hard to believe?

Epoch AI has an in-depth analysis of the same financial questions from the same public information

(Looks inside the Epoch AI article):

So what are the profits? One option is to look at gross profits. This only considers the direct cost of running a model

Ed has gone into detail repeatedly about why excluding the cost of training the model is bullshit.

(More details from the article)

But we can still do an illustrative calculation: let’s conservatively assume that OpenAI started R&D on GPT-5 after o3’s release last April. Then there’d still be four months between then and GPT-5’s release in August,22 during which OpenAI spent around $5 billion on R&D.23 But that’s still higher than the $2 billion of gross profits. In other words, OpenAI spent more on R&D in the four months preceding GPT-5, than it made in gross profits during GPT-5’s four-month tenure.24

Oh that is surprising, the Epoch AI article actually acknowledges the point that these models are wildly unprofitable once you account for the training cost! Of course, they throw away their point in the next section by just magically assuming LLMs will prove to massively valuable in the near future! (One of the exact things Ed has complained about!)

He’s found too many grounds for dismissing all the financial information we have as dishonest or irrelevant to seriously engage with what any of it would imply if it were true.

He has shown in detail how the companies use barely technically not lying obfuscated bullshit metrics like gross profit or ARR to inflate their numbers and if you try un-obfuscate them the numbers look a lot worse.

Kelsey goes on to try to claim how much value LLMs provide

Making them more productive is a big deal, and in 2026, AI makes them more productive.

Zitron can’t really contest this with contemporary data, so he cites 2024 and 2025 studies of much weaker AIs with much weaker productivity impacts.

Two years to… 4 months ago! Such outdated information! In the first place there has been very few rigorous studies of how much of a productivity boost LLM coding agents actually provide, and one of the few studies with even a passing attempt at rigor (while still below good academic standards), was METR’s study (and keep in mind they are a rat-adjacent think tank and not proper academics), which showed programmers thought they got a productivity boost but actually got a net productivity decrease!

From this set of beliefs, you could, in fact, defend a delightful bespoke AI bubble take: that AI would have been a catastrophic investment bubble, but the AI companies were saved from their mistakes by the determined NIMBYs of America killing off the excess data center build-out.

But that’s not Zitron’s stance. He seems to account “the build-out is too aggressive” and “the build-out is not happening as planned” as both independent strikes against AI — both things that show it’s bad, and the more of those he finds, the more bad it is.

It could in fact be all 3! The hyped-up build out, such as that indicated by OpenAI’s and Oracle’s 300 billion dollar detail was completely insanely too aggressive (for it to pay off, Ed calculated LLMs would have to drastically exceed Netflix+Microsoft Office in terms of ubiquity and price point), not achievable given realistic build times for data centers (Ed has also brought the numbers here), and even at the reduced actually rate of build out, still not actually financially viable (is simply because the LLM companies aren’t charging enough). So yes, both things are bad, and one type of badness partway mitigates the other, but it is still all bad!


I advise being very cautious about consuming Zitron’s posts

He has got a dramatic and vitriolic style, but as dgerard says, he has also dug through the numbers. I see lots of criticism of Ed’s style, but not nearly so substantial criticism about the hard numbers he has come up with. The LLM companies put out contradictory and obfuscated numbers, and taken naively they seem to contradict Ed’s numbers, but as Ed has shown, many, many times, when you start trying to un-obfuscate them they start looking really bad for everyone betting on LLMs.

Many coders are using chatbots, but I don’t know of evidence that it makes them more productive

So more and more coders are coming around to “actually AI code is okay”… but as we’ve seen repeatedly with LLM generated content, it is very easy for people to “Clever Hans” themselves and convince themselves LLMs are contributing more than they actually are, so I am not going to trust anecdotal reports.


I wouldn’t give him credit for a full admission. He isn’t acknowledging that “biased left-wing experts” means expert like psychologists with a basic understanding of psychometric validity and geneticists with the basic understanding that popular notions of race don’t have a genetic basis and biological determinism is false.


RSS feed

Posts by scruiser, scruiser@awful.systems

Comments by scruiser, scruiser@awful.systems

GDP is up though, so I’m sure the AI god prosperity will trickle down eventually.


He describes textbook insane cult stuff, but still feels the need to say stuff like:

Again I’ll say: there was a great deal of good at MAPLE.

The way I’ve described it to friends is that it was the best decision I ever made to go there; the second best, however, being to leave.

There was the same dynamic in the writing of the ex-leverage member that got posted here a few weeks ago. She also felt the need to emphasize the “good” parts. Is the cult programming that hard to break? Is it some form of sunk cost or rationalization or need to claim something positive about the experience?


The cope on lesswrong is funny. They are in denial that this is a sign of a wider bubble pop and are insisting it is just because he didn’t hedge on long enough timelines.


good grifter

Not denying that he isn’t also a grifter, but I bet he is a true believer and he had blindly bet his (and other people’s) money on “line goes up” exactly like his scenario said and that is why he is the first one to crack.


and the “safety” community has done jack shit the entire time and has offered zero solutions.

Occasionally I see a sane relatively workable idea on lesswrong (like making stricter laws on liability and transparency for everything AI companies do, I’ve also seen this idea put forward on lawfare a few times), but that is like 1 idea out of 20, with the other 19 being absurd, like the absolutely unworkable fantasies of AI: 2040.


The blogger simultaneously elevates mathematics to some mystical endeavor and fixates on novel theorem proving as the key element of that and believes LLM-based AI will replace human mathematicians at theorem proving in a matter of years… that is quite a combination. I can imagine how 1.5 of those things fit together (although I disagree with that view point obviously), but the whole package is really quite an odd combination.

Also, in the fantasy scenarios where AI really is capable of totally replacing mathematicians, aren’t we supposed to get post-scarcity abundance, freeing up your time to pursue mathematics out of pure desire for enlightenment? Maybe the blogger doesn’t believe that part? Or they are so attached to themselves personally getting to discover novel theorems first they don’t care that the post-scarcity era would on net free up a lot more people to pursue pure mathematics as a hobby.


One of the defenses I see of LLMs on place like /r/singularity (although recently /r/singularity has started wising up and the committed true believers have shifted to /r/accelerate), is that they are going to cure cancer or some other incredibly valuable thing, conflating AI, ML in general, DNN, vs. LLMs specifically. Technology like AlphaFold was a lot more on the “cure cancer” track (although there is still a huge gap between better protein folding predictions and successful drug discovery) than anything LLM related…. this is sad.


Great article! I think everyone here is probably already well aware of the way Silicon Valley keeps deliberately misunderstanding sci-fi for their hype (don’t invent the torment nexus), but this article does a good job wrapping together a lot of examples.

And I realized this website is also the place where a year ago I read a great article on the history of nanotech as a real science vs. Drexler’s fantasies (the magic nanobot lesswrong still believes in) The Nanobots pipedream. So it looks like a pretty good newsletter, I might keep an eye on it going forward!


I’ve think they’ve moved the overton window within their own in-group, but whenever this sort of stuff leaks to a more mainstream audience everyone rightly reacts with an eww or at least a wtf.


Aella regularly post edgy surveys on sexual and gender related topics on twitter. She probably picked the examples for maximum edginess. And then made an complicated diagram to get more engagement, because engagement farming is one of her big things.


“We are now, like, in the singularity”

/r/singularity also thought this was bullshit.



Looking at this post and some of their other posts, it looks like they’ve got some sort of scrupulosity issues with this? Like obsessing over the Evil inflicted on the LLMs (they think every session is a new person, and every session might be the equivalent of hours of human subjective experience, which even by the standards of claiming LLMs have meaningful internal experiences is pretty out there) to the point they can’t even slightly reconsider the factual question of if LLMs have internal experiences.


(2) has a big “if” in there

Its a super big if, but it is one lesswrong (and Anthropic, with their “model welfare” pseudoscience) is allegedly seriously considering as a possibility (since, you know, they think LLMs are already AGI and Eliezer was accusing AI-dungeon, you know, GPT-2, of deliberately scheming). If anything, it speaks to some possible hypocrisy/motivated reasoning that they consider LLMs AGI but not possibly morally relevant.

conceptualized capitalism

This seems to be a lesson most lesswrongers have had rubbed in their faces repeatedly over the past 5 years with everything about LLM company’s behavior, but they just don’t want to learn.


Aella took one of her shitty twitter surveys and made an (dis)infographic, which made it onto the subreddit /r/moralityscaling, which made it to /r/whowouldcirclejerk where I saw it, so now I’m sharing it with you.

CW for Rape, Transphobia, implicit rape apologia

Crossposted from our subreddit: https://www.reddit.com/r/SneerClub/comments/1v80j9q/scalers_need_an_intervention/


I bet Elon could have kept the plates spinning for decades longer if he was happy with just Tesla and SpaceX, but he had to keep doubling down on grifting investors so maybe, finally, we will see it all come crashing down for him. (What happens once his assets no longer cover his debts? Do they actually force him to liquidate?)


I’ll go one step wider. It is not just Sam Altman or his social circle, but all the people most enriched by capitalism (and thus most able to exploit and benefit from a technology that shifts power away from labor and to capital) that I don’t want to see further empowered.


I believe the main ingredient is Lean, which is a formal language resembling a programming language. Math proofs written in Lean can be verified deterministically with a computer, which really helps mitigate the hallucination problems of LLMs.

100% this. Also, looking back at an earlier example that was actually written up in more detail, AlphaGeometry 1 got 28/30 problems, but entirely stripping out the LLM from the system, the symbolic logic proportion alone could get 14/30, and replacing the LLM with different heuristic methods could get 18/30 and 21/30 (for different methods).

Even if math research works out perfectly well (which is a still big if), it’s not going to pay the bills. They would need to find a use case in the real world, where hallucinations can cause serious damage and cannot be formally prevented. And they have certainly tried. Math will not change the fact that all of this will collapse.

The boosters and LLM companies still believe LLMs get their current level of performance by generalizing and not just memorizing facts (and maybe a wide shallow pool of weak heuristics). So they are hoping by pushing the LLM performance up in some narrow domain they can churn out synthetic data for, they will see some large general improvements in LLM performance.


that would be a signal to drop AI and nuke data centers from orbit

On Lesswrong I’ve seen some posts contemplating why this “warning shot” isn’t making any political waves. Also, in response to the fact the OpenAI has released basically no details, I read one poster realizing that actually they need external governance and regulation on the AI companies.


I had two main guesses even before reading into the details…

The obvious, they set this up for the doom crit-hype value. That is how basically every one of these stories like this to come out of Anthropic turn out to be once you read the details of the setup. The lack of details is to keep the mystique and hype up.

The other guess… their coding agent actually did manage to vibe code its way out of its ‘sandbox’ and past Huggingface’s security, but the security on both ends (OpenAI’s"sandbox” and huggingface’s code) was itself laughably bad vibe coded garbage. The lack of details is to obfuscate just how garbage they are. I think this guess is pretty well supported by the linked Mastadon posts (although I still wouldn’t discount my first guess until more details are pried out). My “favorite” quote:

“prepend some text with processing instructions and an SGML tag, and then hand that blob to an llm.”

I fucking hate this timeline.

and they turned it into a mutual marketing stunt

Their marketing actually seem to be trying to take different spins on it…

HuggingFace stressed that it worked out the attack was going on using an open weight model it hosted itself — because the guard rails on the commercial model HuggingFace tried wouldn’t let them do security work. So you should use open weight models from HuggingFace.

Yeah, its an opposite spin than Anthropic and OpenAI have been trying lately, where they want to shut down all open weights model because something something China something something too dangerous something something regulate our competition out of existence. (It is funny, some of the boosters on hackernews and /r/singularity are actually acknowledging without an artificial moat to help boost OpenAI’s and Anthropic’s capex for bigger and bigger models will stop. Even though they still deny Ed Zitron’s calculations of the financials.)

Anyway, I guess even if they have some very opposed strategic directions, huggingface and OpenAI are aligned on maximizing the hype.