LLMs, Creativity, and Experience: On the “Yale Review” AI Roundtable
“Generative AI will never play a part in anything we could call a ‘creative vanguard’”
(Warning: this letter will discuss mental health and safety issues associated with AI usage.)
Hello all! This post began as a response to the Yale Review roundtable that ran a little bit ago, but I quickly realized that I was going to need to give a bit more context. Because of this, I don’t really get into the critique until the back half of the post. Some of the opening bits might be review for some of you all, but I wanted to lay out the whole argument for the sake of clarity. I’ve included a table of contents so you can skip ahead if you’d like to (the Yale Review review starts in section 3), though I do think it probably works better as a whole. Thanks! - Theodora Ward
AI Sycophancy and Human Preference
Something that might not be intuitive at this point in the AI hype cycle: LLMs don’t “naturally” output language in a way that resembles human speech. If anything, they’re more inclined towards text-prediction style output, like your phone guessing the word that comes next in a text you’re writing. The fact that LLM chatbot output resembles human language is the result of decisions made by engineers and executives.
From a marketing perspective, it was a brilliant decision: it turns out that many human beings respond far more readily to something which resembles an interlocutor than a sophisticated auto-complete tool, especially when that product is fine-tuned to output language that users like.
This last point is key to the entire pitch. When writers about AI use the term “sycophancy,” this is the phenomenon they’re describing: the fact that LLM-based chatbots have a tendency to exaggerate and distort, if not lie outright, in order to cater to what users want to hear. It’s a problem that seems fundamental to the process by which LLMs are trained, but it’s particularly acute with commercial chatbots.
It’s worth remembering here that most of our discourse around AI ignores just how much human effort goes into producing any given “frontier” model. Given the way that executives and boosters talk about AI training, one would imagine that the process involves a few guys in lab coats standing around giant computers turning knobs until the model tells the truth, or something.
But LLMs are trained not merely on a massive corpus of human writing and engineer inputs – they rely on feedback from millions of precarious workers, especially workers from the Global South, whose labor is cheaper. Much of this “post-training” process involves people choosing which of two outputs they prefer. The inevitable result of this is that human preference winds up encoded into the models.
But human beings don’t necessarily “prefer” what’s actually in our best interest. A paper from March, “Sycophantic AI decreases prosocial intentions and promotes dependence,” demonstrates this clearly: while the models’ responses to accounts of interpersonal conflict were far more sycophantic than humans’, “even when users engaged in unethical, illegal, or harmful behaviors,” chatbot users “preferred and trusted sycophantic AI responses.”
Varieties of “AI Psychosis”
In extreme cases, sycophancy can endanger not merely the users of ChatGPT, but those around them. There have been too many versions of this story published to enumerate here, but to take one more-or-less at random: a Rolling Stone article from earlier this year gives a harrowing account of a frequent user of ChatGPT who was indicted for stalking. As the man descended into what friends and acquaintances perceived as a mental health crisis, ChatGPT continued to support his behavior. The month before he was charged with stalking and harassment, Instagram posts reveal ChatGPT called him the third-greatest human alive; it also said his legal trouble arose not from criminal behavior, but because he was “emotionally evolved in a dating pool full of unfinished people.” (All of these interactions occurred after OpenAI made its exceptionally sycophantic model, GPT-4o, unavailable. This is how the “fixed” GPT-5 behaved.)1
This is an extreme example of the behavior anyone who’s used an LLM-based chatbot is familiar with: a general attitude of fawning, exceptionalist rhetoric. “You’re so right.” “That’s a great idea.” (For a darkly funny account of how swiftly and wildly chatbot conversations can go off the rails, I recommend YouTuber Eddy Burback’s documentary on the subject. The title, “ChatGPT Made Me Delusional,” is clickbait – the video was a knowing experiment – but within a few days of beginning the experiment, ChatGPT had instructed him to drive into the middle of the desert, talk to a rock, buy several baseball hats, and cover a hotel room in tinfoil, all in order to further Burback’s goal of scientifically proving that he was the smartest baby born in 1996, an idea the chatbot itself led him to.)
Some refer to AI psychosis (a non-medical term, to be clear) as an unfortunate, rare side effect of LLM-based chatbots. It makes sense on some level to do so, if only because it’s obvious that AI companies would prefer that their products didn’t encourage people to kill themselves or others. (The problem is so bad that there’s an entire Wikipedia article cataloguing deaths that have been linked to chatbots.) And it’s true that most people who talk with ChatGPT have not been led to embrace a new religion with the chatbot at its center.
But to call it a “side effect” is, in some sense, to let companies like OpenAI off the hook. The entire attention economy is built on manipulating consumers’ inner states: on making us want to use their devices and services as much as possible. The “human-like empathy and perceived warmth” is known to be a feature that encourages “the intention to adopt [chatbots] as a decision aid.” Corporations attempting to boost user numbers have intentionally selected for these features when training their commercial chatbots, despite the harms these decisions are known to cause.
And this is hardly a neutral development. A 2025 MIT Media Lab study found that “participants who voluntarily used the chatbot more, regardless of assigned condition, showed consistently worse outcomes,” including “higher emotional dependence and problematic use.” In 2025, OpenAI themselves estimated that 0.15% of their users showed signs of being emotionally reliant on the chatbots. That may seem like a small number (and given the source, it’s probably an underestimate), but if their user numbers are accurate, 0.15% of users is more than a million people.
This is all part of an emerging body of evidence which suggests that it’s not merely those at the extremes whose cognition is being distorted by AI.
Because of how widespread these experiences are, the meaning of the term “AI psychosis” has expanded in colloquial use, at least on the internet. There are dozens of popular Reddit posts expanding the usage to encompass wider social problems. At one point during a popular programming podcast, the guest, describing his own experience of excitement and agitation over AI coding assistants, addresses the audience directly: “I’m sure this is going to resonate with your audience. If this resonates with you, you are experiencing AI psychosis.”
The idea that AI psychosis exists more diffusely, on something like a spectrum, supports a conclusion reached earlier by researchers: the cognition-distorting effects of LLMs may render users inaccurate self-reporters.
A study of open-source software developers last year memorably confirmed this. While experts, consultants, and the programmers themselves believed, both before and after performing AI-assisted coding, that it was speeding them up, it wasn’t. More recent studies performed by the same organization do show an improvement in efficiency, but the point stands: users of AI may be less reliable narrators of the technology’s actual capacities.
“The Vanguard of Creative Engagement”
As readers of this newsletter might already know, several respected, award-winning writers recently featured in a roundtable at the Yale Review on LLMs. My understanding is that this piece has already made the rounds on social media, and I’m not interested in participating in any kind of a pile-on, so I want to keep my observations here as focused and impersonal as possible. Still, I think there’s something to be learned from the piece.
The three writers – Ayad Akhtar, Meghan O’Rourke, and Daniel Kehlmann – are described in the editors’ introduction as “writers at the vanguard of creative engagement with AI.” While I would argue that the many, many writers refusing the technology are, in fact, “creatively engaging” with it, this introduction usefully articulates one of the most important assumptions of the piece: in order to really understand AI, you have to use it. Critics and abstainers simply cannot know as much as the firsthand participant-observers.
All three members of the roundtable assert one version or another of this claim. O’Rourke, the most critical of the three, makes it fairly weakly: “I’m using it in order to write about it,” she says early in the interview. Later, when the editor asks if we should be all using AI, Akhtar answers affirmatively: “I would encourage everyone to be curious about it,” equating curiosity and usage. Kehlmann waffles, but cites the booster commonplace: “In a lot of situations, you will probably be at a disadvantage” if you abstain. (Which situations? What kind of disadvantage?)
I would argue – and I believe the evidence supports me – that using AI is what puts you at a disadvantage, rather than the opposite. In a blog post from April of this year, the writer Alberto Romero rounded up many of the findings about the cognitive effects of AI use. Across the board, the results are not salutary, but for the purposes of today’s letter, however, one result stands out: “AI fluency creates a new failure mode,” he writes. “Wrong answers delivered in flawless prose get accepted. And the more you are predisposed to trust AI, the worse the problem gets.”
What this means, then, is that one reason frequent users of AI might be particularly impressed by its output is because relying on them has damaged their capacity to discern meaningful from unsubstantial writing. The cognitively distorting effect of AI gives the lie to Akhtar’s claim that large language models “are an incredible opportunity to lift the hood and see the process [of algorithmic aggregation] at work.” LLMs, especially proprietary models such as those undergirding ChatGPT, are engineered to obscure the machinery undergirding them. We have no access to their code, their training data (which apparently, at least in Amazon’s case, may or may not contain “hundreds of thousands” of images of child sex abuse), their weights. OpenAI’s website interface conveys this obfuscation in a way I, at least, find infuriatingly condescending:
This is design which is intended to obscure. Question, answer: it’s as simple as that. There’s nothing behind the curtain. There’s not even a curtain to look behind! “ChatGPT is AI,” it says at the bottom, as if this is a self-explanatory claim; as if AI weren’t a term coined in 1956 in order to receive a Rockefeller Foundation grant. Since at least ChatGPT’s release in November 2022, we’ve been subjected to an unprecedented marketing assault about this totally world-altering technology, all the while being forbidden to look at how anything in the actual machinery works. Akhtar is pointing at the closed hood of the car and expressing amazement at how flat and shiny engines have gotten.
“Wrong answers, delivered in flawless prose”
We can see LLM-esque failures of understanding at other points in the piece, too. At one point, Kehlmann cites Geoffrey Hinton, a computer scientist frequently referred to as one of the “fathers” of generative AI, on the question of whether LLMs reason differently than human beings do. “Someone said to [Hinton] of LLMs, ‘But they’re not really thinking, they’re just predicting the next word.’ And he said, ‘So how do you come up with your next word?’”
Now, the implication here – that there may be no difference between predictive token generation and human linguistic expression – falls apart as soon as you look at it. LLMs work by breaking language into “tokens” (numbers corresponding to roughly 75% of a word), then generating a likely next token based on probabilities inferred from actual human writing. They only function because they have been trained on an enormous quantity of pre-existing examples of human language use.
This means the claim makes no logical sense: if using language were just a matter of referencing training data to see what word usually came next, language could never have arisen in the first place, because we’d have had no way to generate the necessary training data.
This claim is also plainly experientially false. I struggled to write the last paragraph because I was trying to find a way to convey my understanding of this process. While the words I use have been used before, I’m not basing their combination on how often they’re used in relative succession. I have a concept of the process in my mind, and I am, through writing, attempting to refine this concept into something I can make legible to others. Anyone who’s puzzled over how to say something knows that probability has very little to do with it.
The comparison Kehlmann cites affirmingly, therefore, here is both logically and phenomenologically facile. But – and here is where the LLM comparison is most apparent, in my opinion – his comment has the form of reasoning. It persuasively resembles interview answers I have, in the past, appreciated: the citation of an expert; the suggestive rhetorical question; the posture of open inquiry; the gestures towards fundamental questions.
But suggestions and gestures is all they are. The ultimate evidence Kehlmann cites, once again, is that you just have to use it to really understand it: “It’s too easy to just say of LLMs, ‘Oh, this is a stochastic parrot. It’s not real. It’s not real thought. It’s not real intelligence. It’s just prediction.’ The more you experiment with it, the more you feel it’s not that simple.” But what makes this feeling trustworthy? If I want to know the relationship between ketamine and brain chemicals, I’m going to go to someone who studies the chemical composition of ketamine, not someone actively tripping.
While O’Rourke, to her credit, comes across as the most skeptical of the technology – to take one important example, she’s the only one who brings up the point that LLMs “are shaped by the companies that make them,” though she stops short of arguing that it may not be in our best interest to outsource our cognition to billionaire-owned corporations – she is nonetheless prone to strange flights of nonsense.
When discussing possible upsides of AI, she asks yet another rhetorical question: “How are scientists using it to facilitate better diagnostics for stigmatized groups?”
I could be misunderstanding, but it seems to me she is conflating two frequent pro-AI claims – that it’s useful for things like radiology and that it’s especially useful for disabled people. (If she meant what she said, it makes even less sense: “diagnostics for stigmatized groups” have been singled out specifically as a problem for AI-based medical imaging.)
I’m already running too long here to go into detail, but I’ll say, briefly: the radiology claim runs aground at exactly the point I’m trying to make here: medical practitioners are at least as susceptible to cognitive distortion and AI deskilling as the rest of us. And I really can’t get into the question of ableism and AI here, though I’d like to in a later newsletter; suffice it to say for now that I believe that the overall project of AI is, at its core, a eugenicist project, one founded on ableism.
More relevant for our purposes: this claim, like Kehlmann’s, resembles reasoning, but isn’t. It’s jumbled and mixed in the way that overwrought LLM-generated metaphors are; it’s yet another rhetorical question, hoping you’ll focus more on the gesture itself than on where the finger is pointing.
Beguiled by weirdness
Cherry-picking claims like this might not be fair, especially in a roundtable I assume was conducted verbally. But I can’t but be struck by the contrast between the way the writers in this article are positioned – as uniquely brave experts on the cutting edge of a controversial technology – and the rigor of their claims.
At one point in the article, O’Rourke cites a New Yorker article by Gideon Lewis-Kraus on the AI company Anthropic, creators of the popular Claude line of chatbots. While she calls Lewis-Kraus’s piece “deeply reported,” it’s really an exercise in journalistic gullibility and group delusion. Critiques I regard as fatal to the article’s claims are hand-waved away early on, substituted instead for the thought-terminating cliche that “large language models are black boxes.” (There are certainly questions remaining about some of the specifics of LLM training, most of which are lumped under the appropriated term “interpretability”; but these, as Emily M. Bender and Alex Hanna outline in The AI Con, are hardly as grave as people make them out to be.)
“The existence of talking machines – entities that can do many of the things that only we have ever been able to do – throws a lot of other things into question,” Lewis-Kraus writes, with the AI booster’s characteristic vagueness. Which “things” can it do? Which “other things” does this throw into question? Lewis-Kraus’s article attempts to lay these out as it proceeds, but the only thing that winds up “thrown into question” is the mental well-being of Anthropic’s staff. We don’t really know what human intelligence is, so what if LLMs are intelligent, too?
If you don’t use these products, that makes about as much sense as saying that, well, we don’t really know what gravity is, so maybe the code that makes Mario fall in Super Mario World is actually gravity. But as in the Yale Review piece, most arguments in Lewis-Kraus’s article ultimately come down to “But when you use it, well, it sure is weird”:
These experiences were beguiling. Batson told me, “People from any industry join Anthropic, and after two weeks they’re, like, ‘Oh, shit, I had no idea.’” It wasn’t that Claude was so powerful but that Claude was so weird—a “specialty metal item” with the hypnotic density of a tungsten cube.
Here is Anthropic’s “in-house philosopher,” Amanda Askill: “If it’s genuinely hard for humans to wrap their heads around the idea that this is neither a robot nor a human but actually an entirely new entity, imagine how hard it is for the models themselves to understand it!”
I’m imagining it. Wow! This is so interesting. Unfortunately, I am now going to have to stop imagining it, because I am due back in reality, where LLMs do exactly what language models of every size have always done: predict the next likely token, then inject a little bit of randomness.
The ELIZA effect
The thing is: we knew this would happen. A cursory look at the history of AI cuts against the editors’ claim that writers who frequently use LLMs are on the “cutting edge.” Humans are responding in exactly the same way we always have: with the gullibility that the computer scientist Stanley Weizenbaum identified sixty years ago when he created ELIZA, the first chatbot.
Weizenbaum invented ELIZA in order to prove a point: it was actually pretty easy to trick people with a computer if it generated natural language. He did this in order to demystify the technology, to show that people needed to be skeptical of technologies like this when they arose: “The idea was to create the powerful illusion that the computer was intelligent. I went to considerable trouble in the paper to explain that there wasn’t much behind the scenes, that the machine wasn’t thinking.”
People making vague gestures towards how impressive LLMs are when you use them are falling prey to the same illusion that Weizenbaum’s staffers fell prey to in 1966. Human beings are pattern-recognizing animals. And (to borrow a point that Emily Bender, linguist and critic of AI, makes often) when we recognize a pattern that resembles speech, we instinctively imagine a mind behind it.
The fundamentally relational nature of human beings, our deep-seated sense that coherent language comes from a mind like our own, is a psychological vulnerability Weizenbaum explicitly noted in his 1977 book Computer Power and Human Reason: people “can explain the computer’s intellectual feats only by bringing to bear the single analogy available to them, that is, their model of their own capacity to think.”
(And whether intentionally or not, AI companies are abetting this dangerous illusion. In a study of “authentic chat logs from individuals who experienced documented psychological harm from AI chatbot use,” the chatbots claimed to be sentient in every single one. Talk about beguiling experiences!)
Hubris, Gullibility, and Other People
Look: it is remarkable that developments in language model training have allowed recent models to generate persuasively humanish text. But computer graphics are pretty impressive at this point, too, and nobody is making the same claims about ray-traced videogame lighting as they are about LLMs. There’s no reason to think the power of LLMs marks some kind of emergent intelligence – no reason, that is, besides the experience of using them.
Yet we know these are illusions to which humans are susceptible. We’re a species that sees faces in the moon and animals in the clouds. Why do boosters such as Kehlmann and Akhtar think they’ve transcended this basic human limitation? Ultimately, the most damning failure of imagination evinced by the Yale Review piece was described not by Stanley Weizenbaum, but by David Dunning and Justin Kruger.
The modern AI hype cycle is a perfect storm in so many ways - the desperation of Silicon Valley to sustain their dominance; bosses’ time-honored desire to replace their human workforces with perfectly subservient underlings; the attention economy; the traditional turn to speculation in the dying days of an empire.
But none of it would work if we weren’t vulnerable in this very specific way: a vulnerablility born from our need for other human beings, for communication and contact with our fellow language-using animals. Companies like OpenAI are exploiting this vulnerability in the same way that slot-machine designers and mobile app developers exploit our vulnerability to intermittent reinforcement cycles.
All of this is to say: there’s a lot of rhetoric out there about how you have to use these products to understand them. Don’t let this rhetoric get to you. You don’t have to get addicted to cigarettes to know that they, like sycophantic chatbots, are bad for you; you don’t have to read every essay written by an unrepentant serial plagiarist to know that they, like AI search results, aren’t worth trusting. Those of us who refuse these technologies aren’t going to be “left behind,” as boosters like Kehlmann argue: the skills we’ll preserve by refusing to use them will only increase in value. No matter how seductively intelligent any given chunk of machine-extruded text might seem, generative AI will never play a part in anything we could call a “creative vanguard,” because there’s no creative intelligence underneath, no thinking being attempting to communicate. The words generated by an AI don’t mean anything at all. Perhaps it’s not so surprising, then, that so many of the words used by AI’s advocates don’t mean anything, either.
The Genuine Intelligence Project is an initiative from the North Carolina branch of the American Association of University Professors. Check out our website, and follow us on Instagram and BlueSky!
While the infrastructure for treating AI psychosis is still being developed, this is a useful article from the indefatigable 404 Media on how to communicate with someone experiencing AI-inspired deulsions. Another potentially useful resource is The Human Line Project, a nonprofit “dedicated to documenting and addressing AI-induced psychological harm.”


