Publishing-brain limits people's understanding of AI usefulness
Yep, it's not an answer box. But we figured that out years ago, what else you got?
I’m starting a podcast with a friend called Couch to 4k, where he and I buy films on physical media and talk about them for an hour or so each week. I say the goal of it is to get more people into physical media, but honestly it’s just an excuse to have an interesting conversation about films each week with someone I enjoy talking to, and eventually to invite on as guests people I’ve fallen out of touch with and talk to them as well.
Of course, I’m me, so I can’t do anything without research being a part of it. So with a discussion about Peter Jackson’s The Frighteners coming up, I asked Claude for a summary of the ways in which that film fit in with 1990s comedies in general. And Claude happily complied, centering its discussion on how the film exemplified the “surge of high-concept comedies” in the 1990s. “High-concept” for those who don’t know is a term that designates films where the plot engine can be summarized as a “concept” vs. more mundane character/situation dynamics, often framed as “thought experiment” films. In comedy, they can be described as “what-if” comedies. Happy Gilmore, where a hockey player joins a pro golf tour is high concept. Office Space, as comedy about a man who begins to question which boundaries of his job can be pushed is not. Claude gave me some great information for the podcast — the film is typical of that surge. And this is something I’ve heard from other people too, the 1990s were a surge of the what-if comedy.
Gold, right? But I used my SIFT powers here, and asked for evidence. In particular, I asked Claude to look at all comedies released by three major studios between1986 and 2015 and calculate the percentage of high-concept comedies in each year’s batch, using the queryable Wikidata database as a source.
And it turned out that, well, 1996 wasn’t particularly rife with high-concept comedies. High-concept was core for a long time; but the most interesting trend is a collapse of the high concept comedy 2011-2015, which might be (maybe?) a sign of the rise of the Apatow-style comedy, or could be due to something we’re not thinking about like share of rom-coms or animated offerings.
Now if you’re starting to ask — what if you did this at more granularity, what if you audited the AI decisions on what qualifies as a “what-if” comedy, what happens if we look outside these three studios or eliminate animation — that’s great. If you’re thinking “Maybe the real question is successful high-concept comedies,” that’s great. If you haven’t already, you should learn to use AI in this way, so that you can ask such questions directly as well. If these questions pop into your head, you’ll likely enjoy working with a tool like Claude.
But if you, like some people I see online, are immediately jumping to all the ways in which any understandings derived in this way are provisional, imperfect, and likely to have flaws underneath, and seeing those as an argument against the use of AI, you’ve lost the plot.
Most constructed arguments are not publication
Imagine you get a large flatscreen TV in a box. Let’s say a 90 inch TV. You take out the TV and fill the box with sand. You then take one grain of sand out of the box.
The grain of sand is the amount of arguments (evidence marshalled to support or oppose an assertion) people publish — a research paper, a news article, even a blog post. All the rest of that sand is arguments that are not published.
These numbers are made up of course. They are also radically conservative. Almost every single person on this planet constructs multiple arguments per hour from the age of about two onward. “Let’s have spaghetti with meat sauce tonight, the hamburger is about to go bad” is an argument. “Can I borrow your car tomorrow, mine doesn’t have enough room for the luggage” is an argument. “Why don’t we have Pam run the expense analysis, she did a great job on the toplines last month?” is an argument. And saying “Frighteners was part of a surge of high-concept comedies in the mid-1990s based a rough classification pass” is an argument.1
If you have actually written publications, you know this implicitly. The amount of unpublished, provisional, and often faulty arguments you make as you explore the ideas leading to publication are several orders of magnitude more than what goes into the publication.
So forget the sand-filled box as the metaphor. Fill your house with sand, and take a grain out. Or go to the beach. Publication — and the accuracy people expect from it — is both the bedrock of a literate society and also an outlier to most human experience where we are looking for just a little bit of data to test our assumptions, make a point, inform a decision, or explain how we see the world.
Knowledge may be a thing, but finding it is a process
So what is the point of this?
Early discussions of the accuracy of AI, back in 2023, focused on how laughably bad it was. As we hit the “LLM + tools” era (search, coding, wikidata, etc) I’ve seen much of the critique of the accuracy of these systems shift. These systems get most things right and some things quite wrong. Much of the current critique compares this output against a standard of publication accuracy: what does this get right, and how does that accuracy compare to published results?
But this gets the question wrong, in ways that should be obvious to people in education. The question is not whether the answers you get are perfect (there are no perfect answers to most things anyway). The question is whether a person who learns to use these tools will have more evidence-informed opinions on things than someone who does not.
In almost all cases, neither the AI result nor the decision/statement it informs is going to need a publication-level buttoning-up. All this stuff is the grains of sand in the box.
More importantly, the questions it raises are iterative, like all knowledge exploration. Is “40 Year-Old Virgin” a high concept “what-if” drama, or something else? When you understand information as a process the fact AI gets you to a granularity where you are asking that question and critiquing its decision is not a downside to AI use, but a benefit of using the tool. In engaging with the data you start to move in-and-out of definitional questions, build better understandings, and ultimately unique insights.
I worry writing this that there are small-minded people who will mock what I say next, but the fact that first passes on data have obvious errors is a benefit, not a drawback, from an exploration perspective. If ever you’ve done research with human coders categorizing data you know that much research insight is generated by looking at coding passes that seem wrong and thinking through what definitional questions it raises. You run passes on data, think about what ways the pass didn’t capture data in a way that usefully addresses your question. You redefine your guidance and make another go at it. It’s not necessary to have a perfect pass to get that process started. You just have to have a pass. In fact, an initial bad categorization (and your reaction to it) is most likely to produce your most valuable theoretical insights.
At the risk of pre-butting too many ill-thought out arguments, I know that some will read this and say well, the problem is that the companies have presented this technology as an answer box. To that I’d just say yelling over and over again “It’s a bad answer box!” simply accepts the corporate frame that it is meant to be an answer box, and our argument is over how good at that function it is. Accepting that frame is a horrible disservice to students, who need to conceptualize knowledge discovery and the exploration of ideas in more grown-up ways. It used to be pretty uncontroversial that our job as teachers was not simply to scold bad practice, but model good practice. To do that you do have to engage with these tools and learn how to use them so that you can model that for students. Most of what you need to know is not about the tool, but about thinking through things like classification. But you still must understand how the tools leverage that existing knowledge.
Taking an evidence-based approach to the world is a process, and one that benefits from different levels of precision in different circumstances. It’s hard to break publication-brain and realize that there are many gradations of certainty in between knee-jerk reaction and robust meta-analysis. But it should be familiar; it’s been in the job description forever.
See the comedy data here.
I subscribe to the idea, once debated but now relatively common in epistemology that explanations (here are the reasons I believe something) are not structurally distinct from arguments (here are the reasons you should believe something) and any firm division between the two is a matter of intent and explanations are therefore structurally arguments when they marshal facts.


Fantastic piece. Thank you.
Always interesting, but I can testify from my travels talking to people about LLMs in education that the reality that they aren't an answer box has not penetrated in most places.
There is much more work to be done convincing - in many cases - the people who run education institutions that it is not an answer box before your frame as this kind of tool can even be realized.
I think the idea that one "has to" model these tools to engage in good research practice is overlooking the volume of practice without engaging with them that is required to work with the tools in this way. The knowledge that underpins your explorations is significant and when working with students it is a challenge to get them working from a place of that kind of deep knowledge.