Always interesting, but I can testify from my travels talking to people about LLMs in education that the reality that they aren't an answer box has not penetrated in most places.
There is much more work to be done convincing - in many cases - the people who run education institutions that it is not an answer box before your frame as this kind of tool can even be realized.
I think the idea that one "has to" model these tools to engage in good research practice is overlooking the volume of practice without engaging with them that is required to work with the tools in this way. The knowledge that underpins your explorations is significant and when working with students it is a challenge to get them working from a place of that kind of deep knowledge.
I agree with the point that AI (or anything else) doesn't need to be _accurate enough for publication_ in order to be _accurate enough to be useful_. But I think there's an important effect that means that its unreliability may matter more than one would otherwise expect.
There are only a few top-level AI systems, and they are all rather similar. This means that if everyone uses them as research assistants, the effect is rather as if everyone somehow managed to hire the same research assistant.
This means that the errors and omissions those research assistants perpetrate will be _highly correlated_, much more so than the ones an unassisted human (or a whole lot of actual human research assistants, but most of us can't afford those) would perpetrate. Which means that the amount of harm they can do to the whole intellectual community is greater.
Imagine (this is a deliberately ridiculous example) that all the AIs have a blind spot around leopards, and will tend to fail to find information when it happens to involve leopards. Then in a world where everyone uses the AIs to do a lot of their research, everyone's research will tend to have a similar blind spot around leopards, and that will make its way into everyone's thinking.
In reality, it probably isn't leopards. But it could be something similarly ridiculous -- e.g., many recent OpenAI models have had a tendency to get obsessed with goblins. Or it could be something much more subtle (and hence easy to miss) -- some particular kind of thinking the models are bad at, so that facts that are more easily exposed that way will get neglected everywhere. Or, because AI models' minds are so different from ours, it could be a blind spot of a "shape" that doesn't match anything humans would do, which again could be very easily missed.
And so the whole intellectual community could end up bent out of shape by the _shared_ deficiencies of the AI models, where they wouldn't have been by human research assistants who are no more reliable but are more independently unreliable.
I'm not sure whether this analogy is helpful or not, but: it's a bit like the reason why widely-shared prejudices like racism and sexism are disproportionately harmful. If one employer happens to hate tall people and prefers not to hire them, they might do about as much harm to tall people as a racist employer does to black people. But the next employer is just as likely to favour tall people and hate short people, and on the whole it cancels out. But you might plausibly get a whole town where _all_ the employers are prejudiced against black people, or Jewish people, or women, or gay people, and if so then the situation for those people is very very bad. Which is why there are a lot more laws saying "you aren't allowed to refuse to hire black people" than ones saying "you aren't allowed to refuse to hire tall people". The issue with AIs-as-research-assistants is a bit like this, except that instead of "this group of people is harmed disproportionately" we have something more like "this area of knowledge is neglected/misunderstood disproportionately".
(It is probably also the case that AI models have prejudices -- I have heard people claiming very confidently that they are prejudiced against black people, and others claiming just as confidently that they are prejudiced against white people, and whether or not any of that is true it would be surprising if they managed to have no prejudices at all -- but that _isn't_ the point I'm making here, though one could say about prejudice things similar to what I have been saying about error. I'm using prejudice simply as an analogy.)
The way this insight has occurred to me recently is that after seeking Claude‘s advice about something and thinking it was pretty good, a few minutes later, I thought: wait a minute, it’s just making things up! But it didn’t matter since it helped me think through the issue and gave me much more to work with. You draw out the larger implications of this effectively.
Great piece. Two things I try to keep in mind when I use LLMs is that... 1) they don't have a true semantic understanding of the input or output, and 2) they're pretty terrible at things that require judgment, which humans only get from experience. They are better and better at faking it though.
But yes, that means their proper place is exactly how you described it—surfacing ideas, evidence, arguments, and sources (very efficiently) for us to dig into. The digging is the thinking! I worry about a world where people would even want to hand that over to machines.
This is really just another version of "distant reading", which was a data-driven strategy for trying to talk about what the total space of publication (or cinema) looked like in the past when you took a step away from the strong narratives we have from canon formation via literary (or film) criticism. I don't know if anyone has used generative AI yet to say "give us a narrative hook for all the underrepresented or unrepresented totality of a publication/dissemination space via the rich data records we have of publication over a past time span" but it sounds like a valuable project and yet one that I'm certain will fail, because it's asking generative AI to provide semantically rich descriptions of works that have not been examined in semantically rich ways.
Fantastic piece. Thank you.
Always interesting, but I can testify from my travels talking to people about LLMs in education that the reality that they aren't an answer box has not penetrated in most places.
There is much more work to be done convincing - in many cases - the people who run education institutions that it is not an answer box before your frame as this kind of tool can even be realized.
I think the idea that one "has to" model these tools to engage in good research practice is overlooking the volume of practice without engaging with them that is required to work with the tools in this way. The knowledge that underpins your explorations is significant and when working with students it is a challenge to get them working from a place of that kind of deep knowledge.
I agree with the point that AI (or anything else) doesn't need to be _accurate enough for publication_ in order to be _accurate enough to be useful_. But I think there's an important effect that means that its unreliability may matter more than one would otherwise expect.
There are only a few top-level AI systems, and they are all rather similar. This means that if everyone uses them as research assistants, the effect is rather as if everyone somehow managed to hire the same research assistant.
This means that the errors and omissions those research assistants perpetrate will be _highly correlated_, much more so than the ones an unassisted human (or a whole lot of actual human research assistants, but most of us can't afford those) would perpetrate. Which means that the amount of harm they can do to the whole intellectual community is greater.
Imagine (this is a deliberately ridiculous example) that all the AIs have a blind spot around leopards, and will tend to fail to find information when it happens to involve leopards. Then in a world where everyone uses the AIs to do a lot of their research, everyone's research will tend to have a similar blind spot around leopards, and that will make its way into everyone's thinking.
In reality, it probably isn't leopards. But it could be something similarly ridiculous -- e.g., many recent OpenAI models have had a tendency to get obsessed with goblins. Or it could be something much more subtle (and hence easy to miss) -- some particular kind of thinking the models are bad at, so that facts that are more easily exposed that way will get neglected everywhere. Or, because AI models' minds are so different from ours, it could be a blind spot of a "shape" that doesn't match anything humans would do, which again could be very easily missed.
And so the whole intellectual community could end up bent out of shape by the _shared_ deficiencies of the AI models, where they wouldn't have been by human research assistants who are no more reliable but are more independently unreliable.
I'm not sure whether this analogy is helpful or not, but: it's a bit like the reason why widely-shared prejudices like racism and sexism are disproportionately harmful. If one employer happens to hate tall people and prefers not to hire them, they might do about as much harm to tall people as a racist employer does to black people. But the next employer is just as likely to favour tall people and hate short people, and on the whole it cancels out. But you might plausibly get a whole town where _all_ the employers are prejudiced against black people, or Jewish people, or women, or gay people, and if so then the situation for those people is very very bad. Which is why there are a lot more laws saying "you aren't allowed to refuse to hire black people" than ones saying "you aren't allowed to refuse to hire tall people". The issue with AIs-as-research-assistants is a bit like this, except that instead of "this group of people is harmed disproportionately" we have something more like "this area of knowledge is neglected/misunderstood disproportionately".
(It is probably also the case that AI models have prejudices -- I have heard people claiming very confidently that they are prejudiced against black people, and others claiming just as confidently that they are prejudiced against white people, and whether or not any of that is true it would be surprising if they managed to have no prejudices at all -- but that _isn't_ the point I'm making here, though one could say about prejudice things similar to what I have been saying about error. I'm using prejudice simply as an analogy.)
The way this insight has occurred to me recently is that after seeking Claude‘s advice about something and thinking it was pretty good, a few minutes later, I thought: wait a minute, it’s just making things up! But it didn’t matter since it helped me think through the issue and gave me much more to work with. You draw out the larger implications of this effectively.
Great piece. Two things I try to keep in mind when I use LLMs is that... 1) they don't have a true semantic understanding of the input or output, and 2) they're pretty terrible at things that require judgment, which humans only get from experience. They are better and better at faking it though.
But yes, that means their proper place is exactly how you described it—surfacing ideas, evidence, arguments, and sources (very efficiently) for us to dig into. The digging is the thinking! I worry about a world where people would even want to hand that over to machines.
This is really just another version of "distant reading", which was a data-driven strategy for trying to talk about what the total space of publication (or cinema) looked like in the past when you took a step away from the strong narratives we have from canon formation via literary (or film) criticism. I don't know if anyone has used generative AI yet to say "give us a narrative hook for all the underrepresented or unrepresented totality of a publication/dissemination space via the rich data records we have of publication over a past time span" but it sounds like a valuable project and yet one that I'm certain will fail, because it's asking generative AI to provide semantically rich descriptions of works that have not been examined in semantically rich ways.