Why LLMs give you Olive Garden when you want Carbone
Talking with James Evans of UChicago about AI's effect on science and the literature
One of the people at the absolute forefront of understanding the interaction of artificial intelligence, most recently large language models (LLMs), with knowledge and the scientific literature is computational social scientist James Evans at the University of Chicago. Evans writes a few times a year for us at Science in our Expert Voices feature, and has been publishing at a high rate on the explosion of AI onto the scene of the scientific literature.
Evans was someone I really wanted to talk to about all of this. While he is firmly in the “AI is here to stay” camp, he is still quite skeptical of what it can do. Every time I talk to him, I feel more comforted, not less, that humans will always be indispensable for research. The two main reasons are that (a) LLMs will always make errors and (b) by their very nature, they’re not going to be creative.
When I tracked him down, Evans had just schlepped across the country from Boston to Palo Alto to give two talks at Stanford. Here’s the video of our conversation:
Precise position or velocity, but not both
When I asked him what it was like to be working in this field right now, he talked about what he calls the “Heisenberg uncertainty principle” of AI, which is that at the very moment that you can measure something in the field, an AI agent is simultaneously changing whatever that is. Evans cleverly refers to the AI agents and chatbots that most scientists are using right now as “retail AI” compared to LLMs that are custom-built for specific tasks.
When I asked him if we were right to be worried that retail AI was degrading the scientific literature, he agreed that we were, saying:
This is a major cause for concern because in the same way that humans and human labs are, for the most part, black boxes, this is a new kind of black box. And so it’s really increasingly hard to evaluate the difference between really strong science and weak science, because if [the LLM is] tuned to writing and it’s reinforced against what people expect, then it’s easy for it to generate science that appears solid and strong and novel while dressing up work that’s potentially fragile and weak and particular.
I asked him if there were reasons to be optimistic at the same time, and he pointed to the potential (but not reality) of AI to connect distant fields in ways that would give us new ideas.
Reasons to be interested and optimistic are [that] AI so far is able to read the vast sweep of the literature, and so it’s able to bring together hypotheses from one field to problems in another field, alien patterns it can kind of bring potentially together. I think largely this has been a promise rather than a reality.
Never-ending pasta bowl
This got me excited. One of the things that I’ve been hoping that AI would be able to do is to elevate papers that are important but not published in high-profile journals. Because (and yes, this is good for Science from a business perspective) Google page-rank is always going to surface papers in high-profile journals that get a lot of citations and news stories, but there are many equally important papers that may have very similar results and might even predate the high-profile ones. I wanted to know whether current AI would allow us to do this, and I was told I had to wait. Retail AI currently is most likely to reach to the “center” of fields, not to the edges, which will give you the most common answer, just like a Google search. As Evans says:
If you don’t know anything about Italian restaurants in some city in the US, then it’s going to return Olive Garden. It’s going to return the most voted average center of the distribution, which is not going to be the most interesting Italian restaurant in any US city. I think agentic approaches where you’ve got different agents with different personas and different tools that have specific criteria definitely have the potential to reach out and, and read and surface new parts of the literature that [we] haven’t seen before. On average, that’s not what we see is happening.
Dang. I’m excited for someone to solve this. But I guess I can enjoy the unlimited breadsticks in the meantime.
Where current policies in journalism and scientific publishing stand
As far as searching and analyzing the literature and media are concerned, most organizations have taken the position that using LLMs is permissible as long as the human author takes responsibility for the content. Here’s a handy spreadsheet of Science’s policies for research articles; using AI to search the literature is permissible and authors do not need to declare that they were used. Similarly, the editorial standards for our News from Science division allows use of AI in background research, saying “Reporters may use AI-based tools to transcribe interviews and in their background research, but we expect reporters to check with their editors about other uses of AI.” This is a similar stance to nearly every major news outlet: here is the Associated Press standard, which goes further than News from Science in allowing AI to suggest headlines.
I asked Evans about this, specifically whether we knew where human ideas end and ideas partly attributable to AI begin. He believes we’re well past the point of knowing that. It’s human nature to think that the very clever prompt we gave to the LLM was the thing that created the idea, but the responses that we get undoubtedly shape how we think about the topic. Similarly, even people who are not using LLMs directly are interacting with other people who are, and they’re also reading scientific papers and news stories that are influenced by AI. So, no one who says they never use AI is completely immune to its influence.
There are strong feelings out there about the ethics of all of this, particularly about where the data in the LLMs came from in the first place and the energy demands of the models. Those debates are not going to calm down any time soon. In the meantime, the most important thing we can do is to understand what AI is doing to the knowledge that we are creating, and for this, I hope you’ll watch my whole conversation with James.
Also in the meantime, here’s a New Yorker piece about how to get a table at Carbone New York (they’re not on OpenTable). And here’s their wares, which look better than the OG (which is not, actually, an OG).




