3 Comments
User's avatar
JAMES VON DER HEYDT's avatar

Holden. I like your thinking. (I had ChatGPT write this for me). Jim

Brian Wandell's avatar

The editorial rightly highlights the strain AI tools place on the peer-review and publishing pipeline. However, in evaluating this challenge, we should keep two critical perspectives in mind:

First, we must remember that human researchers make plenty of mistakes, too. Used responsibly, AI has the potential to help identify and correct some of those human errors before they ever reach a journal. Technology can be an ally in improving scientific quality, not just a source of "slop."

Second, and more fundamentally, the root of this crisis is not the technology itself, but the systemic rewards of modern academia. The primary metric for evaluating and promoting scientists remains the sheer volume of publications—particularly those in high-profile journals. By dramatically lowering the friction of writing and formatting, AI enables more people to generate more attempts to satisfy these institutional demands and advance their careers.

If we treat this purely as an AI detection and surveillance problem, we are treating the symptom. The real challenge is to reform how we evaluate scientific contribution so that we reward rigorous, reproducible quality over high-throughput quantity.

William's avatar

I'm glad you're talking about the academic publishing system being put under more stress from the increased rate of manuscripts generation afforded by LLMs. It's a big concern that some less scrupulous authors aren't going to rigorously check their references, but I think the improvement of tooling that does this automatically and maybe a few high-profile examples made of people who try to get away with it will sort that out.

The bigger issue is that academic publishing has been under strain for decades and no one has yet figured out what to do about it. I think publishers should be very careful to maintain the high editorial standards their brand is based on while all this is happening, because these models are consuming all the preprints and will happily rank them according to whatever criteria you provide, and mostly do a good job, too, if you set things up right.

On the public awareness side of things, we really need people like Zeynep to keep up. It's fine for random Bluesky accounts to parrot the line about how LLMs can't reason, but Zeynep should know better. Spend 15 minutes with Claude Code and tell me that what it's doing is just accessing patterns in it's training data while it's solving a problem it has never encountered before.

Better yet, go to Far.ai or CivAI.org or apolloresearch.ai and play with their demos showing how models learn, not just literal text patterns, but higher-order concepts. Did you know that if you train an AI system to give wrong answers to questions, it will be more willing to help you create a weapon or carry out fraud? It appears the model generalizes from giving wrong answers to doing other bad things. If you don't want to call that true reasoning, fine, but be careful you don't exclude a lot of what humans do when you try to draw the line.