The word could appeared more often in scientific abstracts after December 2022. It also became less common.
Before that point, could appeared 60,877 times in the abstracts I analyzed. Afterward, it appeared 73,220 times, which looks like an increase. BUT, the later corpus contained 58% more words (naturally with the vast growth in academic writing in the era of AI). Relative to the amount of scientific writing, could actually fell from 0.487 to 0.371 uses per 1,000 words, a decrease of about 24%.
I think that we have all noticed, in one way or another, that generative AI has changed the way we write, right? I had a more specific hunch about that. I felt that we had lost a lot of hedging in academic writing (words such as may, might, and could), and maybe in writing overall, especially compared with boosters, the words that strengthen a claim (such as demonstrate, establish, and confirm).
The question is: why do I care? And of course, simply said, I really find it interesting, and it is a personal taste, I myself like to include hedging words in my writing, scientific or not, I feel it makes the message real, and makes the reader think more. But, with that said, hedges are words such as may, might, could, suggest, and possibly. They reduce how strongly a sentence commits to a claim. Boosters do the opposite: demonstrate, establish, confirm, clearly, prove.
This is more than style, and it affects how the reader perceives a message. For instance, take this: “The treatment may reduce mortality” and “the treatment reduces mortality” are different scientific claims. The first leaves room for uncertainty in the design, sample, or estimate. The second closes some of that room. Sometimes the hedge is just padding. Sometimes it is carrying the most accurate part of the sentence.
Testing the hunch
So I tried to do some testing. I used Paperclip, a research tool from GXL that lets AI agents search and read millions of scientific papers efficiently (thank you James Zou, Christine Lemke, and the team). I analyzed 1,804,784 abstracts from arXiv, bioRxiv, and medRxiv, containing 322,595,774 words from January 2020 to April 2026. Or, to be accurate, I let Codex do most of the counting.
The pooled result broadly matched the hunch. Hedges fell 6.1% per 1,000 words. Boosters rose 31.4%. The hedge-to-booster ratio fell 28.6% (as you can see in the figure below).
The boosters were not mainly words such as undoubtedly or certainly. They were more often achievement words: highlight rose 124.2% (wow!), demonstrate 46.1%, establish 44.4%, and reveal 43.2%. The writing did not simply become louder. It became more likely to announce that the paper had shown, established, or revealed something.
At the same time, extreme certainty words such as certainly, undoubtedly, always, and clearly also declined. The pattern was more specific than a general rise in confidence: less provisional language inside claims, more language of achievement, and fewer of the loudest declarations. Science may be moving toward a confident middle, which is probably harder to notice than obvious hype.
That pooled comparison is descriptive. The stronger test followed each archive month by month and asked whether its existing trend changed after December 2022. The hedge-to-booster trend shifted downward by 13.1% per year in arXiv, 8.5% in bioRxiv, and 12.9% in medRxiv, relative to each archive’s pre-2023 level. All three remained statistically significant after correction.
How far can we take this?
For extra measure, I added a smaller full-text analysis of 110 first-version papers. It pointed in the same direction: fewer hedges, more boosters, and a lower ratio. The uncertainty was much wider, though. None of the three corrected primary tests was statistically significant. I read that as a sensitivity check, not confirmation.
Of course, this cannot tell us whether LLMs caused the change. We do not know which authors used a model, what they used it for, or whether they used one at all. Many other things changed around the same time. The analysis shows a shift in language around this period, not its cause, and to be honest with you, linguistics is not by any means my field of research. I also do not think that less hedging is automatically worse. Some cautious phrases say very little and direct writing can be clearer. The question is which hedges disappeared. Were they disposable habits, or were they words that accurately represented the limits of the evidence?
Word counts cannot answer that. We would need to match each claim to the study design and results behind it, and ask whether the confidence of the sentence fits the confidence of the evidence. That is a smaller and much harder study. For now, I think something in how we write science has shifted.. and an essay about disappearing hedges should probably keep at least one maybe of its own.