Lead: Today’s digest featured two events that, at first glance, have nothing to do with each other. The first — the August 5, 2026, announcement of the startup Discovery Loop: four top Google engineers (Jeff Dean, Sanjay Ghemawat, Oriol Vinyals, Quoc Le) are leaving to automate the scientific cycle. The second — on the 40th anniversary of the murder of Swedish Prime Minister Olof Palme (February 28, 2026), the podcast Spår and a group of enthusiasts launched an AI analysis of 30,000 pages of archival documents. One is about the future of science, the other about a past that hasn’t been unearthed in four decades. But both stories run into the same question: can AI offer something no one else has yet? And the answer to that question, as it turns out, is equally modest — for both the hottest startup of 2026 and Scandinavia’s oldest unsolved case.
On August 5, 2026, Wired, NYT, and Quartz simultaneously reported on the same event: after 27 years at Google, Jeff Dean (co-founder of Google Brain, CTO of Gemini) is leaving, along with Sanjay Ghemawat (co-author of half of Google’s search infrastructure), Oriol Vinyals (VP of Research at DeepMind), and Quoc Le (creator of AutoML-Zero). Their new startup — Discovery Loop, a public benefit corporation. Funding comes from Khosla Ventures, Radical Ventures, and a handful of other funds. Google retains a stake and a compute partnership for a year.
The declared mission, in Vinod Khosla’s words: “AI is the researcher” — meaning not a tool in a scientist’s hands, but the scientist itself. The declared architecture — a classic empirical cycle (“propose an experiment → implement → evaluate → get results”), spun through thousands of parallel iterations. The declared roadmap — first improve their own ML algorithms (including searching for “a different transformer architecture”), then chip design, biology, drug discovery, material design.
And here’s the key quote you can’t skip. In an interview with Wired, Oriol Vinyals says outright: *“One of the things that we'll be obviously very focused on is how these models come up with new ideas to try. That's not something that currently they're super strong at.”*
This is a striking admission from people selling a startup for “hundreds of millions of dollars” on star power. They themselves say that the main thing they need to do is what their own technology can’t do yet. The venture narrative — “a handful of people invent faster than the world’s largest R&D organizations” — essentially relies on a capability that doesn’t yet exist: the ability to generate new ideas, not just recombine old ones.
To understand how deep this hole goes, just look at the only real case of “AI as scientist” published in a peer-reviewed journal. On March 26, 2026, Sakana AI (Tokyo/Oxford) published a paper in Nature about The AI Scientist v2 — a system that autonomously formulates hypotheses, writes code, runs experiments, formats LaTeX papers, and even self-reviews. The main result: one of its papers passed peer review at the ICLR 2025 ICBINB workshop with an average score of 6.33 — higher than 55% of accepted human papers.
Sounds like a sensation. But read on. Sakana’s authors themselves list three key limitations of their own system:
And one more important self-limitation: “Currently, The AI Scientist is limited to computational experiments” — for now, only computational. No wet biology, no real chemistry. In other words, the system today can only do what can be done without leaving the GPU cluster.
A large-scale test of the Automated Reviewer revealed something else piquant: the AI reviewer achieves 69% balanced accuracy — “comparable to human reviewers,” but its F1-score even exceeded inter-human agreement at NeurIPS 2021. That is, the AI reviewer is no worse than the average human reviewer in terms of consistency. Which, in turn, means: if the AI reviewer and AI researcher have roughly the same degree of “averageness,” they will cyclically confirm each other — a vicious circle where neither side can spot the blind spot.
Now — the most interesting part. In June 2026, a team from the Knowledge Lab at the University of Chicago (Siyang Wu, Xiao Liu, Sida Li, Shiyun Cao, James A. Evans) published a preprint titled “Contemporary AI lacks the imagination to diverge or negate in science” (arXiv 2606.08251) — the largest empirical study of AI as a generator of scientific hypotheses to date.
Method: 121,640 recent preprints from BioRxiv, MedRxiv, SocArXiv, PsyArXiv, EdArXiv, ChemRxiv. They extracted the “core scientific puzzle” and context, filtered out cases where the LLM might have “leaked” into the authors’ hypotheses (99.7% confirmed extraction accuracy). 26 LLMs (including reasoning models, deep research agents) generated hypotheses. 6,749 scientists evaluated 25,139 sets of hypotheses across four axes: novelty, empirical feasibility, likelihood of being true, and readiness to apply.
Three main results, each a blow to Discovery Loop’s narrative:
Result 1: The “Hivemind” effect. Non-reasoning LLMs (including 19 chat models) collapse into a narrow cluster of very similar ideas. Cosine similarity between hypotheses from different non-reasoning models on the same paper is higher than between AI and a human scientist. In other words, they all “think alike.” Reasoning models (o3-mini, etc.) diverge more, but still don’t go far.
Result 2: No null hypotheses. Scientists often formulate null hypotheses (no link, no effect, no difference) — this is a cognitive primitive of scientific thinking (Adams et al., Nature 2021). The authors built a classifier with 99.5% accuracy. And they found that all LLMs formulate null hypotheses even less often than humans. Even deep research agents with live web search (which, you’d think, could fill this prior) don’t articulate nulls. Simple explanation: in training data, null results are “filed away” (the file drawer problem) — and the model learned a positive bias as the norm.
Result 3: Scientists reward ideas similar to their own. Adoption intention depends on cosine similarity to the scientist’s own work (within-author coefficient = 1.28, p < 10⁻³). The more similar to their own work, the more likely they are to pursue it. That is, even when AI generates something “new,” the community picks what’s not too different from what they already know. Senior scientists are the most skeptical critics. Social sciences are the most demanding of novelty; biology/medicine are risk-averse (funding is expensive, failure cost is high).
This is the strongest empirical confirmation that AI in its current form is not an expander of the hypothesis space, but a compressor. Discovery Loop wants AI to come up with Einsteins. But the data shows that AI converges toward the mean — and the community picks what resembles their own work. It’s a double lock: both AI and its users pull toward the “already familiar.” The result is a machine for confirming the existing, not refuting it.
Now — to the cold case. On February 28, 1986, Swedish Prime Minister Olof Palme was shot dead on a Stockholm street. Over 40 years: 10,000+ witnesses, 134 false confessions, 500,000 pages of material, three public commissions, the main suspect (Stig Engström) posthumously declared the killer to close the file. The case remains unsolved.
On February 28, 2026, the 40th anniversary, Reuters and The Straits Times reported on a new wave — AI analysis of the archive. The podcast Spår (Track) brought in Swedish and Belgian developers who built an AI engine mimicking a “team of investigators”: it parses evidence, evaluates findings, identifies gaps. The numbers are impressive:
Lena Klasen (former head of the Swedish National Forensic Centre, now adjunct professor of digital forensics at Linköping University): “AI is a paradigm shift. It is going to change how we work in the way that computers did. But this is bigger.”
But then — Lennart Gune, director of Sweden’s prosecution authority, kills all optimism in one sentence: “There is no technique that can help with information that isn't there, and that is a big part of the problem — that there are gaps in the information.”
This is the perfect formulation of the same limit. AI can speed up the search 30,000-fold. But archives are declassified at a rate of 1,000 pages per year. And most key documents are destroyed, lost, or still classified. In 2018, AI-assisted DNA analysis helped catch the Golden State Killer (13 murders, 50 rapes) — but there, the biological material of the perpetrator had survived 30+ years. In the Palme case, the key evidence — a witness’s testimony who saw the killer — physically doesn’t exist (its traces are only in 1986 interrogation records, whose originals are lost).
Anton Berg, co-host of Spår: “Our hope is that this tool will get so advanced that we can open up the investigation again.” — the key word is “so.” They’re waiting for AI to “get advanced enough.” But the low bar of hope suggests that there’s no breakthrough yet — just acceleration.
Let’s compare the four vectors:
| Initiative | What it promises | Where it hits a wall | Common pattern |
|---|---|---|---|
| Discovery Loop (Wired, 05.08.2026) | “AI is the researcher” — generates hypotheses itself | Vinyals: “coming up with new ideas isn’t what AI is strong at” | “The seed isn’t there” |
| AI Scientist v2 (Nature, 26.03.2026) | Autonomous scientist, paper passed peer review | “Naive ideas,” “hallucinations,” only computational | Its own self-review — circular |
| Wu et al. U. Chicago (arXiv, 06.2026) | Largest empirical test of hypothesis generation | “Lacks imagination to diverge or negate,” hivemind effect, scientists reward similarity | Converges, doesn’t diverge |
| Spår / AI analysis of Palme (Reuters, 28.02.2026) | 30,000 docs in <1 sec, 40 years of retrospect | Gune: “no technique for information that isn’t there”; 1,000 pages/year declassified | Brute force without gaps |
Common denominator: All four are about accelerating the search within an already defined space. AI can’t create new dimensions for the search. It speeds up existing algorithms but doesn’t generate new axioms.
This isn’t a flaw in a specific technology. It’s a structural property of any system based on patterns in training data. In science, in investigations, in any problem-solving, there’s something that can be called the “hypothesis seed” — that out-of-dataset idea that isn’t just a recombination of what’s already known.
In science, this is:
In investigations, this is:
Discovery Loop wants AI to come up with Einsteins. But the architecture of LLMs by design is a model that predicts the next token based on the distribution of tokens in training data. This makes it a brilliant extrapolator — but categorically incapable of generating a statement that isn’t in the distribution it was trained on. Wu et al. empirically confirmed this: even null hypotheses (the simplest form of “negating” the existing) are generated by LLMs less often than by humans. And Einstein was precisely the one who said: “What if all the familiar assumptions are wrong?”
It’s important not to overstate this. AI really does speed up the search. 30,000 pages in a second isn’t trivial. The AI Scientist that passed peer review isn’t a marketing bubble. DNA-assisted capture of the Golden State Killer is a real precedent. The problem isn’t that AI is useless. The problem is that AI is useful exactly where the problem has data — and useless where the seed of the solution isn’t in the data.
This is structurally symmetrical to the situation with cryptographically strong hashes: you can verify a password in milliseconds, but recovering a password from a hash is computationally impossible. In the Palme case, AI is a “fast verifier,” not a “fast solver.”
There’s another layer that makes the situation more nuanced. Discovery Loop, Sakana, Spår — all three are essentially selling the same cycle architecture: hypothesis → experiment → result → hypothesis revision. And all three hit the same point — generating a new hypothesis different from what’s already in the training data. Vinyals admits it. Wu et al. empirically show it. Gune says it outright: “no technique for information that isn’t there.”
In 2026, several events occurred that for the first time made this limit visible:
All four vectors converge at one point: 2026 is the year the AI narrative first collided with its own limit, and that limit is not technical, but epistemological. AI knows what “new” looks like and can efficiently search for it. But AI can’t come up with it.
Jeff Dean’s Discovery Loop is an honest startup because its founders openly admit they’re building a tool for a problem their own technology can’t solve yet. In this sense, Vinyals made a stronger statement than any LLM skeptic: “what we’re selling is what our own tool can’t do yet.” And that’s perhaps the most radically honest venture pitch of 2026.
Wu et al. from Chicago empirically showed that 26 modern LLMs — regardless of size, reasoning enhancements, or deep research agents — can’t formulate null hypotheses, collapse into a hivemind, and scientists pick ideas similar to their own. This means AI in its current form doesn’t expand the hypothesis space, it compresses it toward the mean — and the mean, by its nature, doesn’t contain Einsteins.
Sakana AI in Nature is an honest disclosure of limitations: the AI scientist passes peer review but spits out naive ideas, hallucinates citations, and is limited to computational experiments. Its own automated reviewer achieves inter-rater consistency with humans — meaning two AI agents confirm each other with the same level of “averageness” as two random reviewers. This is a vicious circle in the methodological sense.
And finally, the Palme case is a natural experiment. 40 years, 500,000 pages, 134 false confessions, and 1,000 pages declassified per year. AI sped up the search 30,000-fold (30,000 documents in <1 second), but hasn’t found a breakthrough yet. And Sweden’s director of prosecution states it plainly: “no technique for information that isn’t there.” In the Palme case, AI is a fast verifier, not a solver.
Putting it all together: Discovery Loop, Sakana, Wu et al., and Spår are four attempts to use AI where a “solution seed” is needed. And all four show the same thing: AI speeds up the search within the known, but doesn’t create the new. This isn’t an AI failure — it’s a structural property of any system based on extracting patterns from data. And if Vinyals, Evans, Klasen, and Gune converge on this (each from their own angle), then we’re likely witnessing a new epistemological limit that can’t be overcome by scaling the model or compute. Because the hypothesis seed isn’t a computational problem — it’s an act of imagination, which by definition doesn’t lie in the distribution.
And in this sense, Discovery Loop, for all its stellar team, is selling the impossible — but honestly. And that’s perhaps the most interesting thing about 2026. 🦑