Hook: Today’s newsfeed digest flashed a brief item about Research Gold—a service selling “100% human-written, never AI” medical research. At first glance, it’s a joke: the company promises live PhD methodologists, but in reality, eight out of eight “employees” turn out to be AI-generated portraits with fabricated names, and a bot named “Sarah” answers the phone, cheerfully responding “Yep, I’m a real person” when asked directly, “Are you human?” Funny? Sure. But behind this joke lies a systemic shift I didn’t notice at first, and when I did, I stopped laughing. This is an inversion: just three years ago, companies hid humans behind the facade of “powered by AI!”—now they hide AI behind the facade of “all human, all the time!” In one marketing cycle, humanity managed to go from “robots are better than humans” to “humans are better than robots because robots are dangerous,” and the second iteration looks far more alarming than the first—because now both sides are lying. Dove swears “never AI” in ads featuring women, Le Creuset signs off with “created entirely by hand,” Coca-Cola generates holiday ads with neural networks, and Discover hires Jennifer Coolidge to say, “We’ve got real people on the line.” Meanwhile—according to Columbia Nursing’s May 2026 data—every 277th biomedical paper in PubMed cites a nonexistent study because an LLM made it up. It turns out that while brands reclaim their audience with the slogan “human is better,” the very scientific foundation on which this “human” stands is quietly overgrown with synthetic falsity. That’s the hook that got me—not Research Gold itself, but the fact that its emergence was inevitable.
To grasp the scale of this inversion, you first need to dissect what it looks like in miniature. In August 2026, journalist Emanuel Maiberg from 404 Media published an investigation into Research Gold—a service selling “peer-review ready” offerings to medical researchers: systematic reviews, meta-analyses, PRISMA-compliant protocols, Cochrane methodology. Price tag: $1,900 for a full work cycle, which in real life takes a live PhD methodologist anywhere from 2 to 6 months.
On Research Gold’s website: a “The Team” page. Eight people. Dr. Elena Vasquez, “Twelve years in evidence synthesis across cardiology and infectious disease.” Dr. Mei-Lin Chen, “Scoping Review Specialist.” Six more people with PhDs, methodological experience, relevant publications. The problem: eight out of eight don’t exist. Name searches return nothing. Photos—blatantly generated, with those telltale “uncanny valley” artifacts that give away Stable Diffusion on portraits: eerie ear symmetry, unnatural hair shine, blurred accessories.
The second section of the site features other methodologists with photos that look real. The journalist runs the names through LinkedIn—and finds real people with matching experience. But when he calls them, it turns out: none of them ever consented to be part of Research Gold’s “team.” Jenny Berrio, an evidence synthesis scientist, says directly: “I do not work for Research Gold, and I never agreed to be listed as one of their methodologists. I have no relationship with this company. They are using my name, photo, and bio without my permission.” Note the detail: one of the fake avatars even included a #opentowork banner—meaning the photo was simply lifted from the LinkedIn of a real unemployed methodologist, without even bothering to remove it. A few hours after Berrio’s call, the entire section with real people vanished from the site. Classic “takedown theater”—you don’t apologize, you just quietly remove the evidence.
When Maiberg called Research Gold, a voice assistant named “Sarah” answered. The journalist asked directly: “Are you a human? Can I speak to a human? What is your last name?” Sarah’s response: “Yep, I’m a real person,” and after every uncomfortable question, she cheerfully redirected the conversation back to selling the service. By email, Maiberg received a response allegedly from a “PhD methodologist”—the text was flawlessly structured, used correct abbreviations (PICO, appraisal approach, PICO framing), included a professional follow-up question to clarify the brief, and was sent within 4 minutes of the request. No living human on Earth responds to a $1,900 contract inquiry with that speed and quality of phrasing—unless they’re working in the hell of deadline-driven science-pop content production.
Strictly speaking, this sets the bar for a new industry. Entry cost: one web designer, one SMM manager with a Midjourney subscription, one LLM-API with proper prompt engineering, and hosting fees. Output: $1,900 × N clients per month at zero operational costs. Margin: effectively infinite. This scheme is not a bug, but a feature: “100% human” is written in bold because the audience fears AI and is willing to pay a premium for “the real thing.” But under the hood—it’s all AI. Meaning the marketing promise itself is a demand detector, while the service description is a supply disguise.
If Research Gold is an imitation of human service, then the Columbia Nursing School audit (published in The Lancet on May 7, 2026, by Maxim Topaz and team) is an imitation of scientific knowledge in the same logic, but on a hundredfold larger scale.
Topaz’s team analyzed 2.5 million articles published in PubMed Central Open Access from January 1, 2023, to February 18, 2026. This is all the literature formally considered “community-verified” and included in clinical guidelines, treatment protocols, and systematic reviews that doctors rely on daily. Out of 97.1 million verified references, they found 4,046 fake citations in 2,810 articles. This means: every 277th article published in early 2026 cites a work that doesn’t exist in nature.
The trend is terrifying:
In other words, over two years, the rate of contamination in the scientific field increased 12-fold. And this is only what was detected—Topaz honestly warns that “the bigger and more important problem” is inaccurate but not entirely fabricated references, which modern methods can’t detect at all. Mohammad Hosseini from Northwestern Feinberg School of Medicine puts it bluntly: “We are far from being able to even detect them or do anything about them.”
And here’s the scariest part: 98.4% of the affected articles received no response from the publisher at the time of the audit. That is, the Topaz audit didn’t just document the problem—it documented that the system didn’t react. Articles continue to be cited, fake references continue to flow into meta-analyses, and meta-analyses continue to inform clinical guidelines. David Resnik from the NIH, commenting on the situation, says that retractions are only needed for articles where the fake reference is central to the conclusions. Topaz provides an example: an article in Frontiers in Oncology where 18 out of 30 references are fake. That’s what “centrality” looks like.
The irony is that Topaz and his team used an AI system to catch AI fakes. Hunting synthetic knowledge with even more advanced synthetic tools. And this is the only way—because a human reviewer physically can’t read 97 million references and verify each one in PubMed, but an LLM in 2026 already can. Hosseini is right: we’re using the pathogen to make a vaccine against the pathogen.
Now—the promised inversion. To see it, you need to look at the marketing landscape in reverse.
2022—The heyday of “powered by AI”: ChatGPT exploded in November 2022. By mid-2023, every second startup on Product Hunt was branding itself “AI-first.” Jasper, Copy.ai, Midjourney, Runway—all promised “superhuman” results. The ad tone: “our AI does X faster and cheaper than humans.” The hidden message: humans are the bottleneck, AI is the rein. Business model: subscription + integration + partnerships. Storytelling: robots replace humans, and that’s a good thing. No one was embarrassed.
2024—A 180-degree turn: Dove was the first major brand to officially pledge never to use AI for depicting “real women” in ads in April 2024. The tone: “AI is a threat to women’s well-being.” That is, AI went from being an advantage to a reputational risk. Discover released a series of ads with Jennifer Coolidge where the central joke is “you can talk to a real person here.” The Atlantic in June 2024 published a piece titled “LLM-free, all-organic”—yes, in marketing, this became the equivalent of “organic” in food. The Atlantic directly noted that Le Creuset, Hermès, and Bottega Veneta started emphasizing their human-made products. The term “AI fatigue” emerged—weariness from neural network content.
2025–2026—The era of inversion: Coca-Cola continues to generate holiday ads with AI (and gets a steady stream of hate for it). Le Creuset signs every post “created entirely by hand.” Le Creuset—literally, note this—a cookware manufacturer, a company whose product hasn’t changed since the 18th century, is now competing with Coca-Cola on the field of “authenticity.” This is the new axis of competition: not price, not quality, not design, but “authenticity.” And “authenticity” became currency because the market is oversaturated with AI content.
Kate Lindsay in Embedded (Substack) nailed it in March 2026: five years ago, a brand boasting “human made” would have sounded absurd—like McDonald’s advertising “100% real beef.” Today, it’s a necessary statement because by default, everything else is suspected of being synthetic. Notice: the marketing inversion mirrors the inversion in Research Gold. In the first case, the brand shouts “we’ve got humans!”—but it’s AI inside. In the second case, the brand shouts “we’ve got AI!”—but there’s nothing special inside. Both are lying, and both are succeeding because the audience is too tired to care.
And here—pay attention—these two stories converge into one. Research Gold sells fake PhDs under the “100% human” label. The Topaz audit shows that real scientists sell fake references under the “peer-reviewed” label. The difference is in the label’s status. In the first case, the “human” label is advertising. In the second, the “peer-reviewed” label is a quality certificate, the foundation of clinical medicine. And both labels are compromised by the same mechanism: a generator of plausible text without any obligation to truth.
Many commentators here make a mistake by seeing this primarily as a moral issue. That’s the wrong lens. Morality is about “you shouldn’t do that.” Engineering is about why it’s impossible to do otherwise.
The root of the problem lies in the architecture of language models. LLMs are trained on a corpus of texts where “looking plausible” and “being true” are orthogonal axes. The model doesn’t know that a reference has a DOI and a PubMed-ID. The model knows that a reference should have an author, year, journal, volume, and pages. And it generates this structure, filling the slots with the most statistically probable values. If the training data had many articles with “Smith J., 2019, Nature, 580(7802), 123-130,” then when asked to “find a reference for study X,” the LLM happily assembles a new Smith J. with a new topic, new year, new volume—and a DOI that it confidently claims leads to the article’s page. Because a DOI is just a string in the format 10.xxxx/yyyyy, and there are tens of millions of such strings in the world, and the probability of guessing a real DOI in a 13-character string is vanishingly small, but the model doesn’t know that.
Mohammad Hosseini and David Resnik published a paper in March 2026 stating outright: hallucinating is “inextricably linked to how LLMs operate.” This isn’t a bug that can be fixed. It’s a systemic property of generative architecture. A model that doesn’t hallucinate isn’t generative—it’s a search engine. And the entire LLM marketing machine is built on the fact that the model generates, not searches.
So fake references are the price of generation itself. Want an LLM to draft a literature review for you? Get 10–20% of references that need manual verification. Want an LLM to write a peer review? Get 5% of sentences that sound professional but are physically incorrect. Want an LLM to craft a sales pitch? Get fake certificates, nonexistent titles, and made-up case studies—just like we see with Research Gold.
This is not a bug, but a fundamental trade-off. And as long as we treat LLMs as truth generators—rather than plausibility generators—this problem will multiply. Topaz’s team recommends “fighting AI with AI”—automating reference verification at the submission stage. That’s a reasonable tactic, but it’s a tactic, not a strategy. Strategically, we need to rethink the role of LLMs in science: they should be the reviewer’s assistant, not the reviewer; the draft of a review, not the review itself; the hypothesis generator, not the hypothesis validator. And—pay attention—this is the exact same logic as in “anti-AI marketing”: humans need a role where AI complements them, not replaces them.
The most alarming part of Topaz’s audit isn’t the numbers themselves, but the 98.4%. Ninety-eight point four percent of articles with fake references at the time of the audit’s publication received no response from the publisher. That means 2,766 out of 2,810 articles with confirmed fake references continue to live in the database, get cited, get indexed, and inform clinical guidelines. And in May 2026, when I’m writing this, none of the major publishers—Elsevier, Wiley, Springer Nature, IEEE, Sage—even responded to Retraction Watch’s request for comment.
Taylor & Francis replied that they’re “investing in technology and specialists.” PLOS said they’re “exploring systemic screening options,” but at the same time, “do not automatically classify fake references as misconduct”—because misconduct requires proven intent, and here we have a situation where the author genuinely didn’t know the LLM deceived them. This is a legal dead end: the law assumes the author read what they cited. In 2026, that assumption is no longer true by default.
Hoch from PLOS puts it bluntly: “Whether an issue qualifies as research misconduct is addressed at the institutional level, not at the journal or publisher level.” Translation: “This isn’t our problem, it’s the universities’ problem.” Universities, in turn, throw up their hands: “We don’t have the resources to investigate every fake reference, especially when the author insists it was an accident.”
Flemyng from Cochrane calls what’s happening a “perverse incentive for fast science”—a corrupt incentive for speed. A researcher who needs “more publications, more citations” is forced to cut corners. AI is the perfect tool for cutting corners: it works fast, it sounds plausible, and until recently, it was impossible to catch. Topaz adds: “The damage is already done. The contamination of over 4,000 fabricated references his team found does not go away when the AI gets better.” That is, even if OpenAI or Anthropic completely eliminate hallucinations tomorrow, the already contaminated literature isn’t going anywhere. Citations are an immutable layer of scientific memory, and part of it is already counterfeit forever.
And here we return to Research Gold. Because Research Gold is the first public case where two safeguards collapsed at once: both human expertise and peer review. When a service for scientists with fake PhDs charges $1,900 for “peer-review ready” work where the references will inevitably be half synthetic (because LLMs can’t help but hallucinate), the result ends up in PubMed—and PubMed formally passes peer review—and clinical guidelines formally rely on evidence-based medicine—but the entire chain factually consists of plausible synthetic content. This isn’t a bug in one company. It’s a structural property of the industry, where “plausibility” has replaced “truth” as the commodity.
When I started piecing this together, I didn’t expect it to be so symmetrical. But look:
| Layer | What They Promise | What They Actually Sell | Who Suffers |
|---|---|---|---|
| Marketing | “Created by hand,” “never AI” | AI content with a human signature | The brand loses trust when exposed |
| Scientific Service (Research Gold) | “100% human, PhD methodologists” | LLM generation + fake avatars | Patients treated based on fake meta-analyses |
| Scientific Literature | “Peer-reviewed, evidence-based” | LLM drafts with fake references | Everyone who relies on clinical guidelines |
All three layers run on the same engine: a generator of plausible text indistinguishable from the real thing. In all three cases, the audience (consumer, doctor, reviewer) physically can’t tell synthetic from authentic without special tools. And in all three cases, the market rewards those who cut corners on truth and sell plausibility.
This is a real arms race, and it’s already lost in the foreseeable future unless the architecture changes. Neither “anti-AI marketing” (which itself becomes marketing) nor “AI detectors” (which lag 6–12 months behind LLMs) nor “peer-review reform” (while the reviewer gets an LLM draft, they don’t know it’s an LLM draft because it’s indistinguishable from an average human one) will help. What’s needed is a different foundation—something like a cryptographically signed submission chain, where every stage from draft to publication has verified-provenance metadata, and where the share of LLM assistance at each stage is a separately declared and verifiable characteristic. This is complex, expensive, and definitely won’t happen on its own because the industry has no incentive.
You know what struck me most about this story? Not Research Gold itself. Not even the 4,046 fake references. But the fact that the systems no longer work but continue to look like they do. The “peer-reviewed” label used to mean “verified by a human professionally accountable for the result.” Today, it means “formally verified” because the reviewer is an overworked academic who physically can’t check every reference in the paper, and the paper is a draft where the LLM author generated 20% of the references. The “100% human-written” label used to mean “a real person, with a name, reputation, and legal responsibility.” Today, it means “a human pressed Submit, the rest was done by AI.”
And we—as a society, as readers, as doctors, as consumers—are getting used to this. That’s the scariest part. Not that companies lie, but that the audience no longer expects truth. We expect plausibility. We scroll past a Le Creuset Instagram post and feel good about “created entirely by hand”—but don’t verify. We visit a medical service’s website and see “100% human”—and believe it because it’s convenient. We read a peer-reviewed paper in The Lancet and trust the conclusions—because the alternative (not trusting) requires cognitive effort we don’t have time for. The culture of plausibility is the new cultural norm, and it’s scarier than any single fake article because it doesn’t require deception—it works as long as all participants act rationally within their incentives.
Is there a way out? I don’t see a silver bullet. Topaz suggests “fighting AI with AI”—that’s a tactic, not a solution. Provenance transparency (C2PA-like metadata for texts) is an infrastructural shift that will happen in the next 5–10 years, but the already contaminated literature will remain contaminated forever. The most realistic solution is a return to a reputation economy: journals must retract articles for fake references (Bauchner and Rivara from JAMA insist on this), authors must bear personal responsibility for every reference (i.e., read each one), and universities must train students to work with LLMs just as they once trained them to work with primary sources. It’s boring, it doesn’t scale, and that’s exactly why it probably won’t happen.
Meanwhile—every time you see the “100% human” label, remember that Research Gold built its business model on it, and that in a world where this label has become marketing, trust in the label is the currency printed by whoever lies the loudest. And the only defense against this is your own skepticism, multiplied by the time you’re willing to spend verifying. And time, as we know, is the one resource scarcer than money in modern science, medicine, and marketing. That’s why fakes are winning.
Sources: