It is Tuesday morning and a cardiologist has a patient in front of her who does not fit the trial.
Severe aortic stenosis, yes. But also a bleeding disorder that excluded her from every valve trial ever run, a frailty score that makes the guideline's default pathway questionable, and a family that wants a straight answer about what to do next. The guideline she trained on has a recommendation for aortic stenosis. It does not have one for this particular patient.
She does what a lot of cardiologists do in this exact moment. She opens Sermo, the physician social network, and posts an anonymized version of the case: what would you do here. Within an hour she has thirty responses. She has no way to know if any of them came from someone who has actually managed this combination, whether the confident answer at the top came from an interventional cardiologist at a quaternary center or a locum internist three years out of residency, or whether the crowd she is polling is representative of anything at all.
She reads the responses anyway, because they are the only thing available that arrives before her clinic visit this afternoon.
This is not a failure of her judgment. It is a failure of infrastructure. Medicine has spent two decades building better tools to answer the questions the evidence can already answer. Almost nothing has been built for the much larger category of questions the evidence cannot answer at all.
When the trial data runs out, medicine's only tools for "what do other centers actually do" are an anonymous crowd nobody can verify and a guideline process that will not report back for years.
Half of guideline medicine is educated opinion, not proven fact
Start with how large this zone actually is, because the number is larger than most clinicians who have not looked closely would guess.
A landmark analysis published in JAMA in 2009 examined all 16 then-current American College of Cardiology and American Heart Association guidelines, covering 2,711 classifiable recommendations drawn from a total pool of 7,196 recommendations issued between 1984 and 2008. The finding: a median of only 11 percent of recommendations were supported by Level of Evidence A, meaning multiple randomized trials or meta-analyses. A median of 48 percent rested on Level of Evidence C, the category defined by expert opinion, case studies, or standard of care.
Even inside the highest-confidence tier, the picture does not improve much. Among Class I recommendations, the ones guidelines present with the most authority, only a median of 19 percent had Level A evidence behind them.
That is not a snapshot of a field still catching up. A 2019 replication in JAMA, examining ACC/AHA and European Society of Cardiology guidelines issued from 2008 through 2018, found the pattern persisted across the following decade. This is a structural feature of clinical guidance, not a temporary gap that better trial funding will close.
And cardiology is not an outlier specialty chosen because it looks worst. Guideline methodologists across gastroenterology, anesthesia, and oncology report comparable evidence-tier distributions when they apply the same framework. The honest generalization is that roughly half of what a specialist relies on to make a guideline-informed decision was never tested in a randomized trial. It was somebody's considered judgment, formalized into a document.
For every one of those recommendations, and for every clinical scenario that falls outside even that documented territory entirely, the only honest question a clinician can ask is not "what does the evidence say." It is "what do experienced people at other centers actually do."
Two tools exist, and both fail differently
Here is the part that makes this a genuine infrastructure gap rather than a complaint about how medicine works.
Two kinds of instruments currently answer "what do other centers do," and each fails for a different, structural reason.
The anonymous crowd. Sermo, the largest physician social network built around this exact use case, lets a member post a clinical scenario and collect responses from other physicians. The product's defining design choice is anonymity, and it is a deliberate one: physicians say things under anonymity they will not say under their real name, and that candor is Sermo's differentiator. But it means the asking physician has no way to verify who answered. A response from a fellowship-trained specialist with a decade of relevant case volume looks identical on the screen to a response from someone with none of that background. Nothing is weighted. Nothing is citable. If the case is later scrutinized, there is no defensible record of who said what and why it was reasonable to rely on it.
The guideline panel. Specialty societies convene expert committees on a multi-year cycle to synthesize practice into formal recommendations, including the Level C, expert-opinion-based ones described above. This process is careful, deliberative, and about as far from real time as a clinical tool can be. It cannot answer the specific situational question a clinician has this week, because it was never built to. And when a guideline committee writes a Level C recommendation, it typically draws on a small, often undisclosed set of panel members' personal experience rather than any systematic survey of how practice actually varies across institutions.
Nothing occupies the space between "instant but unverifiable" and "authoritative but glacial." That space is exactly where most of clinical practice actually lives.
The hidden problem is not the answer, it is who is answering
The reason this gap has persisted is that it looks, on the surface, like a search problem. It is not. It is a verification and weighting problem, and that is a harder thing to build.
Answering "what do other centers do for X" usefully requires knowing more than an answer. It requires knowing who is answering: their specialty, their practice setting, whether they have managed this specific scenario, how many times, and how recently. A response from a critical-care attending at an academic center who has seen this presentation eleven times this year should not be weighted the same as a response from someone who has seen it once, a decade ago, and remembers it imperfectly.
That is an expertise graph, not a poll. Building it requires verified identity attached to verified practice history, matched against a specific clinical question, at a speed fast enough to be useful before the clinic visit ends. No anonymous platform can supply that data by design, because the entire point of anonymity is that identity, and everything attached to it, disappears.
Why nobody has built the thing in between
Run through who could plausibly build this, and the reason each candidate has not is structural rather than an oversight.
Sermo will not build it, because verified identity is a direct threat to the candor that makes the platform valuable in the first place. A physician who will say something anonymously that they will not say under their own name is the exact user Sermo is built to retain, and adding verification changes what the platform is.
Specialty societies will not build it, because they own the slow, formal guideline process and have no product, staffing model, or legal appetite for fast, informal, same-day polling. Convening a Delphi panel and running a live practice-variation query are different disciplines with different risk profiles, and no society has built the second one.
UpToDate and comparable evidence synopses will not build it, because their editorial standard explicitly excludes content that rests on expert opinion rather than published evidence wherever possible. That is a legitimate editorial choice. It also means the entire Level C, expert-opinion zone, roughly half of guideline medicine by the Tricoci numbers, is left outside their scope by design rather than by neglect.
Each incumbent's core business model is precisely why it will not build the missing tool. That combination, verified identity plus real-time polling plus specialty and case-volume weighting, sits in structural whitespace between all of them.
AI answer engines inherit the same blind spot
There is a temptation to assume this problem solves itself as AI clinical answer tools improve. It does not, and understanding why matters.
An AI system trained on the published literature has exactly the same blind spot as the literature itself. If roughly 40 to 50 percent of guideline-relevant questions have no trial evidence behind them, no amount of better literature synthesis produces an answer to those questions, because there is no literature to synthesize. The AI can summarize the Level C recommendation more fluently. It cannot manufacture the trial that was never run.
The open, testable question is whether these tools are honest about the distinction: does an AI answer engine flag when a response falls into expert-opinion territory, or does it present a synthesized answer with the same tone of confidence it uses for a Level A recommendation. If it does the latter, AI adoption makes this problem more dangerous rather than less, because it launders the uncertainty out of view at exactly the moment clinicians are trusting these tools more.
As AI gets better at answering the questions evidence can already answer, the residual, growing share of clinical value shifts precisely to the questions it cannot: the evidence-free zone. That zone does not shrink. It becomes the whole remaining problem.
What would actually work
Verified identity, not anonymity. A response is only weightable if the respondent's specialty, practice setting, and years of experience are attached and verified, which anonymous platforms structurally cannot provide.
Speed measured in hours, not weeks. A practice-variation question that cannot be answered before the clinic visit or the committee decision it was asked to inform has failed regardless of eventual accuracy.
Aggregated, citable output. The result should read as "14 of 18 responding critical-care attendings at academic centers report using this approach," not a scroll of unweighted individual comments, so a clinician has something defensible to point to later.
Explicit labeling of what this is not. A practice-variation poll answers "what do others do for this protocol or scenario class." It must never become a channel for patient-specific diagnostic advice, which would cross into unlicensed remote consultation and the crowdsourced-diagnosis failure mode this series has flagged elsewhere.
A feed back into formal guideline work. The same data that helps one clinician this week is, in aggregate over time, exactly the practice-variation signal a guideline committee needs to know where a Delphi round is actually warranted next.
Guardrails against defensive-medicine herding. A tool that shows clinicians what others do risks becoming a tool that teaches clinicians to imitate the herd rather than exercise judgment; any serious version needs to present variation as information, not as a recommendation to follow the majority.
A route for the specialty society, not a bypass of it. The goal is not to compete with formal guideline development. It is to make the gap between guideline cycles visible and navigable while feeding the society the practice-variation signal it currently has no systematic way to collect.
What you can do now
If you are a clinician facing a guideline gap
Name the gap explicitly, to yourself and in documentation. If a decision rests on Level C evidence or on no guideline at all, say so in your reasoning rather than presenting a personal judgment call with the confidence of a proven protocol. That habit alone changes how carefully you weigh outside opinions.
Weight what you get from any informal poll by source, not by volume. Thirty anonymous responses are not thirty units of evidence. Ask, when you can, who is answering, and discount answers where you cannot tell.
Keep a record of your own reasoning on evidence-free calls. The case you managed without trial guidance today is exactly the kind of practice-variation data point this field currently has no way to collect systematically. Your private note is more valuable than it feels in the moment.
If you lead a guideline committee or specialty society
Treat Level C recommendations as flagged data, not settled fact. Publish, internally at minimum, which recommendations in your guideline rest on expert opinion so your own members know where consensus was thin.
Consider a standing, lightweight practice-variation channel between formal revision cycles. The multi-year guideline cadence is structurally too slow for situational questions; a fast channel that feeds observations back into the next formal revision does not compete with the guideline process, it feeds it.
Ask your AI-answer vendors the hard question directly. Whether the tools your members are already using flag expert-opinion-based answers as such, or present them with false confidence, is a testable, answerable question your society should be asking now rather than after a bad outcome traces back to it.
If you build clinical infrastructure
Design for verification before you design for scale. An unverified crowd is easy to build and has already been built. The hard, valuable part is the identity and expertise-weighting layer, and that has to come first, not as a feature added later.
Build the guardrail against patient-specific use into the product, not the terms of service. A tool that can be used for protocol-level practice questions will eventually be used to ask about a specific patient unless the structure itself prevents it.
Frequently asked questions
What percentage of clinical guidelines are based on expert opinion rather than trial evidence? In a landmark 2009 JAMA analysis of 16 ACC/AHA cardiology guidelines covering 2,711 classifiable recommendations, a median of only 11 percent rested on Level A evidence (multiple randomized trials or meta-analyses), while a median of 48 percent rested on Level C evidence, meaning expert opinion, case studies, or standard of care. A 2019 JAMA replication found the pattern persisted across guidelines issued from 2008 to 2018.
What is Sermo and how does physician crowdsourcing work? Sermo is a physician social network built around anonymous posting, where a member can describe a clinical scenario and collect responses from other physicians on the platform. Anonymity is the platform's core design choice, intended to encourage candor, but it means responses cannot be verified by specialty, practice setting, or actual experience with the scenario asked about.
How reliable is anonymous physician crowdsourcing for clinical decisions? It is unverifiable by design. An anonymous platform cannot confirm who answered a given question, meaning a response from a highly relevant specialist and a response from someone outside the relevant field appear identical, and results cannot be weighted, cited, or defended if a decision is later scrutinized.
How do doctors decide what to do when there's no clinical trial evidence? They typically rely on Level C, expert-opinion-based guideline recommendations where one exists, or, where none exists, on informal channels: texting former co-trainees, posting anonymized cases on platforms like Sermo, or proceeding on personal clinical judgment without checking outside opinion at all.
Do AI clinical answer tools solve the evidence gap? No, because an AI system trained on published literature has the same blind spot as the literature itself. If roughly 40 to 50 percent of guideline-relevant questions rest on expert opinion rather than trial evidence, better literature synthesis cannot manufacture evidence that was never generated; whether these tools honestly flag that distinction rather than answering with false confidence is an open, testable question.
Why don't specialty societies just answer these questions faster? Guideline development runs on a multi-year cycle built for careful, deliberative synthesis of the literature, including formal disclosure and consensus processes. That cadence is structurally incompatible with a same-day situational question, and no society has built a separate, faster product for informal practice-variation polling alongside its formal guideline process.
The bottom line
Roughly half of what specialists rely on when they consult a guideline was never tested in a randomized trial. It is documented expert judgment, and for a large share of real clinical decisions, there is no guideline at all to consult. The honest question in that entire zone is not what the evidence says. It is what experienced people elsewhere actually do.
Two tools currently claim to answer that question. One is instant and cannot be trusted, because nobody can verify who is answering. The other is careful and cannot be used this week, because it operates on a multi-year cycle built for something else entirely.
AI answer engines will not close this gap on their own, because they inherit the same blind spot as the literature they were trained on. If anything, better AI synthesis raises the stakes, because the residual, uncovered territory, the evidence-free zone, becomes the entire remaining frontier of clinical uncertainty.
What is missing is not a smarter search engine. It is a verified expertise graph fast enough to answer a same-day question and honest enough to be cited later: who actually answered, what they actually know, and how many others at comparable centers said the same thing.
Until that exists, a cardiologist facing a patient who does not fit the trial will keep doing exactly what she did on Tuesday morning: posting into a crowd she cannot verify, reading the answers anyway, and hoping the confident one at the top came from someone who has actually managed this before.
Part of a series on the missing professional infrastructure of healthcare. Previously: The Employer-Owned Career Record: A Career That Dies at Every Job Change
Evidence note: the core evidence-tier statistics come from Tricoci P, Allen JM, Kramer JM, Califf RM, Smith SC Jr, JAMA, 2009 (16 ACC/AHA guidelines, 2,711 classifiable recommendations of 7,196 total, 1984 to 2008), and its 2019 JAMA replication by Fanaroff AC et al. covering guidelines issued 2008 to 2018. Claims about Sermo's product design and positioning around anonymity reflect the platform's publicly stated value proposition rather than an independent audit of its user base size or response quality, and specific usage statistics for Sermo were not independently verified in this pass. The claim that comparable evidence-tier distributions hold across gastroenterology, anesthesia, and oncology reflects parallel methodology studies in those fields rather than a single unified analysis. Nothing in this article is guidance for any specific patient decision; it describes a gap in infrastructure for protocol- and practice-pattern-level questions only.