Most public health data is free to access. Turning it into interpretable evidence still takes a lot of effort. NHANES is a good example. It is one of the most valuable datasets in the United States: decades of interviews, physical exams and lab results on a nationally representative sample. But getting a correct estimate out of it takes deep familiarity with how the survey was designed, how each file is structured, and how the analysis has to account for both.
A growing set of "chat with your data" tools promises to close that gap, and they take different forms:
- File-upload assistants. General-purpose assistants analyze a spreadsheet or CSV you upload.
- Enterprise text-to-SQL tools. Databricks AI/BI Genie and Snowflake Cortex Analyst translate plain-language questions into queries against a data warehouse.
- Open-source text-to-SQL libraries. Vanna does the same against SQL databases such as Postgres.
- Secure analysis platforms. Many important health datasets can't be downloaded at all. With the UK Biobank Research Analysis Platform or the All of Us Researcher Workbench, you log in to a cloud environment, complete the required training, and learn that platform's tools before you can run an analysis.
These tools are useful, but they have two limits. First, they are generic. They know SQL or Python, but they don't know the methods of a particular dataset: its sampling design, coding conventions, missing-value codes, or how its files fit together. The answers come back fast and look reasonable, and nothing in the output tells you whether they're right. Second, the burden stays on the researcher. Every new dataset means a new platform, new training and a new set of conventions. You still go to the data.
We wanted to test the opposite idea: bring the data to the user, and build the dataset's methodological expertise into the tools the AI uses.
What we built
nhanes-mcp is a proof of concept: an open-source Model Context Protocol (MCP) server that lets an agentic AI assistant analyze NHANES the way NCHS analysts do. By "agentic" we mean an assistant that does more than answer. It plans the analysis and calls a sequence of specialized tools to carry it out: finding variables, assembling the dataset, deriving outcomes, and estimating.
For NHANES, the central methodological issue is the sampling weights. Not every dataset needs survey weights. Many EHR and claims analyses don't. But for NHANES, getting the weights wrong means getting the answer wrong. So the server's tools handle what an experienced NHANES analyst would:
- Find variables and files across survey cycles from 1999 onward.
- Build an analytic dataset, merging the needed components and choosing the correct weight automatically. That means interview, exam or lab-subsample weights, plus the special 4-year weights for 1999–2002. Every choice is logged.
- Derive variables safely, so missing data stays missing. If BMI is missing, "obese" is missing too, not "no."
- Estimate with the full survey design: Taylor-linearized standard errors, domain estimation, Korn-Graubard confidence intervals for proportions, and NCHS reliability flags.
- Age-adjust to the 2000 US standard population.
- Fit survey-weighted regression and Cox models.
The point is that the guardrails live in the server, not in the prompt. The assistant can't easily skip the weights, because the tools won't produce an estimate without them.
Extending it: from query to visualization
We then took one more step. The NHANES Results Explorer is an MCP App, an interactive dashboard that renders inside the chat. Besides querying the data, the assistant can now visualize the results in real time. You can switch between breakdowns by sex, age and survey cycle, and see confidence intervals, reliability flags and the published benchmark side by side.
In one conversation, you go from a research question to a cohort, then to weighted estimates and a visualization.
Does it get the right answer?
A tool like this is only as good as its validation record. So we tested it against published NCHS estimates in three disease areas (Table 1).
Table 1. nhanes-mcp estimates compared with published results.
| Question | nhanes-mcp | Published | Source |
|---|---|---|---|
| Adult obesity prevalence, Aug 2021–Aug 2023 | 40.3% | 40.3% | NCHS Data Brief 508 |
| Hypertension, all estimates in the brief | 17 of 17 exact matches | — | NCHS Data Brief 511 |
| CMV seroprevalence, 1999–2004 | 50.6% (men 45.5, women 55.5) | 50.4% (men 45.2, women 55.5) | Bate et al., Clin Infect Dis 2010 |
We chose CMV on purpose. It's a different disease area, it relies on a lab subsample with its own weight, and it reaches back to the earliest survey cycles. That's exactly where things tend to break.
What testing caught
Testing paid off. The first CMV run exposed two quiet failures:
- The server couldn't find the 1999–2000 serology file, because its name differs from later cycles.
- It didn't recognize the CMV subsample weight, so it would have used the wrong weight.
Neither one caused an error message. Each would have produced a plausible, wrong number. Both are now fixed and covered by automated tests.
Benchmarking against published results is not optional for AI analysis tools. It's how you find the errors that look fine.
What we haven't solved yet
This is a proof of concept, and we want to be clear about its limits:
- In a longer obesity trend series, 18 of 27 estimates match NCHS Health E-Stat exactly. Estimates before 2011 run 0.1–0.7 percentage points low, and we haven't found the cause yet.
- Our CMV sample size (14,198) is smaller than the published one (15,310), even though the estimates agree closely.
- NHANES III (1988–1994) isn't supported yet.
These are listed as known issues in the nhanes-mcp repository. Every other data source will bring its own: different file structures, coding conventions and analytic traps. Each one needs its own validation.
Standing on the shoulders of M3 and M4
We weren't the first to put a clinical dataset behind MCP, and this work builds on people who got there earlier.
M3 (Al Attrach et al., 2025), from researchers working with MIT's Laboratory for Computational Physiology, showed that an MCP server can let clinicians and researchers query MIMIC-IV in plain language with secure, local data access. M4 takes this further, toward agentic clinical research on MIMIC-IV and eICU, with clinician-reviewed clinical concepts, research automation, and interactive apps.
Their work showed us that the pattern holds up and that the community wants it. nhanes-mcp is complementary. M3 and M4 focus on EHR data from the ICU. We focus on complex survey data, where the hard part is the sampling design rather than the clinical definitions. If you work with MIMIC, start there.
Why this matters beyond NHANES
NHANES is the proof of concept. The pattern applies more broadly:
- Put the expertise in the tool layer. Survey weighting for NHANES, cohort logic for EHRs, enrollment and claims windows for insurance data. Each data source has its own traps, and an AI assistant should hit guardrails rather than rely on luck.
- Validate against the published record. Every data server should come with a benchmark set showing it reproduces known results. The validation record, not the code, is what makes a tool trustworthy.
- The data doesn't have to move. An MCP server can run inside a licensed or secure environment, so the analyst asks questions and only aggregate results leave. That opens the same approach to registries, EHR systems and claims databases, including proprietary sources that can never be downloaded.
- Keep the epidemiologist in the loop. These tools don't replace the subject matter expert. They change where the expert's time goes: less on wrangling files and learning platforms, more on framing the question, checking the cohort definition, and judging whether a result makes sense. The agent does the mechanics, and the epidemiologist owns the decisions and the interpretation.
Serious questions remain. How should tools like this be governed? How do we evaluate the AI's interpretation of results, not just its arithmetic? Who signs off on an estimate a model helped produce? But the direction is clear to us: researchers shouldn't have to become data engineers, or learn a new platform for every dataset, to get correct answers from data that already exists.
Try it, and talk to us
- Server (MIT license): github.com/Black-Swan-Causal-Labs/nhanes-mcp. The README has a one-line install with
uvxand a prompt you can give your AI assistant to set it up for you. - Results Explorer: a separate add-on, free for academic and noncommercial use.
- Coming next: a preprint with a systematic benchmark across many published NCHS estimates, and a second proof of concept on an EHR data model.
If you steward a dataset (a registry, a cohort, an EHR or claims source) and want to make it usable this way without the data ever leaving your environment, let's talk.
References
- Emmerich SD, Fryar CD, Stierman B, Ogden CL. Obesity and severe obesity prevalence in adults: United States, August 2021–August 2023. NCHS Data Brief, no 508. Hyattsville, MD: National Center for Health Statistics; 2024. https://www.cdc.gov/nchs/products/databriefs/db508.htm
- Fryar CD, Kit B, Carroll MD, Afful J. Hypertension prevalence, awareness, treatment, and control among adults age 18 and older: United States, August 2021–August 2023. NCHS Data Brief, no 511. Hyattsville, MD: National Center for Health Statistics; 2024. https://www.cdc.gov/nchs/data/databriefs/db511.pdf
- Bate SL, Dollard SC, Cannon MJ. Cytomegalovirus seroprevalence in the United States: the National Health and Nutrition Examination Surveys, 1988–2004. Clin Infect Dis. 2010;50(11):1439-1447.
- Al Attrach R, Moreira P, Fani R, Umeton R, Fiske A, Celi LA. M3: Conversational LLMs simplify secure clinical data access, understanding, and analysis. arXiv:2507.01053. 2025. https://arxiv.org/abs/2507.01053
- M4: clinical database access, research automation, and clinician-reviewed domain knowledge for AI agents. GitHub. https://github.com/hannesill/m4
- Black Swan Causal Labs. nhanes-mcp. GitHub. https://github.com/Black-Swan-Causal-Labs/nhanes-mcp