FHIR Bulk Data Access has become the standard mechanism for moving population-level clinical data out of US EHRs into reporting, analytics, and registry pipelines. The spec has been stable for years now, US EHRs implement it widely, and the surrounding tooling has matured. What separates a smooth bulk data pipeline from a stuck one is mostly the tools you use to drive $export, handle NDJSON streams, and turn the output into something the downstream consumer can use. The list below covers the tools US EHR teams reach for in 2026.
For deeper FHIR walkthroughs, the rest of the site has surrounding material. The complete guide to FHIR-based EHR development for US healthcare in 2026 cornerstone is the right background to read first.
What FHIR Bulk Data Tooling Has to Do
The Bulk Data Access spec defines an asynchronous flow: a client kicks off a $export job, polls for completion, then downloads a set of NDJSON files containing FHIR resources. A bulk data tool has to drive that flow correctly, handle the long-running job state, manage download credentials and retries, and stream the NDJSON output into whatever processing pipeline sits downstream.
It also has to handle the real-world failure modes: slow exports, partial completions, large payloads that need streaming rather than buffering, and the per-EHR quirks in how each implementation hands out completion URLs.
The 6 Tools That Handle Bulk Data Well
These are the bulk data tools US EHR teams actually use in 2026:
- BCDA (CMS Beneficiary Claims Data API) Reference Client, the open-source client originally built for CMS's bulk data work and broadly applicable across other Bulk Data deployments.
- bulk-data-client (SMART Health IT), the open-source reference client maintained by the SMART team, used as a starting point for many US bulk data projects.
- HAPI FHIR Bulk Data Server, the bulk export implementation that ships with HAPI FHIR and is widely used both as a server and as a reference.
- Cumulus, the open-source ETL pipeline from Boston Children's that wraps Bulk Data Access in a research-and-quality-reporting-friendly shell.
- Pathling, the SQL-on-FHIR adjacent toolkit that fits bulk data output into a queryable analytic shape.
- Smile Digital Health Bulk Data Module, the bulk export feature inside the Smile platform, used by US deployments already on Smile.
Each fits a different stage. The reference clients drive $export against any compatible server. The HAPI and Smile pieces are server-side bulk implementations. Cumulus and Pathling sit downstream and turn the NDJSON output into something analytics-friendly.
Where the Tools Differ in Practice
bulk-data-client and BCDA Reference Client are good starting points for any US bulk data integration. They handle the asynchronous flow correctly, retry on the common failure modes, and stream large payloads without exhausting memory. The choice between them is mostly preference and what your team is used to.
HAPI FHIR Bulk Data Server is the workhorse server-side implementation, used both in production and as a reference. Smile's bulk module fits teams already on Smile and integrates cleanly with the rest of the platform.
Cumulus and Pathling solve the downstream problem of turning NDJSON into something a data analyst can query. They are different tools with different shapes: Cumulus leans toward research and quality reporting workflows, Pathling leans toward SQL-friendly analytic access. The best FHIR profiling tools for US EHR vendors in 2026 piece covers some of the validation tooling that fits alongside bulk pipelines.
Where Bulk Data Pipelines Get Stuck
Three patterns trip up US bulk data pipelines regularly. First, building against a sandbox that does not represent real population sizes, then discovering the production export takes hours and the downstream pipeline cannot keep up. Second, handling the asynchronous job state casually, then losing exports when the polling logic breaks. Third, ignoring the per-EHR quirks in completion URL shapes and credential handling, which costs time on the first real production run.
A pre-production bulk export against real-shape data, even if synthetic, surfaces all three of these. The best FHIR sandboxes for EHR developers in 2026 explainer covers the sandboxes that hold up for bulk data testing specifically.
Making the Tooling Pick
A short filter sorts most projects. Client-side, start with bulk-data-client or BCDA Reference Client. Server-side, use HAPI FHIR or Smile depending on the broader stack. Downstream, pick Cumulus for research workflows or Pathling for analytic-style queries.
Test against real-shape data early and the rest of the pipeline holds up.
Sources
- Access Claims Data via FHIR Bulk Data $export (evergreen production reference) - CMS BCDA
- FHIR Bulk Data Implementations registry (evergreen) - SMART Health IT
- production FHIR Bulk Data deployment (evergreen) - CMS Beneficiary Claims Data API


