This page explains how we turn raw hospital price-transparency files into the comparable numbers on this site.
1. Sourcing
We locate each hospital's machine-readable file (MRF) from the price-transparency page it is required to publish. We parse every supported format — CSV, JSON (CMS v3.0.0 schema), and zipped/extension-less endpoints — streaming multi-gigabyte files so nothing is truncated.
2. Code mapping
Each line item carries one or more billing codes (CPT/HCPCS, MS-DRG, etc.). We map these to a curated catalog of shoppable procedures, choosing a canonical code by type so a procedure aggregates consistently across hospitals. For inpatient stays we price from the whole-stay DRG rather than individual professional-fee codes.
3. Aggregation
For each hospital and procedure we compute the median of published amounts by charge type (gross, cash, negotiated), dropping zeros and obvious data errors. The median is robust to the outliers that are common in these files. We then roll medians up to city, state and national levels.
4. Quality filtering
We flag and exclude price rows that fail sanity checks — cash prices far outside the national range for the procedure, or chargemaster prices below cash (a sign of component-code mismatches). Pages with fewer than 11 hospitals are excluded from search engines as too thin to be reliable.
5. Medicare benchmark
For context we compute what Medicare pays using the CMS Physician Fee Schedule (RVUs × conversion factor), Clinical Lab Fee Schedule, and the IPPS DRG weights for inpatient stays.
All prices are estimates derived from public files and may be out of date or incomplete. Always confirm with the hospital.