Project Overview
This project turns the CMS HCAHPS national patient-experience survey, a 100MB+ government dataset covering thousands of hospitals, into labeled sentiment data and state-level insights. HCAHPS reports patient feedback as pre-defined categorical answers (e.g. "Nurses 'always' communicated well") rather than a simple positive or negative score, so the first job is turning that raw taxonomy into something analyzable, then finding out where patients report the most negative experiences once the data is normalized fairly.
Technical Approach
Domain-driven sentiment labeling. All ~70 unique HCAHPS answer categories are mapped to positive, neutral, or negative through a manual, domain-knowledge lookup dictionary, since the answers are pre-categorized rather than free text.
Correcting a real double-counting bug. Each HCAHPS survey batch answers around 18 distinct questions that all share the same respondent base. An earlier version of the pipeline summed estimated responses across all of them together, inflating counts by roughly 18x and blending unrelated questions into one meaningless number. The fix aggregates responses per underlying question first, and a regression test now guards against that bug reappearing.
Per-capita normalization. State-level rankings are joined against U.S. Census Bureau population data and normalized per million residents instead of ranked by raw count. Raw negative-response counts correlate almost perfectly with state population (r = 0.95), since bigger states just have more hospitals and patients. Once normalized per capita, that correlation disappears (r = 0.03), confirming the metric measures care quality rather than population size.
Validation against ground truth. The pipeline's derived "would not recommend" rate is checked against CMS's own official Summary Star Rating for the same states. The two move together strongly (r = -0.86), which is the correct way to sanity-check this kind of derived metric, rather than comparing it against unrelated third-party composite indices.
Key Findings
- Among states and territories with a meaningful sample, Puerto Rico, DC, and Arizona have the highest share of patients who would not recommend their hospital; South Dakota, Minnesota, and Idaho have the lowest. (The U.S. Virgin Islands technically ranks highest at 11%, but on roughly 200 estimated responses, too few to compare fairly.)
- Not all HCAHPS dimensions are equally negative nationally. Measured as negative-response rates (the share of patients answering "sometimes" or "never"): medication side-effect communication has a 32% negative rate and pre-medication communication a 21% negative rate, the weakest-performing measures by far, while nurse and doctor courtesy and respect are consistently strong with only a 3-4% negative rate.
- This project measures patient-reported experience specifically (communication, cleanliness, quietness, discharge information, and whether patients would recommend the hospital). It does not measure, and shouldn't be compared against, broader composite "best/worst state healthcare" rankings that blend in insurance coverage, cost of care, and other unrelated variables.
Data Source
Public CMS (Centers for Medicare & Medicaid Services) "Patient survey (HCAHPS) - Hospital" dataset, available through the CMS Provider Data Catalog. Aggregated and de-identified at the source; no individual patient records are included or accessible.
Testing
Shared aggregation logic (measure-group derivation, response estimation, state aggregation, population join) is covered by 18 unit tests, including a regression test for the original cross-question double-counting bug.
Tech Stack
Python, pandas, matplotlib
Scope & Limitations
Roughly 41% of raw answer-percent values are CMS-suppressed (typically small-sample facilities) and are excluded from aggregates. Population figures are cited U.S. Census Bureau Vintage 2025 estimates for the 50 states, DC, and Puerto Rico, and 2020 Decennial Census counts for Guam, the U.S. Virgin Islands, American Samoa, and the Northern Mariana Islands, since those territories aren't part of the annual postcensal program.