M I MaHBuB

Home / Projects / Healthcare Data

Healthcare Data Exploratory Analysis

What a closer look at 55,000 hospital records revealed.

Completed

Python EDA case study · Monroe University

Exploratory analysis of 55,000 hospital records in Python. Statistical tests showed every variable was evenly distributed and unrelated, pointing to synthetic data and showing why data should be validated before drawing conclusions.

RoleAcademic project TechnologyJupyter • Matplotlib • Pandas • Python • SciPy • Seaborn

What I found

  • Found 4,966 hidden duplicates that matched other rows in everything except age, leaving exactly 50,000 records.
  • Those duplicates created a false billing difference (p = 0.046) that disappeared once they were removed (p = 0.15).
  • Conditions, medications, insurers and test results were all evenly spread, with chi-square tests finding no real differences.
  • Every stay length from 1 to 30 days was equally common, and age, billing and stay length were unrelated.