This site has published 130 items, nearly all of them readings of public files — and after enough of those readings, the files themselves become a subject. This essay is our first attempt at saying what we keep finding, and it is a different kind of page for us: the selection and the through-line are our judgment, but every factual instance below comes from an analysis we already published, is linked to it, and was re-checked against the same archived data before this page was built. Seven habits, then, of Chicago’s own data files — not of Chicago, and not of anyone’s intentions, which these files cannot show and we will not guess at.
A reading across this site’s own published analyses · every instance linked and re-verified against the archived source data · August 18, 20261. Columns that never vary
A column name can promise more than its values deliver. In CPS’s school progress file, the student-attendance column carries 88.3 percent on every one of 649 rows, and the teacher-attendance column carries 94.0 on every row — school-level names, one value. We now test every metric column for constancy before summarizing it, because a distribution computed on those columns would have looked perfectly real.
2. Units nobody states
The city’s bike-rack file has a quantity column that sums to 15,275 across 9,449 records — and no statement of what one unit is, while corral type labels carry their own rack counts. The pothole ledger’s filled-per-block counts are similar: real numbers, house rules unstated. Our practice is to sum such columns as the column’s own total and refuse to rename the unit.
3. Outcomes the categories cannot record
Some files cannot say no. The TIF Investment Committee’s decision file holds 1,698 recorded outcomes in exactly four categories — approved, further review, referred back, on hold — and no denial exists among them; its successor file records decision dates with no outcome column at all. Whether anything was rejected before reaching a recorded decision, the files cannot say, and that absence is itself a fact worth publishing.
4. Sibling files that never merge
Chicago’s bike racks live in two datasets: a current file that begins with 2015 installs and a 5,164-row file last touched in August 2011, with the years between them in neither — the city’s own description says it hopes “eventually to consolidate” them, and it has not. Until it does, no complete count exists in the city’s files, and the honest piece is about the split.
5. Names that will not hold still
Across sixteen budget ordinances, exactly one department label survives every file unchanged — Finance General, the catch-all. The TIF committee’s file spells a district two ways; a library visitor file misspells the system’s own central library. Label drift is why we fold names only deterministically — case-insensitive, variant disclosed — and never fuzzy-match, and why sixteen-year department arcs mostly cannot honestly exist.
6. Geography that fails its own coordinates
When a file carries both a community-area column and coordinates, the two can disagree. The mural registry’s area column matched its own coordinates on 147 of the 283 rows where both existed; the TIF deal file’s disagreed on 13 of 762. The habit cuts both ways: the parcel universe behind our two-flats list passed the same check 600 for 600. So the rule is not distrust — it is check first, then decide, and say which placement the page uses.
7. Words bigger than the file
The transportation-permit dataset’s own description calls it a list of permits granted; its milestone column records 76,446 cancellations and 11,979 denials, and 985,144 of its 1,801,871 applications sit at a status — Open — the file never defines, including 62 percent of those started in 2016. The city’s Pedway dataset describes roughly five miles; its own geometry sums to 6.66 miles. Descriptions and status words are claims about a file, and we have learned to check them against the columns the way we would check anything else.
What this adds up to
None of this is an accusation; record-keeping at a city’s scale is hard, files outlive the systems that made them, and every one of these datasets still supports real findings when read with care. It is a method, stated in public: probe what a row is before summing anything; test columns for constancy; leave units as the file states them or say they are unstated; treat categories, descriptions, and status words as claims to verify; check native geography against its own coordinates; and disclose every blank — the West Nile ledger’s 5,261 resultless rows included. The analyses linked on this page state in their method notes which of these checks they ran. If we have misread a file — or if you know why one of these files is the way it is — corrections come first.
Written by KCM Desk from this site’s published analyses; every cited figure re-verified against the archived source data on August 18, 2026. If you spot an error, corrections come first.

