Colleges, seen differently.
Steve Lattanzio’s 2018 experiment, rebuilt: let the data reveal the structure of American higher education, then examine what that structure can—and cannot—tell us.
Official College Scorecard release: June 10, 2026 · Completed September 5, 2026 · Offline report
A map, before a league table.
Compare a single annual snapshot, trace a college’s path, or follow the leading public and private institutions across two decades.
Annual rankings
VORS is the latest-cohort program-adjusted graduate earnings residual, in 2024 dollars. It is not an annual historical score; it stays attached to its 2022–23 earnings measurement when you change the annual file.
How the annual model works
Dates and default
The explorer covers 2011–12 through 2024–25 using the official June 10, 2026 College Scorecard release. The default is 2024–25 with last observations carried forward. The latest file meeting the original observed-data completeness rule remains 2020–21; that year anchors the fixed model and score scale. Neither label means every field describes the same student cohort. These are revised historical files retrieved in 2026, not archived as-published rankings or a prospective backtest.
Annual eligibility requires observed PREDDEG=3 and observed UGDS>100. Eligibility is never fabricated by carrying enrollment into an absent school-year. Historical schools are included without restricting to current survivors. The 2025–26 file lacks an eligible enrollment population. The six core gates are enrollment ≥90%, completion ≥80%, and admissions rate, graduate debt, ten-year entry earnings and three-year default rate each ≥50% coverage. Observed and carry-forward coverage are reported separately.
Last observation carried forward
Within each UNITID, sort observations in ascending year and carry the last nonmissing value into subsequent missing cells. Use only records from 2011 onward. Never borrow another school's values, fill from a later year, or stitch mergers. No maximum age is imposed. If there is no earlier observation, the value remains missing. Privacy-suppressed current cells may receive a past publicly reported observation; this does not recover the suppressed current value. It is explicitly an old observation.
For monetary variables, first calculate the observed source year's percentile among that year's eligible schools, then carry that percentile forward. This holds the last observed relative position fixed until another observation arrives. Raw monetary values and their source years are also retained in exports. A carried 2020 earnings percentile is not a new 2024 earnings measurement. Student cohorts and dollar vintages still vary by field.
All model columns have source-year provenance in results/feature-source-years.csv. The panel displays equal-domain fractions observed now, carried, and still missing, plus mean age of available inputs. The input mask marks availability after carrying. Freshness and age are audit fields, not additional neural inputs. Remaining missing values become zero after standardization, with a zero mask; their reconstruction targets are ignored. Training randomly hides 15% of available inputs. Rankings require ≥50% available coverage, averaged equally over nine domains, and at least six domains ≥25% available.
Architecture and ensemble
The actual encoded vocabulary has 338 numeric/one-hot columns across nine domains. Each first-stage autoencoder receives its values plus availability masks, uses a 16–64-unit hidden layer, produces six coordinates, and reconstructs its domain. Across those separate inputs there are 676 value/mask channels. The repayment domain has two values, so its six coordinates expand that domain rather than compress it.
The nine six-coordinate domain representations are standardized on reference training institutions and scaled by 1/sqrt(6), giving each domain equal total training variance. Their concatenation supplies 54 signal coordinates. The top autoencoder has architecture 108 → 32 → 2 → 32 → 54: 54 values plus 54 masks at input, with a two-dimensional bottleneck. Hidden layers use tanh; reconstructed values use a linear output.
There are 100 random initializations of the top autoencoder. They share the nine first-stage encoders. Each run is centered and isotropically scaled on reference training institutions, then matched to the first run by orthogonal Procrustes alignment, allowing reflection. One final common rotation aligns the mean plane with equally standardized completion, retention, and ten-year earnings percentile. The perpendicular axis is signed toward enrollment. These explicit anchors mean the direction is an outcome-oriented profile, not an unsupervised discovery of a uniquely correct definition of quality.
All encoders, reference scaling, alignment and score calibration are fixed across years. Score = 50 + 10 × reference-standardized vertical coordinate; the reference training population has mean 50 and population SD 10. Historical monetary percentile normalization is a retrospective annual population operation. The model is not fitted separately each year.
Geography and categories
State/territory is represented by one-hot columns learned from the reference training vocabulary. Latitude φ and longitude λ become three continuous coordinates: cos(φ)cos(λ), cos(φ)sin(λ), sin(φ). This preserves geographic continuity across ±180° longitude. The school domain contains these coordinates, state, and one-hot locale: city, suburb, town, rural. Control, highest degree, distance-only and open-admission fields also use one-hot encoding. Missing or unseen categories mask their entire block. Geography affects the representation and peer relationships; it does not automatically remove geographic effects or establish causal adjustment.
The starting vocabulary is the audited set from the initial reconstruction, plus geography. Selection requires reference training availability ≥25%, variation, and ≥25% observed coverage in at least five annual populations. Unstable classification and test-policy codes are excluded. Numeric clipping and standardization use reference training data. All inclusion decisions and exact input-column names are exported. This choice is a conceptual reconstruction of the original project rather than its unavailable original implementation.
Peers, plots and selectivity
The institution panel lists eight nearest other institutions in the same year's 54-dimensional representation, among rank-eligible schools. Distance is Euclidean after the equal-domain scaling. Peer distances do not come from the 2D map. Clicking a peer selects it in the current year.
The original blue/teal/copper sector palette, cream background, serif headings and gold selected-institution diamond are restored. Bubble area scales with enrollment, with display caps. Spaghetti line width follows selected-year enrollment, also capped. The highlighted cohort can be reset to the top ten public and top ten private nonprofit institutions of the selected year, or modified through search.
Smooth history curves use piecewise cubic Hermite interpolation (PCHIP). They pass through the annual values, preserve local monotonicity, and do not overshoot each interval's endpoint range. This smooths the drawing, not the annual data or its noise. Annual hover markers retain exact scores/ranks; gaps stay gaps. Straight segments are available as an option. Trajectories on the institutional map use the same fixed coordinates.
The quality/admission-rate plot maximizes two objectives: profile score and admission rate. A frontier school has no competitor in the displayed population with at least as high a value on both axes and a strictly higher value on at least one. The frontier is recalculated after sector and freshness filters; its connecting segments are guides rather than feasible interpolated colleges. Current-year admission-rate observations are the default, with carried values optional and source dates visible. Admission rate is not an individual applicant's admission probability. Admissions and SAT also influence the quality profile, so the frontier is a descriptive screen, not causal value added or independent evidence of overperformance. Small score differences and unusual institutional missions can strongly affect frontier membership.
SAT isoclines use a separate ridge-regularized quadratic surface for each year, fitted only to that year's observed SAT averages, never carried SAT values. Isoclines are 50 points apart and clipped to the reporting schools' convex hull and the 400–1600 range. Five-fold validation groups institutions by OPEID6. Since SAT is a model input, the surface is descriptive, not independent validation. Test redesign and changing test-submission composition prevent interpreting contours as an equated SAT time series.
Validation and limitations
Reference institutions are split by OPEID6 reporting group into training, validation and test sets. No reporting group overlaps those partitions. Neural fitting and standardization use the reference training/validation split. Feature availability screening and annual percentiles are retrospective, so held-out reconstruction is not an out-of-time forecast test.
Ensemble convergence compares prefixes of 1, 5, 12, 25, 50, 75 and 100 members with the aligned 100-member consensus on the default year's ranked schools. The first 12 differ by median 7 ranks (90th percentile 20); the first 50 by median 3 (90th percentile 10). Random subsets give median score RMSE of 0.394 points for 12 and 0.135 for 50 relative to all 100. The 100-member comparison is zero by construction, not proof of convergence to an infinite ensemble. Initialization ranges describe member variability, not confidence intervals, and omit first-stage, feature-selection, data and anchor uncertainty.
Carry-forward reduces artifacts caused by a field disappearing, but substitutes persistence for unknown change. Old data can make institutions appear more stable than they were. Differences can still reflect reporting definitions, student composition, pandemic-era outcomes and population changes. Scores describe relative institutional profiles rather than absolute educational improvement.
Source and reproduction
Sources and SHA-256 checksums are pinned in data/manifest.json. Open College-Atlas.html locally in a browser; it needs no server or external script. Run longitudinal.py, annual_extras.py, ivy_vors.py, validate_longitudinal.py and build_annual_report.py in that order to regenerate this revision after obtaining the pinned raw archive. Models, 100-member scores, annual coordinates, peer lists, source years, feature audits, convergence and validation are included. The raw archive can be downloaded by rebuild.py; no API key or GPU is needed. Exact floating-point outputs can vary across BLAS/library versions.
## Hidden Ivies
Hidden Ivy means a non-Ivy school whose mean distance to its three nearest Ivy League institutions, in the selected year's 54-dimensional representation, is no greater than the largest corresponding distance among the eight Ivies themselves. Ivy self-distances are excluded. The threshold is recalculated each year among rank-eligible institutions. Actual Ivy League members are labeled Ivy League, never Hidden Ivy. This restores the earlier reconstruction's inspectable similarity rule; it is not a probability of Ivy membership, an official designation, or a causal quality assessment. Carried data and geography influence distances. The optional black circle outlines use the same flag as the card and table.
VORS — value over replacement school
VORS restores the program-adjusted earnings residual from the earlier reconstruction: observed median earnings four years after graduation minus the out-of-fold model prediction. The model adjusts for measured institutional/student context and broad program mix, with five-fold validation grouped by federal OPEID6 reporting group. It is preserved unchanged, rather than relabeled as a prediction from the new 54D model. Positive values indicate earnings above that model's expectations. It is an observational residual, not a causal institutional contribution.
This VORS estimate describes 2017–18/2018–19 completers, with earnings measured in 2022–23 and expressed in 2024 dollars. It is the latest-cohort estimate available in this reconstruction and does not change when the annual-file selector changes. Historical cards and tables explicitly identify it as latest-cohort, not an as-of historical value. Missing estimates display a dash. The data export retains the earnings cohort, measurement years and dollar year.