What is a CIP code? The federal taxonomy behind every program number in higher ed
A plain-English explainer of CIP codes: the two, four, and six digit structure, who maintains it, the 2020 revision, and every federal system that hangs off it.
Research
The Clema research team publishes original analysis and practical guides for institutional research and institutional effectiveness professionals.
A plain-English explainer of CIP codes: the two, four, and six digit structure, who maintains it, the 2020 revision, and every federal system that hangs off it.
First-time freshmen by state of residence, 2016–2024: New York is down 15,700, Pennsylvania down 13,100, Alaska down a quarter. With in-state reliance for every state, a full exposure ranking.
A practical framework for IR teams drowning in ad hoc requests: a five-criteria scoring rubric, word-for-word scripts for saying no to deans, faculty, and cabinets, and what to say yes to instead.
A plain-English RAG explainer for institutional research: how retrieval-augmented generation works on campus data, why it fits IR work, where it fails, and the questions to ask any vendor.
Institution-level EADA analysis of the decade to AY 2024-25. 32.01 percent of continuously filing colleges cut a team, 261 of those cutters ended the decade flat or larger, tennis and golf supplied 40.5 percent of every drop, and the teams that disappear are small for years first. Two null results on Title IX, and the variable that actually predicts cutting.
On a balanced panel of 1,910 EADA filers, tennis is the only large sport in real decline: minus 97 institutions in ten years and minus 124 in twenty. Softball, golf, football and basketball all look like losers in the raw comparison and all four are positive once the institution set is held fixed. Meanwhile Other Sports went from 60 institutions to 432.
Women are 55.15 percent of undergraduates and 42.10 percent of athletic participants at the 2,037 institutions that filed EADA for AY 2024-25, a 13.05 point gap. It is wider than in 2014-15, football explains roughly half of it, and the 1,161 colleges that sponsor no football at all still carry 10.59 points.
Across 40,829 institution-year records from AY 2005-06 through AY 2024-25, EADA contains zero reported athletics deficits. In AY 2024-25, 1,288 of 2,037 institutions report revenue exactly equal to expense, to the dollar. What that means for the question your provost actually asked.
On a fixed panel of 4,902 colleges, office and administrative support fell from 13.95 to 10.10 percent of all staff between 2012 and 2024, a loss of 95,584 positions, while business and financial operations grew 43.7 percent. The full occupational breakout, plus the honest answer about where IR staff sit in IPEDS.
IPEDS 2024-25 completions divided by BLS 2024-34 annual openings for 17 CIP to SOC pairs, sixteen of them credential gated. Nursing 1.39, paralegal 0.25, 12 of 17 below 1.00, plus the state rankings and every reason to distrust them.
In PPD:2026, 159,461 of 209,321 programs have no published median earnings figure, and the same 159,461 rows carry no federal earnings-test determination. Here is the decision tree for why a cell is blank, the floors we could observe in the file, and how much a second source recovers.
3,804 of 109,790 IPEDS programs went dark after the 2022-23 award year, a 3.46 percent rate that has not moved in nine years. 76.7 percent of them awarded three or fewer credentials in their final year, the fields that churn fastest are mostly the ones adding programs at the same time, and field-level closure rates do not line up with the federal earnings test.
On the 3,361 GE programs in PPD:2026 where both regimes can be scored, 1,002 fail both, 569 fail only the old GE test and 19 fail only the new OBBBA test. The two casualty lists overlap on 63.0 percent of their combined 1,590 programs.
Sort 124 bachelor's fields by graduate earnings and the women's share of completions falls from 75.7% to 44.8%. Take nursing out of the top group and it drops to 27.0%. The field-level baseline IR teams can benchmark a program against.
Bachelor-level registered nursing produced 154,489 graduates in 2018-19 and 154,529 in 2024-25, a net change of 40 awards, while 113 more institutions reported a BSN program. Computing added 50,966. The full field-by-field breakout from IPEDS Completions.
We ran the official OBBBA earnings test across all 209,321 rows of PPD:2026. 1,220 programs fail, out of 44,052 actually tested, a 2.77 percent rate. Visual and Performing Arts carries 27.7 percent of the failures, associate degrees fail at 6.60 percent, and 455 programs miss by less than $2,000.
A general news and updates index for Clema covering product launches, whitepapers, IR benchmark releases, conference appearances, and company news relevant to IR and IE teams in higher ed.
A single hub of every Clema guide, research paper, IR benchmark, conference list, podcast, book, and definitional framework the IR team should bookmark. Updated as new pieces ship.
A when/lifecycle guide for IR teams: which months are peak, moderate, and low for ad-hoc requests, what reports and accreditations hit each season, and what to prepare ahead of time.
How a large IR team (8+ staff) uses Clema's analytics, peer benchmarking, and predictive analytics to scale analyses across schools instead of pulling data one school at a time.
How an IR office uses Clema's data platform to preserve definitions, sources, and terminology across a key-person departure, anchored in the 85% key-person dependency finding.
How an IR director uses Clema to surface the cohort definition on every answer, so "retention" stops meaning whatever the requester had in mind.
How an IE director uses Clema to keep stakeholders out of the email queue and let them ask dashboards in plain English, anchored in the whitepaper finding on dashboard adoption.
How a one-person IR office uses Clema to handle IPEDS pulls and ad-hoc dean requests in parallel during peak season, when capacity utilization hits 90%.
Quantified IR capacity benchmarks from 50+ interviews: 40 to 60% of team capacity consumed by ad hoc requests, 2 to 5 clarification cycles per request, and 220 to 3,200 hours annually reclaimable. By team size.
Where Clema does not fit today: very small institutions, institutions with no data source to connect, real-time pipeline needs, and bespoke statistical research. Honesty content for IR buyers.
A procedural guide for asking IPEDS questions in Clema: the prompt structure, the cohort and year disambiguation, the source check, and four worked examples.
A procedural guide for building a peer group in Clema by Carnegie class, control, region, or designation, and pulling the same metric for each peer in one conversation.
A procedural walkthrough for configuring SSO with Clema using SAML 2.0 or OAuth 2.0, including metadata exchange, attribute mapping, MFA, and the test login that confirms it works.
A week-by-week onboarding checklist for Clema: who does what, what gets signed, what gets connected, and what success looks like at each phase over 2 to 4 weeks.
A quick start for querying IPEDS in plain English with Clema: enrollment, graduation rates, peer comparison, finance, and faculty figures, with the federal source on every answer.
Five plain-English questions to ask Clema in the first session, what each one demonstrates, and how to interpret the sourced answer.
A side-by-side cost comparison of Clema and HelioCampus Theia Analyst for IR and IE teams, covering published pricing, implementation, and what is included.
A side-by-side cost comparison of Clema and EAB Edify for IR and IE teams, covering published pricing, implementation fees, and what is included at each tier.
Why general-purpose AI assistants fail on higher-ed data questions and what a higher-ed-native AI like Clema gets right: domain knowledge, real data, and sourced answers.
The structural mismatch between rising IR data request volume and flat team capacity, and why conversational AI is now the only path that scales without a hiring cycle.
A practical release calendar for the 12 federal datasets IR and IE teams depend on: IPEDS, College Scorecard, Pell, loans, cohort default rates, Clery, EADA, PSEO, PPD, DAPIP, and BLS OEWS.
In a study of 20 IR and IE offices, one term routinely carried several valid definitions at once. Here is why that happens, the eight layers it happens at, and why it breaks reporting even when the data is clean.
A diagnostic for institutional research teams: the Institutional Intelligence Gap, the three tiers (55% Large, 40% Moderate, 5% Small), a five-factor self-assessment, and an illustrative cost model for key-person dependency in higher ed.
Creating data definitions is achievable. Maintenance is where most efforts collapse. Here is a six-step framework, an AI-readiness argument, and tier-by-tier next steps for IR/IE teams.
A directory of 25 ways to learn institutional research: graduate certificates, AIR professional development, free IPEDS and PDP training, executive programs, and self-paced courses, with formats, costs, and links.
IR teams already sit on years of attendance, GPA, LMS, and aid data. A practical guide to turning that record into forward-looking predictions that change what advisors, deans, and provosts can do this term.
A plain-English guide for IR leadership on why classical ML (specifically XGBoost) outperforms LLMs on student-level retention and graduation prediction, what the benchmark literature actually shows, and how to read a per-student risk score.
A straight comparison between the open-source predictive-analytics stack ($40 to $95 per month) and the major incumbent platforms ($30K to $200K per year). Covers infrastructure cost, licensing, where incumbents are strong and where they leave gaps, and a build-vs-buy framework.
A practical look at the features, engineering choices, and timing decisions that drive a real term-to-term retention model. Written for IR leadership who want to understand what the model sees before they hand its output to advisors.
The graduation use case in detail. Features, course-combination effects, plain-English metric reading, the small-group floor that prevents false certainty, and the drift management every long-horizon prediction needs.
A practical guide to the predictive questions Provosts and CFOs own. Enrollment forecasting (ARIMA, Prophet, LSTM), yield prediction, program margin and viability, IPEDS-based screening, and the honest limits of each.
PPD:2026 looks like a single dataset. On disk it's six Excel files joined on opeid6 + credlev + cip4, with mixed-vintage dollars, statistical-noise privacy suppression, and a CIP-4 key that doesn't match the CIP-6 the STATS test will use. Here's where the analyst hours actually go, and how to skip them.
35% of universities have no formal intake process, and 77% of IR offices require extensive clarification before analysis begins. Analysis of 18 university data request forms reveals what the best systems do differently.
IPEDS 2023–2024 data reveals four-year institutions growing 3x faster in dual enrollment than community colleges, while community colleges carry 32.4% structural reliance. A full side-by-side benchmark.
Discover what IR teams actually benchmark most often, from student outcomes and faculty metrics to aspirational peer lists, based on evidence from 50 institutional research interviews.
Research shows 80% of IR data requests arrive missing critical elements. Learn how Request Intelligence (the quality of requests at intake) determines whether your team spends time analyzing or clarifying.
Discover proven strategies from IR leaders to streamline data requests, implement self-service dashboards effectively, and reclaim up to 60% of your team's capacity for strategic work.