Look Upstream: AI, Federal Data, and the Politics of What Gets Counted

Before an algorithm produces a single outcome, institutions have already decided who gets counted, who gets erased, and what questions no longer need an answer.

The US federal government is expanding AI use in high-stakes decisions that affect people’s civil rights, privacy, health and safety, access to education, housing, employment and public benefits, and exposure to law-enforcement and immigration while simultaneously transforming the data, demographic variables, classifications, research programs, government webpages, publications, and risk frameworks through which social inequality can be observed and measured.

These changes are being accelerated by the federal adoption of AI outlined in Executive Order 14179 and President Trump’s AI Action Plan. The latter calls for rapid deployment of AI systems across federal functions with the aim of launching “a new era of human flourishing, economic competitiveness, and national security for the American people.”

These “upstream” changes to AI policy combined with changes to the demographic categories federal agencies choose to recognize and prioritize lead to the “downstream” consequence that specific dimensions of inequality will be measured and researched less, or classified differently.

This raises the question: What happens when the government increases the deployment of AI and automated decision-making, while also changing the mechanisms that allow it to identify affected populations, distinguish among demographic groups, measure inequalities in health, housing, education, employment, safety, and access to public benefits, and assess whether policies affect those groups differently? Groups that are already undercounted or poorly represented in federal data, such as LGBTQ people, racially and ethnically marginalized communities, women, people with disabilities, and immigrants, will be disproportionately impacted by these changes.

What is disappearing—and what stops being produced?

Since January 2025, the scale of federal data and information deleted, terminated, or altered has been unprecedented, especially because these resources benefit “American lives and livelihoods in ways most people never see.” The federal government has deleted entire datasets, removed demographic variables from continuing surveys, wiped content from webpages and public-health resources, and terminated thousands of research grants using a “detection list” with keywords like “gay,” “BIPOC (Black, Indigenous, People of Color),” “indigenous,” “tribal,” “melting pot,” “equality,” and other similar terms.

But not all information loss or disappearance of government data works the same way or has the same impact. The Federal Data Termination Tracker is a new initiative to determine what data the United States has actually stopped measuring. The tracker was created to address what had actually been lost. Previous efforts to address this question often conflated different types of federal information loss and used inconsistent definitions, sometimes grouping deleted webpages and tools, temporary access disruptions, removed variables, altered data products, and the termination of future data collection. The Federal Data Termination Tracker, created by dataindex.us, a collaborative group of data policy experts, developers, and researchers, defines data termination differently.

Different types of data loss are recoverable in different ways. A modified or deleted government webpage might still exist in the Wayback Machine or a data rescue group archive. Removed variables may change what a continuing survey can measure, but a terminated collection stops generating future observations of a population. And a canceled research grant may not produce any evidence at all. These distinctions show that deleting information or data may not be the most consequential loss; the information that is never produced creates an immeasurable gap.

In addition to deleting and terminating data collection, the federal government has also canceled thousands of research grants, meaning evidence of disparities for affected communities may no longer exist, and could lead to less accurate and fair automated decision-making, while widening gaps in opportunities, health, wealth, and overall well-being.

Terminating these research grants has a chilling effect on certain types of research. Grant cancellations have led researchers to reexamine their own proposals to determine whether they use problematic or risky language. When key research terms become associated with risk, researchers are forced to adapt: They reframe research questions or abandon them altogether, leading to evidence that is never produced and, ultimately, changes to the information environment. When upstream changes are made to federal government data collection, such as deleting entire collections and suspending the collection of certain demographic variables on surveys, this also diminishes agencies’ abilities to track future impacts on targeted populations. These changes also affect AI systems by making training data less representative of diverse populations and, in the worst case, erasing entire communities from data, research, and public dialogue.

This absence becomes compounded over time. When data is no longer collected, it can no longer be analyzed, and when data isn’t analyzed, it can no longer inform policy, and inequity becomes institutionalized.

Government is automating more as measurement narrows

Alongside data deletion, research grant cancellations, and the removal of public information from federal websites, the current administration is working to scale the use of AI systems across government agencies. A July 2025 Government Accountability Office (GAO) report found that AI use cases nearly doubled from 571 in 2023 to 1,110 in 2024 across eleven agencies, with generative AI use cases rising from thirty-two to 282 across multiple areas. An April 2026 report released by the Office of Management and Budget found a 105 percent increase in AI use cases from 2024 to 2025 across 56 submitting agencies, with the Department of Health and Human Services reporting the most active use cases at 447.

AI systems already use federal data across agencies to help government staff determine eligibility for federal awards and payments. A June 2026 Government Accountability Office report identified more than one hundred federal data sources that agencies may use for eligibility determinations and found that missing, invalid, or incompatible data can undermine those determinations and limit the use of AI and advanced analytics. Data is therefore needed not only to support automated decisions, but also to measure whether those decisions produce different outcomes across populations.

This is where removing demographic variables and changing measurements start to interact with computational systems that may have consequential outcomes. When the government removes certain variables from a dataset, it becomes nearly impossible for researchers to ask questions about the impacts of automated decision-making on different groups. Excluding variables also undermines researchers’ and watchdog groups’ abilities to hold the government accountable for adverse impacts. As Dhruv Khurana and Andrew Short of UCLA’s Fielding School of Public Health have noted:

When disparities are observable, systems can set targets, monitor performance, and trigger corrective action; when they are not, inequity becomes administratively invisible. Safety-net systems that cannot disaggregate outcomes by SO [sexual orientation] have no denominator for equity accountability; the disparity exists in the population but disappears from the administrative record.

Removing variables for sexual orientation and gender identity (SOGI), and for race, ethnicity, or diversity, equity, and inclusion-related (DEI) categories from data collection reduces the dataset’s capacity to make population-scale disparities visible.

In such cases, data termination ends future collection of deleted or terminated data, but also means previously collected data may no longer be available to train developing large language models (LLMs). This poses serious problems for government agencies that increasingly rely on AI systems to make major decisions that affect people’s lives, including, for example, who receives benefits. Once people are excluded, AI bias audits cannot examine missing variables because they no longer exist.

This shift is occurring alongside the AI Action Plan’s directive that the National Institute of Standards and Technology (NIST) revise the AI Risk Management Framework to eliminate references to misinformation, DEI, and climate change. Removing these subjects from a risk-governance framework does not, by itself, make the associated harms unmeasurable. It can, however, narrow the categories, questions, benchmarks, and reporting expectations through which harms are identified, tested, documented, and mitigated. When combined with the removal of demographic variables and the termination of data collection, these governance changes may make it harder for agencies, researchers, and affected communities to detect when AI-driven policies deny benefits, increase surveillance, misclassify people, or otherwise harm some groups more than others.

Why the information environment matters

Government information, datasets, statistics, reports, and publications are important inputs for developing and training AI systems. Although deleting a webpage or content on a government website doesn’t automatically remove it from a model that has already been trained, AI models change over time as new versions are developed using updated data and methods. A future version of an LLM may be trained on or retrieve from a different information environment. LLMs are not truth engines. They are statistical inference engines. They make correlations based on data, and if the data skews one direction, it will eventually be reflected in the LLM’s outputs.

A report by researchers at the Open Data Institute (ODI) examined how much UK government data sources are represented in and contribute to the performance of AI models. They found that government websites and information were “demonstrably important” data providers for LLMs and are a key source of information, especially for subjects that aren’t widely discussed online. Although this study focused on UK data sources, it demonstrates the importance of access to, and accuracy of, government data and information for developing AI systems.

Faculty from the University of Iceland conducted research confirming that the quality, currency, and trustworthiness of the information an AI system is trained on or can access materially shape the answers it provides. They found that a curated knowledge base might fail if its underlying information is outdated or incomplete, while a system using open-web search could address a wider range of topics but may be less accurate.

This demonstrates that if a corpus of public information, such as data related to climate change, DEI, or transgender communities, is no longer updated because data collection has ended, an AI system operating within that environment faces an information-quality problem: the information may not disappear immediately, but can become increasingly outdated, incomplete, and unrepresentative, weakening the reliability of what the system retrieves or produces.

When content is modified, deleted, or omitted from government websites, when research grants are canceled, and when datasets are removed or no longer collected, the information environment on which AI systems rely becomes less complete and potentially less reliable. Because the upstream effects of these losses are rarely visible or discussed, most users will not know what evidence is missing, whose experiences are no longer represented, or how those absences have shaped the answers, scores, recommendations, and decisions the systems produce.

How information environments are shaped

AI information environments can be shaped by addition as well as subtraction. Much of the discussion of AI information manipulation focuses on actors that might introduce data poisoning, RAG poisoning, propaganda, SEO manipulation, data void exploitation, synthetic content pollution, or injecting other material into the environment AI trains on or retrieves from. We can think of this as introducing information that changes the data environment that AI retrieves sources from.

These mechanisms largely concern addition or amplification: placing more information into the environment or increasing the likelihood that AI systems encounter it. The federal-data changes, research grant cancellations, and public information removal examined here raise the inverse question. What happens when authoritative information is removed, when demographic variables disappear from continuing datasets, or when the government stops producing particular observations altogether?

What becomes harder to observe?

When data or information about a population is deleted, that doesn’t mean that population ceases to exist. But it does mean that the affected population becomes less distinguishable within official government data. Removing SOGI data from ongoing surveys means that identifying gender-diverse respondents will no longer be possible. This can be considered a form of administrative or statistical illegibility, which can obscure “real-world harms and other impacts of policy decisions,” and the types of questions that can no longer be answered.

When a population or demographic group can no longer be distinguished in the data, differences in potential policy outcomes become harder to measure. This means meaningful differences for some groups may disappear into homogenized rates, which can reduce the benefits a population receives as decisions about societal benefits such as health, employment, safety, education, or housing are increasingly aggregated.

The Williams Institute’s February 2026 report on removing SOGI data collection variables emphasizes the importance of this data in helping policymakers, researchers, and service providers identify and address disparities and equitably allocate resources in order to meet the unique needs of a population that has been and continues to be disproportionately impacted. The report further demonstrates what is at stake for vulnerable groups when the data is removed or terminated, stating that these “decisions affect the health and safety of individuals and families, with consequences that span multiple domains and will continue to unfold over time as existing data becomes outdated and new data is not collected.”

When a measurement system stops distinguishing or acknowledging a population, researchers, journalists, policymakers, and computational systems have less direct evidence to describe and understand a population’s conditions.

Critical AI literacy and knowledge resilience

AI systems do not encounter the world directly. Long before a model generates an output, assigns a score, flags a case, recommends an action, or supports a government decision, institutions have already decided what data to collect, which categories to recognize, which research to fund, which sources to preserve, and which risks to measure.

For these reasons, typical approaches to AI literacy, such as asking whether a chatbot hallucinated, whether an answer is biased, or deciding if a source is credible, aren’t enough.

Users of AI systems need to begin “looking upstream” to ask critical questions about who created the evidence, what populations are being measured, what classifications exist, what has stopped being collected, what research has been canceled, and how systems are deemed  “neutral” in the first place.

These questions matter because AI systems inherit information environments that institutions have already structured. The government doesn’t necessarily own the truth, but the data it generates can encode exclusion, surveillance, and classifications that harm populations.

That is why universities, marginalized communities, journalists, libraries, and archives remain essential to preserving existing data and information. Although the coordinated effort to systematically delete official government data is alarming, groups such as the Data Rescue Project have been stepping up to preserve endangered public data. The Data Rescue Project partnered with the Williams Institute, ICPSR/DataLumos, and the Movement Advancement Project to preserve federal datasets that contain measures related to sexual orientation and/or gender identity (SOGI). The Data Rescue Project’s SOGI initiative highlights concerns about preserving this data, considering these targeted categories are especially vulnerable to current changes to federal data collection.

When content is modified, deleted, or omitted from government websites, when research grants are canceled, and when datasets are removed or no longer collected, the information environment on which AI systems rely becomes less complete and potentially less reliable. Because the upstream effects of these losses are rarely visible or discussed, most users will not know what evidence is missing, whose experiences are no longer represented, or how those absences have shaped the answers, scores, recommendations, and decisions the systems produce.

avram anderson is a tenure-track faculty librarian at California State University, Northridge, where they serve as the Collection Management Librarian, teach The Information Ecosystem, and are developing a new course, Critical Thinking in the Age of AI. Their research, teaching, and public scholarship examine critical AI literacy, critical media literacy, censorship, algorithmic accountability, platform power, and the effects of emerging technologies on marginalized communities. They are also a coauthor of The Media and Me: A Guide to Critical Media Literacy for Young People. Read other articles by avram.