From AI Adoption to Public Value: The Measurement and Learning Challenges for States
The most important question in government AI is not whether a tool works. It is whether the tool creates measurable public value. A state may launch a chatbot, deploy document automation, build a sandbox, or train thousands of employees. But the most important question remains: Did it improve public outcomes? The report states the challenge even more directly: by Stage 4, the question is no longer “Can we deploy AI?” but “Is AI delivering measurable public value—and how do we continuously improve it?"
That is the focus of Stage 4: Impact. Impact is about whether states have built the structures, culture, and discipline to measure outcomes, learn from results, and adapt over time.
Code for America’s public report defines Impact as the stage where AI systems are monitored and measured, feedback loops are embedded, governance frameworks are updated based on lessons learned, and training evolves alongside new technology. For states that are already in agile development cycles, these ongoing learning practices will be nothing new. But for those states that are still doing waterfall implementations, they will not be able to keep pace with the fast pace of AI evolution.
Measurement remains underdeveloped
The report is clear: measurement remains the next frontier. Even leading states are still developing systematic ways to measure the public value created by AI.
This is a major finding. Many states are moving quickly on governance, training, and pilots. But fewer have mature systems for evaluating whether AI is actually improving services over time.
The report describes the current landscape as “governance-heavy” and “measurement-light” in many jurisdictions. It also notes that there is limited public reporting on AI impact and that continuous learning infrastructure remains concentrated in a small group of states.
That imbalance is understandable, but it cannot persist. If AI becomes part of core government operations, evaluation must become part of core governance.
Public Value Metrics: What should states measure?
Too often, AI impact is discussed in terms of productivity alone: time saved, documents processed, calls deflected, or tasks completed faster. Those metrics matter, but they are not sufficient. Government should measure AI against public value. That includes:
Efficiency gains
Cost savings
Accuracy improvements
Reduced administrative burden
Faster service delivery
Better language access
Improved resident experience
Reduced error rates
Fairness across populations
Worker satisfaction
Public trust
Appeals and correction rates
Security and privacy performance
Long-term service outcomes
Disparate impacts across populations
Public reporting and transparency quality
Specifically for benefits access AI implementation, states should also measure:
Application completion rates
Time to benefit approval
Churn and procedural denial rates
Resident comprehension and trust
The report’s Road Ahead section predicts that states scaling AI will need clearer evaluation frameworks that track efficiency gains, service improvements, and public value created by AI systems. That is the right standard. AI should not be considered successful simply because it works technically. It should be considered successful if it helps government deliver better, fairer, more accessible services.
Inventories and registries create visibility
One practical step is to create recurring AI system inventories or registries. Phase 3 and Phase 4 both call out these inventories as being critical. The report notes that states with higher maturity in the Impact stage often rely on automated decision system inventories or AI system registries to create visibility into where AI is being used and to establish recurring reporting requirements. This matters because government cannot govern what it cannot see. An AI inventory can answer basic but essential questions:
Where is AI being used?
Which agency owns the system?
What purpose does it serve?
Does it affect residents directly?
What data does it use?
What risks have been assessed?
What human oversight exists?
How is performance monitored?
When was the system last reviewed?
How can residents seek explanation or correction?
Inventories are not just transparency tools. They are management tools for agencies, and they are oversight tools for legislative bodies and the public.
Feedback loops turn AI projects into a learning system
The report’s Road Ahead section emphasizes the need for feedback loops that inform procurement, system design, and policy adjustments. Essentially, agile development practices are needed in AI deployments to keep pace with the the AI industry's fast evoluation. Mature AI deployment and governance should create a cycle:
Set goals
Assess risks
Pilot carefully
Measure outcomes
Engage users and affected communities
Improve the system
Update policy
Decide whether to scale, revise, or sunset
That cycle is especially important as agentic AI and more autonomous tools emerge. The report predicts that agentic AI could dramatically shift the baseline for readiness and influence future stages of the rubric.
As AI systems become more capable, the need for disciplined oversight becomes more urgent.
What the 2026 evaluations show
Impact is the least mature stage nationally. The full report rates 29 states as Early, 15 as Developing, and only 7 as Established. No state is rated Advanced. The critiera for this stage examined demonstrable evidence of:
Quantified impact measurement
Recurring reporting cycles
Monitoring mechanisms such as inventories, audits, and dashboards
Feedback loops that adjust procurement, policy, or models
Cross-agency learning and transparency
This is a major shift. It moves states from asking “Do we have a policy?” to asking “What changed for residents, workers, and public agencies?”
Case study: North Carolina and measurable fiscal impact
North Carolina stands out in the full report because it has stronger evidence of measurable fiscal impact at scale. The report notes that states in the Established tier frequently quantify cost savings, fraud detection recovery, efficiency gains, staff-hours saved, and service-level improvements. North Carolina is singled out for reporting measurable fiscal impact and embedding audit and monitoring processes that reinforce learning over time.
This matters because ROI is not only a budget argument. It is also a governance argument. When states can quantify what AI is doing, they can decide what to scale, what to redesign, and what to stop.
Case study: Vermont and inventory-driven accountability
Vermont is one of the strongest examples of inventory-driven impact governance. The public report notes that Vermont maintains an Automated Decision Systems Inventory requiring annual updates and performance review. Its public reporting and advisory council documentation create recurring transparency and structured evaluation loops.
The full report identifies inventories as the backbone of Stage 4 because they create visibility, establish recurring reporting requirements, and provide a platform for later performance and risk assessment. Vermont, Washington, Connecticut, and New York are cited as examples of how inventories can become the structural backbone for continuous learning.
Case study: Pennsylvania’s measured operational gains
Pennsylvania’s document processing work also belongs in the impact conversation because it includes clear operational metrics: an 80 percent reduction in illegible or incorrect documents and more than 700 staff hours saved.
Those metrics matter because they connect AI to public administration outcomes. They show reduced friction for applicants and reduced workload for staff. The next level of evaluation would ask even more: Did application completion improve? Did processing times fall? Did fewer eligible residents lose benefits because of documentation problems? Did outcomes improve equitably across populations?
That is where state AI measurement needs to go next.
Takeaway
The next frontier for government AI is impact and will be defined by ongoing learning and measuring what matters. The leading states will be those that build feedback loops into procurement, governance, training, and system design. AI systems should not simply be deployed and maintained. They should be evaluated, improved, and held accountable to public outcomes. The strongest AI programs will be those that combine innovation with accountability, experimentation with evidence, and technical progress with human-centered service delivery that rigorously measures public value.
Closing: The future is human-centered AI
The public report closes with a clear vision: states are stepping forward with urgency, but the opportunity is not simply about adopting new technology. It is about shaping technology in ways that are human-centered and grounded in outcomes for communities. That should be the guiding principle for the next phase of government AI.
Human-centered AI in government means:
AI supports residents rather than confusing them
AI strengthens workers rather than replacing judgment
AI reduces burden rather than creating new barriers
AI is transparent enough to be trusted
AI systems are measured against public outcomes
AI governance evolves as evidence accumulates
The goal is not for government be more technologically advanced. The goal is to make government work better for people.