From Pilot to Production: What State AI Implementation Looks Like in Practice
Code for America’s Government AI Landscape Assessment describes Implementation as the stage where AI becomes embedded in government operations and systems. It moves from test environments into real workflows, real systems, and real services. Examples include AI-assisted case management, public-facing chat assistants, predictive tools for benefits administration, fraud detection models, and document automation.
The report frames the core implementation question this way: can government deploy AI reliably and responsibly in real operations? Pilots show what is possible. Implementation shows what government is willing and able to sustain. This is the stage where AI’s promise—and its risks—become most concrete.
Implementation is where complexity increases
A pilot can be narrow, temporary, and closely supervised. An operational system must function reliably over time.
That shift introduces new challenges:
Ongoing model monitoring
Bias mitigation
Cybersecurity
Change management
Resident communication
Staff training
Procurement oversight
Appeals and correction pathways
Integration with legacy systems
Performance measurement
The report notes that once AI becomes operational, it requires capabilities such as ongoing model monitoring, bias mitigation, cybersecurity, and change management in an area where few established protocols exist. This is why moving from pilot to production should be a deliberate decision, not a default next step.
Early implementation often focuses on efficiency
The report finds that many states are prioritizing efficiency gains over full service redesign. Early operational AI systems often focus on backlog reduction, document processing, fraud detection, eligibility triage, and internal workflow automation.
That pattern is understandable. These use cases often have clearer return on investment and can improve internal operations without immediately changing resident-facing decisions. For example, document automation can help agencies process paperwork faster. Fraud detection models can help identify unusual patterns for human review. Chat assistants can answer common questions. These tools may reduce workload and improve responsiveness.
But efficiency alone is not enough. Government AI should also be judged by whether it improves access, fairness, accuracy, dignity, and trust.
Benefits access is a critical test case
The public report explicitly adds a benefits access lens. That is important because benefits systems are among the most consequential areas where AI may be applied. Residents applying for public benefits often face complex forms, documentation requirements, confusing eligibility rules, long wait times, and burdensome renewal processes. AI could help reduce some of those barriers.
Potential use cases include:
Plain-language benefits navigation
Eligibility guidance
Document quality checks
Caseworker support tools
Translation and language access
Call center support
Renewal reminders
Triage for urgent needs
Policy and rules summarization for staff
But benefits systems also raise serious risks. AI must not become a hidden barrier between residents and services. It must not incorrectly discourage eligible people from applying. It must not automate harmful assumptions. It must not make decisions that humans should make with explainability for residents to understand, contest, or correct. The examples of Maryland and Pennsylvania below point toward a human-centered path: use AI to reduce administrative burden, improve clarity, and support workers—while preserving human accountability.
Implementation requires strong data foundations
States with consolidated data platforms, cloud-first strategies, centralized IT governance, and strong interoperability have progressed more quickly into established or advanced operational maturity. This finding should shape how governments think about AI investment. Buying AI tools without improving data infrastructure is unlikely to produce sustainable transformation.
The question is not just whether an agency has access to an AI model as all states have exposed their staff to one or more LLM vendors. It is whether the agency has:
Reliable data
Clear data governance
Secure environments
Interoperable systems
Modern case management tools
Human review workflows
Evaluation capacity
Procurement and vendor oversight
Without those foundations, AI may amplify existing administrative weaknesses rather than solve them.
Implementation should preserve human judgment
Responsible implementation does not mean removing people from public service. In many cases, the best use of AI is to support public servants so they can spend more time on judgment, empathy, and complex problem solving.
That may mean AI drafts a summary, but a caseworker verifies it.
AI flags a document quality issue, but the applicant can correct it.
AI suggests eligibility guidance, but final decisions remain explainable and reviewable.
AI helps identify service gaps, but policy leaders decide how to respond.
The goal is not automation for its own sake. The goal is better government service.
What the 2026 evaluations show
Implementation is where the national picture becomes much more cautious. The full report rates 26 states as Early, 14 as Developing, and 11 as Established in Implementation. No state is rated Advanced.
That is not surprising. Implementation introduces new complexity: monitoring, lifecycle management, procurement controls, vendor oversight, human-in-the-loop workflows, bias mitigation, cybersecurity, staff training, and change management. The public report notes that once AI is operational, states need capabilities such as model monitoring, bias mitigation, cybersecurity practices, and change management in an area where few established protocols exist. The state government trends mirrror and may in face surpass the progress documented in private sector research.
Case study: Maryland’s benefits navigation agent
Maryland is one of the most important implementation case studies because its AI work directly relates to benefits access. The public report highlights Maryland’s partnership with AI providers, including Anthropic, to deploy a generative AI-powered benefits navigation and eligibility guidance agent. The tool helps residents navigate programs such as food assistance, Medicaid, and housing support, identify programs they may qualify for, move through applications, and receive information tailored to household circumstances.
This example shows the promise of resident-facing AI: making complex public benefits easier to understand. But it also shows why implementation requires strong safeguards. Any benefits guidance tool must be accurate, plain-language, accessible, and clear about its limitations.
Case study: Pennsylvania’s intelligent document processing
Pennsylvania’s COMPASS document processing work is one of the clearest examples of AI improving a benefits workflow.
The public report explains that Pennsylvania pioneered an Intelligent Document Processing service that scans uploaded documents for legibility when users submit required documentation through the COMPASS benefits application system. The tool screens for blurriness, image quality, and relevance, allowing clients to resubmit unclear documents immediately. Code for America reports that the tool reduced illegible or incorrect documents by 80 percent and saved county assistance office staff more than 700 hours.
This is a strong implementation case because it does not replace eligibility decision-making. Instead, it reduces friction at a critical point in the application process. It helps applicants submit usable documents and helps workers process information faster.
Case study: Vermont’s ChatVT
Vermont shows a different implementation pattern: public information access. The public report highlights ChatVT, a virtual assistant on the state portal that handles common questions about state programs. It provides 24/7 plain-language responses on topics ranging from health services to DMV information, reducing pressure on call centers.
The value of this case is not that a chatbot exists. It is that Vermont pairs the tool with broader transparency infrastructure, including an Automated Decision Systems Inventory that requires annual updates and performance review.
Operational maturity requires governance and inventories
The full report identifies public AI inventories as one of the strongest signals of operational implementation. Inventories provide a governance mechanism documenting AI tools in use, responsible agencies, and sometimes the purpose or decision impact of those systems. The report notes that public documentation of operational fundamentals—such as service-level agreements, drift monitoring, retraining procedures, and incident response—remains weak even where inventories exist.
That is the next step for implementation maturity: not just knowing what AI systems exist, but knowing how they are governed over time.
Takeaway
Implementation is where AI becomes part of how government works. The strongest state examples use AI to reduce burden, improve service navigation, or support staff—not to remove accountability. The states that succeed will be those that pair deployment with monitoring, human oversight, strong infrastructure, and transparency. States that succeed will be those that scale carefully, focus on public value, build on strong infrastructure, and keep human accountability at the center.