AI adoption in government: building trust and accountability
Public sector organizations are under increasing pressure to improve outcomes for residents while making every taxpayer dollar count. When deployed responsibly, AI presents a practical opportunity to meet that challenge. The technology holds the promise of helping governments reduce wait times, simplify everyday interactions, and free public servants to spend more time solving complex problems for the people they serve.
Achieving real impact, however, requires more than isolated pilots. Fully realizing AI’s potential demands organizations rewire by redesigning processes, ways of working, and operating models to make the technology part of everyday service delivery.
This series explores how leaders can move beyond experimentation and take practical steps to scale AI responsibly, delivering measurable improvements in efficiency, service quality, and the experience of the people they serve. Our ambition is to help the public sector dramatically improve their capacity to deliver for residents by providing leaders with the insights and confidence needed to accelerate that journey.
By Hrishika Vuppala, Roger Roberts, Tim Fountaine, and Tim Ward
Rushing to implement AI may ultimately slow its impact. The solution is for the public sector to build trust into the system rather than bolting it on at the end.
Two things often seem true when it comes to revolutionary technologies: Organizations rush to implement them, and their potential initially exceeds their impact. That’s the case so far with artificial intelligence, which many enterprises have aggressively adopted without yet seeing meaningful results. The solution may be stronger controls.
As Peter Senge famously noted in his seminal systems thinking book The Fifth Discipline (Crown Currency, March 2006), the optimal rate of growth is significantly less than the fastest possible rate. That’s because pushing a complex system such as that needed to fully deploy AI too hard can create bottlenecks and rework, while embedding the right controls early can increase overall speed.
That’s especially true for the public sector, where leaders are typically forced to choose between caution and speed, because slowing down—through the monitoring and management of AI—tends to happen at the end of processes rather than being embedded throughout. The result is stop-and-start activity: reacting to events rather than catching them early, adjusting, and maintaining momentum.
In this second article of our series on rewiring the public sector, we examine the importance of having strong controls that build the trust organizations need to maintain momentum. Our belief is that the public sector agencies most likely to pull ahead in the next decade through the use of AI may be not the boldest or most careful but the most thoughtful—those that have embedded the controls necessary to maintain momentum and accelerate safely over time.
The paradox of permission
The AI mantra for many major governments in the past two years has moved from “be careful” to “go faster, safely.” It’s the right impulse, and public sector agencies in countries from Australia to Britain to the United States have been given permission to accelerate and expand their use of AI. Yet progress does not necessarily flow naturally from permission. AI adoption often still clusters in back-office pilots, stalling when it comes to the services residents actually use and where the value is greatest.
Part of the reason is that government is not a business. Businesses that get AI wrong may lose customers, money, and reputation; when public agencies get AI wrong, they may risk vital services. The result is the loss of something even harder to rebuild: the public’s permission to act. And because people tend to view the public sector as a single entity—“the government”—and can’t take their business elsewhere, a mistake at one agency can effectively become a mistake for all.
That’s why the real limit on how quickly an agency can act is not technology or budget but trust: the degree to which residents, policymakers, and public sector employees believe it can get AI right. In recent surveys, people express concern about AI and a desire for strong oversight—including those remaining ultimately accountable no matter when, where, and how AI is deployed. Trust and confidence are gained when there is transparency about the technology’s use and mechanisms to challenge decisions.
Beyond trust, the balance of the ability to act is shaped by potential consequences. Get AI right, and the credit is slow and shared. Get it wrong, and the backlash may be quick and individual, especially for agency leaders. Governments around the world have seen automated systems fail in public and paid for it in court or at the ballot box.
Given the high risk relative to reward, it’s no wonder many agencies regard the safe move as implementing a long committee-driven process or a series of hurdles to clear in sequence. Yet while that looks like prudence and diligence, it often ultimately changes very little. And, in many cases, the committee itself becomes the problem rather than the answer by impeding the innovation and iteration it was created to enable.
Designing controls from the start
When it comes to balancing the need to act quickly but safely with AI, a committee at the end of the process is a weak control mechanism. It is typically slow, because everything waits for approval, which often arrives late. The committee process also misses too much. By the time anyone reviews a finished system, the choices that mattered—which data to use, what it may decide, and what recourse residents have—are likely to have been settled months earlier (for leaders wanting to evaluate where their agency stands, see sidebar, “Six questions for the leadership team”).
The alternative is controls designed at the start and carried through the life of the system, with degrees of intensity depending on whether AI is being used for everyday or mission-critical tasks. In our experience, this ultimately delivers faster performance. Problems are surfaced while they are still cheap to fix, with reusable controls and a final approval stage that confirms work already done rather than uncovering work that was not. In this mode, leaders provide oversight not by attending committee meetings but by being actively engaged in the design of the control system to ensure it is tailored to an agency’s mission, ethos, and objectives.
Doing this well requires two components (Exhibit 1). The first is a set of controls following every AI use case from design to launch to ongoing monitoring, in which governance supports an environment that allows AI to scale safely. The second is the shared machinery making those controls practical: people, ways of working, technology, data, and change management—all supported by risk governance. The controls define what must happen. The shared machinery lets an agency do it repeatedly and at scale.
Image description: Two illustrative charts show the two components of managing IA risk through design and deployment. The first shows risk management through design and deployment in four steps: 1. Scope, 2. Get data, build, and evaluate, 3. Industrialize, and 4. Monitor and maintain. Within these steps are more-specific actions. Scope includes designing the solution. Get data, build, and evaluate includes obtain training data, building an AI app that solves scoped problems, and evaluating performance and business fit. Industrialize includes moving the model to the production environment, deploying the model for business use after final approvals, and managing inventory of all deployed tools. Monitor and maintain includes conducting live monitoring in production.
The second chart lists underlying capabilities to underpin risk management. These are the technical base, defined by technology and data; the people who run it, defined by talent and the operating model; and what keep it honest, defined by governance and change management.
End of image description.
Embedding risk management across the process
In the first part of the process, a single use case travels from idea to live service through five stages, with three approval gates along the way. The trick is to treat each control as an early design choice, not as a box to check at the end.
Design. Before anyone writes code, two questions should shape everything that follows. First, should the agency use AI here at all? The need to adopt the technology should be judged against publicly defined AI ethics principles and whether the same outcome can be achieved more simply or inexpensively in another way (for example, deterministic automation using rules-based systems). Second, how much freedom should the system have to act? The answer depends on the complexity and harm a wrong decision may cause (Exhibit 2). The best early uses of AI are often simple, low-risk, high-volume tasks for which it can handle most of the work. Decisions that materially affect individuals are different. In those cases, a person should remain in control. This is also the time to set clear rules for when the system must refuse, escalate to a person, or rely only on approved sources.
Image description: A matrix diagram depicts AI autonomy based on complexity of decision and potential consequences, from low to high, with decision complexity on the x-axis and the degree of consequence of a bad action on the y-axis. For decisions of low consequence and low decision complexity, AI can run them. For decisions of medium consequence and medium decision complexity, when AI recommends decisions, a human should approve it. For decisions of low consequence but high complexity, AI can propose a decisions, but a human should have the final say. For decisions of high or medium consequence and respective medium or high complexity or decisions of high consequence and complexity, a human should decide, while AI should only assist in making that decision.
End of image description.
Get data. AI-enabled workflows are only as trustworthy as the information they use. Agencies should use AI governance processes to verify that every source can be legally used, handle personal information carefully, and feed the AI system clean, approved material rather than a messy shared drive. Difficult test prompts should also be created early, allowing the system to be tested later against realistic failure modes.
Build and evaluate. At this stage, many controls are engineering choices. An AI-enabled workflow needs clear instructions, limits on the form of answers, tight permissions for any tools used, and filters on what goes in and comes out. The harder step? Proving that the system works well enough for real service. This requires repeatable tests, expert review, stress tests for misuse or data leaks, red-team tests for security and harmful outputs, and checks for fair and equitable performance across affected groups. Gate two in the process, approval to implement, should depend on this evidence, not on the project deadline.
Industrialize. Moving into production requires another set of controls, including locking in the AI-enabled workflow and prompt versions, setting limits on use and cost, adding emergency stop points, and documenting what the system can and cannot do. A sponsoring leader should name an owner, train staff on safe use and escalation, and keep a record of what the system does. Last, the finished use case should be registered and key artifacts retained, including its prompts, data sets, and test results, so that the next team can build on the work rather than starting again. At this point in the process, gate three is reached: approval to go live.
Monitor. Once the system is live, the controls still have to work. In our experience, more than half of all AI oversight work happens after production release. Alerts should flag unsafe behavior and attempts to trick the system. The agency should track who is using the AI solutions, what those solutions are doing, how well they are working, and what they cost to run and maintain. Tests should keep running, because AI-enabled workflows and supplier systems can change in ways that affect behavior. Someone should also ask periodically whether using AI is still appropriate. If the answer has changed, the agency needs to be willing to roll back and update the system. The objective is to create an environment that is both safe and fast, enabling AI applications to scale as efficiently and effectively as possible.
Building capabilities to make the process repeatable
Controls applied to individual use cases will not scale on their own. Public sector agencies need shared capabilities so the same discipline can be applied again and again, faster each time. That is the second part of the process detailed in Exhibit 1: It cannot be bought off the shelf or delegated to a vendor. Rather, it lies at the center of what an agency must build. Three broad capability areas individually and collectively reinforce the control and governance of AI-enabled workflows to strengthen trust.
The technical base. Technological capabilities are critical because overseeing AI use cases with logs in spreadsheets will not scale, especially as systems begin to take actions, call other systems, and link steps together. Agencies should have a common, governed platform connecting use cases with the models, tools, and data on which they rely (Exhibit 3). The platform should control which tools an AI system may use, manage identity and permissions, record activities for auditing, enforce policy, support testing, and maintain a register of approved models. It should also provide reusable building blocks such as hosting, orchestration, and shared agents.
Image description: A flow chart depicts the function of a capability gateway relative to use cases and models, tools, and data. First is use cases, including chatbots, fraud triage, drafting, and case summary. These flow into the capability gateway, which is governed AI platform, where each new use case inherits the guardrails of previous use cases. This includes identity and tool access, testing and assurance, guardrails and evaluation framework, logging and audits, and model and agent registry. These flow into AI building blocks, which are models and agents, tools, and data.
End of image description.
In addition, while newer risks may be less visible, they are just as important. Agencies should protect systems from malicious prompts or corrupted context. They also need to control wasteful resource consumption, which drives costs up as AI systems run. A wide range of innovative start-ups and major cloud providers are already moving in this direction as they enhance their governance layers, and the benefits compound: Once guardrails are built into the platform, every new use case inherits them and the cost curve of control bends with scale, creating greater AI delivery productivity and velocity.
A platform is only as good as the data beneath it. Reliable AI needs current, approved information to draw clear rules for the systems it connects to. It also needs privacy protections built in from the start. In the public sector, data is often sensitive, and misuse can cause serious harm. That means data quality and protection are not just housekeeping but what makes the system work—and acceptable to the public.
The people who run it. The second capability area involves creating truly multidisciplinary teams to build and run AI safely. This includes a product owner who understands the AI-enabled workflow well enough to redesign it and who owns the risk as well as the features; people with skills the public sector is often short of, such as evaluation, red teaming (cybersecurity testing), and AI safety; and emerging roles such as AI trust architects, who ensure that control technologies and techniques are continually advancing and that they are reused to the greatest degree possible. Agencies also need a workforce confident enough to work alongside agentic tools and enhance them, which is more than simple “human in the loop” oversight. Many governments today rely on contractors for these skills, but as AI becomes more important, they may increasingly choose to build expertise in-house. Doing so may require more than hiring; it likely will require rethinking roles, job classifications, operating models, and workforce policies to attract, develop, and retain AI talent.
Product ownership, perhaps the most important choice, is often where agencies go wrong. Making a chief AI officer responsible for trust can recreate the same end-stage checkpoint under a new title. For that reason, accountability should rest with the service leader affected by the AI system. Responsibility lies with the chief AI officer and central AI function, which together provide the platform, standards, shared testing, approved patterns, and scarce technical expertise. Delivery should still sit with the service team that owns the outcome. In practice, this means there should be cross-functional teams that own value and risk and a light central layer for common tools and guardrails. It also means measuring performance not by the number of pilots launched but by outcomes such as accuracy, safety, resolution time, and resident experience.
What keeps it honest and effective. As the use of AI-enabled workflows grows, two things prevent standards from slipping. The first is governance that fits into an agency’s existing risk approach rather than sitting apart from it: clear ownership, clear risk rules, and clear paths for escalation and exceptions. For example, a cross-functional forum can coordinate oversight without becoming a bottleneck. The second is change management, which is critical given how AI alters the nature of work. Leaders should be clear about where AI supports staff and where it removes work, building capability through hands-on coaching. They should also be open about the technology’s limits, model expected behavior, and measure and share real impact, such as time saved, errors avoided, and outcomes improved. While AI’s potential stems from data, algorithms, and talent, it must be supported by trust to achieve real impact. If trust declines, so does impact. Trust defined by strong controls is critical to accelerating adoption and achieving impact from AI investments.
Controls accelerate AI
The public sector plays a unique role in providing critical services to residents, and the cost of failure can be high. But as the opportunity AI presents to drive public sector efficiency and effectiveness becomes clearer, so does the need to find paths to deliver AI-enabled services safely and at speed.
Sustainable speed comes from designing the critical control system to support it. The agencies likely to deliver on the promise of AI in the decade to come will not necessarily be the boldest or the most careful. They will be the ones with the best engineering discipline—the ones that have thoughtfully implemented checks and balances across AI-enabled workflows rather than relying on committees at the end of a process. Strong controls help maximize impact by ensuring issues can be caught earlier, before large sunk costs, and by building a system that can accelerate adoption by mitigating risk, building public confidence, and assuring trust in the ability of governments to deploy AI.
Six questions for the leadership team
Public sector leadership teams wondering where they stand can start by asking six questions. In each case, the desired answer is not just “yes” but “yes—and here is the evidence.”
- Judgment before capability. Do we decide whether using AI is appropriate before we decide whether it is possible?
- Built in, not bolted on. Are the controls designed into how we build each system from the start rather than added at sign-off at the end?
- Demonstrable reliability. Can we show with evidence rather than assurances that each live system still does what we promised?
- Capability to repeat. Are we building shared capability—platform, data, and governance—so the next use case is faster and safer, not built from scratch?
- Future-ready workforce. Are we evolving our workforce, roles, and operating model so the organization can succeed in an AI-enabled world?
- Clear ownership. Does the leader of each affected service own its AI, with the center enabling rather than deciding?
ABOUT THE AUTHOR(S)
Hrishika Vuppala and Tim Ward are senior partners in McKinsey’s Southern California office, Roger Roberts is a partner in the Bay Area office, and Tim Fountaine is a senior partner in the Sydney office.
By Hrishika Vuppala, Tim Fountaine, Tim Ward, and Tony D’Emidio
AI has the potential to rewire how governments work, transforming public sector efficiency and effectiveness. But seizing the moment takes much more than a bolt-on approach.
July 15, 2026 – Governments around the world face mounting challenges that threaten their ability to effectively deliver essential services. Barriers ranging from fiscal constraints to skilled-workforce shortages and high consumer expectations are forcing agencies to do more with less, while declining public trust reduces their ability to act. The need for efficiency and innovation has arguably never been greater.
Technology has long been heralded as a potential solution, promising modernization, streamlined processes, and productivity gains across the private and public sectors. Recent breakthroughs in AI give reason for optimism—especially the application of generative and agentic AI—around how the technology can improve productivity and other measures. Yet its impact on services to residents remains elusive and many AI deployments seem stuck in “pilot purgatory,” not yet achieving results that reach residents or frontline workers as organizations struggle with data access, workflow integration, model risk, and high ongoing operating costs (Exhibit 4).
Image description:
A bar chart shows that the public sector trails the global average on AI quotient score. The global average score is 35 out of 100, while the public sector’s score is just 26. Other sector scores range from 32 for consumers to 44 for retail and high tech.
End of image description.
Can governments capture value from AI? It’s rarely, if ever, about the tools or technology alone. Unleashing the potential of AI in the public sector requires rewiring operations by redesigning workflows, adopting fundamentally new ways of working, and engaging the workforce to drive and scale adoption. Even the most advanced tools will underdeliver if they do not address structural issues, including outdated processes, fragmented decision-making, and misaligned workforce models.
That demands a proven, four-part approach undertaken simultaneously, not sequentially (Exhibit 5):
- Crafting a strategy that leads with the mission outcome. The focus should be on meaningful outcomes for residents at an appropriate cost, not on tools or technologies.
- Reimagining required workflows end to end. The time for piecemeal experiments or use cases is over.
- Building the operating system around the technology. Careful change management will be needed to drive adoption that scales.
- Keeping humans in the loop for any consequential action. There should be clear lines around what humans do and what AI does.
Image description:
A table highlights four moves that can be made in parallel to rewire government.
The first action is crafting a strategy that leads with the mission outcome. This includes prioritizing more than just cost-cutting, being bold about the ambition, starting with a clean sheet, and designing for continuous evolution.
The second action is reimagining required workflows end to end. This means ensuring teams can execute and innovate safely; adopting agile, product-based operating models; deploying scalable technology to reengineer workflow; building a robust data foundation to scale AI; and building financial operations discipline.
The third action is building the operating system around the technology. This means focusing on organizational change management; integrating human-centered design; providing real-time, data-driven, ethical AI governance; and evolving procurement and contracting for scalable outcomes.
The fourth action is keeping humans in the loop for any consequential action. This means defining human sign-off by consequence, not category, and making human-in-the-loop decisions a political safety net.
End of image description.
While AI is already part of people’s lives and has the potential to be a powerful force for good in the public sector, governments have very different regulatory regimes and degrees of organizational and public support. The road to rewiring can therefore be variable and lengthy. But the four-part approach detailed in this article can, when taken together, be a starting point to unlock AI’s potential to deliver better services to the public while building trust and resilience.
Understanding the public sector’s unique challenges
Governments face a structural mismatch between today’s demands and a workforce and technology architecture built decades ago. Fiscal pressures are mounting, with most advanced economies facing balance sheet constraints. Public sector headcount is well below historical peaks in many countries. And confidence in public sector institutions continues to decline.
In addition to these challenges, the public sector is simply different. Government agencies exist to serve residents by delivering critical services without a profit motive while balancing access, fraud prevention, due process, cost, and fairness. This is exactly where AI can do real harm if it is not applied correctly. Structural characteristics make AI in public sector services meaningfully different from AI in the private sector, as exemplified by the following:
- Most governments’ procurement rules were not designed for outcomes-based contracting, and multiyear appropriations with rigid line items are often incompatible with AI capabilities that evolve in months.
- Unified views of data common in private sector AI are much harder to replicate in public agencies, where resident, mission, and operational data are split across agencies for legitimate privacy and statutory reasons.
- Workforce and role rigidity that can come from job classifications, collective bargaining, and merit system rules can stretch role redesign cycles to a year or more, even when leadership is aligned.
- Decisions affecting individuals require explainability, reviewability, and auditability, raising the bar on rigor and human-in-the-loop oversight well above private sector norms.
What’s driving successful public sector AI efforts
The good news is that while these differences impose real constraints on the ability of governments to realize AI’s potential, they are surmountable. Examples throughout this article will show governments delivering on AI-enabled transformation today, not just in theory. And learning from these success stories is critical as an increasing number of countries seek to improve efficiency and effectiveness through technology: 95 countries today have national data and AI strategies, compared with less than 20 in 2020.
In our experience, public sector AI programs that deliver value do the following:
- Reimagine entire cross-functional, end-to-end processes (or domains) from the resident’s perspective. The agencies pulling ahead are choosing a whole cross-functional, end-to-end process workflow (such as a benefits journey, a licensing process, an inspection workflow, or an eligibility determination) and redesigning it from the resident’s perspective, with AI as one tool among several. McKinsey research finds that about 70 percent of domain-based programs reach production, compared with just 30 percent of programs led by individual use cases.
- Start work without hesitation. The most common stall pattern is a two-year program to “get the data ready” before moving. Agencies making progress do the data work concurrently with the AI work, not sequentially, and apply AI to the deterministic parts of a workflow. In our experience, there is ample opportunity to make significant gains with data already available.
- Make it clear AI is not about reducing head count. Workforces in many governments are already shrinking while demands and expectations increase. AI is a lever for existing workforces to meet the moment, empowering them to handle larger volumes while spending more of their time on judgment-intensive work only humans can do. This framing matters because public sector workforce skepticism is real: Only one in five public sector employees expects AI to meaningfully affect their daily work, and just 31 percent trust their employer to develop AI safely, compared with 71 percent across industries.
- Measure AI’s progress by outcomes, not by announcements. Beyond the dollars spent on the underlying technology, leaders should equally plan for a multiple of that on adoption, training, and capability building. Almost no public sector budget reflects this balance. And leaders measure success not by publicity but whether residents can tell the difference, through tangible progress such as shorter wait times, fewer rejections, faster emergency response, and clearer answers.
Where to start may matter as much as how. In a forthcoming companion piece, we will examine which government processes have the highest value at stake, as well as where the largest clusters of value lie. But what’s important for governments to understand today is the uniqueness of this moment in terms of the opportunity to rewire their operations—and how to go about it.
Seizing the moment: Four steps to rewiring government
Powerful advances driven by AI and growing openness to change have created a rare moment for government agencies to truly rewire processes at scale, unlocking meaningful productivity gains while improving service quality. Work could evolve into an AI-enabled partnership between people, agents, and robots as modern technologies, especially agentic AI, grow increasingly capable.
In the public sector, that holds the promise of improving operational agility, unlocking innovation, and fundamentally transforming the resident experience by accelerating execution, enabling parallel processing, enhancing adaptability and personalization, introducing elasticity to operations, and strengthening resilience. Their capabilities will likely continue to improve rapidly as models iterate, data scales, and infrastructure matures. But these capabilities will only translate into more effective and efficient services when governments take a holistic approach to rewiring that prioritizes strategy and people together with technology. Otherwise, governments risk AI following a familiar pattern of investment failing to deliver expected or potential impact.
Encouragingly, policy signals are fostering innovation, and there is a growing demand to procure and implement AI and to follow best commercial practices. Governments worldwide are actively pursuing AI initiatives to enhance productivity and the quality of their services, from national initiatives in Singapore, the United Arab Emirates, and the United Kingdom to initiatives in US states such as California and Pennsylvania. A full government rewiring effort augments and builds on these policy steps with a comprehensive approach balancing cost, efficiency, and effectiveness while leveraging both people and technology. It also demands strong leadership, cross-sector collaboration, and a focus on resident-centered outcomes.
In our experience, four steps can maximize the odds of a successful transformation:
Step one: Crafting a strategy that leads with the mission outcome
Technological investments are often made without a clear vision of the outcomes they are meant to achieve. Effective rewiring instead begins by defining mission-critical goals, whether it’s improving public health outcomes, reducing response times for resident services, or increasing operational efficiency. By focusing on outcomes first, governments can ensure that technology enables rather than distracts. This involves the following:
- Prioritizing more than just cost cutting. Investments should focus on achieving an agency’s mission-critical goals, such as improving resident well-being or enhancing public safety, rather than solely reducing expenses. Leading with outcomes helps ensure technology is deployed thoughtfully, rather than as a “hammer looking for a nail.”
- Being bold. Organizations with ambitious AI agendas are seeing the most benefit. Conversely, organizations that set low expectations often see only incremental change.
- Starting with a clean sheet. Unlocking the full potential of AI requires more than plugging agents into existing workflows. Rather than just digitalizing what already exists, agencies can take a hard look at workflows to identify inefficiencies, redundancies, and bottlenecks. Rewiring requires reimagining workflows from the ground up, often leveraging human-centered design principles to ensure they meet the real (not just perceived) needs of both employees and the public.
- Designing for continuous evolution. In an environment where AI capabilities evolve daily, waiting months to perfect a strategy before moving forward risks obsolescence. Leading organizations set guardrails and priorities early, then prototype quickly, learn in the field, and refine in waves. Governments can pair clear strategic direction with rapid pilots, continuous learning, and the flexibility to adapt as capabilities evolve, fueled by a funding mechanism that frees up investments concurrent with each wave.
Step two: Reimagining required workflows end to end
Generative and agentic AI may fundamentally reshape how public and private sector organizations advance their missions in the largest organizational paradigm shift since the industrial and digital revolutions. AI technologies are not simply new tools; they are redefining how work is structured, how decisions are made, and how services are delivered—with humans and agents working side by side. And because the capabilities themselves are evolving rapidly, the way government agencies operate and the skills people need must also change dramatically.
We see four key levers to drive organizational capabilities and ways of working:
- Ensuring teams—inside and outside of IT—have the skills, capabilities, and opportunities to execute and innovate safely. Effective AI adoption requires more than access to tools. As agentic AI becomes embedded in day-to-day operations and agents increasingly support service delivery, humans could move from completing tasks to delivering outcomes. This would require specific skills, such as AI fluency, agent orchestration, deep problem-solving, and quality control. That may demand agencies rethink every aspect of their talent systems—from roles to career paths to incentives to leadership models—providing consistent, on-demand training to build strong digital skills at every level.
- Adopting agile, product-based operating models with flat, cross-functional teams built for speed and scale. Agencies have an opportunity to move away from rigid, hierarchical structures to embrace agile, product-based operating models that concentrate skills, decision rights, and accountability. Critically, this is where the redesign of the work itself happens: McKinsey research shows roughly 60 percent of AI value comes from workflow redesign rather than layering models on top of existing processes.
- Deploying scalable technology to reengineer workflows. Digital factories integrating physical production with digital technologies such as AI agents can be launched within 12–18 months to accelerate delivery, enable rapid experimentation, and reduce execution risk. As agentic systems mature, agencies will need integration patterns that go beyond traditional APIs alone. Emerging standards (such as agent-to-agent) can help agents collaborate, while protocols, such as model context protocol, and existing APIs can connect agents to enterprise systems, data, and tools.
- Building a robust data foundation to scale AI across the organization. Research across the public and private sectors shows that data issues are a leading cause of AI project failure, whether due to poor data cleaning and management or simply insufficient data. This can be especially acute in the public sector, where sharing data across siloed, privacy-sensitive agencies presents a cultural and regulatory challenge. Establishing modern data architectures with clear ownership and federated governance can enable real-time decision-making. Critically, data work should occur concurrently with AI work, not sequentially.
- Building financial operations discipline. AI run costs can rise quickly in agencies that don’t track them, and many agencies don’t. Government CTOs are increasingly asking how to forecast, allocate, and contain inference costs before they become the line item that ends the program. Building financial operations capabilities around AI early is cheaper than retrofitting them after the bill arrives.
Step three: Building the operating system around the technology
Responsible adoption of AI captures intended value through an intentional organizational change-management program that scales digital solutions while building organizational skills such as procurement, training, culture, and oversight. Risk and ethics are embedded into the foundation of the work, ensuring the government’s unique position and responsibility are upheld in ethical principles from ideation to post-deployment monitoring. This transformation requires the following:
- Focusing on organizational change management. Building organizational capability is essential to capturing value from any technology investment, especially AI. This includes attracting skilled talent, ensuring public service remains compelling, systematically upskilling the existing workforce, and redefining roles to meet modern demands. It applies not only to IT but also to functions such as procurement, finance, legal, and hiring, where new tools and ways of working must be integrated into daily operations. Fully realizing the benefits is not inexpensive or automatic: McKinsey research shows that for every $1 spent on technology, $5 must be spent on change management to drive capability building, adoption, and buy-in successfully.
- Integrating human-centered design. Processes and systems should be designed in conjunction with employees to ensure usability, accessibility, and the ability to deliver on behalf of residents. Adoption will take time as humans learn how to partner with agents side by side, although a more intuitive design will ease the transition. Agencies can also co-develop new systems with the people and organizations who rely on their services. Agency leaders and administrators understand policy deeply, but they are rarely the end users—designing from a resident’s perspective improves clarity, reduces friction, yields more-effective outcomes, and strengthens legitimacy. Employees are also learning to revise work habits entirely, from utilizing AI to transcribe video calls and create meeting notes to embedding the technology to monitor their calendars, emails, and other common work tools. The intent is to use AI collaboration to both ease the administrative burden on employees and strengthen their higher-value tasks.
- Providing real-time, data-driven governance, ensuring AI is implemented responsibly, ethically, and transparently. Public trust is foundational to effective governance. As with other technologies, leaders should prioritize data privacy, security, and fairness to build confidence in new systems. Government agencies should also build monitoring and evaluation directly into workflows, and leaders should set clear expectations for human accountability and oversight (such as technology leaders validating code outputs while program leaders verify sources and policy interpretations). Without intentional governance, agencies risk either accumulating unapproved tools with unverified outputs or absorbing added compliance, quality, and reputational risk as adoption grows.
- Evolving procurement to avoid free-pilot traps and contracting for scalable outcomes. Government procurement was built to buy inputs—such as licenses, seats, and tidy demonstrations—but AI’s value only shows up when a redesigned workflow runs at scale. The “free-pilot trap” may deliver a no-cost or low-cost proof of concept that dazzles in a demonstration but then has to be rebuilt almost entirely to work in a scaled environment. The solution is to make procurement evolve in step with the technology—contracting for measurable, scalable outcomes and shared delivery risk so that the government contracts for impact that sustains rather than for tools that never leave the pilot stage.
Step four: Keeping humans in the loop for any consequential action
Even with a strong strategy, capable teams, and the right technology, risk issues can stall AI deployment in government. Agency leaders are right to ask hard questions about over-automation, accountability, and unintended consequences, and agencies pulling ahead design for risk management from the start rather than retrofit governance after deployment. Two moves matter:
- Defining what requires human sign-off—by consequence, not category. A benefit denial, a license revocation, or a public safety dispatch is a consequential decision requiring a human in the loop. A grammar check on a draft response is not. Most agencies default to a human in the loop for everything because not enough work has been done to draw clear lines around what humans do and what AI does (Exhibit 6).
- Making the human-in-the-loop decision a political safety net. Public trust in government AI rests on the ability of an affected resident to know that a human can review, override, and explain a decision. Designing for that capacity from the start—and saying so publicly—may be the single biggest determinant of whether an agency retains public trust.
Image description:
Two pie charts illustrate challenging trends for government investment in technology and states that AI alone cannot solve the gap between investment and impact.
The left-hand chart shows that more than 70% of federal technology programs remain over budget or behind schedule. The right-hand chart shows that more than 80% of organizations that deployed AI reported no tangible enterprise impact.
End of image description.
The path forward
The frustrations of navigating government bureaucracy have become a cliché for a reason. In our experience, public agency leaders are often equally frustrated by the challenges they face in transforming the business of government: legacy systems, the sheer scale of change required, the unique responsibilities of public service, and the risks inherent in any misstep.
AI presents a potential inflection point. But rewiring government is not just about modernizing systems or bolting on AI tools to existing processes. This moment is an opportunity for public agencies to truly transform, becoming more agile and efficient while better meeting people’s needs. Exhibit 7 shows how residents may experience a truly AI-native government.
Image description:
A table shows a series of government processes by level of AI integration. The table identifies five levels of integration and provides a description and example for each.
At the low end of AI integration are exclusively human processes such as policy approval. The second-lowest level contains processes that are led and performed by humans while AI is deployed on key selected tasks. An example is case review. The third level is where AI executes and humans decide. For example, AI might conduct a permit inspection and prepare a report for human approval.
The fourth level is AI-led processes with human oversight. This might include initial screening of license applications. Finally, the fifth level comprises processes that are led by AI end to end. These include case intake and triage.
End of image description.
We’re not suggesting it’s easy. Meeting the demands of this modern era is complex and challenging, but the rewards are clear. Governments can create agile, transparent, and resident-focused systems by starting with the mission, reimagining workflows from the ground up, building organizational capabilities to deliver advanced technologies, and empowering both people and technology (for a tactical leader’s to-do list, see below, “Ninety-day actions: Immediate steps for government leaders”). In the process, they can move beyond incremental improvements to achieve significant improvements in outcomes.
Ninety-day actions: Immediate steps for government leaders
The first step is often the hardest. Here are recommended actions public sector leaders may take to kick-start the transformation process and gain immediate momentum for change.
- Publish the resident-facing outcomes you will measurably improve with AI by year-end. This includes wait and response times as well as accuracy and reduced fraud—not the number of pilots launched or executive orders signed.
- Map two end-to-end workflows from a resident or frontline perspective. Look for steps that exist only because of legacy constraints, not because of mission need or legislative requirement.
- Stand up one cross-functional pod. This pod should combine policy, technology, operations, and frontline staff to own a single end-to-end process. Give this group of five to eight people real decision rights.
- Review your AI procurement pipeline for outcomes-based clauses. If every contract still pays for inputs (hours, licenses, or seats), you cannot share risk with vendors or capture value.
- Preposition your data governance for agents, not just analytics. Identify the three data flows that an agent would need to act on and resolve the access, lineage, and audit questions now before the agent becomes the bottleneck.
- Commit to a change management investment ratio. For every dollar approved for technology, secure a parallel commitment for adoption, training, and capability building.
ABOUT THE AUTHOR(S)
Hrishika Vuppala and Tim Ward are senior partners in McKinsey’s Southern California office, Tim Fountaine is a senior partner in the Sydney office, and Tony D’Emidio is a partner in the Washington, DC, office.
The authors wish to thank Ali Ustun, Anne Neville-Bonilla, Deidre Harrison, and Kelly Ungerman for their contributions to this article.