Most AI programmes at incumbent banks end up in one of two places. A vendor-led transformation that produces a great deal of activity and no measurable change to the bottom line. Or a chatbot pilot that ships into production, attracts a press release, and gets retired quietly twelve months later when the eval framework that was never properly built fails to catch the regression.
There is a third path. Very few firms are on it.
The third path is what “AI-native” actually means, stripped of the marketing. It is what some banks have, in pockets and without naming it as such, already started to do. The next material competitive move for incumbent banks is to recognise that work for what it is, and accelerate it deliberately. This article argues for the discipline, the practice, and the conditions that make it work.
The technology is not the work. The operating layer is the work.
The market has been moving fast in a direction that, until very recently, was easy to miss. By May 2026 the OpenAI Deployment Company existed as a $10B vehicle anchored by TPG and Brookfield with Bain, McKinsey and Capgemini on the cap table, and a 17.5% guaranteed return for principal investors. Anthropic shipped a similar shape a month earlier with Goldman Sachs, Blackstone and Hellman & Friedman, $1.5B aimed at deploying its model into PE portfolio companies. Google and Accenture announced joint forward-deployed engineering pairings on customer sites in April.
The vehicles are not coincidences. They are the structural admission, in the form most credible to capital markets, that the value created by AI in an enterprise is not the model. It is the operating layer the model works inside. The bank’s policies, workflows, decision rules, customer records, and auditable controls, all rendered legible enough for an agent to act on. The labs have set up institutional vehicles to industrialise the work of producing that operating layer because they cannot do it at scale themselves, and because the consultancies, left to it alone, deliver activity rather than outcome.
The frontier labs are right about where the value is. The consultancies are right that the operating layer work is hard. The market structure tells us both.
It also tells us something else. The window for incumbents to do this work themselves is short. Two years, perhaps three. After that, the gap between firms that did this well and firms that bought AI features will compound for a decade.
Goldratt told us forty years ago how this ends.
The goal of a business is the bottom line. Improving a non-constraint is wasted effort, and the technology of the moment is never the goal.
AI is the moment, not the goal. The goal here is the same as it has always been: return on capital, cost-to-income ratio, decision velocity, regulator response time, complaints aged, claims time-to-decision, the production rate of good lending decisions, customer NPS, share of revenue from products launched in recent years, growth and retention of customer relationships. The work pays back when those numbers move. It does not pay back when the bank announces an AI initiative.
The trap that catches most companies is treating AI as the goal. They announce AI initiatives, appoint AI leads, and bolt AI onto existing workflows. The investment does not pay back in business terms because they never named what they were trying to lift.
AI agents can lift constraints no prior wave could touch. On the cost side: decision capacity in small expert teams, the cost of an auditable decision, the latency between customer event and response, the cost-to-serve of regulated high-touch operations. On the revenue side: the speed of bringing new products to market, and the depth of engagement we can give every customer rather than only the ones a salesperson reaches. Workflow engines, data warehouses, and analytics did not reach these constraints. AI agents can, provided the operating layer is in place.
And here is the focus of this article.
AI is a flywheel. It is both the destination and the mechanism that gets you there.
The same large language models that need clean, structured inputs to operate are the most capable tools available for producing them from messy ones. They can interview the senior people who carry the rules in their heads and produce documented policy. They will read fifteen years of SharePoint and extract structure. They will propose a business ontology from existing artefacts and let humans correct rather than invent. And they can document decisions as they happen.
The operating layer work that has failed inside banks for thirty years (codified policies, ontologies, decision rules, structured customer records, auditable workflows) is the work AI itself can now do faster than any prior approach. Each turn of AI-assisted operating layer work makes the next agent deployment more capable. The deployment in turn produces more of the operating layer as a by-product of how it operates. Each gain makes the next round of AI work cheaper and more accurate.
That is what makes the moment different from every prior digital transformation wave: the means are now sufficient for the work.
What an agent actually needs.
An AI-native bank is a bank an agent can operate inside. The structural properties this requires are old in enterprise architecture and operations theory. The consumer that finally cares about them is new.
An AI-native bank has seven properties an agent can rely on:
- A single source of truth at the level of business facts: customer, product, exposure, policy, control.
- Capability and decision rights documentation that is consumed, not filed.
- Workflows where the decisions are visible and the exceptions are bounded.
- Evaluation loops that gate production.
- Access and audit at agent grain, so every action by an agent leaves a defensible trail.
- A formal ontology of the business and its rules that exists somewhere other than senior people’s heads.
- An internal capability layer (Skills, MCP servers, sub-agents) that agents can call, exposing the business as a structured surface rather than a tribal one.
None of this is novel as architectural ambition. What is novel is that there is now an external actor that depends on every one of these properties to do useful work. For thirty years, enterprise architecture has not delivered the impact it promised or the business desired, because the artefacts get filed rather than consumed. Agents are the consumer that finally cares.
Where most banks actually are.
Most banks, like most incumbents in financial services, are closer to pen-and-paper underneath the digital veneer than the language of digital transformation admits. Data sits scattered across legacy systems, semi-transformed into a data platform that is still incomplete. Useful knowledge is buried in years of SharePoint and shared drives, and standard operating procedures, though documented, are hard to surface at the point of need and ambiguous enough to leave room for interpretation. Manual processes glue these together with humans acting as integration layers. The bank has no formal ontology of what it does or by what rules, except in pockets.
Those pockets matter, and we should be specific about them.
Some parts of any incumbent bank have been pushed by regulation toward more structure than the rest. Treasury and market risk run more on structured data and codified policy. Credit decisioning runs partly on legible rules and audited models, wholesale trading systems carry audit at every step, and model risk management has built one of the more disciplined legibility practices in financial services. None of this is uniform across the pockets, the discipline is patchier in practice than its reputation suggests, and none of it is ready for agents in its current form. The policy lives in regulatory submissions, the structured data in modelling environments, the audited rules in trading systems, each surface built for a human reader, a regulator, or a downstream machine that is not an agent. These are candidate areas, not a default starting point. The case for any one pocket has to be made on the evidence inside it, not on the reputation of the discipline.
The rest of the bank is the harder problem, and it is where most of the value sits. Creating products customers want to buy, engaging the customers you have, serving them well enough that they stay and bring more business, designing distribution and pricing well enough that the margin holds: none of this runs on codified rules. The judgment lives in people and the relationships in conversations; the records sit across CRMs, shared drives, and the slides from last quarter’s strategy off-site. Operational risk, complaints, customer vulnerability handling, change management, and internal knowledge management share the same characteristic: the policy lives in PDFs and the judgment lives in people. McKinsey’s 2025 State of AI survey found that only 5.5% of firms see more than 5% of EBIT attributable to AI; BCG’s 10-20-70 rule says 70% of the work in any AI programme aiming for impact is process and people. The numbers triangulate. The operating layer work is the bottleneck, and always has been.
Some banks have base ingredients in place, not a ready meal. A target operating model. Federated capability ownership of some kind. Governance frameworks for data products that sometimes add more friction than value. The beginnings of a business ontology built to control grade. Operational disciplines in the platform team that, where they hold, can deliver. The platform is rarely complete and the picture is rarely uniform, but the components exist to be assembled.
The data strategy a bank has been running for years is not a different programme from the AI-native strategy. It has been producing the ingredients the AI-native strategy was always going to need. Naming that connection clearly, and continuing the work and extending to unstructured data, is the next stage.
The work that pays back.
Forward-deployed engineering is the discipline frontier labs, the new AI-native services firms, and ultimately Palantir twenty years ago all converged on. It decomposes into four phases: diagnostic, codification, build-and-bound, and hand-off.
Codification is where the value sits and where internal teams stall. It is the phase the in-house technologists running this work are best placed to own, and it is what this article is principally about.
A few specific properties of the discipline that matters:
The deliverables get named. Anthropic’s forward-deployed engineer job spec lists the artefacts shipped: MCP servers, sub-agents, agent Skills, eval suites. The reason is that “we will help you with AI” is a vague promise, while “we will ship you these named artefacts in this order” is a deliverable. Every bank doing this work needs its own version of this list. These are the deliverables that matter.
Evaluation is the gate. “No eval, no ship” is now consensus across every credible AI deployment vendor and every research lab. The reason is the Klarna pattern. Klarna replaced 700 customer service agents with an AI system in February 2024, claimed $40M of profit improvement, and quietly reversed course a year later. Their CEO admitted that “cost unfortunately seems to have been a too predominant evaluation factor… what you end up having is lower quality.” A system that replaces humans on cost grounds without a defended regression metric will be publicly walked back. Most banks already have a model risk function that defines the shape of this discipline in the credit risk and market risk domains. The work is to extend it to all agent deployments.
Outcomes are the unit of measurement. The new AI-native services firms price by outcome: per resolved customer interaction, per conversation, or per seat with a high floor that reflects the operating layer value their products carry. This is the discipline an in-house technology community applies to itself and to any vendor on the other side of the table. Every workflow brought onto the operating layer is tied to a named bottom-line metric before it ships. A vendor that cannot price an engagement against a named bottom-line outcome is selling classical consulting.
Hand-off is a design decision, not an afterthought. Most public AI-native playbooks cover hand-off thinly because their commercial model presumes the vendor stays in the loop. If a bank wants to become AI-native, it has to be able to run this work itself. A capability the bank cannot operate without an external partner is not a capability the bank has. Operationalising and owning the capability requires answers to questions like: who operates the eval, who owns the dashboard, who is the named accountable officer per the model risk framework, who has the override authority. The design surfaces these at design time so the answers can be agreed before launch, not deferred to a run-book written after.
And then the two steps the public playbooks miss entirely.
Step zero: find the more structured pockets the bank already has, and start there. Every regulated bank has them. Starting in greenfield where the operating layer is absent makes a transformation programme look slow and expensive; starting where regulation has already paid for the foundations makes a programme look fast and credible.
Step six: regulatory and risk binding. Model risk management, third-party risk, data residency, audit trail completeness, override authority, the named accountable officer per the NIST AI RMF and ISO 42001. Most banks are at the same place as peers on this discipline, sometimes ahead and sometimes behind, and the work is non-negotiable in either case.
Codification does not compile, but it does compound. The first deployment is expensive. The second is cheaper. The fifth pays for the first four and changes the bank’s relationship with its own operating model.
The translation problem.
The most difficult and central piece of work in this strategy is not the strategy itself. It is the work of socialising it with the executive committee and stakeholders, so they internalise it deeply enough to carry it in their own voice. To investors, across the bank, to the regulator, and to the board. That has to start with the technologists, the data and AI practitioners, the people closest to the work, internalising it first. Their internalisation is a means to that end, not the end. The strategy is irrelevant if it lives only on the page that contains it.
Translation is the part of the work that tends to be done last and worst. It also tends to be the part that is delegated to communications functions or to design teams, which is the move that ensures the translation fails. The translation is the strategy, not a deliverable of it.
Executive stakeholders are smart, literate, time-poor, and emotionally exposed to AI hype. They are fluent in business language. They are not fluent in the vocabulary of capabilities agents can call, evaluation discipline, structured operating layers, codified policy artefacts, ontologies, or hand-off design. If the strategy is carried forward in that vocabulary, it will be translated by the listener into something resembling the vendor pitch they already heard last week. The specific argument will collapse into the noise.
What follows is the translation approach proposed. It is credible and possibly viable, but it needs testing. Lead with the bottom-line metric the strategy is moving and defer the mechanism. Never use the word “ontology” unless the audience asked for it; say “the bank’s record of what it does, how, and by what rules” instead. Use the bank’s existing data strategy programme as the bridge from familiar ground to new ground, because the data strategy is already an accepted investment and a known quantity. Quote the McKinsey finding that only 5.5% of firms see EBIT impact from AI, rather than describing what codification is, because the failure rate of the industry is more persuasive than the discipline that prevents it. Use the Klarna reversal to illustrate the eval point because a memorable failure carries the lesson better than the methodology does.
Vendor immunity is one of the outputs of the translation work. The pitches mostly land in front of business stakeholders, not the technology community; the job of the practitioners is to ground-check what the business has been told. Business peers need the same immunity. The sales pitch is consistent. “AI will transform the business within twelve to twenty-four months without the operating model itself being touched.” The risk is that it directs the spend at the wrong layer of the stack. Vendors selling AI features hope the bank does not do the operating layer work, because the operating layer work is what creates leverage, and leverage is what reduces dependence on the vendor. The discipline applied when the next vendor lands a meeting is one question: what business metric will you commit to move, by when, in what currency, what is the rebate if it does not move, and what is the hand-off that makes us independent? Vendors who can answer that question are the ones worth the time. The ones who cannot are selling activity.
Translation is the work that will take most effort on the strategy side. The technical diagrams, the deployment plans, and the eval frameworks are the easier parts. The message has landed, and the bank is pulling in the same direction.
What the work looks like.
The shape of a credible AI-native programme is straightforward to describe and not trivial to execute.
An in-house community takes the lead on drafting the bank’s AI-native strategy and the supporting roadmap, with strong executive sponsorship. The strategy paper goes to senior leadership, written from the start in language executives can carry forward. Translation work is treated as the principal deliverable, not as a polishing step at the end. The roadmap is concrete to twelve months, indicative to twenty-four.
Inside that frame, the work that pays back is the operating layer work. Two things get named: the deliverable taxonomy (what the practitioners ship into the rest of the bank when they ship AI-native architecture) and the maturity baseline (where the bank actually starts from, in language the business uses rather than vendor terms). From the more structured pockets the bank already has, two or three candidate workflows get identified, where an AI-native deployment can ship into production quickly against a measurable bottom-line outcome. The evaluation discipline that gates production gets stood up. So does the framework by which any AI vendor or MSP pitch is evaluated against a standard. The first deployment ships as an operating layer build with an agent on top, not as a vendor pilot. What that first deployment teaches becomes the reference pattern for the second and third.
What an AI-native programme needs to succeed.
Three conditions, in roughly decreasing order of importance.
A shared view of the strategy at the top of the technology and data leadership, written in language the executive committee can carry forward. Without it, the executive conversation drifts to vendor framing.
Named executive ownership of AI governance at executive committee level. The strongest correlate with AI delivering EBIT impact at the bank level is named executive ownership, not a steering committee.
Agreement that one named in-house community owns both codification and translation, with regulatory binding as a partnership rather than a hand-off. Codification is where the value sits and where teams stall. Translation is where the strategy thrives or dies in the executive layer. Both are too important to outsource.
Closing.
Many banks have spent years attempting to build substantial parts of the foundation that AI-native transformation requires. The opportunity is to recognise this work for what it is, name the data strategy that has been running for years for what it has always been, and accelerate the foundation so agents can call it. The disciplines this work needs are largely the disciplines incumbent banks already know. What is new is that AI itself now lets us do this work faster than at any prior point, and that, in turn, is the basis on which the rest of the strategy turns.