← Return to the book

FUTURECENTRAL PRESS · BOOK SAMPLE

AI in Financial Services

From Machine Learning to Generative and Agentic AI

Chapter 22: Designing, Sourcing, and Adopting Agentic Systems

From Opportunity Prioritization to Operating Model and Vendor Strategy

Learning Outcomes

By the end of this chapter, the reader will be able to:

Identify the agentable work in a financial institution and prioritize it, and explain why no published, testable method for doing so exists despite the scale of the portfolios being committed.

Apply the experimental evidence on knowledge-worker productivity to role design, including the jagged frontier finding, the skill-leveling effect measured in customer support, and the conditions under which human oversight reduces accuracy.

Distinguish automation bias from selective adherence, and explain why the second is the more relevant control problem in credit and collections.

Evaluate adoption evidence in banking, separating licenses issued from regular use, and account for the absence of any disclosed change-management budget across the sector.

Assess build, buy, and partner decisions against the regulators’ own findings on vendor concentration, including the shift from single-provider dependence to a three-provider oligopoly and the absence of any designated artificial-intelligence provider in either the European or United Kingdom critical third-party regimes.

Apply the Agentic Opportunity Prioritization Matrix across its four axes and the Human-Agent-System Swimlane Method across its five steps to any candidate deployment.

Map the sourcing perimeter: provider and deployer obligations under the European artificial-intelligence regime and the conduct by which a bank becomes a provider, the oversight regime for critical technology providers, and the Reserve Bank of India’s outsourcing directions as they reach model suppliers.

Opening Vignette: The Answers That Sounded Better When They Were Wrong

In a randomized controlled trial run with 758 consultants at Boston Consulting Group, and published in Organization Science in 2026, researchers gave some participants access to a frontier language model and set them tasks of two kinds. On tasks inside the model’s capability, the results were what the enthusiasm predicted. Quality rose by roughly 34 percent for consultants who had the model and some prompt training, and by 30 percent for those who had the model alone. Completion rates rose. Time to finish fell by more than a fifth.

The second kind of task had been constructed to sit just outside the model’s capability. There the direction reversed. Consultants working without the model answered correctly 84.5 percent of the time. Those working with it answered correctly between 60 and 71 percent of the time, a fall of between 14 and 24.5 percentage points.

The people with the better tool produced worse work, and they produced it faster, which is a combination that should trouble anyone who has approved an agentic business case on the strength of a cycle-time improvement without asking what happened to the error rate underneath it.

The finding that should hold a designer’s attention is neither of those. Independent raters scored the answers for coherence on a ten-point scale. On the tasks outside the frontier, the answers produced with the model scored about one to one and a half points higher on that scale than the answers produced without it. Narrow the comparison to the answers that were wrong and the gap widens, running from one and a half to almost two points. The wrong answers were the better-written ones. They read as more structured, more confident, and more complete, and they were wrong.

That is the mechanism by which agentic output defeats review. A reviewer looking at agent work is not comparing a good answer to a bad one that announces itself. The reviewer is comparing two answers, one of which is more coherent and less correct, under time pressure, with a queue behind it. Every design in this chapter has to survive that fact, and most designs in circulation have not been tested against it.

The trial also found that the consultants in the bottom half of the prior performance distribution gained proportionally more than those in the top half. That result recurs across the literature and it has a direct consequence for how roles should be drawn. An institution deploying agents is not making everyone somewhat better. It is compressing the distribution, which changes what its most experienced people are for.1

This chapter examines how to find the agentable work and rank it, how to draw the boundary between human and agent inside a process, what the adoption evidence in banking actually shows, and how to source the capability in a market three providers dominate.

1. Finding the agentable work: prioritization without a published method

Every institution in this book is running a selection problem. It has more candidate use cases than it can fund, no reliable way to estimate the return on any of them, and a portfolio decision to make anyway.

The published evidence on how that decision is being made, and on how it is turning out, is unflattering and worth examining before any framework is offered.

Announcements and evidence. Agentic deployment is now a substantial share of what banks announce, and a small share of what they can show for it. The Evident AI Index found that nearly one in three artificial-intelligence use cases newly reported by the banks it tracks in the first quarter of 2026 were agentic, up from 15 percent in the preceding quarter, with specialist vendors beyond the major cloud providers accounting for 68 percent of deployments. Against that, only fifteen of the fifty banks in the index disclose an aggregate return on investment at all, and 38 percent of the quarter’s use cases reported any outcome. The tracker counts public announcements, so it understates unannounced work and cannot separate a pilot from a production system.

The survey picture. The cross-industry survey evidence shows adoption rising and attributed financial impact flat. McKinsey’s survey of 1,719 respondents across ninety-seven countries, published in August 2026, found 89 percent of organizations using artificial intelligence in at least one function and 44 percent scaling it enterprise-wide, up from 38 percent a year earlier. The proportion attributing any impact on earnings before interest and taxes was 37 percent, essentially unchanged year on year, and 6 percent qualified as high performers. Deloitte’s survey of 1,854 senior executives across Europe and the Middle East, conducted in August and September 2025 among organizations that already had artificial intelligence operational, found 57 percent using agentic systems and 10 percent reporting significant measurable return. 6 percent saw payback inside a year, against a typical seven to twelve months for conventional technology. Agentic systems are paying back more slowly.

The one variable. One variable separates the organizations reporting returns from those that are not, and it has nothing to do with the technology. In the same McKinsey survey, nearly three-quarters of high performers had redesigned workflows end to end, against a quarter of everyone else. That is a gap of nearly three times, and it is the best available answer to the question of what distinguishes use cases that scale. It is also survey correlation, and it cannot establish that redesign caused the performance. The honest reading is that redesigning a workflow and getting a return from artificial intelligence travel together, and that institutions buying a tool to accelerate an unchanged process are in the group that reports nothing. The redesign is what moves the result.2

Reading the Gartner number. Gartner’s prediction that more than 40 percent of agentic projects will be canceled by the end of 2027 is widely quoted. Its other finding is more useful to a buyer. The same analysis estimated that of the thousands of vendors describing their products as agentic, roughly 130 offered genuine agentic capability.

That practice has a name in the industry, which is agent washing, and it puts a specific burden on the sourcing process described later in this chapter: the first question about a candidate vendor is whether the product does what the category name implies, and that question has to be answered technically rather than commercially.

One bank’s criteria. One bank has disclosed its selection criteria, and they are prescriptions from experience. Bank of America’s chief technology and information officer described in April 2026 a prioritization approach resting on four commitments: end-to-end process transformation in preference to task-level pilots, scale and reuse across roughly 3,000 processes, governance paired with a clear view of return before initiation, and a data foundation carrying the bank’s compute strategy and model economics. The bank runs an annual technology budget of $13.5 billion, of which about 30 percent goes to new initiatives including artificial intelligence. The same executive named the tension the criteria are meant to manage, observing that the technology is very hard to govern, that overdoing governance stalls innovation, and that underdoing it introduces a great deal of risk.3

What does not exist is a method. No peer-reviewed or methodologically documented framework for selecting and prioritizing artificial-intelligence use cases in banks could be located. The influential ten, twenty, seventy allocation of effort across algorithms, technology and data, and people and process is a consultancy prescription and has been widely miscited as a measured finding. Gartner publishes a proprietary assessment whose method is not public. No credible published figure exists for the cost of an agentic pilot at a bank, with every number in circulation traceable to a firm selling implementation services.

Institutions are therefore committing multi-billion-dollar technology portfolios using criteria that have never been tested against outcomes, which is the condition the framework later in this chapter is offered into and the reason it is offered with limits attached.

2. Designing workflows and roles without code

The design question inside a process is not which tasks a model can perform. It is which tasks a model can perform where a human can still tell whether it did. Those are different sets, and the gap between them is where most agentic design fails.

Jaggedness, and what follows. The jagged frontier is jagged, which means capability cannot be inferred from apparent similarity between tasks. The consulting trial described at the opening of this chapter placed two task types either side of a boundary that was invisible to the participants. Inside it, performance improved substantially. Outside it, performance fell below the unaided baseline. The participants could not see which side of the line a given task sat on, and neither, in most institutions, can the person specifying an agentic workflow. The practical consequence is that capability must be established empirically for each task class in the actual process, on the institution’s own data, and re-established when the model version changes.

The leveling effect. The skill-leveling effect is the best-evidenced result in this literature and it should drive how roles are redrawn. The phrase is this book’s own coinage; the authors describe their result as substantial heterogeneity in the gains across workers, cutting against the skill-biased pattern of earlier computing waves. A study of 5,172 customer support agents, published in the Quarterly Journal of Economics in 2025, found that access to a generative assistant raised issues resolved per hour by 15 percent on average. The average conceals the finding. Novice and lower-skilled workers gained about 30 percent, while the most experienced and highest-skilled saw minimal gains and a small decline in conversation quality. Customer sentiment improved and attrition fell, particularly among newer staff.

The mechanism the authors identify is that the system captures and disseminates the behaviors of the most productive agents, converting tacit expertise into a shared asset.4

What follows for the experts. That mechanism has an uncomfortable implication for the people whose expertise was captured. If the system distributes what the best performers know, the marginal value of being a best performer at the task falls, and the marginal value of being able to specify, review, and improve the system rises. Role design should follow. The senior analyst in an agentic process is a designer of specifications, an adjudicator of the cases the agent escalates, and the owner of the taxonomy the agent works against. An institution that deploys agents while leaving senior roles defined by throughput has created a job whose measured contribution is about to decline.

Oversight as a control. Human oversight is the most commonly proposed control in agentic design, and the experimental evidence on it is genuinely unwelcome. In a study published in PLOS ONE in 2024, adding a human to the loop raised willingness to use algorithmic recommendations from 66 to 73 percent while worsening accuracy, with mean absolute deviation rising from 17.4 to 18 percentiles. Participants adjusted the recommendation 63 percent of the time. The damaging detail sits in the distribution of those adjustments: participants corrected large errors less often than small ones, and by smaller margins. Oversight was weakest exactly where it mattered most. The sample was small and the subjects were not professionals, so the result should not be over-read.

It is nonetheless direct evidence that a human in the loop can be simultaneously legitimating and value-destroying, and that adding one is a design decision requiring justification.

Automation bias, revisited. The received account of automation bias is also more complicated than the design literature assumes. Three survey experiments spanning more than 2,800 participants, including over 1,300 civil servants, found no general tendency to defer to algorithmic advice. Across the two general-population studies pooled, 11.1 percent followed a low algorithmic score against 10.5 percent for the same advice from a human expert, a difference well inside noise. The civil servant study replicated that result. What the same experiments did find was selective adherence. Advice was followed substantially more readily when it aligned with an ethnic stereotype about the person being assessed. The effect was not uniform across the three experiments, and the authors present the mechanism rather than a stable effect size as the finding, so the direction matters here more than the magnitude. The setting was Dutch public administration, so transfer to banking must be argued. The control problem it describes is not indiscriminate deference to machines. It is the use of machine advice to ratify a conclusion the reviewer already held, and in credit assessment and collections that is the more dangerous failure by some distance.5

Rule against evidence. The regulatory requirement and the empirical evidence pull against each other, and a designer has to hold both. The European artificial-intelligence regime requires deployers of high-risk systems to assign human oversight to natural persons who have the necessary competence, training, and authority, together with the support to exercise it. It requires logs to be retained for at least six months, and workers’ representatives and the affected workers to be informed before a high-risk system is put into use in the workplace. The obligation is to provide competent oversight. The evidence is that competent oversight sometimes reduces accuracy and reliably increases the confidence placed in the output. Satisfying the rule is therefore the floor.

Designing oversight that actually catches the errors that matter is a separate piece of work, and it begins with knowing which errors the reviewer can detect at all.

3. Experience design and adoption

Adoption is the variable most likely to determine whether an agentic program produces anything, and the least likely to be measured honestly. The sector reports licenses issued far more often than it reports use, and it reports use far more often than it reports what the use displaced.

Co-signed figures. The most frequently cited adoption figures in banking come from releases the vendor co-signed, and should be read accordingly. BBVA announced in December 2025 that it was deploying a commercial assistant to more than 120,000 employees, having run an earlier phase covering 11,000. The usage figures belong to that earlier phase: more than 80 percent of those users accessed the assistant daily, and they saved an average of nearly 3 hours a week on routine tasks. The distinction matters. Read as a property of the full rollout, which is how the figures are almost always quoted, they describe a population more than ten times the one they were measured on. The release was issued jointly with the model provider, both parties benefited from a high number, and no measurement method was published for the 3 hours.

Signals that cost something. The more credible adoption signals are the ones that cost the employee something. NatWest disclosed in February 2026 that all of its roughly 60,000 colleagues had access to artificial-intelligence tools, and that more than 12,000 coders were using them. Around 35 percent of the bank’s code was being written with them, automated call summaries had saved more than 70,000 hours, and relationship managers had gained 30 percent more time for customer conversations. The figure worth isolating is that more than half of employees had opted in to additional training. An opt-in is a revealed preference. A license count is not.6

Two adoption funnels. Two banks show what the adoption funnel actually looks like when an institution reports its own intermediate stages. Bank of America disclosed that 200,000 of its employees were using artificial-intelligence capabilities, generating over 400,000 prompts a day, with more than 90 percent adoption of its internal employee assistant. Within that, the bank reported more than 300 approved use cases, of which 114 were live generative implementations and thirty-four were fully implemented. Approval to full implementation is roughly one in nine.

Citi required 175,000 employees across eighty locations to complete prompt-engineering training, adapted to take under 10 minutes for an expert and about 30 for a beginner, following more than 6.5 million prompts submitted through internal tools in the nine months to October 2025.

More conservative reporting. The Asian disclosures are more conservative and probably more representative. Axis Bank reported that around 40 percent of its nearly 100,000 employees regularly used its internal artificial-intelligence knowledge platform, and that 30 to 40 percent of coding was automated. DBS reported its generative assistant in use by two-thirds of employees, and a coding assistant reducing data scientists’ coding time by up to 20 percent on certain coding tasks. The hedging in that last phrase is instructive about how a bank discloses a benefit it has measured carefully. 40 percent regular use against BBVA’s 80 percent daily use is the range within which an institution should set its own expectations, and the lower figure is the more candid measure because regular use is a harder test than daily access.

Anxiety among staff. A peer-reviewed study of artificial-intelligence anxiety among bank employees reports a pattern that complicates the usual change-management prescription. A survey of 858 employees of a major Korean commercial bank found that job anxiety induced by artificial intelligence was positively associated with perceived effectiveness of the bank’s initiatives, and had no significant relationship with intrinsic motivation, with knowledge acquisition about the technology, or with domain knowledge. The authors describe an appraisal and engagement gap: the employees who feel most threatened rate the technology as most effective and do not invest in learning it. The study is cross-sectional, single-country, single-bank, and framed around sustainability initiatives, so its transfer to a general adoption setting is limited. Within those limits it suggests that communicating capability persuasively may raise perceived threat without raising engagement. That is the opposite of what most adoption programs assume.7

No bank in this chapter disclosed what it spent on any of this. Citi, NatWest, BBVA, Bank of America, and Axis Bank all disclose activity. None discloses a training or change-management budget line. The widely repeated claim that change management should absorb 70 percent of transformation effort originates in a consultancy prescription about effort allocation, and no measured spend ratio could be located.

An institution building a business case therefore has no external benchmark for the largest and least visible cost in the program, and should treat any figure offered to it as unsourced until the source is produced.

4. Build, buy, or partner: the operating model and the concentration question

The sourcing decision in agentic artificial intelligence has an unusual property. For most enterprise technology, build and buy are genuine alternatives. Here the regulators have said plainly that one of them is foreclosed for almost everyone, and the market data shows the dependency moving from one provider to three.

Intention against outcome. The intention data says hybrid, and the outcome data does not exist. Asked how they would allocate artificial-intelligence investment, 38 percent of executives favored a combination of in-house and external work, 32 percent leaned toward vendor-built solutions, and 24 percent planned to build internally. No study measuring the outcomes of build-versus-buy decisions in bank artificial intelligence could be located. The available evidence establishes what institutions choose and never whether the choice paid off, which is a material gap in a decision this expensive.

Where the boundary sits. The boundary is moving toward build for software and staying with buy for models. McKinsey found in August 2026 that 32 percent of organizations had decided against purchasing software because they could build it themselves using artificial-intelligence coding tools, a direct effect of the coding-assistant adoption described in the preceding section and consistent with NatWest writing 35 percent of its code that way. That shift does not reach the model layer. The Financial Stability Board judged in November 2024 that the cost and complexity of training a large language model from scratch are generally prohibitive for firms that are not themselves technology specialists.

When a standard-setting body tells the sector that an option is unavailable, the sourcing question narrows to which external dependency an institution prefers and how it governs it.

The concentration findings. The regulators’ concentration findings are specific, and more balanced than the commentary around them. The Financial Stability Board found the markets for accelerated computing chips and cloud services dominated by a limited number of entities, with vertical integration across parts of the supply chain, and noted that concentration vulnerabilities intensify where many firms rely on a limited set of providers or on the same provider for several services. It observed that some level of concentration is difficult to avoid given increasing returns to scale. The Bank of England’s Financial Policy Committee warned in April 2025 that reliance on a small number of providers could generate systemic risks where rapid migration to alternatives is not feasible, and separately that widespread use of a small number of models or convergence on very similar model designs could drive correlated positioning across the market. The Committee also noted the offsetting force. Widely available open-source models could reduce concentration even as cost and vertical integration increase it.8

Those are two different risks and they should not be merged. Operational concentration is the risk that a provider fails and many institutions cannot switch. Model monoculture is the risk that many institutions make similar decisions because they are running similar models on similar data, which produces correlated behavior without any provider failing at all. The first is addressed by exit planning and multisourcing. The second is not addressed by either, because two institutions that each hold contracts with three providers may still be running the same model on the same benchmark-tuned prompts.

Agentic Opportunity Prioritization Matrix. A four-axis method for ranking candidate agentic opportunities, scoring value at stake, structural feasibility, consequence of error, and strategic fit and reuse, and plotting the first two against each other to sort a portfolio into quick wins, moonshots, traps, and discards.

Human-Agent-System Swimlane Method. A five-step method for designing an agentic process without code, producing a diagram with four lanes, human, agent, system, and fallback, in which each stage is assigned by the Role Allocation Test and every boundary crossing carries a specified handshake.

What the market data says. The market data cuts against the assumption that concentration is increasing. Evident’s tracking of fifty major banks found the leading model provider’s share of deployments falling from around half to a third over the eighteen months to January 2026, with two other providers gaining and specialist vendors accounting for most deployments. Several leading banks are deliberately model-agnostic so they can switch as capabilities evolve.

The accurate description of the market is therefore a shift from single-provider dependence to a three-provider oligopoly, which reduces the risk that any one failure is catastrophic and does nothing about the monoculture risk.

An unapplied architecture. The regulatory architecture for this dependency exists and has not yet been applied to any artificial-intelligence provider. The European resilience regime empowers the supervisory authorities to designate critical technology providers against the four criteria at Article 31(2): systemic impact of a large-scale operational failure, the systemic character or importance of the financial entities that rely on the provider, reliance in relation to critical or important functions whether direct or indirect through subcontracting, and the degree of substitutability. Where a designated provider does not comply with a required measure, Article 35 gives the Lead Overseer a periodic penalty payment of up to 1 percent of the provider’s average daily worldwide turnover in the preceding business year, imposed daily until compliance and for no longer than six months, and available only after at least thirty calendar days from notification of the measure. The first tranche of designations was made in November 2025 and contained no artificial-intelligence or foundation-model specialist. The major cloud providers are captured in their capacity as cloud providers, so model capability is supervised only incidentally. In the United Kingdom the position is starker. A parliamentary committee concluded in January 2026 that financial services firms are overly reliant on a small number of United States technology firms and recommended that major artificial-intelligence and cloud providers be designated as critical third parties by the end of 2026. In responses published in April 2026 the Treasury told the committee it was still gathering evidence and expected initial decisions that year, and the Bank of England confirmed on April 1 that none had been made. HM Treasury made the first designations by statutory instrument on July 8, 2026 and announced them on July 10. The Critical Third Parties (Designation) Regulations 2026 came into force on July 13, 2026 and designated four cloud and technology providers with effect from that date, with no artificial-intelligence or foundation-model specialist among them. They reach the cloud providers in that capacity, which leaves the gap the committee identified where the European regime leaves it: the supervision of model capability as such.

Rules for another dependency. The outsourcing rules that do apply were written for a different kind of dependency. European outsourcing guidance requires institutions to assess aggregate exposure to the same service provider and the potential cumulative impact, to notify sub-outsourcing of critical or important functions in advance with termination rights, to maintain documented exit plans, and to keep a register. It is technology-neutral and predates generative artificial intelligence. It bites on an artificial-intelligence vendor where the arrangement constitutes outsourcing of a critical or important function, and many tool deployments arguably do not meet that threshold even where the institution’s dependency on them is real.

The aggregate-exposure and exit-plan duties formally apply to a model dependency and were not designed for one, which is the analytically interesting gap in this area and the reason the Indian position examined later in this chapter is worth comparing.

One further obligation changes the sourcing calculus for any institution that customizes a model. The European regime separates provider obligations, which cover quality management, technical documentation, conformity assessment, registration, and postmarket monitoring, from deployer obligations, which cover use according to instructions, competent human oversight, input data, log retention, and worker notification. A deployer becomes subject to the heavier provider obligations in three circumstances. It puts its own name or trademark on a high-risk system already placed on the market. It makes a substantial modification to a high-risk system already on the market in a way that leaves the system high-risk. It changes the intended purpose of a system that was not classified as high-risk so that the system becomes high-risk. A bank that fine-tunes a vendor model and brands the result has to consider whether it has become a provider by conduct, and that question belongs in the sourcing decision, ahead of any compliance review. The timetable for these obligations moved in 2026, with high-risk obligations for stand-alone systems deferred to December 2027 and for systems embedded in regulated products to August 2028. Credit scoring and creditworthiness assessment sit within the high-risk category, so December 2027 is the operative date for the largest bank use case, and any plan built on the previously published August 2026 date is working to a deadline that no longer exists. That date has gone.9

5. Framework: The Agentic Opportunity Prioritization Matrix

Chapter 18’s Autonomy Spectrum classifies a deployment once chosen and Chapter 20’s Suitability Ladder classifies the work. This matrix chooses. It scores candidate opportunities across four axes, and it is built to resist the failure most visible in the survey evidence, which is collapsing risk into a single number. The second failure, funding a tool for a process nobody intends to redesign, sits outside the matrix for the reason given below, and is tested separately.

Axis One: Value at Stake. Value is decomposed into three factors and never estimated as a single figure: the value per unit of work, the volume of units, and the share of that value the institution can realistically capture. The third factor is where most business cases inflate. A process that consumes 40,000 hours a year does not release 40,000 hours of value, because the released time is distributed in fragments across many people, and fragments convert to value only where a role is redesigned or a queue clears. Scoring the capture rate explicitly forces the question of what happens to the time, and it is the axis on which a business case should be rejected most often.

Axis Two: Structural Feasibility. Four conditions are scored separately: whether the data the agent needs exists and is of adequate quality, whether the process is describable well enough to specify, whether the integration surface is reachable, and whether the work passes the reversal test set out in Chapter 20. An opportunity that scores well on three and fails on one fails outright. Feasibility is closer to a conjunction than to an average, and a matrix that averages the four will keep recommending the use case whose data does not exist.

Axis Three: Consequence of Error, scored as three separate quantities. Severity, reversibility, and detectability are not the same thing and collapsing them into a risk score destroys the information a designer needs. A high-severity error that is instantly detectable and cheaply reversible is a manageable proposition. A low-severity error that is undetectable and accumulates is the one that produces a remediation program three years later. Detectability deserves particular weight in agentic work for the reason the opening vignette established: agent output is more coherent than unaided work even when it is wrong, so the reviewer’s ability to spot an error is lower than the institution’s intuition suggests.

Axis Four: Strategic Fit and Reuse. This axis asks whether the opportunity builds a capability the institution will use again. An integration layer, a break taxonomy, an evidence-pack standard, or an escalation architecture built for one use case and reusable across ten is worth materially more than its own business case shows.

Bank of America’s stated criterion of scale and reuse across some 3,000 processes is this axis expressed as policy, and it is the axis that distinguishes a portfolio strategy from a collection of pilots.

Scoring and the four quadrants. Score each axis from one to five, then plot value at stake against structural feasibility, using consequence of error to set the governance tier and strategic fit as the tiebreaker within a quadrant. High value and high feasibility are the quick wins that should be sequenced first and used to build the reusable assets. High value and low feasibility are the moonshots, which belong in a separate multiyear track with the feasibility blockers named as their own projects. Low value and high feasibility are the traps: they are easy, they get built, and they consume the delivery capacity that the quick wins needed. Low value and low feasibility are discards, and a portfolio that cannot name what it discarded has not run a prioritization.

When the matrix misleads. The matrix scores the opportunities someone thought to list, and in most institutions the list comes from whoever attended the workshop. It also has no axis for the variable most strongly associated with realized returns in the survey evidence, which is willingness to redesign the workflow, because that is a property of the organization and not of the opportunity. An institution that scores a portfolio honestly and then funds the winners without redesigning anything will land in the 63 percent of organizations not reporting a positive earnings contribution from artificial intelligence. Score the opportunities, then ask separately whether the business units in question are prepared to change how the work is done, and treat a negative answer as disqualifying regardless of the score.

6. Framework: The Human-Agent-System Swimlane Method

Prioritization selects the process. This method designs it, without code, at a level of detail a business audience can argue about and a technology team can build from. It runs in five steps and produces a diagram with four lanes: human, agent, system, and fallback.

Step One: Stage Decomposition. Break the process into stages at the level where the output of one stage is a thing a person could inspect. Too coarse a decomposition hides the handoffs where failure occurs; too fine a decomposition produces a diagram nobody reads. The test of the right granularity is whether each stage has a nameable output and a nameable owner.

Step Two: the Role Allocation Test. For each stage, three questions decide the lane, and all three must be answered before the stage is assigned. Is this task inside the demonstrated capability frontier for this model on this institution’s data, established by measurement rather than by analogy to a similar task? If the agent performs it and gets it wrong, can the reviewer detect the error from what the reviewer will actually see? Does the reviewer have the authority and the time to act on a detected error? A stage that fails the second question belongs in the human lane whatever the model can do, because a review that cannot detect an error is a control in name only.

Step Three: Handshake Specification. For every boundary crossed between lanes, specify what passes, in what form, with what completeness guarantee, and what the receiving lane does when the handshake is incomplete. This is where most process designs are silent and most implementations fail. A handshake specification names the fields, the evidence attached, the confidence representation if any, and the explicit statement of what the sending lane did not check.

Step Four: the Fallback Lane. The fourth lane exists because the first three describe the process working. It carries what happens when the agent cannot proceed, when a system is unavailable, when the handshake fails validation, and when volume exceeds capacity. Each fallback names who absorbs the work and at what service level.

A process design without this lane is a design for the good case, and in a regulated process the bad case is the one carrying the statutory deadline, the customer who is already unhappy, and the volume spike that arrived on the day the upstream system was unavailable.

Step Five: Instrumentation and KPI Placement. Metrics are placed at the handshakes. End-to-end cycle time tells an institution that something has changed and never where. Measuring at each boundary yields the stage-level data needed to move the human-agent line later, which is the point of drawing the line explicitly in the first place. Place at minimum: volume and latency at each handshake, the rate at which each handshake fails validation, the rate at which the reviewer changes the agent’s proposal, and the rate at which changed proposals turn out to have been right.

When the method misleads. A swimlane assumes the process is knowable in advance and stable enough to draw, which holds for reconciliation and fails for an investigation that goes where the evidence leads. The deeper limitation sits in the human lane. Drawing a review box does not create review, and the experimental evidence in this chapter shows oversight failing precisely on the large errors it was placed there to catch. A diagram that shows a human reviewing every agent output satisfies a governance committee and tells that committee nothing about whether the errors that matter will be found. The honest use of this method is to place the reviewer where the second question in Step Two has been answered affirmatively, and to say plainly where it has not.

7. India Cases

India offers three things this chapter needs and the global record supplies less clearly: a bank that has described its operating model in structural detail, a vendor layer whose commercial substance can be independently verified in at least some places, and a national policy response to concentration risk that treats indigenous model capability as a prudential matter.

Axis Bank’s operating model, described at a level of detail few banks match. Axis Bank established an Enterprise AI Centre of Excellence in the second half of 2025 to set strategic direction, sitting alongside a Business Intelligence Unit organized into five divisions covering data engineering, data science, business analytics, reporting analytics, and data governance, and managing some three petabytes of data across three layers. The unit tests multiple large and small language models alongside banking-specific models, a deliberately vendor-agnostic posture that matches the model-agnostic approach several global banks have adopted. Adoption is reported candidly: around 40 percent of nearly 100,000 employees regularly use the internal knowledge platform, and 30 to 40 percent of coding is automated. Six strategic themes organize the portfolio, running from knowledge management and internal workflow automation through software development and conversational servicing to fully artificial-intelligence-driven customer propositions. The bank’s executives have projected artificial intelligence contributing roughly 1 percent of profit and loss in the 2027 financial year, scaling to about 10 percent by the 2029 financial year. That projection was given in a trade interview and not to investors, carries no stated basis, and should be read as management’s own model. Its shape is the teachable part. A tenfold increase in two years describes a late inflection, and a student should be able to say what would have to be true for that curve to materialize.

The Indian vendor layer, and the limits of what can be verified about it. Perfios, a software provider to the Indian financial sector, reported operating income of ₹669.5 crore (₹6.7 billion) in the 2025 financial year, growing 20 percent year on year with operating margins of 23.3 percent, up from 19 percent, and repeat business above 90 percent across a client list that includes State Bank of India, ICICI Bank, Axis Bank, and HDFC Bank. It was upgraded by a credit rating agency in February 2026. During 2025 it acquired three companies for roughly ₹580 crore, covering fraud risk management, artificial-intelligence-based debt collection, and healthcare insurance data. A domestic platform that buys artificial-intelligence capability it could have built is the Indian expression of the build-versus-buy question, and it is legible only because a rating agency publishes on the company. Sarvam AI raised $234 million in June 2026 in the first close of a $300 million round at a $1.5 billion valuation, led by an Indian technology services firm, and disclosed financial-sector deployments including a sales platform supporting a 350,000 person sales force and a voice campaign covering 45 million policyholders. Those deployments name no client and cannot be independently checked. For several other frequently named Indian vendors, no funding, client count, or deployment figure from a primary or independent source could be located at all.

The Indian agentic supply chain is materially less transparent than the banks that buy from it, which is an uncomfortable position for institutions whose regulator holds them accountable for their vendors.

Sovereign capability treated as a concentration-risk control. The IndiaAI Mission, with an outlay of ₹10,372 crore (₹103.72 billion), had onboarded more than 38,000 graphics processing units for its common compute facility, up from just over 34,000 in May 2025, supplied by fourteen empanelled providers at a subsidized average rate of about ₹65 per graphics processing unit hour. Its foundation-model program had selected its first four developers by May 2025 to build sovereign models ranging from 14 to 120 billion parameters, including multilingual and voice capabilities, and had selected twelve across two phases by the end of that year. The policy logic connecting this to banking is explicit in the Reserve Bank’s FREE-AI report, whose recommendations include indigenous financial-sector-specific models developed as public goods, financial-sector data infrastructure treated as digital public infrastructure, and capacity building inside both regulated entities and the supervisor. The report also records that concentration risk arises from reliance on few dominant providers. The position was stated more directly by the then Governor in October 2024, who said that heavy reliance on artificial intelligence can lead to concentration risks where a small number of technology providers dominate the market, that opacity makes algorithms difficult to audit, and that banks have to ride on the advantages of artificial intelligence and big technology and not allow the latter to ride on them.10

The shared lesson. The Indian outsourcing directions reach the model layer more explicitly than their European counterparts. They require a bank to assess concentration risk from multiple outsourcing arrangements to the same service provider, to create an inventory of outsourced IT services including the key entities involved in their suppliers’ supply chains, to maintain an exit strategy, and to obtain prior consent before a provider subcontracts. The supply-chain inventory requirement is more explicit about dependency beyond the first tier than the European guidance. The Reserve Bank reissued these obligations in consolidated form on November 28, 2025, and the consolidated text makes no reference to artificial intelligence, machine learning, or generative systems anywhere in it. Taken together the three cases describe a jurisdiction that has a clearer legal handle on vendor dependency than most, a policy program aimed at reducing that dependency at the national level, and a domestic supplier base whose disclosure is thin enough that the legal handle is hard to exercise in practice.

8. Global Cases

Three global cases cover the operating model, the leadership structures now forming around it, and the only credible published statement of what any of this is worth.

Citi’s Arc and the gated rollout. Citi launched an internal platform for building and scaling artificial-intelligence agents on April 30, 2026, describing agents that would be “monitored, auditable and governed,” and restricting initial access to its own developers building agents for well-defined internal use cases. More than 80 percent of the 180,000 colleagues with access to the bank’s artificial-intelligence tools were reported to use them regularly. The disclosure is notable for what it withholds: no agent count, no headcount for the platform team, no governance committee structure, and no named model providers. The teaching content is the rollout design itself. A platform gated to developers, restricted to internal use cases with defined boundaries, and instrumented for audit before it is opened more widely is a deliberate answer to the pilot-proliferation problem, and it stands in contrast to the enterprise-wide assistant deployments described earlier in this chapter.

Commonwealth Bank of Australia and the dual-role leadership pattern. The bank appointed a chief artificial intelligence officer from early 2026, recruited from a United Kingdom bank, and then in May 2026 appointed its first chief artificial intelligence scientist to lead a team of distinguished scientists. It names four technology partners across model providers and cloud, and operates technology hubs in San Francisco and Seattle. NatWest has adopted a similar structure with a chief artificial intelligence research officer alongside its technology leadership. The pattern separates delivery from research, which is a recognizable response to a market where capability changes faster than a delivery organization can absorb it. Neither bank has disclosed headcount, reporting lines, or budget for these functions, so the pattern is visible and its substance is not. The four named partners are themselves a concentration datapoint: a bank actively managing provider diversity still ends up naming the same small set that every other institution names.

Santander and the only credible number. At its investor day in February 2026 Santander stated a target of more than €1 billion of annual business value from data and artificial-intelligence initiatives by 2028, and said that this contributes approximately 1 percentage point to the improvement in its cost-to-income ratio, within an overall efficiency target of about 36 percent. That is the most useful disclosure in this chapter, and its usefulness lies in its precision. Read carefully, the percentage point is a contribution to the improvement in the ratio and not to its level. Against a cost-to-income ratio standing above 40 percent in 2025, that makes artificial intelligence responsible for something on the order of a fifth of the efficiency gain the bank is promising its shareholders, which is a material share and a long way from the transformational claims made on the technology’s behalf. The bank disclosed no headcount reduction plan, no artificial-intelligence spend figure, and no separate operating model, describing the technology as fully embedded in the businesses. DBS reported comparable ambition on a different measure, with more than 2,000 models and over 430 use cases delivering economic value of approximately 1 billion Singapore dollars in 2025, on a bank-defined construct with no published methodology.11

Against these, no major bank discloses artificial-intelligence spend as a separate audited line, so every investment figure in circulation is either a total technology budget or an executive’s characterization of a portion of one.

9. Regulatory Landscape

The general perimeter was set out in Chapters 1, 14, 18, and 23. This section covers what bears on designing, sourcing, and adopting: the division between provider and deployer, the oversight of critical technology dependencies, and the outsourcing regimes that reach model suppliers.

Global Layer

The European regime separates provider from deployer obligations, and a bank can move between them by its own conduct. Provider obligations cover quality management, technical documentation, conformity assessment, registration, and postmarket monitoring. Deployer obligations cover use in accordance with instructions, assignment of human oversight to “natural persons who have the necessary competence, training and authority,” representative input data where the deployer controls it, retention of logs for at least six months, and notification of workers’ representatives and the affected workers before a high-risk system enters workplace use. A deployer assumes provider obligations on three triggers, each keyed to high-risk status: branding a high-risk system already on the market, modifying one substantially so that it stays high-risk, and repurposing a system that was not high-risk so that it becomes one. The timetable moved in 2026. Regulation (EU) 2026/1744, in force from July 27, 2026, sets the application of high-risk obligations for stand-alone systems at December 2, 2027 and for systems embedded in regulated products at August 2, 2028. The transparency obligations were not deferred. They applied from August 2, 2026 to systems placed on the market from that date, with a four-month transitional period running to December 2, 2026 for content marking on systems already on the market before August 2, 2026. Credit scoring and creditworthiness assessment are within the high-risk category.

The critical third-party regimes exist on both sides of the Channel and have not yet reached any artificial-intelligence provider. The European resilience regulation empowers the supervisory authorities to designate critical technology providers on the four Article 31(2) criteria: systemic impact, the systemic character or importance of the financial entities that rely on the provider, reliance in relation to critical or important functions whether direct or indirect through subcontracting, and substitutability. Article 35 gives the Lead Overseer powers to request information, conduct general investigations and inspections, and issue recommendations. It backs them with a periodic penalty payment of up to 1 percent of average daily worldwide turnover in the preceding business year, imposed daily until compliance and capped at six months. The payment is periodic. The first designations were made in November 2025 and included no artificial-intelligence or foundation-model specialist, so model providers are supervised only where they also supply cloud services. In the United Kingdom a parliamentary committee recommended in January 2026 that the major artificial-intelligence and cloud providers be designated by the end of that year, and the first four designations were made in July 2026, covering cloud and technology providers only.12

Outsourcing guidance applies to model suppliers awkwardly, because it was written for a different dependency. European outsourcing guidelines require institutions to assess aggregate exposure to a single provider and the cumulative impact of multiple arrangements, to notify sub-outsourcing of critical or important functions in advance with termination rights, to document exit plans, and to maintain a register, with competent authorities addressing sector-level concentration. The guidelines are technology-neutral and predate generative artificial intelligence, and they engage only where an arrangement amounts to outsourcing a critical or important function. An institution can therefore hold a material dependency on a model provider that its outsourcing register does not capture.

The financial stability bodies have made the concentration finding themselves. The Financial Stability Board recorded in November 2024 that the markets for accelerated computing and cloud services are dominated by a limited number of entities, and that vertical integration exists across the supply chain. Vulnerabilities intensify where firms rely on a limited set of providers or on one provider for several services, and “the cost and complexity of training [large language models] from scratch” is “generally prohibitive” for nonspecialist firms. Its October 2025 follow-up carried a dedicated case study on monitoring artificial-intelligence third-party dependency.

The Bank of England’s Financial Policy Committee set out in April 2025 both the operational concentration risk and the distinct risk that convergence on similar models drives correlated market behavior, while noting that open-source availability pulls in the opposite direction.

India Layer

The Indian outsourcing directions reach the supply chain behind a model more explicitly than the European equivalent. The Reserve Bank of India consolidated its outsourcing instructions on November 28, 2025, replacing the Master Direction on Outsourcing of Information Technology Services of April 10, 2023 with a set of entity-class Directions. For a commercial bank the operative instrument is the Reserve Bank of India (Commercial Banks – Managing Risks in Outsourcing) Directions, 2025, RBI/DOR/2025-26/171, which carries the outsourcing of information technology services in its Chapter IV. Paragraph 62 requires a bank to assess the impact of concentration risk posed by multiple outsourcing arrangements to the same service provider. An inventory of outsourced IT services, including key entities involved in their supply chains, together with a mapping of the bank’s dependency on third parties, comes from paragraph 77. A clear exit strategy is mandated at paragraph 79. Paragraph 69 sets the agreement minimums for IT outsourcing and incorporates paragraph 34, whose subparagraph (iv) requires prior approval or consent of the bank before a service provider engages a subcontractor. The supply-chain inventory requirement addresses dependency beyond the immediate counterparty, which is where a model provider sitting behind an application vendor is otherwise invisible. As Chapter 20 set out, the same regime makes accountability nondelegable at paragraphs 9 and 17, extends audit rights and the Reserve Bank’s inspection powers to subcontractors at paragraphs 69 and 71, and requires at paragraph 56 that cyber incidents be reported through the vendor chain so that the bank reports to the Reserve Bank within six hours of detection by the service provider.

The FREE-AI recommendations treat capability and concentration as policy problems and not only as firm-level risks. The report of August 2025 organizes 26 recommendations across six pillars under an innovation-enablement framework covering infrastructure, policy, and capacity, and a risk-mitigation framework covering governance, protection, and assurance. Those bearing on sourcing include financial-sector data infrastructure as digital public infrastructure, indigenous financial-sector models developed as public goods, capacity building within regulated entities and within the supervisor, an artificial-intelligence inventory inside regulated entities alongside a sector-wide repository, and an incident reporting and sectoral risk intelligence framework. The report observes that existing outsourcing guidance already addresses accountability and third-party risk, and suggests enhancements including clauses addressing artificial-intelligence-specific risks and vendor disclosure requirements. The gap it identifies is narrow. The report is recommendatory, as Chapter 20 noted, and the binding instruments are now the consolidated Directions the Reserve Bank issued on November 28, 2025.13

National compute and model capability form the policy answer to the dependency the regulator has identified. The IndiaAI Mission’s compute program and its foundation-model program, described in the India cases above, are the operational expression of the recommendation that indigenous models be developed as public goods, and the clearest instance in this book of a jurisdiction treating sovereign model capability as a mitigant for a concentration risk its own central bank has named.

The Practitioner’s Lens: Chief Information Officer, Midsized European Bank

Picture the chief information officer of a European bank with around 30,000 employees, retail and commercial, operating in four countries and large enough to matter to its supervisor without being large enough to build anything at the model layer. This composite figure, drawn from practice across the sector, is two years into an agentic program and has changed her mind about most of it.14

The first year’s portfolio was chosen the way most are, and it did not survive contact with the delivery organization. Thirty-one use cases came out of a series of business-unit workshops, each sponsored by the unit that proposed it, each with a business case built on hours released. Nineteen were funded. Four reached production. The postmortem found a pattern nobody had scored for: the four that landed were all in units that had already agreed to change how the work was organized, and most of the fifteen that stalled were in units that had understood the tool as something that would make the existing process faster.

The second year’s portfolio introduced a capture-rate column, and the effect was uncomfortable and useful. Requiring each sponsor to state what would happen to released time, and to name the person accountable for realizing it, removed about a third of the candidates before scoring began. Several sponsors withdrew at that point.

The chief information officer regards that as the single most valuable control the program has, and notes that no part of it is technological.

The human oversight question forced a change in how the bank writes its designs. The original template placed a reviewer after every agent step. That satisfied the risk committee and, on inspection, meant very little. The revised template requires each review point to state what error the reviewer is expected to catch and how the reviewer would see it. Two designs failed that test outright and were rebuilt so the agent’s output arrived with the specific evidence a reviewer needed to falsify it. One design was abandoned because no such evidence could be produced, which the chief information officer regards as the clearest result the exercise delivered, since the alternative had been to deploy the thing with a review step that would have caught nothing and satisfied everyone.

Vendor strategy settled into a position the bank would not have chosen on principle. Contracts run with three model providers behind an internal abstraction layer, which is expensive and duplicative and exists so the bank can move. The supervisor’s interest lies in whether the bank could leave, and the honest answer for the first eighteen months was that it could not, because the prompts, the evaluation sets, and the tool definitions had all been built against one provider’s interfaces. Portability turned out to be an artifact of engineering discipline sustained over time.

The adoption numbers the bank reports internally are deliberately harsher than the ones it could report. Licenses issued is close to universal, daily active use is a fraction of that, and the measure the executive committee actually reviews is the proportion of eligible work items that passed through an agentic path in the month, which is smaller again. The chief information officer’s argument for the third measure is that the first two describe access and enthusiasm, and only the third describes whether the bank’s work has changed. It is also the only one of the three she would be comfortable defending to a supervisor.

Applied Exercise: The Agentic Strategy Snapshot and Implementation Roadmap

Five to eight hours, and suitable as a group capstone. Deliverable: a one-page agentic strategy snapshot for a chosen institution, supported by a completed Agentic Opportunity Prioritization Matrix for at least six candidate opportunities, a Human-Agent-System Swimlane Method diagram for the priority opportunity, a sourcing and operating-model recommendation, and an 18-month roadmap with named owners.

Choose a financial institution you can study in depth. Public disclosure is sufficient where an employer is not available, and the Indian and global cases in this chapter and in Chapters 18 to 21 provide the raw material for several.

Step 1: Build the opportunity list, and say where it came from. Identify at least six candidate agentic opportunities across the institution’s value chain. Record how each was identified, because the provenance of a list determines what is missing from it. Name at least one opportunity that no business unit would have proposed.

Step 2: Score the portfolio on the Prioritization Matrix. Apply all four axes, decomposing value into unit value, volume, and capture rate, scoring feasibility as a conjunction of its four conditions, separating severity, reversibility, and detectability in the consequence-of-error axis, and scoring strategic fit and reuse against the capabilities each opportunity would leave behind. Plot the quadrants. State explicitly what you are discarding and why.

Step 3: Test organizational willingness separately. For the top three opportunities, assess whether the owning business unit is prepared to redesign the workflow, and name the evidence for that judgment. Treat an unwilling unit as disqualifying and reorder accordingly. This step exists because it is the variable most strongly associated with realized returns and it is absent from the matrix by design.

Step 4: Draw the swimlane for the priority opportunity. Decompose the process into stages, apply the three-question Role Allocation Test to each, specify every handshake, populate the fallback lane, and place instrumentation at the boundaries.

Then subject every human review point to the falsification test set out in the Practitioner’s Lens, and move or redesign any stage that fails it.

Step 5: Make the sourcing and operating-model recommendation. Decide build, buy, or partner for each layer of the stack and justify it against the concentration findings in this chapter. State how the institution would exit a model provider, and how long that would take today. Specify the operating model: centralized, federated, or a center of excellence with delivery in the business, with the reporting line named. Identify whether any planned customization would make the institution a provider under the European regime, which turns on the three triggers the European regime sets, each of them keyed to whether the system is or becomes high-risk, and which carries the heavier obligation set if the answer is yes.

Step 6: Write the one-page snapshot and the roadmap. The snapshot states the institution’s agentic position in one page: where it will compete, the three opportunities it will fund, the capability it is building that outlasts them, its sourcing posture, and the two or three metrics its executive committee will review. The roadmap sequences eighteen months with named owners, decision points, and the metric that would cause each initiative to be stopped. State what the institution will disclose publicly, and what it will not.

Executive Briefing

Chapter 22: Designing, Sourcing, and Adopting Agentic Systems: Key Takeaways

The variable most strongly associated with getting a return is willingness to redesign the work, which no vendor supplies. Nearly three-quarters of the organizations reporting high performance had redesigned workflows end to end, against a quarter of everyone else, while the proportion of all organizations attributing any earnings impact to artificial intelligence stayed flat at 37 percent as adoption rose. Institutions buying tools to accelerate unchanged processes are the ones reporting nothing, and no prioritization framework can compensate for a business unit that will not change how the work is done.

Capability is jagged, and the errors it produces are more persuasive than the work it replaces. In a controlled trial of 758 consultants, access to a frontier model raised quality by around a third on tasks inside its capability and cut correctness by up to 24.5 percentage points on a task outside it. The answers produced outside the frontier were rated more coherent than the unaided ones even though they were less correct. Any design that relies on a reviewer noticing a problem must account for the fact that agent output presents better than human work at the moment it is wrong.

Human oversight is a design decision requiring justification, and it is not free. Experimental evidence shows that adding a human to the loop raises willingness to use a system while reducing its accuracy, with reviewers correcting large errors less often and by smaller margins than small ones. Separate evidence finds no general tendency to defer to machine advice, but a marked tendency to follow it selectively when it confirms an existing stereotype. For credit and collections that second finding is the material one, and it is not addressed by placing a reviewer in the process.

Sourcing has narrowed to choosing which external dependency to hold, and the regulators have said so. The Financial Stability Board has called training a large language model from scratch generally prohibitive for firms that are not technology specialists. Concentration at the leading provider is falling, from around half of tracked bank deployments to a third in eighteen months, which describes a shift from single-provider dependence to a three-provider oligopoly. Operational concentration and model monoculture are different risks, and multisourcing addresses only the first.

The regulatory architecture for this dependency exists and has not been applied. The European resilience regime designated its first tranche of critical technology providers in November 2025 with no artificial-intelligence or foundation-model specialist among them, so model capability is supervised only where the provider also sells cloud. In the United Kingdom, a parliamentary committee demanded designations by the end of 2026, and the first four were made in July 2026. All four are cloud providers. Meanwhile the obligation that most affects a customizing bank is the one that converts a deployer into a provider by conduct, and it sits in the sourcing decision rather than in a later compliance review.

Key Terms

Jagged frontier. The uneven boundary of a model’s capability, along which tasks of apparently similar difficulty fall on opposite sides, and which cannot be inferred from resemblance between tasks.

Capture rate. The share of theoretically released value an institution can realistically convert into cost or revenue, distinct from the value nominally freed by an intervention.

Selective adherence. The tendency to follow algorithmic advice more readily where it confirms a pre-existing belief or stereotype, distinct from automation bias, which describes indiscriminate deference.

Agent washing. The marketing of conventional automation or assistant products as agentic, identified as widespread in the vendor population.

Provider and deployer. The two roles the European artificial-intelligence regime distinguishes, carrying different obligations, with a deployer assuming provider obligations on three triggers keyed to high-risk status: branding, substantial modification, and a change of intended purpose that makes a system high-risk.

Critical third party. A technology provider designated by an authority as systemically important to the financial sector, bringing it within a direct oversight regime.

Model monoculture. The risk that widespread use of a small number of models or of very similar model designs produces correlated behavior across institutions, distinct from the operational risk of provider failure.

Handshake. The specified transfer of work between lanes in a process design, defined by what passes, in what form, with what completeness guarantee, and what the receiving lane does when it is incomplete.

Fallback lane. The lane in a process design carrying what happens when the agent cannot proceed, a system is unavailable, a handshake fails validation, or volume exceeds capacity.

Center of excellence. An organizational unit holding artificial-intelligence capability centrally and supporting delivery in the business units, as distinct from fully centralized or fully federated operating models.

Regular use. An adoption measure counting employees who use a system on a recurring basis, materially harder to satisfy than licenses issued and more informative than daily access among licensed users.

Outcome-based contracting. A commercial model pricing a service against delivered results, in place of effort or headcount, still a small minority of information technology services revenue as of 2026.

Reflection Questions

List six agentic opportunities in your institution and record how each was identified. What kind of opportunity does your identification process systematically fail to surface, and who would have to be in the room for it to appear?

Apply the capture-rate test to a business case you have seen. If the released hours are distributed in fragments across many people, who is accountable for converting them into anything, and what happens to the business case if nobody is?

For one agentic design in your institution, state for each human review point what error the reviewer is expected to catch and how they would see it. Which review points cannot survive that question?

The evidence suggests no general automation bias but a clear tendency to follow machine advice selectively when it confirms an existing view. Where in your credit or collections process would that pattern be hardest to detect, and what control would detect it?

If your institution had to leave its principal model provider, how long would it take and what would break? Answer for the prompts, the evaluation sets, the tool definitions, and the integration layer separately.

Santander attributes roughly 1 percentage point of the improvement in its cost-to-income ratio to artificial intelligence. If your institution had to state an equivalent figure to investors, what would it be, and what would you need to measure between now and then to be able to defend it?

Further Reading

On the empirical foundations of role design, three papers repay reading in full. The Organization Science study of consultants at Boston Consulting Group, published in 2026, established the jagged frontier and the coherence trap. The Quarterly Journal of Economics study of customer support agents, published in 2025, established the skill-leveling effect and identified the mechanism. The PLOS ONE study of human-in-the-loop decision-making, published in 2024, is short and should be read by anyone designing an oversight control.

On adoption and returns, McKinsey’s state of artificial intelligence survey for 2026 and Deloitte’s study of artificial-intelligence return, published in October 2025, should be read together and against each other, since both are self-reported and both screen their samples in ways that flatter the technology. The Evident AI Index tracking of disclosed bank use cases is the most useful independent measure available, with the caveat that it counts announcements.

On concentration and sourcing, the Financial Stability Board’s report on the financial stability implications of artificial intelligence, November 2024, and the Bank of England’s Financial Stability in Focus of April 2025 are the two primary regulator statements and they are balanced in a way most commentary on them is not. The House of Commons Treasury Committee report of January 2026 and the responses published in April 2026 document the gap between a designation regime and its use.

On the Indian position, the Reserve Bank of India (Commercial Banks – Managing Risks in Outsourcing) Directions, 2025 and the FREE-AI report of August 2025 should be read together, the first for the binding obligations that reach a model provider through the supply chain and the second for the policy logic connecting concentration risk to sovereign model capability. The IndiaAI Mission’s published compute and model-program updates supply the current state of that program.

References and notes

  1. Dell’Acqua, Fabrizio, Edward McFowland III, Ethan Mollick, Hila Lifshitz, Katherine C. Kellogg, Saran Rajendran, Lisa Krayer, François Candelon and Karim R. Lakhani, “Navigating the Jagged Technological Frontier: Field Experimental Evidence of the Effects of Artificial Intelligence on Knowledge Worker Productivity and Quality,” Organization Science 37(2), 2026, 403–423, for the sample of 758 consultants, the quality improvements inside the frontier, the fall in correctness outside it against an 84.5 percent control baseline, the coherence ratings on incorrect answers, which run to 1.45 points for the model-plus-training condition and 1.78 points for the model-only condition on a ten-point scale, and the distribution of gains across prior performance.

  2. Evident Insights, AI Use Case Trends in Banking Q1 2026, May 2026, for the agentic share of newly disclosed use cases, the proportion of banks disclosing a return target, the proportion of use cases reporting outcomes, and the specialist-vendor share; and Evident use case tracker data reported January 2026 for the change in the leading provider’s share of tracked deployments. McKinsey and Company, The State of AI in 2026, August 2026, for the adoption and earnings-impact figures, the workflow redesign comparison, and the build-versus-buy software finding. Deloitte, AI ROI: The Paradox of Rising Investment and Elusive Returns, October 22, 2025, for the agentic usage and measurable return figures, the payback comparison, and the sourcing intention split.

  3. Gartner press release, June 25, 2025, for the cancellation prediction and the estimate of genuine agentic vendors. Bank of America, chief technology and information officer interview reported April 14, 2026, for the prioritization criteria and the technology budget; Bank of America press release of April 8, 2025, for the internal employee assistant adoption figure; and Bank of America disclosures reported July 14, 2026, for the employee usage figures and the approved, live generative and fully implemented use case counts. Citigroup, Introducing AI Agents, April 30, 2026, for the Arc platform, its gated rollout and the access and usage figures; Citi training disclosures reported October 2025.

  4. Brynjolfsson, Erik, Danielle Li and Lindsey R. Raymond, “Generative AI at Work,” Quarterly Journal of Economics 140(2), 2025, 889–942, for the 5,172 support agents, the 15 percent average productivity effect, the distribution of gains by skill, the customer sentiment and attrition findings, and the tacit-knowledge dissemination mechanism.

  5. Sele, Daniela, and Marina Chugunova, “Putting a human in the loop: Increasing uptake, but decreasing accuracy of automated decision-making,” PLOS ONE 19(2), February 9, 2024, e0298037, for the uptake and accuracy findings and for the distribution of corrections across error sizes. Alon-Barkat, Saar, and Madalina Busuioc, “Human-AI Interactions in Public Sector Decision Making: Automation Bias and Selective Adherence to Algorithmic Advice,” Journal of Public Administration Research and Theory 33(1), January 2023, 153–169, for the absence of general automation bias and the selective adherence finding, in a Dutch public-sector setting whose transfer to financial services must be argued.

  6. BBVA, announcement of its alliance with OpenAI, December 12, 2025, for the deployment scale, and for the daily usage and time saved, which are reported for the eleven-thousand-user phase rather than the full rollout; issued jointly with the model provider and flagged as such in the text. NatWest Group, technology and artificial intelligence disclosure, February 13, 2026, for access, coding, hours saved, relationship manager time and voluntary training take-up. Ankush Das, “Inside Axis Bank’s Six-Point GenAI Strategy,” Inc42, February 16, 2026, for the center of excellence, the Business Intelligence Unit structure, the data estate, adoption and coding automation, and the profit and loss projection, which was given in a trade interview and is flagged in the text as management’s own unaudited model. DBS Group Holdings, Annual Report 2025, for assistant adoption, the coding time reduction, the model and use case counts and the economic value figure, which is a bank-defined construct.

  7. Kim, Bowon, “AI-Induced Job Anxiety and the Perceived Effectiveness of AI-Enabled ESG Initiatives: Evidence from Bank Employees,” Sustainability 18(9), April 28, 2026, 4353, for the survey of 858 employees of a major Korean commercial bank and the appraisal and engagement findings, with the cross-sectional, single-bank and sustainability-framed limitations noted in the text.

  8. Financial Stability Board, The Financial Stability Implications of Artificial Intelligence, November 14, 2024, for the concentration findings, the vertical integration observation and the statement on the prohibitive cost of training large models; and Monitoring Adoption of Artificial Intelligence and Related Vulnerabilities in the Financial Sector, October 10, 2025. Bank of England, Financial Stability in Focus: Artificial intelligence in the financial system, April 9, 2025, for the provider reliance and model convergence warnings and the offsetting open-source observation.

  9. Regulation (EU) 2024/1689, Articles 16, 25, 26 and 72, for provider and deployer obligations, the post-market monitoring obligation on providers, and the conduct by which a deployer assumes provider obligations. Regulation (EU) 2026/1744 amending Regulation (EU) 2024/1689, in force July 27, 2026, for the revised application dates for high-risk obligations and for the transitional period for content-marking transparency. Regulation (EU) 2022/2554, Articles 31 and 35, for the designation criteria and the penalty provision; European Supervisory Authorities, designation of the first critical ICT third-party service providers, November 18, 2025, listing nineteen providers and containing no artificial-intelligence or foundation-model specialist. European Banking Authority, Guidelines on outsourcing arrangements, EBA/GL/2019/02, for aggregate exposure, sub-outsourcing notification, exit plans and the register.

  10. ICRA rating rationale for Perfios Software Solutions Private Limited, February 10, 2026, for the operating income, growth, margins, repeat business, client list and the 2025 acquisitions. Sarvam AI funding announcement, June 15, 2026, for the round size, valuation and lead investor, and for the financial-sector deployment figures, which name no client and are flagged in the text as unverifiable. Ministry of Electronics and Information Technology press release of May 30, 2025, for the compute capacity as at that date and the foundation-model program selections; the Minister for Electronics and Information Technology’s reply in the Lok Sabha of July 2025, for the count of empanelled providers and the subsidised average rate; the Minister of State for Electronics and Information Technology’s statement to the Lok Sabha of March 25, 2026, for the outlay and the updated compute capacity; and the Press Information Bureau press note of December 30, 2025, for the twelve foundation-model developers selected across the first two phases.

  11. Commonwealth Bank of Australia, appointment announcements for its Chief Artificial Intelligence Officer and its first Chief Artificial Intelligence Scientist, 2026, for the dual-role structure and the named technology partners. Banco Santander, 2026 Investor Day, February 25, 2026, for the business value target, the contribution to the improvement in the cost-to-income ratio and the efficiency target. Business Standard, August 20, 2026, for the outcome-based contracting shares reported by Coforge, Cognizant and Infosys, which sit alongside the repricing of delivery away from headcount described in Chapter 20, where the model reported is intellectual-property licensing with minimum volume commitments.

  12. House of Commons Treasury Committee, Artificial intelligence in financial services, Fifteenth Report of Session 2024-26, HC 684, January 20, 2026, for the reliance finding and the designation recommendation; and AI in financial services: Responses to the Committee’s Fifteenth report, Seventh Special Report of Session 2024-26, April 16, 2026, for the position as at that date. HM Treasury, first designations of critical third parties, July 10, 2026, with oversight commencing July 13, 2026, for the four designated providers, which are Amazon Web Services, Google Cloud, Microsoft and Oracle.

  13. Reserve Bank of India, Managing Risks in Outsourcing Directions for commercial banks, RBI/DOR/2025-26/171, DOR.ORG.REC.No.90/21-04-158/2025-26, November 28, 2025, paragraphs 62, 77, 79 and 69 read with paragraph 34(iv), for the concentration risk assessment, the supply-chain inventory, the exit strategy and the subcontracting consent requirements. They replaced the Master Direction on Outsourcing of Information Technology Services of April 10, 2023, which is on the Reserve Bank’s Circulars Withdrawn register, and the consolidation issued parallel Directions for eight further classes of regulated entity. Reserve Bank of India, Report of the Committee on Framework for Responsible and Ethical Enablement of Artificial Intelligence, August 13, 2025, for the pillar structure and the recommendations bearing on data infrastructure, indigenous models, capacity building, inventory and incident reporting, with the report’s recommendatory status noted in the text. Shaktikanta Das, then Governor of the Reserve Bank of India, “Central Banking at Crossroads,” RBI@90 High-Level Conference, New Delhi, October 14, 2024, on concentration risk and algorithmic opacity.

  14. The Practitioner’s Lens in this chapter is composite and illustrative and does not refer to any specific named executive or institution. Its framing draws on the operating-model and sourcing disclosures cited above and on the published record of agentic deployment at midsized European banks.