Blog
August 17, 2026
2
MIN READ
Your AI Model Is Now an Attack Surface

AI models are now an attack surface. Learn how model targeting, prompt injection and AI-driven attacks are reshaping security risks across BFSI.

Share this post

TABLE OF CONTENT

AI systems as active decision-makers

For most of the last decade, the security question about AI in banking was a data question. Is the training data protected? Is the model encrypted at rest? Who can access the pipeline? Those questions assumed the model was a passive asset — something you locked in a vault and monitored like any other sensitive file.

That assumption is now the vulnerability.

Across our forensic and red-team work this year, a different picture emerged, documented across SISA’s Digital Threat Report 2025-26. The AI systems making credit, fraud, and onboarding decisions in BFSI are no longer passive. They are active decision-makers running at machine speed — and increasingly, active participants in the attack itself. The model is being targeted as a model, exploited as an execution environment, and turned into a weapon pointed back at the institution that deployed it. Three different things are happening at once, and most security programs are only watching for none of them.

The three ways in which AI turns on its owner

1. The model as target

Start with the most direct shift. When a model approves or declines a transaction, its decision boundary is a surface an attacker can map.

The technique is unglamorous and effective: submit slightly varied inputs, watch what passes and what doesn't, and reverse-engineer the logic. Enough iterations and an adversary learns exactly what a fraudulent application needs to look like to clear the model at minimum cost. No malware. No breached credential. Just patient probing of a system that was assumed to be internal and therefore trustworthy.

The uncomfortable part is what makes this possible. Many models helpfully expose their own confidence scores and decision rationale on applicant-facing surfaces — the very signals an attacker needs to calibrate the attack. The model is teaching its adversary how to beat it.

And here is the sentence that should worry a CISO most: a model that was never red-teamed for adversarial robustness can quietly perform worse than the rule set it replaced. You didn't upgrade. You traded a legible rulebook for an opaque one that fails in ways you can't see.

2. The model as execution environment

The second shift is newer and stranger. Agentic AI — systems that read documents, act on instructions, and complete tasks — has moved into production. These agents typically operate with broad, admin-adjacent privilege, but under governance designed for a piece of software, not for a privileged user.

That gap is the opening. Prompt injection has stopped being a lab curiosity. Malicious instructions hidden in an email, a document, or a web page get ingested by an enterprise agent and executed as if they were legitimate. The document is the carrier. The agent is the executor. And the agent has the access.

The mental model to abandon: the agent is not a feature. It is an identity — one with privilege, one that acts, and one that currently answers to almost no one.

3. The model as weapon

The third shift closes the loop. The same AI tooling your developers use to move faster is available to your attackers, and they are using it to generate polymorphic malware and fraudulent documents that are structurally unique on every iteration — which is precisely what defeats signature-based detection. Synthetic identities engineered to pass automated KYC. Deepfake voices authorizing wire transfers. Grammatically flawless, context-aware phishing at a scale no human team could produce.

Detection logic built to catch clumsy, human-authored attacks loses its footing here, because the tell it was trained on — the imperfection — is gone.

Why compliance doesn't see any of this

The reason these three shifts stay invisible is structural, and it's the point worth taking to a board.

This isn't hypothetical. In a recent audit engagement across a payments environment, we repeatedly found controls that passed attestation but dissolved under adversarial conditions: cardholder data encrypted at rest, yet fully readable by database administrators; sensitive authentication data that should have been destroyed after authorization sitting in application logs and database tables; encryption keys that hadn't been rotated in years. Every one of those environments could produce a clean point-in-time attestation. None of them would survive contact.

The reason these AI shifts stay invisible is the same structural reason. Compliance validates that AI workloads are encrypted at rest. It confirms the pipeline has access controls. It attests, at a point in time, that the control exists. None of that touches whether the model holds up when someone probes its decision boundary, injects instructions into its inputs, or attacks it with another model. Attestation asks does the control exist. Adversarial reality asks does it survive contact. Between those two questions sits the entire attack surface described above.

You can be fully compliant and fully exposed at the same time. In AI, that's not an edge case. It's the default.

What actually changes on Monday

Three moves, in order of leverage:

  • Treat adversarial robustness testing as a deployment gate, not a research project. Before a fraud, credit, AML, or KYC model goes live, it should have to survive decision-boundary probing and adversarial inputs — with a defined cadence to re-test, because a model that was robust at launch drifts. And stop broadcasting confidence scores and decision rationale to applicant-facing surfaces. You are handing attackers the calibration data.
  • Govern AI agents as privileged identities. Same rigor you apply to a human with admin rights: continuous behavioral monitoring, least privilege, an audit trail. Add prompt-injection-aware controls at the point where the agent ingests content — because that ingestion boundary is now an attack boundary.
  • Move detection from content to behavior. When attacks are AI-generated, the content looks perfect. What still gives an attacker away is behavior — the rhythm of probing, the velocity of queries, the systematic variation of inputs. That's the signal that survives when content quality stops being a tell.

The old question was whether your AI was protected. The real question now is whether it can be trusted — under pressure, against an adversary using the same tools you are, while acting with privileges you granted and forgot to watch. In BFSI, where a model's decision is instant and often irreversible, that's not a technical concern filed under data security. It's a business risk that belongs on the risk register, with an owner, next to the others. Your model is no longer just something you defend. It's something that can be turned against you. Govern it accordingly.

The three shifts described here: the model as target, as execution environment, and as weapon, are drawn from SISA's forensic and red-team casework across BFSI this year. The full analysis, including the other structural shifts reshaping the sector, is in SISA's Digital Threat Report 2025–26.

SHARE THIS POST

AI Security
Sappers DFIR

Lorem ipsum dolor sit amet, consectetur adipiscing elit. Suspendisse varius enim in eros elementum tristique. Duis cursus, mi quis viverra ornare, eros dolor interdum nulla, ut commodo diam libero vitae erat. Aenean faucibus nibh et justo cursus id rutrum lorem imperdiet. Nunc ut sem vitae risus tristique posuere.