Company verification
AI Vendor Due Diligence for Financial Institutions: What to Check Beyond the Demo
AI vendor due diligence for banks and insurers: six kinds of evidence no demo can give you, DORA Article 30 clauses to insist on, and how to test on your data.

A vendor demo is a curated sample: the vendor’s data, the vendor’s prompts, the vendor’s happy path — and, more often than procurement teams assume, a human quietly sitting in the loop. None of that is evidence. Due diligence begins where the demo ends, and in a regulated firm what it has to produce is a file that survives supervisory review, not a good feeling in the room.
The short answer: what to check beyond the demo
AI vendor due diligence at a financial institution comes down to proving six things a demo cannot: model provenance (whose model, running where, on whose infrastructure, with which sub-processors), evaluation evidence produced on your own data rather than the vendor’s benchmark, change management that tells you when the model version underneath you changes, data rights covering training, retention and residency, contractual audit and access rights that meet your regulator’s standard, and a tested exit path. Under the Digital Operational Resilience Act (Regulation (EU) 2022/2554, applicable since 17 January 2025), most of those stop being good practice and become mandatory contractual terms as soon as the service supports a critical or important function — Article 30(3). Everything else, certifications and reference calls included, is context rather than proof.
The practical test for any due diligence pack is simple: could a second-line reviewer who never saw the demo reconstruct why you believed the system works? If the answer rests on a screen recording, you have a sales artefact, not an assessment.
Why demos systematically mislead
The failure mode is usually scope, not dishonesty. The demo runs on data the vendor curated, at a volume that never touches a rate limit, on a model version that will be silently replaced two quarters after you sign.
Sometimes it is dishonesty. In March 2024 the SEC settled charges against two investment advisers, Delphia (USA) Inc. and Global Predictions Inc., over false and misleading statements about their use of AI, with civil penalties of $225,000 and $175,000 respectively. In April 2025 the U.S. Attorney’s Office for the Southern District of New York and the SEC brought charges against Albert Saniger, founder of the shopping app Nate, alleging that transactions marketed as AI-automated had in fact been completed manually by contractors. Both cases are worth citing in an internal paper for the same reason: the claims were checkable, and nobody checked.
The control that carries over to every deal is a written question about the automation rate in production: what share of transactions complete with no human intervention, measured over the last 90 days, and what happens when the fallback path fires. A vendor who cannot answer that from telemetry does not have telemetry.
Anchor the assessment to the rulebook you actually sit under
The evidence you can demand, and the leverage you have to demand it, depend on your regime. Settle that first, because a supervisor will read your file against their own rulebook, not against a generic vendor scorecard.
EU. DORA governs ICT third-party risk end to end: a pre-contractual assessment of concentration risk (Article 29), mandatory contractual terms (Article 30), a register of information on all contractual arrangements reported to your competent authority (Article 28(3)), documented exit strategies (Article 28(8)), and an oversight framework for providers designated as critical (Article 31). Separately, the EU AI Act (Regulation (EU) 2024/1689) treats AI used to evaluate the creditworthiness of natural persons or establish credit scores as high risk under Annex III, point 5(b), with fraud detection carved out. The application dates for the high-risk obligations were deferred by the Digital Omnibus on AI, formally adopted in June 2026: stand-alone Annex III systems take effect on 2 December 2027, and AI embedded in Annex I regulated products on 2 August 2028 — even so, check the consolidated text on EUR-Lex and your national supervisor’s latest communication rather than trusting a date in any article, this one included.
United States. The Interagency Guidance on Third-Party Relationships: Risk Management (Federal Reserve, FDIC and OCC, June 2023) sets the lifecycle: planning, due diligence and third-party selection, contract negotiation, ongoing monitoring, termination. Layer SR 11-7 / OCC Bulletin 2011-12 on top of it. The model risk guidance is explicit that vendor models belong inside your own model risk management framework, that you are expected to validate your use of a vendor product, and that you should have contingency plans ready for the day the product is no longer available.
United Kingdom. PRA SS2/21 covers outsourcing and third party risk management, and the critical third parties regime (PRA SS6/24 and FCA PS24/16) took effect on 1 January 2025 for providers designated by HM Treasury. The Bank of England and the FCA also publish a periodic survey, Artificial intelligence in UK financial services, which is a useful benchmark for where peers actually are — read the current edition rather than quoting a figure from an older one.
The evidence pack: what to demand in writing
Model provenance and the sub-outsourcing chain
Ask whether the vendor trains its own weights, fine-tunes an open-weight model, or wraps a frontier API. None of the three is disqualifying, but each leaves you with a different set of risks to manage. A wrapper means the model you are validating is controlled by a fourth party you have no contract with — precisely the sub-outsourcing chain DORA Article 30(2) requires you to pin down, and a concentration risk under Article 29 if three of your vendors turn out to sit on the same upstream model.
Request the sub-processor list, the regions where inference runs, and where prompts, outputs and logs are stored. “The cloud” is not an answer; a region name is.
Evaluation evidence on your data
A vendor benchmark proves the vendor can build a slide. Ask instead for the evaluation harness, the method used to build the held-out set, and the raw predictions — then re-run the whole thing against your own historical population. For anything touching credit, pricing, fraud or surveillance, run a shadow or champion-challenger period against the incumbent process before the system makes a single live decision.
Name the metrics up front, including the ones that will embarrass everyone: false negative rate at your operating threshold, performance on segments where your book is thin, and stability across two years rather than one quarter.
Change management and version pinning
The most under-negotiated clause in AI contracts is the silent model update. Your validated system changes underneath you, and the first indication is a drift alert in production — assuming you built one.
Contract for advance written notice of model version changes, the ability to pin a version for a defined window, regression results accompanying each change, and the right to trigger revalidation. Tie the vendor’s notice period to your revalidation cycle, not the other way round.
Data rights, retention and residency
Pin down in the contract whether your inputs and outputs are used for training, fine-tuning, evaluation or abuse monitoring, and by whom: the vendor, the upstream model provider, or both. Get retention windows for prompts, outputs and logs in days, plus deletion mechanics you can verify rather than a policy statement you cannot.
Audit and access rights
For services supporting critical or important functions, DORA Article 30(3) requires unrestricted rights of access, inspection and audit, alongside cooperation with threat-led penetration testing under Article 26. Many US-headquartered AI vendors resist this and offer a SOC 2 report instead. That is a gap to escalate to your risk committee with a written rationale, not a box to tick as “mitigated by certification”.
Exit and reversibility
Ask what you take with you: prompt libraries, fine-tuning artefacts, embeddings, labelled feedback data, evaluation results. Then ask how long migration takes and who pays for it. DORA Article 28(8) requires exit strategies with an adequate transition period; the honest reading of that requirement is a rehearsed plan, not a paragraph.
What the certificates actually prove
| Artefact | What it proves | What it does not prove | How to close the gap |
|---|---|---|---|
| SOC 2 Type II | Stated controls operated over a defined period, within the auditor’s scope | Model accuracy, drift, bias, or which model version serves you | Read the scope section and the exceptions, not the cover page; confirm the AI systems are in scope |
| ISO/IEC 42001:2023 certificate | An AI management system exists and has been audited | That any particular model meets your risk appetite | Request the scope statement and the statement of applicability |
| ISO/IEC 27001:2022 certificate | An information security management system is in place and certified | AI-specific control over training data or the model supply chain | Ask which Annex A controls were excluded, and why |
| Model card or system card | Intended use, stated limitations, and the evaluation sets the vendor used | Performance on your population | Re-run the evaluations on your own held-out data |
| Penetration test report | Application-layer security at a single point in time | LLM-specific attack surface: prompt injection, or exfiltration through tool use | Require testing mapped to the OWASP Top 10 for LLM Applications and MITRE ATLAS |
| Reference customer call | Somebody is live in production | Comparability to your scale or regulatory tier | Ask for a reference at your tier, and ask about incidents rather than satisfaction |
The NIST AI Risk Management Framework (AI RMF 1.0, January 2023) and its Generative AI Profile (NIST AI 600-1, July 2024) are useful for structuring the questionnaire itself. Both are voluntary and neither is certifiable, so treat “NIST-aligned” in a vendor deck as an opening for questions rather than an answer to them.
The vendor viability question that gets skipped
An AI vendor can pass every control test and still be gone in eighteen months. Ask for audited accounts where they exist and, where they do not — which covers most venture-funded vendors — ask directly, under NDA, for months of runway and the date of the last raise. UK entities are checkable on Companies House, EU entities in national business registers, and US registered advisers on the SEC’s IAPD.
Then ask the follow-up: if the vendor fails, what survives? Source code and model escrow are worth exactly as much as the rehearsal behind them. An escrow deposit you have never tried to deploy is a filing cabinet.
Contract clauses that do the actual work
DORA Article 30 reads like a compliance chore, but it works as a first-rate procurement checklist in any jurisdiction. Insist on a precise description of functions and services; the locations where services are provided and data processed; data protection and access provisions; service levels with quantitative targets; incident notification on a clock strictly faster than your own regulatory reporting deadline; assistance during ICT incidents; termination rights; and exit support.
Two clauses deserve separate negotiation. Liability first: a standard SaaS cap of twelve months’ fees bears no relation to the exposure created by a system making credit or surveillance decisions at volume. Insurance second: request the technology E&O and cyber certificates and the exclusions pages, because AI-related exclusions are increasingly common and a broker’s summary will not show them.
Running the assessment so it survives challenge
Sequence matters: a structured RFI first, then evidence review against the six buckets, then technical validation on your own data, then contract negotiation with the gaps found in stages one to three as the agenda, and finally ongoing monitoring with named metrics and a fixed review cadence.
Two governance habits separate a file that holds from one that does not. Record the residual risks you accepted and who accepted them by name; supervisors are far more comfortable with a documented accepted risk than with a gap nobody noticed. And decide early whether your deployment makes you a provider rather than a deployer under the EU AI Act — putting your own brand on a high-risk system, substantially modifying it, or changing its intended purpose can shift Article 25 obligations onto you. That determination belongs in the due diligence file before signature, not in the first supervisory dialogue after go-live.
We report facts with their date and caveats. We never label a named company as fraudulent or "AI-washing" as a statement of fact — we present verifiable data and the questions an investor should ask.
Frequently asked questions
Is a SOC 2 Type II report enough for AI vendor due diligence?+
No. A SOC 2 Type II opinion covers the controls the auditor was asked to test over a stated period, and says nothing about model accuracy, drift, bias or version control. Read the scope section and the exceptions rather than the cover page, and confirm that the AI systems you are buying were actually in scope. Treat it as one input alongside evaluation evidence on your own data.
Does DORA apply if our AI vendor is based outside the EU?+
DORA binds the financial entity, not the vendor, so a non-EU provider has no direct obligation — but your contract must still carry the Article 30 content, which means the vendor has to accept it commercially. The arrangement must specify where services are provided and data is processed. Note that a non-EU provider designated as a critical ICT third-party provider under Article 31 is required to establish a subsidiary in the Union before oversight can be exercised.
What do we do if the vendor refuses on-site audit rights?+
For services supporting a critical or important function, DORA Article 30(3) requires unrestricted rights of access, inspection and audit, so a refusal is a material gap rather than a negotiating detail. Pooled audits or an audit conducted by an appointed third party can be acceptable alternatives; a SOC 2 report alone generally is not. If the vendor will not move, document the residual risk and take the decision to the accountable committee by name.
Do we still have to validate the model if the vendor has already validated it?+
Yes, if you are a US bank subject to SR 11-7 / OCC Bulletin 2011-12, which places vendor models inside your own model risk management framework and expects you to validate your use of the product. Vendor validation evidence is an input; it is not a substitute for testing on your population at your operating thresholds. The same logic holds in the EU, where the deployer carries the operational risk regardless of who built the system.
How should we assess a vendor that is really a wrapper on a frontier model?+
Being a wrapper is not disqualifying, but it changes what you are buying and what you must document. Identify the upstream provider as part of the sub-outsourcing chain, check whether several of your vendors depend on the same model (a concentration risk you must pre-assess under DORA Article 29), and negotiate for version pinning plus advance notice of model changes the vendor cannot unilaterally control. Ask explicitly whether any upstream IP indemnity is passed through to you and on what conditions.
When does buying an AI system make us a provider rather than a deployer under the EU AI Act?+
Article 25 shifts provider obligations onto you if you put your own name or trademark on a high-risk system, make a substantial modification to it, or change its intended purpose so that it becomes high risk. White-labelling a vendor's credit-decisioning tool under your brand is the common trigger in financial services. Make this determination during due diligence and record it, because the obligations that follow are considerably heavier than a deployer's.
Related reading