
Before buying an AI product, ask 25 questions across five areas: training-data use, model provenance, data handling, security and accountability. The single most important one is whether your data is used to train the vendor's models, and whether you can opt out contractually. That answer separates enterprise-ready vendors from the rest.
What are the five question areas?
The 25 split five ways: training-data use (questions 1 to 5), model provenance (6 to 10), data handling (11 to 15), security (16 to 20) and accountability (21 to 25). Ask them in writing. Verbal assurances from a sales engineer do not survive contract negotiation.
- Is our data used to train your models, and can we opt out contractually?
- Does the opt-out cover fine-tuning and "product improvement", or only foundation model training?
- Are our prompts and outputs retained for training by any sub-processor or underlying model provider?
- Can you evidence the opt-out technically, through configuration and logs, and warrant it in the contract?
- If we terminate, what happens to any of our data already absorbed into a model?
- Which underlying models power the product, and are they built in-house or licensed from a third party?
- When you swap or upgrade an underlying model, how and when are we told?
- What can you tell us about the data your models were trained on, and will you stand behind its lawfulness?
- Which open-source components are in the product, and under what licences?
- How is a new model version tested before it reaches our users?
- Where is our data processed and stored, and can processing be confined to a region we choose?
- How long are prompts, outputs and logs retained, and can we set the retention period?
- Who are your sub-processors, including model providers, and how are we notified of changes?
- How is our data segregated from other customers' data?
- At termination, how is our data deleted, and will you certify deletion?
- Which certifications and audit reports can you share, such as ISO 27001 or SOC 2?
- How do you defend against prompt injection and the other OWASP risks for LLM applications?
- How is your staff's access to our data controlled and logged?
- What is your breach notification commitment, in hours, and who gets the call?
- Have the AI features specifically been penetration tested by an independent party?
- Who is liable when the model outputs something wrong, infringing or defamatory?
- Do you indemnify us for IP infringement claims arising from model outputs?
- Which accuracy or performance commitments will you put in the contract?
- Can we audit, or receive independent assurance over, your AI-specific controls?
- How are you preparing for incoming AI regulation in the markets we operate in?
Good answers share a texture: crisp and checkable. On training data, "no by default, and here is the clause". On provenance, named models and a notification commitment for swaps. On data handling, a published sub-processor list and configurable retention. On security, current audit reports and a breach commitment measured in hours. On accountability, a vendor willing to discuss indemnity at all.
Which answers are dealbreakers?
The training-data question is the litmus test. A vendor that trains on customer data by default, offers no contractual opt-out, or cannot say what its underlying model provider does with your prompts has answered every other question for you. The engineering to keep customer data out of training pipelines is table stakes for enterprise AI vendors now, and a vendor that has not built it is telling you where you sit in its priorities.
Watch for the soft dealbreakers too. "The opt-out is on the roadmap" means no. "We can't disclose our sub-processors" means your data's route is a secret, from you, its owner. A refusal to put any breach notification window in writing means the window is however long their lawyers want it to be. Any one of these should stop a purchase involving confidential or personal data.
Vagueness is itself the signal. Vendors with good answers give them quickly.
What goes in the contract?
An AI addendum, a short schedule sitting alongside the data processing agreement, carrying four clause families.
Data use: no training on your data without express written consent, worded to cover fine-tuning, product improvement and the underlying model provider, and matching answers 1 to 4 above. Retention: defined periods for prompts, outputs and logs, your right to set them, and certified deletion at exit. Indemnity: cover for third-party IP infringement claims arising from outputs, with the cap and carve-outs negotiated rather than defaulted. Sub-processors: a current list, advance notice of changes, and an objection right, explicitly including swaps of the underlying model.
The addendum matters because standard DPAs predate the training-data question. A DPA written for a CRM vendor says nothing about whether your support tickets become someone else's model weights.
How does this fit your vendor-risk process?
Extend the third-party risk management process you already run. Do not build a parallel one for AI. The 25 become an AI module appended to your existing questionnaire, triggered whenever a vendor or renewal involves AI features, and the answers feed the same risk tiers, the same approval chain and the same annual review cycle as every other vendor.
A parallel AI process drifts. It gets a different owner, a different template, then a different risk scale, and two years later nobody can say which register is authoritative. One process, one register, one extra module.
One wrinkle worth planning for: AI features arrive inside vendors you already approved. The document platform you cleared in 2023 ships an assistant in 2026. Renewal reviews, and contract clauses requiring notice of new AI features, are where the module catches those.
Zavior stores each vendor's answers to the 25 as evidence in your vendor register, so a renewal review starts from last year's answers instead of a blank questionnaire.
Frequently asked questions
Do standard security questionnaires cover AI?
Partly. They cover hosting, access control, encryption and incident response, and those answers still matter. They are silent on training use, model provenance and output liability, because they were written before those risks existed. Keep the questionnaire and add the AI module rather than replacing anything.
What's an AI addendum?
A short contract schedule, layered onto the main agreement and the DPA, that pins down the four clause families above: data use for training, retention, indemnity for outputs, and sub-processor transparency. It exists because standard vendor paper predates AI-specific risks and stays silent on them by default.
How do you assess model quality claims?
Do not buy on leaderboard benchmarks. Ask for evaluation results on tasks shaped like yours, then run a time-boxed pilot on your own data with pass criteria agreed before it starts. Whatever performance the vendor claims in the deck, ask them to commit to it in the contract, and watch how quickly the claim shrinks.
Zavior · AI Governance
Good answers are only useful if they reach the contract and stay checkable. Zavior folds the AI addendum and the 25 into the third-party risk process you already run, links each answer to the cyber-security control it depends on, and catches the moment an approved tool switches on a new AI feature. One questionnaire, one register, no parallel AI process quietly going stale in a corner.
Book a free 30-minute business assessment →This is general information, not legal advice.