← Back to blog

Legal, Procurement & Compliance: NIST+FTC AI Vendor Due Diligence

October 5, 2026
Legal, Procurement & Compliance: NIST+FTC AI Vendor Due Diligence

Before signing, we require explicit data-use limits, verifiable artifacts such as a SOC 2 Type II report and a sub-processor list, and a short pilot run on representative data. These three commitments convert marketing claims into enforceable terms. The Federal Trade Commission has signaled that privacy promises carry enforcement weight, and the NIST AI Risk Management Framework gives procurement teams a structure for turning those promises into measurable controls.


TL;DR:

  • Vendors must provide recent, verifiable artifacts like SOC 2 Type II reports and sub-processor lists, with a short pilot run on representative data.
  • Responses should be mapped to the NIST AI Risk Management Framework and FTC guidance, ensuring measurable controls and enforceable contractual terms.
  • Contract clauses should explicitly limit data use, specify exit procedures, and require vendor notification and indemnity measures to mitigate legal and operational risks.
  • Artifacts such as SOC 2 reports and pen test summaries must be recent, scope-specific, and show remediation evidence, not just summary pass/fail statements.
  • A structured, risk-based due diligence process, tied to clear metrics and renewal triggers, provides ongoing oversight aligned with industry standards and regulations.

Ixcommunities
Strengthen Your Talent Leadership Decisions
Join secure peer networking and benchmarking communities where corporate talent leaders share and learn from one another.
Visit Ixcommunities

Table of Contents

A reusable AI vendor due diligence checklist by risk domain

A due diligence questionnaire works best when it is organized by risk domain rather than by department, since vendor answers in one area often reveal gaps in another. Use the following structure as a base for an RFP or vendor questionnaire.

Data handling

  • Does the vendor use customer data, prompts, or outputs to train or fine-tune models, and can that use be disabled contractually?
  • What are the retention, deletion, and export policies for prompts, logs, and generated outputs?

Security

  • Is there a current SOC 2 Type II report, and what exceptions were noted?
  • What is the penetration test cadence, and how are findings remediated?

Privacy and compliance

  • Is a Data Processing Agreement (DPA) available, and does it name all international transfer mechanisms used?
  • What regulatory attestations apply to the vendor's industry and the buyer's industry?

Model governance and behavior

  • Can the vendor document data provenance, bias testing results, and hallucination mitigation steps?
  • Do outputs carry confidence scores or other signals a reviewer can act on?

Operational

  • What SLAs cover uptime, support response, and rollback to a prior model version?
  • How much advance notice precedes a material product or model change?

Commercial and legal

  • Who owns the output, and does the vendor indemnify against claims arising from training-data misuse?
  • What happens to data, logs, and access on termination, and within what timeline?

When evaluating responses, accept artifacts that are recent, specific, and independently verifiable. Treat vague references to "industry-standard security" or undated certifications as unresolved, not satisfied.

  1. Send the questionnaire before commercial negotiation starts.
  2. Score each domain separately rather than averaging across categories.
  3. Flag any domain scoring below threshold for legal and security sign-off before contracting.

Pro Tip: Ask for the sub-processor list and the SOC 2 report in the same email thread. A vendor that stalls on one usually stalls on both.

Mapping vendor answers to NIST AI RMF and FTC expectations

Once questionnaire responses arrive, mapping them to a recognized framework turns a procurement file into audit-ready evidence. The NIST AI Risk Management Framework organizes risk management into four functions, and vendor diligence items map cleanly onto each one:

  • MAP: data provenance, architecture diagrams, and intended use statements establish what the system does and where data flows.
  • MEASURE: benchmark reports, bias metrics, and penetration test results quantify how the system performs against stated claims.
  • MANAGE: SLAs, incident response commitments, and rollback procedures govern ongoing operation after deployment.
  • GOVERN: defined roles, audit rights, and escalation paths establish accountability inside both the vendor and the buying organization.

NIST's framework explicitly calls for supplier risk assessments and contract language addressing third-party components, which gives legal teams a citable basis for requiring these clauses rather than treating them as optional asks.

On the regulatory side, the FTC has stated that companies are accountable for the privacy and confidentiality commitments they make, including promises about not using customer data for model training, and that violations can lead to enforcement remedies such as mandated deletion of models trained on improperly obtained data. That guidance means a vendor's marketing page claim about data use is not sufficient; the same promise needs to appear in the contract with an audit right attached.

Vendor claim mapped to contract audit rights

Technical due diligence protects against operational failure. Contract language protects against the legal and financial fallout when something goes wrong anyway.

  1. Data training and secondary use: require an explicit clause limiting use of customer data to the services purchased, with contractual opt-out from model training as the default position unless separately negotiated.
  2. Exit and portability: require exportable data, metadata, and logs in defined formats, with a handover timeline specified in days, not left open-ended; for critical systems, consider an escrow arrangement for key deliverables.
  3. Incident notification and remediation: require immediate notification for critical security incidents and a short, defined window for serious AI-related incidents, paired with a remediation SLA rather than a vague "commercially reasonable efforts" standard.
  4. Liability and indemnity: allocate responsibility clearly when a vendor relies on a third-party model provider, and require indemnity language covering claims arising from training-data misuse or output infringement.
  5. Audit and inspection rights: secure the right to remote or third-party assessment, with a defined process for handling confidentiality and redaction so the vendor cannot use confidentiality as a blanket refusal.

A legal-focused review of AI vendor risk notes that counsel should treat these clauses as part of the core risk evaluation, not as boilerplate added after commercial terms are settled.

Pro Tip: Negotiate the exit clause before the pilot, not after. Leverage shifts once a system is embedded in daily operations.

Reading security and privacy artifacts correctly

Not every artifact a vendor hands over carries the same weight. A SOC 2 Type II report matters only if it is recent and the exceptions section is reviewed, not just the cover letter. A pen test summary matters only if it states scope and shows remediation evidence, not just a pass or fail conclusion.

  • SOC 2 Type II: confirm the report covers the past twelve months and read the exceptions noted by the auditor.
  • Penetration testing: ask for scope, date, and remediation status rather than accepting a summary statement.
  • Encryption and key management: ask who holds the encryption keys and whether the buyer can rotate or revoke them independently.
  • Multi-tenancy and redaction: ask how tenant data is isolated and what redaction happens before a prompt reaches an external model.
  • Sub-processors: require a current list and a contractual notification or approval step before any change.

Industry guidance for professional-practice due diligence recommends requesting SOC 2 Type II reports, sub-processor lists, and DPAs as a baseline before any production use, treating each new AI feature as a fresh diligence event rather than an extension of a prior approval. When a vendor cannot produce a current sub-processor list on request, that gap alone justifies escalation to security and legal before proceeding.

Designing pilots that actually validate vendor claims

A pilot only tells you something useful if it is scoped before it starts. Define the data scope, the anonymization approach, and the exit criteria in writing, and agree on success metrics with the vendor before data moves.

  1. Select a representative or anonymized dataset that mirrors production conditions, not a cherry-picked sample.
  2. Request benchmark results on accuracy, confidence thresholds, and any available bias metrics, with the sample size stated alongside the figures.
  3. Define where human review is mandatory, so low-confidence outputs route to a person before they reach a customer or a decision.
  4. Set a re-benchmarking schedule after launch, since model behavior can shift after a vendor-side update.

Professional-practice guidance for evaluating AI tools recommends trial periods and validation testing against the buyer's own data before full deployment rather than relying on vendor-supplied demo results. An enterprise evaluation guide for AI platforms makes a similar case for structuring pilot criteria before commercial commitment rather than after.

Pro Tip: Write the pilot's exit criteria before the pilot starts. A pilot with no defined failure condition rarely ends on schedule.

Building AI vendor diligence into the procurement process

Diligence holds up over time only when it is tied to a repeatable process rather than a one-time review. A risk-tiering approach helps route effort to where it matters.

  • Low risk: internal tools with no customer data exposure, light-touch questionnaire, annual review.
  • Medium risk: customer-facing tools with limited data exposure, full questionnaire, security and legal sign-off.
  • High risk: systems processing sensitive or regulated data, full questionnaire plus pilot, executive and privacy sign-off, renewal tied to any sub-processor or model change.
Scoring fieldWhat it capturesWho reviews it
Artifact recencyAge and relevance of SOC 2, pen test, DPASecurity and legal
Test outcomesPilot accuracy, bias, and confidence resultsBusiness owner and security
Contractual protectionsData-use limits, exit terms, indemnityLegal
Operational readinessSLA terms, rollback, support responseProcurement and IT

Set renewal triggers around model updates, sub-processor changes, and any reported incident, not just the contract's calendar date.

Why this checklist holds up under scrutiny

This framework draws on the NIST AI Risk Management Framework and FTC guidance on privacy and confidentiality commitments, both of which set the standard regulators and auditors apply when reviewing vendor oversight programs.

  • Talent and recruiting leaders who want to compare how peer organizations structure vendor governance can draw on benchmarking resources through TLIX Membership, ESIX Membership, and DSIX Membership, which bring recruiting and talent functions together in a vendor-free setting.
  • Framework mappings in this article follow published NIST functions and FTC enforcement positions rather than vendor marketing claims.

A pragmatic view on where to hold the line

We push hardest on data-use limits and exit terms, since those cause the most damage if left vague. Compensating controls, such as added monitoring, can substitute for some contractual guarantees, but never for an undisclosed training-data clause, a missing sub-processor list, or a refusal to allow any audit rights. Those three should stop procurement outright.

— Simon

How our peer communities help you build this capability internally

Evaluating AI vendors well takes practice, and our training courses and peer communities give talent and recruiting leaders a place to build that skill alongside others doing the same work.

Ixcommunities

Explore the IX Academy recruiter training courses to start building this capability on your own team.

FAQ

How is AI used in due diligence?

AI tools can speed up document review and flag anomalies in vendor-supplied artifacts, but the underlying verification of SOC 2 reports, contract terms, and pilot results still requires human review. Legal and procurement teams use AI as a triage step, not a substitute for the checklist itself.

How do you use AI in vendor management?

Vendor management teams use AI to track contract renewal dates, sub-processor changes, and SLA performance against logged data, flagging deviations for human follow-up. The underlying governance decisions, such as whether to renew or escalate, remain with legal, security, and the business owner.

What is included in vendor due diligence?

Vendor due diligence for AI systems typically covers data handling and training-use limits, security artifacts like a SOC 2 Type II report, contractual protections including a DPA and exit terms, and pilot testing against representative data. Domains align with functions in the NIST AI Risk Management Framework.

What are the "Four P's" of due diligence?

Definitions of the "Four P's" vary across industries and are not standardized for AI vendor assessment specifically. A common version used in general vendor diligence covers people, process, performance, and protection, which maps loosely onto governance roles, operational SLAs, benchmark testing, and security controls covered in this checklist.

Sources