You send the security questionnaire. The vendor returns a policy PDF, a model card, and a line that reads "we take AI safety seriously." None of that is evidence. It is a claim about a claim. The decision in front of you is not whether the vendor sounds responsible. It is whether they can produce artifacts that an outsider can inspect and grade, and what you do when they cannot.
This post is about that decision. It leads with how to grade the evidence a vendor can hand you, and it closes on the fallback for the common case where the honest answer is that they have none.
Start with the decision, not the questionnaire
The question worth asking a vendor is narrow: can they show you artifacts, produced by someone external to their build team, that let you judge whether their safety and security claims are accurate? Everything else is preamble.
Most buyer-side AI due diligence gets this backwards. It collects attestations and treats the volume of paper as a proxy for rigor. The research on frontier auditing is blunt about why that fails: public transparency alone cannot close the trust gap, because many safety-relevant details are legitimately confidential and require expert interpretation to read at all (the frontier auditing team makes exactly this point). A stack of self-published documents does not become assurance by getting taller.
The same body of work names the failure mode you are trying to avoid. Vendors grading their own homework is an approach with a checkered track record across industries, and removing that conflict of interest is a central reason to prefer third-party assessment (the frontier auditing paper gives this as one reason among several). A vendor who assesses their own system is not lying to you. They are answering a different question than the one you need answered.
Distinguish an audit from assurance
Two words get used interchangeably in vendor decks, and the difference decides what the evidence covers. The FAccT prototype for third-party assurance separates them along three axes. An audit looks at outcomes, assurance looks at both outcomes and the process that produced them. An audit usually lands after deployment, assurance can run throughout development. An audit is more punitive, assurance is more constructive (the third-party assurance prototype draws this distinction directly).
Why this matters for a buyer: an outcomes-only artifact tells you the model passed a benchmark on a given day. It tells you nothing about whether the process that built it can be trusted to hold up when the data shifts or the team turns over. If a vendor hands you a single evaluation result and calls it assurance, they have handed you an audit and mislabeled it.
The same prototype has four components: a responsibility assignment matrix that maps each stakeholder's level of involvement at each stage of the AI lifecycle, an interview protocol that asks each stakeholder whether they discharged their responsibilities, a maturity matrix that grades both the process and the resulting outcomes against best practice at every stage, and a template for the assurance report itself (the framework spells out these four components). Those are the shapes of real assurance. A model card is not one of them.
Grade the evidence you get
When a vendor does produce artifacts, grade them on independence and depth of access rather than on page count. The frontier auditing framework gives you a vocabulary for this. It defines AI Assurance Levels from AAL-1 to AAL-4, ranging from time-bounded system audits at the low end to continuous, deception-resilient verification at the high end (the AAL scale is defined here). The paper recommends AAL-1 as a baseline for frontier AI generally and AAL-2 as a near-term goal for the most advanced developers (these are the recommended thresholds).
Figure 1: The paper recommends AAL-1 as a baseline for frontier AI and AAL-2 as a near-term goal for the most advanced developers. It states that AAL-3 and AAL-4 are not yet technically or organizationally feasible.
Most vendors you evaluate are not frontier labs, so do not import the labels wholesale. Import the axis. The thing that separates a weak artifact from a strong one is whether the assessor drew on non-public information under secure access, or whether they read the same marketing page you did. The research is clear that even the best current third-party assessments lack the independent, non-public-information access that is standard in established industries (the paper says this plainly).
There is a second axis that vendor evidence routinely fails on: scope. A rigorous evaluation of one model in isolation is not an evaluation of the organization. The frontier auditing work warns against the abstraction error of treating a single component as sufficient to judge overall risk, because risk emerges from the interaction of digital systems, computing hardware, and governance practices, and harm can arise even when a model is never externally deployed (the organizational-perspective argument is here). When a vendor scopes their evidence down to one endpoint, they have narrowed the frame to the part that looks best.
The compliance trap: a framework name is not evidence
Here is where buyers get fooled most often. A vendor points to ISO/IEC 42001 certification or an EU AI Act readiness memo and treats the framework name as the evidence. Those frameworks were built for a different job. The FAccT authors group ISO/IEC 42001 and the NIST AI Risk Management Framework with the other guidance that organizations use internally to govern the systems they develop and use, and they wrote an external assurance framework because internal processes cannot establish the trustworthiness a third party can (the paper draws the internal-versus-external line here). A framework tells an organization what to govern. It does not hand you the proof.
Figure 2: Six of the seven items sit on the internal side, ISO/IEC 42001 and the NIST framework among them. The seventh, the paper's own prototype, is the only one on the external side.
The standards the field is converging on are broad. The same paper reports an emerging convergence around the standards set by the EU AI Act and the NIST framework, then says those standards remain quite broad (the breadth problem is named here). Breadth leaves a vendor room to claim alignment with a framework and still be unable to answer the question you asked.
The last piece of the trap is timing. An assessment describes the system as it stood on the day it ran, and the frontier auditing work names the decay directly: configuration drift, meaning outdated or incomplete audit results caused by system-level changes, where a change as modest as enabling a new tool, relaxing a filter threshold, or swapping in a different safety classifier can materially alter misuse potential (configuration drift is one of at least four ways the paper says abstraction errors arise). A certificate dated last year describes a system that has been reconfigured since. Ask for the date of the underlying evidence, not the date of the certificate.
What counts as real evidence, ranked
Put the artifacts on a ladder from strongest to weakest, and grade what a vendor gives you against it.
At the top: a third-party assessment with secure access to non-public information, covering both process and outcomes at the organization level. That is assurance in the sense the FAccT prototype and the frontier auditing framework both mean it. Almost no vendor in your pipeline will clear this bar today.
Below that: a claims-based assurance case, where the vendor states explicit assurance claims and backs each with evidence, from training-data provenance to test reports. The DOD framework for AI-enabled systems builds exactly this structure, capturing evidence for each claim and requiring an assurance plan that defines how assurance is maintained over the system's lifecycle (the claims-based framework is described here). A vendor who can show you a claim, the evidence under it, and the plan to keep it true is far ahead of one who cannot.
At the bottom: model cards and factsheets. The FAccT authors list model cards among the responsible-AI tools that organizations use internally, which is the gap their external assurance framework exists to fill (the paper puts these tools on the internal side of the line). Treat them as documentation, not assurance.
When the honest answer is none
Most of the time, on most vendors, the honest answer to your evidence question is that they have neither of the top two. This is the case the frameworks quietly admit. Third-party AI auditing is a proposed vision, not a market you can buy from today. The frontier auditing paper itself calls for rapid growth in auditor capacity and credible oversight of auditors as things that do not yet exist at scale (the paper lists these as open requirements). Waiting for AAL-2 evidence from a mid-market vendor means waiting past your deployment date.
The fallback is not to reject the vendor. It is to manufacture the evidence you cannot buy, and to hold the risk explicitly rather than pretend it is covered.
First, run your own version of the interview protocol. The FAccT prototype's questions are designed for exactly this: to ask each stakeholder whether they discharged their responsibilities at each lifecycle stage (the interview protocol is one of the four components). You do not need the vendor to have hired an auditor. You need their engineers, their data owners, and their governance lead to answer structured questions, on the record, so gaps surface.
Second, write the risk down as an accepted risk with a named owner, not as a resolved control. The DOD framework centers on explicit assurance claims, and backs them with tangible completion criteria, explicit approval gates, and an assurance plan maintained over the system's lifecycle (the framework's organizing idea is described here). Borrow the discipline. If you cannot get evidence, the compensating artifact is a signed risk acceptance that says who owns the exposure and what would trigger a reassessment.
Third, make freshness a contractual term. Duration is one of the dimensions the AAL scale uses to separate its levels, running from a time-bounded assessment of a few weeks at AAL-1 to continuous verification at AAL-4 (the level details set out duration and access side by side). You will not buy continuous verification from a mid-market vendor, so buy the closest thing you can enforce. Require the vendor to regenerate evidence on a cadence tied to retraining, and keep the right to re-ask.
The one-line rule
Grade the artifact by who produced it and what they could see, not by how much it weighs. When the artifact does not exist, do not launder the gap into a green checkbox. Name it, own it, and put a date on when you will look again. A vendor with no evidence and an honest risk acceptance is a safer counterparty than a vendor with a thick binder and a six-month-old certificate.
