A control that reports is not a control that works. The two look identical on a slide. Both show a green tile, a policy number, a coverage percentage. The difference shows up only when something goes wrong, and by then the slide is no longer the artifact anyone cares about. This post is about how a board tells the two apart before the incident, not after.
The distinction matters more with AI than with the systems that came before it. A firewall rule either passes traffic or blocks it, and you can test the rule directly. An AI control sits on top of a system whose behavior shifts with its inputs, its data, and the social context it operates in. NIST makes this point plainly: AI systems are socio-technical, and their risks emerge from the interplay of technical behavior with how a system is used and who operates it, so what you measure depends on the purpose and audience of the evaluation (NIST AI RMF Measure playbook). A control that was effective at commissioning can quietly stop working under data drift or model drift without anyone touching it.
Start with the question, not the dashboard
The first thing a board should ask is not "is the control in place" but "how do we know it is doing anything." Those are different questions, and most reporting answers the first while pretending to answer the second.
Here is the shape of what a board should demand, in order.
First, what risk is this control supposed to reduce, and by how much. A control with no target has no way to fail, which means it has no way to succeed either. The information security discipline has known this for a long time: you run a risk analysis to determine which controls to implement, then define the critical variables that show your exposure level, then set the metrics that measure whether the control moves those variables (SANS, Measuring effectiveness in Information Security Controls). AI does not change that chain. It only makes each link harder to observe.
Second, what would we see if the control were failing. If the answer is "the same dashboard, still green," the control reports and does not work. A real control has a failure signal that is distinct from its success signal. NIST frames this as defining acceptable limits for system performance and building in course correction for when the system performs beyond those limits (NIST AI RMF Measure playbook).
Third, when did we last test that the signal fires. A control whose alarm has never been tripped in a controlled exercise is a control whose alarm you are trusting on faith.
Fourth, does the measurement survive contact with the real deployment, or was it taken in a lab. This is the gap that gets organizations hurt, and it is worth its own section.
Figure 1: A credible effectiveness claim decomposes into distinct measurements. Four of these are dimension scores and one is a risk reduction against a baseline (source: Governance Controls for AI-Generated Test Artifacts).
The reality gap is where reporting hides
The reason a reporting control can look like a working control for months is that most AI measurement happens inside the stack. Benchmarks measure abstract capability. MLOps dashboards measure system stability. Neither tells a decision-maker outside the AI stack how the system behaves under real user variability and real constraints. The CIRCLE framework names this the reality gap between model-centric performance metrics and the outcomes that materialize in deployment (CIRCLE, Evaluating AI from a Real-World Lens). A control validated only against a benchmark is a control validated against a proxy for the thing you care about.
The advertising study makes the gap concrete. Researchers ran a large-scale audit of Meta's "See less" ad control, randomly assigning participants to mark either Body Weight Control or Parenting topics as unwanted, then measuring the ads delivered before and after. Participants who marked "See less" did not see a significant decrease in the proportion of ads on those topics over time (Why am I Still Seeing This). The control existed. It had a button, a setting, a stored preference. It reported that the user had expressed a choice. It did not reduce the risk it appeared to govern. That is the entire distinction in one experiment, and the only reason anyone knows is that someone measured the outcome rather than the button.
The same study found something a board should sit with: the majority of targeting explanations for local ads made no reference to location-specific criteria, and did not explain why the flagged topics kept appearing (Why am I Still Seeing This). The reporting layer and the control layer had drifted apart. The system told a story about itself that its own behavior did not support. A control that explains itself well and works poorly is more dangerous than one that does neither, because the explanation buys it trust it has not earned.
What a good answer looks like
When you ask whether a control works, a good answer has a particular texture. It does not lead with coverage. It leads with a measured effect on a named risk, and it tells you how the measurement was taken.
A good answer distinguishes what the model does from what happens downstream. CIRCLE separates primary effects, the immediate outputs of a model, from secondary effects, the near-term impacts those outputs produce in everyday use, and tertiary effects on productivity and return over longer horizons (CIRCLE, Evaluating AI from a Real-World Lens). Most control reporting stops at the primary effect because that is the easiest thing to measure. The risk usually lives in the secondary and tertiary layers, where over-reliance and cumulative interaction patterns show up, none of which come from any single output (CIRCLE, Evaluating AI from a Real-World Lens).
A good answer names its own blind spots. NIST is explicit that risks or trustworthiness characteristics that will not or cannot be measured should be documented, with justification for why they are not measured (NIST AI RMF Measure playbook). A control owner who can tell you what they are not measuring, and why, is telling you the truth. One who reports full coverage is either lucky or hiding the gaps.
A good answer shows the measurement has a refresh cadence tied to change in the system. NIST calls for assessing the effectiveness of existing metrics and controls on a regular basis across the lifecycle, and for developing new metrics when the existing ones prove insufficient (NIST AI RMF Measure playbook). Data drift and model drift are named reasons the appropriateness of a metric decays. A number measured once at commissioning and never again is a number that describes a system that no longer exists.
Figure 2: Standard reporting answers one of the four board questions this post asks and leaves the other three unanswered.
Effectiveness is not one number
Boards want a single score, and there is not one. The evidence a control works is assembled from several kinds of measurement that check each other. Governance research on AI-generated test artifacts reports the shape of this: a governance-aware framework was evaluated across governance accuracy, artifact reliability, compliance accuracy, and explainability, not a single headline figure, and it reduced governance-related risks by a measured amount against a conventional baseline (Governance Controls for AI-Generated Test Artifacts). The lesson for a board is not the specific numbers. It is that a credible claim of effectiveness decomposes into distinct measured dimensions, and each one can fail on its own.
Where the control governs a human working with the AI rather than the model alone, the measurement gets harder still. Evaluating human-AI collaboration cannot rely on task performance and response time alone, because the effectiveness of the pairing depends on interaction quality, trust, and adaptability, which traditional metrics miss (Evaluating Human-AI Collaboration). A control that assumes a human reviewer will catch the model's error is only as good as that human's actual behavior under load, and that is a measurable thing that most control reporting never measures.
The board's standing questions
Four questions separate a working control from a reporting one, and three of them are the ones control owners least want to answer.
What risk does this reduce, and by how much. What would failure look like on this same dashboard. When did we last test that the failure signal fires. Was this measured in deployment or in a lab. The owner who can answer all four has a control. The owner who answers the first and stalls on the rest has a report.
NIST itself concedes the hardest part. Its own framework says that ways to measure bottom-line improvements in the trustworthiness of AI systems are still future work, to be developed with the community (NIST AI RMF effectiveness). The people who wrote the standard are honest that the measurement problem is unsolved. A vendor or an internal team that claims to have solved it deserves more scrutiny, not less. Ask them to show you the deployment measurement, the failure signal, and the last date it was tested. If those exist, you have a control. If they do not, you have a slide.
