Operational & Evidence Signals
Cannot define alert rate or secondary review workload from a real deployment
Alert rate drives your operational cost. No answer = no comparable deployment.
Performance claims are entirely self-reported from controlled demo conditions
Require DHS MSR data, third-party validation, or documented comparable deployment.
Uses 'best in class' or 'industry-leading' without defining the measurement
Ask: compared to what, by whom, under what conditions? No answer = marketing only.
Refuses to pilot under realistic conditions — actual volume, real benign scenarios
Any vendor who won't test under your conditions is hiding performance data.
Pilot success criteria are not defined in writing before go-live
Post-hoc success criteria are vendor-defined success criteria. Define them now.
Modifies model or detection parameters during the pilot without prior written notice
Any change during a pilot invalidates your baseline data. This is a disqualifier.
Governance & Contract Signals
Cannot explain how the AI makes decisions in plain language
Black-box answers create liability you own. You must explain this in any incident review.
No documented model update protocol — changes happen without customer notice
Undocumented model changes mean your operational risk is entirely unmanaged.
Contract language does not match demo representations
If specific metrics are replaced with 'best efforts,' those numbers no longer exist.
No contractual responsibility for false positive operational burden
If the contract is silent on who bears this cost, you do — by default.
Audit logs are not customer-accessible or not in litigation-ready format
You must produce complete, tamper-evident logs within 24 hours of an incident inquiry.
No healthcare-specific deployment references
Sports venues and courthouses are not comparable. References must be hospital deployments.