Thanks for joining the session

From the IAHSS 2026 session: AI in Healthcare Physical Security. Here's everything you need to keep the conversation going and start putting it to work.

AI Physical Security Vendor Scorecard

When an AI tool fails in a healthcare setting, everyone points at someone else — and patients and regulators don't care. Six questions every compliance, legal, and IT leader should be able to answer before going live with an AI deployment.

Logo for the 2026 IAHS Annual Conference & Exhibition with colorful overlapping shapes in purple, green, orange, and yellow.

Or use the interactive version below

AI Physical Security Vendor Scorecard — JHarris Advisory
Total Score
0 / 40
⚑ RED FLAG
A Claims
0
/8
B Ops
0
/8
C Gov
0
/8
D Integ
0
/6
E Pilot
0
/10

Why This Matters

81.6%Nurses reporting WPV
in the past year (2024 survey)
4–5×Higher WPV risk vs. any
other private industry (OSHA)
38+States with active hospital
WPV bills introduced or passed
  • AI security vendors routinely make performance claims that don't hold in real hospital conditions.
  • California, Ohio, and Texas now require written security plans — and more states are moving fast.
  • A bad AI deployment isn't just a budget problem. It's liability, staff trust, and patient safety.

How to Use This Guide

1. Score each criterion 0–2. Require written documentation for any score above 0.
2. Record the vendor's exact words in the Notes field — you'll need them in contract negotiations.
3. Total the score. Below 30 = significant gaps. Below 20 = do not proceed to contract.
4. Run the Red Flags list in real time during demos. Any hit = pause the evaluation.
0–19
DO NOT
PROCEED
20–29
SIGNIFICANT
GAPS
30–34
CONDITIONAL
35–39
ADVANCE
TO PILOT
40
STRONG
CONTENDER

Evaluation Details

Score Key:
0
Cannot demonstrate
1
Partial or verbal only
2
Fully documented, comparable conditions
⚑ Require written evidence for any score above 0
A
Claims & Evidence
8 points possible — 4 criteria
0 / 8
A1 Performance data is specific and independently verified — not self-reported from a controlled demo. Vendor identifies source, conditions, and date of all claims.
/2
A2 Alert rate (alerts per 1,000 entrants or events) is defined in writing from a comparable deployment. Industry averages are not a substitute.
/2
A3 Detection or event capture rates are broken down by threat type or scenario — not reported as a single aggregate accuracy figure.
/2
A4 All performance claims are tied to conditions comparable to your facility: patient volume, layout, staff presence, and device interference.
/2
B1 Secondary review rate — % of alerts requiring human action — is documented from a live comparable deployment, not from a demo environment.
/2
B2 Staffing model is explicit: how many personnel, what role, what shift coverage is required to operate this system at your facility's projected volume.
/2
B3 Throughput impact is measured: vendor provides queue time and wait data from a comparable facility at peak volume, not average daily conditions.
/2
B4 Full alarm workflow is defined: who receives the alert, what action follows, how clinical care or facility access continues through resolution.
/2
C1 Vendor explains how the AI engine makes decisions in plain language. No black-box answers. You must be able to explain this in any incident or legal review.
/2
C2 Training data provenance is disclosed: what data was used, how it was labeled, what environments and demographics it represents.
/2
C3 Model update protocol is documented: who authorizes changes, how customers are notified in advance, and what testing occurs before any production push.
/2
C4 Audit logs are customer-accessible, exportable, and tamper-evident — covering all alerts, human overrides, model changes, and tuning adjustments.
/2
D1 Integration with existing VMS, access control, and incident management systems is documented via API or certified connector — not a future roadmap item.
/2
D2 Incident workflow is fully mapped: how alerts connect to dispatch, documentation, post-incident review, and regulatory reporting obligations.
/2
D3 HIPAA and privacy obligations are assessed: what data is collected, who accesses it, how retention is managed, and PHI exposure is defined in writing.
/2
E1 Vendor provides a written pilot test plan with specific pass/fail thresholds defined before you commit to full deployment — not after.
/2
E2 Pilot is conducted under realistic conditions: actual patient volume, realistic benign scenarios, your specific environmental constraints.
/2
E3 Performance representations in the contract match what was demonstrated. No 'best efforts' language substitutes for specific metrics.
/2
E4 Change management clause: no model update or tuning change without 14-day prior written notice. 'Material change' is defined quantitatively in the contract.
/2
E5 Healthcare-specific deployment references — not venues or airports — are provided and contactable before you sign.
/2
0 / 40
Not Yet Scored
Score all criteria to see your recommendation.
Red Flag Register — Any of These = Pause or Walk
12 disqualifying signals. These are not concerns to monitor. Document each and share with legal before proceeding.
0
Operational & Evidence Signals
Cannot define alert rate or secondary review workload from a real deployment
Alert rate drives your operational cost. No answer = no comparable deployment.
Performance claims are entirely self-reported from controlled demo conditions
Require DHS MSR data, third-party validation, or documented comparable deployment.
Uses 'best in class' or 'industry-leading' without defining the measurement
Ask: compared to what, by whom, under what conditions? No answer = marketing only.
Refuses to pilot under realistic conditions — actual volume, real benign scenarios
Any vendor who won't test under your conditions is hiding performance data.
Pilot success criteria are not defined in writing before go-live
Post-hoc success criteria are vendor-defined success criteria. Define them now.
Modifies model or detection parameters during the pilot without prior written notice
Any change during a pilot invalidates your baseline data. This is a disqualifier.
Governance & Contract Signals
Cannot explain how the AI makes decisions in plain language
Black-box answers create liability you own. You must explain this in any incident review.
No documented model update protocol — changes happen without customer notice
Undocumented model changes mean your operational risk is entirely unmanaged.
Contract language does not match demo representations
If specific metrics are replaced with 'best efforts,' those numbers no longer exist.
No contractual responsibility for false positive operational burden
If the contract is silent on who bears this cost, you do — by default.
Audit logs are not customer-accessible or not in litigation-ready format
You must produce complete, tamper-evident logs within 24 hours of an incident inquiry.
No healthcare-specific deployment references
Sports venues and courthouses are not comparable. References must be hospital deployments.
§
Require All of These in Writing Before Signing
6 non-negotiable contract provisions
Performance Representations
Pilot alert rates and detection benchmarks are incorporated by reference. No 'best efforts' language replaces specific metrics.
Change Management
No model update, tuning change, or hardware modification without 14-day prior written notice. 'Material change' is defined: >5% shift in alert or detection rate.
Audit Log Access
Customer has direct, perpetual access to exportable, tamper-evident logs: all alerts, overrides, model changes, and tuning adjustments.
False Positive Liability
Vendor bears cost of secondary review workload exceeding the contractual baseline rate by more than 15%.
Post-Incident Support
Full technical support for legal holds and regulatory inquiries within 24 hours. Data export in customer-specified format required.
Healthcare References
All performance representations are warranted to be based on comparable hospital deployments — not venues, airports, or courthouses.
📋
Applicable to any AI security deployment
Negotiate and document all six before pilot starts — not after
Alert / Capture Rate
Alerts per 1,000 entrants or events — define acceptable range before pilot starts.
Secondary Review Rate
% of alerts requiring human action — your staffing cost baseline.
Time-to-Clear
Minutes from alert to resolution — measure peak and off-peak separately.
Throughput Impact
Wait time delta vs. pre-deployment baseline — measured at peak volume.
False Positive Breakdown
By category: benign carry, device interference, environmental — not aggregate.
Staff Burden Hours
Hours per shift on alert response — vendor projections must match pilot actuals.

Want to go deeper?

Whether you're evaluating a specific vendor, building a governance framework, or figuring out where to start, JHarris Advisory is happy to talk through what you're facing.