Operational AI Evaluation Kit

Evaluate the evidence, not the demo.

The Operational AI Evaluation Kit is a vendor-neutral, 21-point scorecard for healthcare operations teams. It turns seven high-stakes decision areas into 28 evidence checks, a consistent score, clear gaps, and an exportable review.

Version 1.0Published September 2, 2026Prepared by Linear Health
Method

A score is useful only when the evidence standard is shared.

Score the evidence you can inspect, not the confidence of the presentation. Adapt the rubric to your risk, procurement, legal, privacy, security, and clinical governance requirements.

0

Not addressed

No clear answer or usable evidence was provided.

1

Claimed

The answer is stated, but the evidence cannot yet be checked.

2

Demonstrated

The capability works in a controlled scenario or pilot.

3

Verified

The evidence is observable in production and can be independently checked.

  1. 01

    Run a real workflow

    Use one representative case with normal complexity, missing data, and at least one exception. A scripted happy path is not enough.

  2. 02

    Score only what you can inspect

    Move from claimed to demonstrated or verified only when the evidence is visible, attributable, and relevant to your environment.

  3. 03

    Decide with gates and gaps

    Record unresolved evidence, assign owners, and treat privacy, security, safety, and authority issues as gates that a high total cannot offset.

Interactive scorecard

One review. Seven decisions. Explicit gates.

Check the evidence you received, score each dimension from 0 to 3, and record the gap. Export a CSV for the decision record or print the completed worksheet.

Review details

Name the decision record.

Optional. This tool does not save a decision record. Do not enter PHI or confidential information.

01
Awaiting score

Workflow completion

Does the system complete the operational job, or only create another task?

A polished interaction is not the same as a completed workflow. Evaluate the full path from intake to a confirmed operational outcome.

Evidence checks0 of 4
Evidence to request
  • Timestamped workflow trace
  • Written in-scope and out-of-scope list
  • Exception queue with ownership
  • Completion definition and audit record
Score this dimension
02
Awaiting score

Integration and write-back

Can the workflow act safely inside the systems your team already uses?

Integration quality determines whether automation removes work or shifts it into reconciliation, duplicate entry, and exception cleanup.

Evidence checks0 of 4
Evidence to request
  • Production integration architecture
  • Field-level read and write map
  • Failure and retry log
  • Audit-history example
Score this dimension
03
Awaiting scoreRequired gate

Human oversight and boundaries

Can people see, pause, correct, and escalate the system at the right moments?

Operational AI should make ownership clearer. Review who approves sensitive actions, who receives exceptions, and how corrections are recorded.

Evidence checks0 of 4
Evidence to request
  • Human oversight map
  • Role and permission matrix
  • Escalation service levels
  • Correction and rollback history
Score this dimension
04
Awaiting scoreRequired gate

Privacy and security

Are data responsibilities, safeguards, access, and retention explicit before sensitive data is used?

Healthcare buyers need a documented data flow and accountable controls. Treat this dimension as a review gate, not a point total to average away.

Evidence checks0 of 4
Evidence to request
  • Data-flow diagram
  • Applicable agreement and BAA materials
  • Security control documentation
  • Retention, deletion, and incident procedures
Score this dimension
05
Awaiting score

Implementation readiness

Is there a credible path from signed agreement to a stable operating workflow?

Implementation risk often sits outside the interface. Surface dependencies, owners, acceptance criteria, training, and support before selection.

Evidence checks0 of 4
Evidence to request
  • Implementation plan with owners
  • Dependency register
  • Acceptance and rollback criteria
  • Training and support plan
Score this dimension
06
Awaiting score

Measurement and economics

Can you tell whether the workflow improved, for whom, and at what cost?

A percentage without a baseline, denominator, timeframe, and scope is not decision-grade evidence. Agree on measurement before the intervention starts.

Evidence checks0 of 4
Evidence to request
  • Baseline extract and data dictionary
  • Outcome measurement plan
  • Raw counts and denominators
  • Total-cost model with assumptions
Score this dimension
07
Awaiting score

Commercial terms and exit

Are the operating commitments and the path out as clear as the path in?

Commercial fit includes usage growth, service obligations, data portability, transition support, and what happens when the relationship ends.

Evidence checks0 of 4
Evidence to request
  • Order form and usage model
  • Service-level and support terms
  • Sample data export
  • Termination and deletion procedure
Score this dimension
Scope and limitation

This tool supports operational vendor evaluation. It is not legal, regulatory, privacy, security, clinical, or procurement advice. It does not certify compliance, product quality, or fitness for a specific use. Your accountable reviewers must verify evidence and make the final decision.

Stay updated

Get the latest on AI healthcare coordination.