# AI Adoption Capability Check

### Diagnostic

Distinguish AI access and activity from usable workforce capability.

---

## Purpose

Tool licenses issued, logins recorded, and courses completed measure exposure to AI, not the ability to use it well. This diagnostic separates the two. Use it to find out whether your team can actually produce reliable work with AI, not just whether they have accounts.

---

## What Capability Is Not

Before running this check, rule out the three most common false signals of readiness:

- **Access is not capability.** A seat license or admin-granted tool account shows procurement happened, not that anyone can use the tool on real work.
- **Activity is not capability.** Prompt volume, login frequency, or chat history length show engagement, not quality of judgment.
- **Course completion is not capability.** A training certificate shows exposure to concepts at one point in time. It does not show whether the person applies those concepts under real deadline pressure, on real client or company material.

If your current adoption metrics are limited to these three, you do not yet have a capability measurement. You have an activity log.

---

## The Six Capability Behaviors

A team has usable AI capability when people can demonstrate all six of the following, in real work, without prompting from someone else.

### 1. Select an approved tool for the task

What to look for: the person chooses a tool appropriate to the task's data sensitivity, complexity, and output requirements, from an approved list, without defaulting to whatever is open.

Red flag: the same tool is used for every task regardless of fit, or unapproved tools are used because no one checked.

### 2. Provide the context and constraints needed for reliable output

What to look for: inputs to the AI system include the relevant background, audience, format requirements, and known constraints, not a bare instruction.

Red flag: output quality is inconsistent and traced back to thin or missing input context, not the tool itself.

### 3. Detect weak, inaccurate, or unsafe output

What to look for: the person can identify when an output is wrong, incomplete, biased, fabricated, or unsafe to use, without being told by a reviewer after the fact.

Red flag: errors are caught only downstream, by a manager, client, or customer, rather than by the person who generated the output.

### 4. Apply the defined quality standard before handing work over

What to look for: a known, documented standard exists for the task, and the person checks output against it before passing work along.

Red flag: no documented standard exists, or the person cannot describe what they checked for.

### 5. Know when to stop, escalate, or seek human approval

What to look for: the person recognizes tasks or outputs that exceed their authority or the tool's reliability, and routes those to a human reviewer instead of proceeding.

Red flag: the person proceeds on judgment calls, sensitive data, or client-facing output without knowing that escalation was an option.

### 6. Complete the workflow without relying on one expert to correct every result

What to look for: the workflow functions when the single most AI-fluent person on the team is unavailable.

Red flag: one person is silently the bottleneck who fixes everyone else's AI output, and the workflow has no documented path without them.

---

## How to Measure This

Do not measure these behaviors through self-report surveys or training quizzes. Measure them in real work, using one or more of the following methods:

- **Work sample review.** Pull five to ten recent AI-assisted outputs from the person or team. Score each against the six behaviors above.
- **Shadowed task.** Observe one real task from tool selection through handoff. Note where the person skipped a step or asked someone else to make the call.
- **Incident trace.** For any recent AI-related error or rework, trace which of the six behaviors broke down. This is often more revealing than a clean sample.

[UNVERIFIED] Any specific pass rate, benchmark score, or industry baseline for these behaviors should come from your own measured samples, not an assumed target. Set your own baseline from the first round of measurement before comparing future rounds against it.

---

## Capability Rating

Score each of the six behaviors per person or per team, using a consistent scale:

| Rating | Definition |
|---|---|
| 0 | Behavior not observed in the sample |
| 1 | Behavior observed inconsistently, dependent on task or reminder |
| 2 | Behavior observed consistently across the sample, without prompting |

A total score of 10 to 12 indicates usable independent capability. A score of 5 to 9 indicates partial capability, with specific gaps to close. A score below 5 indicates the person or team is operating on access and activity alone and should not yet be relied on for unsupervised AI-assisted work.

---

## Action by Outcome

- **Usable capability (10 to 12):** Confirm the workflow can run without the single most fluent expert present. Document this person's practice as a reference example.
- **Partial capability (5 to 9):** Identify the specific behaviors scoring 0 or 1 and target them directly. Generic refresher training rarely closes a specific behavioral gap. Pair the person with a reviewer on the exact behavior missing, not the whole workflow.
- **Access and activity only (below 5):** Do not expand this person's or team's unsupervised scope. Keep a mandatory human review checkpoint on all output until a follow-up measurement shows improvement.

---

## Common Mistakes When Running This Check

- Measuring at the moment of training completion instead of weeks later, in live work
- Asking people to self-assess instead of reviewing actual output
- Treating one strong work sample as proof of consistent capability
- Conflating enthusiasm for the tool with skill in using it
- Rolling out wider AI access based on capability shown by one person, without checking whether the rest of the team shows the same behaviors

---

## Notes and Assumptions

[ASSUMPTION] This check assumes your organization already has an approved-tool list and at least one documented quality standard for the task being measured. If neither exists, establish those first. Behaviors 1 and 4 cannot be measured against a standard that does not exist.

---

### Related WenceStudio Resources

Turn observable workplace behaviors into repeatable operational capability:

1. **Useful Work Per Dollar: Custom-Agent ROI** (Executive Report & ROI Model)
   A 24-page research guide on measuring completed business outcomes, unit economics, dependability rates, and return on compute.
   - [Read ROI Whitepaper (PDF)](https://drive.google.com/file/d/1St1yw4nQ2EaZD_923UuiE28lbj4wdW-N/view?usp=sharing)

2. **Prompt Shortcuts Playbook: 72 Command Codes** (Field Reference & Playbook)
   A master catalog of 72 slash-style command shortcuts across 11 categories to structure persona, strategy, and editing passes without prompt drift.
   - [Open Shortcuts Playbook](https://docs.google.com/document/d/1bGqhUjd5DZsGtaV_IrCG700kel3LbpiaD0AkJb_cUS4/edit?usp=sharing)

3. **Model Context Protocol (MCP) Server Starter Pack** (Architecture Guide)
   Production server implementations for Claude Desktop and Cursor. Connect frontier models to private databases and file tools with Docker configs.
   - [Open Architecture Guide](https://docs.google.com/document/d/1ChPsXZ41yI4Jm8Z7vOCuE4G1-yTk5qD0-_Na2FiHnx8/edit?usp=sharing)
