Assistant pilot research · Research
Designing acceptance criteria for a virtual assistant pilot
A prospective method for deciding what counts as ready, returned, escalated, or accepted before a staffing pilot begins.
Headline statistic
One declared buyer decision, one traceable observation unit, and zero assumed outcomes.
Methodology: Structured desk review of five named primary or official sources, checked September 22, 2026, followed by a proposed local decision protocol. Research question: What observable evidence will let the buyer decide whether a bounded assistant pilot should continue, narrow, pause, or stop? Unit of analysis: one eligible pilot item with its input state, expected output, acceptance checks, reviewer decision, return reason, exception state, elapsed states, and final disposition. The method separates retained facts, analysis, inference, and uncertainty. It has not been applied to private client outcomes and makes no universal claim about price, savings, performance, location, classification, or business results.
Key stats
- Decision: whether the tested task lane meets its declared finish conditions under the access, review, and exception controls used during the pilot.
- Observation unit: one eligible pilot item with its input state, expected output, acceptance checks, reviewer decision, return reason, exception state, elapsed states, and final disposition.
- Evidence base: five named primary or official sources with URLs and checked dates.
Key takeaways
- A pilot can test a specified lane under stated conditions. It cannot prove universal assistant performance, justify unrelated access, establish a permanent staffing ratio, or guarantee future results.
- Blind-review a small but representative set against the same rubric, retain disagreements and exclusions, then apply the written go, narrow, pause, or stop rule without rewriting it after outcomes appear.
- Accountable owner: the buyer’s named pilot reviewer who has authority to accept work and change or stop the lane.
Define the decision before collecting convenient numbers
What observable evidence will let the buyer decide whether a bounded assistant pilot should continue, narrow, pause, or stop?
The decision in scope is whether the tested task lane meets its declared finish conditions under the access, review, and exception controls used during the pilot. Write that decision, its owner, and the date it must be made before asking for metrics. This prevents a familiar reversal in which an attractive number appears first and the team invents a question it seems to answer. A provider comparison, pilot score, coverage test, or cost model is useful only when it changes a named choice.
Use one eligible pilot item with its input state, expected output, acceptance checks, reviewer decision, return reason, exception state, elapsed states, and final disposition as the observation unit. Keep the original record beside any category or score. A ticket, spreadsheet row, calendar event, quote, or interview answer is a source; it becomes decision evidence only when its definition, date, scope, provenance, and relationship to the buyer’s question are recorded.
O*NET lists varied tasks and work contexts for administrative occupations. That breadth is a discovery aid, not a ready-made role for this buyer. The SBA likewise places hiring among wider management, finance, compliance, cybersecurity, and continuity responsibilities. The buyer still has to define the actual lane and its limits.
| Item | Finding | Source note |
|---|---|---|
| Buyer decision | whether the tested task lane meets its declared finish conditions under the access, review, and exception controls used during the pilot | Pre-specified local protocol |
| Observation unit | one eligible pilot item with its input state, expected output, acceptance checks, reviewer decision, return reason, exception state, elapsed states, and final disposition | Pre-specified local protocol |
Assemble evidence that another reviewer can reconstruct
The minimum evidence set is versioned examples, eligibility rules, a scoring guide, all pilot items including open and cancelled work, independent review notes, corrections, exceptions, and the pre-committed decision rule. Use consecutive or otherwise reproducibly selected records from a declared observation window. Retain normal, difficult, cancelled, returned, waiting, and still-open cases when they satisfy the eligibility rule. Record every exclusion with its reason and approver.
Label where each field came from: system event, signed document, provider response, manager note, participant recollection, or later reconstruction. Preserve unknown values as unknown. Missing review time is not zero; an absent exception note is not proof that no exception occurred; a sales statement is not an implemented control.
GAO frames data reliability in relation to the intended use. Apply that principle field by field. A rough task count might support early discovery but be inadequate for a staffing schedule. A current quote might be precise but incomplete if it excludes tools, management, or exit work. State which decisions the evidence can and cannot support.
| Item | Finding | Source note |
|---|---|---|
| Evidence set | versioned examples, eligibility rules, a scoring guide, all pilot items including open and cancelled work, independent review notes, corrections, exceptions, and the pre-committed decision rule | Local records and authoritative-source review |
| Reliability rule | Assess each field against its intended decision use | U.S. GAO data-reliability guidance |
Retain variation instead of averaging it away
Important sources of variation are input completeness, task class, consequence, reviewer, correction type, customer visibility, waiting dependencies, novelty, and whether the item followed the normal path. Declare these dimensions before inspecting outcomes. Report counts, ranges, and distributions where the sample supports them; otherwise show the individual cases. An average that hides peaks, exceptions, open work, or unlike tasks can create false confidence.
Separate arrival, active preparation, waiting, owner review, correction, escalation, acceptance, cancellation, and closure. These states represent different resource demands. Waiting is not active labour. Escalation can be correct performance. A reopened item may reflect new information rather than an earlier defect. Preserve the state history before interpreting it.
Compare like with like. Hold the task lane, finish condition, observation period, decision rights, and service level constant before comparing options. When those conditions differ, show the difference as part of the result instead of forcing a single rank. Sensitivity cases are more honest than a precise answer built from unstable assumptions.
| Item | Finding | Source note |
|---|---|---|
| Variation to retain | input completeness, task class, consequence, reviewer, correction type, customer visibility, waiting dependencies, novelty, and whether the item followed the normal path | Niche-specific study design |
| Comparison rule | Normalize the lane or disclose the material difference | Local analysis protocol |
Map responsibility and access to the work
The accountable owner is the buyer’s named pilot reviewer who has authority to accept work and change or stop the lane. Record who prepares, recommends, approves, acts, verifies, receives an exception, and removes access. One person may hold several roles, but the responsibilities should remain distinct so a tool permission or job title does not silently become approval authority.
NIST CSF 2.0 treats governance, roles, policy, oversight, and supply-chain risk as parts of risk management. NIST SP 800-53 provides more detailed concepts for account management, least privilege, separation of duties, logging, external services, and contingency. Neither source selects a provider or staffing model; both support explicit and reviewable responsibility.
Connect each permission to a current task, resource, approved action, business purpose, owner, evidence threshold, review point, and removal trigger. Keep money movement, account ownership changes, legal or regulated judgment, sensitive personnel action, broad data export, and customer commitments on the specifically authorised path.
| Item | Finding | Source note |
|---|---|---|
| Accountable owner | the buyer’s named pilot reviewer who has authority to accept work and change or stop the lane | Buyer governance record |
| Access rule | Task-specific, least-privilege, approved, logged, reviewed, and removable | NIST CSF 2.0 and SP 800-53 |
Run a bounded test with pre-committed outcomes
Blind-review a small but representative set against the same rubric, retain disagreements and exclusions, then apply the written go, narrow, pause, or stop rule without rewriting it after outcomes appear. Define eligibility, start state, finish condition, review sample, exception route, stop rule, and end point before live work begins. The test should expose uncertainty while limiting consequence; it should not be used to imply a production guarantee.
Use realistic but safe records. Minimise or mask personal and confidential information when the decision does not require it. Have reviewers apply the declared rule independently where feasible, then retain their original decisions and the reason for disagreement. If the rule cannot be applied consistently, revise the rule before increasing volume or access.
Pre-commit to proceed, narrow, pause, and stop states. Proceed means only that the tested lane may continue under the tested controls. Narrow when one task class is ready and another is not. Pause when a recoverable dependency has a named owner and review date. Stop when the safe boundary is crossed or reliable evaluation is unavailable.
| Item | Finding | Source note |
|---|---|---|
| Bounded test | Blind-review a small but representative set against the same rubric, retain disagreements and exclusions, then apply the written go, narrow, pause, or stop rule without rewriting it after outcomes appear. | Prospective local protocol |
| Decision states | Proceed, narrow, pause, or stop with evidence and owner | Buyer decision record |
Separate facts, analysis, inference, and uncertainty
A central distortion risk is measuring speed without acceptance, dropping returned or unfinished items, allowing coaching to erase the original result, or expanding scope because a few easy cases passed. Counter it by preserving the eligible population, original records, criteria, exclusions, missing fields, reviewer disagreements, corrections, and changes in operating conditions. Do not improve the apparent result by redefining success after outcomes appear.
Facts are retained events, documents, and source statements. Analysis applies declared definitions to those facts. Inference proposes why a pattern occurred or what might happen next. Uncertainty includes missing data, ambiguous categories, small samples, changing conditions, conflicts, and plausible alternative explanations. Label each layer where the reader encounters it.
A pilot can test a specified lane under stated conditions. It cannot prove universal assistant performance, justify unrelated access, establish a permanent staffing ratio, or guarantee future results. The five cited sources supply occupational, small-business, measurement, governance, and control concepts. None evaluates this buyer, provider, candidate, assistant, work lane, cost model, or pilot. Recommendations here are proposed applications of those principles, not observed client results or testimonials.
| Item | Finding | Source note |
|---|---|---|
| Known distortion | measuring speed without acceptance, dropping returned or unfinished items, allowing coaching to erase the original result, or expanding scope because a few easy cases passed | Niche-specific limitation analysis |
| Claim boundary | A pilot can test a specified lane under stated conditions. It cannot prove universal assistant performance, justify unrelated access, establish a permanent staffing ratio, or guarantee future results. | Explicit research limitation |
Produce a dated decision packet and learning loop
The decision packet should include the question, owner, scope, eligible population, observation period, source register, field definitions, raw-record references, exclusions, missing-data note, comparisons, exceptions, reviewer decisions, limitations, and next action. Version the packet used for approval and preserve later corrections with a truthful modification date.
A second reviewer should be able to reconstruct the conclusion without a private conversation. That does not require publishing sensitive material. Use stable internal identifiers, minimise personal information, and disclose only what the decision requires. Route unresolved legal, tax, employment, privacy, security, financial, or regulated issues to qualified owners or advisers.
The decision-grade conclusion remains bounded: A pilot can test a specified lane under stated conditions. It cannot prove universal assistant performance, justify unrelated access, establish a permanent staffing ratio, or guarantee future results. The next test is equally specific: Blind-review a small but representative set against the same rubric, retain disagreements and exclusions, then apply the written go, narrow, pause, or stop rule without rewriting it after outcomes appear. Repeat the definitions after any change, retain contrary cases, and compare only equivalent work. This creates an honest learning loop while the buyer retains scope, access, budget, and consequential authority.
| Item | Finding | Source note |
|---|---|---|
| Packet owner | the buyer’s named pilot reviewer who has authority to accept work and change or stop the lane | Named buyer decision record |
| Next test | Blind-review a small but representative set against the same rubric, retain disagreements and exclusions, then apply the written go, narrow, pause, or stop rule without rewriting it after outcomes appear. | Prospective repeat with stable definitions |
Use the record in a staffing conversation
Use the completed record to discuss a bounded assistant pilot. Bring the task examples, source records, exceptions, access boundaries, schedule constraints, open questions, and the name of the person who will accept the work.
The buyer retains responsibility for consequential business decisions and should involve qualified advisers for legal, employment, privacy, security, tax, financial, or regulated questions.
Related Research
A readiness test for the first week with a Philippines-based assistant
An evidence-based go, narrow, or wait test for examples, access, review coverage, and escalation before a first-week pilot begins.
Assistant quality sampling plans: check evidence before volume
A defensible sampling routine for recurring administrative and research work.
Queue denominator integrity in assistant quality research
Why incomplete, reopened, and excluded cases must remain visible when recurring work is measured.
Questions people ask
What is the first question for designing acceptance criteria for a virtual assistant pilot?
What observable evidence will let the buyer decide whether a bounded assistant pilot should continue, narrow, pause, or stop? Name the decision owner and observation unit before choosing a score or comparison.
Does this method prove that outsourced assistant support will save money or improve performance?
No. It structures a local decision from declared evidence and uncertainty. It makes no causal, price, savings, capacity, classification, geographic, or performance promise.
Who approves the resulting staffing decision?
The accountable owner is the buyer’s named pilot reviewer who has authority to accept work and change or stop the lane. Qualified specialists should review matters within their legal, employment, tax, privacy, security, financial, or regulated authority.
Sources
- 1. O*NET OnLine, Executive Secretaries and Executive Administrative Assistants — Official U.S. Department of Labor occupational data used only to identify task, work-context, and skill dimensions that a buyer should verify locally. Checked September 22, 2026.
- 2. U.S. Small Business Administration, Manage Your Business — Official small-business guidance covering planning, employment, compliance, finance, cybersecurity, and continuity responsibilities. Checked September 22, 2026.
- 3. U.S. GAO, Assessing Data Reliability — Primary audit-method guidance for deciding whether operational records are reliable enough for a stated staffing decision. Checked September 22, 2026.
- 4. NIST Cybersecurity Framework 2.0 — Primary governance framework used for responsibility, risk context, supplier, and review concepts; it does not endorse a staffing model. Checked September 22, 2026.
- 5. NIST SP 800-53 Rev. 5, Security and Privacy Controls — Primary control catalogue used for least privilege, separation of duties, external-service, contingency, logging, and account-management concepts. Checked September 22, 2026.
Explore research briefing support · Review the SOP handoff checklist