Why should a custom-system pilot involve more than one superuser?
A pilot run by one knowledgeable superuser proves that a well-informed person with broad access can navigate the system. It does not prove that the operating team can use it. Before launch, representative initiators, operators, reviewers and support staff should complete normal and exception journeys with their own roles, realistic data, ordinary devices and actual handoffs. Record outcomes, elapsed and waiting time, assistance, workarounds and data discrepancies. The superuser may prepare the run and help diagnose failures, but should not click through other roles or make the launch decision alone.
Separate demonstration, testing, pilot and acceptance
These activities are often compressed into a single meeting, even though they answer different questions.
| Activity | Question it answers | Primary participants | What it cannot replace |
|---|---|---|---|
| Product demonstration | What has been implemented and how is it intended to work? | Product and delivery teams, business representatives | Independent work by real roles |
| System testing | Does the solution meet functional, integration, permission and failure specifications? | Test and engineering teams | Business comprehension and adoption |
| User acceptance testing (UAT) | Can business users achieve the agreed result? | Representative operational roles | Sustained production operation |
| Limited pilot | Can the workflow, handoffs and support model operate with real work? | Pilot group, operations, support and delivery | Contractual acceptance |
| Contract acceptance | Has the agreed scope, quality and asset handover been met? | Authorised client and supplier representatives | Continuing product improvement |
Microsoft's implementation testing guidance places UAT inside a broader pre-deployment strategy and relates scope to business processes, data, geography, security, compliance and end-to-end operation. The guidance is written for Dynamics 365, so its product details are not a universal contract. The transferable principle is that UAT is a business-process exercise, not a partner's final sales demonstration.
This article addresses UAT and a limited operational pilot. Contract acceptance still follows the parties' statement of work and acceptance criteria; see How should software acceptance criteria be defined? for that separate decision.
Why the superuser creates false confidence
A superuser has usually attended discovery meetings, understands why fields exist and may hold privileges unavailable to ordinary staff. They remember hidden navigation, know which spreadsheet fixes incomplete data and can ask the delivery team to correct a record directly. Those advantages conceal five common gaps:
- Role gaps: the requester can submit, but the operator cannot see an attachment or support cannot see the refund state.
- Handoff gaps: one state changes correctly, but the next role receives no signal or cannot tell who owns the work.
- Permission gaps: the journey succeeds only because the same account can request, approve, execute and export.
- Environment gaps: the office laptop and network work, while a shop-floor phone, an older browser, a scanner or a weak connection fails.
- Exception gaps: demonstration data is complete; production brings duplicates, missing evidence, wrong amounts, returns, cancellations and timeouts.
The GOV.UK introduction to user research calls for studying an end-to-end service, including offline steps and the staff or third parties who provide and support it. Its public-service standards do not set a private company's acceptance thresholds. They do directly support a general product conclusion: testing only the customer-facing screen omits the operational people and handoffs that determine the outcome.
Build the minimum pilot around four responsibilities
These are responsibilities, not mandatory job titles. A small organisation may assign several to one person, but the person should still change account and permission context when the workflow changes.
| Responsibility | Representative work | Failures it should expose | Evidence to retain |
|---|---|---|---|
| Initiator | Create a request, order, case or service job; provide evidence; follow progress | Unclear fields, long forms, duplicate submission, poor mobile use | Independent completion, time, missing evidence, duplicate records |
| Operator | Accept, assign, execute and update the work | Ambiguous queues, wrong priority, weak-network loss, missed notifications | Acceptance time, state history, files and activity record |
| Reviewer | Check amount, quality, authority or completion evidence; approve, return or reject | Decisions without source evidence, unsafe combined permissions, dead-end returns | Review basis, decision, return reason and resubmission |
| Support | Reconstruct the journey, explain status and escalate an exception | Fragmented history, unclear ownership, unrecorded administrative changes | First response, diagnosis, escalation and recovery |
Add finance, warehouse, field engineers, external partners or administrators where those handoffs carry risk. Do not mechanically recruit one person per role. If novice and experienced staff, head office and branches, desk and field work, or different devices can change the result, cover those conditions explicitly.
Write scenario cards from business trigger to verifiable result
“Log in, open the menu and save” is not an acceptance scenario. A useful scenario card names:
- the business outcome, rather than a page to visit;
- starting account, permissions, source data and operational state;
- every role that receives or transfers responsibility;
- device, network, timing and external-system constraints;
- the normal path from trigger to result;
- one or more meaningful exceptions, such as missing evidence, duplication, return, cancellation, timeout or dependency failure;
- the expected state, amount, notification, audit event, report and customer-visible result;
- the rule for blocking launch versus accepting a time-bound residual issue.
For a field-service case, do not stop after “support created a ticket.” Let a customer submit an asset number and photograph, support verify entitlement and assign the job, a technician accept it on a phone and upload work evidence, and a reviewer close or return it. Then remove a required document or make the part unavailable. Verify reassignment, customer notification, support visibility and recovery. Every participant uses their own role; an administrator appears only through a recorded support procedure.
Digital.gov's usability-test model starts with an explicit study purpose and observes task completion before follow-up questions. The GOV.UK moderated-testing guide recommends tasks with a clear, relevant and believable goal that does not reveal the answer. These resources concern usability research, not the whole of UAT. A pilot can still use the method to prevent a facilitator from teaching the participant through the journey and then calling the coached result a pass.
Make data and environment realistic without losing control
A database containing three perfect records will not expose duplicates, incomplete history, boundary values or dirty imports. Prepare authorised, de-identified representative data covering routine volume, important historical states and high-consequence exceptions. A migration pilot should reconcile source and target totals, critical fields, money and sampled records; “import completed” is not a business result.
Realism still needs safeguards:
- use an isolated pilot environment or an explicitly limited production cohort, not an unknown build deployed to everyone;
- create real role accounts and never share an administrator password;
- place payments, messages, external notifications and device control behind sandboxes, allowlists or controlled limits;
- document data provenance, access, retention and end-of-pilot deletion;
- fix the build, configuration, integrations and data snapshot for each round so a failure remains reproducible.
Microsoft's data-migration UAT guidance combines business users, an integrated test environment, customer and migrated data, and the current solution version. It is product-specific implementation guidance, not proof that every project needs the same tooling. It directly supports testing with business users under integrated, representative conditions instead of using a synthetic interface demonstration.
Measure operating evidence, not only defect count
Ten wording defects and one incorrect settlement amount do not carry equal risk. Group evidence into five areas:
- Task outcome: whether each role completed the work independently and whether critical data and amounts were correct.
- Flow efficiency: completion time, waiting time, re-entry and number of handoffs.
- Assistance and workarounds: prompts required, spreadsheets or chat used outside the system, and administrator intervention.
- Exception recovery: whether returned, cancelled, failed and interrupted work can resume safely.
- Operational readiness: whether help, account provisioning, monitoring, escalation and support ownership are usable.
Thresholds come from the business baseline, consequence and contract. There is no universal “95% completion” target. Errors affecting money, authority, privacy or physical equipment may block launch; low-risk copy problems may enter a dated residual list. Every result should retain its scenario, build, participant role, expected and actual outcome, evidence, severity and decision.
Pass three gates before launch
Gate 1: the operational loop closes
Representative roles complete every critical journey to a verifiable result. No key step depends on an undocumented phone call, database edit or superuser proxy.
Gate 2: exceptions recover safely
Representative returns, duplicates, timeouts, integration failures and denied permissions do not create wrong balances, lost work or ownerless states. High-risk actions can be paused, reversed or handed to a responsible person.
Gate 3: the organisation can support the release
Account provisioning, training, issue intake, monitoring, backup, escalation and rollback have named owners. “No severe software defect” is insufficient when nobody can support the service.
The decision can be a restricted pilot, repair and retest, reduced scope, deferred integration or no launch—not only an unconditional pass or fail. A conditional pass must name the cohort, duration, owner, residual risk and rollback trigger.
Common mistakes
“The department head can accept it for everyone.” A sponsor can decide business priority but may not perform mobile capture, batch processing or support searches. Decision authority and operational evidence are different.
“Training will solve every failure.” Training explains a legitimate business rule. If representative users repeatedly fail because the task entry, field meaning or recovery message is unclear, improve the product rather than converting a design defect into permanent training cost.
“Production will tell us what the real problems are.” A reversible, limited pilot can produce real evidence. Money, permissions, personal data, bulk notification and device control should not receive their first test in an unrestricted launch.
“The superuser had no complaint, so the system is ready.” Ask whether ordinary roles completed the task without prompting, whether the next role could verify the result and whether a failure had a known owner and recovery route.
Nine questions for the launch owner
- Which end-to-end failures would affect payment, fulfilment, customers, privacy or safety?
- Who initiates, operates, reviews, supports and owns each journey?
- Which differences in experience, location, device, network or assistive technology may change the outcome?
- Which build, configuration, integrations and data snapshot are in this round?
- Which normal and exception scenarios must pass, and where is evidence retained?
- How will facilitators avoid revealing the answer or taking control?
- Which findings block launch, and which can receive a dated remediation plan?
- Who monitors, supports, pauses and rolls back the release?
- For a conditional pass, what are the cohort, duration and exit conditions?
Wavesteam can prepare the role–task matrix, end-to-end scenarios, data, accounts, evidence structure and launch gates as part of solution planning and delivery. Business participants can then test throughout milestones instead of clicking through every feature on the final day. A good pilot does not show that the most knowledgeable person can operate the system. It shows that ordinary roles can complete accountable work under real constraints—and that the organisation knows how to recover when they cannot.
Sources
- Digital.gov: How to conduct a usability test, covering study purpose, task observation, follow-up and success indicators; accessed 11 October 2026.
- GOV.UK Service Manual: User research for government services and Using moderated usability testing, covering real users, delivery staff, end-to-end services and believable tasks; accessed 11 October 2026.
- Microsoft Learn: Test your Dynamics 365 solution before deployment and Guide for User Acceptance Test after Data Migration, covering end-to-end business scope, business users, integrated environments and representative data; accessed 11 October 2026.