Revision: revised with a proposed order-portal study, neutral task prompts and clear boundaries between research, code checks and acceptance.

“User testing” is an ambiguous label; usability testing describes a more specific activity. People may use user testing to mean usability sessions, prototype feedback or broader evaluation with customers. Agree the research question and method rather than assuming those words describe two universally separate processes.

In usability testing, representative participants try to complete tasks while researchers examine what happens. GOV.UK's moderated usability guidance describes observing participants using a service. The evidence can include qualitative observations and, with a suitable study design, quantitative measures. “Qualitative versus quantitative” is not the defining difference between these labels.

Start with the question you need to answer

Different methods can inform the same project. The table gives suggested matches for a B2B order portal; it is a planning framework, not an exclusive classification.

Match the uncertainty to evidence that can address it
QuestionSuitable methodEvidence and its limit
How do staff currently resolve order enquiries?User interviews and observation of existing work.Needs, constraints and workarounds; not proof that a proposed interface works.
Does the proposed concept make sense?Prototype discussion or concept feedback.Reactions and misunderstandings; stated preference is not task completion.
Can someone find and explain an order's status?Task-based usability testing.Observed actions, interpretation, errors and assistance under stated conditions.
Does a status filter return the correct records?Functional checks, including automated regression tests.Specified behavior against test data; not human understanding.
Can the business accept the agreed workflow?User acceptance testing, or UAT, with business representatives.Acceptance against agreed requirements and recorded exceptions.
What accessibility barriers remain?Accessibility evaluation plus research with disabled participants.Technical findings and lived task experience; neither a small study nor a tool alone covers everything.

For a broader project brief, our software QA and testing services guide explains how to scope the work and its deliverables.

Define a small study around a real decision

Suppose a fictional distributor is redesigning its order portal. The team wants to know whether staff can distinguish dispatched items from cancelled items and explain whether anything remains to ship. This is a proposed study plan, not a report of research we conducted.

Choose participants who do, or are likely to do, this work. Include relevant differences such as experience with the business and access needs. Developers who already know the navigation are not substitutes for the intended users. GOV.UK's research planning guidance connects questions, user groups and methods before recruitment.

Use a prototype or test environment with fictional order SO-104: 12 units ordered, eight dispatched and four cancelled. Nothing remains to ship. Ensure the interface actually contains the information needed to reach that conclusion. This scenario does not require real customer accounts, transactions or messages.

Give a goal without giving away the route

A neutral prompt could be: “A colleague asks what happened to order SO-104 and whether they should expect another shipment. Use the portal to find out, then explain what you would tell them.”

A leading alternative would be: “Open the Completed tab, select SO-104 and read the shipment and cancellation panels.” That supplies navigation and vocabulary the study is meant to evaluate. GOV.UK recommends tasks with a clear, believable goal that do not reveal the solution.

Try the prompt before recruiting the study participants. Check that the task is understandable and achievable in the prototype. Document unsupported interactions so a prototype limitation is not later presented as a participant's error.

Agree the completion rule before observing

For this proposed task, unassisted completion means finding the correct order and explaining that eight units shipped, four were cancelled and none remain to ship, without directional help. Reaching a screen or saying “done” is insufficient if the participant concludes that another four units will arrive.

Record assisted completion separately when a facilitator supplies a hint. Record incomplete or incorrect outcomes when the participant stops or reaches the wrong conclusion. A prototype failure is a separate cause to investigate. Do not change these categories afterward to improve the apparent result.

If timing matters, define the start, end and treatment of interruptions in advance. Think-aloud discussion and facilitator assistance affect what a duration means. Use the same conditions for comparisons, and retain the qualitative explanation behind any number.

Keep a practical facilitation and evidence record

Explain the session, any recording and how the material will be used before starting. Establish the participant's agreement. Clarify that the interface is being evaluated rather than their ability. Observe before offering help, and record any intervention so it remains visible in the findings.

For each task, keep a compact record: participant code and relevant experience, prototype version, device and access setup, exact prompt, observed actions, errors, outcome category, assistance and follow-up questions. Separate direct observations from your interpretation.

For example, “opened the order total twice and did not inspect cancellations” would be an observation if it occurred. “The cancellation information is difficult to discover” would be an interpretation to examine. These are examples of recording practice, not invented findings from this portal. GOV.UK's analysis guidance likewise separates observations, findings and subsequent actions.

Report small studies as bounded evidence

A small qualitative round can reveal useful problems without estimating their frequency across the whole customer population. There is no universal participant count that proves an interface usable or exposes every issue. Choose coverage around the questions, user differences and decisions at stake, then identify what remains unexamined.

Report the tasks, participant characteristics, conditions and observed counts with their denominator. Avoid turning a convenience sample into a precise population success rate. A severe obstacle seen once can deserve attention even when it is not the most frequent observation.

Prioritize a finding by its consequence and the available evidence. Propose a change, identify the behavior it should improve and evaluate it again. Keep uncertainty visible when several explanations remain plausible.

Combine human observation with other checks

The companion automated regression testing example covers an authenticated workspace order list and its status filter. It checks specified software behavior; it does not observe a person finding an order or interpreting a partial dispatch. This proposed usability study covers a different question.

Accessibility also needs distinct evidence. W3C recommends combining evaluation with disabled users and evaluation against accessibility standards. One participant cannot represent every disability, device or assistive-technology setup, and successful task completion alone does not establish conformance.

Use the QA and testing brief to record the question, method, environment, evidence and decision owner. A useful testing plan says what each activity can establish and connects its findings to a concrete product decision.