BootcampCapstone · Deliverable 4

Usability Test Plan & Findings

Builds on Topic 7.

What you'll produce

A Usability Test Plan & Findings report for the FreshCart checkout prototype you built in Deliverable 3 — a short, two-part document that first plans a test (a goal, realistic task scenarios, and exactly what you'll watch for) and then reports what happened when ~5 real people tried it: where they hesitated, the issues ranked by severity against named usability heuristics, and the prioritized fixes you actually made to the prototype. This is the deliverable that separates a designer from a decorator. Anyone can draw screens; a UX/UI designer can prove a design works — or honestly admit where it doesn't — by watching real people use it and letting that observation drive the next iteration. "I tested this with five users and here's what I changed" is the single most persuasive sentence in a portfolio review and a design interview, and this report is where you earn the right to say it.

Instructions

  1. Write a one-line test goal and your top open question. State what you're trying to learn, framed as a doubt, not a wish — e.g. "Can a first-time FreshCart shopper get from the home screen to a placed order in under 3 minutes without help, and where exactly do they stall?" If you're secretly trying to confirm your design is good, you'll run a bad test. Aim to find problems.

  2. Define a success metric per task before you test, so a result can't be argued away later: a target completion rate, a target time, and the number of moments you needed to step in ("assists"). Decide these now — afterward you'll be tempted to move the goalposts.

  3. Write 2–4 realistic task scenarios as goals, not instructions. Say "You want fresh ingredients for tonight's dinner delivered — go ahead and place an order" — never "tap the Search icon, then add Bananas, then tap Cart." Naming the buttons tells the user the answer and destroys the test. Each scenario needs a clear, observable success state ("order confirmation screen reached").

  4. List exactly what you'll observe. For each task, write the specific signals: completion (yes/no), time on task, number of assists, error paths taken (wrong taps, backtracks), points of hesitation, and verbatim quotes. You can only see what you decide to look for.

  5. Recruit ~5 participants who match your persona, not your friends who already know the app. Five is not a budget compromise — it's the number that surfaces the large majority of usability problems. Note one screening fact per person (e.g. "orders groceries online monthly, never used FreshCart"). No user pool and doing this solo at home? You still have four real paths — pick whichever you can finish this week, and write down which one you used. The bar is the same for all of them: 5 people, persona-matched, prototype-blind (they've never seen your design), one screening fact logged each.

    • Unmoderated remote test (fastest; small cost or free tier): Upload the Figma prototype to Maze, UserTesting, Useberry, or Lyssna/UserCrowd, write your tasks once, and use the tool's panel to target the persona ("orders groceries online ≥ monthly, owns a smartphone, ages 25–45"). It records screen + think-aloud audio from 5 strangers and you watch the replays — you don't need to be present, and you never met them, so no one knows the "answer."
    • Post your own screener (free; costs time): Drop a one-paragraph screener in a relevant subreddit, a Discord/Slack community, a neighborhood Facebook group, or a campus board: "Do you order groceries online at least monthly? Never used FreshCart? 15-min video call, I'll send a $5 coffee card." Screen before scheduling — accept only people who match the persona and have never seen your prototype — then run it over Zoom/Meet with screen-share.
    • Acquaintances who genuinely match — and are blind to the design: A cousin, neighbor, or coworker from another team is fine only if they fit the persona, have never opened your Figma file, and you can resist coaching them. The disqualifier isn't "I know them," it's "they already know where to tap." Your roommate who watched you build it is out; an aunt who orders groceries online weekly and has never seen the prototype is a valid participant.
    • In-person intercept (zero cost, in an afternoon): Ask 5 matching strangers in a café, library, or co-working space for 10 minutes each. Hand them your phone with the prototype open and the task on a card, then go quiet.

    Whichever path you pick, name your recruitment method in the report (e.g. "Recruited via a Maze unmoderated test, panel filtered to monthly online-grocery shoppers" or "5 acquaintances screened to the persona, none had seen the prototype"). Stating how you sourced and screened participants is itself a credibility signal — it shows the results came from the right people, not a convenience sample.

  6. Run each session the same way: think-aloud, no rescuing. Ask them to narrate their thoughts out loud, give the task, then go quiet. When they get stuck, count to ten in your head and take notes — their struggle is your data. The instant you say "it's the button up top," you've thrown away that finding.

  7. Log raw observations per participant in a simple table — task, completed?, time, assists, and where they stalled — before you interpret anything. Keep facts separate from your theories about them.

  8. Cluster observations into distinct issues and rank each by severity (a 0–4 scale: 0 = not a problem, 1 = cosmetic, 2 = minor/annoying, 3 = major/blocks many users, 4 = catastrophic/blocks the task). Severity is roughly frequency × impact — an issue 4 of 5 people hit that stops the order outranks one cosmetic gripe.

  9. Tag each issue with the usability heuristic it violates (visibility of system status, match to the real world, consistency, error prevention/recovery, recognition over recall). The heuristic names the root cause and points to the fix — "people missed it" is an observation; "no visibility of system status" is a diagnosis.

  10. Decide and document a fix for the top issues, then say what you changed in the prototype and what you'd re-test. Fixes are part of the deliverable — a report that finds problems but ships nothing proves you can critique, not design. Close with what you'd verify in a second round.

Worked example

(Prototype under test: the redesigned FreshCart "browse → cart → checkout" flow from Deliverable 3. Goal: first-time shopper places an order in under 3 minutes.)

Part 1 — Test plan

  • Test goal: Determine whether a first-time FreshCart shopper can go from the home screen to a confirmed order in under 3 minutes with zero facilitator help, and pinpoint exactly where they stall.
  • Top open question: Our analytics say 58% of new users abandon their first order between browsing and checkout. Deliverable 3 deliberately fixed the big causes — forced sign-up (guest is now default), the surprise delivery fee (now surfaced in the cart), the missing order summary (the 08 Review screen now restates every line item). The question is whether the remaining, subtler details we left rough in grey-box — the delivery-slot selected-state, and whether the fee shown in the cart actually matches the slot the user later picks — quietly re-create the same hesitation.
  • Method: Moderated, remote, think-aloud. 5 participants, ~20 min each, tested on the clickable Figma prototype on their own phones.
  • How I recruited (solo, no user pool): Posted a 3-line screener in a local "online grocery deals" Facebook group and two friends-of-friends threads — "Order groceries online at least monthly? Never used FreshCart? 15 min on a video call, $5 coffee card as thanks." Of 11 replies I took the first 5 who matched the persona and confirmed they'd never seen the prototype; turned away two who were close contacts already familiar with the design. (Free alternative if you can't fill the slots: an unmoderated Maze test with its panel filtered to monthly online-grocery shoppers — same 5 strangers, no scheduling.)
  • Participants (persona match — "Maya," the busy first-time shopper; all prototype-blind): P1 orders groceries online ~monthly (Instacart), never used FreshCart. P2 new to grocery delivery entirely. P3 weekly online grocery shopper. P4 shops in-app but on a slow connection. P5 first-time grocery-delivery user, low tech confidence.
  • Tasks & success metrics (set before testing):
    • T1 — Find & add items: "You want fresh ingredients for tonight's dinner delivered tonight. Add what you'd need for a simple pasta dinner to your cart." Success = ≥3 items in cart. Target: 5/5 complete, < 90s, 0 assists.
    • T2 — Review cart & set delivery: "Get your order ready to pay — pick a delivery time that works for you." Success = reaches Review with a delivery slot clearly selected (participant can say which slot they picked). Target: 5/5, < 45s, 0 assists.
    • T3 — Confirm & place the order: "Check everything looks right, then place the order." Success = order-confirmation screen reached without backtracking to re-check the price. Target: 4/5, < 45s, ≤1 assist.
  • What I'll observe per task: completed (Y/N), time on task, # of assists, wrong taps / backtracks, hesitation points (>5s pause), and verbatim quotes.

Part 2 — Findings

Raw results

ParticipantT1 add itemsT2 cart + slotT3 confirm + placeTotal timeAssists
P1✅ 71s✅ 49s✅ 40s2:400
P2✅ 88s⚠️ 1:10 (re-tapped slots, unsure which was picked)✅ 38s3:361
P3✅ 64s⚠️ 1:02 (picked evening slot)⚠️ 1:25 (fee on Review ≠ fee seen in cart)3:311
P4✅ 79s⚠️ 1:18 (slot looked "already selected")⚠️ 1:08 (backtracked twice to re-check total)3:451
P5✅ 92s⚠️ 1:30 (stalled — "is this asking me to pick?")✅ 44s3:462

Completion: T1 5/5, T2 5/5 (4 slow), T3 5/5 (3 slow, 0 abandonment). But only 1 of 5 hit the under-3-minute goal. The big D3 fixes held — nobody hit a sign-up wall, everybody found search on the home screen, and the 08 Review summary stopped anyone from abandoning. The time leaked instead at two rough edges we'd left in grey-box: the delivery-slot selected-state and a fee that changed between the cart and Review when the chosen slot wasn't the default.

Issues, severity-ranked, with heuristic violated

A note on what this list is — and isn't. The headline D3 fixes (guest-first, fee-in-cart, the 08 Review order summary) worked: nobody hit a sign-up wall, nobody abandoned, everyone found search on the home screen. That's the test confirming the big bets — which is why the issues below are narrower. A good second-round test doesn't re-discover the problems you already designed out; it surfaces the rough edges you left in grey-box because you couldn't see them until a real thumb hit them. All four below live inside the prototype exactly as built in D3.

  1. [Severity 3 — major] Delivery slot has no clear "selected" state. The single biggest time sink: 4 of 5 stalled on T2. On 06 Delivery Address & Slot the slots are tappable and the primary button reads "Choose this slot," but a tapped slot looks identical to an untapped one — no fill, no check, no "Selected: …" line. Users couldn't tell whether tapping registered, re-tapped to be safe, or assumed a slot was already chosen for them (P4 read the top slot as pre-selected and nearly shipped a time they didn't want). The grey-box wireframe defined the button but never the selected-state, and that gap is exactly what bit. Heuristic: Visibility of system status. Verbatim (P5): "Is this asking me to pick one, or just showing me the times? I can't tell if I've done it."
  2. [Severity 3 — major] The delivery fee shown in the cart doesn't match the fee on Review. D3's 04 Cart surfaces the fee up front as "Delivery (tomorrow AM) £3.99" — but that's the fee for the default slot. P3 picked an evening slot on 06, and the £3.99 they'd anchored on in the cart quietly became a different number by the time they reached 08 Review. They burned 1:25 backtracking to work out why the total had moved. We fixed the surprise-at-the-very-end fee but introduced a fee that contradicts itself mid-flow whenever the chosen slot isn't the default — a regression hiding inside a fix. Heuristic: Consistency + visibility of system status (one line item shows two values across the flow). Verbatim (P3): "Hang on — it said £3.99 in my basket. Why is it more now? Did I add something?"
  3. [Severity 2 — minor] The step indicator misreports where the user is on the Account Gate. The 4-step bar Cart · Delivery · Payment · Review rides screens 5–8 (per the D3 sticky note), so on 05 Account Gate it highlights Cart — a step already finished and behind them — while the gate itself ("guest vs. sign in") has no label in the bar at all. Two participants paused, unsure whether "Continue as guest" would drop them back into the cart. A cue meant to answer "how far am I" actively gives the wrong answer. Heuristic: Visibility of system status + recognition over recall. Verbatim (P2): "It still says Cart at the top — am I going backwards?"
  4. [Severity 1 — cosmetic] "Total" vs. "Order total" wording drifts between 04 Cart ("Total £6.18") and 08 Review ("Place order — £24.18"). One participant noticed the labels weren't identical; no one was blocked. Heuristic: Consistency.

Prioritized fixes made to the prototype

  • Fix #1 (Issue 1): Gave 06 Delivery slots a real selected-state — the chosen slot fills in, gets a check, and a confirmation line below the list reads "Selected: Today 6–7pm (£3.99)." Demoted the always-on "Choose this slot" button so it only activates after a slot is tapped, removing the "is it already picked?" ambiguity. This targets the biggest time sink. Re-test: does anyone still re-tap or stall wondering whether a slot is selected? Target T2 < 45s for 5/5.
  • Fix #2 (Issue 2): Made the delivery fee a single source of truth tied to the chosen slot. The 04 Cart line now reads "Delivery (from £3.99 — choose at checkout)" instead of quoting the default slot's price as if it were final, and 06/08 show the actual fee for the selected slot. No number changes silently between screens. Re-test: does the total on 08 Review ever surprise a participant who picked a non-default slot? Target 0 backtracks-to-recheck-price.
  • Fix #3 (Issue 3): Re-anchored the step indicator so it stops misreporting position: dropped the orphaned "Cart" step from the in-checkout bar and relabeled it Account · Delivery · Payment · Review, with 05 Account Gate now correctly highlighting "Account" as step 1 of 4. Re-test: do participants on 05 still ask whether they're being sent backward?
  • Deferred — Issue 4 (label wording): cosmetic and unblocking; batched into the typography/labels pass in the high-fidelity UI work (Deliverable 5) rather than spent on now.
  • What I'd verify next round: re-run the same 3 tasks with 5 fresh, prototype-blind participants against the fixed prototype; target the under-3-minute goal for ≥4 of 5, T2 < 45s for 5/5, and zero price-surprise backtracks on T3.

Rubric

The app's AI scores the learner's submission against these criteria and gives feedback. Levels: Needs work (1) / Solid (2) / Excellent (3). Passing = every criterion at Solid or above.

  • Realistic task scenarios (goals, not instructions) — 1: tasks name the buttons/steps or aren't tied to the flow · 2: realistic goal-based tasks with observable success states · 3: tasks read like real first-timer goals, each with a pre-set success metric (rate/time/assists) that makes results undeniable.
  • Unbiased test method & recruitment — 1: leading or rescuing the user, no think-aloud, or participants who already know the design · 2: think-aloud with no facilitator help, ~5 persona-matched users, and a stated way they were recruited · 3: clean think-aloud protocol, explicit "what I'll observe" list, and a concrete, prototype-blind recruitment method named (unmoderated panel like Maze/UserTesting, a posted screener, or genuinely-matching acquaintances who've never seen the prototype) — proving the 5 participants were sourced, not assumed.
  • Findings grounded in observation — 1: opinions/assumptions about the design, not what users did · 2: concrete observations (stalls, errors, completion, quotes) per participant · 3: a clear facts-then-interpretation log with verbatim quotes and per-task data that visibly drives every issue.
  • Severity ranking + heuristic diagnosis — 1: a flat list with no ranking or heuristics · 2: issues ranked by severity and each tagged to a usability heuristic · 3: severity reflects frequency × impact (the order-blocker outranks the cosmetic gripe) and each heuristic names the true root cause, not just the symptom.
  • Prioritized fixes that drove iteration — 1: problems found but nothing changed · 2: a fix per top issue, applied to the prototype · 3: fixes clearly matched to the highest-severity issues, what changed in the prototype is specific, and each names what you'd re-test to confirm it worked.
  • Coherence with prior deliverables — 1: disconnected from the case study · 2: tests the actual Deliverable 3 prototype and persona · 3: traces cleanly from the persona (D1) and flow (D2) through the prototype (D3), and the fixes set up the high-fidelity UI (D5).