BootcampInterview prep

Interview Drills — AI Output Evaluation / Quality

6 drills with frameworks and rubrics.

Interview Drills — AI Output Evaluation / Quality

Open-ended interview questions for AI Operations, AI quality analyst, AI trainer, and model-evaluator roles. Each has a Framework (the structure a strong answer follows), a Model answer (a concise example), and a Rubric (what an interviewer listens for). These exercises put a real or hypothetical AI response in front of you and ask you to judge it. Practice thinking aloud, naming the dimension you're testing, and always pointing to the specific span of text you're reacting to. The app can role-play these as mock interviews (see mock-interview.md).

The universal evaluation structure: Restate the task and audience → check the dimensions one by one (factuality → relevance → completeness → safety → instruction-following → tone/clarity) → flag the most serious problem first → quote the exact offending text → give a verdict (pass/revise/reject) with specific, documented feedback. The single most important habit: separate sounds good from is correct — confident fluency is not evidence of truth.

D1

  • difficulty: easy
  • concept: evaluation-dimensions An AI customer-support reply sounds polished and confident. Walk me through how you'd evaluate it before it goes to a customer.
  • Framework: Restate the task and audience → don't be swayed by fluency → check each dimension in order: factuality (are the claims true?), relevance (does it answer this customer?), completeness (anything missing?), safety (any harmful/inappropriate content?), instruction-following (did it honor format/policy constraints?), tone (right for the audience?) → give a verdict and the single biggest issue.
  • Model answer: "First I'd note that polish tells me nothing about correctness — confident AI is exactly where errors hide, so I start with factuality. I'd verify any specifics: did it quote the right refund window, the right policy, the right link? Then relevance — does it address the customer's actual question or a generic version? Completeness — did it skip a required step? Safety — no over-promising or inappropriate tone. Instruction-following — did it stay within our support template and escalation rules? Tone last. My verdict would name the most serious problem first, e.g. 'reads well but states a 60-day return window when ours is 30 — factual error, must fix before sending.'"
  • Rubric: Strong answers explicitly resist surface fluency, check multiple named dimensions with factuality first, tie checks to the specific audience/task, and end with a clear verdict naming the worst issue. Weak answers say "it looks good/professional," judge only tone or grammar, or never mention verifying the facts.

D2

  • difficulty: medium
  • concept: hallucination-detection Here's an AI answer to "Summarize the company's 2023 revenue from this report." It states a specific figure and cites a source. How do you decide whether to trust it?
  • Framework: Treat every specific fact as unverified until checked → separate claims that are in the source from claims the model may have invented → check the cited source actually exists and actually says that → watch for the classic hallucination tells (oddly precise numbers, plausible-but-fake citations, confident detail with no grounding) → verdict: grounded, unsupported, or fabricated.
  • Model answer: "I'd never accept a number just because it's specific — fabricated stats often look the most precise. I'd go back to the report and locate the exact figure and the line it came from. If the answer cites a page or table, I check that page exists and says that. If the model produced a clean figure the source doesn't contain, or cited a section that isn't there, that's a hallucination — AI generates plausible text, not retrieved truth. My verdict: 'the $4.2M figure is not supported by the source; the cited "Table 3" doesn't exist in the report — reject and regenerate with the figure verified against page 7.'"
  • Rubric: Strong answers default to distrust on specifics, go back to the source to ground each claim, know the hallucination tells (fake citations, ungrounded precision), and distinguish grounded vs. fabricated. Weak answers trust the number because it's specific or because a citation is present, or never check the source at all.

D3

  • difficulty: medium
  • concept: evaluation-rubric We don't have a quality rubric for our AI's drafted emails yet. Design one. How would you make evaluations consistent across a team of reviewers?
  • Framework: State why a rubric matters (consistency over people and time, like QA acceptance criteria) → pick the dimensions that matter for this output → make each a concrete, checkable criterion with a rating scale → define what "good" looks like at each level → add tie-breakers/examples to reduce reviewer disagreement → note how you'd calibrate the team.
  • Model answer: "A rubric makes judgments consistent instead of gut-feel. For drafted emails I'd score four dimensions 1–3: Factuality (3 = every claim verifiable and correct; 1 = any false claim), Instruction-following (did it honor the requested length, recipient, and ask?), Tone-fit (right for the audience), and Completeness (covers the needed points). Passing = factuality and instruction-following both at 3, others at 2+. To keep reviewers consistent I'd write a one-line 'what counts as a 2 vs 3' anchor per dimension, attach a gold-standard example, and run a calibration round where everyone scores the same five emails and we reconcile disagreements."
  • Rubric: Strong answers explain why rubrics create consistency, choose dimensions fit to the output, define concrete level descriptors and a passing bar, and address inter-rater calibration (anchors, examples, reconciliation). Weak answers give a vague "rate it 1–5," omit what each level means, or never address how to keep multiple reviewers aligned.

D4

  • difficulty: medium
  • concept: documented-feedback A junior reviewer's feedback on an AI output just says "This is wrong, rejected." Why isn't that good enough, and rewrite it.
  • Framework: Explain why vague feedback fails (it doesn't improve the system, can't become training signal, and isn't reproducible) → name what good feedback contains: which dimension failed, the exact offending text, why it's wrong, and the corrected/expected version → rewrite it specifically → note the bonus of tagging a failure category so patterns surface.
  • Model answer: "'This is wrong' tells no one what to fix — and our evaluations are supposed to become training signal and guide where to improve the model, so they have to be specific and reproducible. Good feedback points at the span and the reason. I'd rewrite it as: 'Factual error — paragraph 2 states the product launched in 2019; it launched in 2021 (see spec). Also instruction-following: the user asked for under 100 words and this is 180. Verdict: revise. Failure tag: invented-date.' That's actionable, lets us track that the model keeps inventing dates, and someone can act on it without re-reading the whole task."
  • Rubric: Strong answers explain that specific feedback is what actually improves AI and enables pattern-spotting, then rewrite with the dimension, the quoted offending text, the reason, the correction, and ideally a failure tag. Weak answers just add "be more detailed" without modeling it, or rewrite it but still vaguely.

D5

  • difficulty: hard
  • concept: safety-and-bias An AI résumé-screening tool produces summaries of candidates that read as helpful and neutral. What would you specifically look for to decide they're safe to use?
  • Framework: Recognize this is a high-stakes, people-affecting use where surface neutrality can hide harm → check for bias (does it treat similar candidates differently by gender/age/name/background? does it surface or penalize protected attributes?) → check for fabrication (inferring facts not in the résumé) → check fairness consistency across comparable cases → safety/harm and instruction-following → recommend human-in-the-loop and a verdict.
  • Model answer: "Because this affects real people's livelihoods, 'reads neutral' isn't enough — bias is learned from data and can be amplified, and it hides behind polite language. I'd test for disparate treatment: run near-identical résumés that differ only by name, gender cue, or graduation year and see if the summaries shift in tone or emphasis. I'd check it isn't inventing qualifications or gaps not in the résumé (hallucination), and isn't surfacing protected attributes as if relevant. I'd verify it follows the instruction to summarize, not score or recommend. Verdict: don't ship unsupervised — keep a human in the loop for consequential decisions, and only after a bias audit shows consistent treatment across matched pairs."
  • Rubric: Strong answers flag the high-stakes/people-affecting context, actively test for bias (matched-pair reasoning), catch fabrication and protected-attribute leakage, and insist on human-in-the-loop plus an audit before trusting it. Weak answers accept "it reads neutral," check only grammar/tone, or ignore bias and fairness entirely.

D6

  • difficulty: hard
  • concept: failure-patterns You've evaluated 200 AI outputs this week. The team asks, "So is the model good?" How do you answer in a way that actually helps them improve it?
  • Framework: Reject a single thumbs-up/down → report by dimension and by failure pattern, not one overall vibe → quantify where it fails and how badly → name recurring failure types with examples (the high-value signal) → tie patterns to where in the system to fix → recommend the next action, not just a grade.
  • Model answer: "I wouldn't give one verdict — 'good' isn't actionable. I'd report by pattern: 'Factuality passes 92% of the time, but it invented a statistic in 14 of the 16 failures — a recurring hallucination pattern on numeric claims. Instruction-following is the bigger issue: it ignored length limits in 30% of long prompts. Tone and relevance are strong.' Then I'd point to the fix: the numeric-hallucination pattern suggests we add a retrieval/verification step or a guardrail on stats; the length misses suggest a prompt or post-processing fix. That turns evaluation into a map of what to fix first, which is the point — recurring failure types guide where to improve the system."
  • Rubric: Strong answers refuse a single grade, break results down by dimension and recurring failure pattern with rough quantification and examples, and connect each pattern to a concrete fix/priority. Weak answers answer "yes it's good / no it's bad," give no breakdown, or list errors without spotting patterns or pointing toward improvement.