Case Study
Hallucination detection and testing platform for enterprise LLM teams.
8 months
Product Leader · Sole UX Designer (me) · 4 UI Designers · 2 Front-end Engineers · 2 ML Engineers · Marketing Team
UX Design · Design System (0→1) · Marketing Landing Page
Figma, FigJam, Framer
INTRODUCTION
~75% of hallucinations caught by users before customers were affected
2.5× team efficiency after design system adoption
1–2 days → 5 min time to validate a prompt change
CONTEXT
What I had when I joined
WHAT I HAD
Lo-fi wireframes, no connected user flow, a 30-person team, and no UX designer.
WHAT I HAD TO FIGURE OUT
Who this is for, what problem we’re actually solving, and how the end-to-end flow connects.
CHALLENGE
Learning about users I couldn’t directly access
1:1s with the team — Engineers, ML, frontend, and UI were the closest proxies I had.
Internal proxy users — PMs and engineers whose day-to-day matched Maya’s profile.
Leader as information bridge — Enterprise buyer conversations synthesized through the product leader.
MEET THE USER

Maya — Product Manager of AI features at an insurance tech company
OWNS
Multiple LLMs in production. Detects hallucination only after customers complain.
CORE PROBLEM
Trust. She is accountable for AI outcomes she can’t fully trust.
KNOWS
Product, business, and customer context. Works closely with engineers, marketing, and leadership.
DOESN’T KNOW
ML internals. Configurations like “temperature” and “coefficient” aren’t intuitive to her.
MOST-MENTIONED ISSUES
The core insight: Maya is accountable for AI outcomes she can’t fully trust
01
Lack of visibility
Maya can’t see what’s happening with her models.
“I detect hallucinations only after a customer complains.”
02
Lack of validation
Maya ships model changes without a way to test them first.
“I push the change, then I wait, then I find out from the customer.”
03
Navigation discontinuity
Maya loses context when switching between features.
“I’m three clicks deep and I forgot which model I was looking at.”
PROBLEM STATEMENT
SOLUTION 01
Dashboard
Make hallucination visible.
SOLUTION 02
Playground
Provide a safe way to validate.
SOLUTION 03
Two-layer navigation
Preserve context while navigating.
DESIGN SOLUTION
Monitoring dashboard — improves visibility
1ST ITERATION
Model detail — improves diagnosis

2ND ITERATION
Playground — provides confidence in testing
Alert threshold sets Maya’s risk tolerance — no more guessing. Hallucinations are marked inline, so she knows exactly which sentence to worry about. Revised answers make it possible to compare the original and improved response before shipping.
USABILITY & SCALABILITY
Finding the right navigation
Two-layer navigation won: top navigation handles feature-level movement; the sidebar maintains model-level context. Usability testing confirmed that users frequently switch between Dashboard and Playground.
FINAL SOLUTION
Testing with real tasks, then closing the gap
With three internal proxy users, we asked participants to test whether changing the Alert Threshold improved a model’s output. The interface wasn’t the barrier — ML-facing language was. The next sprint would translate configuration terms into plain questions: “How sensitive should detection be?” and “How strictly should uncertain answers be flagged?”
DESIGN SYSTEM
Extracting a system from existing, scattered work
I extracted a scalable system from scattered components, migrated existing screens with four UI designers without a hard cutover, and aligned with front-end engineering in one working session.
IMPACT
Earning trust through visibility, validation, and navigation
~75%
Hallucinations caught by users before customers were negatively affected.
2.5×
Team efficiency after design-system adoption cut meeting and onboarding time.
1–2 days → 5 min
Time to validate a prompt change — from waiting to knowing.
Retention
Clients who can trust the output stay. Trust is the retention strategy.
KEY LEARNINGS
Think, make, iterate
No direct user access meant treating internal proxies and buyer signals as data. Building relationships with the ML team became a design resource — enabling faster iteration and better decisions.
Trust is not a feature. It is visibility, clarity, and transparency, translated into decisions users can actually act on.




