Case Study

Repartee AI

Repartee AI

Hallucination detection and testing platform for enterprise LLM teams.

Timeline

Timeline

8 months

Team

Team

Product Leader · Sole UX Designer (me) · 4 UI Designers · 2 Front-end Engineers · 2 ML Engineers · Marketing Team

My Role

My Role

UX Design · Design System (0→1) · Marketing Landing Page

Tools

Tools

Figma, FigJam, Framer

Repartee AI monitoring dashboard
Repartee AI monitoring dashboard

INTRODUCTION

Enterprise teams ship LLMs they can’t fully trust. I designed the layer that makes hallucination visible, testable, and explainable — before it reaches a customer.

Enterprise teams ship LLMs they can’t fully trust. I designed the layer that makes hallucination visible, testable, and explainable — before it reaches a customer.

~75% of hallucinations caught by users before customers were affected

2.5× team efficiency after design system adoption

1–2 days → 5 min time to validate a prompt change

CONTEXT

What I had when I joined

WHAT I HAD

Lo-fi wireframes, no connected user flow, a 30-person team, and no UX designer.

WHAT I HAD TO FIGURE OUT

Who this is for, what problem we’re actually solving, and how the end-to-end flow connects.

CHALLENGE

Learning about users I couldn’t directly access

1:1s with the team — Engineers, ML, frontend, and UI were the closest proxies I had.

Internal proxy users — PMs and engineers whose day-to-day matched Maya’s profile.

Leader as information bridge — Enterprise buyer conversations synthesized through the product leader.

MEET THE USER

Maya — Product Manager of AI features at an insurance tech company

OWNS
Multiple LLMs in production. Detects hallucination only after customers complain.

CORE PROBLEM
Trust. She is accountable for AI outcomes she can’t fully trust.

KNOWS
Product, business, and customer context. Works closely with engineers, marketing, and leadership.

DOESN’T KNOW
ML internals. Configurations like “temperature” and “coefficient” aren’t intuitive to her.

MOST-MENTIONED ISSUES

The core insight: Maya is accountable for AI outcomes she can’t fully trust

01
Lack of visibility
Maya can’t see what’s happening with her models.

“I detect hallucinations only after a customer complains.”

02
Lack of validation
Maya ships model changes without a way to test them first.

“I push the change, then I wait, then I find out from the customer.”

03
Navigation discontinuity
Maya loses context when switching between features.

“I’m three clicks deep and I forgot which model I was looking at.”

PROBLEM STATEMENT

How might we help Maya trust the models she owns?

How might we help Maya trust the models she owns?

SOLUTION 01
Dashboard
Make hallucination visible.

SOLUTION 02
Playground
Provide a safe way to validate.

SOLUTION 03
Two-layer navigation
Preserve context while navigating.

DESIGN SOLUTION

Monitoring dashboard — improves visibility

Repartee AI model monitoring dashboard
Repartee AI model monitoring dashboard

1ST ITERATION

Model detail — improves diagnosis

Repartee model detail and report table

2ND ITERATION

Playground — provides confidence in testing

Alert threshold sets Maya’s risk tolerance — no more guessing. Hallucinations are marked inline, so she knows exactly which sentence to worry about. Revised answers make it possible to compare the original and improved response before shipping.

Repartee AI testing playground
Repartee AI testing playground

USABILITY & SCALABILITY

Finding the right navigation

Two-layer navigation won: top navigation handles feature-level movement; the sidebar maintains model-level context. Usability testing confirmed that users frequently switch between Dashboard and Playground.

Comparison of navigation directions
Comparison of navigation directions

FINAL SOLUTION

Testing with real tasks, then closing the gap

With three internal proxy users, we asked participants to test whether changing the Alert Threshold improved a model’s output. The interface wasn’t the barrier — ML-facing language was. The next sprint would translate configuration terms into plain questions: “How sensitive should detection be?” and “How strictly should uncertain answers be flagged?”

DESIGN SYSTEM

Extracting a system from existing, scattered work

I extracted a scalable system from scattered components, migrated existing screens with four UI designers without a hard cutover, and aligned with front-end engineering in one working session.

Repartee design-system work
Repartee design-system work

IMPACT

Earning trust through visibility, validation, and navigation

~75%
Hallucinations caught by users before customers were negatively affected.

2.5×
Team efficiency after design-system adoption cut meeting and onboarding time.

1–2 days → 5 min
Time to validate a prompt change — from waiting to knowing.

Retention
Clients who can trust the output stay. Trust is the retention strategy.

KEY LEARNINGS

Think, make, iterate

No direct user access meant treating internal proxies and buyer signals as data. Building relationships with the ML team became a design resource — enabling faster iteration and better decisions.

Trust is not a feature. It is visibility, clarity, and transparency, translated into decisions users can actually act on.

2026. Designed with ♡ in Seattle.

2026. Designed with ♡ in Seattle.

2026. Designed with ♡

in Seattle.