Voice cloning is increasingly used for audiobooks, podcasts, and digital content creation.
Yet current systems remain black boxes. It’s hard for people to understand what’s happening or customize the outcome.
How can break this black box and make this creative workflow more intuitive & human-centered?
2026 / PRODUCT DESIGN
Product Design
Product Design
Development
Project
CLIENT
Salesforce AI
Time
Jan - June 2026
Team
2 Researchers, 2 Designers
Role
Product Designer
Tools
Git, Claude, Vercel
Role
This project was University of Washington's final capstone project sponsored by Salesforce AI team based in San Francisco.
As product designer of the team, I led the product prototyping and interaction design from concept to final experience.
My biggest contributions were:
Setting up the initial Git repository and Vercel deployment workflow.
Creating the branded design system.
Leading the design and implementation of the voice editing experience.
Solution
An AI Playground for Voice Cloning
Our final product brings recording, evaluation, and fine-tuning into a single interactive workspace.
Rather than accepting whatever the model produces, users can experiment, compare alternatives, and iteratively refine their voice clone.
Transparent, guided Input
A quality voice clone starts from the recording process.
Clear feedback and progress provide actionable feedback, giving users the right information rather than just showing progress.

Realtime Voice Coverage
Users receive real time feedback on each voice parameter analyzed by the voice model API.
Structured
Evaluation
When there's only one clone, it's hard to tell whether its good or bad.
To support decision making, we provide a structured comparison.
This establishes a baseline and encourages an iterative process of re-recording or adding new recordings.
Editing starts with context
Clear goals drive better iteration.
By choosing a content type, users provide context for both themselves and the model. This creates a shared baseline, enabling context-aware presets, editing controls, and recommendations.
Content aware presets
Selecting a content type presets the recommended tones

Realistic scripts
Realistic scripts were written to match different use cases.
Global and granular editing
Users can pinpoint the parts that sound "off."
No need to describe technical voice attributes. AI translates high-level intent into precise model edits.
Project Context
This project was sponsored by Salesforce AI Research, where voice cloning is an emerging area of interest for B2B enterprise applications, related to creating brand-aligned voices for sales and customer support.
My research explored how voice cloning systems could become more interpretable and human-centered for everyday users with no AI expertise.

Problem
Current voice cloning tools are powerful but opaque.
While platforms like Hume and ElevenLabs offer user-friendly experiences, they provide little transparency into how voices are generated.

Goal
Current voice cloning tools are powerful but opaque.
For example, while platforms like Hume AI and ElevenLabs prioritize generating voices, they themselves don't know what's happening in the model level.
What we did
This project was sponsored by Salesforce AI Research, where voice cloning is an emerging area of interest for B2B enterprise applications, related to creating brand-aligned voices for sales and customer support.
My research explored how voice cloning systems could become more interpretable and human-centered for everyday users with no AI expertise.

Speak naturally, not technically
Voices are complex to describe. AI translates their feedback into model-editable voice attributes.

Iteration becomes collaboration
Rather than making hidden model changes, AI explains how it interprets each request. This makes every edit transparent.
Impact
Our final prototype was delivered and showcased at UW's HCDE capstone showcase in June 2026.
We introduced the system to 100+ visitors and put it in the hands of 30+ people. The response was overwhelmingly positive.



Reflections
This project was sponsored by Salesforce AI Research, where voice cloning is an emerging area of interest for B2B enterprise applications, related to creating brand-aligned voices for sales and customer support.
My research explored how voice cloning systems could become more interpretable and human-centered for everyday users with no AI expertise.





