SALESFORCE AI

YOUR VOICE CLONE FOR YOUR CONTENT

Voice cloning is increasingly used for audiobooks, podcasts, and digital content creation.


Yet current systems remain black boxes. It’s hard for people to understand what’s happening or customize the outcome.

How can break this black box and make this creative workflow more intuitive & human-centered?

2026 / PRODUCT DESIGN

Product Design

Product Design

Development

SALESFORCE AI

YOUR VOICE CLONE FOR YOUR CONTENT

SALESFORCE AI

YOUR VOICE CLONE FOR YOUR CONTENT

Project

CLIENT

Salesforce AI

Time

Jan - June 2026

Team

2 Researchers, 2 Designers

Role

Product Designer

Tools

Git, Claude, Vercel

Role

Alex

Designer

Soyun

Designer

Gahui

Researcher

Mingjin

Researcher

Alex

Designer

Gahui

Researcher

Soyun

Designer

Mingjin

Researcher

Alex

Designer

Mingjin

Researcher

Soyun

Designer

Gahui

Researcher

This project was University of Washington's final capstone project sponsored by Salesforce AI team based in San Francisco.

As product designer of the team, I led the product prototyping and interaction design from concept to final experience.

My biggest contributions were:

  • Setting up the initial Git repository and Vercel deployment workflow.

  • Creating the branded design system.

  • Leading the design and implementation of the voice editing experience.

Solution

An AI Playground for Voice Cloning
Our final product brings recording, evaluation, and fine-tuning into a single interactive workspace.

Rather than accepting whatever the model produces, users can experiment, compare alternatives, and iteratively refine their voice clone.

Transparent, guided Input

A quality voice clone starts from the recording process.

Clear feedback and progress provide actionable feedback, giving users the right information rather than just showing progress.

Actionable Feedback

Instead of simply telling users what's fulfilled or not, we give them actionable feedback.

Actionable Feedback

Instead of simply telling users what's fulfilled or not, we give them actionable feedback.

Realtime Voice Coverage

Users receive real time feedback on each voice parameter analyzed by the voice model API.

Structured
Evaluation

When there's only one clone, it's hard to tell whether its good or bad.

To support decision making, we provide a structured comparison.
This establishes a baseline and encourages an iterative process of re-recording or adding new recordings.

Editing starts with context

Clear goals drive better iteration.

By choosing a content type, users provide context for both themselves and the model. This creates a shared baseline, enabling context-aware presets, editing controls, and recommendations.

Content aware presets

Selecting a content type presets the recommended tones

Realistic scripts

Realistic scripts were written to match different use cases.

Global and granular editing

Users can pinpoint the parts that sound "off."

No need to describe technical voice attributes. AI translates high-level intent into precise model edits.

Speak naturally, not technically

AI translates plain user feedback into model parameters.

Speak naturally, not technically

AI translates plain user feedback into model parameters.

Iteration becomes collaboration

AI explains how it interprets each request. This makes every edit transparent and collaborative.

Iteration becomes collaboration

AI explains how it interprets each request. This makes every edit transparent and collaborative.

Project Context

This project was sponsored by Salesforce AI Research, where voice cloning is an emerging area of interest for B2B enterprise applications, related to creating brand-aligned voices for sales and customer support.

My research explored how voice cloning systems could become more interpretable and human-centered for everyday users with no AI expertise.

Problem

Current voice cloning tools are powerful but opaque.
While platforms like Hume and ElevenLabs offer user-friendly experiences, they provide little transparency into how voices are generated.

Goal

Current voice cloning tools are powerful but opaque.
For example, while platforms like Hume AI and ElevenLabs prioritize generating voices, they themselves don't know what's happening in the model level.

What we did

This project was sponsored by Salesforce AI Research, where voice cloning is an emerging area of interest for B2B enterprise applications, related to creating brand-aligned voices for sales and customer support.

My research explored how voice cloning systems could become more interpretable and human-centered for everyday users with no AI expertise.

Speak naturally, not technically

Voices are complex to describe. AI translates their feedback into model-editable voice attributes.

Iteration becomes collaboration

Rather than making hidden model changes, AI explains how it interprets each request. This makes every edit transparent.

Impact

Our final prototype was delivered and showcased at UW's HCDE capstone showcase in June 2026.

We introduced the system to 100+ visitors and put it in the hands of 30+ people. The response was overwhelmingly positive.

Reflections

This project was sponsored by Salesforce AI Research, where voice cloning is an emerging area of interest for B2B enterprise applications, related to creating brand-aligned voices for sales and customer support.

My research explored how voice cloning systems could become more interpretable and human-centered for everyday users with no AI expertise.

Alex Yeha Chung, product designer who finds joy in scaling products with visual taste and intention

Craft

+

Strategy

Currently in Seattle

Local time 1:53 PM

Alex Yeha Chung, product designer who finds joy in scaling products with visual taste and intention

Craft

+

Strategy

Currently in Seattle

Local time 1:53 PM

Alex Yeha Chung, product designer who finds joy in scaling products with visual taste and intention

Craft

+

Strategy

Currently in Seattle

Local time 1:53 PM