Interactive speech editing research prototype

Vowel-Level Emotion Editing with ProMoNet

A controllable pipeline for transferring pitch, loudness, periodicity, duration, and phonetic features from an emotional donor to selected vowels in neutral speech.

Why vowel-level editing?

Make emotional prosody measurable—and controllable.

Emotion emerges from coordinated changes in pitch, loudness, timing, and voice quality. Conventional recordings make these cues difficult to isolate because they naturally change together.

This project uses ProMoNet to manipulate selected cues inside aligned vowels while retaining surrounding consonants and sentence content.

Project at a glance

Local control without rewriting the utterance.

Input

Sentence-matched speech

Neutral, happy, and sad recordings with manually corrected Montreal Forced Aligner TextGrids.

Control

Five donor features

Select individual vowels and transfer pitch, loudness, periodicity, duration, or PPG.

Output

Auditable synthesis

Export an edited WAV, feature tensors, and a per-vowel record of every transferred cue.

Donor-transfer workflow

One neutral base. One emotional donor. Five controlled steps.

The recordings must share the same ordered vowel sequence, preventing transfer between unrelated speech segments.

  1. 01

    Align

    Read vowel intervals from corrected TextGrids and verify sequence correspondence.

  2. 02

    Extract

    Convert boundaries to frames and extract pitch, loudness, periodicity, duration, and PPG.

  3. 03

    Select

    Choose the vowels and donor features to transfer in the notebook interface.

  4. 04

    Transfer

    Replace selected neutral segments, resampling trajectories when neutral timing is retained.

  5. 05

    Synthesize

    Reconstruct with the chosen ProMoNet voice and save the waveform and audit artifacts.

PitchLoudnessPeriodicityDurationPPG

Boundary behavior: an optional taper blends continuous acoustic features into the neutral context. PPG uses a direct splice because it represents categorical phonetic information.

Two guided interfaces

Adapt a voice, then design an edit.

01

Speaker adaptation

Choose recordings, set a training target, resume managed runs, and confirm the resulting checkpoint.

Speaker-adaptation notebook interface with recording and training controls.
02

Vowel editor

Pair neutral and donor recordings, select aligned vowels, choose donor features, and synthesize the result.

Vowel-level editing notebook interface with aligned recordings and donor-feature controls.

Adaptation in four steps

  1. Name the speaker run.
  2. Select varied recordings.
  3. Set target and checkpoint steps.
  4. Train or resume, then refresh the model list.

Editing controls

Duration
Adopts donor timing
Pitch, loudness, periodicity
Support tapered boundaries
PPG
Uses a direct splice

Sentence-matched demonstration

Listen before and after.

All examples use sentence 47_01. Each player is paired with the same-scale spectrogram and its DNSMOS Pro predicted speech-quality MOS.

Reference

Neutral original

DNSMOS Pro MOS: 2.08 / 5
Spectrogram of the neutral original recording.
Donor

Happy original

DNSMOS Pro MOS: 2.56 / 5
Spectrogram of the natural happy donor recording.
Baseline

Neutral reconstruction

DNSMOS Pro MOS: 2.95 / 5
Spectrogram of the ProMoNet neutral reconstruction.

Controlled edits

Change one cue—or transfer them together.

These examples use the same happy donor, all aligned vowels, and a 0 ms boundary taper.

Acoustic

Pitch only

DNSMOS Pro MOS: 2.84 / 5

Transfers the donor fundamental-frequency trajectory.

Spectrogram of the pitch-only happy edit.
Acoustic

Loudness only

DNSMOS Pro MOS: 2.61 / 5

Transfers the donor energy trajectory.

Spectrogram of the loudness-only happy edit.
Acoustic

Periodicity only

DNSMOS Pro MOS: 2.87 / 5

Transfers the donor degree of voiced excitation.

Spectrogram of the periodicity-only happy edit.
Timing

Duration only

DNSMOS Pro MOS: 2.76 / 5

Adopts the duration of each donor vowel.

Spectrogram of the duration-only happy edit.
Phonetic

PPG only

DNSMOS Pro MOS: 3.06 / 5

Transfers the donor phonetic posteriorgram.

Spectrogram of the PPG-only happy edit.

Praat comparison

Two conventional resynthesis baselines.

Both files use the matching sentence 47_01 and transfer donor pitch, loudness, and duration contours.

Praat baseline

Neutral to happy

DNSMOS Pro MOS: 1.49 / 5
Spectrogram of the Praat neutral-to-happy edit.
Praat baseline

Neutral to sad

DNSMOS Pro MOS: 2.06 / 5
Spectrogram of the Praat neutral-to-sad edit.

Predicted speech quality

DNSMOS Pro across every condition.

Higher predicted MOS indicates better perceived speech quality on the model’s 1–5 scale. These values describe quality, not transcription accuracy or emotion strength.

PipelineNeutralHappySad
Original audio2.3812.7182.563
Vowel-edit pipeline2.1202.1692.144
Praat pipeline—2.0872.123

Means use 50 original trials per emotion; 50 neutral reconstructions; 49 happy and 50 sad edits per pipeline. Happy sentence 39 is excluded because its aligned vowel sequence does not match.

Method: the NISQA-trained DNSMOS Pro checkpoint was pinned to commit 72f0fa4. Audio was converted to mono 16 kHz, repetitively extended or cropped to 10 seconds, and scored with the model’s recommended STFT input. The displayed score is the predicted MOS mean; model variance is retained in the CSV files.

Feature trajectories

See exactly what changed.

Orange shows the donor, blue the neutral base, and green the edited output. Shaded bands mark aligned vowel regions.

Pitch-only transfer trajectories for pitch, loudness, and periodicity.
Pitch-only edit: edited pitch follows the donor inside selected vowels.

Objective validation

Original-trial check.

Before evaluating the edits, the 150 natural recordings establish how well audEERING represents the three original emotion conditions.

Arousal-valence scatter plot of 150 original recordings with emotion centroids and 50 percent covariance data ellipses.
Fifty original trials per emotion. Diamonds mark centroids; shaded covariance ellipses enclose 50% of each group under a bivariate-normal contour.

Strong signal, with a useful limit.

The happy originals form a clearly elevated arousal cluster, showing that audEERING captures the strongest emotion contrast well in these recordings.

Neutral and sad centroids also move in the expected direction, but their ellipses overlap. The model is therefore informative for this validation without implying perfect categorical classification.

Edited-trial comparison

Sentence-matched movement in emotion space.

All 398 recordings were scored with audEERING’s dimensional emotion model. Each edit was paired by speaker and sentence with its natural happy or sad target.

ProMoNet · Happy+0.128

median target gain

Praat · Happy+0.108

median target gain

ProMoNet · Sad−0.002

median target gain

Praat · Sad−0.011

median target gain

Three-panel validation plot with happy target gain, sad target gain, and arousal-valence space.
Left and center: sentence-level target gains with median/bootstrap summaries and paired pipeline tests. Right: audEERING arousal-valence space for all systems.

Happy targets

Both systems moved consistently toward natural happy speech. ProMoNet showed the larger median gain from its reconstruction baseline; Praat had the smaller absolute distance to the target.

Sad targets

Neither system showed consistent movement toward natural sad speech. Both bootstrap confidence intervals included zero.

TargetSystemNMedian gain95% bootstrap CIPositive gainMedian target distance
HappyProMoNet490.1283[0.0836, 0.1521]91.8%0.1484
HappyPraat490.1077[0.0812, 0.1229]95.9%0.0846
SadProMoNet50−0.0022[−0.0271, 0.0086]48.0%0.1007
SadPraat50−0.0105[−0.0336, 0.0012]38.0%0.1168

Paired pipeline comparison

Two-sided paired t-tests were conducted to compare mean target gain between systems after matching the same sentence. Neither emotion shows a statistically significant ProMoNet–Praat difference.

TargetPairsProMoNet medianPraat medianMean paired differencet(df)Two-sided p
Happy490.12830.1077+0.01301.27 (48)0.209
Sad50−0.0022−0.0105+0.01300.90 (49)0.375

Interpretation limit: this validation includes one female native speaker of American English. The bootstrap intervals describe variation across this speaker’s sentences and do not support speaker-general conclusions.

Reproduce the workflow

Code, notebooks, and analysis.

The public repository contains the two interfaces, the Praat baseline, DNSMOS Pro quality analysis, audEERING scoring and statistics, and the environment specification.

Open the GitHub repository