← Research

Leveraging Generative AI to Synthesize the Lost Voice of Laryngectomy Patients

Aug 2023 – Jun 2024View on GitHub →

Fine-tuned Matcha-TTS (an encoder-decoder TTS model trained with Optimal-Transport Conditional Flow Matching) on a single target speaker's voice to let laryngectomy patients regain their voice identity. The 14M-parameter model reached a Mean Opinion Score of 3.66/5 and a real-time factor of 0.038 — fast enough for live conversation — then was exported to ONNX for lightweight, on-device inference.

Leveraging Generative AI to Synthesize the Lost Voice of Laryngectomy Patients — poster board