Leveraging Generative AI to Synthesize the Lost Voice of Laryngectomy Patients
Fine-tuned Matcha-TTS (an encoder-decoder TTS model trained with Optimal-Transport Conditional Flow Matching) on a single target speaker's voice to let laryngectomy patients regain their voice identity. The 14M-parameter model reached a Mean Opinion Score of 3.66/5 and a real-time factor of 0.038 — fast enough for live conversation — then was exported to ONNX for lightweight, on-device inference.
