On-device text-to-speech model personalization

On-device speech sample generation with criteria-based selection addresses resource and privacy challenges in personalized text-to-speech model training, enhancing user voice matching on portable devices.

WO2026107241A1PCT designated stage Publication Date: 2026-05-21QUALCOMM INC
View PDF 3 Cites 0 Cited by

Patent Information

Authority / Receiving Office
WO · WO
Patent Type
Applications
Current Assignee / Owner
QUALCOMM INC
Filing Date
2025-11-13
Publication Date
2026-05-21

AI Technical Summary

Technical Problem

Portable devices lack the processing and memory resources to support personalized text-to-speech model training, which typically requires several hours of user speech samples and fine tuning, leading to security and privacy issues and increased latency when performed off-device.

Method used

On-device speech sample generation and criteria-based selection to train a personalized text-to-speech model, using confidence, loss, and lexicon diversity checks to select high-quality user speech samples for training, thereby reducing resource burden and privacy concerns.

Benefits of technology

Enables convenient and effective personalization of text-to-speech models on-device, improving voice and vocal characteristics matching without requiring extensive user recording and avoiding network overhead and privacy issues.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure US2025055400_21052026_PF_FP_ABST
    Figure US2025055400_21052026_PF_FP_ABST
Patent Text Reader

Abstract

A device includes a memory configured to store a set of speech samples and one or more processors coupled to the memory. The one or more processors are configured to obtain, during normal operation of the device, one or more audio signals that include user speech and perform a sequence of sample criteria checks on the speech samples associated with the one or more audio signals. The sequence of sample criteria checks includes a check whether a confidence value associated with an automatic speech recognition (ASR) transcription of a sample exceeds a transcription confidence threshold. The sequence of sample criteria checks also includes a check whether a loss value associated with a personalized text-to-speech (TTS) output of the sample exceeds a loss threshold, whether the ASR transcription satisfies a lexicon diversity criterion, or both.
Need to check novelty before this filing date? Find Prior Art