A personalized voice cloning method and system based on individual voice habits

By separating nonverbal voice features from the target speaker's historical audio data, and using acoustic-semantic association analysis and generative adversarial networks, nonverbal voices that conform to personalized vocal habits are generated and seamlessly integrated with speech content. This solves the problem of insufficient simulation of nonverbal vocal habits in existing technologies and achieves a more natural voice cloning effect.

CN122116873APending Publication Date: 2026-05-29CLOUD ATTACK NETWORK TECH HEBEI CO LTD

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
CLOUD ATTACK NETWORK TECH HEBEI CO LTD
Filing Date
2026-01-20
Publication Date
2026-05-29

AI Technical Summary

Technical Problem

Existing voice cloning technology struggles to accurately capture and simulate the unique nonverbal vocalization habits of the target speaker, resulting in cloned voices lacking realism and naturalness in interactions.

Method used

By separating nonverbal voice features from the target speaker's historical audio data, and using acoustic-semantic association analysis and generative adversarial networks, nonverbal voices that conform to personalized vocal habits are generated and seamlessly integrated with the speech content, thus achieving anthropomorphic voice cloning.

Benefits of technology

The generated cloned voice closely resembles the target speaker in timbre and pronunciation habits, enhancing the naturalness and realism of the interaction and providing a more authentic and natural communication experience.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN122116873A_ABST
    Figure CN122116873A_ABST
Patent Text Reader

Abstract

The application discloses a personification voice cloning method and system based on personalized voice habits, which comprises the following steps: extracting non-verbal sound features from historical audio of a target speaker, analyzing the semantic association with the dialogue text, and determining personalized voice habit parameters; clustering the parameters to obtain a classified voice habit dataset; when the emotional category in the dataset matches the current dialogue context, generating and selecting the optimal non-verbal sound variant matching the context by using a generative adversarial network; fusing the acoustic features of the optimal variant with the acoustic features of the speech content, reconstructing the waveform through a vocoder, and outputting the personification cloned voice. The application solves the problem of mechanical and single non-verbal sound simulation in the prior art, makes the cloned voice closer to a real person in terms of voice habits and emotional expression, and significantly improves the realism and naturalness of voice interaction.
Need to check novelty before this filing date? Find Prior Art