声音克隆方法、系统、电子设备及介质

By segmenting and grading reference speech to determine its stability, a timbre anchor segment library is constructed. High-risk boundary windows are identified and locally corrected, solving the problem of timbre drift accumulation in existing technologies and achieving efficient, low-latency long text audio cloning.

CN121983069BActive Publication Date: 2026-07-17CHONGQING MALYA MEDIA CO LTD

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
CHONGQING MALYA MEDIA CO LTD
Filing Date
2026-04-07
Publication Date
2026-07-17

AI Technical Summary

Technical Problem

Given the overlapping constraints of limited reference speech duration and uneven quality, multiple pause boundaries in the target text, and the need for near real-time output, existing voice cloning technologies struggle to identify high-risk timbre drift windows and cannot effectively perform local corrections, leading to the accumulation and spread of timbre deviations in long text synthesis.

Method used

By segmenting the reference speech into segments, calculating the stability score, and constructing a stable timbre anchor segment library, local corrections are made for high-risk boundary windows. Local backfilling and resynthesis are performed using boundary risk scores and continuous residuals to prevent bias propagation. The threshold is adjusted through an adaptive update mechanism.

Benefits of technology

It effectively identifies and corrects high-risk tone drift windows, avoids high latency in whole-sentence regeneration, ensures output tone stability, adapts to different recording quality conditions, and is suitable for mobile and online applications.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121983069B_ABST
    Figure CN121983069B_ABST
Patent Text Reader

Abstract

本发明涉及音频数据处理技术领域,具体是声音克隆方法、系统、电子设备及介质,包括对输入的参考语音形成稳定音色锚片段库;对目标文本进行前端处理,对各锚片段加权后叠加至该窗口的基础说话人条件向量,形成增强说话人条件向量,并基于增强说话人条件向量合成候选声学结果;从候选声学结果中提取说话人状态表征,计算说话人状态表征与状态一致锚中心向量之间的连续性残差;连续性记忆以衰减方式参与后续窗口的说话人条件注入;将经局部回填后的声学结果输入声码器,生成目标语音波形并输出。本发明解决参考语音时长有限且质量不均匀、目标文本包含多处停顿边界、且需要近实时输出的情况下的声音克隆难题。
Need to check novelty before this filing date? Find Prior Art