Voice conversion method and device, electronic equipment and storage medium

By generating acoustic features through posterior probability features of speech and historical memory information, the limitations of multi-source speaker conversion in existing technologies are overcome, enabling low-cost, real-time speech conversion that is suitable for scenarios such as live streaming and instant messaging.

CN115547349BActive Publication Date: 2026-07-21出门问问(苏州)信息科技有限公司
View PDF 2 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
出门问问(苏州)信息科技有限公司
Filing Date
2022-09-19
Publication Date
2026-07-21

AI Technical Summary

Technical Problem

Existing speech conversion technologies require parallel corpus training, which limits the application of speech conversion from multiple source speakers to the same target speaker, and is difficult to apply in real-time scenarios, resulting in high costs.

Method used

Acoustic features are generated using posterior probability features and historical memory information. Speech blocks are received via WebSocket and converted in real time using a pre-trained acoustic model, achieving many-to-one speech conversion.

Benefits of technology

It achieves low-cost, real-time speech conversion, supports conversion from multiple source speakers to the same target speaker, and is suitable for scenarios such as live streaming and instant messaging.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN115547349B_ABST
    Figure CN115547349B_ABST
Patent Text Reader

Abstract

The present disclosure provides a voice conversion method and device, electronic equipment and storage medium. The voice conversion method of the present disclosure comprises: obtaining a current source speech block; extracting a speech posterior probability feature of the current source speech block; generating an acoustic feature corresponding to the current source speech block by using the speech posterior probability feature of the current source speech block and historical memory information of a previously obtained previous source speech block; and synthesizing the acoustic feature corresponding to the current source speech block into a target speech block; wherein the target speech block has a timbre feature of a target speaker, and a speaker of the source speech block is a source speaker. The present disclosure can realize real-time voice conversion, has low cost, fast execution speed, and supports one-to-many voice conversion.
Need to check novelty before this filing date? Find Prior Art