Voice conversion method and device, electronic equipment and storage medium
By generating acoustic features through posterior probability features of speech and historical memory information, the limitations of multi-source speaker conversion in existing technologies are overcome, enabling low-cost, real-time speech conversion that is suitable for scenarios such as live streaming and instant messaging.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- 出门问问(苏州)信息科技有限公司
- Filing Date
- 2022-09-19
- Publication Date
- 2026-07-21
AI Technical Summary
Existing speech conversion technologies require parallel corpus training, which limits the application of speech conversion from multiple source speakers to the same target speaker, and is difficult to apply in real-time scenarios, resulting in high costs.
Acoustic features are generated using posterior probability features and historical memory information. Speech blocks are received via WebSocket and converted in real time using a pre-trained acoustic model, achieving many-to-one speech conversion.
It achieves low-cost, real-time speech conversion, supports conversion from multiple source speakers to the same target speaker, and is suitable for scenarios such as live streaming and instant messaging.
Smart Images

Figure CN115547349B_ABST