一种目标说话人语音获取方法和系统

By designing a target speaker speech acquisition model in multi-speaker scenarios and using hybrid and reference corpora for feature separation and comparison, the problem of low voiceprint recognition accuracy in multi-speaker scenarios is solved, achieving more efficient target speaker recognition and application expansion.

CN115881093BActive Publication Date: 2026-07-17XIAMEN KUAISHANGTONG TECH CORP LTD

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
XIAMEN KUAISHANGTONG TECH CORP LTD
Filing Date
2022-10-26
Publication Date
2026-07-17

AI Technical Summary

Technical Problem

In multi-speaker scenarios, existing voiceprint recognition technology struggles to accurately identify the target speaker, resulting in low recognition accuracy and limiting the application scenarios of voiceprint recognition systems.

Method used

By acquiring mixed corpora, reference corpora, and single-speaker corpora, and utilizing the target speaker's speech acquisition model, including a mixed speech interface module, a speech encoding module, a speaker extraction module, a reference speech interface module, a speaker encoding module, and a speaker comparison module, one-to-one feature scoring and speech decoding are performed to improve the model's training effect and recognition accuracy.

Benefits of technology

It effectively improves the accuracy of voiceprint recognition in multi-speaker scenarios, expands the application scenarios of voiceprint recognition, and enhances the robustness and recognition efficiency of the model.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN115881093B_ABST
    Figure CN115881093B_ABST
Patent Text Reader

Abstract

本发明实施例提供一种目标说话人语音获取方法,包括:获取混合语料、参考语料以及多个单人语料,语音编码模块获取混合语料的混合声学特征,说话人提取模块分离出混合声学特征中不同说话人的单人声学特征;说话人编码模块获取参考语料中的参考声学特征,得到参考声学特征集;说话人比对模块将得到的单人声学特征分别到参考声学特征集中进行一对一特征打分,确定出目标说话人;根据目标人声学特征还原为目标人说话语音,得到训练好的目标说话人语音获取模型;将目标说话人的参考语音和含有目标说话人的混合语音,输入到训练好的目标说话人语音获取模型中,得到目标说话人语音;本发明提供的方法,能够有效提升多说话人场景下的声纹识别准确率。
Need to check novelty before this filing date? Find Prior Art