A method, apparatus and system for processing audio data

By combining the sound source location and identity recognition results of audio data with short-term and long-term voiceprint features, the problem of recognizing multiple speakers in multi-point video conferences was solved, achieving accurate audio data classification and recognition.

CN114333853BActive Publication Date: 2026-04-03HUAWEI TECH CO LTD
View PDF 3 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2020-09-25
Publication Date
2026-04-03

AI Technical Summary

Technical Problem

In multi-point video conferencing, existing voiceprint recognition systems struggle to accurately identify multiple speakers, especially in free-flowing conversation scenarios. Furthermore, channel differences during voiceprint registration and recognition contribute to insufficient accuracy.

Method used

By acquiring the sound source location information and identity recognition results of audio data, and combining them with facial recognition methods, accurate classification of audio data can be achieved. Speaker identification is performed by combining short-term and long-term voiceprint features, without the need for pre-registration of voiceprint features.

Benefits of technology

It achieves accurate audio data classification in multi-person speaking scenarios, improves the accuracy and efficiency of speech data recognition, and reduces dependence on channel differences.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN114333853B_ABST
    Figure CN114333853B_ABST
Patent Text Reader

Abstract

This application provides an audio data processing method, device, and system for classifying conference audio data according to the speaker's identity. Specifically, this application includes: the conference recording processing device acquiring audio data from a first conference room, sound source location information corresponding to the audio data, and an identity recognition result, wherein the additional field information includes the sound source location information corresponding to the audio data, and the identity recognition result is used to indicate the correspondence between the speaker's identity information obtained through facial recognition and the speaker's speaking time information; then, the conference recording processing device performs speech segmentation on the audio data to obtain first segment audio data; finally, the conference recording processing device determines the speaker corresponding to the first segment audio data based on the voiceprint features of the first segment audio data and the identity recognition result.
Need to check novelty before this filing date? Find Prior Art

Citation Information

Patent Citations

  • Method, device and system for sorting voice conference minutes

    CN102968991A

  • Method and device for generating conference record and conference terminal

    CN110232925A

  • Voice user interface display method and conference terminal

    CN111258528A