An audio information processing method and system

By combining spectral and temporal coding with global semantic and local structural representation models, the problem of low accuracy in voiceprint recognition in existing technologies has been solved, achieving highly robust and high-precision voiceprint recognition.

CN122369472APending Publication Date: 2026-07-10INNOVATION & INNOVATION CENT OF STATE GRID ZHEJIANG ELECTRIC POWER CO LTD +2
View PDF 0 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
INNOVATION & INNOVATION CENT OF STATE GRID ZHEJIANG ELECTRIC POWER CO LTD
Filing Date
2026-01-30
Publication Date
2026-07-10

AI Technical Summary

Technical Problem

In existing technologies, audio information processing methods are easily affected by environmental noise, speaker pronunciation characteristics, and excessively short speech duration, which leads to a decrease in the accuracy of voiceprint recognition. Furthermore, they fail to fully exploit spectral features, and the discriminative power of single time-domain features is insufficient.

Method used

By extracting spectral features from audio data and performing spectral and temporal coding, spectral vectors and temporal vectors are generated. These are then combined with global semantic and local structural representation models to form a set of voiceprint embeddings. Max pooling is then performed to generate voiceprint feature embedding vectors.

Benefits of technology

It achieves joint representation of audio signals in the frequency and time domains, enhances the comprehensiveness and discriminative ability of voiceprint features, and improves the accuracy of voiceprint recognition.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN122369472A_ABST
    Figure CN122369472A_ABST
Patent Text Reader

Abstract

This invention discloses an audio information processing method and system, applied in the field of audio signal processing technology. The method includes extracting spectral features from audio data to be processed and performing spectral encoding on the spectral features to obtain a spectral vector representation; performing temporal encoding on the audio data to be processed to obtain a temporal vector representation; obtaining a first voiceprint result based on the spectral vector representation and the temporal vector representation; and obtaining a second voiceprint result based on the spectral vector representation and the temporal vector representation; integrating the first and second voiceprint results into a voiceprint embedding set; performing max pooling on the voiceprint embedding set to obtain a voiceprint feature embedding vector; and determining the voiceprint recognition result of the audio data to be processed based on the analysis results of the voiceprint feature embedding vector. The method of this invention improves the accuracy of voiceprint recognition results.
Need to check novelty before this filing date? Find Prior Art