An audio information processing method and system
By combining spectral and temporal coding with global semantic and local structural representation models, the problem of low accuracy in voiceprint recognition in existing technologies has been solved, achieving highly robust and high-precision voiceprint recognition.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- INNOVATION & INNOVATION CENT OF STATE GRID ZHEJIANG ELECTRIC POWER CO LTD
- Filing Date
- 2026-01-30
- Publication Date
- 2026-07-10
AI Technical Summary
In existing technologies, audio information processing methods are easily affected by environmental noise, speaker pronunciation characteristics, and excessively short speech duration, which leads to a decrease in the accuracy of voiceprint recognition. Furthermore, they fail to fully exploit spectral features, and the discriminative power of single time-domain features is insufficient.
By extracting spectral features from audio data and performing spectral and temporal coding, spectral vectors and temporal vectors are generated. These are then combined with global semantic and local structural representation models to form a set of voiceprint embeddings. Max pooling is then performed to generate voiceprint feature embedding vectors.
It achieves joint representation of audio signals in the frequency and time domains, enhances the comprehensiveness and discriminative ability of voiceprint features, and improves the accuracy of voiceprint recognition.
Smart Images

Figure CN122369472A_ABST