Speaker recognition methods, systems, storage media, and devices based on ordinary pronunciation
By generating spectral masks and fusing spectral features through the UNET network, the problem of insufficient stability in ordinary pronunciation recognition is solved, and efficient and accurate speaker recognition is achieved.
CN115762535BActive Publication Date: 2026-05-26NANJING INST OF INTELLIGENT TECH INST OF MICROELECTRONICS OF THE CHINESE ACAD OF
View PDF 2 Cites 0 Cited by
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- NANJING INST OF INTELLIGENT TECH INST OF MICROELECTRONICS OF THE CHINESE ACAD OF
- Filing Date
- 2022-11-18
- Publication Date
- 2026-05-26
Smart Images

Figure CN115762535B_ABST
Abstract
This invention discloses a speaker recognition method, system, storage medium, and device based on ordinary pronunciation. The method includes: acquiring real-time audio data and extracting spectral features based on the real-time audio data to obtain spectral features corresponding to the real-time audio data; inputting the spectral features corresponding to the real-time audio data into a trained UNET network to generate a spectral mask corresponding to the real-time audio data, and detecting whether the real-time audio data is an ordinary pronunciation based on the spectral mask; if the real-time audio data is an ordinary pronunciation, fusing the spectral mask and spectral features to obtain an enhanced spectrum corresponding to the real-time audio data; inputting the enhanced spectrum corresponding to the real-time audio data into a trained speaker embedding layer network to obtain a real-time speaker embedding layer corresponding to the real-time audio data; and comparing the real-time speaker embedding layer with the registered speaker embedding layer to identify the speaker corresponding to the real-time audio data.
Need to check novelty before this filing date? Find Prior Art