Speaker Recognition Feature Extraction via DNN Segmentation
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Existing speaker recognition systems in teleconferencing increase the computational load on the reception side by performing speaker recognition processing, which includes calculating speaker features and comparing them with registered features.
Innovation Solution
An information transmission device that calculates acoustic features and speaker features using a deep neural network (DNN), and transmits these features along with condition information to an information reception device, allowing it to perform speaker recognition with reduced computational load.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Measurement precision
If speaker recognition processing is performed on the reception side, then speaker recognition accuracy is improved, but computational load on the reception device increases
Solution Approach 1:
The speaker recognition processing is segmented into two parts: acoustic feature extraction performed on the transmission side, and speaker feature calculation performed on the reception side. This division allows the reception device to avoid the computationally intensive acoustic feature extraction while still achieving accurate speaker recognition through the transmitted acoustic features.
Solution Approach 2:
The acoustic feature extraction is performed in advance on the transmission side before transmission. By preparing the acoustic features beforehand and transmitting them along with the audio data, the reception device receives pre-processed information that reduces its computational burden while maintaining recognition accuracy.
2Use of energy by moving object
If acoustic feature calculation is performed on the transmission side, then computational load on reception device is reduced, but transmission data volume increases
Solution Approach 1:
Only the essential acoustic features are extracted and transmitted from the transmission side, rather than transmitting all raw audio data or complete processing results. This selective extraction of necessary features reduces the data volume that needs to be transmitted while still providing sufficient information for accurate speaker recognition on the reception side.
Data Source
AI summary
An information transmission device according to the present disclosure includes: an acoustic feature calculator that calculates an acoustic feature of a spoken voice; a speaker feature calculator that calculates a speaker feature from the acoustic feature using a deep neural network (DNN), the speaker feature being a feature unique to a speaker of the spoken voice; an analyzer that analyzes condition information indicating a condition to be used in calculating the speaker feature, based on the spoken voice; and an information transmitter that transmits the speaker feature and the condition information to an information reception device that performs speaker recognition processing on the spoken voice, as information to be used by the information reception device to recognize the speaker of the spoken voice.


