Speaker Recognition Feature Extraction via DNN Segmentation

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Existing speaker recognition systems in teleconferencing increase the computational load on the reception side by performing speaker recognition processing, which includes calculating speaker features and comparing them with registered features.

Innovation Solution

An information transmission device that calculates acoustic features and speaker features using a deep neural network (DNN), and transmits these features along with condition information to an information reception device, allowing it to perform speaker recognition with reduced computational load.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Measurement precision

If speaker recognition processing is performed on the reception side, then speaker recognition accuracy is improved, but computational load on the reception device increases

Engineering Contradiction:
Improvespeaker recognition accuracyVSAvoidcomputational load on reception device
Core Design Contradiction:
Measurement precisionVSUse of energy by moving object

Solution Approach 1:

The speaker recognition processing is segmented into two parts: acoustic feature extraction performed on the transmission side, and speaker feature calculation performed on the reception side. This division allows the reception device to avoid the computationally intensive acoustic feature extraction while still achieving accurate speaker recognition through the transmitted acoustic features.

Inventive Principle:
Principle #1Segmentation

Solution Approach 2:

The acoustic feature extraction is performed in advance on the transmission side before transmission. By preparing the acoustic features beforehand and transmitting them along with the audio data, the reception device receives pre-processed information that reduces its computational burden while maintaining recognition accuracy.

Inventive Principle:
Principle #10Preliminary action

2Use of energy by moving object

If acoustic feature calculation is performed on the transmission side, then computational load on reception device is reduced, but transmission data volume increases

Engineering Contradiction:
Improvecomputational load on reception deviceVSAvoidtransmission data volume
Core Design Contradiction:
Use of energy by moving objectVSQuantity of substance

Solution Approach 1:

Only the essential acoustic features are extracted and transmitted from the transmission side, rather than transmitting all raw audio data or complete processing results. This selective extraction of necessary features reduces the data volume that needs to be transmitted while still providing sufficient information for accurate speaker recognition on the reception side.

Inventive Principle:
Principle #2Taking out (Extraction)

Data Source

PatentUS20250069603A1Information transmission device, information reception device, information transmission method, recording medium, and system
Publication Date: 2025.02.27 PANASONIC INTELLECTUAL PROPERTY CORP OF AMERICA
  • US20250069603A1 patent drawing
  • US20250069603A1 patent drawing
  • US20250069603A1 patent drawing

AI summary

An information transmission device according to the present disclosure includes: an acoustic feature calculator that calculates an acoustic feature of a spoken voice; a speaker feature calculator that calculates a speaker feature from the acoustic feature using a deep neural network (DNN), the speaker feature being a feature unique to a speaker of the spoken voice; an analyzer that analyzes condition information indicating a condition to be used in calculating the speaker feature, based on the spoken voice; and an information transmitter that transmits the speaker feature and the condition information to an information reception device that performs speaker recognition processing on the spoken voice, as information to be used by the information reception device to recognize the speaker of the spoken voice.