Speaker Identification Without Noise Suppression

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Conventional speaker recognition techniques face accuracy issues due to noise suppression methods that distort personal characteristics and require high calculation amounts, leading to lower recognition accuracy.

Innovation Solution

A speaker identification method that calculates similarity between voice data and registered voice data, determines suitability for identification based on similarity thresholds, and outputs identification results without performing noise suppression, thereby improving accuracy without increasing calculation complexity.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Object-affected harmful factors

If noise suppression is applied to input voice, then noise is reduced, but personal characteristics of the speaker are distorted and recognition accuracy is lowered

Engineering Contradiction:
ImprovenoiseVSAvoidspeaker recognition accuracy
Core Design Contradiction:
Object-affected harmful factorsVSMeasurement precision

Solution Approach 1:

The patent extracts and removes the noise suppression processing step from the speaker recognition system. Instead of applying noise suppression that distorts speaker characteristics, the system directly uses the raw acoustic feature amounts for speaker recognition, thereby eliminating the harmful effect of characteristic distortion while maintaining noise tolerance through alternative means.

Inventive Principle:
Principle #2Taking out (Extraction)

Solution Approach 2:

The patent inverts the conventional approach by not suppressing noise but rather adapting the speaker recognition system to tolerate noise. The system calculates similarity between acoustic feature amounts without prior noise suppression, and determines speaker recognition results based on similarity thresholds that account for noisy conditions.

Inventive Principle:
Principle #13The other way round (Inversion)

2Object-affected harmful factors

If conventional noise suppression methods are used, then noise is reduced, but calculation amount increases

Engineering Contradiction:
ImprovenoiseVSAvoidcalculation amount
Core Design Contradiction:
Object-affected harmful factorsVSPower

Solution Approach 1:

The patent removes the computationally intensive noise suppression processing from the system workflow. By extracting this unnecessary processing step, the system achieves both reduced calculation amount and maintained noise tolerance, as the similarity-based recognition approach naturally handles noisy inputs without requiring complex signal processing.

Inventive Principle:
Principle #2Taking out (Extraction)

3Power

If similarity calculation is performed without noise suppression, then calculation amount is reduced, but noise affects recognition accuracy

Engineering Contradiction:
Improvecalculation amountVSAvoidspeaker recognition accuracy
Core Design Contradiction:
PowerVSMeasurement precision

Solution Approach 1:

The patent changes the recognition criterion from direct acoustic feature comparison to similarity-based comparison with threshold determination. By calculating similarity between acoustic feature amounts and comparing against dynamically determined thresholds, the system maintains high recognition accuracy even in noisy conditions while avoiding computationally intensive noise suppression processing.

Inventive Principle:
Principle #35Parameter changes

Data Source

PatentUS20250022470A1Speaker identification method, speaker identification device, and non-transitory computer readable recording medium storing speaker identification program
Publication Date: 2025.01.16 PANASONIC INTELLECTUAL PROPERTY CORP OF AMERICA
  • US20250022470A1 patent drawing
  • US20250022470A1 patent drawing
  • US20250022470A1 patent drawing

AI summary

A speaker identification device acquires voice data to be identified, acquires a plurality of pieces of registered voice data that are registered in advance, calculates a similarity between the voice data to be identified and each of the plurality of pieces of registered voice data, selects a registered speaker of registered voice data corresponding to a highest similarity from among a plurality of calculated similarities, determines, based on the plurality of calculated similarities, whether or not the voice data to be identified is suitable for speaker identification, determines, based on the highest similarity, whether or not to identify the selected registered speaker as a speaker to be identified of the voice data to be identified in a case where the voice data to be identified is determined to be suitable for the speaker identification, and outputs the identification result.