Speech Processing Apparatus Acoustic Diversity Feature Compensation

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Existing speech processing systems face challenges in maintaining accurate speaker recognition due to variations in sound types and noise or distortion in speech signals, leading to decreased recognition accuracy.

Innovation Solution

A speech processing apparatus that calculates an acoustic diversity degree to compensate for recognition feature values, using a combination of acoustic diversity degree calculation and feature compensation processors to enhance the accuracy of speaker recognition by adjusting for variations in sound types and noise.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Adaptability or versatility

If speech signals with noise or distortion are input to speaker recognition apparatus, then the system can process real-world speech data, but accuracy of speaker recognition decreases due to distortion in acoustic features

Engineering Contradiction:
Improveability to process real-world speech signalsVSAvoidaccuracy of speaker recognition
Core Design Contradiction:
Adaptability or versatilityVSMeasurement precision

Solution Approach 1:

The patent applies preliminary action by calculating acoustic diversity degree before speaker recognition processing. The system pre-processes speech signals to determine their acoustic diversity characteristics, then uses this information to adjust recognition thresholds or weighting factors in advance, thereby compensating for expected distortions before they affect recognition accuracy

Inventive Principle:
Principle #10Preliminary action

Solution Approach 2:

The patent implements parameter changes by dynamically adjusting speaker recognition parameters based on acoustic diversity degree. When speech signals exhibit high acoustic diversity (indicating noise or distortion), the system modifies recognition thresholds, similarity criteria, or feature weighting to maintain accurate speaker identification despite degraded signal quality

Inventive Principle:
Principle #35Parameter changes

2Adaptability or versatility

If speech signals lacking certain sound types are input, then the system can handle varied speech content, but difference between acoustic features and speaker model increases, reducing recognition accuracy

Engineering Contradiction:
Improveability to handle varied speech contentVSAvoidmatch between acoustic feature and speaker model
Core Design Contradiction:
Adaptability or versatilityVSMeasurement precision

Solution Approach 1:

The system performs preliminary analysis of acoustic diversity degree to identify speech signals lacking certain sound types before speaker recognition. This early detection allows the system to prepare appropriate compensation strategies, such as adjusting feature extraction parameters or modifying speaker model comparison criteria, to maintain recognition accuracy despite incomplete sound representations

Inventive Principle:
Principle #10Preliminary action

Solution Approach 2:

The patent changes recognition parameters based on acoustic diversity characteristics. When speech signals show low diversity in certain sound types, the system adjusts feature weighting to emphasize available sound types and modifies similarity calculation methods to account for missing components, thereby maintaining accurate speaker recognition across varied speech content

Inventive Principle:
Principle #35Parameter changes

Data Source

PatentUS10490194B2Speech processing apparatus, speech processing method and computer-readable medium
Publication Date: 2019.11.26 NEC CORP
  • US10490194B2 patent drawing
  • US10490194B2 patent drawing
  • US10490194B2 patent drawing

AI summary

A speech processing apparatus, method and non-transitory computer-readable storage medium are disclosed. A speech processing apparatus may include a memory storing instructions, and at least one processor configured to process the instructions to calculate an acoustic diversity degree value representing a degree of variation in types of sounds included in a speech signal representing a speech, on a basis of the speech signal, and compensate for a recognition feature value calculated to recognize specific attribute information from the speech signal, using the acoustic diversity degree value.