Speech Processing Device with Position-Based Feature Adjustment

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Speech recognition rates are lowered due to changes in sound source position under reverberation conditions, as the positional relationship between the sound source and recording point affects the degree of reverberation, leading to decreased accuracy in speech recognition processing.

Innovation Solution

A speech processing device that includes a sound source localization unit, a reverberation suppression unit, a feature quantity calculation unit, and a feature quantity adjustment unit, which determines the sound source position, suppresses reverberation components, calculates adjusted feature quantities, and performs speech recognition using these adjusted quantities to maintain recognition accuracy despite changes in sound source position.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Adaptability or versatility

If speech recognition processing is performed for recorded speech under reverberation conditions, then speech recognition can be performed in indoor environments, but speech recognition rate becomes lower than original sound

Engineering Contradiction:
Improvespeech recognition capability under reverberationVSAvoidspeech recognition rate
Core Design Contradiction:
Adaptability or versatilityVSReliability

Solution Approach 1:

The invention extracts and removes the reverberation component from the recorded speech signal. The reverberation suppression unit separates the direct sound from the reverberation component and eliminates the reverberation, keeping only the clean direct sound for speech recognition processing. This extraction approach resolves the contradiction by removing the harmful reverberation while preserving the useful speech content.

Inventive Principle:
Principle #2Taking out (Extraction)

Solution Approach 2:

The invention uses the captured reverberation component to generate position information about the sound source. By analyzing the characteristics of the reverberation, the system determines the positional relationship between the sound source and recording point, then uses this position information to adjust the acoustic model. This converts the harmful reverberation into a useful source of positional information that improves recognition accuracy.

Inventive Principle:
Principle #22Blessing in disguise (Convert harm into benefit)

2Adaptability or versatility

If the position of sound source or recording point changes, then mobility and flexibility are improved, but speech recognition rate is lowered due to different degrees of reverberation

Engineering Contradiction:
Improveposition flexibilityVSAvoidspeech recognition rate
Core Design Contradiction:
Adaptability or versatilityVSReliability

Solution Approach 1:

The invention dynamically adapts the acoustic model based on the determined sound source position. The system adjusts the acoustic model parameters according to the positional relationship between sound source and recording point, allowing the recognition system to optimize for each specific configuration. This dynamic adaptation resolves the contradiction by making the system flexible to position changes while maintaining high recognition rates.

Inventive Principle:
Principle #15Dynamics

Solution Approach 2:

The invention changes the parameters of the acoustic model based on the sound source position information. By adjusting acoustic model parameters according to the determined position, the system compensates for variations in reverberation characteristics that occur with different positional relationships. This parameter adjustment maintains speech recognition accuracy despite changes in sound source or recording point position.

Inventive Principle:
Principle #35Parameter changes

3Device complexity

If reverberation suppression is performed without position information, then processing complexity is reduced, but speech recognition accuracy decreases due to position-dependent reverberation characteristics

Engineering Contradiction:
Improveprocessing complexityVSAvoidspeech recognition accuracy
Core Design Contradiction:
Device complexityVSMeasurement precision

Solution Approach 1:

The invention performs preliminary determination of sound source position information from the reverberation component before the speech recognition process. By extracting position information in advance and using it to adjust the acoustic model beforehand, the system prepares the optimal recognition parameters for the specific situation. This preliminary action resolves the contradiction by enabling accurate position-based adaptation without adding significant complexity to the main recognition process.

Inventive Principle:
Principle #10Preliminary action

Data Source

PatentUS9972315B2Speech processing device, speech processing method, and speech processing system
Publication Date: 2018.05.15 HONDA MOTOR CO LTD
  • US9972315B2 patent drawing
  • US9972315B2 patent drawing
  • US9972315B2 patent drawing

AI summary

A speech processing device includes a sound source localization unit configured to determine a sound source position from acquired speech, a reverberation suppression unit configured to suppress a reverberation component of the speech to generate dereverberated speech, a feature quantity calculation unit configured to calculate a feature quantity of the dereverberated speech, a feature quantity adjustment unit configured to multiply the feature quantity by an adjustment factor corresponding to the sound source position to calculate an adjusted feature quantity, and a speech recognition unit configured to perform speech recognition using the adjusted feature quantity.