Voice Recognition Noise Superimposition for Multi-Speaker Precision

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Conventional voice recognition systems face challenges in improving recognition precision when multiple speakers utter simultaneously, as interference voices can lead to distorted input and improper recognition results due to inadequate noise suppression, especially when ambient noise is weaker than the utterance voice.

Innovation Solution

A voice recognizing apparatus that includes a superimposition amount determining unit to estimate and reduce interference components by superimposing a reduction signal, specifically white noise, onto the input voice based on the position and distance of speakers, thereby improving recognition precision and reducing improper recognition results.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Measurement precision

If white noise is superimposed onto the input voice to suppress ambient noise influence, then voice recognition precision is improved, but when the power of white noise is increased excessively, distortion in the input voice increases and recognition precision deteriorates

Engineering Contradiction:
Improvevoice recognition precisionVSAvoidinput voice distortion
Core Design Contradiction:
Measurement precisionVSManufacturing precision

Solution Approach 1:

The patent dynamically adjusts the white noise superimposition power based on the voice-likeness level of each frame. When voice-likeness is low (indicating ambient noise dominance), higher noise power is applied to suppress noise. When voice-likeness is high, the noise power is reduced to avoid distortion. This parameter adaptation resolves the contradiction between noise suppression effectiveness and voice distortion.

Inventive Principle:
Principle #35Parameter changes

Solution Approach 2:

The patent implements dynamic control of the white noise superimposition process by continuously monitoring voice-likeness levels across different frames. The noise power is not fixed but varies in real-time based on the detected speech characteristics, allowing the system to adapt to changing acoustic environments and maintain optimal recognition precision without excessive distortion.

Inventive Principle:
Principle #15Dynamics

2Object-affected harmful factors

If conventional white noise superimposition is applied to multi-speaker environments, then ambient noise suppression is achieved, but interference utterance voices cannot be distinguished and improper recognition results occur

Engineering Contradiction:
Improveambient noise suppressionVSAvoidinterference voice distinction capability
Core Design Contradiction:
Object-affected harmful factorsVSMeasurement precision

Solution Approach 1:

The patent applies different noise superimposition strategies to different temporal regions based on voice-likeness detection. By analyzing the spectral characteristics of each frame, the system identifies whether the dominant sound is ambient noise or interference voice, and applies appropriate noise power accordingly. This local differentiation allows the system to suppress ambient noise while preserving interference voice information for proper recognition.

Inventive Principle:
Principle #3Local quality

Solution Approach 2:

The patent incorporates feedback from voice-likeness detection and interference voice identification into the noise superimposition process. The system continuously monitors the input signal, detects the presence and characteristics of interference voices, and adjusts the noise power in real-time. This feedback mechanism prevents the system from misinterpreting interference voices as ambient noise, thereby avoiding improper recognition results.

Inventive Principle:
Principle #23Feedback

Data Source

PatentUS8150688B2Voice recognizing apparatus, voice recognizing method, voice recognizing program, interference reducing apparatus, interference reducing method, and interference reducing program
Publication Date: 2012.04.03 NEC CORP
  • US8150688B2 patent drawing
  • US8150688B2 patent drawing
  • US8150688B2 patent drawing

AI summary

A voice recognizing apparatus includes a microphone 12 which inputs an input voice including speech voice uttered by a user speaker and interference voice uttered by an interference speaker other than the user speaker, superimposition amount determining unit 14 which determines a noise superimposition amount for the input voice on the basis of a speech voice and an interference voice separately input as the input voice, a noise superimposing unit 16 which superimposes noise according to the noise superimposition amount onto the input voice and outputs the resultant voice as noise-superimposed voice; and a voice recognizing unit 18 which recognizes the noise-superimposed voice.