Frequency Band Weighting for Speech Discrimination
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Existing speech recognition systems face challenges in accurately discriminating user speech from disturbance sounds, such as system noises, as current methods often exclude frequency bands containing both user speech and system sounds, leading to reduced accuracy in speech/non-speech discrimination.
Innovation Solution
An apparatus that assigns weights to each frequency band based on the amplitude of acoustic signals from a main microphone and a sub microphone, excluding frequency bands with high disturbance sound presence, allowing for feature extraction that minimizes the exclusion of user speech elements, using a weight assignment unit, feature extraction unit, and speech/non-speech discrimination unit.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Reliability
If the frequency band including the main element of the system sound is excluded, then the robustness for system sound is improved, but the main element of user speech is also excluded, causing accuracy to discriminate speech/non-speech to fall
Solution Approach 1:
The patent applies local quality by assigning different weights to different frequency bands based on their characteristics. Instead of uniformly excluding entire frequency bands, the system calculates weights for each frequency band individually, allowing selective suppression of system sound frequencies while preserving user speech frequencies. This is achieved through the weight calculation unit that computes weights based on the ratio of system sound power to total power in each frequency band.
Solution Approach 2:
The patent changes the parameter of frequency band selection from binary exclusion/inclusion to continuous weighting. By introducing weight values that can vary continuously based on the power ratio of system sound to total sound in each frequency band, the system can dynamically adjust the degree of suppression for each frequency band, thereby maintaining speech discrimination accuracy while improving robustness against system sounds.
2Object-affected harmful factors
If a frequency band is excluded based on system sound spectrum, then the influence of disturbance sound is reduced, but user speech elements in the same frequency band are also removed
Solution Approach 1:
The patent applies local quality by treating each frequency band independently with its own weight calculation. Instead of applying a global exclusion policy, the system evaluates each frequency band locally based on the ratio of system sound power to total power in that specific band. This allows the system to suppress disturbance sounds in frequencies where they dominate while preserving frequencies where user speech is present.
Solution Approach 2:
The patent replaces the mechanical binary exclusion system with a weighted filtering system. Instead of simply excluding or including entire frequency bands, the system uses continuous weight values to gradually suppress or preserve different frequency components. This substitution allows for more nuanced control over which frequency components are suppressed and which are preserved.
Data Source
AI summary
According to one embodiment, an apparatus for discriminating speech/non-speech of a first acoustic signal includes a weight assignment unit, a feature extraction unit, and a speech/non-speech discrimination unit. The weight assignment unit is configured to assign a weight to each frequency band, based on a frequency spectrum of the first acoustic signal including a user's speech and a frequency spectrum of a second acoustic signal including a disturbance sound. The feature extraction unit is configured to extract a feature from the frequency spectrum of the first acoustic signal, based on the weight of each frequency band. The speech/non-speech discrimination unit is configured to discriminate speech/non-speech of the first acoustic signal, based on the feature.


