User Speech Detection with Human-Labeled Background Speech Training
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Voice-driven systems struggle to distinguish user speech from background speech, particularly in noisy environments with public address systems, leading to interference and erroneous interpretations.
Innovation Solution
A voice-data model is created in a training environment that includes both desired user speech and unwanted background sounds, using human listeners to tag and train a processor-based learning system to differentiate between user and background speech.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Measurement precision
If the system captures all speech in the environment, then the system can detect user speech, but background speech interferes and causes erroneous interpretations
Solution Approach 1:
The patent segments the audio signal into multiple components using source separation techniques. The processor divides the mixed audio signal into a first audio signal (user speech) and a second audio signal (background speech), allowing independent processing and analysis of each component to improve detection accuracy while eliminating interference.
Solution Approach 2:
The patent introduces an intermediary classification process that analyzes acoustic characteristics, spatial information, and temporal patterns to mediate between the raw audio signal and the final speech recognition output. This intermediary layer filters and categorizes different speech sources before processing, preventing background speech from causing erroneous interpretations.
2Adaptability or versatility
If the system uses traditional speech recognition, then the system can process voice input, but it cannot distinguish user speech from background speech in noisy environments
Solution Approach 1:
The patent implements dynamic adaptation by continuously learning and updating speaker profiles based on environmental acoustic characteristics. The system adjusts its speech separation and recognition parameters in real-time according to the specific noise environment, enabling reliable operation across diverse noisy conditions while maintaining high accuracy.
Solution Approach 2:
The patent changes multiple parameters including acoustic feature extraction parameters, spatial filtering parameters, and temporal analysis parameters to optimize performance in different noise environments. By dynamically adjusting these parameters based on environmental conditions, the system achieves both noise tolerance and interpretation accuracy.
3Ease of operation
If the system processes all audio signals equally, then the system simplifies processing, but it cannot identify which speech originates from the user
Solution Approach 1:
The patent performs preliminary classification and segmentation of audio signals before full speech recognition processing. By pre-identifying and separating potential user speech from background speech using acoustic fingerprinting and spatial analysis, the system simplifies subsequent processing while enabling accurate source identification.
Solution Approach 2:
The patent adds multiple dimensions to speech analysis including spatial dimension (microphone array positioning), temporal dimension (speech timing patterns), and acoustic characteristic dimension (frequency spectrum analysis). This multi-dimensional approach enables source identification without excessively increasing processing complexity, as each dimension provides complementary information.
Data Source
AI summary
A device, system, and method whereby a speech-driven system can distinguish speech obtained from users of the system from other speech spoken by background persons, as well as from background speech from public address systems. In one aspect, the present system and method prepares, in advance of field-use, a voice-data file which is created in a training environment. The training environment exhibits both desired user speech and unwanted background speech, including unwanted speech from persons other than a user and also speech from a PA system. The speech recognition system is trained or otherwise programmed to identify wanted user speech which may be spoken concurrently with the background sounds. In an embodiment, during the pre-field-use phase the training or programming may be accomplished by having persons who are training listeners audit the pre-recorded sounds to identify the desired user speech. A processor-based learning system is trained to duplicate the assessments made by the human listeners.


