User Speech Detection with Human-Labeled Background Speech Training

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Voice-driven systems struggle to distinguish user speech from background speech, particularly in noisy environments with public address systems, leading to interference and erroneous interpretations.

Innovation Solution

A voice-data model is created in a training environment that includes both desired user speech and unwanted background sounds, using human listeners to tag and train a processor-based learning system to differentiate between user and background speech.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Measurement precision

If the system captures all speech in the environment, then the system can detect user speech, but background speech interferes and causes erroneous interpretations

Engineering Contradiction:
Improvespeech detection accuracyVSAvoidbackground speech interference
Core Design Contradiction:
Measurement precisionVSObject-affected harmful factors

Solution Approach 1:

The patent segments the audio signal into multiple components using source separation techniques. The processor divides the mixed audio signal into a first audio signal (user speech) and a second audio signal (background speech), allowing independent processing and analysis of each component to improve detection accuracy while eliminating interference.

Inventive Principle:
Principle #1Segmentation

Solution Approach 2:

The patent introduces an intermediary classification process that analyzes acoustic characteristics, spatial information, and temporal patterns to mediate between the raw audio signal and the final speech recognition output. This intermediary layer filters and categorizes different speech sources before processing, preventing background speech from causing erroneous interpretations.

Inventive Principle:
Principle #24Intermediary (Mediator)

2Adaptability or versatility

If the system uses traditional speech recognition, then the system can process voice input, but it cannot distinguish user speech from background speech in noisy environments

Engineering Contradiction:
Improvenoise environment toleranceVSAvoidspeech interpretation accuracy
Core Design Contradiction:
Adaptability or versatilityVSReliability

Solution Approach 1:

The patent implements dynamic adaptation by continuously learning and updating speaker profiles based on environmental acoustic characteristics. The system adjusts its speech separation and recognition parameters in real-time according to the specific noise environment, enabling reliable operation across diverse noisy conditions while maintaining high accuracy.

Inventive Principle:
Principle #15Dynamics

Solution Approach 2:

The patent changes multiple parameters including acoustic feature extraction parameters, spatial filtering parameters, and temporal analysis parameters to optimize performance in different noise environments. By dynamically adjusting these parameters based on environmental conditions, the system achieves both noise tolerance and interpretation accuracy.

Inventive Principle:
Principle #35Parameter changes

3Ease of operation

If the system processes all audio signals equally, then the system simplifies processing, but it cannot identify which speech originates from the user

Engineering Contradiction:
Improvespeech source identificationVSAvoidsignal processing complexity
Core Design Contradiction:
Ease of operationVSDevice complexity

Solution Approach 1:

The patent performs preliminary classification and segmentation of audio signals before full speech recognition processing. By pre-identifying and separating potential user speech from background speech using acoustic fingerprinting and spatial analysis, the system simplifies subsequent processing while enabling accurate source identification.

Inventive Principle:
Principle #10Preliminary action

Solution Approach 2:

The patent adds multiple dimensions to speech analysis including spatial dimension (microphone array positioning), temporal dimension (speech timing patterns), and acoustic characteristic dimension (frequency spectrum analysis). This multi-dimensional approach enables source identification without excessively increasing processing complexity, as each dimension provides complementary information.

Inventive Principle:
Principle #17Another dimension (Dimensionality change)

Data Source

PatentUS12400678B2Distinguishing user speech from background speech in speech-dense environments
Publication Date: 2025.08.26 VOCOLLECT INC
  • US12400678B2 patent drawing
  • US12400678B2 patent drawing
  • US12400678B2 patent drawing

AI summary

A device, system, and method whereby a speech-driven system can distinguish speech obtained from users of the system from other speech spoken by background persons, as well as from background speech from public address systems. In one aspect, the present system and method prepares, in advance of field-use, a voice-data file which is created in a training environment. The training environment exhibits both desired user speech and unwanted background speech, including unwanted speech from persons other than a user and also speech from a PA system. The speech recognition system is trained or otherwise programmed to identify wanted user speech which may be spoken concurrently with the background sounds. In an embodiment, during the pre-field-use phase the training or programming may be accomplished by having persons who are training listeners audit the pre-recorded sounds to identify the desired user speech. A processor-based learning system is trained to duplicate the assessments made by the human listeners.