SVD Voice Recognition Subspace Segmentation for Multi-Talker Noise

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Voice recognition systems in vehicles face challenges in accurately recognizing user voice commands in noisy environments with multiple speakers, particularly when background voices, such as children's voices, interfere with the recognition process.

Innovation Solution

A method and system that decompose digitized speech energy into features and project them into a feature space with multiple speaker subspaces, allowing for speech recognition operations to be performed on a selected subspace to resolve the utterance, enhancing accuracy and robustness in noisy conditions.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Measurement precision

If traditional voice recognition systems are used in vehicle environments, then the system can operate with basic functionality, but the recognition accuracy deteriorates due to background noise and multiple speakers

Engineering Contradiction:
Improvevoice recognition accuracyVSAvoidbackground noise interference
Core Design Contradiction:
Measurement precisionVSObject-affected harmful factors

Solution Approach 1:

The patent segments the mixed speech signal into multiple speaker-specific subspaces using singular value decomposition (SVD). The feature space is divided into N subspaces corresponding to N different speakers, allowing the system to separate and process each speaker's contribution independently. This segmentation enables accurate identification of the target speaker's voice even in noisy multi-talker environments.

Inventive Principle:
Principle #1Segmentation

Solution Approach 2:

The patent introduces an intermediary processing layer between signal reception and recognition that projects the mixed speech features into speaker-specific subspaces. This intermediary transformation uses SVD to create a bridge that separates the mixed signal components, allowing the recognition system to operate on purified speaker-specific features rather than the raw mixed signal.

Inventive Principle:
Principle #24Intermediary (Mediator)

2Measurement precision

If the system processes all speech energy equally, then the processing is simple, but the system cannot distinguish between multiple speakers and background noise

Engineering Contradiction:
Improvespeaker distinction capabilityVSAvoidsignal processing complexity
Core Design Contradiction:
Measurement precisionVSDevice complexity

Solution Approach 1:

The patent applies segmentation by dividing the feature space into multiple speaker-specific subspaces using SVD. Each subspace corresponds to a different speaker and is characterized by its own singular vectors. This mathematical segmentation transforms the complex mixed signal into separable components that can be individually processed with simpler algorithms.

Inventive Principle:
Principle #1Segmentation

Solution Approach 2:

The patent changes the parameter representation of the speech signal by transforming it from the original feature space into a decomposed subspace representation using SVD. This parameter transformation reveals the underlying speaker-specific structures in the data, enabling the system to distinguish between multiple speakers through mathematical decomposition rather than complex pattern recognition.

Inventive Principle:
Principle #35Parameter changes

3Adaptability or versatility

If voice recognition is performed in child voice environments, then the system must handle vast frequency differences, but recognition accuracy deteriorates due to frequency variation

Engineering Contradiction:
Improvefrequency adaptation capabilityVSAvoidrecognition accuracy with child voices
Core Design Contradiction:
Adaptability or versatilityVSMeasurement precision

Solution Approach 1:

The patent applies parameter changes by using SVD to transform the speech features into a new representation space that is invariant to speaker-specific frequency characteristics. The singular value decomposition separates the frequency variations into orthogonal components, allowing the system to recognize speech content regardless of whether the speaker is an adult or child with different fundamental frequencies.

Inventive Principle:
Principle #35Parameter changes

Solution Approach 2:

The patent creates a universal processing framework that handles multiple speaker types (adults, children, different genders) through a single SVD-based subspace decomposition approach. The system does not require separate processing paths for different speaker categories; instead, the mathematical decomposition automatically adapts to handle the vast frequency differences across all speaker types uniformly.

Inventive Principle:
Principle #6Universality (Multi-functionality)

Data Source

PatentUS9177557B2Singular value decomposition for improved voice recognition in presence of multi-talker background noise
Publication Date: 2015.11.03 GENERAL MOTORS LLC
  • US9177557B2 patent drawing
  • US9177557B2 patent drawing
  • US9177557B2 patent drawing

AI summary

A system and method for providing speech recognition functionality offers improved accuracy and robustness in noisy environments having multiple speakers. The described technique includes receiving speech energy and converting the received speech energy to a digitized form. The digitized speech energy is decomposed into features that are then projected into a feature space having multiple speaker subspaces. The projected features fall either into one of the multiple speaker subspaces or outside of all speaker subspaces. A speech recognition operation is performed on a selected one of the multiple speaker subspaces to resolve the utterance to a command or data.