Voice Authentication Noise Separation via Speech Extraction

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Voice authentication systems face errors due to intrinsic mismatches caused by noise differences between enrollment and authentication phases, particularly in high-noise environments.

Innovation Solution

Generating clean speech statistics during enrollment and using them to separate speech and noise data during authentication, or generating noisy speech statistics by combining noise data with enrollment speech data to compare with noisy input, thereby reducing authentication errors.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Adaptability or versatility

If voice authentication is performed by comparing speech models generated in different noise environments, then the system can operate in various conditions, but authentication accuracy deteriorates due to intrinsic mismatches caused by noise differences

Engineering Contradiction:
Improveoperational conditionsVSAvoidauthentication accuracy
Core Design Contradiction:
Adaptability or versatilityVSReliability

Solution Approach 1:

The patent segments the speech signal into speech components and noise components separately. By dividing the noisy speech signal into distinct speech and noise parts, the system can process each component independently, allowing authentication to be performed on the clean speech portion while maintaining adaptability to various noise environments.

Inventive Principle:
Principle #1Segmentation

Solution Approach 2:

The patent extracts the speech signal from the noisy speech signal by removing the noise component. This extraction process isolates the clean speech data needed for authentication, eliminating the harmful noise effects while preserving the ability to operate in noisy conditions.

Inventive Principle:
Principle #2Taking out (Extraction)

2Reliability

If clean speech statistics are used for authentication comparison, then authentication accuracy improves, but the system becomes more complex requiring additional processing steps

Engineering Contradiction:
Improveauthentication accuracyVSAvoidprocessing complexity
Core Design Contradiction:
ReliabilityVSDevice complexity

Solution Approach 1:

The patent performs preliminary noise estimation and speech-noise separation during the enrollment phase to generate clean speech statistics. By preparing the clean speech model in advance under controlled conditions, the system establishes a reliable reference for later authentication without requiring complex real-time processing during the actual authentication event.

Inventive Principle:
Principle #10Preliminary action

Solution Approach 2:

The patent introduces noise estimation as an intermediary step that bridges the noisy speech signal and the clean speech model. This intermediary noise model allows the system to account for environmental noise effects while maintaining the simplicity of comparing against a pre-established clean speech reference.

Inventive Principle:
Principle #24Intermediary (Mediator)

Data Source

PatentUS10720165B2Keyword voice authentication
Publication Date: 2020.07.21 QUALCOMM INC
  • US10720165B2 patent drawing
  • US10720165B2 patent drawing
  • US10720165B2 patent drawing

AI summary

A method of authenticating a user based on voice recognition of a keyword includes generating, at a processor, clean speech statistics. The clean speech statistics are generated from an audio recording of the keyword spoken by the user during an enrollment phase. The method further includes separating speech data and noise data from noisy input speech using the clean speech statistics during an authentication phase. The method also includes authenticating the user by comparing the speech data to the clean speech statistics or by comparing the noisy input speech to noisy speech statistics. The noisy speech statistics are based at least in part on the noise data.