Trigger Phrase Enrollment Quality Assessment for Speech Recognition

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Current speech recognition systems, particularly in always-on audio devices, face challenges in accurately recognizing trigger phrases due to background noise and poor enrollment quality, leading to degraded recognition accuracy.

Innovation Solution

An electronic device evaluates the quality of a spoken trigger phrase by measuring characteristics such as background noise level, length, and noise variability, and rejects the phrase if it does not meet predetermined thresholds, ensuring high-quality enrollment for improved recognition models.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Ease of operation

If the device continuously listens for trigger phrases in always-on mode, then user convenience is improved by eliminating manual activation, but speech recognition accuracy deteriorates due to background noise and false triggers

Engineering Contradiction:
Improveuser convenienceVSAvoidspeech recognition accuracy
Core Design Contradiction:
Ease of operationVSReliability

Solution Approach 1:

The system performs preliminary voice activity detection and background noise assessment before full speech recognition processing. The device evaluates whether ambient noise levels are acceptable for accurate recognition, and only then activates the full speech recognition engine, preventing wasted processing on noisy segments

Inventive Principle:
Principle #10Preliminary action

Solution Approach 2:

A noise assessment module acts as an intermediary between the always-on microphone and the speech recognition engine. This intermediary evaluates background noise characteristics and controls when the speech recognition engine should be activated, filtering out poor-quality audio segments before they reach the recognition system

Inventive Principle:
Principle #24Intermediary (Mediator)

2Reliability

If the trigger phrase recognizer is trained with multiple utterances, then recognition accuracy improves through model adaptation, but the enrollment time and complexity increase

Engineering Contradiction:
Improvetrigger phrase recognition accuracyVSAvoidenrollment time
Core Design Contradiction:
ReliabilityVSLoss of time

Solution Approach 1:

The system collects more utterances than the minimum required for model training. By gathering excessive samples during enrollment, the system ensures sufficient data for robust model adaptation while maintaining a user-friendly process that doesn't feel overly lengthy to the end user

Inventive Principle:
Principle #16Partial or excessive action

Solution Approach 2:

The enrollment process is designed to be autonomous and self-guided, with the system automatically managing recording, evaluation, and model training without requiring user intervention beyond providing the voice samples. The system self-assesses recording quality and requests re-recordings only when necessary

Inventive Principle:
Principle #25Self-service

3Productivity

If the system accepts all recorded trigger phrases for training, then enrollment process simplifies and speeds up, but model quality deteriorates due to inclusion of noisy or poor-quality recordings

Engineering Contradiction:
Improveenrollment speedVSAvoidmodel quality
Core Design Contradiction:
ProductivityVSManufacturing precision

Solution Approach 1:

The system implements real-time feedback during the enrollment process by continuously assessing recording quality metrics such as signal-to-noise ratio, voice activity duration, and noise variability. Based on this feedback, the system immediately determines whether each recording should be accepted or rejected, allowing users to re-record problematic samples without delaying the overall process unnecessarily

Inventive Principle:
Principle #23Feedback

Solution Approach 2:

The system dynamically adjusts evaluation thresholds and parameters based on the specific recording conditions. For example, it modifies the minimum voice activity duration requirement or noise level tolerance depending on the ambient environment detected during each recording attempt, optimizing the balance between acceptance rate and quality

Inventive Principle:
Principle #35Parameter changes

4Adaptability or versatility

If the device processes speech recognition in noisy environments, then system versatility improves, but recognition accuracy deteriorates due to background noise interference

Engineering Contradiction:
Improveenvironmental adaptabilityVSAvoidrecognition accuracy
Core Design Contradiction:
Adaptability or versatilityVSReliability

Solution Approach 1:

The system dynamically adjusts its operation mode based on real-time noise assessment. In low-noise environments, it performs full speech recognition processing, while in high-noise environments, it either requests the user to speak again or activates noise suppression algorithms, creating a dynamic response that adapts to changing acoustic conditions

Inventive Principle:
Principle #15Dynamics

Data Source

PatentUS11676581B2Method and apparatus for evaluating trigger phrase enrollment
Publication Date: 2023.06.13 GOOGLE TECHNOLOGY HOLDINGS LLC
  • US11676581B2 patent drawing
  • US11676581B2 patent drawing
  • US11676581B2 patent drawing

AI summary

An electronic device includes a microphone that receives an audio signal that includes a spoken trigger phrase, and a processor that is electrically coupled to the microphone. The processor measures characteristics of the audio signal, and determines, based on the measured characteristics, whether the spoken trigger phrase is acceptable for trigger phrase model training. If the spoken trigger phrase is determined not to be acceptable for trigger phrase model training, the processor rejects the trigger phrase for trigger phrase model training.