Selective Noise Reduction for Invocation Phrase Detection

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Existing automated assistants often struggle with accurately detecting spoken invocation phrases in environments with strong background noise, leading to poor robustness and accuracy in invocation phrase detection.

Innovation Solution

The implementation of a noise reduction technique that selectively adapts based on audio data frames, using a trained machine learning model to generate output indications for each frame, which are then used to determine whether the audio data frames contain trigger, near-trigger, or noise indications, allowing for adaptive noise cancellation and improved detection of invocation phrases.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Measurement precision

If local processing of audio data is performed without noise reduction, then device complexity and processing speed are maintained, but detection accuracy and robustness deteriorate in noisy environments

Engineering Contradiction:
Improveinvocation phrase detection accuracyVSAvoidnoise reduction processing complexity
Core Design Contradiction:
Measurement precisionVSDevice complexity

Solution Approach 1:

The system performs preliminary classification of audio frames into noise frames and non-noise frames before applying noise reduction. This preliminary action allows the system to prepare noise reduction filters in advance and apply them selectively, improving detection accuracy without unnecessarily increasing processing complexity for all audio data.

Inventive Principle:
Principle #10Preliminary action

Solution Approach 2:

The patent applies noise reduction processing locally only to specific audio frames that are classified as containing potential invocation phrases. Instead of applying noise reduction to all audio data, the system identifies frames with near-trigger indications and applies noise reduction selectively to those frames, thereby improving detection accuracy while minimizing additional processing complexity.

Inventive Principle:
Principle #3Local quality

2Reliability

If noise reduction technique is applied to all audio data, then detection robustness improves, but processing time and computational resources increase

Engineering Contradiction:
Improveinvocation phrase detection robustnessVSAvoidaudio processing time
Core Design Contradiction:
ReliabilityVSLoss of time

Solution Approach 1:

The system applies noise reduction processing partially rather than excessively. It classifies audio frames and applies noise reduction only to frames that contain near-trigger indications, rather than applying noise reduction to all audio data. This partial application maintains detection robustness while reducing overall processing time and computational resource consumption.

Inventive Principle:
Principle #16Partial or excessive action

Solution Approach 2:

The patent segments audio data into individual frames and classifies each frame as either noise or non-noise. This segmentation allows the system to apply noise reduction processing only to specific segments (frames with near-trigger indications) rather than processing the entire audio stream, thereby maintaining robustness while reducing processing time.

Inventive Principle:
Principle #1Segmentation

3Measurement precision

If threshold for invocation phrase detection is lowered to reduce false negatives, then detection sensitivity improves, but false positive rate increases in noisy environments

Engineering Contradiction:
Improveinvocation phrase detection sensitivityVSAvoidfalse positive detections
Core Design Contradiction:
Measurement precisionVSObject-affected harmful factors

Solution Approach 1:

The patent introduces noise reduction processing as an intermediary step between raw audio input and invocation phrase detection. This intermediary process reduces background noise before the detection algorithm processes the audio, allowing the system to use lower detection thresholds without increasing false positives. The noise reduction acts as a mediator that cleans the input signal, improving both sensitivity and accuracy.

Inventive Principle:
Principle #24Intermediary (Mediator)

Solution Approach 2:

The system performs preliminary anti-action by reducing noise before detection. By applying noise reduction processing to audio frames before attempting to detect invocation phrases, the system counteracts the harmful effect of background noise in advance. This allows for more accurate detection with reduced false positives, as the noise has already been attenuated before the detection threshold is applied.

Inventive Principle:
Principle #9Preliminary anti-action

Data Source

PatentUS12260857B2Selective adaptation and utilization of noise reduction technique in invocation phrase detection
Publication Date: 2025.03.25 GOOGLE LLC
  • US12260857B2 patent drawing
  • US12260857B2 patent drawing
  • US12260857B2 patent drawing

AI summary

Techniques are described for selectively adapting and/or selectively utilizing a noise reduction technique in detection of one or more features of a stream of audio data frames. For example, various techniques are directed to selectively adapting and/or utilizing a noise reduction technique in detection of an invocation phrase in a stream of audio data frames, detection of voice characteristics in a stream of audio data frames (e.g., for speaker identification), etc. Utilization of described techniques can result in more robust and/or more accurate detections of features of a stream of audio data frames in various situations, such as in environments with strong background noise. In various implementations, described techniques are implemented in combination with an automated assistant, and feature(s) detected utilizing techniques described herein are utilized to adapt the functionality of the automated assistant.