Selective Noise Reduction for Invocation Phrase Detection
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Existing automated assistants often struggle with accurately detecting spoken invocation phrases in environments with strong background noise, leading to poor robustness and accuracy in invocation phrase detection.
Innovation Solution
The implementation of a noise reduction technique that selectively adapts based on audio data frames, using a trained machine learning model to generate output indications for each frame, which are then used to determine whether the audio data frames contain trigger, near-trigger, or noise indications, allowing for adaptive noise cancellation and improved detection of invocation phrases.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Measurement precision
If local processing of audio data is performed without noise reduction, then device complexity and processing speed are maintained, but detection accuracy and robustness deteriorate in noisy environments
Solution Approach 1:
The system performs preliminary classification of audio frames into noise frames and non-noise frames before applying noise reduction. This preliminary action allows the system to prepare noise reduction filters in advance and apply them selectively, improving detection accuracy without unnecessarily increasing processing complexity for all audio data.
Solution Approach 2:
The patent applies noise reduction processing locally only to specific audio frames that are classified as containing potential invocation phrases. Instead of applying noise reduction to all audio data, the system identifies frames with near-trigger indications and applies noise reduction selectively to those frames, thereby improving detection accuracy while minimizing additional processing complexity.
2Reliability
If noise reduction technique is applied to all audio data, then detection robustness improves, but processing time and computational resources increase
Solution Approach 1:
The system applies noise reduction processing partially rather than excessively. It classifies audio frames and applies noise reduction only to frames that contain near-trigger indications, rather than applying noise reduction to all audio data. This partial application maintains detection robustness while reducing overall processing time and computational resource consumption.
Solution Approach 2:
The patent segments audio data into individual frames and classifies each frame as either noise or non-noise. This segmentation allows the system to apply noise reduction processing only to specific segments (frames with near-trigger indications) rather than processing the entire audio stream, thereby maintaining robustness while reducing processing time.
3Measurement precision
If threshold for invocation phrase detection is lowered to reduce false negatives, then detection sensitivity improves, but false positive rate increases in noisy environments
Solution Approach 1:
The patent introduces noise reduction processing as an intermediary step between raw audio input and invocation phrase detection. This intermediary process reduces background noise before the detection algorithm processes the audio, allowing the system to use lower detection thresholds without increasing false positives. The noise reduction acts as a mediator that cleans the input signal, improving both sensitivity and accuracy.
Solution Approach 2:
The system performs preliminary anti-action by reducing noise before detection. By applying noise reduction processing to audio frames before attempting to detect invocation phrases, the system counteracts the harmful effect of background noise in advance. This allows for more accurate detection with reduced false positives, as the noise has already been attenuated before the detection threshold is applied.
Data Source
AI summary
Techniques are described for selectively adapting and/or selectively utilizing a noise reduction technique in detection of one or more features of a stream of audio data frames. For example, various techniques are directed to selectively adapting and/or utilizing a noise reduction technique in detection of an invocation phrase in a stream of audio data frames, detection of voice characteristics in a stream of audio data frames (e.g., for speaker identification), etc. Utilization of described techniques can result in more robust and/or more accurate detections of features of a stream of audio data frames in various situations, such as in environments with strong background noise. In various implementations, described techniques are implemented in combination with an automated assistant, and feature(s) detected utilizing techniques described herein are utilized to adapt the functionality of the automated assistant.


