Voice Assistant Audio Sampling for Recorded Audio Detection
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Voice recognition systems often confuse previously recorded audio played through a speaker with real voice input, leading to unintended triggering of voice assistants, which can result in undesirable commands, security breaches, and resource consumption.
Innovation Solution
Implementing high-frequency sampling of audio signals beyond the expected finite sampling frequency of previously recorded audio to differentiate between real voice input and recorded audio, analyzing the sampled signal for artifacts to determine if it was played through a speaker or spoken directly, and adjusting the voice assistant's activation accordingly.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Measurement precision
If voice assistants use standard sampling frequency (44.1kHz or 48kHz) for audio capture, then the system complexity is reduced and processing is simplified, but the ability to distinguish between real voice and recorded audio played through speakers deteriorates
Solution Approach 1:
The patent applies parameter changes by switching from standard sampling frequencies (44.1kHz or 48kHz) to high-frequency sampling (greater than 96kHz). This parameter change enables the voice assistant to detect artifacts in recorded audio that are invisible at lower sampling rates, thereby improving the ability to distinguish between real voice and played-back audio without significantly increasing system complexity
2Ease of operation
If voice assistants continuously monitor for trigger words to provide responsive service, then user convenience and responsiveness are improved, but energy consumption and system resource usage increase
Solution Approach 1:
The patent applies preliminary action by pre-filtering audio signals using high-frequency sampling and artifact detection before the voice assistant fully processes and responds to commands. This preliminary action identifies recorded audio artifacts early in the processing pipeline, allowing the system to avoid unnecessary full processing cycles for played-back audio, thus reducing overall energy consumption while maintaining responsiveness for genuine voice commands
3Productivity
If the voice assistant triggers functions based on detected keywords, then the system responds to user commands efficiently, but false triggering from recorded audio leads to errors and security issues
Solution Approach 1:
The patent applies feedback by continuously monitoring audio characteristics during the listening phase and using high-frequency sampling to detect artifacts that indicate recorded audio. This feedback mechanism allows the voice assistant to adjust its triggering behavior in real-time, suppressing false triggers from recorded audio while maintaining efficient response to genuine commands, thereby improving both productivity and reliability
Applied Scientific Principles
This section explains which scientific principles are used to turn an abstract innovation direction into a practical engineering solution.
Function Achieved in This Case
Effectively prevents unintended triggering of voice assistants by accurately distinguishing between real voice and recorded audio, reducing errors and resource consumption while enhancing security systems.
Implementation Method 1
an audio signal may be captured by sampling a microphone of the voice assistant at a sampling frequency that is higher than an expected finite sampling frequency
Data Source
AI summary
Systems and methods are provided herein for avoiding inadvertently trigging a voice assistant with audio played through a speaker. An audio signal is captured by sampling a microphone of the voice assistant at a sampling frequency that is higher than an expected finite sampling frequency of previously recorded audio played through the speaker to generate a voice data sample. A quality metric of the generated voice data sample is calculated by determining whether the generated voice data sample comprises artifacts resulting from previous compression or approximation by the expected finite sampling frequency. Based on the calculated quality metric, it is determined whether the captured audio signal is previously recorded audio played through the speaker. Responsive to the determination that the captured audio signal is previously recorded audio played through the speaker, the voice assistant refrains from being activated.


