Voice Assistant Audio Sampling for Recorded Audio Detection

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Voice recognition systems often confuse previously recorded audio played through a speaker with real voice input, leading to unintended triggering of voice assistants, which can result in undesirable commands, security breaches, and resource consumption.

Innovation Solution

Implementing high-frequency sampling of audio signals beyond the expected finite sampling frequency of previously recorded audio to differentiate between real voice input and recorded audio, analyzing the sampled signal for artifacts to determine if it was played through a speaker or spoken directly, and adjusting the voice assistant's activation accordingly.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Measurement precision

If voice assistants use standard sampling frequency (44.1kHz or 48kHz) for audio capture, then the system complexity is reduced and processing is simplified, but the ability to distinguish between real voice and recorded audio played through speakers deteriorates

Engineering Contradiction:
Improveaudio differentiation precisionVSAvoidsampling system complexity
Core Design Contradiction:
Measurement precisionVSDevice complexity

Solution Approach 1:

The patent applies parameter changes by switching from standard sampling frequencies (44.1kHz or 48kHz) to high-frequency sampling (greater than 96kHz). This parameter change enables the voice assistant to detect artifacts in recorded audio that are invisible at lower sampling rates, thereby improving the ability to distinguish between real voice and played-back audio without significantly increasing system complexity

Inventive Principle:
Principle #35Parameter changes

2Ease of operation

If voice assistants continuously monitor for trigger words to provide responsive service, then user convenience and responsiveness are improved, but energy consumption and system resource usage increase

Engineering Contradiction:
Improvevoice assistant responsivenessVSAvoidprocessor energy consumption
Core Design Contradiction:
Ease of operationVSUse of energy by moving object

Solution Approach 1:

The patent applies preliminary action by pre-filtering audio signals using high-frequency sampling and artifact detection before the voice assistant fully processes and responds to commands. This preliminary action identifies recorded audio artifacts early in the processing pipeline, allowing the system to avoid unnecessary full processing cycles for played-back audio, thus reducing overall energy consumption while maintaining responsiveness for genuine voice commands

Inventive Principle:
Principle #10Preliminary action

3Productivity

If the voice assistant triggers functions based on detected keywords, then the system responds to user commands efficiently, but false triggering from recorded audio leads to errors and security issues

Engineering Contradiction:
Improvecommand execution efficiencyVSAvoidtrigger accuracy
Core Design Contradiction:
ProductivityVSReliability

Solution Approach 1:

The patent applies feedback by continuously monitoring audio characteristics during the listening phase and using high-frequency sampling to detect artifacts that indicate recorded audio. This feedback mechanism allows the voice assistant to adjust its triggering behavior in real-time, suppressing false triggers from recorded audio while maintaining efficient response to genuine commands, thereby improving both productivity and reliability

Inventive Principle:
Principle #23Feedback

Applied Scientific Principles

This section explains which scientific principles are used to turn an abstract innovation direction into a practical engineering solution.

Function Achieved in This Case

Effectively prevents unintended triggering of voice assistants by accurately distinguishing between real voice and recorded audio, reducing errors and resource consumption while enhancing security systems.

Implementation Method 1

an audio signal may be captured by sampling a microphone of the voice assistant at a sampling frequency that is higher than an expected finite sampling frequency

Methodology Applied
Scientific EffectSound wave sampling:

Data Source

PatentUS20240282308A1Systems and methods for avoiding inadvertently triggering a voice assistant
Publication Date: 2024.08.22 ADEIA GUIDES INC
  • US20240282308A1 patent drawing
  • US20240282308A1 patent drawing
  • US20240282308A1 patent drawing

AI summary

Systems and methods are provided herein for avoiding inadvertently trigging a voice assistant with audio played through a speaker. An audio signal is captured by sampling a microphone of the voice assistant at a sampling frequency that is higher than an expected finite sampling frequency of previously recorded audio played through the speaker to generate a voice data sample. A quality metric of the generated voice data sample is calculated by determining whether the generated voice data sample comprises artifacts resulting from previous compression or approximation by the expected finite sampling frequency. Based on the calculated quality metric, it is determined whether the captured audio signal is previously recorded audio played through the speaker. Responsive to the determination that the captured audio signal is previously recorded audio played through the speaker, the voice assistant refrains from being activated.