Adversarial Audio Perturbation for Smart Speaker Attacks
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Automated speech recognition devices, particularly those using neural networks, are vulnerable to inconspicuous audio adversarial attacks, where malicious inputs can deceive the system without being perceived by humans, posing threats such as unauthorized device control or transaction activation.
Innovation Solution
A method for over-the-air black-box attacks on smart speakers involves sending an audio file, retrieving information from the device, perturbing the file using an evolutionary algorithm to meet predetermined criteria, and iteratively adjusting the audio to achieve the desired attack, utilizing hardware and software coordination to optimize the adversarial perturbations.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Reliability
If adversarial perturbations are added to audio input to mislead speech recognition, then the attack effectiveness is improved, but the imperceptibility to humans deteriorates
Solution Approach 1:
The patent transforms the audio perturbation from the time domain to the frequency domain using Fourier transform, and applies psychoacoustic masking thresholds to control the magnitude of perturbations at different frequencies. This parameter transformation allows the attack to remain effective while staying below human perception thresholds.
Solution Approach 2:
The patent introduces psychoacoustic masking as an intermediary mechanism that mediates between the adversarial perturbation and human perception. The masking threshold acts as a filter that allows perturbations to pass through without being perceived by humans, while still being effective against the speech recognition system.
2Measurement precision
If gradient-based optimization is used to craft adversarial examples, then the attack precision is improved, but the computational complexity increases
Solution Approach 1:
The patent replaces the traditional gradient-based optimization mechanism with an evolutionary algorithm that uses mutation and selection operations. This substitution maintains attack precision while reducing computational complexity and avoiding the need for gradient calculations.
Solution Approach 2:
The patent introduces dynamic mutation rates and adaptive selection pressures in the evolutionary algorithm, allowing the optimization process to adapt its behavior based on the current state of the population. This dynamic approach achieves high attack precision without requiring exhaustive computational resources.
3Object-affected harmful factors
If multi-objective optimization is applied to balance text dissimilarity and acoustic similarity, then the imperceptibility is improved, but the optimization difficulty increases
Solution Approach 1:
The patent uses psychoacoustic masking thresholds as an intermediary constraint that automatically enforces imperceptibility without requiring direct optimization of multiple conflicting objectives. The masking threshold acts as a hard constraint that simplifies the optimization landscape.
Solution Approach 2:
The patent segments the optimization problem into two independent parts: (1) generating adversarial perturbations that are effective against the speech recognition system, and (2) filtering these perturbations through the psychoacoustic masking threshold to ensure imperceptibility. This segmentation reduces optimization difficulty by decoupling the conflicting objectives.
Data Source
Figure 1
Figure 2~3
Figure 4~5
AI summary
The invention provides a method for over-the-air attacks on an automated speech recognition device, the method comprising: a) sending an audio file to a speaker; b) receiving retrieving information associated with an over-the-air recording from an automated speech recognition device, wherein the recording is associated with the audio file sent to the speaker; c) evaluating the received recording against a predetermined criterion; d) pertubating the audio file to generate a new audio file if the predetermined criterion is not reached; and e) repeating steps a) to e) with the new audio file until the predetermined criterion is reached.