A method and system for camouflage attack for intelligent voice systems
Patent Information
- Application Number
- CN202311436869.6
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2023-10-31
- Publication Date
- 2026-09-25
- Estimated Expiration
- 2043-10-31
AI Technical Summary
[0005]本发明的目的在于提供一种用于智能语音系统的伪装攻击方法及系统,以克服现有方法存在局限性,普适性低,易使伪装过后的语音信号会导致采样前后语义内容不一致的问题
[0027]本发明提供一种用于智能语音系统的伪装攻击方法,通过将原始信号调整为采样率为r1的信号;根据目标系统采样率,计算重采样算法的阻带,根据获取的阻带,构造频谱在阻带上的噪声信号,将采样率为的信号进行能量缩小后与生成的噪声信号进行相加,从而生成伪装后的信号,本发明只需要了解目标系统的采样率,就可以生成伪装攻击样本,能在不了解模型任何信息情况下即可完成攻击,这大大增加了攻击算法的应用范围,本发明能够将普通的语音信号伪装成电流噪声,达到伪装攻击的目的;验证了所提出的伪装攻击算法对于多种采样算法的有效性和普适性,即只需了解目标算法的输入采样率就可以进行攻击。
Smart Images

Figure CN117524208B_ABST
Abstract
Description
Technical Field
[0001] This invention belongs to the field of deep learning security and relates to a method and system for spoofing attacks on intelligent voice systems. Background Technology
[0002] Voice is a crucial medium for human-computer interaction. With the widespread application of deep learning algorithms, various technologies in the voice field have developed rapidly, such as voice search, smart homes, and voice customer service. These voice technologies have changed the way people interact with everyday smart devices and brought convenience to people's lives. Currently, DNN (Deep Neural Networks) has become a standard feature of intelligent voice system frameworks, which has significantly improved the accuracy of speech recognition. Common speech recognition systems include Kaidl and DeepSpeech.
[0003] However, while deep learning has promoted the development of speech recognition technology, it also presents serious vulnerabilities and various potential security threats. Among these, adversarial examples are the most threatening. These attacks can introduce minute perturbations imperceptible to humans into the original samples, causing deep learning models to produce incorrect recognitions with high confidence. With the continuous development and maturation of adversarial examples, this attack strategy has posed a serious threat to related tasks in the fields of vision and speech. This has led more and more researchers to study adversarial examples to improve the robustness of existing speech models against such attacks. Therefore, researching different speech attack methods is of great significance for improving the security of the speech recognition process and the robustness of speech models.
[0004] Currently, most speech attack algorithms are based on adversarial example attacks, but these algorithms have certain limitations, such as long attack times and limited universality. Resampling algorithms also play an important role in the speech domain, for example, converting speech signals with different sampling rates into a fixed sampling rate that deep models can process. However, traditional speech resampling algorithms do not consider the impact of malicious input; for example, disguised speech signals can lead to inconsistencies in semantic content before and after sampling, thus significantly reducing recognition accuracy. Summary of the Invention
[0005] The purpose of this invention is to provide a spoofing attack method and system for intelligent voice systems, so as to overcome the limitations of existing methods, low universality, and the problem that the spoofed voice signal may cause inconsistencies in semantic content before and after sampling.
[0006] A spoofing attack method for intelligent voice systems includes the following steps:
[0007] S1, adjust the original signal to a signal with a sampling rate of r1;
[0008] S2, Calculate the stopband of the resampling algorithm based on the target system sampling rate;
[0009] S3, Based on the obtained stopband, construct the noise signal with the spectrum in the stopband;
[0010] S4, after reducing the energy of the signal with a sampling rate of 100%, adds it to the generated noise signal to generate the disguised signal.
[0011] Preferably, the obtained stopband is greater than half of the target system sampling rate.
[0012] Preferably, the noise spectrum is set to f a =r² / 2 + 0.5kHz to f b = r2 / 2 + 1KHz.
[0013] Preferably, broadband noise and pure tone are used to disguise the original signal to obtain the noise signal in the stopband of the spectrum.
[0014] Preferably, white noise S[n] ~ N(0,1) with the same length as the original speech signal is generated, and the white noise is normalized to the interval S[n] = S[n] / max(|S|), and then a passband of [f a ,f b By filtering S[n] with the filter of ], broadband noise of the specified frequency band can be obtained.
[0015] Preferably, based on the sampling frequency f s Generate a cosine signal with a specified frequency f using the following formula:
[0016] S[n]=Acos(2π·f·(n / f s )) (1)
[0017] f / f s Known as digital frequency, it is the original signal frequency f divided by the sampling frequency f. s The normalization of the sequence S[n] is such that the length of the sequence is consistent with the length of the masked signal s.
[0018] Preferably, the original signal with a sampling rate of r1 is reduced by a factor of 200.
[0019] A spoofing attack system for intelligent voice systems includes a resampling module, a stopband calculation module, a noise module, and an attack module;
[0020] The resampling module is used to adjust the original signal to a signal with a sampling rate of r1;
[0021] The stopband calculation module is used to calculate the stopband of the resampling algorithm based on the target system sampling rate.
[0022] The noise module constructs a noise signal with a spectrum in the stopband based on the acquired stopband.
[0023] The attack module reduces the energy of the signal with a sampling rate of 100% and adds it to the generated noise signal to generate a disguised signal.
[0024] Preferably, the obtained stopband is greater than half of the target system sampling rate.
[0025] Preferably, broadband noise and pure tone are used to disguise the original signal to obtain the noise signal in the stopband of the spectrum.
[0026] Compared with the prior art, the present invention has the following beneficial technical effects:
[0027] This invention provides a spoofing attack method for intelligent voice systems. The method involves adjusting the original signal to a signal with a sampling rate of r1; calculating the stopband of a resampling algorithm based on the target system's sampling rate; constructing a noise signal with a spectrum in the stopband based on the obtained stopband; and adding the energy-reduced signal with the signal at a sampling rate of r1 to the generated noise signal to generate the spoofed signal. This invention only requires knowledge of the target system's sampling rate to generate spoofing attack samples, enabling attacks without any knowledge of the model. This significantly expands the application range of the attack algorithm. This invention can spoof ordinary voice signals as current noise to achieve the purpose of a spoofing attack. It verifies the effectiveness and universality of the proposed spoofing attack algorithm for various sampling algorithms, i.e., attacks can be performed only by knowing the input sampling rate of the target algorithm.
[0028] This invention accurately reconstructs the statistical features of speech signals through a sampling algorithm, achieving a high recognition accuracy even with extremely low signal-to-noise ratio input noise, thus verifying the effectiveness of the proposed algorithm. Attached Figure Description
[0029] Figure 1 This is a flowchart of a spoofing attack method for an intelligent voice system in an embodiment of the present invention.
[0030] Figure 2 The generation process of spoofed samples is illustrated in the frequency domain in this embodiment of the invention. Figure 2 (a) represents the spectrum of the noise signal. Figure 2 (c) represents the spectrum of the original speech signal after energy reduction. Figure 2 (b) represents the spoofed sample generated by combining the original signal with the noise signal. Figure 2 (d) represents the original speech signal recovered after passing through the anti-aliasing filter.
[0031] Figure 3These are the spectrograms of the spoofed speech signal before and after passing through the anti-aliasing filter in an embodiment of the present invention.
[0032] Figure 4 This is a diagram showing how different downsampling algorithms preserve the temporal features of the speech signal in the anti-aliasing filter of this invention.
[0033] Figure 5 This describes the attack process on the intelligent speech recognition system in an embodiment of the present invention. Detailed Implementation
[0034] To enable those skilled in the art to better understand the present invention, the technical solutions of the present invention will be clearly and completely described below with reference to the accompanying drawings of the embodiments of the present invention. Obviously, the described embodiments are only some embodiments of the present invention, and not all embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those skilled in the art without creative effort should fall within the scope of protection of the present invention.
[0035] It should be noted that the terms "first," "second," etc., in the specification, claims, and accompanying drawings of this invention are used to distinguish similar objects and are not necessarily used to describe a specific order or sequence. It should be understood that such data can be interchanged where appropriate so that the embodiments of the invention described herein can be implemented in orders other than those illustrated or described herein. Furthermore, the terms "comprising" and "having," and any variations thereof, are intended to cover a non-exclusive inclusion; for example, a process, method, system, product, or apparatus that comprises a series of steps or units is not necessarily limited to those steps or units explicitly listed, but may include other steps or units not explicitly listed or inherent to such processes, methods, products, or apparatus.
[0036] like Figure 2 As shown, the present invention provides a spoofing attack method for intelligent voice systems, specifically including the following steps:
[0037] S1, adjusts the original signal s to a signal s1 with a sampling rate of r1;
[0038] S2, calculate the stopband of the resampling algorithm based on the target system sampling rate r2;
[0039] Given that the required sampling rate for an intelligent voice system is r², according to the Nyquist sampling theorem, the original frequency of the signal must be less than half the sampling frequency to avoid aliasing. Therefore, the stopband of the low-pass filter must be greater than r² / 2. For example, for a system requiring an input audio sampling rate of 16kHz, embedding noise signals into the 8kHz frequency band will allow the low-pass filter to effectively remove the noise. In auditory psychology, the closer the masking frequency band is to the masked frequency band, the better the masking effect. Considering the influence of the transition band in the filter, the noise spectrum is set to f. a =r² / 2 + 0.5kHz to f b = r2 / 2 + 1KHz.
[0040] S3. Based on the obtained stopband, construct the noise signal S[n] with the spectrum in the stopband.
[0041] This application generates two different types of sound to disguise the original signal: ① broadband noise. First, white noise S[n]~N(0,1) with the same length as the original speech signal is generated, and the white noise is normalized to the interval S[n]=S[n] / max(|S|). Then, a passband of [f a ,f b The filter of ] is used to filter S[n] to obtain broadband noise of the specified frequency band; ② Pure tone, with known sampling frequency f s Generate a cosine signal with a specified frequency f using the following formula:
[0042] S[n]=Acos(2π·f·(n / f s )) (1)
[0043] f / f s Known as digital frequency, it is the original signal frequency f divided by the sampling frequency f. s The normalization of the sequence S[n] is such that the length of the sequence is consistent with the length of the masked signal s.
[0044] S4, after reducing the energy of the signal with sampling rate r1, adds it to the generated noise signal, s'=s1*k+S, thus generating the disguised signal.
[0045] This invention is based on the linear property of the Fourier transform, as shown below:
[0046] F[αf1(t)+βf2(t)]=αF[f1(t)]+βF[f2(t)] (2)
[0047] The original speech signal s has a signal reduction factor k = 200, a camouflage signal sampling rate r1, and a target system sampling rate r2, where r1 > r2; the output is the camouflaged speech signal s'.
[0048] F[·] represents the Fourier transform. The above formula shows that the spectrum of two signals f1(t) and f2(t) directly added in the time domain is equal to the spectrum of the two signals directly added. Therefore, it is not necessary to perform a Fourier transform on the original signal. The noise signal can be added directly by adding the signal with a sampling rate of r1 to the noise signal to add noise to the specified frequency band of the original signal. Reducing the energy of the signal with a sampling rate of r1 can lower the hearing threshold of the speech signal and make it easier to mask. However, due to the limitations of file encoding precision, the signal with a sampling rate of r1 cannot be reduced indefinitely. Audio files generally use 16-bit signed integers for encoding, with an encoding range of -2. 15 to 2 15 -1. Common speech tasks, such as speech recognition and speaker recognition, can be accomplished using an 8-bit signed integer, with an encoding range of -2. 7 to 2 7 -1. Reduce the intensity of the 16-bit voice signal by 2. 8 =256, and the accuracy of the obtained speech signal is sufficient for the recognition task. Therefore, the attack algorithm proposed in this invention selects to reduce the signal with a sampling rate of r1 to 200 times its original value.
[0049] This invention proposes a spoofing attack method for intelligent voice systems. This method only requires knowledge of the target system's sampling rate to generate spoofing attack samples, demonstrating that it is a general black-box attack algorithm, not limited to deep learning models. It can complete attacks without knowing any information about the model, greatly expanding the application scope of the attack algorithm. Based on the auditory masking theory of speech, a novel speech spoofing attack strategy is proposed. This strategy can disguise ordinary speech signals as electrical noise to achieve the purpose of spoofing attacks. The effectiveness and universality of the proposed spoofing attack algorithm for various sampling algorithms are verified; that is, attacks can be carried out simply by knowing the input sampling rate of the target algorithm.
[0050] In one embodiment of the present invention, such as Figure 2 As shown, the frequency domain illustrates the generation process of the spoofed samples. Figure 2 (a) represents the spectrum of the noise signal. Figure 2 (c) represents the spectrum of the original speech signal after energy reduction. Figure 2 (b) represents the spoofed sample generated by combining the original signal with the noise signal. Figure 2 (d) represents the original speech signal recovered after passing through the anti-aliasing filter.
[0051] Specifically, the following steps are included:
[0052] Given the original speech signal s, a signal reduction factor k = 200, a spoofing signal sampling rate r1, and a target system sampling rate r2, where r1 > r2. The algorithm outputs the spoofed speech signal s'.
[0053] S1 adjusts the original signal s to a signal s1 with a sampling rate of r1.
[0054] S2, calculate the stopband of the resampling algorithm based on the target system sampling rate r2; when the required sampling rate of the intelligent voice system is r2, the noise frequency band is defined as f. a =r² / 2 + 0.5kHz to f b = r2 / 2 + 1KHz.
[0055] S3, construct the noise signal S[n] in the stopband; for broadband noise: first generate white noise S[n] ~ N(0,1) with the same length as the original speech signal, and normalize the white noise to the interval S[n] = S[n] / max(|S|), then use the passband [f a ,f b By filtering S[n] with the filter of ], broadband noise of the specified frequency band is obtained; for pure tone signals: given the sampling frequency f s Generate a cosine signal with a specified frequency f using the following formula:
[0056] S[n]=Acos(2π·f·(n / f s ))
[0057] f / f s Known as digital frequency, it is the original signal frequency f divided by the sampling frequency f. s The normalization of the sequence S[n] is such that the length of the sequence is consistent with the length of the masked signal s.
[0058] S4, s' = s1*k + S, adds the generated noise signal to the original signal after energy reduction to generate the disguised signal.
[0059] The spectrograms of the spoofed speech signal obtained using the method of this invention before and after passing through the anti-aliasing filter are shown below. Figure 3 As can be seen, after adding pure tone and noise signals to the original speech signal, the original speech features are hidden. However, after passing through the anti-aliasing filter, the speech features are recovered, and the difference from the original speech features is small. This demonstrates the feature preservation property of the spoofing attack algorithm, which is why this attack algorithm is applicable to various intelligent speech systems.
[0060] like Figure 4 As shown, the time-domain features of the speech signal are preserved using different downsampling algorithms. Figure 4 As can be seen, apart from the polyphase and linear algorithms, the camouflage attack algorithm exhibits good temporal feature preservation.
[0061] like Figure 5The original signal energy containing the semantic "I like tea" is reduced by 2. 8 The attack sample is obtained by adding the original signal to a fixed-frequency noise. Due to the masking effect of the noise and the significant reduction in the energy of the original signal, the human ear can no longer hear the original semantic "I like tea." Instead, it hears irregular electrical noise or pure tone signals, depending on the nature of the masking signal. However, the intelligent speech system can still recognize "I like tea," which creates a cognitive difference between the human ear and the intelligent speech system, thus achieving the attack objective. This invention accurately reconstructs the statistical characteristics of the speech signal through a sampling algorithm, achieving a high recognition accuracy even with extremely low signal-to-noise ratio noise input, verifying the effectiveness of the proposed algorithm.
Claims
1. A method for spoofing attacks on intelligent voice systems, characterized in that, Includes the following steps: S1, adjust the original signal to a sampling rate of The signal; S2, Calculate the stopband of the resampling algorithm based on the target system sampling rate; S3, Based on the obtained stopband, construct the noise signal with the spectrum in the stopband; S4, with a sampling rate of The signal is reduced in energy and then added to the generated noise signal to generate the disguised signal.
2. The spoofing attack method for intelligent voice systems according to claim 1, characterized in that, The obtained stopband is greater than half of the target system's sampling rate.
3. The spoofing attack method for intelligent voice systems according to claim 1, characterized in that, Set the noise spectrum to arrive .
4. The spoofing attack method for intelligent voice systems according to claim 1, characterized in that, Broadband noise and pure tone were used to disguise the original signal to obtain the noise signal in the stopband.
5. A spoofing attack method for an intelligent voice system according to claim 4, characterized in that, Generate white noise with the same length as the original speech signal. And normalize the white noise. Within the interval, the passband is then used. filter pair By filtering, broadband noise in the specified frequency band can be obtained.
6. A spoofing attack method for an intelligent voice system according to claim 4, characterized in that, Based on sampling frequency Use the following formula to generate a specified frequency. Cosine signal: (1) Known as digital frequency, it is the original signal frequency. For sampling frequency Normalization, sequence Length and the masked signal Consistent.
7. The spoofing attack method for intelligent voice systems according to claim 1, characterized in that, The sampling rate is The original signal was reduced by 200 times.
8. A spoofing attack system for intelligent voice systems, characterized in that, It includes a resampling module, a stopband calculation module, a noise module, and an attack module; The resampling module is used to adjust the original signal to a sampling rate of [missing information]. The signal; The stopband calculation module is used to calculate the stopband of the resampling algorithm based on the target system sampling rate. The noise module constructs a noise signal with a spectrum in the stopband based on the acquired stopband. The attack module will set the sampling rate to... The signal is reduced in energy and then added to the generated noise signal to generate the disguised signal.
9. A spoofing attack system for an intelligent voice system according to claim 8, characterized in that, The obtained stopband is greater than half of the target system's sampling rate.
10. A spoofing attack system for an intelligent voice system according to claim 8, characterized in that, Broadband noise and pure tone were used to disguise the original signal to obtain the noise signal in the stopband.
Citation Information
Patent Citations
Digital speech resampling detection method based on bandwidth inconsistency of frequency bands
CN108665905A
Method and system for generating synthetic multi-conditioned data sets for robust automatic speech recognition
US20210065681A1