Voice confrontation defense method and system for speaker recognition system
Through feature filtering based on F-ratio feature distribution and spectrum recovery of U-Net model, combined with phase estimation of the Griffin-Lim algorithm, the vulnerability problem of the speaker recognition system in the face of adversarial attacks is solved, and the efficient defense effect of plug and play is achieved.
Patent Information
- Application Number
- CN202510022873.0
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-01-07
- Publication Date
- 2025-05-06
- Estimated Expiration
- 2045-01-07
AI Technical Summary
Existing speaker recognition systems are vulnerable to confronting adversarial attacks and are difficult to defend effectively, and existing defense methods often require retraining or fine-tuning, which is costly and unsuitable for plug-and-play deployments.
A fine feature filtering method based on the F-ratio speaker feature distribution is adopted. The non-rootive feature filtering module is used to remove non-rootive features, compress the living space against noise, and reconstruct the deleted features through the U-Net model. Phase estimation is performed in combination with the Griffin-Lim algorithm to obtain the purified voice waveform.
It realizes effective resistance to various attacks without fine-tuning or retraining, improves the security and robustness of the speaker recognition system, and has plug-and-play characteristics, and is suitable for speaker recognition systems of different architectures.
Smart Images

Figure CN119943057A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of speech recognition technology, and in particular to a speech confrontation defense method and system for a speaker recognition system. Background Art
[0002] Voiceprint recognition technology has become a widely used biometric identification method due to its seamless integration with Voice User Interfaces (VUI). Advances in deep learning have enabled Speech Recognition Systems (SRS) to be widely used in hardware devices (such as Google Home, Amazon Alexa) and software applications (such as HSBC, WeChat). Although the current state-of-the-art SRS have made significant progress in performance, recent studies have revealed that they are extremely sensitive to small perturbations from adversarial attacks, which may lead to erroneous recognition results and pose serious security risks. This issue has led to increasing concerns about the security of speech recognition systems.
[0003] Given the serious threat posed by adversarial attacks, developing effective countermeasures to protect speaker recognition systems has become an urgent task. Existing adversarial defense methods can be divided into three categories: adversarial training, random smoothing, and input reconstruction methods. Since the first two methods are difficult to directly apply to mature commercial SRSs due to the need for retraining and huge computational requirements, we mainly focus on methods based on input reconstruction. Among the input reconstruction methods, one technical route is to defend against adversarial samples through signal / feature processing methods, such as quantization, resampling, low-pass filters, and FeCo. However, these methods usually sacrifice the recognition performance of normal samples and require the SRS to be fine-tuned or retrained to adapt to the corresponding defense strategies. Another method is to purify adversarial noise by training auxiliary networks such as Variational Autoencoder (VAE), which have the characteristics of plug-and-play. However, this method is often difficult to effectively defend against adaptive attacks.
[0004] Considering the high cost of further fine-tuning or retraining speaker recognition systems in actual deployment scenarios, we propose a novel plug-and-play speech adversarial defense paradigm as a line of defense for any deployed SRS to mitigate the vulnerability of SRS in the face of attacks. To implement a plug-and-play defense system, the challenges are as follows: (1) how to separate adversarial perturbations from speech samples to ensure the defense performance of adversarial samples; (2) how to restore the pre-processed and corrupted features so that SRS can correctly recognize them without fine-tuning or retraining. Therefore, to implement a plug-and-play adversarial defense system, balancing the adversarial defense performance and the restored speech recognition accuracy is a key issue. Summary of the invention
[0005] The present invention provides a voice confrontation defense method and system for a speaker recognition system, which are used to solve the defects existing in the prior art when the speaker recognition system is deployed.
[0006] In a first aspect, the present invention provides a method for defending against speech adversarial attacks on a speaker recognition system, comprising: Acquire input speech data, perform feature extraction on the input speech data, and obtain an amplitude spectrogram of the input speech data; The amplitude spectrum is processed by a preset fine-grained non-robust feature filtering module to obtain a masked amplitude spectrum; Reconstructing the masked amplitude spectrum through a U-Net model to obtain a reconstructed amplitude spectrum graph; Based on the Griffin-Lim algorithm, phase estimation is performed from the reconstructed amplitude spectrogram to obtain a reconstructed speech signal.
[0007] According to a speech adversarial defense method for a speaker recognition system provided by the present invention, input speech data is obtained, feature extraction is performed on the input speech data, and an amplitude spectrogram of the input speech data is obtained, including: The input voice data is framed, and each frame signal after the framing is subjected to a fast Fourier transform to obtain a frequency domain signal of each frame; The absolute value represented by the frequency domain signal of each frame is calculated to obtain the amplitude spectrum.
[0008] According to a speech adversarial defense method for a speaker recognition system provided by the present invention, the amplitude spectrogram is processed by presetting a fine-grained non-robust feature filtering module to obtain a masked amplitude spectrum, including: Obtain a speech data set formed by the amplitude spectrograms corresponding to multiple speakers, randomly select a number of people M from the speech data set, select a vector N for each person, obtain an audio sample from the librosa python library, and truncate and lengthen the audio sample to a fixed size sampling point; The F-ratio statistics are used to calculate the speaker weights of different frequency bands in the truncated and padded audio samples to obtain high speaker-related frequency bands and low speaker-related frequency bands. Randomly generate uniform noise with the same length as the input audio and a preset range to simulate adversarial noise, add the simulated adversarial noise to the speech data set to obtain a simulated adversarial sample, and obtain a high speaker-related frequency band masking threshold and a low speaker-related frequency band masking threshold according to the simulated adversarial sample; The number of masks required for each frequency band is calculated according to the F-ratio value, and the position of random masking in each frequency band is determined.
[0009] According to a speech adversarial defense method for a speaker recognition system provided by the present invention, the speaker weights of different frequency bands in the truncated and lengthened audio sample are calculated using F-ratio statistics to obtain high speaker-related frequency bands and low speaker-related frequency bands, including: According to any amplitude spectrum vector of any speaker and the vector N, an average feature vector of any speaker is calculated; Calculate the average feature vector of all speakers according to the average feature vector of any speaker and the number of people M; Calculating an F-ratio value based on the average feature vector of any speaker, the average feature vector of all speakers, the vector N, and the number of people M; The average F-ratio value corresponding to any frequency band is determined, and combined with the total number of frequency bands to obtain a division threshold, wherein the division threshold is used to divide the high speaker correlation frequency band set and the low speaker correlation frequency band set.
[0010] According to a speech adversarial defense method for a speaker recognition system provided by the present invention, a high speaker-related frequency band masking threshold and a low speaker-related frequency band masking threshold are obtained according to the simulated adversarial sample, including: Respectively calculating the simulated adversarial sample amplitude value of the simulated adversarial sample in any frequency band, and the original sample amplitude value of the corresponding original sample in any frequency band; Taking the absolute value of the difference between the amplitude value of the simulated adversarial sample and the amplitude value of the original sample and then taking the maximum value to obtain the high speaker correlation frequency band masking threshold; The pitch amplitude value of the simulated adversarial sample amplitude value is calculated and then averaged to obtain the low speaker-related frequency band masking threshold.
[0011] According to a speech adversarial defense method for a speaker recognition system provided by the present invention, the number of masks required for each frequency band is calculated according to the F-ratio value, and the position of random masking in each frequency band is determined, including: Get the amplitude value corresponding to any frequency band after Fratio calculation, the Fratio value corresponding to any frequency band and the masking threshold corresponding to the frequency band; For all the frequency bands smaller than the masking threshold corresponding to the frequency band, the corresponding amplitude values after Fratio calculation are summed and then multiplied by the Fratio value corresponding to the frequency band to obtain the number of each frequency band that needs to be masked; The original sample amplitude values that are smaller than the masking threshold corresponding to the frequency band and the number of each frequency band that needs to be masked are randomly selected to obtain the random masking position in each frequency band.
[0012] According to a speech adversarial defense method for a speaker recognition system provided by the present invention, the masked amplitude spectrum is reconstructed by a U-Net model to obtain a reconstructed amplitude spectrum graph, including: Determine to use a feature recovery module of the U-Net model to reconstruct the masked amplitude spectrum, wherein the feature recovery module includes a four-layer encoder and a four-layer decoder; The masked amplitude spectrum is input into the encoder, a one-dimensional convolution is performed only along the frequency axis, and a frequency transform is connected before each encoding layer. Each encoder layer includes two compressed residual branches, a Snake activation function is used instead of a ReLU activation function, and a long short-term memory network and a time-based attention module are added to the internal layer of the encoder. The bottleneck feature is converted back to a complete amplitude spectrum of the same size as the encoder by the decoder; During training, the amplitude spectrum loss and multi-resolution short-time fast Fourier transform loss based on the F-ratio speaker feature distribution are calculated separately, and the total loss function is obtained by weighted summation.
[0013] According to a speech adversarial defense method for a speaker recognition system provided by the present invention, the amplitude spectrum loss based on the F-ratio speaker feature distribution includes: Obtain the original speech amplitude spectrum and the restored speech amplitude spectrum of any frequency band in the amplitude spectrum, as well as the F-ratio value of any frequency band in the amplitude spectrum; Based on the original speech amplitude spectrum, the restored speech amplitude spectrum and the F-ratio value of any frequency band in the amplitude spectrum, the amplitude spectrum loss based on the F-ratio speaker characteristic distribution is calculated; Correspondingly, the multi-resolution short-time fast Fourier transform loss includes; Obtaining a multi-resolution short-time Fourier transform value of the original speech audio and a multi-resolution short-time Fourier transform value of the restored speech audio; Based on the original speech audio multi-resolution short-time Fourier transform value and the restored speech audio multi-resolution short-time Fourier transform value, obtaining any resolution short-time fast Fourier transform loss; Different resolutions are formed by combining different fast Fourier transform frequencies, jump lengths and window sizes, and the short-time fast Fourier transform losses of different resolutions are summed and averaged to obtain the multi-resolution short-time fast Fourier transform loss.
[0014] In a second aspect, the present invention further provides a voice adversarial defense system for a speaker recognition system, comprising: An extraction module, used to obtain input speech data, perform feature extraction on the input speech data, and obtain an amplitude spectrogram of the input speech data; A masking module, used for processing the amplitude spectrum through a preset fine-grained non-robust feature filtering module to obtain a masked amplitude spectrum; A reconstruction module, used for reconstructing the masked amplitude spectrum through a U-Net model to obtain a reconstructed amplitude spectrum graph; The estimation module is used to perform phase estimation from the reconstructed amplitude spectrogram based on the Griffin-Lim algorithm to obtain a reconstructed speech signal.
[0015] In a third aspect, the present invention further provides an electronic device, comprising a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein when the processor executes the program, a speech adversarial defense method for a speaker recognition system as described in any one of the above is implemented.
[0016] The speech adversarial defense method and system for speaker recognition system provided by the present invention, by applying the guiding idea of "subtracting first and then adding" at the feature level, the main purpose of subtraction is to remove non-robust features with low correlation with speaker features, while retaining backbone features to compress the living space of adversarial noise. The present invention proposes a fine feature filtering method based on F-ratio speaker feature distribution to remove non-robust features. As for the addition part, the present invention proposes a spectrum recovery module based on U-Net and an amplitude loss based on -ratio speaker feature distribution, the purpose of which is to reconstruct the robust features of real speech based on the retained speech backbone features, so as to obtain the purified speech waveform from the adversarial sample, so that it can be correctly recognized by the speaker recognition system without fine-tuning or retraining. Compared with the existing method technology, the advantage of the present invention is that it can be used as a plug-and-play pre-defense means, which can not only effectively resist various attackers, but also can be applied to speaker recognition systems of different architectures without the need for retraining or fine-tuning of the recognition system. Compared with other existing methods, the method proposed by the present invention not only performs outstandingly in defending against different attack methods, but also has achieved significant improvements in plug-and-play and versatility. BRIEF DESCRIPTION OF THE DRAWINGS
[0017] In order to more clearly illustrate the technical solutions in the present invention or the prior art, the drawings required for use in the embodiments or the description of the prior art will be briefly introduced below. Obviously, the drawings described below are some embodiments of the present invention. For ordinary technicians in this field, other drawings can be obtained based on these drawings without paying creative work.
[0018] Figure 1 It is a flow chart of a voice adversarial defense method for a speaker recognition system provided by the present invention; Figure 2It is a principle block diagram of the voice adversarial defense method for a speaker recognition system provided by the present invention; Figure 3 It is a structural schematic diagram of a speech adversarial defense system for a speaker recognition system provided by the present invention; Figure 4 It is a structural schematic diagram of the electronic device provided by the present invention. DETAILED DESCRIPTION
[0019] In order to make the purpose, technical solution and advantages of the present invention clearer, the technical solution of the present invention will be clearly and completely described below in conjunction with the drawings of the present invention. Obviously, the described embodiments are part of the embodiments of the present invention, not all of the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by ordinary technicians in this field without creative work are within the scope of protection of the present invention.
[0020] With the increasing importance of voiceprint technology in biometric authentication, the security vulnerabilities of speaker recognition systems (SRS) have attracted widespread attention due to the threat of adversarial samples. In order to address the limitations of existing methods in meeting plug-and-play requirements and defending against adaptive attacks, this paper proposes a novel adversarial purification framework, SA-Net, whose key idea is to adopt a "subtract first and then add" strategy at the feature level. In order to suppress the existence of adversarial perturbations, in the "addition" part, this paper proposes a fine feature filtering method based on F-ratio speaker feature distribution to remove non-robust features. In addition, in order to ensure the plug-and-play capability of the method of the present invention, in the "subtraction" part, this paper proposes a spectrum recovery module based on U-Net and an amplitude loss based on F-ratio speaker feature distribution to reconstruct the deleted features, thereby achieving accurate speaker recognition without fine-tuning or retraining.
[0021] Figure 1 FIG. 1 is a flow chart of a method for defending against speech adversarial attacks on a speaker recognition system provided by an embodiment of the present invention. Figure 1 As shown, including: Step 100: Acquire input speech data, perform feature extraction on the input speech data, and obtain an amplitude spectrogram of the input speech data; Step 200: Processing the amplitude spectrum through a preset fine-grained non-robust feature filtering module to obtain a masked amplitude spectrum; Step 300: reconstructing the masked amplitude spectrum through a U-Net model to obtain a reconstructed amplitude spectrum graph; Step 400: Perform phase estimation from the reconstructed amplitude spectrogram based on the Griffin-Lim algorithm to obtain a reconstructed speech signal.
[0022] The technical solution adopted by the present invention is a plug-and-play adversarial defense method for speaker recognition systems based on filtering and reconstruction, such as Figure 2 As shown, the following steps are included: Step 1: Extract features of the input speech to obtain the amplitude spectrogram of the audio data.
[0023] Step 2: The amplitude spectrum is processed through a fine-grained non-robust feature filtering module to remove non-robust features and compress the living space against disturbances.
[0024] Step 3: The masked amplitude spectrum is passed through the feature recovery module to reconstruct the complete amplitude spectrum.
[0025] Step 4: Estimate the phase from the reconstructed amplitude spectrogram using the Griffin-Lim algorithm to finally obtain the reconstructed speech signal.
[0026] In one embodiment, in step 1, a short-time Fourier transform (STFT) is performed on the input speech signal. The speech signal is first divided into frames, and then a fast Fourier transform (FFT) is applied to each frame signal to convert it from the time domain to the frequency domain. Next, the absolute value of the frequency domain representation of each frame is calculated to obtain an amplitude spectrum. The amplitude spectrum represents the amplitude information of each frequency component and reflects the strength of the speech signal at different frequencies.
[0027] In one embodiment, step 2 includes: Obtain a speech data set formed by the amplitude spectrograms corresponding to multiple speakers, randomly select a number of people M from the speech data set, select a vector N for each person, obtain an audio sample from the librosa python library, and truncate and lengthen the audio sample to a fixed size sampling point; The F-ratio statistics are used to calculate the speaker weights of different frequency bands in the truncated and padded audio samples to obtain high speaker-related frequency bands and low speaker-related frequency bands. Randomly generate uniform noise with the same length as the input audio and a preset range to simulate adversarial noise, add the simulated adversarial noise to the speech data set to obtain a simulated adversarial sample, and obtain a high speaker-related frequency band masking threshold and a low speaker-related frequency band masking threshold according to the simulated adversarial sample; The number of masks required for each frequency band is calculated according to the F-ratio value, and the position of random masking in each frequency band is determined.
[0028] Specifically, the embodiment of the present invention randomly selects M people from a multi-speaker speech data set, and each person selects N sentences. The librosa python library reads the audio samples and truncates and lengthens them to fixed-size sampling points to ensure that all training samples have the same length.
[0029] Use the F-ratio statistical method to calculate the speaker weights of different frequency bands in the audio data and obtain high speaker-related frequency bands and low speaker related frequency band area .
[0030] Define Fratio as follows:
[0031]
[0032] in, denote the average feature vector of speaker i and all speakers respectively. Denotes the jth magnitude spectrum vector of speaker i. Based on the calculation of F-ratio, the speaker weights of different frequency bands in the magnitude spectrum can be determined.
[0033] Calculate the split thresholds for high and low speaker-dependent frequency bands:
[0034] in, represents the average Fration value corresponding to the bth frequency band, and B represents the total number of frequency bands. and partition threshold The size of is used to obtain the high speaker-related frequency band and the low speaker-related frequency band set respectively.
[0035] Randomly generate a string with the same length as the input audio and in the range The uniform noise is used to simulate the adversarial noise, and the adversarial noise is added to the original speech sample to obtain the simulated adversarial sample. According to the simulated adversarial sample, the masking thresholds of the high speaker-related frequency band and the low speaker-related frequency band are obtained, which are expressed as and .
[0036]
[0037]
[0038] in, and They represent the amplitude values of the simulated adversarial sample x' and the original sample x in the bth frequency band respectively. pitch() represents the amplitude value of the calculated fundamental pitch.
[0039] The number num of masks required for each frequency band is calculated according to the F-ratio value, and the position of random masking in each frequency band is determined, so that the subsequent recovery step can better recover the adjacent frequency components.
[0040]
[0041]
[0042]
[0043] in, Indicates the Fratio value corresponding to the b-th frequency band, It indicates the amplitude value corresponding to the b-th frequency band after Fratio calculation. represents the masking threshold corresponding to the frequency band, represents a random selection function.
[0044] In one embodiment, step 3 includes: Determine to use a feature recovery module of the U-Net model to reconstruct the masked amplitude spectrum, wherein the feature recovery module includes a four-layer encoder and a four-layer decoder; The masked amplitude spectrum is input into the encoder, a one-dimensional convolution is performed only along the frequency axis, and a frequency transform is connected before each encoding layer. Each encoder layer includes two compressed residual branches, a Snake activation function is used instead of a ReLU activation function, and a long short-term memory network and a time-based attention module are added to the internal layer of the encoder. The bottleneck feature is converted back to a complete amplitude spectrum of the same size as the encoder by the decoder; During training, the amplitude spectrum loss and multi-resolution short-time fast Fourier transform loss based on the F-ratio speaker feature distribution are calculated separately, and the total loss function is obtained by weighted summation.
[0045] Specifically, the masked amplitude spectrogram is reconstructed into a complete amplitude spectrum through a feature recovery module based on the U-Net model. This module includes an encoder and a decoder part, each with four layers. The encoder takes the masked amplitude spectrum as input and uses a 1D convolution that operates only along the frequency axis. In order to improve the performance of the model, we add a frequency transformation layer (FTL) before each encoder layer, which is designed to capture the global correlation on the frequency axis, expand the receptive field of the model, and enhance the reconstruction ability of the amplitude spectrum. Each layer contains two compressed residual branches and uses the Snake activation function instead of the ReLU activation function. In addition, we also add a long short-term memory network (LSTM) and a time-based attention module to the internal layers of the encoder. After the encoder, the decoder module converts the bottleneck features back to a complete amplitude spectrum of the same size as the input encoder.
[0046] During the training process, the amplitude spectrum loss based on the F-ratio speaker feature distribution is calculated separately and multi-resolution STFT loss , and weighted to get the total loss function .
[0047]
[0048] Among them, the amplitude spectrum loss based on the F-ratio speaker feature distribution for:
[0049] in, Indicates the F-ratio value of the i-th frequency band of the magnitude spectrum x. Represents the restored amplitude spectrum. and Represent the amplitude spectra of the original and restored speech respectively.
[0050] Multi-resolution STFT loss for:
[0051]
[0052] Among them, STFT(x) and STFT( ) represent the multi-resolution short-time Fourier transform (STFT) of the original audio and the reconstructed audio. The FFT frequency bins are 512, 1024, and 2048, the jump lengths are 50, 120, and 240, and the window sizes are 240, 600, and 1200. The multi-resolution STFT loss is calculated by summing and averaging the losses at different resolutions. .
[0053] In one embodiment, step 4 includes: The Griffin-Lim algorithm is used to estimate the phase information of the waveform instead of the original phase features because the phase information may also contain adversarial perturbations, which may have a negative impact on the final decision of the target model. The reconstructed amplitude spectrum and the estimated phase spectrum are then synthesized into the final purified speech waveform.
[0054] Furthermore, the present invention is further illustrated by experiments, which use the LibriSpeech dataset, specifically the "train-clean-100" subset. This subset contains 100 hours of English speech from 251 unique speakers. For each speaker, the present invention randomly selects 90% of the speech as training data, and the remaining 10% is used for inference testing. The target speaker recognition system (SRS) is the widely adopted ECAPA-TDNN speaker recognition model, which is pre-trained on VoxCeleb to extract speaker embedding. The similarity calculation uses cosine similarity as the similarity function. In order to comprehensively evaluate the defense effect, 4 white-box attacks, including FGSM, PGD, CW∞ and CW2, and 3 black-box attacks, including FAKEBOB (FB), SirenAttack (SA) and Kenansville (KS), are implemented. The defense capabilities of different defense methods against different attack methods are shown in Table 1.
[0055] Table 1 Comparison of experimental results
[0056] Table 1 shows the average defense performance results of different defense methods under non-adaptive attacks. Input reconstruction methods such as quantization, noise addition, smoothing, downsampling, low-pass filtering, and AAC compression perform poorly in classification accuracy for both benign and adversarial samples. These methods degrade the quality of the original audio, and the speaker recognition system is not fine-tuned or retrained for these defense methods. The present invention achieves 99.1% accuracy on benign samples and performs well under white-box attacks, with an accuracy of 93.6% for FGSM attacks, 92.3% for PGD attacks, 90.3% for CWinf attacks, and 99.1% for CW2 attacks. In terms of black-box attacks, the present method outperforms other methods, achieving 98.5% accuracy for FAKEBOB attacks, 92.5% for SirenAttack attacks, and 88.3% for Kenansville attacks. These results reveal the effectiveness and universality of the present method, demonstrating its ability to resist various types of attacks, as well as its plug-and-play nature, without any adjustment to the SRS.
[0057] The speech adversary defense system for a speaker recognition system provided by the present invention is described below. The speech adversary defense system for a speaker recognition system described below and the speech adversary defense method for a speaker recognition system described above can be referred to each other.
[0058] Figure 3 FIG. 1 is a schematic diagram of the structure of a voice confrontation defense system for a speaker recognition system provided by an embodiment of the present invention. Figure 3 As shown, it includes: an extraction module 31, a masking module 32, a reconstruction module 33 and an estimation module 34, wherein: The extraction module 31 is used to obtain input speech data, perform feature extraction on the input speech data, and obtain an amplitude spectrogram of the input speech data; the masking module 32 is used to process the amplitude spectrogram through a preset fine-grained non-robust feature filtering module to obtain a masked amplitude spectrum; the reconstruction module 33 is used to reconstruct the masked amplitude spectrum through a U-Net model to obtain a reconstructed amplitude spectrogram; the estimation module 34 is used to perform phase estimation from the reconstructed amplitude spectrogram based on the Griffin-Lim algorithm to obtain a reconstructed speech signal.
[0059] Figure 4 An example of a physical structure diagram of an electronic device is shown in FIG. Figure 4As shown, the electronic device may include: a processor 410, a communication interface 420, a memory 430 and a communication bus 440, wherein the processor 410, the communication interface 420 and the memory 430 communicate with each other through the communication bus 440. The processor 410 may call the logic instructions in the memory 430 to execute a voice adversarial defense method for a speaker recognition system, the method comprising: obtaining input voice data, extracting features from the input voice data, and obtaining an amplitude spectrogram of the input voice data; processing the amplitude spectrogram through a preset fine-grained non-robust feature filtering module to obtain a masked amplitude spectrum; reconstructing the masked amplitude spectrum through a U-Net model to obtain a reconstructed amplitude spectrogram; and performing phase estimation from the reconstructed amplitude spectrogram based on the Griffin-Lim algorithm to obtain a reconstructed voice signal.
[0060] In addition, the logic instructions in the above-mentioned memory 430 can be implemented in the form of a software functional unit and can be stored in a computer-readable storage medium when it is sold or used as an independent product. Based on this understanding, the technical solution of the present invention, in essence, or the part that contributes to the prior art or the part of the technical solution, can be embodied in the form of a software product, and the computer software product is stored in a storage medium, including a number of instructions for a computer device (which can be a personal computer, a server, or a network device, etc.) to perform all or part of the steps of the method described in each embodiment of the present invention. The aforementioned storage medium includes: U disk, mobile hard disk, read-only memory (ROM, Read-Only Memory), random access memory (RAM, Random Access Memory), disk or optical disk, etc. Various media that can store program codes.
[0061] The device embodiments described above are merely illustrative, wherein the units described as separate components may or may not be physically separated, and the components displayed as units may or may not be physical units, that is, they may be located in one place, or they may be distributed on multiple network units. Some or all of the modules may be selected according to actual needs to achieve the purpose of the scheme of this embodiment. Ordinary technicians in this field can understand and implement it without paying creative labor.
[0062] Through the description of the above implementation methods, those skilled in the art can clearly understand that each implementation method can be implemented by means of software plus a necessary general hardware platform, and of course, can also be implemented by hardware. Based on this understanding, the above technical solution is essentially or the part that contributes to the prior art can be embodied in the form of a software product, and the computer software product can be stored in a computer-readable storage medium, such as ROM / RAM, a disk, an optical disk, etc., including a number of instructions for a computer device (which can be a personal computer, a server, or a network device, etc.) to execute the methods described in each embodiment or some parts of the embodiments.
[0063] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention, rather than to limit it. Although the present invention has been described in detail with reference to the aforementioned embodiments, those skilled in the art should understand that they can still modify the technical solutions described in the aforementioned embodiments, or make equivalent replacements for some of the technical features therein. However, these modifications or replacements do not deviate the essence of the corresponding technical solutions from the spirit and scope of the technical solutions of the embodiments of the present invention.
Claims
1. A method for defending against speech adversarial attacks on a speaker recognition system, characterized in that: include: Acquire input speech data, perform feature extraction on the input speech data, and obtain an amplitude spectrogram of the input speech data; The amplitude spectrum is processed by a preset fine-grained non-robust feature filtering module to obtain a masked amplitude spectrum; Reconstructing the masked amplitude spectrum through a U-Net model to obtain a reconstructed amplitude spectrum graph; Based on the Griffin-Lim algorithm, phase estimation is performed from the reconstructed amplitude spectrogram to obtain a reconstructed speech signal.
2. The method for voice adversarial defense against a speaker recognition system according to claim 1, characterized in that: Acquiring input speech data, performing feature extraction on the input speech data, and obtaining an amplitude spectrogram of the input speech data, including: The input voice data is framed, and each frame signal after the framing is subjected to a fast Fourier transform to obtain a frequency domain signal of each frame; The absolute value represented by the frequency domain signal of each frame is calculated to obtain the amplitude spectrum.
3. The method for defending against speech adversarial attacks on a speaker recognition system according to claim 1, characterized in that: The amplitude spectrum is processed by a preset fine-grained non-robust feature filtering module to obtain a masked amplitude spectrum, including: Obtain a speech data set formed by the amplitude spectrograms corresponding to multiple speakers, randomly select a number of people M from the speech data set, select a vector N for each person, obtain an audio sample from the librosa python library, and truncate and lengthen the audio sample to a fixed size sampling point; The F-ratio statistics are used to calculate the speaker weights of different frequency bands in the truncated and padded audio samples to obtain high speaker-related frequency bands and low speaker-related frequency bands. Randomly generate uniform noise with the same length as the input audio and a preset range to simulate adversarial noise, add the simulated adversarial noise to the speech data set to obtain a simulated adversarial sample, and obtain a high speaker-related frequency band masking threshold and a low speaker-related frequency band masking threshold according to the simulated adversarial sample; The number of masks required for each frequency band is calculated according to the F-ratio value, and the position of random masking in each frequency band is determined.
4. The method for defending against speech adversarial attacks on a speaker recognition system according to claim 3, characterized in that: The F-ratio statistics are used to calculate the speaker weights of different frequency bands in the truncated and padded audio samples to obtain high speaker-related frequency bands and low speaker-related frequency bands, including: According to any amplitude spectrum vector of any speaker and the vector N, an average feature vector of any speaker is calculated; Calculate the average feature vector of all speakers according to the average feature vector of any speaker and the number of people M; Calculating an F-ratio value based on the average feature vector of any speaker, the average feature vector of all speakers, the vector N, and the number of people M; The average F-ratio value corresponding to any frequency band is determined, and combined with the total number of frequency bands to obtain a division threshold, wherein the division threshold is used to divide the high speaker correlation frequency band set and the low speaker correlation frequency band set.
5. The method for defending against speech adversarial attacks on a speaker recognition system according to claim 4, characterized in that: According to the simulated adversarial sample, a high speaker-related frequency band masking threshold and a low speaker-related frequency band masking threshold are obtained respectively, including: Respectively calculating the simulated adversarial sample amplitude value of the simulated adversarial sample in any frequency band, and the original sample amplitude value of the corresponding original sample in any frequency band; Taking the absolute value of the difference between the amplitude value of the simulated adversarial sample and the amplitude value of the original sample and then taking the maximum value to obtain the high speaker correlation frequency band masking threshold; The pitch amplitude value of the simulated adversarial sample amplitude value is calculated and then averaged to obtain the low speaker-related frequency band masking threshold.
6. The method for defending against speech adversarial attacks on a speaker recognition system according to claim 5, characterized in that: Calculate the number of masks required for each frequency band according to the F-ratio value, and determine the random masking position in each frequency band, including: Get the amplitude value corresponding to any frequency band after Fratio calculation, the Fratio value corresponding to any frequency band and the masking threshold corresponding to the frequency band; For all the frequency bands smaller than the masking threshold corresponding to the frequency band, the corresponding amplitude values after Fratio calculation are summed and then multiplied by the Fratio value corresponding to the frequency band to obtain the number of each frequency band that needs to be masked; The original sample amplitude values that are smaller than the masking threshold corresponding to the frequency band and the number of each frequency band that needs to be masked are randomly selected to obtain the random masking position in each frequency band.
7. The method for defending against speech adversarial attacks on a speaker recognition system according to claim 1, characterized in that: The masked amplitude spectrum is reconstructed by a U-Net model to obtain a reconstructed amplitude spectrum diagram, including: Determine to use a feature recovery module of the U-Net model to reconstruct the masked amplitude spectrum, wherein the feature recovery module includes a four-layer encoder and a four-layer decoder; The masked amplitude spectrum is input into the encoder, a one-dimensional convolution is performed only along the frequency axis, and a frequency transform is connected before each encoding layer. Each encoder layer includes two compressed residual branches, a Snake activation function is used instead of a ReLU activation function, and a long short-term memory network and a time-based attention module are added to the internal layer of the encoder. The bottleneck feature is converted back to a complete amplitude spectrum of the same size as the encoder by the decoder; During training, the amplitude spectrum loss and multi-resolution short-time fast Fourier transform loss based on the F-ratio speaker feature distribution are calculated separately, and the total loss function is obtained by weighted summation.
8. The method for defending against speech adversarial attacks on a speaker recognition system according to claim 7, characterized in that: The amplitude spectrum loss based on the F-ratio speaker feature distribution includes: Obtain the original speech amplitude spectrum and the restored speech amplitude spectrum of any frequency band in the amplitude spectrum, as well as the F-ratio value of any frequency band in the amplitude spectrum; Based on the original speech amplitude spectrum, the restored speech amplitude spectrum and the F-ratio value of any frequency band in the amplitude spectrum, the amplitude spectrum loss based on the F-ratio speaker characteristic distribution is calculated; Correspondingly, the multi-resolution short-time fast Fourier transform loss includes; Obtaining a multi-resolution short-time Fourier transform value of the original speech audio and a multi-resolution short-time Fourier transform value of the restored speech audio; Based on the original speech audio multi-resolution short-time Fourier transform value and the restored speech audio multi-resolution short-time Fourier transform value, obtaining any resolution short-time fast Fourier transform loss; Different resolutions are formed by combining different fast Fourier transform frequencies, jump lengths and window sizes, and the short-time fast Fourier transform losses of different resolutions are summed and averaged to obtain the multi-resolution short-time fast Fourier transform loss.
9. A voice adversarial defense system for a speaker recognition system, characterized in that: include: An extraction module, used to obtain input speech data, perform feature extraction on the input speech data, and obtain an amplitude spectrogram of the input speech data; A masking module, used for processing the amplitude spectrum through a preset fine-grained non-robust feature filtering module to obtain a masked amplitude spectrum; A reconstruction module, used for reconstructing the masked amplitude spectrum through a U-Net model to obtain a reconstructed amplitude spectrum graph; The estimation module is used to perform phase estimation from the reconstructed amplitude spectrogram based on the Griffin-Lim algorithm to obtain a reconstructed speech signal.
10. An electronic device comprising a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein: When the processor executes the program, the speech adversarial defense method for a speaker recognition system as claimed in any one of claims 1 to 8 is implemented.
Citation Information
Patent Citations
Voiceprint recognition confrontation sample defense method based on cosine similarity and voice denoising
CN115188384A
Method and system for defending against sample attacks for speech recognition
CN115457939A
Confrontation sample construction method for voiceprint recognition defense module
CN116013318A
Robust speaker recognition method based on spectrogram denoising and adversarial learning
CN116469394A
Confrontation and defense method and system for voiceprint recognition system based on F-ratio self-adaptive masking
CN117219085A