SWM-based playback attack detection method and device, electronic equipment and storage medium

By using a SWM-based method and a DNN model, the feature differences between replayed speech and real speech are amplified, solving the misidentification problem of replay attack detection in existing technologies and achieving higher detection accuracy.

CN117219123BActive Publication Date: 2025-11-18GUANGDONG INST OF SCI & TECH
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202311179306.3
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2023-09-12
Publication Date
2025-11-18
Estimated Expiration
2043-09-12

AI Technical Summary

Technical Problem

In existing technologies, playback attack detection methods are prone to misidentification due to the similarity between the audio domain features and the reference frequency domain features of legitimate users, making it difficult to effectively distinguish between real speech and playback speech.

Method used

A method based on short-time-window average fuzzy quantization continuous wavelet transform (SWM) is adopted. By calculating the frequency domain sample spacing and sample weights, the target SWM CQCC features are extracted and combined with a deep neural network (DNN) model for detection, amplifying the feature differences between the replayed speech and the real speech.

Benefits of technology

It improves the accuracy of replay attack detection, enhances the recognition accuracy of DNN models, and effectively distinguishes between real speech and replay speech.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN117219123B_ABST
    Figure CN117219123B_ABST
Patent Text Reader

Abstract

The application provides a SWM-based playback attack detection method and device and electronic equipment, the method comprising: performing CQT transformation on a to-be-detected voice signal, obtaining a target amplitude spectrum by calculating the average value of all frame amplitude spectra, calculating the sample weight of each frequency domain sample according to the frequency domain sample interval of the target amplitude spectrum, obtaining a target SWM according to the sample weight and the target amplitude spectrum, extracting a target SWM CQCC from the target SWM, obtaining a preset DNN model, pre-labeled training data and test data, extracting a test SWM CQCC from the test data, inputting the target SWM CQCC, the training data and the test SWM CQCC into the DNN model, and obtaining a playback attack detection result output by the DNN model. According to the technical scheme of the embodiment of the application, the feature difference between the amplified playback voice and the real voice is amplified by the SWM, the recognition accuracy of the DNN model is improved, and the accuracy of the playback attack detection is improved.
Need to check novelty before this filing date? Find Prior Art

Description

TECHNICAL FIELD

[0001] The present application relates to the technical field of voice security, and in particular to a SWM-based playback attack detection method and device, electronic equipment and a storage medium. BACKGROUND

[0002] Voiceprint recognition is to compare the voiceprint features of different speakers with the corresponding voiceprint templates through an automatic speaker verification (ASV) system, so as to identify the identity of the speaker, and is widely used in various biometric recognition scenarios. However, when deploying an automatic speaker verification (ASV) system, three types of spoofing attacks may be encountered, including voice synthesis, voice conversion and playback attack. The playback attack mainly plays the recording of the actual voice of a legal client to the ASV system. In the related art, the playback attack is mainly detected by extracting the frequency domain features of the input voice. However, the sound source of the playback attack comes from a real legal user, so the directly extracted voice frequency domain features are similar to the reference frequency domain features of the legal user, which is easy to cause misrecognition. SUMMARY

[0003] The present application aims to at least solve one of the technical problems existing in the prior art. To this end, the present application provides a SWM-based playback attack detection method and device, electronic equipment and a storage medium, which can amplify the feature difference between the input voice and the sample voice, and improve the accuracy of playback attack detection.

[0004] In a first aspect, the present application provides a SWM-based playback attack detection method, comprising:

[0005] Obtaining a to-be-detected voice signal, performing CQT transformation on the to-be-detected voice signal to obtain a frequency domain voice signal;

[0006] Extracting the frame amplitude spectrum of each frame of the frequency domain voice signal, calculating the average value of all the frame amplitude spectrums, and obtaining a target amplitude spectrum according to a preset target frequency interval;

[0007] Calculating the sample weight of each frequency domain sample according to the frequency domain sample interval of the target amplitude spectrum, and obtaining a target SWM according to the sample weight and the target amplitude spectrum;

[0008] Extracting a target SWM CQCC from the target SWM;

[0009] Obtaining a preset DNN model, pre-labeled training data and test data;

[0010] Extracting a test SWM CQCC from the test data;

[0011] input the target SWM CQCC, the training data and the test SWM CQCC into the DNN model, and obtain a playback attack detection result of an output of the DNN model, where the output result is used to indicate whether the to-be-detected voice signal is real voice or playback voice.

[0012] According to some embodiments of the present application, the inputting of the target SWM CQCC, the training data and the test SWM CQCC into the DNN model comprises:

[0013] aligning the real sample voice signal and the playback sample voice signal in frequency;

[0014] obtaining real sample amplitude spectrum according to the target frequency interval by calculating the average value of the amplitude spectrum of each frame of the real sample voice signal;

[0015] obtaining playback sample amplitude spectrum according to the target frequency interval by calculating the average value of the amplitude spectrum of each frame of the playback sample voice signal;

[0016] determining real sample SWM corresponding to the real sample amplitude spectrum, and determining playback sample SWM corresponding to the playback sample amplitude spectrum;

[0017] extracting real sample SWM CQCC from the real sample SWM, and extracting playback sample SWM CQCC from the playback sample SWM;

[0018] training the DNN model according to the real sample SWM CQCC, the playback sample SWM CQCC and the test SWM CQCC;

[0019] inputting the target SWM CQCC into the trained DNN model to obtain the playback attack detection result.

[0020] According to some embodiments of the present application, the frequency range of the frequency domain voice signal, the real sample voice signal and the playback sample voice signal is 0Hz to 8000Hz, the target frequency interval is 0Hz to 2600Hz, and the training data and the test data are from an ASVspoof 2017 training set.

[0021] According to some embodiments of the present application, the training of the DNN model according to the real sample SWM CQCC, the playback sample SWM CQCC and the test SWM CQCC comprises:

[0022] inputting the real sample SWM CQCC and the playback sample SWM CQCC into the DNN model for training.

[0023] The error rate of the DNN model is determined based on the test SWM CQCC. When the false alarm rate and the error rate represented by the error rate are the same, the DNN model is considered to have completed training.

[0024] According to some embodiments of the present invention, the DNN model includes an input layer, an output layer and two hidden layers. The input layer consists of 11 frames of context windows for input feature vectors. The output layer has 2 nodes and the hidden layer has 512 nodes.

[0025] According to some embodiments of the present invention, the sample weight of each frequency domain sample is calculated based on the frequency domain sample spacing of the target amplitude spectrum, and the target SWM is obtained based on the sample weight and the target amplitude spectrum, using the following formula:

[0026]

[0027] SWM={(w)y1, (w)y2,...(w)y κ , (w)y κ+1 , ...y K-1 y K};

[0028] Where w is the sample weight, k is the number of frequency domain sample intervals, n is a preset positive integer, and Y is the target amplitude spectrum, satisfying Y = {y1, y2, ... y}. κ y κ+1 , ...y K-1 y K}y k This is the k-th frequency domain sample.

[0029] According to some embodiments of the present invention, extracting the target SWM CQCC from the target SWM includes:

[0030] After squaring and logarithmic processing of the target SWM, uniform resampling is performed to obtain the linear scale SWM.

[0031] The target SWM CQCC is obtained by performing a discrete cosine transform on the linear scale SWM.

[0032] Secondly, embodiments of the present invention provide a replay attack detection device based on SWM, including at least one control processor and a memory for communicatively connecting with the at least one control processor; the memory stores instructions executable by the at least one control processor, the instructions being executed by the at least one control processor to enable the at least one control processor to perform the SWM-based replay attack detection method as described in the first aspect above.

[0033] Thirdly, embodiments of the present invention provide an electronic device including a SWM-based replay attack detection device as described in the second aspect above.

[0034] Fourthly, embodiments of the present invention provide a computer-readable storage medium storing computer-executable instructions for executing the SWM-based replay attack detection method as described in the first aspect above.

[0035] The playback attack detection method based on SWM according to embodiments of the present invention has at least the following beneficial effects: acquiring a speech signal to be detected; performing CQT transformation on the speech signal to be detected to obtain a frequency domain speech signal; extracting the frame amplitude spectrum of each frame of the frequency domain speech signal; calculating the average value of all the frame amplitude spectra; obtaining a target amplitude spectrum according to a preset target frequency range; calculating the sample weight of each frequency domain sample according to the frequency domain sample spacing of the target amplitude spectrum; obtaining a target SWM according to the sample weight and the target amplitude spectrum; extracting a target SWM CQCC from the target SWM; acquiring a preset DNN model, pre-labeled training data and test data; extracting a test SWM CQCC from the test data; inputting the target SWM CQCC, the training data and the test SWM CQCC into the DNN model; obtaining the playback attack detection result output by the DNN model, wherein the output result is used to indicate whether the speech signal to be detected is real speech or playback speech. According to the technical solution of the present invention, the feature differences between replayed speech and real speech are amplified by SWM, thereby improving the recognition accuracy of the DNN model and thus improving the accuracy of replay attack detection. Attached Figure Description

[0036] Figure 1 This is a flowchart of a replay attack detection method based on SWM provided in an embodiment of the present invention;

[0037] Figure 2 This is a flowchart of training a DNN model provided in another embodiment of the present invention;

[0038] Figure 3This is a full-frequency waveform diagram of training data provided in another embodiment of the present invention;

[0039] Figure 4 This is a waveform diagram of training data in the target frequency range provided by another embodiment of the present invention;

[0040] Figure 5 This is a flowchart for determining training completion provided in another embodiment of the present invention;

[0041] Figure 6 This is a flowchart for extracting SVM CQCC provided in another embodiment of the present invention;

[0042] Figure 7 This is a schematic diagram of an SVM CQCC extraction module provided in another embodiment of the present invention;

[0043] Figure 8 This is a structural diagram of a SWM-based replay attack detection device provided in another embodiment of the present invention. Detailed Implementation

[0044] Embodiments of the present invention are described in detail below. Examples of these embodiments are shown in the accompanying drawings, wherein the same or similar reference numerals denote the same or similar elements or elements having the same or similar functions throughout. The embodiments described below with reference to the accompanying drawings are exemplary and are only used to explain the present invention, and should not be construed as limiting the present invention.

[0045] In the description of this invention, it should be understood that the orientation descriptions, such as up, down, front, back, left, right, etc., are based on the orientation or positional relationship shown in the accompanying drawings. They are only for the convenience of describing this invention and simplifying the description, and do not indicate or imply that the device or element referred to must have a specific orientation, or be constructed and operated in a specific orientation. Therefore, they should not be construed as limiting this invention.

[0046] In the description of this invention, "several" means one or more, "more than" means two or more, "greater than," "less than," and "exceeding" are understood to exclude the stated number, while "above," "below," and "within" are understood to include the stated number. The use of "first" and "second" in the description is merely for distinguishing technical features and should not be construed as indicating or implying relative importance, or implicitly indicating the number of indicated technical features, or implicitly indicating the order of the indicated technical features.

[0047] In the description of this invention, unless otherwise explicitly defined, terms such as "set up," "install," and "connect" should be interpreted broadly, and those skilled in the art can reasonably determine the specific meaning of the above terms in this invention in conjunction with the specific content of the technical solution.

[0048] This invention provides a method, apparatus, and electronic device for detecting playback attacks based on Swing Wires (SWM). The method includes: acquiring a speech signal to be detected; performing a CQT transform on the speech signal to obtain a frequency domain speech signal; extracting the frame amplitude spectrum of each frame of the frequency domain speech signal; calculating the average value of all frame amplitude spectra; obtaining a target amplitude spectrum based on a preset target frequency range; calculating the sample weight of each frequency domain sample based on the frequency domain sample interval of the target amplitude spectrum; obtaining a target Swing Wire based on the sample weight and the target amplitude spectrum; extracting a target Swing Wire CQCC from the target Swing Wire; acquiring a preset DNN model, pre-labeled training data, and test data; extracting a test Swing Wire CQCC from the test data; inputting the target Swing Wire CQCC, the training data, and the test Swing Wire CQCC into the DNN model; and obtaining the playback attack detection result output by the DNN model, wherein the output result indicates whether the speech signal to be detected is real speech or playback speech. According to the technical solution of the present invention, the feature differences between replayed speech and real speech are amplified by SWM, thereby improving the recognition accuracy of the DNN model and thus improving the accuracy of replay attack detection.

[0049] The control method of the present invention will be further described below with reference to the accompanying drawings.

[0050] Reference Figure 2 , Figure 2 The flowchart illustrates a replay attack detection method based on SWM, which includes, but is not limited to, the following steps:

[0051] S11, acquire the speech signal to be detected, perform CQT transformation on the speech signal to be detected to obtain the frequency domain speech signal;

[0052] S12, extract the frame amplitude spectrum of each frame of the frequency domain speech signal, and obtain the target amplitude spectrum according to the preset target frequency range after calculating the average value of all frame amplitude spectra.

[0053] S13, calculate the sample weight of each frequency domain sample based on the frequency domain sample spacing of the target amplitude spectrum, and obtain the target SWM based on the sample weight and the target amplitude spectrum;

[0054] S14, Extract the target SWM CQCC from the target SWM:

[0055] S15, obtain the preset DNN model, pre-labeled training data and test data;

[0056] S16, Extract the test SWM CQCC from the test data:

[0057] S17, input the target SWM CQCC, training data and test SWM CQCC into the DNN model, and obtain the playback attack detection result of the DNN model output. The output result is used to indicate whether the speech signal to be detected is real speech or playback speech.

[0058] It should be noted that CQT can convert speech signals from the time domain to the frequency domain, providing a data basis for calculating amplitude spectrum values ​​in the mass spectrometry stage.

[0059] It should be noted that after extracting the frequency domain speech signal, the frame amplitude spectrum of each frame is extracted first, and then the values ​​at the corresponding positions of the frame amplitude spectrum are averaged to obtain the target amplitude spectrum.

[0060] In one embodiment, after obtaining the target amplitude spectrum Y = {y1, y2, ... y} κ y κ+1 , ...y K-1 y K}y k Next, assuming Y is the amplitude spectrum of a given frame, K is the total number of frequency sample intervals, K+1 represents the frequency sample interval above K, and n is a positive integer, then... SWM={(w)y1, (w)y2,...(w)y κ , (w)y κ+1 , ...y K-1 y K}, where w is the sample weight, y k Let be the k-th frequency domain sample. By calculating the target SWM, the amplitude of the spectrum can be weighted, thereby amplifying the difference between the replayed speech and the real speech and improving recognition efficiency.

[0061] It should be noted that in the target SWM CQCC, the SWM function is used to amplify the difference between the playback and the real speech, thereby improving the accuracy of model recognition.

[0062] It should be noted that DNN models can adopt common structures, and those skilled in the art are familiar with how to configure DNN models, so we will not go into details here.

[0063] It should be noted that the test data can be the ASVspoof 2017 dataset. Using ASVspoof 2017 for the challenge will result in higher performance of the trained model.

[0064] It should be noted that after the DNN model training is completed, the target SWM CQCC is input into the DNN model, and the DNN model's classifier outputs the playback attack detection result, thereby determining whether the speech signal to be detected belongs to the playback speech or the real speech, thus realizing the automatic detection of playback attacks.

[0065] Additionally, in one embodiment, reference is made to Figure 2 , Figure 1 Step S17 shown also includes, but is not limited to, the following steps:

[0066] S21, frequency alignment between the real sample speech signal and the playback sample speech signal;

[0067] S22, after calculating the average value of the amplitude spectrum of each frame of the real sample speech signal, the amplitude spectrum of the real sample is obtained according to the target frequency range.

[0068] S23, after calculating the average value of the amplitude spectrum of each frame of the playback sample speech signal, obtain the amplitude spectrum of the playback sample according to the target frequency range;

[0069] S24, determine the true sample SWM corresponding to the true sample amplitude spectrum, and determine the replay sample SWM corresponding to the replay sample amplitude spectrum;

[0070] S25, extract the real sample SWM CQCC from the real sample SWM, and extract the replay sample SWM CQCC from the replay sample SWM;

[0071] S26, Train a DNN model based on real sample SWM CQCC, replay sample SWM CQCC and test SWM CQCC;

[0072] S27. Input the target SWM CQCC into the trained DNN model to obtain the replay attack detection result.

[0073] It should be noted that the ASVspoof 2017 training set contains 3016 utterances, including 1508 real speech samples and 1508 playback samples. Due to the playback recording process, a pair of real and playback speech utterances were not aligned. To analyze the spectral amplitude differences between the real and playback speech, it is necessary to first align the real and playback speech utterance pairs.

[0074] It should be noted that in this embodiment, the amplitude spectrum (MS) of each frame of the 1508 real sample speech signals and the playback sample speech signals were calculated. Then, we calculated the average MS value of the real speech and the playback speech. Assuming K is the frequency domain sample spacing of our MS, as shown... Figure 3As shown, the average MS of the real speech and the replayed speech can be represented by Yg and Yp, respectively. Figure 3 This shows a comparison of Yg and Yp across the entire training data from 0 to 8000Hz. Figure 3 As can be seen, there is a significant spectral difference between Yg and Yp. The pronounced bass may be attributed to the overcompensation of the weaker low end by the mid-bass cone. The enhanced mid-high frequencies may be attributed to conventional pre and post compressor filtering, the weak bass response characteristics, and the interpretability considerations in speaker design. The sharp attenuation near the highest frequencies is likely due to Nyquist filtering. To amplify the differences between the two and distinguish them, this embodiment introduces SWM for weighted calculations to ensure the training accuracy of the DNN model.

[0075] In one embodiment, the frequency range of the frequency domain speech signal, the real sample speech signal, and the playback sample speech signal is 0 Hz to 8000 Hz, the target frequency range is 0 Hz to 2600 Hz, and the training data and test data are from the ASVspoof 2017 training set.

[0076] It should be noted that, from Figure 3 As can be seen, this applies to the frequency spectrum between 0Hz and 2600Hz. Figure 3 It can be seen that the difference between Yg and Yp is very small. To amplify their difference, their sum can be used as the weight for normalization of the relevant frequency domain sample intervals within the frequency band. In this embodiment, real and replayed speech in the region between 0Hz and 2600Hz are used. However, the region between 2000Hz and 8000Hz is not weighted. Since the weighting is derived from the statistical average, it is called the statistically weighted amplitude spectrum (SWM). Assuming Y is the amplitude spectrum of a given frame, K is the total number of frequency domain sample intervals, which in this embodiment is the frequency domain sample interval at 2600kHz, K+1 represents the frequency domain sample interval above K, and n is a positive integer, the SWM calculation formula described above is used: SWM={(w)y1,(w)y2,...(w)y κ , (w)y κ+1 , ...y K-1 y K}, we can obtain the SWM corresponding to Yg. g SWM corresponding to Yp p Its waveform can be referenced. Figure 4 As shown, from Figure 4 As can be seen, the difference between SWMg and SWMp is amplified between 0 and 2600 Hz, making them easily distinguishable in this region. Furthermore, from... Figure 3 As can be seen, Yg and pY can be distinguished in the 2600Hz to 8000Hz range. Therefore, it is sufficient to say that SWM allows playback and real speech to be easily distinguished across the entire frequency band.

[0077] Additionally, in one embodiment, reference is made to Figure 5 , Figure 2 Step S26 shown also includes, but is not limited to, the following steps:

[0078] S51, Input the real sample SWM CQCC and the replay sample SWM CQCC into the DNN model for training;

[0079] S52, determine the equal error rate of the DNN model based on the test SWM CQCC. When the false alarm rate and the error rate represented by the equal error rate are the same, the DNN model is considered to have completed training.

[0080] It should be noted that the ASVspoof 2017 challenge database was collected from 15 playback devices and 16 recording devices at 4 different locations. It consists of three subsets: training data, development data, and evaluation data. Table 1 lists the details of ASVspoof2017.

[0081] Set #Genuine #Replay #Total Training data 1508 1508 3016 Development data 760 950 1710 Evaluation data 1298 12008 13306

[0082] Table 1. ASVSPOOF 2017 Corpus

[0083] Two types of models were trained according to the rules of the ASVspoof 2017 challenge. The first method evaluated the proposed algorithm based on performance on an evaluation set. This model used 4,726 utterances from both the training and development sets. The second method evaluated the proposed algorithm based on performance on a development set. This model used 3,016 utterances from the training set. Additionally, the equal error rate (EER) was used as an evaluation metric, representing the percentage of false alarms and missed alarms that are equal at a certain threshold.

[0084] In CQT, several important parameters affect the final performance. These are the frequency domain sample interval per octave, the number of octaves, the sampling period, and the gamma value. For consistency and comparison, in this embodiment, the frequency domain sample interval per octave is set to 96, the number of octaves is set to 9, the sampling period is set to 16, and the gamma value is set to 3.3026. Additionally, the static dimension of SWM-CQCC is set to 20. Different combinations of features, S, D, and A, represent static, incremental, and acceleration, respectively.

[0085] The Computational Network Toolkit (CNTK) is used to train the DNN model, employing stochastic gradient descent (SGD) during training. This embodiment provides a four-layer DNN classifier with two hidden layers (512 nodes each), an output layer with two nodes, and an input layer consisting of 11-frame context windows of the input feature vector. The feature combinations of SWM-CQCC require very different input layers. For example, for SWM-CQCC-S, the input layer consists of 20 x 11 nodes, with 5 frames each on the left and right sides, while for SWM-CQCC-SDA, the input layer is 60 x 11.

[0086] Referring to Table 2, which presents the recognition results (EER) of the ASVspoof2017 development set using SWM-CQCC under different feature combinations and n conditions, it can be seen from Table 2 that: when N=1, SWM-CQCC-DA performs best, followed by SWM-CQCC-D; when n=3, SWM-CQCC-SDA performs best, followed by SWM-CQCC-DA; when n=5, SWM-CQCC-D performs best, followed by SWM-CQCC-DA. Based on this, this embodiment determines SWM-CQCC-DA and SWM-CQCC-D as relatively suitable features for PSD.

[0087]

[0088]

[0089] Table 2: Schematic diagram of results for different feature combinations

[0090] Furthermore, Table 3 compares the performance of the SWM-CQCC-DA algorithm with existing systems on the ASVspoof 2017 evaluation set. As can be seen from Table 3, the performance of SWM-CQCC-DA significantly outperforms the features used in existing systems. This is due to the discriminative quality of SWM, which can amplify the differences between real and replayed speech in SWM-CQCC-DA.

[0091] Features EER LPCCres(6-8KHz) 27.61 Cepstrum(6-8KHz) 22.24 I-vector based on LPCC 12.54 CQCC-DA 19.18 HFCC 23.90 MFCC 16.26 SFFCC-D 20.20 IFCC 35.19 VESA-IFCC 15.50 SWM-CQCC-DAA 10.27

[0092] Table 3 compares the performance of SWM-CQCC-DA with existing systems.

[0093] In another embodiment, the DNN model includes an input layer, an output layer, and two hidden layers. The input layer consists of 11 frames of context windows for the input feature vector, the output layer has 2 nodes, and the hidden layer has 512 nodes.

[0094] Additionally, in one embodiment, reference is made to Figure 6 , Figure 1Step S16 shown also includes, but is not limited to, the following steps:

[0095] S61, after squaring and logarithmic processing of the target SWM, uniform resampling is performed to obtain the linear scale SWM;

[0096] S62, perform discrete cosine transform on the linear scale SWM to obtain the target SWM CQCC.

[0097] It should be noted that, from Figure 7 As shown in the diagram, the SVM CQCC extraction module comprises six modules: Constant q-Transform (CQT), Amplitude Spectrum (MS), Squaring, Logarithmic, Uniform Resampling, and DCT. CQT first converts the speech from the time domain to the frequency domain. Then, the amplitude spectrum values ​​are calculated in the mass spectrum stage. After squaring and logarithmic calculations, uniform resampling is performed. SWM, performed after the MS stage, converts the frequency representation from a non-linear octave scale to a linear scale. Finally, the Discrete Cosine Transform (DCT) method is used to extract the main information.

[0098] It's important to note that the Discrete Cosine Transform (DCT) is essentially a DCT of approximately twice its length. This DCT is performed on a real even function. The spectrum obtained from the Fourier transform of a real function is mostly complex, while the Fourier transform of an even function results in a real function. Based on this, by making the signal function an even function and removing the imaginary part of the spectral function, a set of light intensity data can be converted into frequency data through the cosine transform, allowing us to understand the intensity variations.

[0099] like Figure 8 As shown, Figure 8 This is a structural diagram of a SWM-based replay attack detection device according to an embodiment of the present invention. The present invention also provides a SWM-based replay attack detection device, comprising:

[0100] The processor 801 can be implemented using a general-purpose central processing unit (CPU), microprocessor, application specific integrated circuit (ASIC), or one or more integrated circuits, and is used to execute relevant programs to implement the technical solutions provided in the embodiments of this application.

[0101] The memory 802 can be implemented as a read-only memory (ROM), static storage device, dynamic storage device, or random access memory (RAM). The memory 802 can store the operating system and other applications. When the technical solutions provided in the embodiments of this specification are implemented through software or firmware, the relevant program code is stored in the memory 802 and is called and executed by the processor 801 to implement the SWM-based replay attack detection method of this application embodiment.

[0102] The 803 input / output interface is used to implement information input and output.

[0103] The communication interface 804 is used to enable communication and interaction between this device and other devices. Communication can be achieved through wired means (such as USB, Ethernet cable, etc.) or wireless means (such as mobile network, WIFI, Bluetooth, etc.).

[0104] Bus 805 transmits information between various components of the device (e.g., processor 801, memory 802, input / output interface 803, and communication interface 804);

[0105] The processor 801, memory 802, input / output interface 803, and communication interface 804 are connected to each other within the device via bus 805.

[0106] This application also provides an electronic device, including the SWM-based replay attack detection device described above.

[0107] This application embodiment also provides a storage medium, which is a computer-readable storage medium, storing a computer program that, when executed by a processor, implements the above-described SWM-based replay attack detection method.

[0108] Memory, as a non-transitory computer-readable storage medium, can be used to store non-transitory software programs and non-transitory computer-executable programs. Furthermore, memory may include high-speed random access memory, and may also include non-transitory memory, such as at least one disk storage device, flash memory device, or other non-transitory solid-state storage device. In some embodiments, memory may optionally include memory remotely located relative to the processor, and these remote memories can be connected to the processor via a network. Examples of such networks include, but are not limited to, the Internet, intranets, local area networks, mobile communication networks, and combinations thereof. The device embodiments described above are merely illustrative, and the units described as separate components may or may not be physically separate, and may be located in one place or distributed across multiple network units. Some or all of the modules can be selected to achieve the purpose of this embodiment according to actual needs.

[0109] It will be understood by those skilled in the art that all or some of the steps and systems in the methods disclosed above can be implemented as software, firmware, hardware, and suitable combinations thereof. Some or all of the physical components can be implemented as software executed by a processor, such as a central processing unit, digital signal processor, or microprocessor, or as hardware, or as an integrated circuit, such as an application-specific integrated circuit. Such software can be distributed on a computer-readable medium, which can include computer storage media (or non-transitory media) and communication media (or transient media). As is known to those skilled in the art, the term computer storage media includes volatile and non-volatile, removable and non-removable media implemented in any method or technology for storing information (such as computer-readable instructions, data structures, program modules, or other data). Computer storage media includes, but is not limited to, RAM, ROM, EEPROM, flash memory or other memory technologies, CD-ROM, digital versatile disc (DVD) or other optical disc storage, magnetic cartridges, magnetic tape, disk storage or other magnetic storage devices, or any other medium that can be used to store desired information and is accessible to a computer. Furthermore, as is known to those skilled in the art, communication media typically include computer-readable instructions, data structures, program modules, or other data in modulated data signals such as carrier waves or other transmission mechanisms, and may include any information delivery medium.

[0110] The above provides a detailed description of the preferred embodiments of the present invention. However, the present invention is not limited to the above embodiments. Those skilled in the art can make various equivalent modifications or substitutions without departing from the spirit of the present invention. All such equivalent modifications or substitutions are included within the scope defined by the claims of the present invention.

Claims

1. A replay attack detection method based on SWM, characterized in that, include: Acquire the speech signal to be detected, and perform CQT transformation on the speech signal to be detected to obtain the frequency domain speech signal; Extract the frame amplitude spectrum of each frame of the frequency domain speech signal, and obtain the target amplitude spectrum according to the preset target frequency range after calculating the average value of all the frame amplitude spectra. The sample weight of each frequency domain sample is calculated based on the frequency domain sample spacing of the target amplitude spectrum, and the target SWM is obtained based on the sample weight and the target amplitude spectrum. Extract the target SWM CQCC from the target SWM; Obtain the preset DNN model, pre-labeled training data, and test data; Extract the test SWM CQCC from the test data; The target SWM CQCC, the training data, and the test SWM CQCC are input into the DNN model to obtain the playback attack detection result of the output of the DNN model, wherein the playback attack detection result is used to indicate whether the speech signal to be detected is real speech or playback speech; The sample weight of each frequency domain sample is calculated based on the frequency domain sample spacing of the target amplitude spectrum. The target SWM is obtained based on the sample weight and the target amplitude spectrum, using the following formula: ; ; Where w is the sample weight, k is the number of frequency domain sample intervals, n is a preset positive integer, and Y is the target amplitude spectrum, and satisfies , This is the k-th frequency domain sample.

2. The SWM-based replay attack detection method according to claim 1, characterized in that, The step of inputting the target SWM CQCC, the training data, and the test SWM CQCC into the DNN model includes: Frequency alignment is performed between the real sample speech signal and the replay sample speech signal; After calculating the average value of the amplitude spectrum of each frame of the real sample speech signal, the amplitude spectrum of the real sample is obtained according to the target frequency range. After calculating the average value of the amplitude spectrum of each frame of the playback sample speech signal, the amplitude spectrum of the playback sample is obtained according to the target frequency range. Determine the true sample SWM corresponding to the true sample amplitude spectrum, and determine the replay sample SWM corresponding to the replay sample amplitude spectrum; Extract the real sample SWM CQCC from the real sample SWM, and extract the replay sample SWM CQCC from the replay sample SWM; The DNN model is trained based on the real sample SWM CQCC, the replay sample SWM CQCC, and the test SWM CQCC. The target SWM CQCC is input into the trained DNN model to obtain the replay attack detection result.

3. The SWM-based replay attack detection method according to claim 2, characterized in that, The frequency range of the frequency domain speech signal, the real sample speech signal, and the playback sample speech signal is 0Hz to 8000Hz, the target frequency range is 0Hz to 2600Hz, and the training data and test data are from the ASVspoof 2017 training set.

4. The SWM-based replay attack detection method according to claim 2, characterized in that, The step of training the DNN model based on the real sample SWM CQCC, the replay sample SWM CQCC, and the test SWM CQCC includes: The real sample SWM CQCC and the replay sample SWM CQCC are input into the DNN model for training; The error rate of the DNN model is determined based on the test SWM CQCC. When the false alarm rate and the error rate represented by the error rate are the same, the DNN model is considered to have completed training.

5. The SWM-based replay attack detection method according to claim 1, characterized in that, The DNN model includes an input layer, an output layer, and two hidden layers. The input layer consists of 11 frames of context windows for the input feature vector. The output layer has 2 nodes, and the hidden layer has 512 nodes.

6. The SWM-based replay attack detection method according to claim 1, characterized in that, Extracting the target SWM CQCC from the target SWM includes: After squaring and logarithmic processing of the target SWM, uniform resampling is performed to obtain the linear scale SWM. The target SWM CQCC is obtained by performing a discrete cosine transform on the linear scale SWM.

7. A replay attack detection device based on SWM, characterized in that, It includes at least one control processor and a memory for communicatively connecting to the at least one control processor; The memory stores instructions that can be executed by the at least one control processor, which, when executed, enables the at least one control processor to perform the SWM-based replay attack detection method as described in any one of claims 1 to 6.

8. An electronic device, characterized in that, Includes the SWM-based replay attack detection device as described in claim 7.

9. A computer-readable storage medium, characterized in that, The computer-readable storage medium stores computer-executable instructions for causing a computer to perform the SWM-based replay attack detection method as described in any one of claims 1 to 6.

Citation Information

Patent Citations

  • Training method of emotion recognition model, emotion recognition method, device, equipment, and storage medium

    CN109817246A

  • Playback attack detection method and training method of corresponding detection model

    CN110718229A