A method and device for detecting voice spoofing attacks

The method addresses voice spoofing detection in voice recognition systems by employing linear frequency cepstral coefficients and a deep learning model with Res2Net and coordinate attention, enhancing detection accuracy and model generalization.

CN119049479BActive Publication Date: 2025-07-15HUAZHONG UNIV OF SCI & TECH
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202411211329.2
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2024-08-30
Publication Date
2025-07-15
Estimated Expiration
2044-08-30

AI Technical Summary

Technical Problem

The prior art is difficult to effectively extract acoustic features in voice spoofing attack detection, resulting in low detection accuracy and insufficient universality and generalization capabilities of the model.

Method used

The linear frequency cepspectral coefficient is used as the acoustic feature of speech, combined with the deep learning model and coordinate attention mechanism, the frequency domain characteristics of speech signals are extracted through linear filtering and Fourier transform, and the detection models of the convolutional layer, Res2Net module, attention module and full connection layer are built to perform speech spoof attack detection.

Benefits of technology

It improves the accuracy of voice spoofing attack detection and the universality of the model, weakens the impact of environmental noise, and enhances the ability of the detection model to adapt to different types of voice spoofing attacks.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119049479B_ABST
    Figure CN119049479B_ABST
Patent Text Reader

Abstract

The present invention discloses a method for detecting voice spoofing attacks, comprising the following steps: Step S1, loading voice data in a voice training set, sampling the voice data at a preset sampling rate, and converting the voice data into a digital voice signal; Step S2, inputting the digital voice signal into a linear filter to extract linear frequency cepstral coefficients describing the frequency domain characteristics of the digital voice signal; Step S3, building a detection model for voice spoofing attacks, and training the detection model using the linear frequency cepstral coefficients of the digital signal; Step S4, inputting the voice to be detected into the trained detection model, and judging whether the voice to be detected is spoofed voice according to the output of the detection model. The present invention has the effect of high accuracy in voice spoofing detection.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of Internet of Things security, and in particular to a method and device for detecting voice spoofing attacks. Background Art

[0002] Voice spoofing attack detection is a core technical issue in the development of speech recognition and a basic requirement for a speech recognition system to operate properly and provide reliable verification services. Voice spoofing attacks, including voice synthesis attacks, voice conversion attacks, and voice replay attacks, will greatly reduce the reliability, credibility, and availability of speech recognition systems and affect network security. The task of voice spoofing attack detection is to detect the voice input to the speech recognition system and determine whether there is a voice spoofing attack in as little time as possible. Therefore, it is particularly important to perform real-time and effective spoofing attack detection on voices.

[0003] The difficulties in voice spoofing attack detection are mainly reflected in three aspects. First, there are differences between different types of spoofed voices, and the choice of model will directly affect the generality and generalization ability of the voice spoofing attack detection method; second, select a suitable attention mechanism for the model to pay more attention to the parts beneficial to attack detection during the information extraction process to improve the accuracy of the voice spoofing attack detection model; third, before the voice is input into the attack detection model, it is necessary to select and extract effective acoustic features to contain more voice information and increase the detection accuracy of the model. Summary of the Invention

[0004] In view of this, it is necessary to provide a method and device for detecting voice spoofing attacks to effectively solve the technical problems that effective acoustic features cannot be extracted in voice attack detection and the detection accuracy is low.

[0005] The present invention provides a method for detecting voice spoofing attacks, including the following steps:

[0006] Step S1: Load the voice data in the voice training set, sample the voice data at a preset sampling rate, and convert the voice data into a digital voice signal;

[0007] Step S2: Input the digital voice signal into a linear filter to extract the linear frequency cepstral coefficients describing the frequency domain characteristics of the digital voice signal;

[0008] Step S3: Build a detection model for voice spoofing attacks and train the detection model using the linear frequency cepstral coefficients of the digital signal;

[0009] Step S4: Input the voice to be detected into the trained detection model and determine whether the voice to be detected is a spoofed voice according to the output of the detection model.

[0010] Preferably, the step S2 is specifically as follows:

[0011] Step S21: Pre-emphasize the digital voice signal;

[0012] Step S22: Frame the digital voice signal to obtain multiple short-time frames, and connect adjacent frames;

[0013] Step S23: Multiply each short-time frame by a Hamming window function;

[0014] Step S24: Perform Fourier transform on each short-time frame to obtain the spectrum of the digital voice signal;

[0015] Step S25: Take the modulus square of the spectrum of the digital voice signal to obtain the power energy spectrum;

[0016] Step S26: Extract the frequency band characteristics of the power energy spectrum through a group of linear scale filters, and apply discrete cosine transform to retain the cepstrum coefficients to obtain the linear frequency cepstrum coefficients of the digital voice signal.

[0017] Preferably, the step S24 is specifically as follows: When performing N-point Fourier transform on each short-time frame, if the signal length of the short-time frame is insufficient, it is padded with zeros.

[0018] Preferably, in the step S21, the pre-emphasis coefficient for pre-emphasizing the digital voice signal is taken as 0.95 or 0.97.

[0019] Preferably, the step S3 is specifically as follows:

[0020] Step S31: Build a detection model for voice spoofing attack, and the detection model includes a convolutional layer, a Res2Net module, an attention module, a fully connected layer, and an activation function;

[0021] Step S32: Input the linear frequency cepstrum coefficients of the digital signal into the detection model, and adjust the parameters of the detection model through backpropagation until a preset end condition is reached.

[0022] Preferably, the step S32 is specifically as follows:

[0023] Step S321: Input the linear frequency cepstrum coefficients into the convolutional layer of the detection model for sliding window processing to extract initial feature information;

[0024] Step S322: Input the initial feature information into the Res2Net module and the attention module of the detection model to obtain a feature matrix;

[0025] Step S323: Obtain a feature description vector by performing global average pooling on the feature matrix, perform a flattening operation on the feature description vectors on each channel, input the flattened feature description vectors into the fully connected layer of the detection model, and after the output of the fully connected layer passes through an activation function, obtain the final prediction vector;

[0026] Step S324: According to the data and labels of the prediction vector, use the binary cross-entropy loss function to backpropagate and calculate the gradient, and adjust the weights and biases of the detection model.

[0027] Preferably, the attention module of the detection model uses global average pooling to respectively extract the position information of the input signal in the vertical and horizontal directions, calculates the attention weights in the vertical and horizontal directions respectively according to the position information in the vertical and horizontal directions, and obtains the output of the attention module.

[0028] Preferably, step S4 is specifically:

[0029] Step S41: Load the speech to be detected in the database, extract the linear frequency cepstral coefficients of the speech to be detected, and input them into the trained detection model;

[0030] Step S42: Judge whether the speech to be detected is spoofed speech according to the prediction vector output by the detection model, and output the detection result.

[0031] Preferably, step S41 is specifically:

[0032] Input the linear frequency cepstral coefficients of the speech to be detected into the convolutional layer, Res2Net module, attention module, fully connected layer and activation function of the detection model in sequence to obtain the prediction vector of the speech to be detected.

[0033] The present invention also provides a voice spoofing attack detection device, including a memory and a processor, wherein a computer program is stored on the memory, and when the computer program is executed by the processor, the voice spoofing attack detection method is implemented.

[0034] Compared with the prior art, the beneficial effects of the present invention are as follows: In the present invention, linear frequency cepstral coefficients are used as the acoustic features of speech, which are closer to the frequency resolution perceived by the human ear and can more accurately reflect the frequency characteristics of speech signals. Due to the selection of the linear frequency scale, the linear frequency cepstral coefficients can weaken the influence of environmental noise on the features to a certain extent, can be used for different types of voice spoofing attack detections, and improve the generality of the detection model. Description of the Drawings

[0035] The accompanying drawings described herein are used to provide a further understanding of the present invention and form a part of this application. The schematic embodiments of the present invention and their descriptions are used to explain the present invention and do not constitute an improper limitation of the present invention. In the drawings:

[0036] Figure 1 It is a flowchart of an embodiment of a voice spoofing attack detection method provided by the present invention;

[0037] Figure 2 is Figure 1 The detection principle diagram of an embodiment of the voice spoofing attack detection method in the illustrated embodiment;

[0038] Figure 3 is Figure 1 The structural principle diagram of an embodiment of the detection model in the illustrated embodiment;

[0039] Figure 4 is Figure 1 The detection result example diagram of an embodiment of the voice spoofing attack detection method in the illustrated embodiment. Specific Embodiments

[0040] The following will specifically describe the preferred embodiments of the present invention in conjunction with the accompanying drawings. Among them, the accompanying drawings form a part of this application and are used together with the embodiments of the present invention to explain the principles of the present invention, and are not used to limit the scope of the present invention.

[0041] The following will first explain and illustrate the technical terms of the present invention:

[0042] Res2Net: A convolutional neural network structure that enhances the network's perception ability of multi-scale information and improves the performance of speech recognition tasks by introducing multi-scale residual connections and grouped feature extraction.

[0043] Coordinate Attention, abbreviated as CA in English: An attention mechanism that can assign different weights to the features at different positions according to the position information of the input data to enhance the model's ability to model position-related information.

[0044] Linear Frequency Cepstral Coefficients: A representation method used for feature extraction in speech signal processing. By performing short-time Fourier transform on the speech signal, then linearly filtering and discrete cosine transform on the spectrum, the coefficients representing speech features are obtained.

[0045] Global Average Pooling, abbreviated as GAP in English: Global average pooling is a feature extraction operation that averages the feature values at all positions on the feature map to obtain a fixed-length feature vector to represent the overall feature information.

[0046] Softmax activation function: A function used for multi-class classification problems that converts a set of real numbers into a vector representing a probability distribution, where each element represents the probability of the corresponding class.

[0047] Example 1

[0048] Please refer to Figure 1 , a voice spoofing attack detection method in this embodiment specifically includes the following steps:

[0049] Step S1: Load the voice data in the voice training set, sample the voice data at a preset sampling rate, and convert the voice data into a digital voice signal;

[0050] Step S2: Input the digital voice signal into a linear filter to extract the linear frequency cepstral coefficients that describe the frequency domain characteristics of the digital voice signal;

[0051] Step S3: Build a detection model for voice spoofing attacks and train the detection model using the linear frequency cepstral coefficients of the digital signal;

[0052] Step S4: Input the voice to be detected into the trained detection model and determine whether the voice to be detected is a spoofed voice according to the output of the detection model.

[0053] Most existing methods use Gaussian classifiers or linear classifiers for simple classification, and the performance and classification ability of the models are poor. The detection model adopted in this embodiment is a deep learning model, which can increase the equivalent receptive field by modifying the structure of the bottleneck block, so as to perceive information from different scales and improve the representation ability of the detection model. At the same time, existing voice spoofing attack detection models usually use traditional acoustic features as the feature representation of voices, but these features all have problems of overly complex operations and a lot of redundant information. Using linear frequency cepstral coefficients as features in this embodiment can improve the compactness and computational efficiency of the features by selecting appropriate linear frequency scales.

[0054] Specifically, as Figure 2As shown in the figure, in this embodiment, the voice data of the training set in the database is first loaded, the sampling rate is set according to a preset standard according to the time length of the voice, the analog signal is converted into a digital signal, and the sampled continuous amplitude values are mapped into discrete amplitude levels to obtain a digital voice signal that is discrete in both time and amplitude. Feature extraction module: The digital voice signal is processed and then input into a linear filter to extract the linear frequency cepstral coefficients that describe the frequency domain characteristics of the digital voice signal. Model training module: Build a voice spoofing attack detection model, and use the linear frequency cepstral coefficients of the training set voice to train the model. Load the voice to be detected, and use the trained voice spoofing attack detection model to perform attack detection, and judge whether it is spoofed voice according to the prediction vector output by the model.

[0055] This embodiment uses linear frequency cepstral coefficients, which are closer to the frequency resolution perceived by the human ear and can more accurately reflect the frequency characteristics of the voice signal. Due to the selection of the linear frequency scale, the linear frequency cepstral coefficients can weaken the influence of environmental noise on the features to a certain extent and can be used for different types of voice spoofing attack detection, improving the generality of the detection model. At the same time, the structures of different modules of the detection model in this embodiment are clear and complete, and can be flexibly applied and adjusted in other deep learning tasks. The detection method of the present invention has high detection accuracy, strong generality, high efficiency, and is easy to integrate.

[0056] Further, as Figure 2 shown, step S2 is specifically:

[0057] Step S21, pre-emphasize the digital voice signal;

[0058] In order to balance the amplitudes in the high-frequency and low-frequency regions of the digital voice signal, the digital voice signal is first pre-emphasized, and the expression is as follows:

[0059] y(t) = x(t) - αx(t - 1)

[0060] where y(t) is the digital voice signal after pre-emphasis, t is the time node of the digital voice signal, x(t) represents the intensity of the digital voice signal at time t, x(t - 1) represents the intensity of the digital voice signal at time t - 1, and α is the pre-emphasis coefficient;

[0061] Step S22, frame the digital voice signal to obtain a plurality of short-time frames, and connect adjacent frames;

[0062] After pre-emphasis, the digital voice signal needs to be divided into short-time frames, and a good approximation of the signal frequency profile is obtained by connecting adjacent frames;

[0063] Step S23, multiply each short-time frame by a Hamming window function;

[0064] Multiply each short-time frame into which the digital voice signal is divided by a Hamming window function to increase the continuity at the left and right ends of the window. The formula for the Hamming window function is:

[0065]

[0066] where w(n) is the value of the window function in the discrete time series, n is the index in the frame sequence, N is the window length, and a0 is a constant coefficient used to adjust the shape of the window function;

[0067] Step S24: Perform a Fourier transform on each of the short-time frames to obtain the spectrum of the digital voice signal;

[0068] Use the Fourier transform to convert the digital voice signal in the time domain into the energy distribution in the frequency domain. Perform an N-point Fourier transform on each short-time frame signal after frame division and windowing to obtain the spectrum of the digital voice signal. The calculation formula is:

[0069]

[0070] where, S i (k) represents the complex function of the digital voice signal in the frequency domain, s i (n) is the function of the digital voice signal in the time domain, n is the index in the frame sequence, N is the window length, N takes 512 in this embodiment, and k is the frequency variable, representing the frequency of the digital voice signal in the frequency domain;

[0071] Step S25: Take the modulus square of the spectrum of the digital voice signal to obtain the power energy spectrum;

[0072] Take the modulus square of the spectrum of the digital voice signal to obtain the power energy spectrum of the digital voice signal. The calculation formula is:

[0073]

[0074] where P is the power energy spectrum of the digital voice signal, S i (k) represents the complex function of the digital voice signal in the frequency domain, and N is the window length;

[0075] Step S26: Extract the band features of the power energy spectrum through a group of linear scale filters, and apply the discrete cosine transform to retain the cepstral coefficients to obtain the linear frequency cepstral coefficients of the digital voice signal;

[0076] Extract the band features of the power energy spectrum through a group of linear scale filters, and apply the discrete cosine transform to retain the obtained 13 cepstral coefficients to obtain the linear frequency cepstral coefficients of the digital voice signal. The formula for the discrete cosine transform is:

[0077]

[0078] Among them, C(m) represents the m-th order linear cepstral coefficients of the digital speech signal, M represents the order of the linear cepstral coefficients; s(n) represents the frequency band characteristics of the short-time frame, n is the index in the frame sequence, N is the window length, and L is the number of linear scale filters.

[0079] Further, the step S24 is specifically: when performing N-point Fourier transform on each short-time frame, if the signal length of the short-time frame is insufficient, it is padded with zeros, and N is taken as 512.

[0080] Further, in the step S21, the pre-emphasis coefficient for pre-emphasizing the digital speech signal is taken as 0.95 or 0.97.

[0081] Further, as Figure 2 shown, the step S3 is specifically:

[0082] Step S31: Build a detection model for voice spoofing attack. The detection model includes a convolutional layer, a Res2Net module, an attention module, a fully connected layer, and an activation function. The activation function is a Softmax activation function;

[0083] Step S32: Input the linear cepstral coefficients of the digital signal into the detection model, and adjust the parameters of the detection model through backpropagation until a preset end condition is reached.

[0084] As Figure 3 shown, in this embodiment, the detection model is composed of a convolutional layer Conv_1, four Res2Net modules, four attention modules, a global average pooling layer, a fully connected layer, and a Softmax activation function; input the linear cepstral coefficients of the voice into the voice spoofing attack detection model, and adjust the weights and biases of the model through backpropagation to realize the training of the detection model; repeat the training until the training end condition is reached.

[0085] Further, as Figure 3 shown, the step S32 is specifically:

[0086] Step S321: Input the linear cepstral coefficient matrix with the dimension of C×H×W into the convolutional layer Conv_1 of the detection model, perform sliding window processing, and extract initial feature information, where C, H, and W are the channels, height, and width of the linear cepstral coefficient matrix respectively;

[0087] Step S322: Input the initial feature information into the Res2Net module and the attention module of the detection model, and repeat the Res2Net module and the attention module until the initial information features input to the model pass through all the Res2Net modules and attention modules, obtaining the final feature matrix Z;

[0088] Step S323: Obtain a feature description vector by global average pooling of the feature matrix Z, perform a flattening operation on the feature description vector on each channel, and input the flattened feature description vector into the fully connected layer of the detection model. The fully connected layer consists of multiple neurons, each neuron is connected to an element in the feature vector, and operations of weight multiplication and weighted summation are performed; after the output of the fully connected layer passes through an activation function, a final prediction vector is obtained; the prediction vector represents the distribution probabilities of real speech and spoofed speech;

[0089] Step S324: According to the data and labels of the prediction vector, use the binary cross-entropy loss function to backpropagate and calculate the gradient, and adjust the weights and biases of the detection model;

[0090] The expression of the binary cross-entropy loss function is:

[0091] Loss=-(k·log(p)+(1-k)·log(1-p))

[0092] where Loss represents the binary cross-entropy loss function, k is the true label value; p is the prediction result of the model, representing the probability of belonging to the positive class.

[0093] Further, the Res2Net module obtains multiple groups of feature sub-matrices based on the input signal, applies a convolution operation to each of the feature sub-matrices, extracts the features of the feature sub-matrices, adds the input features of the feature sub-matrices to the features after the convolution operation to obtain residual features; the residual features are passed to the next layer for convolution operation, summation operation, channel connection, and channel number adjustment to obtain the output of the Res2Net module;

[0094] The attention module uses global average pooling to extract the position information of the input signal in the vertical and horizontal directions respectively, calculates the attention weights in the vertical and horizontal directions respectively according to the position information in the vertical and horizontal directions to obtain the output of the attention module;

[0095] Input the initial feature information into multiple of the Res2Net modules and multiple of the attention modules in sequence to obtain the feature matrix.

[0096] Specifically, as Figure 3As shown, the initial feature information x extracted from the previous layer is input into the Res2Net module, and x is divided into s groups of sub-matrices x1, x2..., x s , in this embodiment, it is divided into 4 groups of sub-matrices x1, x2, x3, x4. Each sub-matrix has the same height and width, but the number of channels may be different; then a convolution operation Γ is applied to each sub-matrix i , and the features of the sub-matrix are extracted through a series of convolutional layers; these convolutional layers have different depths and convolutional kernel sizes to capture features of different scales; inside each sub-matrix, residual connections are used to transmit information; the input features of the sub-matrix are added to the features obtained after the convolution operation to obtain residual features; then the residual features are passed to the next layer of convolution operation or other processing; the calculation formula is expressed as:

[0097]

[0098] where, where y i is the output of the i-th group of convolution operations Γ i (), for each group of sub-matrices except x1, a convolution with a convolutional kernel size of 3 is performed, denoted as Γ i (); except for x1 and x2, first, the i-th group of x i is added to the output Γ i-1 () of the previous group, and the operation of Γ i () is the summation result; finally, the sub-outputs of the s groups are concatenated in the channel dimension, and then the number of channels of the features is adjusted through convolution.

[0099] As Figure 3 shown, the output y of the Res2Net module is input into the corresponding attention module. First, global average pooling is used to extract the coordinate information in the vertical and horizontal directions respectively; for the input y with dimensions C×H×W, the outputs in the two spatial directions of channel C are expressed as:

[0100]

[0101] where, are the outputs of the attention module in the h and w dimensions respectively, and y c (h,i), y c (j,w) are the coordinate information in the vertical and horizontal directions extracted by global average pooling respectively.

[0102] Next, the weights are calculated respectively from the obtained vertical and horizontal position information, and the formula is expressed as:

[0103]

[0104] where, h and w represent the vertical dimension and the horizontal dimension respectively, and k h, k w are the attention weights in the h and w dimensions respectively, σ() is the sigmoid activation function that maps variables to the range [0, 1] to extract weights, and φ h , φ w are 1×1 convolutions, and f h , f w are the intermediate feature maps of the position information in the h and w dimensions respectively, expressed as f = Ψ(F({z h , z w})), where in this formula, F is a convolution operation with a convolution size of 1×1, {} represents the concatenation operation, and Ψ is a non-linear activation function;

[0105] Repeat the operations of the above Res2Net module and attention module until the initial information features of the input pass through all the Res2Net modules and attention modules to obtain the final feature matrix Z.

[0106] Most existing methods choose the channel attention mechanism. Since these attention mechanisms only consider the attention in the channel dimension, they cannot capture the attention in the spatial dimension. The coordinate attention mechanism adopted in this embodiment can simultaneously focus on the channel information and position information of the speech, which is beneficial to extracting the differences between spoofed speech and genuine speech, thereby accurately distinguishing synthetic speech and replayed speech. The present invention also provides a corresponding speech spoofing attack detection device based on coordinate attention. The coordinate attention mechanism adopted in this embodiment has relatively high computational efficiency. Compared with the traditional attention mechanism, it does not require global attention calculation, but only focuses on the channel information and the position information of the features. Coordinate attention reduces the computational complexity and improves the running speed of the model.

[0107] Further, step S4 is specifically as follows:

[0108] Step S41: Load the speech to be detected in the database, extract the linear frequency cepstral coefficients of the speech to be detected, and input them into the trained detection model;

[0109] Step S42: Determine whether the speech to be detected is spoofed speech according to the prediction vector output by the detection model, and output the detection result.

[0110] Further, step S41 is specifically as follows:

[0111] Input the linear frequency cepstral coefficients of the speech to be detected into the convolutional layer, Res2Net module, attention module, fully connected layer, and activation function of the detection model in sequence to obtain the prediction vector of the speech to be detected.

[0112] Specifically, the linear frequency cepstral coefficients of the speech to be detected are processed through the convolutional layer Conv_1 with a sliding window to extract the initial feature information in the input;

[0113] The initial feature information x extracted in the previous layer is input into the Res2Net module, and x is divided into s groups of sub-matrices x1, x2..., x s , each sub-matrix having the same height and width, but possibly different numbers of channels; then a convolution operation is applied to each sub-matrix, and the features of the sub-matrix are extracted through a series of convolutional layers; these convolutional layers have different depths and convolutional kernel sizes to capture features at different scales; within each sub-matrix, residual connections are used to transmit information; the input features of the sub-matrix are added to the features obtained after the convolution operation to obtain residual features; then, the residual features are passed to the next layer of convolution operation or other processing; the calculation formula is expressed as:

[0114]

[0115] where, where y i is the output of the i-th group of convolution operations Γ i (), except for x1, each group of sub-matrices performs a convolution with a convolutional kernel size of 3, denoted as Γ i (), except for x1 and x2, first the i-th group of x i is added to the output Γ i-1 (), and the operation of Γ i () is the summation result; finally, the sub-outputs of the s groups are concatenated in the channel dimension, and then the number of channels of the features is adjusted through convolution;

[0116] For the output y of the Res2Net module, first use global average pooling to extract the coordinate information in the vertical and horizontal directions respectively; given an input y with dimensions C×H×W, the outputs of the channel C in the two spatial directions are expressed as:

[0117]

[0118] Next, calculate the weights from the obtained vertical and horizontal position information respectively; the formula is expressed as:

[0119]

[0120] where, h and w represent the vertical and horizontal dimensions, k h , k w are the attention weights in the h-dimensional and w-dimensional respectively, where φ is a 1×1 convolution, f is the intermediate feature map of the position information, expressed as f = Ψ(F({z h , z w})) where F is a convolution operation with a convolution size of 1×1, {} represents a concatenation operation, and Ψ is a non-linear activation function;

[0121] Repeat the operations of the Res2Net module and the attention module until the initial information features of the model input pass through all the Res2Net modules and attention modules to obtain the final feature matrix Z;

[0122] The feature matrix Z is passed through global average pooling to obtain a feature description vector. Then, the feature description vector composed of the average values on each channel is flattened. Finally, this flattened feature description vector is input into a fully connected layer for processing; the fully connected layer consists of multiple neurons, each neuron is connected to an element in the feature vector, and performs operations of weight multiplication and weighted summation; after the output of the fully connected layer passes through the Softmax activation function, the final prediction vector is obtained, representing the probability distributions of real speech and spoofed speech.

[0123] The voice spoofing attack detection method provided in this embodiment aims to accurately and quickly detect voice spoofing attacks during the voice recognition process. This voice attack detection method judges the authenticity of voice through a deep learning model, uses linear frequency cepstral coefficients as the acoustic features of voice, and guides the attention area of the model based on the coordinate attention mechanism to improve the detection accuracy of the model, thereby solving the technical problem of voice spoofing attack detection.

[0124] Figure 4 This is an example of the visualization result of the voice spoofing attack detection method in this embodiment on the ASVspoof2019 dataset. In this figure, the vertical axis represents the spoofed speech recognition accuracy of the detection model, and the horizontal axis represents different tracks and sets of the dataset. For the voice spoofing attack detection model in this embodiment, CA-Res2Net in the figure represents the detection model in this embodiment, and its recognition accuracy on each track of the ASVspoof2019 dataset is better than the other two baseline models. Res2Net and SE-Res2Net in the figure respectively represent the other two baseline models.

[0125] Embodiment 2

[0126] This embodiment provides a voice spoofing attack detection device, including a memory and a processor. A computer program is stored on the memory, and when the computer program is executed by the processor, it implements the voice spoofing attack detection method described in Embodiment 1.

[0127] The voice spoofing attack detection device provided in this embodiment is used to implement the voice spoofing attack detection method. Therefore, the technical effects possessed by the voice spoofing attack detection method are also possessed by the voice spoofing attack detection device, which will not be elaborated here.

[0128] As described above, it is only the preferred specific implementation manner of the present invention, but the protection scope of the present invention is not limited thereto. Any changes or substitutions that can be easily thought of by those skilled in the art within the technical scope disclosed by the present invention should be covered by the present invention.

Claims

1. A method for detecting voice spoofing attacks, characterized in that, It includes the following steps: Step S1: Load the speech data in the speech training set, sample the speech data at a preset sampling rate, and convert the speech data into a digital speech signal; Step S2: Input the digital speech signal into a linear filter to extract the linear frequency cepstral coefficients that describe the frequency domain characteristics of the digital speech signal; Step S3: Build a detection model for speech spoofing attacks, and use the linear frequency cepstral coefficients of the digital speech signal to train the detection model; Step S4: Input the speech to be detected into the trained detection model, and determine whether the speech to be detected is spoofed speech according to the output of the detection model; The specific content of step S3 is as follows: Step S31: Build a detection model for speech spoofing attacks. The detection model includes a convolutional layer, a Res2Net module, an attention module, a fully connected layer, and an activation function; Step S32: Input the linear frequency cepstral coefficients of the digital signal into the detection model, and adjust the parameters of the detection model through backpropagation until a preset end condition is reached; The specific content of step S32 is as follows: Step S321: Input the linear frequency cepstral coefficients into the convolutional layer of the detection model for sliding window processing to extract initial feature information; Step S322: Input the initial feature information into the Res2Net module and the attention module of the detection model to obtain a feature matrix; Step S323: Obtain a feature description vector by global average pooling of the feature matrix, perform a flattening operation on the feature description vectors on each channel, input the flattened feature description vectors into the fully connected layer of the detection model, and after the output of the fully connected layer passes through the activation function, obtain the final prediction vector; Step S324: According to the data and labels of the prediction vector, use the binary cross-entropy loss function to backpropagate and calculate the gradient, and adjust the weights and biases of the detection model; The attention module of the detection model uses global average pooling to extract the position information of the input signal in the vertical and horizontal directions respectively, calculates the attention weights in the vertical and horizontal directions respectively according to the position information in the vertical and horizontal directions, and obtains the output of the attention module; Input the initial feature information extracted in the previous layer into the Res2Net module; Input the output of the Res2Net module into the corresponding attention module; repeat the operations of the Res2Net module and the attention module above until the input initial information features pass through all the Res2Net modules and attention modules to obtain the final feature matrix.

2. The voice spoofing attack detection method according to claim 1, wherein The specific content of step S2 is as follows: Step S21: Perform pre-emphasis on the digital speech signal; Step S22: Frame the digital speech signal to obtain multiple short-time frames, and connect adjacent frames; Step S23: Multiply each short-time frame by a Hamming window function; Step S24: Perform Fourier transform on each short-time frame to obtain the spectrum of the digital speech signal; Step S25: Take the modulus square of the spectrum of the digital speech signal to obtain the power energy spectrum; Step S26: Extract the band features of the power energy spectrum through a set of linear scale filters, apply the discrete cosine transform to retain the cepstrum coefficients, and obtain the linear frequency cepstrum coefficients of the digital speech signal.

3. The voice spoofing attack detection method according to claim 2, wherein The specific operation of step S24 is as follows: When performing the N-point Fourier transform on each short-time frame, if the signal length of the short-time frame is insufficient, it is padded with zeros.

4. The voice spoofing attack detection method according to claim 2, wherein, In step S21, the pre-emphasis coefficient for pre-emphasizing the digital speech signal is taken as 0.95 or 0.

97.

5. The voice spoofing attack detection method according to claim 1, wherein The specific operation of step S4 is as follows: Step S41: Load the speech to be detected in the database, extract the linear frequency cepstrum coefficients of the speech to be detected, and input them into the trained detection model. Step S42: Judge whether the speech to be detected is spoofed speech according to the prediction vector output by the detection model, and output the detection result.

6. The voice spoofing attack detection method according to claim 5, characterized in that, The specific operation of step S41 is as follows: Input the linear frequency cepstrum coefficients of the speech to be detected into the convolutional layer, Res2Net module, attention module, fully connected layer, and activation function of the detection model in sequence to obtain the prediction vector of the speech to be detected.

7. A voice spoofing attack detection device, characterized in that, It includes a memory and a processor. A computer program is stored on the memory. When the computer program is executed by the processor, it implements the speech spoofing attack detection method according to any one of claims 1-6.

Citation Information

Patent Citations

  • Intelligent voice forgery attack detection method based on attention mechanism

    CN116416997A

  • Voiceprint recognition method and device for complex scene, electronic equipment and storage medium

    CN116758921A