Counterfeit voice detection method fusing multi-scale fundamental frequency features and enhancing attention

By integrating multi-scale fundamental frequency features and enhancing attention, the problem of insufficient feature extraction and modeling capabilities in deep forged speech detection is solved, and more efficient forged speech detection is achieved, which improves the performance and generalization of the model.

CN120340528APending Publication Date: 2025-07-18NANJING UNIV OF POSTS & TELECOMM +1
View PDF 0 Cites 2 Cited by

Patent Information

Application Number
CN202510665625.8
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-05-22
Publication Date
2025-07-18

AI Technical Summary

Technical Problem

The existing deep fake speech detection methods have shortcomings in feature extraction and modeling capabilities, especially the lack of global dependency modeling capabilities of convolutional neural networks, which leads to poor generalization of the model and it is difficult to effectively detect various fake speeches.

Method used

Fusion of multi-scale fundamental frequency features and enhancing attention, extract global and local information of features in the deep extraction module through global and local enhancement of attention, and fuses the two features using fine-grained fusion module to improve detection performance.

Benefits of technology

The performance and generalization capabilities of the deep forged speech detection model are improved, and multiple forged speech can be more effectively recognized, which enhances the detection ability of speech tone and emotional expression.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120340528A_ABST
    Figure CN120340528A_ABST
Patent Text Reader

Abstract

The method comprises the following steps: obtaining training data of a voice data construction model, preprocessing the obtained training data to obtain original waveform features and multi-scale fundamental frequency features, inputting the features into the model to carry out model training, and obtaining an attention enhancement model; in the training process of the detection model, hyper-parameters of the detection model are set, a loss function is continuously reduced, the trained detection model is obtained when the set training frequency is reached, test data of the voice data construction model are obtained, and preprocessing is carried out to obtain original waveform features and multi-scale fundamental frequency features; the original waveform features and the multi-scale fundamental frequency features are input into a trained model for model testing, in the testing process, a voice authenticity classification result is obtained, the performance of the model is evaluated, and the performance of the detection model can be effectively improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention belongs to the technical field of deepfake voice detection, and specifically relates to a deepfake voice detection method that fuses multi-scale fundamental frequency features and enhanced attention. Background Art

[0002] With the continuous development of deep learning technology, the voice quality generated by voice synthesis and voice conversion algorithms is getting higher and higher, which can deceive the ASV (Automatic Speaker Verification) system and achieve the effect of being indistinguishable from the real one. The ASV system is vulnerable to attacks by forged voices. To solve this problem, deepfake voice detection technology has emerged and become a popular research direction.

[0003] The construction of a deepfake voice detection system mainly depends on two aspects: one is the input features, and the other is the deep extraction module. In terms of input features, early research mainly focused on manually designed features such as CQCC (Constant QCepstral Coefficients) features and LFCC (Linear Frequency Cepstral Coefficients) features, which have been proven to be able to effectively capture clues in forged voices; later, some scholars proposed to directly learn general features from the original voice waveform for forged voice detection, breaking through the limitations of manually designed features; in recent years, models using the wav2vec 2.0 pre-trained model as the feature extraction front-end have achieved the best performance and generalization.

[0004] In terms of the deep extraction module, most early models used ResNet (Residual Network) and improved Res2Net. The residual network can effectively extract useful information from the input features; to further improve the feature extraction ability, networks based on the attention mechanism were introduced, such as SENet (tSqueeze-Excitation Network) and GAT (Graph Attention Networks), and good results were obtained; based on the graph attention network, the spectral-temporal graph attention network captures clues in the spectral and time domains from the input features, greatly improving the feature extraction ability; some scholars also used the Transformer model and position encoding in the field of natural language processing to model the long-term dependencies in the input features.

[0005] Currently, most methods use CNN (Convolutional Neural Network) to construct a deep extraction module, which can effectively model local dependencies. However, due to the limitation of the convolutional receptive field, CNN lacks the ability to extract global features and requires a sufficiently deep network structure to model long-term dependencies. When the input is sequence data, modeling long-term dependencies is particularly important. Only using CNN will result in insufficient feature extraction ability of the model. In addition, most methods only use one type of feature as input, which will lead to insufficient generalization ability of the model. Because the information contained in one type of feature is limited, it is difficult for the model to effectively detect all types of forged speech, and there are always some forged speeches that are difficult to detect. Summary of the Invention

[0006] The problem to be solved by the present invention is to provide a method for detecting deepfake speech that fuses multi-scale fundamental frequency features and enhanced attention. Aiming at the problem of insufficient original waveform feature information, multi-scale fundamental frequency features are input to provide rich prosodic information; aiming at the problem that the convolutional neural network lacks the ability to model global dependencies, global enhanced attention is used in the deep extraction modules of both the original waveform features and multi-scale fundamental frequency features to extract the global information of the features and model global dependencies; further, in order to extract more prosodic information from the multi-scale fundamental frequency features, local enhanced attention is used in the deep extraction module of the multi-scale fundamental frequency features to extract the local detailed information of the features; finally, a fine-grained fusion module is used to fully fuse the two types of features that have passed through the deep extraction module to achieve effective deepfake speech detection.

[0007] The present invention is a method for detecting deepfake speech that fuses multi-scale fundamental frequency features and enhanced attention, including a training stage and a testing stage;

[0008] The training stage includes the following steps:

[0009] Step S1, obtain speech data to construct the training data of the model, and the training data includes real speech and forged speech generated by speech synthesis and speech conversion algorithms;

[0010] Step S2, perform different preprocessings on the obtained training data to respectively obtain the original waveform feature X Raw and the multi-scale fundamental frequency feature X F0 ;

[0011] Step S3, input the original waveform feature and the multi-scale fundamental frequency feature into the model respectively for model training;

[0012] Step S4, during the training process of the detection model, set the hyperparameters of the detection model to make the loss function continuously decrease, and obtain the trained detection model when the set number of training times is reached;

[0013] The testing phase includes the following steps:

[0014] Step S5: Obtain the test data for building the model from the voice data. The test data includes real voices and forged voices generated by voice synthesis and voice conversion algorithms. Preprocess them separately to obtain the original waveform features X Raw and multi-scale fundamental frequency features X F0 ;

[0015] Step S6: Input the original waveform features and multi-scale fundamental frequency features into the trained model for model testing respectively. During the testing process, obtain the voice authenticity classification results and evaluate the performance of the model.

[0016] Preferably, in step S2, different preprocessings are performed on the obtained training data. The training data is sampled at a sampling rate of 16 kHz to obtain the original waveform features. The fundamental frequency envelope is extracted from the training data and the multi-scale fundamental frequency features are obtained through continuous wavelet transform, as shown in formulas (1) and (2).

[0017]

[0018] Among them, W(F0)(τ, t) represents the fundamental frequency spectrum, F0(x) represents the fundamental frequency value at position x, ψ represents the wavelet mother function, τ and t represent the scale and position of the wavelet respectively. When τ0 = 5 ms, let i = 1,..., 10. The fundamental frequency envelopes of a total of 10 scales form the multi-scale fundamental frequency feature X F0 .

[0019] Preferably, in step S3, the model consists of a Sinc module, a deep extraction module for original waveform features, a deep extraction module for multi-scale fundamental frequency features, a fine-grained fusion module, and a classification module:

[0020] The Sinc module performs shallow feature extraction on the input original waveform features; the deep extraction module for original waveform features further extracts the features extracted by the Sinc module to obtain deep features. The deep extraction module for multi-scale fundamental frequency features extracts the input multi-scale fundamental frequency features to obtain deep features. The fine-grained fusion module fuses the deep features output by the two deep extraction modules to obtain fusion features. The classification module processes the input fusion features to obtain classification results.

[0021] Preferably, the Sinc module consists of 1 one-dimensional convolutional layer, 1 max-pooling layer, 1 batch normalization layer, and 1 SELU activation function.

[0022] The deep extraction module for original waveform features consists of 6 one-dimensional convolutional blocks and 6 global enhanced attention mechanisms; the one-dimensional convolutional block consists of 2 one-dimensional convolutional layers, 1 batch normalization layer, 1 LeakyReLU activation function, and 1 max pooling layer, and the one-dimensional convolutional blocks and global enhanced attention mechanisms are placed alternately;

[0023] The global enhanced attention mechanism consists of 1 one-dimensional convolutional layer, 2 normalization layers, 5 linear layers, 1 linear attention mechanism, 2 SiLU activation functions, and 1 multi-layer perceptron;

[0024] The deep extraction module for multi-scale fundamental frequency features consists of 6 two-dimensional convolutional blocks, 1 local enhanced attention mechanism, and 1 global enhanced attention mechanism; the two-dimensional convolutional block consists of 1 two-dimensional convolutional layer, 1 batch normalization layer, and 1 ReLU activation function;

[0025] The local enhanced attention mechanism consists of 3 linear layers, 2 Dropout layers, and 1 average pooling layer;

[0026] The fine-grained fusion module consists of 1 spatial attention mechanism, 1 channel attention mechanism, 1 Sigmoid activation function, and 1 two-dimensional convolutional layer. The spatial attention mechanism consists of 1 two-dimensional convolutional layer, and the channel attention mechanism consists of 1 adaptive average pooling layer, 2 two-dimensional convolutional layers, and 1 ReLU activation function;

[0027] The classification module consists of 1 gated recurrent unit, 2 linear layers, 1 batch normalization layer, and 1 SELU activation function.

[0028] Preferably, the global enhanced attention mechanism is used to extract the global information of the feature vector, and the method is as follows:

[0029] First, the input feature X passes through the normalization layer to obtain the normalized feature X norm , which is expressed as follows:

[0030] X norm = Norm(X) (3),

[0031] where Norm(·) represents the normalization operation. Then, X norm successively passes through the linear layer, convolutional layer, and SiLU activation function to highlight the useful information and obtain the input feature X in of the linear attention mechanism, which is expressed as follows:

[0032] X in = SiLU(Conv(Linear(X norm ))) (4),

[0033] Among them, Linear(·) represents a linear layer, Conv(·) represents a stack of convolutional layers, SiLU(·) represents the SiLU activation function. Linear attention processes the input as a sequence of N features with the same number of channels C, so as to model the relationship between each feature in the sequence. According to the attention mechanism, first, the query vector the key vector and the value vector are calculated respectively. The calculation formulas are as follows:

[0034] Q = φ(X in W Q ) (5),

[0035] K = φ(X in W K ) (6),

[0036] V = X in W V (7),

[0037] where and represent the weight vectors corresponding to three linear layers; C represents the number of channels of the vector, d represents the dimension of the vector, N represents the length of the vector, φ(·) represents the kernel function operation. Then, the query vector Q and the key vector K are subjected to a dot product operation to obtain the attention weights. Finally, the value vector V is weighted according to the attention weights to highlight the important information in the input features, and the output X att of the linear attention is obtained, which is expressed as:

[0038]

[0039] where is a vector of all 1s, which is used to normalize the attention weights. Finally, the output Y of the global enhanced attention is obtained through a linear layer, a normalization layer, and a multi-layer perceptron, as follows:

[0040] Y = X + Linear(X att X in ) + Mlp(Norm(X + Linear(X att X in ))) (9),

[0041] where Mlp(·) represents a multi-layer perceptron. The global enhanced attention has a global receptive field and aggregates context information, so it can model global dependencies and extract more global information.

[0042] Preferably, local enhanced attention is used to extract the local detailed information of the feature vector, and the method is as follows:

[0043] Local enhanced attention focuses on the input features at each spatial position (i, j), and calculates the similarity with all adjacent spatial positions within a local window of size K×K centered at (i, j). First, according to the attention mechanism, the attention weight vector and the value weight vector are calculated respectively, and the calculation formulas are as follows:

[0044] A = XW A (10),

[0045] V = XW V (11),

[0046] where and represent the weight vectors corresponding to two linear layers; H and W represent the height and width of the vector, C represents the number of channels of the vector, and K represents the side length of the local window; calculate all the values within the local window centered at (i, j) which is specifically expressed as:

[0047]

[0048] where {·} represents a set, represents rounding up;

[0049] Reshape the attention weight vector at position (i, j) into and obtain the aggregated value attention weights through the Softmax function. Then, multiply the value attention weights at position (i, j) by all the values V i,j within the window to obtain the weighted output Y i,j at position (i, j), which can be expressed as:

[0050] Y i,j = MatMul(Softmax(A i,j ), V i,j ) (13),

[0051] where MatMul(·) represents the matrix multiplication operation, and Softmax(·) represents the Softmax function;

[0052] Then, add the different weighted outputs Y i,j from the same position (i, j) of different K×K local windows to obtain the local enhanced attention output at position (i, j), which is expressed as:

[0053]

[0054] Finally, replace the value X of all positions (i, j) of the input feature X i,j with the output of the local enhanced attention at that position to obtain the output Y of the local enhanced attention;

[0055] The local enhanced attention aggregates the detailed information of each local window, thereby capturing the local information in the input feature.

[0056] Preferably, the fine-grained fusion module is used to fuse the extracted feature vectors, and the method is as follows:

[0057] First, additively fuse the input feature and to obtain the fused input feature which is expressed as:

[0058] X = X1 + X2 (15),

[0059] where H and W represent the height and width of the feature vector, and C represents the number of channels of the feature vector;

[0060] To obtain an attention weight vector with the same dimension and channels as the fused input feature calculate the corresponding channel attention weight vector W c and spatial attention weight vector W s respectively through channel attention and spatial attention. The formulas are as follows:

[0061]

[0062] where represents a convolutional layer with a kernel size of k×k, max(0, x) represents the ReLU activation function, [·] represents the channel-level concatenation operation, GAP c (·) represents the global average pooling operation across the spatial dimension, GAP s (·) represents the global average pooling operation across the channel dimension, GMP s (·) represents the feature processed by the global max pooling operation across the channel dimension. In channel attention, to reduce the number of model parameters, the first 1×1 convolution reduces the channel dimension from C to (r is set to ), and the second 1×1 convolution expands its channel dimension back to C;

[0063] Then, use the addition operation to fuse the channel attention weight vector W c and the spatial attention weight vector W s together to obtain the coarse-grained attention weight vector which is specifically expressed as:

[0064] W coa = W c + W s (18),

[0065] Next, adjust the coarse-grained attention weight vector W according to the input feature X coa for each channel, and use the channel shuffle operation to rearrange each channel of X and W coa in an alternating manner to obtain the fine-grained attention weight vector denoted as:

[0066]

[0067] where σ represents the Sigmoid operation, represents a grouped convolutional layer with a kernel size of k×k, CS(·) represents the channel shuffle operation, [·] represents the channel-level concatenation operation. Finally, according to the fine-grained attention weight W, the input features X1 and X2 are weighted and fused to obtain the fused output X Fuse , denoted as:

[0068]

[0069] The fine-grained fusion module calculates the attention weight for each channel, focuses on the important information in each channel, thus effectively fusing the two input features and retaining more useful information.

[0070] Preferably, the training process in step S4 is consistent with the testing process in step S6 and includes the following sub-steps:

[0071] Step S4-1, input the original waveform feature X Raw into the Sinc module, and after being processed by the Sinc function filter, obtain the waveform time-frequency feature X RawTF ;

[0072] Step S4-2, input the waveform time-frequency feature X RawTF into the depth extraction module of the original waveform feature, and successively pass through a network composed of 6 one-dimensional convolutional blocks and 6 global enhanced attention overlaps to obtain the deep feature vector X′ Raw ; input the multi-scale fundamental frequency feature X F0 into the depth extraction module of the multi-scale fundamental frequency feature, and successively pass through 6 two-dimensional convolutional blocks, 1 local enhanced attention, and 1 global enhanced attention to obtain the deep feature vector X′ F0 ;

[0073] Step S4-3, the deep feature vectors X′ Raw and X′ F0Input it into the fine-grained fusion module for feature fusion, and then input the fused feature X Fuse into the classification module to obtain the classification result output by the detection model.

[0074] Preferably, the loss function of the detection model in step S4, that is, the cross-entropy loss function, is expressed as follows:

[0075]

[0076] where y represents the true label, represents the predicted probability of the true label output by the detection model; log(·) represents the logarithmic function.

[0077] Compared with the prior art, the beneficial effects of the present invention are as follows:

[0078] 1) In the deepfake voice detection method of the present invention, the input of the model includes not only the original waveform features, but also multi-scale fundamental frequency features, and a fine-grained fusion module is used to fuse the two features after the deep extraction module. The information contained in the original waveform features is limited, and the detection ability for some types of forged voices is insufficient. The fundamental frequency is related to the pitch of the voice and also related to the emotional expression in the voice. Most voice conversion and synthesis algorithms are imperfect, resulting in unnatural pitch and lack of emotional expressiveness in the generated voice. Therefore, introducing multi-scale fundamental frequency features as the model input can compensate for the lack of information contained in the original waveform features. The fine-grained fusion module uses spatial attention and channel attention to calculate the attention weights, and fuses the two features after the deep extraction module by weighting, achieving the effect of information complementarity. In summary, using multi-scale fundamental frequency features as the model input and using the fine-grained fusion module for feature fusion can increase the effective information for forged voice detection, thereby improving the performance and generalization of the detection model;

[0079] 2) The deepfake voice detection method of the present invention uses global enhanced attention and local enhanced attention in the deep extraction module to enhance the feature extraction ability. The global enhanced attention has a global receptive field, aggregates context information, and can effectively model global dependencies. As the main component of the deep extraction module, the convolutional network can effectively model local dependencies, but lacks the ability to extract global features. When the input is sequence data, it is crucial to model long-term dependencies. Therefore, global enhanced attention is added to both the deep extraction modules of the original waveform features and the multi-scale fundamental frequency features to improve the global modeling ability. Then, the local enhanced attention focuses on the detailed information of each local position and encodes the detailed information. The multi-scale fundamental frequency features contain multi-scale fundamental frequency spectrograms, which contain rich detailed information. Therefore, local enhanced attention is further added to the deep extraction module of the multi-scale fundamental frequency features to improve the local modeling ability. In summary, using global enhanced attention and local enhanced attention can enhance the feature extraction ability, thereby effectively improving the performance of the detection model. BRIEF DESCRIPTION OF THE DRAWINGS

[0080] Figure 1 is a flowchart of the deepfake voice detection method of the present invention;

[0081] Figure 2 is a network structure diagram of the Sinc module described in the implementation of the present invention;

[0082] Figure 3 is a network structure diagram of the deep extraction module of the original waveform features described in the implementation of the present invention;

[0083] Figure 4 is a network structure diagram of the deep extraction module of the multi-scale fundamental frequency features described in the implementation of the present invention;

[0084] Figure 5 is a network structure diagram of the global enhanced attention described in the implementation of the present invention;

[0085] Figure 6 is a network structure diagram of the local enhanced attention described in the implementation of the present invention;

[0086] Figure 7 is a network structure diagram of the fine-grained fusion module described in the implementation of the present invention;

[0087] Figure 8 is a network structure diagram of the classification module described in the implementation of the present invention. DETAILED DESCRIPTION OF THE INVENTION

[0088] The following further clarifies the present invention in conjunction with the accompanying drawings and specific embodiments. It should be understood that the following specific embodiments are only used to illustrate the present invention and are not used to limit the scope of the present invention.

[0089] Example: In one embodiment of the present invention, a deepfake voice detection method that fuses multi-scale fundamental frequency features and enhanced attention includes a training phase and a testing phase. The training phase is used to obtain the detection network model and its parameters required for deepfake voice detection, while the testing phase is used to implement the detection of deepfake voices.

[0090] The present invention is a deepfake voice detection method that fuses multi-scale fundamental frequency features and enhanced attention, including a training phase and a testing phase;

[0091] As Figure 1 shown, the training phase includes the following steps:

[0092] Step S1: Obtain speech data to construct the training data of the model. The training data includes real voices and forged voices generated by speech synthesis and voice conversion algorithms;

[0093] In this example, the training data comes from the ASVspoof2019LA dataset, including real voices and 6 types of forged voices generated by speech synthesis and voice conversion algorithms; the duration of most voices is 1 to 8 seconds, there are 25,380 voices in the training data, and 24,986 voices in the validation data.

[0094] Step S2: Perform different preprocessings on the obtained training data to obtain the original waveform feature X Raw and the multi-scale fundamental frequency feature X F0 ;

[0095] In this instance, the training data is sampled at a sampling rate of 16 kHz, and the input time is set to 4 seconds (i.e., 64,000 samples) to obtain the original waveform feature X Raw ; Extract the fundamental frequency envelope from the training data and perform continuous wavelet transform to obtain the multi-scale fundamental frequency feature X F0 , with the length set to 1800, and the method is as follows:

[0096]

[0097] Among them, W(F0)(τ, t) represents the fundamental frequency spectrum, F0(x) represents the fundamental frequency value at position x, ψ represents the wavelet mother function, τ and t represent the scale and position of the wavelet respectively. When τ0 = 5 ms, let i = 1,..., 10, and the fundamental frequency envelopes of 10 scales in total form the multi-scale fundamental frequency feature X F0 .

[0098] Step S3: Input the original waveform feature and the multi-scale fundamental frequency feature into the model respectively for model training;

[0099] In this example, the model mainly consists of a depth extraction module, a feature fusion module, and a classification module. The original waveform features are first passed through the Sinc module to extract preliminary features, and then further deep features are obtained through the depth extraction module of the original waveform features; at the same time, the multi-scale fundamental frequency features are passed through the depth extraction module of the multi-scale fundamental frequency features to extract deep features; then, the deep features output by the two depth extraction modules are subjected to feature fusion through the fine-grained fusion module; finally, the classification result is obtained through the classification module.

[0100] Furthermore, the detection model includes five parts: the Sinc module, the depth extraction module of the original waveform features, the depth extraction module of the multi-scale fundamental frequency features, the fine-grained fusion module, and the classification module:

[0101] In the Sinc module, as Figure 2 shown, the original waveform features pass through a one-dimensional convolutional layer, a max-pooling layer, a batch normalization layer, and the SELU activation function to obtain the waveform time-frequency features.

[0102] In the depth extraction module of the original waveform features, the waveform time-frequency features pass through the global enhancement attention and the one-dimensional convolutional block in sequence. The structure of the depth extraction module of the original waveform features is as Figure 3 shown, including 6 one-dimensional convolutional blocks and 6 global enhancement attentions, which are overlapped in sequence in the order of global enhancement attention first and one-dimensional convolutional block second. Among them, the one-dimensional convolutional block consists of 2 one-dimensional convolutional layers, 1 batch normalization layer, 1 LeakyReLU activation function, and 1 max-pooling layer.

[0103] In the global enhancement attention, the global information of the feature vector is extracted.

[0104] The network structure of the global enhancement attention is as Figure 5 shown, consisting of 1 one-dimensional convolutional layer, 2 batch normalization layers, 5 linear layers, 1 linear attention, 2 SiLu activation functions, and 1 multi-layer perceptron.

[0105] First, the input feature X passes through the normalization layer to obtain the normalized feature X norm , which is expressed as follows:

[0106] X norm = Norm(X) (3),

[0107] where Norm(·) represents the normalization operation. Then, X norm successively passes through the linear layer, the convolutional layer, and the SiLu activation function to highlight the useful information and obtain the input feature X in of the linear attention, which is expressed as follows:

[0108] X in= SiLU(Conv(Linear(X norm )) (4),

[0109] where Linear(·) represents a linear layer, Conv(·) represents a stack of convolutional layers, SiLU(·) represents the SiLU activation function, and linear attention processes the input as a sequence of N features with the same number of channels C, thereby modeling the relationships between each feature in the sequence. According to the attention mechanism, first, the query vector key vector and value vector are calculated separately. The calculation formulas are as follows:

[0110] Q = φ(X in W Q ) (5),

[0111] K = φ(X in W K ) (6),

[0112] V = X in W V (7),

[0113] where and represent the weight vectors corresponding to three linear layers; C represents the number of channels of the vector, d represents the dimension of the vector, N represents the length of the vector, φ(·) represents the kernel function operation. Then, the query vector Q and the key vector K are dot - producted to obtain the attention weights. Finally, the value vector V is weighted according to the attention weights to highlight the important information in the input features, and the output X att of the linear attention is obtained, which is expressed as:

[0114]

[0115] where is a vector of all 1s, used to normalize the attention weights. Finally, the output Y of the global enhanced attention is obtained through a linear layer, a normalization layer, and a multi - layer perceptron, as follows:

[0116] Y = X + Linear(X att X in ) + Mlp(Norm(X + Linear(X att X in ))) (9),

[0117] where Mlp(·) represents the multi - layer perceptron. The global enhanced attention has a global receptive field and aggregates context information, so it can model global dependencies and extract more global information.

[0118] In the deep extraction module of multi-scale fundamental frequency features, the multi-scale fundamental frequency features pass through a two-dimensional convolutional block, a global enhancement attention, and a local enhancement attention in sequence. The structure of the deep extraction module of multi-scale fundamental frequency features is as Figure 4 shown, including 6 two-dimensional convolutional blocks, 1 local enhancement attention, and 1 global enhancement attention. Among them, the two-dimensional convolutional block consists of 1 two-dimensional convolutional layer, 1 batch normalization layer, and 1 ReLU activation function.

[0119] Furthermore, in the local enhancement attention, the local detailed information of the feature vector is extracted.

[0120] The network structure of the local enhancement attention is as Figure 6 shown, consisting of 3 linear layers, 2 Dropout layers, and 1 average pooling layer.

[0121] The local enhancement attention focuses on each spatial position (i, j) of the input feature and calculates the similarity with all adjacent spatial positions within a local window of size K×K centered at (i, j). First, according to the attention mechanism, the attention weight vector and the value weight vector are calculated respectively. The calculation formulas are as follows:

[0122] A = XW A (10),

[0123] V = XW V (11),

[0124] where and represent the weight vectors corresponding to the two linear layers; H and W represent the height and width of the vector, C represents the number of channels of the vector, and K represents the side length of the local window; calculate all the values within the local window centered at (i, j), which is specifically expressed as:

[0125]

[0126] where {·} represents a set, represents rounding up;

[0127] Reshape the attention weight vector at position (i, j) into and obtain the aggregated value attention weight through the Softmax function. Then, multiply the value attention weight at position (i, j) by all the values V i,j within the window to obtain the weighted output Y i,j at position (i, j), which can be expressed as:

[0128] Y i,j = MatMul(Softmax(A i,j ), V i,j ) (13),

[0129] where MatMul(·) represents the matrix multiplication operation and Softmax(·) represents the Softmax function;

[0130] Then, the different weighted outputs Y at the same position (i, j) from different K×K local windows i,j are added together to obtain the local enhanced attention output at position (i, j) which is expressed as:

[0131]

[0132] Finally, the values X of all positions (i, j) of the input feature X i,j are replaced with the local enhanced attention output at that position to obtain the output Y of the local enhanced attention;

[0133] The local enhanced attention aggregates the detailed information of each local window, thereby capturing the local information in the input feature.

[0134] Furthermore, in the fine-grained fusion module, the two input features are first fused by addition, then the attention weights are calculated through spatial attention and channel attention, and finally the two input features are fused with weights.

[0135] The fine-grained fusion module has a structure as Figure 7 shown, and is composed of 1 spatial attention, 1 channel attention, 1 Sigmoid activation function, and 1 two-dimensional convolutional layer; the spatial attention is composed of 1 two-dimensional convolutional layer; the channel attention is composed of 1 adaptive average pooling layer, 2 two-dimensional convolutional layers, and 1 ReLU activation function.

[0136] First, the input features and are fused by addition to obtain the fused input feature which is expressed as follows:

[0137] X = X1 + X2 (15),

[0138] where H and W represent the height and width of the feature vector, and C represents the number of channels of the feature vector;

[0139] To obtain the attention weight vector with the same dimension and channels as the fused input feature Calculate the corresponding channel attention weight vector \(W_c\) and spatial attention weight vector \(W_s\) through channel attention and spatial attention respectively. c and the spatial attention weight vector \(W_s\) s , the formulas are as follows:

[0140]

[0141] where represents a convolutional layer with a kernel size of \(k\times k\), \(\max(0, x)\) represents the ReLU activation function, \([\cdot]\) represents the channel-level concatenation operation, and GAP c (\(\cdot\)) represents the global average pooling operation across the spatial dimensions, and GAP s (\(\cdot\)) represents the global average pooling operation across the channel dimensions, and GMP s (\(\cdot\)) represents the features processed by the global max pooling operation across the channel dimensions. In channel attention, to reduce the number of model parameters, the first \(1\times1\) convolution reduces the channel dimension from \(C\) to (\(r\) is set to ), and the second \(1\times1\) convolution expands its channel dimension back to \(C\);

[0142] Then, use the addition operation to fuse the channel attention weight vector \(W_c\) c and the spatial attention weight vector \(W_s\) s together to obtain the coarse-grained attention weight vector Specifically, it is expressed as:

[0143] \(W = W_c + W_s\) (18), coa c s (18),

[0144] Next, adjust each channel of the coarse-grained attention weight vector \(W\) according to the input feature \(X\), and use the channel shuffle operation to rearrange each channel of \(X\) and \(W\) in an alternating manner to obtain the fine-grained attention weight vector coa Expressed as: coa Expressed as:

[0145]

[0146] where \(\sigma\) represents the Sigmoid operation, represents a grouped convolutional layer with a kernel size of \(k\times k\), \(CS(\cdot)\) represents the channel shuffle operation, \([\cdot]\) represents the channel-level concatenation operation. Finally, according to the fine-grained attention weight \(W\), the input features \(X_1\) and \(X_2\) are weighted and fused to obtain the fused output \(X\) Fuse , expressed as:

[0147]

[0148] The fine-grained fusion module calculates the attention weights for each channel, focusing on the important information in each channel. Thus, it effectively fuses the two input features and retains more useful information.

[0149] Further, in the classification module, the fused features output the true / false labels through a gated recurrent unit and a linear layer.

[0150] The classification module has a structure as Figure 8 shown, consisting of 1 gated recurrent unit, 2 linear layers, 1 batch normalization layer, and 1 SELU activation function.

[0151] Step S4: During the training process of the detection model, set the hyperparameters of the detection model to continuously minimize the loss function. When the set number of training times is reached, a trained detection model is obtained.

[0152] The loss function of the detection model in step S4, i.e., the cross-entropy loss function, is expressed as follows:

[0153]

[0154] where y represents the true label (0 for false and 1 for true), represents the predicted probability of the true label output by the detection model; log(·) represents the logarithmic function.

[0155] Specifically, in this embodiment, the number of training times is set to 100 times.

[0156] Further, the training process in step S4 includes the following sub-steps:

[0157] Step S4-1: Input the original waveform feature X Raw extracted from the training data into the Sinc module. After being processed by the Sinc function filter, the waveform time-frequency feature X RawTF is obtained;

[0158] Step S4-2: Input the waveform time-frequency feature X RawTF into the depth extraction module of the original waveform feature. It passes through a network composed of 6 one-dimensional convolutional blocks and 6 global enhanced attention overlaps in sequence to obtain the deep feature vector X′ Raw ; Input the multi-scale fundamental frequency feature X F0 into the depth extraction module of the multi-scale fundamental frequency feature. It passes through 6 two-dimensional convolutional blocks, 1 local enhanced attention, and 1 global enhanced attention in sequence to obtain the deep feature vector X′ F0 ;

[0159] Step S4-3: In the embodiment, the deep feature vectors X′ Raw and X′ F0Input it into the fine-grained fusion module for feature fusion, and then input the fused feature X Fuse into the classification module to obtain the classification result output by the detection model.

[0160] The test phase includes the following steps:

[0161] Step S5, obtain speech data to construct the test data of the model. The test data includes real speech and forged speech generated by speech synthesis and voice conversion algorithms, and perform preprocessing on them respectively to obtain the original waveform feature Z Raw and the multi-scale fundamental frequency feature Z F0 ;

[0162] In this example, the preprocessing process is the same as that in the training process; the test data comes from the ASVspoof2019LA dataset, including real speech and 13 types of forged speech generated by speech synthesis and voice conversion algorithms; the duration of most speech is 1 to 8s, and there are 71933 pieces of speech in the test data.

[0163] Step S6, input the original waveform feature and the multi-scale fundamental frequency feature into the trained model for model testing. During the testing process, obtain the speech authenticity classification result and evaluate the performance of the model.

[0164] Furthermore, the testing process in step S6 includes the following sub-steps:

[0165] Step S6-1, input the original waveform feature Z extracted from the test data Raw into the Sinc module, and after processing by the Sinc function filter, obtain the waveform time-frequency feature Z RawTF ;

[0166] Step S6-2, input the waveform time-frequency feature Z RawTF into the depth extraction module of the original waveform feature, and successively pass through a network composed of 6 one-dimensional convolutional blocks and 6 global enhanced attention overlaps to obtain the deep feature vector Z RawTF ; Input the multi-scale fundamental frequency feature Z extracted from the test data F0 into the depth extraction module of the multi-scale fundamental frequency feature, and successively pass through 6 two-dimensional convolutional blocks, 1 local enhanced attention, and 1 global enhanced attention to obtain the deep feature vector Z' F0 ;

[0167] Step S6-3, input the deep feature vectors Z' Raw and Z' F0 into the fine-grained fusion module for feature fusion, and then input the fused feature Z Fuse into the classification module to obtain the classification result output by the detection model.

[0168] It should be noted that the above content only illustrates the technical idea of the present invention and cannot be used to limit the protection scope of the present invention. For those of ordinary skill in the art, without departing from the principle of the present invention, several improvements and refinements can be made, and these improvements and refinements all fall within the protection scope of the claims of the present invention.

Claims

1. A method for detecting forged speech that fuses multi-scale fundamental frequency features and enhanced attention, characterized in that, It includes a training stage and a testing stage; The training stage includes the following steps: Step S1, obtain speech data to construct the training data of the model, where the training data includes real speech and forged speech generated by speech synthesis and speech conversion algorithms; Step S2: Perform different preprocessings on the obtained training data to respectively obtain the original waveform feature X Raw and the multi-scale fundamental frequency feature X F0 ; Step S3, input the original waveform features and multi-scale fundamental frequency features into the model respectively for model training; Step S4, during the training process of the detection model, set the hyperparameters of the detection model to continuously minimize the loss function, and obtain a trained detection model when the set number of training times is reached; The testing stage includes the following steps: Step S5: Obtain the test data for building the model from the voice data. The test data includes real voices and forged voices generated by voice synthesis and voice conversion algorithms. Respectively, perform preprocessing to obtain the original waveform feature X Raw and the multi-scale fundamental frequency feature X F0 ; Step S6, input the original waveform features and multi-scale fundamental frequency features into the trained model respectively for model testing. During the testing process, obtain the speech authenticity classification result and evaluate the performance of the model.

2. The method for detecting forged speech by fusing multi-scale fundamental frequency features and enhancing attention according to claim 1, wherein Step S2 preprocesses the obtained training data in different ways. Sample the training data to obtain the original waveform features, extract the fundamental frequency envelope from the training data and obtain the multi-scale fundamental frequency features through continuous wavelet transform, and the methods are as shown in formula (1) and formula (2). Among them, \(W(F_0)(\tau, t)\) represents the fundamental frequency spectrum, \(F_0(x)\) represents the fundamental frequency value at position \(x\), \(\psi\) represents the mother wavelet function, and \(\tau\) and \(t\) represent the scale and position of the wavelet respectively. When \(\tau_0 = 5\mathrm{ms}\), let \(i = 1,\cdots,10\). The fundamental frequency envelopes of a total of 10 scales form the multi-scale fundamental frequency feature \(X\). F0 。 3. The forged speech detection method that fuses multi-scale fundamental frequency features and enhances attention according to claim 1, wherein In step S3, the model consists of a Sinc module, a deep extraction module for original waveform features, a deep extraction module for multi-scale fundamental frequency features, a fine-grained fusion module, and a classification module: The Sinc module performs shallow feature extraction on the input original waveform features; the deep extraction module for original waveform features further extracts the features extracted by the Sinc module to obtain deep features, the deep extraction module for multi-scale fundamental frequency features extracts the input multi-scale fundamental frequency features to obtain deep features, the fine-grained fusion module fuses the deep features output by the two deep extraction modules to obtain fused features, and the classification module processes the input fused features to obtain the classification result.

4. The forged speech detection method that fuses multi-scale fundamental frequency features and enhances attention according to claim 3, wherein The Sinc module consists of 1 one-dimensional convolutional layer, 1 max pooling layer, 1 batch normalization layer, and 1 SELU activation function. The deep extraction module for original waveform features consists of 6 one-dimensional convolutional blocks and 6 global enhancement attention modules; each one-dimensional convolutional block consists of 2 one-dimensional convolutional layers, 1 batch normalization layer, 1 LeakyReLU activation function, and 1 max pooling layer, and the one-dimensional convolutional blocks and global enhancement attention modules are placed alternately. The global enhancement attention module consists of 1 one-dimensional convolutional layer, 2 normalization layers, 5 linear layers, 1 linear attention, 2 SiLU activation functions, and 1 multi-layer perceptron. The deep extraction module for multi-scale fundamental frequency features consists of 6 two-dimensional convolutional blocks, 1 local enhancement attention module, and 1 global enhancement attention module; each two-dimensional convolutional block consists of 1 two-dimensional convolutional layer, 1 batch normalization layer, and 1 ReLU activation function. The local enhancement attention module consists of 3 linear layers, 2 Dropout layers, and 1 average pooling layer. The fine-grained fusion module consists of 1 spatial attention module, 1 channel attention module, 1 Sigmoid activation function, and 1 two-dimensional convolutional layer. The spatial attention module consists of 1 two-dimensional convolutional layer, and the channel attention module consists of 1 adaptive average pooling layer, 2 two-dimensional convolutional layers, and 1 ReLU activation function. The classification module consists of 1 gated recurrent unit, 2 linear layers, 1 batch normalization layer, and 1 SELU activation function.

5. The method for detecting forged speech by fusing multi-scale fundamental frequency features and enhancing attention according to claim 4, wherein Global enhanced attention is used to extract the global information of the feature vector as follows: First, the input feature X passes through the normalization layer to obtain the normalized feature X norm , which is expressed as follows: X norm = Norm(X) (3), Among them, Norm(·) represents the normalization operation, and then, X norm successively passes through a linear layer, a convolutional layer, and the SiLU activation function to highlight useful information and obtain the input feature X of the linear attention in , which is expressed as follows: X in = SiLU(Conv(Linear(X norm )) (4), Among them, Linear(·) represents a linear layer, Conv(·) represents a stack of convolutional layers, SiLU(·) represents the SiLU activation function. The linear attention processes the input as a sequence of N features with the same number of channels C, so as to model the relationship between each feature in the sequence. According to the attention mechanism, first calculate the query vector the key vector and the value vector respectively, and the calculation formulas are as follows: Q = φ(X in W Q ) (5) K = φ(X in W K ) (6) V = X in W V (7), Among them, and represent the weight vectors corresponding to three linear layers; C represents the number of channels of the vector, d represents the dimension of the vector, N represents the length of the vector, and φ(·) represents the kernel function operation. Then, the dot product operation is performed on the query vector Q and the key vector K to obtain the attention weights. Finally, the value vector V is weighted according to the attention weights to highlight the important information in the input features, and the output X of the linear attention is obtained att , which is expressed as: wherein, is a vector of all 1s, which is used to normalize the attention weights. Finally, the output Y of the global enhanced attention is obtained through a linear layer, a normalization layer, and a multi-layer perceptron as follows: Y = X + Linear(X att X in ) + Mlp(Norm(X + Linear(X att X in ))) (9) Among them, Mlp(·) represents a multi-layer perceptron. The global enhanced attention has a global receptive field, aggregates context information, and thus can model global dependencies and extract more global information.

6. The forged speech detection method that fuses multi-scale fundamental frequency features and enhances attention according to claim 1, wherein Local enhanced attention is used to extract the local detailed information of the feature vector as follows: Local enhanced attention focuses on the input features at each spatial position (i, j), and calculates the similarity with all adjacent spatial positions within a local window of size K×K centered at (i, j). First, according to the attention mechanism, the attention weight vector and the value weight vector are calculated respectively, and the calculation formulas are as follows: A = XW A (10), V = XW V (11), Among them, and represent the weight vectors corresponding to two linear layers; H and W represent the height and width of the vector, C represents the number of channels of the vector, and K represents the side length of the local window; calculate all the values within the local window centered at (i, j) Specifically, it is expressed as: where {·} represents a set, represents rounding up; The attention weight vector at position (i, j) is reshaped into The aggregated value attention weights are obtained through the Softmax function, and then the value attention weights at position (i, j) are multiplied by all values V within the window i,j to obtain the weighted output Y at position (i, j) i,j , which can be expressed as: Y i,j = MatMul(Softmax(A i,j ), V i,j ) (13), Among them, MatMl(·) represents a matrix multiplication operation, and Softmax(·) represents the Softmax function; Then, different weighted outputs Y at the same position (i, j) from different K×K local windows i,j are added together to obtain the local enhanced attention output at position (i, j). It is expressed as: Finally, replace the value X of all positions (i, j) of the input feature X i,j with the output of local enhanced attention at that position to obtain the output Y of local enhanced attention; The local enhanced attention aggregates the detailed information of each local window, thereby capturing the local information in the input features.

7. The method for detecting forged speech by fusing multi-scale fundamental frequency features and enhancing attention according to claim 4, characterized in that The fine-grained fusion module is used to fuse the extracted feature vectors as follows: First, fuse the input features and by addition to obtain the fused input feature as shown below: X = X1 + X2 (15), where H and W represent the height and width of the feature vector, and C represents the number of channels of the feature vector; To obtain an attention weight vector with the same dimension and number of channels as the fused input feature The corresponding channel attention weight vector W c and spatial attention weight vector W s are calculated through channel attention and spatial attention respectively, and the formula is as follows: Among them, denotes a convolutional layer with a kernel size of k×k, max(0, x) denotes the ReLU activation function, [·] denotes the channel-level concatenation operation, and GAP c (·) denotes the global average pooling operation across the spatial dimensions, and GAP s (·) denotes the global average pooling operation across the channel dimensions, and GMP s (·) denotes the features processed by the global max pooling operation across the channel dimensions. In channel attention, to reduce the number of model parameters, the first 1×1 convolution reduces the channel dimension from C to (r is set to ), and the second 1×1 convolution then expands its channel dimension back to C; Then, the channel attention weight vector W c and the spatial attention weight vector W s are fused together to obtain a coarse-grained attention weight vector Specifically, it is expressed as: W coa = W c + W s (18), Next, adjust each channel of the coarse-grained attention weight vector W according to the input feature X, and use the channel shuffle operation to rearrange each channel of X and W in an alternating manner to obtain the fine-grained attention weight vector coa and represent it as: coa ​​ where, σ represents the Sigmoid operation, represents a grouped convolutional layer with a kernel size of k×k, CS(·) represents the channel shuffle operation, [·] represents the channel-level concatenation operation, and finally, according to the fine-grained attention weight W, the input features X1 and X2 are weighted and fused to obtain the fused output X Fuse , which is expressed as: The fine-grained fusion module calculates the attention weights for each channel, focuses on the important information in each channel, and thus effectively fuses the two input features and retains more useful information.

8. The forged speech detection method that fuses multi-scale fundamental frequency features and enhances attention according to claim 1, characterized in that The training process in step S4 is the same as the testing process in step S6 and includes the following sub-steps: Step S4-1, input the original waveform feature X Raw into the Sinc module. After being processed by the Sinc function filter, the waveform time-frequency feature X RawTF is obtained; Step S4-2, input the waveform time-frequency feature X RawTF into the deep extraction module of the original waveform feature, and successively pass through a network composed of 6 one-dimensional convolutional blocks and 6 global enhanced attention overlaps to obtain the deep feature vector X' Raw ; Input the multi-scale fundamental frequency feature X F0 into the deep extraction module of the multi-scale fundamental frequency feature, and successively pass through 6 two-dimensional convolutional blocks, 1 local enhanced attention, and 1 global enhanced attention to obtain the deep feature vector X' F0 ; Step S4-3, input the deep feature vectors X′ Raw and X′ F0 into the fine-grained fusion module for feature fusion, and then input the fused feature X Fuse into the classification module to obtain the classification result output by the detection model.

9. The forged speech detection method that fuses multi-scale fundamental frequency features and enhances attention according to claim 1, characterized in that The loss function of the detection model in step S4, that is, the cross-entropy loss function, is expressed as follows: where y represents the true label, represents the predicted probability of the true label output by the detection model; log(·) represents the logarithmic function.

Citation Information

Cited By

  • Adaptive deep forgery detection method based on modal interaction analysis

    CN120599287A

  • An adaptive deepfake detection method based on modal interaction analysis

    CN120599287B