Voice authentic identification method based on dual-track difference modeling

Through the speech pseudo-detection method based on binaural differential modeling, the use of texture enhancement and multi-head attention modules to process the speech features is solved, and the problem of insufficient accuracy and robustness of speech pseudo-detection in the prior art is achieved, and higher accuracy of pseudo-detection and model strength is achieved.

CN120279938AActive Publication Date: 2025-07-08INNER MONGOLIA UNIVERSITY

Patent Information

Application Number
CN202311223079.X
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2023-09-20
Publication Date
2025-07-08
Estimated Expiration
2043-09-20

AI Technical Summary

Technical Problem

The existing speech depth forgery detection methods are insufficient in the accuracy and robustness of complex speech situations, making it difficult to meet the high demands of widespreadness and intensiveness of speech detection.

Method used

The speech pseudo-detection method based on two-channel differential modeling is adopted, and the single and double-channel speech conversion model is pre-trained, mono audio is converted into stereo, left and right channels Mel spectral features are extracted, and feature extraction and fusion are extracted and fusion are used for texture enhancement and multi-head attention modules, and discrimination is performed in combination with the attention pooling layer.

Benefits of technology

It improves the accuracy and robustness of speech and false recognition, enhances the migration and generalization performance of the model, solves the problem of the disappearance of fine-grained differences in deep neural networks, and improves the accuracy and the strength of false recognition.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120279938A_ABST
    Figure CN120279938A_ABST
Patent Text Reader

Abstract

The invention discloses a voice authentic identification method based on dual-track difference modeling, and belongs to the technical field of voice authentic identification. The method comprises the following steps: converting monaural audio into stereophonic sound through a pre-trained monaural and dual-channel voice conversion model, extracting left and right channel Mel spectrum features, and performing texture enhancement optimization processing on an absolute value of a difference between the left and right channel Mel spectrum features and an original monaural Mel spectrum feature. The method comprises the following steps: firstly, optimizing the voice authenticity, respectively inputting an optimization result into a left double-branch feature extractor and a right double-branch feature extractor as input of double-branch feature extractors, carrying out extraction, processing and information fusion on feature information to obtain a final attention map, and inputting the final attention map into an attention pooling layer and a final dichotomy layer to obtain a voice authenticity discrimination result. According to the invention, two-track differential modeling, a fine-grained texture enhancement method and a multi-head attention feature fusion method are adopted to carry out authentic identification on the voice, so that the method has higher accuracy, mobility and generalization.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention belongs to the technical field of voice anti-forgery, and particularly relates to a voice anti-forgery method based on dual-channel difference modeling. Background Art

[0002] Deepfake, that is, deep forgery, uses deep learning technology to generate audio that people have never spoken, things that have never been done, or video content that has never existed. As the threshold and difficulty for the public to use voice deepfake software gradually decrease, once criminals use these software for illegal activities, it will pose huge challenges to aspects such as social trust, news authenticity, monitoring, and judicial evidence collection in our country.

[0003] Currently, the main voice deepfake detection methods are voice anti-forgery based on robust feature extraction and voice anti-forgery based on effective model design. With the continuous evolution of voice forgery technology, various complex voice situations emerge one after another on the Internet, posing higher requirements for the extensiveness and intensiveness of voice anti-forgery. The present invention starts from the mutual conversion between mono-channel and stereo, providing a unique perspective on determining the authenticity of voice. Summary of the Invention

[0004] Aiming at the above deficiencies in the prior art, the present invention provides a voice anti-forgery method based on dual-channel difference modeling, which greatly improves the accuracy and robustness of voice anti-forgery, and has better migration and generalization capabilities.

[0005] In order to achieve the above invention purpose, the technical solution adopted by the present invention is: a voice anti-forgery method based on dual-channel difference modeling, including the following steps:

[0006] S1: Convert the mono-channel audio into stereo through a pre-trained mono-stereo voice conversion model, and respectively extract the left-channel Mel spectrogram feature and the right-channel Mel spectrogram feature

[0007] S2: Obtain the absolute values of the differences between the left-channel and right-channel Mel spectrogram features and and the Mel spectrogram feature of the original mono-channel voice respectively and wherein, is the absolute value of the difference between the left-channel Mel spectrogram feature and the Mel spectrogram feature of the original mono-channel voice, is the absolute value of the difference between the right-channel Mel spectrogram feature and the Mel spectrogram feature of the original mono-channel voice;

[0008] S3: Take the absolute values and Perform texture enhancement processing, and respectively use two dual-branch feature extractors to extract and fuse information from the texture enhancement processing results to obtain the final attention map G'. L / R ;

[0009] S4: According to the final attention map G' L / R , use the attention pooling layer and the binary classification layer to obtain the discrimination result of speech authenticity.

[0010] The beneficial effects of this technical solution are as follows: In the present invention, dual-channel information is differentially modeled, and the difference between the Mel spectrum information after conversion to stereo and the original mono-channel information is modeled for speech anti-spoofing, which has higher accuracy, greater robustness, better transferability, and generalization performance.

[0011] The beneficial effects of the present invention are as follows: The present invention provides a speech anti-spoofing method based on dual-channel difference modeling, and uses a fine-grained texture enhancement method to process the feature map, which solves the problem of the disappearance of fine-grained differences in deep neural networks and improves the accuracy of speech anti-spoofing; uses a multi-head attention module to model the correlation of information in the left and right channels, and extracts more detailed and useful information in the dual channels from a more fine-grained perspective, improving the accuracy of anti-spoofing; uses an attention pooling module to further perform feature fusion, and at the same time uses a variety of data augmentation methods and a variety of different types of data sets to increase the transferability and generalization of the model, making the model more powerful and more robust.

[0012] Furthermore: The specific steps of the S1 pre-training are as follows:

[0013] A1: Obtain a paired audio data set as the training data set, where the paired audio data set includes a conditional time signal C representing the position and direction of the source and the listener 1:T ;

[0014] A2: Use a neural time warping module to read the conditional time signal C 1:T , predict the neural warping field ρ, and use a time convolution module to modify and simulate the predicted neural warping field ρ to obtain the optimal neural warping field ρ';

[0015] A3: According to the neural warping field ρ and the optimal neural warping field ρ ′ Calculate the cochlear signals of the left and right ears through a recursive activation function and

[0016] A4: According to the cochlear signals of the left and right ears and the time signal C 1:T , use a time convolution network including a wave net to reconstruct the Mel spectrum features of the left and right channels and

[0017] A5: Repeat steps A1 to A4 for a total of N times to obtain the parameters of the single - and dual - channel voice conversion models with N different training times and the left - and right - channel Mel - spectrum features. and

[0018] A6: Compare the left - and right - channel Mel - spectrum features of the single - and dual - channel voice conversion models with N different training times and Select the parameters of the single - and dual - channel voice conversion model with the optimal Mel - spectrum feature effect as the parameters of the final pre - trained model.

[0019] Furthermore, in step S3, the dual - branch feature extractor includes a convolutional filter layer, a residual network layer, a graph attention network layer, and a graph pooling layer;

[0020] The convolutional filter layer and the residual network layer are used to perform feature extraction on the absolute value and to obtain the features of the left - and right - channel information and

[0021] The graph attention network layer is used to perform attention weight calculation on the spatial attention graph G to obtain the aggregation result of adjacent nodes;

[0022] The graph pooling layer is used to select the subset of nodes with the largest amount of information in the aggregation result of adjacent nodes and combine them into a new attention graph to obtain the final attention graph G'. L / R 。

[0023] The above - mentioned further beneficial effect is: Using the dual - branch feature extractor, perform feature extraction on the absolute value and and perform attention weight calculation to aggregate node information to obtain the final attention graph G'. L / R 。

[0024] Furthermore, the specific steps of step S3 are:

[0025] S31: Perform texture enhancement operation processing on the absolute value and and input the processing results into two dual - branch feature extractors respectively;

[0026] S32: Use the convolutional filter layer and the residual network layer to perform feature extraction on the absolute value and respectively to obtain and where is the feature of the left-channel information, is the feature of the right-channel information;

[0027] S33: According to the left and right channel feature information and respectively use the multi-head attention module to predict the spatial attention map, and obtain and Among them, is the spatial attention map predicted by the left channel, is the spatial attention map predicted by the right channel;

[0028] S34: Through the attention module, the spatial attention maps predicted by the left and right channels and are subjected to feature information fusion to obtain the spatial attention map G;

[0029] S35: Calculate the attention weights for the spatial attention map G through the graph attention network layer, aggregate adjacent nodes, and obtain the aggregation result;

[0030] S36: Use the graph pooling layer to select a subset of the nodes with the most information in the aggregation result and combine them into a new attention map to obtain the final attention map G′ L / R .

[0031] The beneficial effects of the above further solution are as follows: The dual-branch architecture is used to process the left and right channel feature information respectively. Finally, the attention feature fusion of the dual-channel information is performed for the final decision-making; After combining the dual-channel stereo information processing, the forged audio is more easily exposed and successfully detected by the model, improving the accuracy and generalization of the model.

[0032] Furthermore: The specific steps for performing the texture enhancement operation in step S31 are as follows:

[0033] B1: Use the local average pooling layer to downsample the feature maps of the absolute values and to obtain the pooled feature map;

[0034] B2: Use the residual module to perform residual operation processing on the pooled feature map to obtain the channel texture information;

[0035] B3: Use the convolutional block to enhance the channel texture information to obtain the Mel spectrogram information of the left and right channels after texture enhancement.

[0036] The beneficial effects of the above further solution are as follows: By using the fine-grained texture enhancement method to average the global information, the problem of the disappearance of fine-grained differences in deep neural networks is solved, achieving an efficient fine-grained classification effect and improving the accuracy of voice anti-forgery.

[0037] Furthermore: The expression for the residual operation in step B2 is as follows:

[0038] T L / R = f(a L / R ) - D L / R

[0039] where T L / R is the channel texture information obtained after the residual operation, L and R are respectively the absolute values of the differences in Mel spectrogram information between the left and right channels of the stereo, f(a L / R ) is the feature map of the absolute value information of the differences in Mel spectrogram information between the left and right channels, and D L / R is the pooled feature map.

[0040] The above further beneficial effect is: Using the residual module to perform residual operation processing on the pooled feature map to obtain the channel texture information, which is convenient for texture enhancement processing of the channel information.

[0041] Furthermore: The expressions for the spatial attention maps and in step S33 are respectively as follows:

[0042]

[0043]

[0044] where is the absolute value of the difference in Mel spectrogram information between the left channel and the original mono channel, is the absolute value of the difference in Mel spectrogram information between the right channel and the original mono channel, f L (.) and f R (.) are respectively the process functions for feature extraction of the Mel spectrogram information of the left and right channels, and are respectively the features of the left and right channels extracted.

[0045] The beneficial effect of the above further solution is: Extracting features from the information of the left and right channels and predicting the spatial attention map, which is convenient for subsequent analysis and modeling.

[0046] Furthermore: The expression for the spatial attention map G in step S34 is as follows:

[0047] G = G(N, ε, h ′ )

[0048] where N is the number of nodes included in the spatial attention map, ε is the connection edges between all nodes including the self-connection, and h ′ is the feature representation.

[0049] The above further beneficial effect is: fusing the spatial attention maps of the left and right channels to obtain

[0050] Furthermore: the expression of the attention weight in step S35 is as follows:

[0051]

[0052] where u and n are both nodes, and α u,n is the attention weight after aggregation of nodes u and n, exp(.) is the exponential function, M(n) is the adjacent nodes of node n, W is the learnable weight, h ′ n is the feature vector of node n, h ′ u is the feature vector of node u, ⊙ is the element-wise multiplication, and h ′ w is the feature vector of node w.

[0053] The beneficial effect of the above further solution is: using the learnable weight to aggregate adjacent nodes through the self-attention mechanism to obtain more detailed and useful information, and improving the accuracy of forgery detection.

[0054] Furthermore: the expression of the node information of the n nodes in the final attention map G′ L / R is as follows:

[0055] o n = ReLU(BN(m n + h ′ n ))

[0056]

[0057] where o n is the node information of the n nodes in the final attention map G′ L / R , m n is the node information after aggregating the n nodes, ReLU is the activation function, BN is batch normalization, m n is the aggregation information of the nth node, h ′ n is the feature vector of node n, M(n) is the set of adjacent nodes of node n, α u,n is the attention weight between nodes u and n, and h ′ u is the feature vector of node u.

[0058] The above further beneficial effects are as follows: By using the ReLU function and the BN function, batch normalization processing and activation are performed on the n aggregated node information to obtain the final attention map G'. L / R The n node information in L / R . BRIEF DESCRIPTION OF THE DRAWINGS

[0059] Figure 1 It is a flowchart of a voice forgery detection method based on binaural difference modeling. DETAILED DESCRIPTION OF THE INVENTION

[0060] The following describes the specific embodiments of the present invention to facilitate those skilled in the art of this technology to understand the present invention. However, it should be clear that the present invention is not limited to the scope of the specific embodiments. For those of ordinary skill in the art of this technology, as long as various changes are within the spirit and scope of the present invention defined and determined by the appended claims, these changes are obvious, and all inventions created using the concept of the present invention are within the scope of protection.

[0061] As Figure 1 shown, the present invention provides a voice forgery detection method based on binaural difference modeling, including the following steps:

[0062] S1: Convert the monaural audio into stereo through a pre-trained monaural-to-binaural voice conversion model, and respectively extract the left-channel Mel spectrum features and the right-channel Mel spectrum features

[0063] S2: Obtain the absolute values of the differences between the left and right channel Mel spectrum features and and the Mel spectrum features of the original monaural voice respectively and wherein, is the absolute value of the difference between the left-channel Mel spectrum feature and the Mel spectrum feature of the original monaural voice, is the absolute value of the difference between the right-channel Mel spectrum feature and the Mel spectrum feature of the original monaural voice;

[0064] S3: Perform texture enhancement processing on the absolute values and , and respectively use two dual-branch feature extractors to extract and fuse the texture enhancement processing results to obtain the final attention map G'. L / R ;

[0065] S4: According to the final attention map G' L / R , use the attention pooling layer and the binary classification layer to obtain the discrimination result of the authenticity of the voice.

[0066] The present invention uses a large amount of speech data to pre-train a mono-stereo speech conversion model, and at the same time uses a texture enhancement module to solve the problem of the disappearance of fine-grained differences in deep neural networks; by differentially modeling the stereo information, it is proposed to model the difference between the Mel spectrogram information after conversion to stereo and the original mono information to obtain a voice anti-spoofing model method with higher accuracy and better generalization performance.

[0067] The specific steps of the pre-training in S1 are as follows:

[0068] A1: Obtain a paired audio dataset as the training dataset, where the paired audio dataset includes a conditional time signal C representing the position and direction of the source and the listener 1:T ;

[0069] A2: Use a neural time warping module to read the conditional time signal C 1:T to predict the neural warping field ρ, and use a temporal convolutional module to modify and simulate the predicted neural warping field ρ to obtain the optimal neural warping field ρ';

[0070] A3: Calculate the cochlear signals of the left and right ears through a recursive activation function according to the neural warping field ρ and the optimal neural warping field ρ ′ ; and

[0071] A4: According to the cochlear signals of the left and right ears and the time signal C 1:T , use a temporal convolutional network including a WaveNet to reconstruct the Mel spectrogram features of the left and right channels and

[0072] A5: Repeat steps A1 to A4 a total of N times to obtain the parameters of the mono-stereo speech conversion model with N different training times and the Mel spectrogram features of the left and right channels and

[0073] A6: Compare the Mel spectrogram features of the left and right channels of the mono-stereo speech conversion models with N different training times and , and select the parameters of the mono-stereo speech conversion model with the optimal Mel spectrogram feature effect as the parameters of the final pre-training model.

[0074] Before pre-training the mono-stereo voice conversion model, paired audio data with a fixed length of <mono, stereo> is required. Large-scale voice data comes from 20 hours of paired mono and stereo data of 14 different speakers, including 7 males and 7 females. Specifically, a human model equipped with binaural microphones on the ears is used as the listener. The participants are required to walk in a circle with a radius of 1.5 meters around the human model and have an unscripted conversation with it. The adjusted time signal C of each audio 1:T is obtained by recording from real scenarios.

[0075] In step S3, the dual-branch feature extractor includes a convolutional filter layer, a residual network layer, a graph attention network layer, and a graph pooling layer, and introduces a multi-head attention module and an attention module;

[0076] The convolutional filter layer and the residual network layer are used to calculate the absolute value of the difference between the mel-spectrogram features of the left and right channels and the mel-spectrogram features of the original mono voice and perform feature extraction to obtain the features of the left and right channel information and

[0077] The multi-head attention module is used to obtain the predicted left and right spatial attention maps according to the features and ; and

[0078] The attention module is used to fuse the predicted left and right spatial attention maps and to obtain the spatial attention map G;

[0079] The graph attention network layer is used to perform attention weight calculation on the spatial attention map G to obtain the aggregation result of adjacent nodes;

[0080] The graph pooling layer is used to select the subset of nodes with the largest amount of information in the aggregation result of adjacent nodes and combine them into a new attention map to obtain the final attention map G'; L / R .

[0081] In the embodiment of the present invention, the specific steps of step S3 for extracting and processing the dual-channel mel-spectrogram features are as follows:

[0082] S31: Perform texture enhancement operation processing on the absolute value and , and input the processing results into two dual-branch feature extractors respectively;

[0083] S32: Use the convolutional filter layer and the residual network layer to process the absolute value respectively With feature extraction is performed to obtain and Among them, is the feature of the left-channel information, is the feature of the right-channel information;

[0084] S33: According to the left and right channel feature information and respectively use the multi-head attention module to predict the spatial attention map, and obtain and Among them, is the spatial attention map predicted by the left channel, is the spatial attention map predicted by the right channel;

[0085] S34: Through the attention module, the spatial attention maps predicted by the left and right channels and are fused with feature information to obtain the spatial attention map G;

[0086] S35: Calculate the attention weights for the spatial attention map G through the graph attention network layer, aggregate adjacent nodes, and obtain the aggregation result;

[0087] S36: Use the graph pooling layer to select a subset of the nodes with the largest amount of information in the aggregation result and combine them into a new attention map to obtain the final attention map G′ L / R .

[0088] The specific steps of performing the texture enhancement operation in step S31 are as follows:

[0089] B1: Use the local average pooling layer to downsample the feature maps of the absolute values and to obtain the pooled feature map;

[0090] B2: Use the residual module to perform residual operation processing on the pooled feature map to obtain the channel texture information;

[0091] B3: Use the convolutional block to enhance the channel texture information to obtain the mel spectrogram information of the left and right channels after texture enhancement.

[0092] The expression of the residual operation processing in step B2 is as follows:

[0093] T L / R = f(a L / R ) - D L / R

[0094] Among them, T L / Ris the channel texture information obtained after the residual operation, L and R are respectively the absolute values of the differences between the Mel spectrum information of the left and right channels of the stereo, and f(a L / R ) is the feature map of the absolute value information of the difference between the Mel spectrum information of the left and right channels, and D L / R is the pooled feature map.

[0095] In step S33, the spatial attention maps and have the following expressions respectively:

[0096]

[0097]

[0098] Among them, is the absolute value of the difference between the Mel spectrum information of the left channel and the original mono channel, is the absolute value of the difference between the Mel spectrum information of the right channel and the original mono channel, f L (.) and f R (.) are respectively the process functions for feature extraction of the Mel spectrum information of the left and right channels, and are respectively the features of the left and right channels extracted.

[0099] The expression of the spatial attention map G in step S34 is as follows:

[0100] G = G(N, ε, h ′ )

[0101] Among them, N is the number of nodes included in the spatial attention map, ε is the connection edges between all nodes including the self-connection, and h ′ is the feature representation.

[0102] The expression for calculating the attention weight in step S35 is as follows:

[0103]

[0104] Among them, u and n are both nodes, α u,n is the attention weight after aggregation of nodes u and n, exp(.) is the exponential function, M(n) is the adjacent node of node n, W is the learnable weight, h ′ n is the feature vector of node n, h ′ u is the feature vector of node u, ⊙ is the element-wise multiplication, and h ′ w is the feature vector of node w. The expression of the node information of the n nodes in the final attention map G′ L / R in S36 is as follows:

[0105] o n = ReLU(BN(m n + h ′ n ))

[0106]

[0107] where o n is the node information of n nodes in the final attention map G', L / R m n is the node information after aggregating n nodes, ReLU is the activation function, BN is batch normalization, m n is the aggregation information of the nth node, h ′ n is the feature vector of node n, M(n) is the set of adjacent nodes of node n, α u,n is the attention weight between node u and n, h ′ u is the feature vector of node u.

[0108] The working principle and process of the present invention: First, a large-scale speech data is used to pre-train the mono-stereo speech conversion model, and at the same time, a texture enhancement module is used to solve the problem of the disappearance of fine-grained differences in deep neural networks; the mono-channel audio is converted into stereo by the trained mono-stereo speech conversion model, and the dual-channel Mel spectrogram features are extracted; the absolute value of the difference between the dual-channel Mel spectrogram information and the Mel spectrogram information of the original mono-channel speech is used as the input of the audio anti-forgery model, and is respectively input into two dual-branch feature extractors to extract and process the feature information, obtaining the final information G' L / R ; the final information is input into the attention pooling layer and the binary classification layer to obtain the discrimination result of the authenticity of the speech.

[0109] The beneficial effects of the present invention are as follows: The present invention provides a speech anti-forgery method based on dual-channel difference modeling, providing a new perspective for the research in the field of speech anti-forgery and improving the accuracy of anti-forgery; adopting a fine-grained texture enhancement method to solve the problem of the disappearance of fine-grained differences in deep neural networks and improving the accuracy of speech anti-forgery; adopting a multi-head attention mechanism to model the correlation of information in the left and right channels, extracting more detailed and useful information in the dual-channel from a more fine-grained perspective and improving the accuracy of anti-forgery; using the final attention module and attention pooling module to further perform feature fusion, and at the same time using a variety of data augmentation methods and a variety of different types of data sets to increase the transferability and generalization of the model, making the anti-forgery model more powerful and more robust.

Claims

1. A voice anti-spoofing method based on dual-channel difference modeling, characterized in that, Including the following steps: S1: Convert the monophonic audio to stereo through a pre-trained monophonic-to-stereo voice conversion model, and separately extract the left-channel Mel spectrogram features and the right-channel Mel spectrogram features S2: Obtain the mel-spectrum features of the left and right channels and the absolute value of the difference between the mel-spectrum features of the left and right channels and those of the original single-channel speech respectively and where is the absolute value of the difference between the mel-spectrum features of the left channel and those of the original single-channel speech, is the absolute value of the difference between the mel-spectrum features of the right channel and those of the original single-channel speech; S3: Take the absolute value and perform texture enhancement processing, and respectively use two dual-branch feature extractors to extract and fuse information from the texture enhancement processing results to obtain the final attention map G' L / R ; S4: According to the final attention map G' L / R , the discriminant result of the speech authenticity is obtained by using the attention pooling layer and the binary classification layer.

2. The voice forgery detection method based on binaural difference modeling according to claim 1, wherein, The specific steps of the pre-training in step S1 are as follows: A1: Obtain a paired audio dataset as the training dataset, where the paired audio dataset includes a conditional time signal C representing the position and orientation of the source and the listener 1:T ; A2: Reading the conditional time signal C using a neural time warping module, predicting the neural warping field ρ, and modifying and simulating the predicted neural warping field ρ using a temporal convolutional module to obtain the optimal neural warping field ρ'; 1:T Performing reading, predicting the neural warping field ρ, and modifying and simulating the predicted neural warping field ρ using a temporal convolutional module to obtain the optimal neural warping field ρ'; A3: The curling signals of the left and right ears are calculated from the neural distortion field ρ and the optimal neural distortion field ρ′ through a recursive activation function. and A4: According to the curling signals of the left and right ears and the time signal C 1:T , the left and right channel Mel-spectrum features are reconstructed by using a temporal convolutional network including a wave grid and A5: Repeat steps A1 to A4 for a total of N times to obtain the parameters of the single-channel and dual-channel voice conversion models with N different training times and the left and right channel mel-spectrum features and A6: The left and right channel Mel-spectrum features of single-channel and dual-channel speech conversion models with N different training times are compared with each other, and the parameters of the single-channel and dual-channel speech conversion model with the optimal Mel-spectrum feature effect are selected as the parameters of the final pre-trained model.

3. The voice anti-counterfeiting method based on dual-channel difference modeling according to claim 1, characterized in that In step S3, the dual-branch feature extractor includes a convolutional filter layer, a residual network layer, a graph attention network layer, and a graph pooling layer; The convolutional filter layer and the residual network layer are used for the absolute value and to perform feature extraction to obtain the features of the left and right channel information and The graph attention network layer is used to perform an attention weight operation on the spatial attention graph G to obtain an aggregation result of adjacent nodes; The said graph pooling layer is used to select a subset of nodes with the largest amount of information from the aggregation results of adjacent nodes, combine them into a new attention graph, and obtain the final attention graph G′ L / R 。 4. The voice anti-counterfeiting method based on binaural difference modeling according to claim 1, wherein, The specific steps of step S3 are as follows: S31: Perform texture enhancement operation on the absolute value and and input the processing results into two double-branch feature extractors respectively; S32: Use a convolutional filter layer and a residual network layer to separately perform feature extraction on the absolute value and to obtain and wherein is the feature of the left channel information, is the feature of the right channel information; S33: According to the left and right channel feature information and respectively use the multi-head attention module to predict the spatial attention map, obtaining and wherein is the spatial attention map predicted for the left channel, is the spatial attention map predicted for the right channel; S34: Spatial attention maps predicted for the left and right channels through the attention module and perform feature information fusion to obtain the spatial attention map G; S35: Perform an attention weight calculation on the spatial attention graph G through the graph attention network layer, aggregate adjacent nodes, and obtain an aggregation result; S36: Use the graph pooling layer to select a subset of nodes with the most information in the aggregation result, combine them into a new attention map, and obtain the final attention map G' L / R .

5. The voice forgery detection method based on binaural difference modeling according to claim 4, wherein The specific steps of performing the texture enhancement operation in step S31 are as follows: B1: Using a local average pooling layer to downsample the feature map of the absolute value and to obtain a pooled feature map; B2: Use a residual module to perform a residual operation on the pooled feature map to obtain channel texture information; B3: Use a convolutional block to enhance the channel texture information to obtain the mel spectrogram information of the left and right channels after texture enhancement.

6. The voice anti-counterfeiting method based on binaural difference modeling according to claim 5, characterized in that The expression of the residual operation in step B2 is as follows: T L / R = f(a L / R ) - D L / R Among them, T L / R is the channel texture information obtained after the residual operation, L and R are respectively the absolute values of the differences between the Mel spectrum information of the left and right channels of the stereo, f(a L / R ) is the feature map of the absolute value information of the difference between the Mel spectrum information of the left and right channels, D L / R is the pooled feature map.

7. The voice forgery detection method based on binaural difference modeling according to claim 4, wherein The spatial attention map in step S33 and have the following expressions respectively: Wherein, is the absolute value of the difference between the left channel and the Mel-spectrum information of the original monophonic sound, is the absolute value of the difference between the right channel and the Mel-spectrum information of the original monophonic sound, f L (.) and f R (.) are the process functions for extracting the Mel-spectrum information features of the left and right channels respectively, and are the features of the left and right channels extracted respectively.

8. The voice forgery detection method based on binaural difference modeling according to claim 4, wherein The expression of the spatial attention graph G in step S34 is as follows: G = G(N, ε, h′) where N is the number of nodes included in the spatial attention graph, ε is the connection edges between all nodes including self-connections, and h′ is the feature representation.

9. The voice anti-counterfeiting method based on binaural difference modeling according to claim 4, wherein The expression of the attention weight in step S35 is as follows: Among them, both u and n are nodes, and α u,n is the aggregated attention weight of nodes u and n, exp(.) is the exponential function, M(n) is the adjacent node of node n, W is the learnable weight, h′ n is the feature vector of node n, h′ u is the feature vector of node u, ⊙ is the element-wise multiplication, and h′ w is the feature vector of node w.

10. The voice anti-counterfeiting method based on binaural difference modeling according to claim 4, characterized in that The node information expressions of the n nodes in the final attention map G' in step S36 L / R are as follows: o n = ReLU(BN(m n + h′ n )) Among them, o n is the node information of n nodes in the final attention map G', L / R m is the node information after aggregating n nodes, ReLU is the activation function, BN is batch normalization, n m is the aggregation information of the nth node, h' n is the feature vector of node n, M(n) is the set of adjacent nodes of node n, α n is the attention weight between node u and n, h' u,n is the feature vector of node u. u ​

Citation Information

Patent Citations

  • Voice emotion recognition method and device of multi-channel auto-encoder based on attention feature fusion

    CN115472182A

  • Voice authentic identification method and system based on residual attention network

    CN115831099A

  • Intelligent voice forgery attack detection method based on attention mechanism

    CN116416997A

  • Multi-style audio synthesis method, apparatus and device, and storage medium

    WO2022116432A1

Cited By

  • Breathing sound and additional sound automatic identification method based on multi-label deep learning

    CN122201357A