Voice traceability evidence obtaining method and device, electronic equipment and storage medium

By introducing the channel attention mechanism and gradient inversion network layer into the voice traceability forensics model, the accuracy of the audio traceability task generated by the audio large-scale model is solved, and the fast and accurate results of voice traceability for evidence for evidence for evidence is achieved.

CN120581034AActive Publication Date: 2025-09-02THE FIRST RES INST OF MIN OF PUBLIC SECURITY +1
View PDF 5 Cites 0 Cited by

Patent Information

Application Number
CN202510738971.4
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-06-04
Publication Date
2025-09-02
Estimated Expiration
2045-06-04

AI Technical Summary

Technical Problem

The accuracy of existing audio traceability tasks is difficult to ensure when facing the audio generated by the audio big model, especially due to the improvement of naturalness and fidelity of speech synthesis, which makes it difficult to trace the traceability tasks between different generation methods.

Method used

The channel attention mechanism is added to the speech traceability forensics model, and a gradient inversion network layer is introduced during the training process. The robustness and fine-grained audio feature learning ability of the model are improved through self-supervised pre-training and adversarial training.

Benefits of technology

It improves the accuracy and robustness of speech traceability and evidence collection, and can quickly and accurately determine the source or method of speech generation.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120581034A_ABST
    Figure CN120581034A_ABST
Patent Text Reader

Abstract

The invention provides a voice traceability evidence obtaining method and device, electronic equipment and a storage medium, and the method comprises the steps: processing a to-be-detected voice based on an audio band-pass filter, and determining an effective audio of the to-be-detected voice; inputting the effective audio into a pre-trained voice traceability evidence obtaining model, extracting audio advanced features of the effective audio, performing convolution processing on the audio advanced features based on a convolution pooling network layer of the voice traceability evidence obtaining model, and determining the audio advanced features after convolution processing; processing the audio advanced features after convolution processing based on a channel attention mechanism of a voice traceability evidence obtaining model and a full-connection network layer, determining target advanced features, and calculating probability distribution of each traceability category corresponding to the target advanced features based on an activation function, and predicting a voice traceability evidence obtaining result of the to-be-detected voice. And a voice traceability evidence obtaining result corresponding to the voice is quickly and accurately determined through the voice traceability evidence obtaining model.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present application relates to the field of speech recognition technology, and in particular to a speech source tracing and evidence collection method, device, electronic device and storage medium. Background Art

[0002] Existing audio attribution tasks still use raw audio features as input, such as acoustic features such as MFCC, Fbank, and LFCC. These features are widely used in some audio classification tasks. However, with the development of large audio models, the exponential growth of training data and model parameters has led to significant improvements in the naturalness and realism of speech synthesis. Previous speech synthesis and speech conversion algorithms often suffer from problems such as a strong mechanical feel, loss of sound quality, lack of rhythm, and lack of emotion. However, large audio generation models, through deep learning and complex acoustic modeling, can more accurately simulate the characteristics and patterns of human speech, generate more natural and fluent audio content, and accurately synthesize or imitate the speaker's timbre, rhythm, emotion, pronunciation habits, and other information. With the powerful learning ability of large audio models, the difference between the audio generated by different generation methods and the real audio is getting smaller and smaller, which also makes the attribution task between different generation methods more difficult. Therefore, how to improve the accuracy of audio attribution determination has become a technical issue that cannot be underestimated. Summary of the Invention

[0003] In view of this, the purpose of this application is to provide a speech tracing forensics method, device, electronic device and storage medium, by adding a channel attention mechanism to the speech tracing forensics model to improve the model's learning ability for fine-grained audio features, and introducing a gradient reversal network layer during the training process to improve the robustness of the speech tracing forensics model, so that the speech tracing forensics result corresponding to the speech can be quickly and accurately determined through the speech tracing forensics model.

[0004] The present invention provides a method for collecting voice evidence, which includes:

[0005] Acquire a speech to be detected, process the speech to be detected based on an audio bandpass filter, and determine a valid audio of the speech to be detected;

[0006] Inputting the valid audio into a pre-trained speech forensics model, extracting high-level audio features of the valid audio, and performing convolution processing on the high-level audio features based on the convolutional pooling network layer of the speech forensics model to determine the high-level audio features after convolution processing; wherein the speech forensics model is determined by adding a channel attention mechanism and a gradient reversal network layer to the neural network model to simulate generative adversarial training;

[0007] Based on the channel attention mechanism and the fully connected network layer of the speech tracing forensics model, the high-level audio features after convolution processing are processed to determine the target high-level features, and the probability distribution of each tracing category corresponding to the target high-level features is calculated based on the activation function to predict the speech tracing forensics result of the speech to be detected; wherein, tracing forensics refers to verifying the generation source or generation method of the generated speech.

[0008] In one possible implementation, the channel attention mechanism and the fully connected network layer based on the speech tracing forensics model process the high-level audio features after convolution processing to determine the target high-level features, calculate the probability distribution of each tracing category corresponding to the target high-level features based on the activation function, and predict the speech tracing forensics result of the speech to be detected, including:

[0009] Performing attention processing on the high-level audio features after the convolution processing based on a channel attention mechanism to determine the high-level audio features after the attention processing;

[0010] Processing the audio high-level features after the attention processing based on the fully connected network layer to determine the target high-level features;

[0011] Based on the activation function, the probability distribution of each tracing category corresponding to the target high-level feature is calculated to determine multiple probability values, and the tracing category corresponding to the maximum probability value is determined as the voice tracing forensics result of the voice to be detected.

[0012] In one possible implementation, the voice tracing forensics model is trained by the following steps:

[0013] Screening out sample audio data from the training data after data expansion, performing filtering on the sample audio data, and determining sample valid audio of the sample audio data;

[0014] Inputting the sample valid audio into a neural network model, extracting sample audio high-level features of the sample valid audio based on a self-supervised pre-training method, and performing adversarial training on the sample audio high-level features until the neural network model converges, and stopping the training of the neural network model;

[0015] The model parameters of each layer of a preset number of neural network models with the best performance during the model training process are averaged, and the voice tracing forensics model is constructed based on the averaged model parameters of each layer.

[0016] In one possible implementation, the training data after data expansion is determined by:

[0017] Add noise with different signal-to-noise ratios to the training data to determine the training data after data expansion;

[0018] Alternatively, a gradient-based adversarial example generation algorithm generates different adversarial examples to determine the training data after data augmentation.

[0019] In one possible implementation, inputting the sample valid audio into a neural network model, extracting sample audio high-level features of the sample valid audio based on self-supervised pre-training, and performing adversarial training on the sample audio high-level features until the neural network model converges, includes:

[0020] Performing convolution pooling processing and channel attention mechanism processing on the sample audio high-level features to determine the sample audio high-level features after the channel attention mechanism processing;

[0021] Processing the sample audio high-level features after the channel attention mechanism based on the first fully connected network layer, and outputting the first sample audio high-level features and the predicted tracing label corresponding to the first sample audio high-level features;

[0022] Add a gradient reversal network layer after the first fully connected network layer, input the first sample audio high-level feature into the gradient reversal network layer for processing, output the processed first sample audio high-level feature, input the processed first sample audio high-level feature into the second fully connected network layer after the gradient reversal network layer for processing, output the first sample audio high-level feature processed by the second fully connected network layer, input the first sample audio high-level feature processed by the second fully connected network layer into the third fully connected network layer after the second fully connected network layer, output the second sample audio high-level feature and the adversarial label of the second sample audio high-level feature;

[0023] Based on the first sample audio high-level features, the second sample audio high-level features, the predicted tracing label, the adversarial label and the true tracing label of the sample audio data, the total loss value of the neural network model is determined, and the training of the neural network model is stopped until the neural network model converges.

[0024] In one possible implementation, performing convolution pooling processing and channel attention mechanism processing on the sample audio high-level features to determine the sample audio high-level features after the channel attention mechanism processing includes:

[0025] Performing global maximum pooling processing and global average pooling processing on the sample audio high-level features in a spatial dimension to determine global average pooling features and maximum pooling features;

[0026] Inputting the global average pooling feature and the maximum pooling feature into the multi-layer perceptron network layer respectively for channel dimension feature learning, and outputting the global average pooling feature and the maximum pooling feature processed by the multi-layer perceptron network layer;

[0027] Splicing the global average pooling feature processed by the multi-layer perceptron network layer and the maximum pooling feature processed by the multi-layer perceptron network layer to determine the spliced ​​feature;

[0028] Mapping the concatenated features based on an activation function to determine a channel attention weight matrix;

[0029] Based on the channel attention weight matrix, the global average pooling features and the maximum pooling features, the high-level features of the sample audio after processing by the channel attention mechanism are determined.

[0030] In one possible implementation, determining a total loss value of the neural network model based on the first sample audio high-level features, the second sample audio high-level features, the predicted source label, the adversarial label, and the true source label of the sample audio data, and stopping training the neural network model until the neural network model converges, includes:

[0031] Calculating the first sample audio high-level features, the predicted source tracing label, and the true source tracing label based on a weighted cross entropy loss function to determine a classification loss value;

[0032] Calculating the second sample audio high-level features and the adversarial label based on a binary cross entropy loss function to determine an adversarial loss value;

[0033] Based on the classification loss value and the adversarial loss value, a total loss value of the neural network model is determined, and training of the neural network model is stopped until the neural network model converges.

[0034] The present application also provides a voice source tracing and evidence collection device, which includes:

[0035] A filtering processing module is used to obtain a speech to be detected, process the speech to be detected based on an audio bandpass filter, and determine a valid audio of the speech to be detected;

[0036] A first processing module is configured to input the valid audio into a pre-trained speech forensics model, extract high-level audio features of the valid audio, perform convolution processing on the high-level audio features based on a convolutional pooling network layer of the speech forensics model, and determine the high-level audio features after convolution processing; wherein the speech forensics model is determined by adding a channel attention mechanism and a gradient reversal network layer to a neural network model to simulate generative adversarial training;

[0037] The second processing module is used to process the high-level audio features after convolution processing based on the channel attention mechanism and the fully connected network layer of the speech tracing forensics model, determine the target high-level features, calculate the probability distribution of each tracing category corresponding to the target high-level features based on the activation function, and predict the speech tracing forensics results of the speech to be detected; wherein, tracing forensics refers to verifying the generation source or generation method of the generated speech.

[0038] An embodiment of the present application also provides an electronic device, comprising: a processor, a memory, and a bus, wherein the memory stores machine-readable instructions executable by the processor. When the electronic device is running, the processor and the memory communicate through the bus, and when the machine-readable instructions are executed by the processor, the steps of the voice tracing and forensics method as described above are performed.

[0039] An embodiment of the present application further provides a computer-readable storage medium having a computer program stored thereon. When the computer program is executed by a processor, the steps of the voice tracing and evidence collection method as described above are executed.

[0040] The embodiments of the present application provide a speech source forensics method, device, electronic device and storage medium, the speech source forensics method comprising: obtaining a speech to be detected, processing the speech to be detected based on an audio bandpass filter, and determining the effective audio of the speech to be detected; inputting the effective audio into a pre-trained speech source forensics model, extracting the high-level audio features of the effective audio, and convolving the high-level audio features based on the convolutional pooling network layer of the speech source forensics model to determine the high-level audio features after the convolution processing; wherein the speech source forensics model is determined by simulating generative adversarial training by adding a channel attention mechanism and a gradient reversal network layer to a neural network model; processing the high-level audio features after the convolution processing based on the channel attention mechanism and the fully connected network layer of the speech source forensics model to determine the target high-level features, calculating the probability distribution of each tracing category corresponding to the target high-level features based on the activation function, and predicting the speech source forensics result of the speech to be detected; wherein source forensics refers to verifying the generation source or generation method of the generated speech. By adding a channel attention mechanism to the speech tracing forensics model, the model's ability to learn fine-grained audio features is improved, and the gradient reversal network layer is introduced during the training process to improve the robustness of the speech tracing forensics model, so that the speech tracing forensics model can quickly and accurately determine the speech tracing forensics results corresponding to the speech.

[0041] In order to make the above-mentioned objects, features and advantages of the present application more obvious and easy to understand, preferred embodiments are given below and described in detail with reference to the accompanying drawings. BRIEF DESCRIPTION OF THE DRAWINGS

[0042] In order to more clearly illustrate the technical solutions of the embodiments of the present application, the following is a brief introduction to the drawings required for use in the embodiments. It should be understood that the following drawings only show certain embodiments of the present application and therefore should not be regarded as limiting the scope. For ordinary technicians in this field, other relevant drawings can be obtained based on these drawings without creative work.

[0043] Figure 1 A flowchart of a voice tracing and evidence collection method provided in an embodiment of the present application;

[0044] Figure 2 This is one of the structural diagrams of a voice tracing and evidence collection device provided in an embodiment of the present application;

[0045] Figure 3 This is a second structural diagram of a voice tracing and evidence collection device provided in an embodiment of the present application;

[0046] Figure 4A schematic diagram of the structure of an electronic device provided in an embodiment of the present application. DETAILED DESCRIPTION

[0047] In order to make the purpose, technical solutions and advantages of the embodiments of the present application clearer, the technical solutions in the embodiments of the present application will be clearly and completely described below in conjunction with the drawings in the embodiments of the present application. Obviously, the described embodiments are only a part of the embodiments of the present application, rather than all of the embodiments. The components of the embodiments of the present application generally described and shown in the drawings here can be arranged and designed in various different configurations. Therefore, the following detailed description of the embodiments of the present application provided in the drawings is not intended to limit the scope of the application for which protection is claimed, but merely represents the selected embodiments of the present application. Based on the embodiments of the present application, each other embodiment obtained by those skilled in the art without making creative work falls within the scope of protection of the present application.

[0048] First, the application scenarios to which this application is applicable are introduced. This application can be applied in the field of speech recognition technology.

[0049] Research has found that existing audio attribution tasks still use raw audio features as input, such as acoustic features such as MFCC, Fbank, and LFCC. These features are widely used in some audio classification tasks. However, with the development of large audio models, the exponential growth of training data and model parameters has led to significant improvements in the naturalness and realism of speech synthesis. Previous speech synthesis and speech conversion algorithms often suffered from problems such as a strong mechanical feel, loss of sound quality, lack of rhythm, and lack of emotion. However, large audio generation models, through deep learning and complex acoustic modeling, can more accurately simulate the characteristics and patterns of human speech, generating more natural and fluent audio content. They can accurately synthesize or imitate the speaker's timbre, rhythm, emotion, pronunciation habits, and other information. Thanks to the powerful learning capabilities of large audio models, the audio generated by different generation methods is becoming increasingly different from the real audio, which also makes the attribution task between different generation methods more difficult. Therefore, how to improve the accuracy of audio attribution determination has become a technical issue that cannot be underestimated.

[0050] Based on this, an embodiment of the present application provides a speech tracing forensics method, which improves the model's learning ability for fine-grained audio features by adding a channel attention mechanism to the speech tracing forensics model, and introduces a gradient reversal network layer during the training process to improve the robustness of the speech tracing forensics model, so that the speech tracing forensics result corresponding to the speech can be quickly and accurately determined through the speech tracing forensics model.

[0051] See also Figure 1 , Figure 1 This is a flow chart of a voice tracing and evidence collection method provided by an embodiment of the present application. Figure 1 As shown in , the voice source tracing and evidence collection method provided by the embodiment of the present application includes:

[0052] S101: Acquire a speech to be detected, process the speech to be detected based on an audio bandpass filter, and determine a valid audio of the speech to be detected.

[0053] In this step, the speech to be detected is acquired, and the speech to be detected is filtered using an audio bandpass filter to determine the effective audio of the speech to be detected.

[0054] Here, in addition to the voice to be detected, the audio to be detected can also be detected, and this part is not specifically limited.

[0055] Here, the bandpass filter used is a fifth-order Butterworth bandpass filter with a pass frequency range of [50Hz, 6000Hz]. This is because below 50Hz, it is generally low-frequency noise of some devices themselves, which does not contain the effective information required for audio tracing. If this frequency band information is retained, the tracing task will converge slowly. In addition, with the development of large-model speech generation technology, the difference between different generation algorithms at high frequencies is very small, and it is difficult for the tracing engine to trace the audio based on high-frequency features. Therefore, the present invention proposes to filter the frequency domain features above 6000Hz and only retain the effective audio information in the [50Hz, 6000Hz] interval for the audio tracing task.

[0056] S102: Input the valid audio into a pre-trained speech tracing forensics model, extract the high-level audio features of the valid audio, perform convolution processing on the high-level audio features based on the convolutional pooling network layer of the speech tracing forensics model, and determine the high-level audio features after convolution processing.

[0057] In this step, the valid audio is input into the speech tracing forensics model, the high-level audio features of the valid audio are extracted, and the high-level audio features are convolved according to the convolutional pooling network layer of the speech tracing forensics model to determine the high-level audio features after convolution processing.

[0058] Among them, the speech tracing forensics model is determined by adding a channel attention mechanism and a gradient reversal network layer to the neural network model to simulate generative adversarial training.

[0059] Here, usually the audio tracing algorithm uses 1 or 2 audio features as input, but it often has certain limitations in the tracing task of generating audio for large models. In addition, the supervised training of neural network models is very dependent on large labeled data sets. Generally speaking, the larger the data set, the better the generalization performance of the model. However, it costs a lot to collect corresponding large labeled data sets for different tasks, and low-quality labels may cause the model training process to not converge. In order to alleviate the dependence on supervised training, this patent proposes to use self-supervised pre-training and supervised fine-tuning strategies to improve the performance of the tracing model, and extract representation vectors through the self-supervised pre-training model to make up for the problem of insufficient feature extraction caused by insufficient data in supervised training. The high-level audio features here can be extracted using models based on large-scale unsupervised training such as Wav2vec, CPC, Wav2vec2.0, HuBERT, Whisper, and WavLM.

[0060] S103: Based on the channel attention mechanism and the fully connected network layer of the speech tracing forensics model, the high-level audio features after convolution processing are processed to determine the target high-level features, and the probability distribution of each tracing category corresponding to the target high-level features is calculated based on the activation function to predict the speech tracing forensics result of the speech to be detected.

[0061] In this step, the high-level audio features after convolution processing are processed according to the channel attention mechanism and the fully connected network layer of the speech tracing forensics model to determine the target high-level features. Then, the probability distribution of each tracing category corresponding to the target high-level features is calculated according to the activation function to predict the speech tracing forensics results of the speech to be detected.

[0062] Among them, the traceability category can be traceability between different large audio models, such as Sora, Audiobox, VASA-1, Mega-TTS, chatTTS, etc.; it can also be traceability of different commercial software, such as ACE studio, Suno, FreeTTS, TTSMaker, Uberduck, etc.; it can trace the audio forgery methods, such as speech synthesis, timbre conversion, speech splicing, speech replay, and replay after forgery; it can trace the source of different vocoders, such as those based on autoregression, Flows flow, generative adversarial networks, VAE, diffusion, etc.; it can trace the source of different audio synthesis manufacturers, such as iFlytek, Baidu, Alibaba, Tencent, Byte, etc.

[0063] In one possible implementation, the channel attention mechanism and the fully connected network layer based on the speech tracing forensics model process the high-level audio features after convolution processing to determine the target high-level features, calculate the probability distribution of each tracing category corresponding to the target high-level features based on the activation function, and predict the speech tracing forensics result of the speech to be detected, including:

[0064] A: Perform attention processing on the high-level audio features after the convolution processing based on the channel attention mechanism to determine the high-level audio features after the attention processing.

[0065] Here, the high-level audio features after convolution processing are subjected to attention processing according to the channel attention mechanism to determine the high-level audio features after attention processing.

[0066] B: Processing the audio high-level features after the attention processing based on the fully connected network layer to determine the target high-level features.

[0067] Here, the high-level audio features after attention processing are processed according to the fully connected network layer to determine the target high-level features.

[0068] C: Based on the activation function, the probability distribution of each tracing category corresponding to the target high-level feature is calculated to determine multiple probability values, and the tracing category corresponding to the maximum probability value is determined as the voice tracing forensics result of the voice to be detected.

[0069] Here, the probability distribution of each tracing category corresponding to the target high-level feature is calculated according to the activation function, and multiple probability values ​​are determined. The tracing category corresponding to the maximum probability value is determined as the speech tracing forensics result of the speech to be detected.

[0070] In one possible implementation, the voice tracing forensics model is trained by the following steps:

[0071] (1): Sample audio data is selected from the training data after data expansion, and the sample audio data is filtered to determine the sample effective audio of the sample audio data.

[0072] Here, sample audio data is screened out from the training data after data expansion, and filtering is performed on the sample audio data to determine sample valid audio of the sample audio data.

[0073] In one possible implementation, the training data after data expansion is determined by:

[0074] Noise with different signal-to-noise ratios is added to the training data to determine the training data after data expansion; alternatively, a gradient-based adversarial sample generation algorithm generates different adversarial samples to determine the training data after data expansion.

[0075] Here, different audio generation algorithms often generate audio with varying degrees of reverberation or metallic sounds, but it is difficult to cover a variety of generated audio in the training data for audio tracing. In addition, some large audio model generation algorithms can generate AI songs online, and some even include synthesized accompaniment. The generated accompaniment can mask the subtle differences in the synthesized songs, which also increases the difficulty of the audio tracing task. The present invention proposes to add noise with different signal-to-noise ratios to the training data online to simulate various possible situations, thereby achieving the purpose of expanding the data. The noise data set can select music and some accompaniment songs, and the signal-to-noise ratio is set according to [5, 10, 15, 20, 25, 30] dB, where the signal-to-noise ratio is calculated as follows:

[0076]

[0077] Among them, in the case of noise, the signal-to-noise ratio (SNR) is often used to measure the noise level of speech. r The signal-to-noise ratio represents the power of the pure speech signal (P s ) and the noise speech signal power (P n ), where i is the sampling point number, s(i) is the clean signal, n(i) is the noise signal, and N is the length of the speech sequence. Then the noisy speech signal y(i) = s(i) + n(i), and N is the total number of sampling points.

[0078] To protect the provenance model from attacks, this paper proposes using adversarial examples to train the provenance model for robustness. Gradient-based adversarial example generation algorithms are very common, including the Fast Gradient Signed Method (FGSM), PGD, and BIM, as well as gradient-optimized adversarial example generation methods like C&W and L-BFGS, and black-box attack methods like JSMA and One Pixel.

[0079] (2): Input the sample effective audio into the neural network model, extract the sample audio high-level features of the sample effective audio based on the self-supervised pre-training method, and perform adversarial training on the sample audio high-level features until the neural network model converges and stops training the neural network model.

[0080] Here, the sample valid audio is input into the neural network model, the sample audio high-level features of the sample valid audio are extracted according to the self-supervised pre-training method, and the sample audio high-level features are subjected to adversarial training until the neural network model converges and the training of the neural network model is stopped.

[0081] In one possible implementation, inputting the sample valid audio into a neural network model, extracting sample audio high-level features of the sample valid audio based on self-supervised pre-training, and performing adversarial training on the sample audio high-level features until the neural network model converges, includes:

[0082] I: Perform convolution pooling processing and channel attention mechanism processing on the sample audio high-level features to determine the sample audio high-level features after the channel attention mechanism processing.

[0083] Here, the high-level audio features are subjected to convolution pooling processing and channel attention mechanism processing to determine the high-level audio features of the sample after channel attention mechanism processing.

[0084] During the model training phase, the high-level features of the sample audio are reduced in feature dimension through convolutional pooling operations, and then each channel of the feature map is regarded as a feature detector through the channel attention mechanism, so that the channel features focus on the effective information in the feature map.

[0085] In one possible implementation, performing convolution pooling processing and channel attention mechanism processing on the sample audio high-level features to determine the sample audio high-level features after the channel attention mechanism processing includes:

[0086] i: performing global maximum pooling processing and global average pooling processing in the spatial dimension on the sample audio high-level features to determine global average pooling features and maximum pooling features.

[0087] Here, global maximum pooling and global average pooling are performed on the high-level features of the sample audio in the spatial dimension to determine the global average pooling features and the maximum pooling features.

[0088] Among them, global maximum pooling and global average pooling are performed on a sample audio high-level feature map F of size H×W×C in the spatial dimension to obtain two 1×1×C feature maps, which are the feature maps corresponding to the global average pooling features and the maximum pooling features respectively.

[0089] ii: Input the global average pooling feature and the maximum pooling feature into the multi-layer perceptron network layer respectively for channel dimension feature learning, and output the global average pooling feature processed by the multi-layer perceptron network layer and the maximum pooling feature processed by the multi-layer perceptron network layer.

[0090] Here, the global average pooling features and the maximum pooling features are respectively input into the multi-layer perceptron network layer for channel dimension feature learning, and the global average pooling features processed by the multi-layer perceptron network layer and the maximum pooling features processed by the multi-layer perceptron network layer are output.

[0091] The global maximum pooling features and global average pooling features are fed into a shared multi-layer perceptron (MLP) for learning. The number of neurons in the first layer of the MLP is C / r, the activation function is ReLU, and the number of neurons in the second layer is C.

[0092] iii: splicing the global average pooling features processed by the multi-layer perceptron network layer and the maximum pooling features processed by the multi-layer perceptron network layer to determine the spliced ​​features; mapping the spliced ​​features based on the activation function to determine the channel attention weight matrix.

[0093] Here, the global average pooling features processed by the multi-layer perceptron network layer and the maximum pooling features processed by the multi-layer perceptron network layer are spliced ​​to determine the spliced ​​features, and the spliced ​​features are mapped according to the activation function to determine the channel attention weight matrix.

[0094] Among them, the result of MLP output is added, and then mapped by Sigmoid activation function to finally obtain the channel attention weight matrix M c .

[0095] M c ∈R C×1×1

[0096] In order to reduce the calculation parameters, a dimension reduction coefficient r is used in MLP. c ∈R C / r×1×1 , where R is the dimension.

[0097] iv: Based on the channel attention weight matrix, the global average pooling feature and the maximum pooling feature, determine the high-level features of the sample audio after the channel attention mechanism is processed.

[0098] Here, the high-level features of the sample audio after the channel attention mechanism is determined by the following formula:

[0099]

[0100] in, is the global average pooling feature, is the maximum pooling feature, W1 and W2 are the weights of MLP, σ(·) is the sigmoid function, M c (F) is the high-level features of the sample audio after processing by the channel attention mechanism.

[0101] II: Processing the sample audio high-level features after the channel attention mechanism based on the first fully connected network layer, outputting the first sample audio high-level features and the predicted tracing label corresponding to the first sample audio high-level features.

[0102] Here, the high-level features of the sample audio processed by the channel attention mechanism are processed according to the first fully connected network layer, and the high-level features of the first sample audio and the predicted tracing label corresponding to the high-level features of the first sample audio are output.

[0103] Among them, the traceability label is the label corresponding to the audio traceability category corresponding to the predicted high-level features of the first sample audio.

[0104] III: Add a gradient reversal network layer after the first fully connected network layer, input the first sample audio high-level features into the gradient reversal network layer for processing, output the processed first sample audio high-level features, input the processed first sample audio high-level features into the second fully connected network layer after the gradient reversal network layer for processing, output the first sample audio high-level features processed by the second fully connected network layer, input the first sample audio high-level features processed by the second fully connected network layer into the third fully connected network layer after the second fully connected network layer, output the second sample audio high-level features and the adversarial label of the second sample audio high-level features.

[0105] Here, a gradient reversal network layer is added after the first fully connected network layer, the first sample audio high-level features are input into the gradient reversal network layer for processing, and the processed first sample audio high-level features are output. The processed first sample audio high-level features are input into the second fully connected network layer after the gradient reversal network layer for processing, and the first sample audio high-level features processed by the second fully connected network layer are output. The first sample audio high-level features processed by the second fully connected network layer are input into the third fully connected network layer after the second fully connected network layer, and the second sample audio high-level features and the adversarial labels of the second sample audio high-level features are output.

[0106] IV: Based on the first sample audio high-level features, the second sample audio high-level features, the predicted tracing label, the adversarial label and the true tracing label of the sample audio data, determine the total loss value of the neural network model until the neural network model converges and stops training the neural network model.

[0107] Here, the total loss value of the neural network model is determined based on the first sample audio high-level features, the second sample audio high-level features, the predicted tracing label, the adversarial label and the true tracing label, and the training of the neural network model is stopped until the neural network model converges.

[0108] In one possible implementation, determining a total loss value of the neural network model based on the first sample audio high-level features, the second sample audio high-level features, the predicted source label, the adversarial label, and the true source label of the sample audio data, and stopping training the neural network model until the neural network model converges, includes:

[0109] Based on the weighted cross entropy loss function, the first sample audio high-level features, the predicted tracing label and the true tracing label are calculated to determine the classification loss value; based on the binary cross entropy loss function, the second sample audio high-level features and the adversarial label are calculated to determine the adversarial loss value; based on the classification loss value and the adversarial loss value, the total loss value of the neural network model is determined, and the training of the neural network model is stopped until the neural network model converges.

[0110] Among them, the first sample audio high-level feature F is output through the first fully connected network layer s , F s The classification loss L is calculated by the weighted cross-entropy (WCE) loss function with the traceability label y and the true traceability label norm A gradient reversal layer (GRL) is added after the first fully connected network layer. The gradient passed to the GRL is multiplied by a negative number λ, so that the training objectives of the network before and after the GRL are opposite. In the initial stage of the model, λ approaches 0. As the model is trained more times, its value approaches -1. Its calculation formula is as follows:

[0111]

[0112] Where C iis the current iteration number, K is the total number of iterations set for model training, and T is the empirical coefficient. After accessing GRL, the fully connected network layer will have two goals to meet. The first is that the fully connected network layer needs to generate features that can predict the correct label, and the second is that the features extracted by the fully connected layer need to be as difficult to determine which task domain they come from as possible. The two goals compete with each other. When calculating the adversarial loss, the features after the GRL layer will pass through two fully connected network layers in succession, and then the output will be compared with the adversarial label. (All labels are 0 or all labels are 1) The adversarial loss L is calculated through the binary cross entropy loss (BCE) function adv , and finally the classification loss L norm and adversarial loss L adv The loss function of the model obtained by weighted average is:

[0113] L=L norm +α×L adv

[0114] Where α is the adjustment factor used to adjust the impact of adversarial loss on model loss, and L is the total loss value.

[0115] Here, by adding a GRL layer to the model, during the model training process, GRL can achieve an identity transformation in forward propagation. During the backward propagation process, the gradient uses a negative coefficient and propagates to the upper layer. This adds constraints to the model training, allowing the trained speech tracing forensics model to better identify unknown tracing algorithms.

[0116] c: averaging the model parameters of each layer of a preset number of neural network models with the best performance during the model training process, and constructing the speech tracing forensics model based on the averaged model parameters of each layer.

[0117] During model training, the ten best-performing model parameters on the validation set are saved. Because traceability performance varies among different model parameters, to improve the stability of the traceability model, the parameters of each layer of the top ten models are averaged to produce the final voice traceability model. During the operational phase, test audio is input into the voice traceability model to obtain scores for each audio traceability category. The softmax function is then applied to obtain the probability distribution for each category. The category corresponding to the column with the maximum value is then identified as the audio traceability category, compared to the training labels.

[0118] In order to solve the problem that the feature input of the existing tracing algorithm cannot meet the large-model audio generation tracing task, the present invention proposes to use an end-to-end network structure to extract high-level representation of audio from the original audio sequence, which simplifies the tracing task and can more effectively extract feature parameters for the tracing task; in order to solve the problem that the amount of data of some tracing types is small, it is proposed to use adversarial sample methods to expand training data. Compared with the previous method of adding noise to expand data, this method can more effectively improve the generalization and robustness of the model; to solve the problem that the existing tracing model has poor generalization for the same generation method and its variants, the present invention adjusts the traditional tracing model structure, proposes a self-supervised model method using large-scale data pre-training, and adds a channel attention mechanism to the model to improve the model's learning ability for fine-grained audio features; in addition, a gradient reversal layer is introduced during the model training process to further improve the robustness of the speech tracing forensics model;

[0119] In the specific embodiment, step 1: the original audio signal is processed through a unified codec and a unified frequency format; step 2: various noises are randomly added to the audio, and different signal-to-noise ratios need to be controlled; step 3: different adversarial sample generation methods are used to construct an adversarial sample data set; the methods include gradient-based FGSM, PGD, BIM and optimization-based adversarial sample generation methods C&W, L-BFGS and black box attack methods JSMA, One Pixel, etc.; step 4: all data are passed through a [50Hz, 6000Hz] bandpass filter to retain the effective information of the audio; step 5: the filtered audio is input into the end-to-end model, the high-level representation of the audio is extracted and adversarial training is performed until the model converges; step 6: the Top 10 model parameters of the model are averaged to obtain a speech tracing forensics model; step 7: the test audio is input into the speech tracing forensics model to obtain the scores of each tracing category, and the probability distribution of each category is obtained through the Softmax function. The category corresponding to the column where the maximum value is located is the audio tracing category according to the training label.

[0120] Here, when training the model, it is proposed to use a combination of adversarial samples and noise to expand the traceability data, and then use the audio preprocessing module to remove low-frequency and high-frequency interference (audio signals in this frequency band often have a negative effect on audio tracing tasks), and input it into the end-to-end traceability network. The intermediate nodes of the large speech model based on self-supervised pre-training are used as high-level representations of audio features. During the training process, the gradient reversal layer (GRL) is used to simulate the generative adversarial mechanism and enhance the generalization of the model during training. Finally, the Top 10 models of the traceability model are parameter-weighted fused. The present invention solves the problems of small traceability data samples, slow traceability tasks, poor results, and poor generalization performance in the context of large model generation.

[0121] An embodiment of the present application provides a speech source forensics method, which includes: obtaining a speech to be detected, processing the speech to be detected based on an audio bandpass filter, and determining the effective audio of the speech to be detected; inputting the effective audio into a pre-trained speech source forensics model, extracting the audio high-level features of the effective audio, and convolving the audio high-level features based on the convolutional pooling network layer of the speech source forensics model to determine the audio high-level features after the convolution processing; wherein, the speech source forensics model is determined by adding a channel attention mechanism and a gradient reversal network layer to a neural network model to simulate generative adversarial training; processing the audio high-level features after the convolution processing based on the channel attention mechanism and the fully connected network layer of the speech source forensics model to determine the target high-level features, calculating the probability distribution of each tracing category corresponding to the target high-level features based on the activation function, and predicting the speech source forensics result of the speech to be detected; wherein, source forensics refers to verifying the generation source or generation method of the generated speech. By adding a channel attention mechanism to the speech tracing forensics model, the model's ability to learn fine-grained audio features is improved, and the gradient reversal network layer is introduced during the training process to improve the robustness of the speech tracing forensics model, so that the speech tracing forensics model can quickly and accurately determine the speech tracing forensics results corresponding to the speech.

[0122] See also Figure 2 、 Figure 3 , Figure 2 This is one of the structural diagrams of a voice tracing and evidence collection device provided in an embodiment of the present application; Figure 3 This is a second structural diagram of a voice tracing and evidence collection device provided in an embodiment of the present application. Figure 2 As shown in , the voice tracing and evidence collection device 200 includes:

[0123] The filtering processing module 210 is used to obtain the speech to be detected, process the speech to be detected based on the audio bandpass filter, and determine the valid audio of the speech to be detected;

[0124] A first processing module 220 is configured to input the valid audio into a pre-trained speech forensics model, extract high-level audio features of the valid audio, perform convolution processing on the high-level audio features based on a convolutional pooling network layer of the speech forensics model, and determine the high-level audio features after convolution processing; wherein the speech forensics model is determined by adding a channel attention mechanism and a gradient reversal network layer to a neural network model to simulate generative adversarial training;

[0125] The second processing module 230 is used to process the audio high-level features after convolution processing based on the channel attention mechanism and the fully connected network layer of the speech tracing forensics model, determine the target high-level features, calculate the probability distribution of each tracing category corresponding to the target high-level features based on the activation function, and predict the speech tracing forensics result of the speech to be detected; wherein, tracing forensics refers to verifying the generation source or generation method of the generated speech.

[0126] Furthermore, when the second processing module 230 processes the high-level audio features after convolution processing using the channel attention mechanism and the fully connected network layer based on the speech tracing forensics model, determines the target high-level features, calculates the probability distribution of each tracing category corresponding to the target high-level features based on the activation function, and predicts the speech tracing forensics result of the speech to be detected, the second processing module 230 is specifically used to:

[0127] Performing attention processing on the high-level audio features after the convolution processing based on a channel attention mechanism to determine the high-level audio features after the attention processing;

[0128] Processing the audio high-level features after the attention processing based on the fully connected network layer to determine the target high-level features;

[0129] Based on the activation function, the probability distribution of each tracing category corresponding to the target high-level feature is calculated to determine multiple probability values, and the tracing category corresponding to the maximum probability value is determined as the voice tracing forensics result of the voice to be detected.

[0130] Further, such as Figure 3 As shown, the voice tracing forensics model 200 further includes a model training module 240, which trains the voice tracing forensics model through the following steps:

[0131] Screening out sample audio data from the training data after data expansion, performing filtering on the sample audio data, and determining sample valid audio of the sample audio data;

[0132] Inputting the sample valid audio into a neural network model, extracting sample audio high-level features of the sample valid audio based on a self-supervised pre-training method, and performing adversarial training on the sample audio high-level features until the neural network model converges, and stopping the training of the neural network model;

[0133] The model parameters of each layer of a preset number of neural network models with the best performance during the model training process are averaged, and the voice tracing forensics model is constructed based on the averaged model parameters of each layer.

[0134] Further, such as Figure 3 As shown, the voice tracing evidence model 200 further includes a data expansion module 250, which is used to:

[0135] Add noise with different signal-to-noise ratios to the training data to determine the training data after data expansion;

[0136] Alternatively, a gradient-based adversarial example generation algorithm generates different adversarial examples to determine the training data after data augmentation.

[0137] Furthermore, when the model training module 240 is used to input the sample valid audio into the neural network model, extract the sample audio high-level features of the sample valid audio based on the self-supervised pre-training method, and perform adversarial training on the sample audio high-level features until the neural network model converges and stops training the neural network model, the model training module 240 is specifically used to:

[0138] Performing convolution pooling processing and channel attention mechanism processing on the sample audio high-level features to determine the sample audio high-level features after the channel attention mechanism processing;

[0139] Processing the sample audio high-level features after the channel attention mechanism based on the first fully connected network layer, and outputting the first sample audio high-level features and the predicted tracing label corresponding to the first sample audio high-level features;

[0140] Add a gradient reversal network layer after the first fully connected network layer, input the first sample audio high-level feature into the gradient reversal network layer for processing, output the processed first sample audio high-level feature, input the processed first sample audio high-level feature into the second fully connected network layer after the gradient reversal network layer for processing, output the first sample audio high-level feature processed by the second fully connected network layer, input the first sample audio high-level feature processed by the second fully connected network layer into the third fully connected network layer after the second fully connected network layer, output the second sample audio high-level feature and the adversarial label of the second sample audio high-level feature;

[0141] Based on the first sample audio high-level features, the second sample audio high-level features, the predicted tracing label, the adversarial label and the true tracing label of the sample audio data, the total loss value of the neural network model is determined, and the training of the neural network model is stopped until the neural network model converges.

[0142] Furthermore, when the model training module 240 is used to perform convolution pooling processing and channel attention mechanism processing on the sample audio high-level features and determine the sample audio high-level features after the channel attention mechanism processing, the model training module 240 is specifically used to:

[0143] Performing global maximum pooling processing and global average pooling processing on the sample audio high-level features in a spatial dimension to determine global average pooling features and maximum pooling features;

[0144] Inputting the global average pooling feature and the maximum pooling feature into the multi-layer perceptron network layer respectively for channel dimension feature learning, and outputting the global average pooling feature and the maximum pooling feature processed by the multi-layer perceptron network layer;

[0145] Splicing the global average pooling feature processed by the multi-layer perceptron network layer and the maximum pooling feature processed by the multi-layer perceptron network layer to determine the spliced ​​feature;

[0146] Mapping the concatenated features based on an activation function to determine a channel attention weight matrix;

[0147] Based on the channel attention weight matrix, the global average pooling features and the maximum pooling features, the high-level features of the sample audio after processing by the channel attention mechanism are determined.

[0148] Furthermore, when the model training module 240 is used to determine the total loss value of the neural network model based on the first sample audio high-level features, the second sample audio high-level features, the predicted traceability label, the adversarial label, and the true traceability label of the sample audio data, and stops training the neural network model when the neural network model converges, the model training module 240 is specifically used to:

[0149] Calculating the first sample audio high-level features, the predicted source tracing label, and the true source tracing label based on a weighted cross entropy loss function to determine a classification loss value;

[0150] Calculating the second sample audio high-level features and the adversarial label based on a binary cross entropy loss function to determine an adversarial loss value;

[0151] Based on the classification loss value and the adversarial loss value, a total loss value of the neural network model is determined, and training of the neural network model is stopped until the neural network model converges.

[0152] An embodiment of the present application provides a speech source forensics device, which includes: a filtering processing module for obtaining a speech to be detected, processing the speech to be detected based on an audio bandpass filter, and determining the effective audio of the speech to be detected; a first processing module for inputting the effective audio into a pre-trained speech source forensics model, extracting high-level audio features of the effective audio, convolving the high-level audio features based on a convolutional pooling network layer of the speech source forensics model, and determining the high-level audio features after the convolution processing; wherein the speech source forensics model is determined by adding a channel attention mechanism and a gradient reversal network layer to a neural network model to simulate generative adversarial training; a second processing module for processing the high-level audio features after the convolution processing based on the channel attention mechanism and the fully connected network layer of the speech source forensics model, determining the target high-level features, calculating the probability distribution of each tracing category corresponding to the target high-level features based on an activation function, and predicting the speech source forensics result of the speech to be detected; wherein source forensics refers to verifying the generation source or generation method of the generated speech. By adding a channel attention mechanism to the speech tracing forensics model, the model's ability to learn fine-grained audio features is improved, and the gradient reversal network layer is introduced during the training process to improve the robustness of the speech tracing forensics model, so that the speech tracing forensics model can quickly and accurately determine the speech tracing forensics results corresponding to the speech.

[0153] See also Figure 4 , Figure 4 This is a schematic diagram of the structure of an electronic device provided in an embodiment of the present application. Figure 4 As shown in FIG, the electronic device 400 includes a processor 410 , a memory 420 and a bus 430 .

[0154] The memory 420 stores machine-readable instructions executable by the processor 410. When the electronic device 400 is running, the processor 410 communicates with the memory 420 via the bus 430. When the machine-readable instructions are executed by the processor 410, the above-mentioned Figure 1 The steps of the voice tracing and evidence collection method in the method embodiment shown are specifically implemented in accordance with the method embodiment and will not be described in detail here.

[0155] The embodiment of the present application also provides a computer-readable storage medium on which a computer program is stored. When the computer program is executed by a processor, the computer program can execute the above-mentioned Figure 1 The steps of the voice tracing and evidence collection method in the method embodiment shown are specifically implemented in accordance with the method embodiment and will not be described in detail here.

[0156] Those skilled in the art will clearly understand that, for the convenience and brevity of description, the specific working processes of the systems, devices and units described above can refer to the corresponding processes in the aforementioned method embodiments and will not be repeated here.

[0157] In the several embodiments provided in this application, it should be understood that the disclosed systems, devices and methods can be implemented in other ways. The device embodiments described above are merely schematic. For example, the division of the units is merely a logical function division. There may be other division methods in actual implementation. For example, multiple units or components may be combined or integrated into another system, or some features may be ignored or not executed. Another point is that the mutual coupling or direct coupling or communication connection shown or discussed may be through some communication interface, indirect coupling or communication connection of devices or units, which may be electrical, mechanical or other forms.

[0158] The units described as separate components may or may not be physically separate, and the components shown as units may or may not be physical units, that is, they may be located in one place or distributed across multiple network units. Some or all of these units may be selected to achieve the purpose of this embodiment according to actual needs.

[0159] In addition, each functional unit in each embodiment of the present application may be integrated into one processing unit, or each unit may exist physically separately, or two or more units may be integrated into one unit.

[0160] If the functions are implemented in the form of software functional units and sold or used as independent products, they can be stored in a non-volatile computer-readable storage medium that is executable by a processor. Based on this understanding, the technical solution of the present application, or the part that contributes to the prior art, or the part of the technical solution, can be embodied in the form of a software product. The computer software product is stored in a storage medium and includes several instructions for enabling a computer device (which can be a personal computer, a server, or a network device, etc.) to execute all or part of the steps of the method described in each embodiment of the present application. The aforementioned storage medium includes various media that can store program codes, such as a USB flash drive, a mobile hard disk, a read-only memory (ROM), a random access memory (RAM), a magnetic disk or an optical disk.

[0161] Finally, it should be noted that the above-described embodiments are only specific implementation methods of the present application, which are used to illustrate the technical solutions of the present application, rather than to limit them. The scope of protection of the present application is not limited thereto. Although the present application has been described in detail with reference to the above-mentioned embodiments, those skilled in the art should understand that any person skilled in the art can modify or easily conceive of changes to the technical solutions described in the above-mentioned embodiments within the technical scope disclosed in the present application, or perform equivalent replacements for some of the technical features thereof. These modifications, changes, or replacements do not deviate from the spirit and scope of the technical solutions of the embodiments of the present application, and should be included in the scope of protection of the present application. Therefore, the scope of protection of the present application should be based on the scope of protection of the claims.

Claims

1. A voice tracing and evidence collection method, characterized in that: The voice source tracing and evidence collection method includes: Acquire a speech to be detected, process the speech to be detected based on an audio bandpass filter, and determine a valid audio of the speech to be detected; Inputting the valid audio into a pre-trained speech forensics model, extracting high-level audio features of the valid audio, and performing convolution processing on the high-level audio features based on the convolutional pooling network layer of the speech forensics model to determine the high-level audio features after convolution processing; wherein the speech forensics model is determined by adding a channel attention mechanism and a gradient reversal network layer to the neural network model to simulate generative adversarial training; Based on the channel attention mechanism and the fully connected network layer of the speech tracing forensics model, the high-level audio features after convolution processing are processed to determine the target high-level features, and the probability distribution of each tracing category corresponding to the target high-level features is calculated based on the activation function to predict the speech tracing forensics result of the speech to be detected; wherein, tracing forensics refers to verifying the generation source or generation method of the generated speech.

2. The voice source tracing and evidence collection method according to claim 1 is characterized in that: The channel attention mechanism based on the speech tracing forensics model and the fully connected network layer processes the high-level audio features after convolution processing, determines the target high-level features, calculates the probability distribution of each tracing category corresponding to the target high-level features based on the activation function, and predicts the speech tracing forensics result of the speech to be detected, including: Performing attention processing on the high-level audio features after the convolution processing based on a channel attention mechanism to determine the high-level audio features after the attention processing; Processing the audio high-level features after the attention processing based on the fully connected network layer to determine the target high-level features; Based on the activation function, the probability distribution of each tracing category corresponding to the target high-level feature is calculated to determine multiple probability values, and the tracing category corresponding to the maximum probability value is determined as the voice tracing forensics result of the voice to be detected.

3. The voice source tracing and evidence collection method according to claim 1 is characterized in that: The speech tracing forensics model is trained by the following steps: Screening out sample audio data from the training data after data expansion, performing filtering on the sample audio data, and determining sample valid audio of the sample audio data; Inputting the sample valid audio into a neural network model, extracting sample audio high-level features of the sample valid audio based on a self-supervised pre-training method, and performing adversarial training on the sample audio high-level features until the neural network model converges, and stopping the training of the neural network model; The model parameters of each layer of a preset number of neural network models with the best performance during the model training process are averaged, and the voice tracing forensics model is constructed based on the averaged model parameters of each layer.

4. The voice tracing and evidence collection method according to claim 2 is characterized in that: The training data after data expansion is determined in the following way: Add noise with different signal-to-noise ratios to the training data to determine the training data after data expansion; Alternatively, a gradient-based adversarial example generation algorithm generates different adversarial examples to determine the training data after data augmentation.

5. The voice source tracing and evidence collection method according to claim 3 is characterized in that: The step of inputting the sample valid audio into the neural network model, extracting sample audio high-level features of the sample valid audio based on a self-supervised pre-training method, and performing adversarial training on the sample audio high-level features until the neural network model converges, and stopping the training of the neural network model, includes: Performing convolution pooling processing and channel attention mechanism processing on the sample audio high-level features to determine the sample audio high-level features after the channel attention mechanism processing; Processing the sample audio high-level features after the channel attention mechanism based on the first fully connected network layer, and outputting the first sample audio high-level features and the predicted tracing label corresponding to the first sample audio high-level features; Add a gradient reversal network layer after the first fully connected network layer, input the first sample audio high-level feature into the gradient reversal network layer for processing, output the processed first sample audio high-level feature, input the processed first sample audio high-level feature into the second fully connected network layer after the gradient reversal network layer for processing, output the first sample audio high-level feature processed by the second fully connected network layer, input the first sample audio high-level feature processed by the second fully connected network layer into the third fully connected network layer after the second fully connected network layer, output the second sample audio high-level feature and the adversarial label of the second sample audio high-level feature; Based on the first sample audio high-level features, the second sample audio high-level features, the predicted tracing label, the adversarial label and the true tracing label of the sample audio data, the total loss value of the neural network model is determined, and the training of the neural network model is stopped until the neural network model converges.

6. The voice source tracing and evidence collection method according to claim 5 is characterized in that: The performing convolution pooling processing and channel attention mechanism processing on the sample audio high-level features to determine the sample audio high-level features after the channel attention mechanism processing includes: Performing global maximum pooling processing and global average pooling processing on the sample audio high-level features in a spatial dimension to determine global average pooling features and maximum pooling features; Inputting the global average pooling feature and the maximum pooling feature into the multi-layer perceptron network layer respectively for channel dimension feature learning, and outputting the global average pooling feature and the maximum pooling feature processed by the multi-layer perceptron network layer; Splicing the global average pooling feature processed by the multi-layer perceptron network layer and the maximum pooling feature processed by the multi-layer perceptron network layer to determine the spliced ​​feature; Mapping the concatenated features based on an activation function to determine a channel attention weight matrix; Based on the channel attention weight matrix, the global average pooling features and the maximum pooling features, the high-level features of the sample audio after processing by the channel attention mechanism are determined.

7. The voice source tracing and evidence collection method according to claim 5 is characterized in that: The method further comprises determining a total loss value of the neural network model based on the first sample audio high-level feature, the second sample audio high-level feature, the predicted source tracing label, the adversarial label, and the true source tracing label of the sample audio data, and stopping training the neural network model until the neural network model converges, including: Calculating the first sample audio high-level features, the predicted source tracing label, and the true source tracing label based on a weighted cross entropy loss function to determine a classification loss value; Calculating the second sample audio high-level features and the adversarial label based on a binary cross entropy loss function to determine an adversarial loss value; Based on the classification loss value and the adversarial loss value, a total loss value of the neural network model is determined, and training of the neural network model is stopped until the neural network model converges.

8. A voice tracing and evidence collection device, characterized in that: The voice source tracing and evidence collection device comprises: A filtering processing module is used to obtain a speech to be detected, process the speech to be detected based on an audio bandpass filter, and determine a valid audio of the speech to be detected; A first processing module is configured to input the valid audio into a pre-trained speech forensics model, extract high-level audio features of the valid audio, perform convolution processing on the high-level audio features based on a convolutional pooling network layer of the speech forensics model, and determine the high-level audio features after convolution processing; wherein the speech forensics model is determined by adding a channel attention mechanism and a gradient reversal network layer to a neural network model to simulate generative adversarial training; The second processing module is used to process the high-level audio features after convolution processing based on the channel attention mechanism and the fully connected network layer of the speech tracing forensics model, determine the target high-level features, calculate the probability distribution of each tracing category corresponding to the target high-level features based on the activation function, and predict the speech tracing forensics results of the speech to be detected; wherein, tracing forensics refers to verifying the generation source or generation method of the generated speech.

9. An electronic device, characterized in that: include: A processor, a memory and a bus, wherein the memory stores machine-readable instructions executable by the processor. When the electronic device is running, the processor and the memory communicate through the bus. When the processor is running, the machine-readable instructions execute the steps of the voice tracing and forensics method as described in any one of claims 1 to 7.

10. A computer-readable storage medium, characterized in that The computer-readable storage medium stores a computer program, and when the computer program is executed by a processor, the steps of the voice source tracing and evidence collection method according to any one of claims 1 to 7 are executed.

Citation Information

Patent Citations

  • Voice traceability evidence obtaining method and device, equipment and storage medium

    CN115083422A

  • Speaker-related speech synthesis attack prevention method and system

    CN115910022A

  • Intelligent voice forgery attack detection method based on attention mechanism

    CN116416997A

  • Fake voice detection method based on dual-track differential modeling

    US20250095669A1

  • Method and apparatus for speech endpoint detection based on neural network, device, and medium

    WO2021208728A1