Deep forged audio detection protection method based on large-scale pre-training model Whisper

By using a large-scale pre-trained model Whisper and LoRA fine-tuning strategy based on Transformer architecture, combined with audio transcription text as prompt information, the problem of deep fake audio detection is solved, and high accuracy and low cost fake audio detection effect is achieved.

CN120126481AActive Publication Date: 2025-06-10YANGTZE DELTA REGION INST OF TSINGHUA UNIV ZHEJIANG
View PDF 10 Cites 0 Cited by

Patent Information

Application Number
CN202510297989.5
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-03-13
Publication Date
2025-06-10
Estimated Expiration
2045-03-13

AI Technical Summary

Technical Problem

It is difficult for the prior art to effectively detect and protect deep-forged audio, especially in deep-forged audio generated based on codec systems, there are problems such as the number of fake audios, enhanced fraudulent ability and enhanced audio adaptability.

Method used

Using a large-scale pre-trained model Whisper based on the Transformer architecture, through transfer learning and innovative fine-tuning strategies, the complete transcription text of the audio is embedded into the decoder as prompt information, and combined with the LoRA fine-tuning strategy, it reduces computing costs and resource consumption.

Benefits of technology

It realizes high accuracy detection of deep fake audio, breaks through the input length limit of the 30-second audio clip of the Whisper model, significantly reduces computing cost and resource consumption, and has stronger robustness and generalization capabilities.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120126481A_ABST
    Figure CN120126481A_ABST
Patent Text Reader

Abstract

The invention discloses a deep forged audio detection protection method based on a large-scale pre-training model Whisper, and the method comprises the steps: S1, inputting an audio to be detected into the pre-training model Whisper, and employing a transcription text of the whole audio as a prompt in the fine tuning process of each audio segment, thereby enabling the model Whisper to use global text information; s2, in the training and decoding process, preprocessing the audio data to adapt to a model Whisper; and S3, evaluating the performance of the model Whisper in the training process by designing a cross entropy loss detection function. According to the deep forged audio detection protection method based on the large-scale pre-training model Whisper, audio authenticity identification is realized through transfer learning. Different from an existing voice classification or recognition task, a complete transcription text of the audio is used as prompt information to be embedded into a decoder through an innovative fine tuning strategy, and the problem that traditional acoustic features are single in dependence is solved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention belongs to the technical field of forged audio detection and protection, and specifically relates to a method for detecting and protecting deepfake audio based on the large-scale pre-trained model Whisper. Background Art

[0002] With the rapid development of deep learning technology, its application in the field of machine audition has become increasingly widespread. Many efficient and accurate models have emerged continuously, surpassing traditional technologies and greatly improving the production and creation efficiency of media content. Especially in the field of audio synthesis, such as text-to-speech (TTS) and voice conversion (VC), these technologies can generate highly realistic target speaker voices and even deceive human ears and human-computer interaction devices.

[0003] However, the abuse of audio deepfake technology brings potential risks that cannot be ignored. The potential harm objects of deepfake audio include the human auditory system and the machine auditory system, which pose a serious threat to the people's livelihood, economy, politics and social security around the world. Therefore, there is an urgent need to develop an automated and highly accurate method for detecting deepfake audio to effectively address this threat. Summary of the Invention

[0004] The main purpose of the present invention is to provide a method for detecting and protecting deepfake audio based on the large-scale pre-trained model Whisper. Based on the large-scale pre-trained model Whisper of the Transformer architecture, audio authenticity identification is achieved through transfer learning. Different from existing speech classification or recognition tasks (such as tasks based on Hubert, Wav2vec, etc.), the present invention not only utilizes the powerful multilingual speech recognition ability of the Whisper model, but also through an innovative fine-tuning strategy, embeds the complete transcribed text of the audio as a prompt into the decoder, overcoming the problem of single dependence on traditional acoustic features.

[0005] To achieve the above object, the present invention provides a method for detecting and protecting deepfake audio based on the large-scale pre-trained model Whisper, including the following steps:

[0006] Step S1: Input the audio to be detected into the pre-trained model Whisper (hereinafter referred to as the model Whisper), and use the transcribed text of the entire audio as a prompt during the fine-tuning process of each audio segment, so that the model Whisper can utilize the global text information;

[0007] Step S2: During the training and decoding processes (the model Whisper requires that the audio be padded or trimmed into segments not exceeding a preset time (30 seconds). For long-form transcription tasks, Whisper provides a continuous transcription interface that utilizes the context information of previous segments. However, during the fine-tuning process of transfer learning, the audio segments must still meet the 30-second limit), preprocess the audio data to adapt it to the model Whisper;

[0008] Step S3: In terms of the design of the loss function, based on the default cross-entropy loss function, evaluate the performance of the model Whisper during the training process by designing a detection cross-entropy loss function (paying particular attention to the accuracy of the predicted content between specific tokens identified as <|transcribe|> and <|endoftext|>).

[0009] As a further preferred technical solution of the above technical solution, in Step S1:

[0010] Define the audio as x and the full text transcribed by the model Whisper as transcript x , in this framework, the prediction process of the model Whisper for each segment is expressed as:

[0011] P(y|x,transcript x )→result;

[0012] where y represents the input token sequence (token sequence) of the decoder in the model, and result is the classification result of the model Whisper. During the training process, the prompt of the transcribed text should be used as a reference rather than the target of prediction. Therefore, in the input and target of the decoder, from | <startofprev>|The position of the prompt marking the start remains unchanged, and the model Whisper directly outputs the same content based on the input. When encountering| <startoftranscript>|Marker, start predicting subsequent content;

[0013] (This invention plans to use the officially provided Whisper-large model to transcribe audio to obtain its corresponding text information) During the testing phase, not only the prefix needs to be provided, but also the complete text needs to be provided as a prompt to ensure that the Whisper model refers to the complete text information when generating the output.

[0014] As a further preferred technical solution of the above technical solution, in step S2:

[0015] For long audio exceeding the preset time, use the Whisper model to split it into smaller decoding segments and record the end time of each segment (such as T 0 , T 1 , …, T k ), then, adopt the greedy merging strategy to merge the decoding segments into merged segments in sequence to make their lengths (as much as possible) close to but not exceed the preset time, and the remaining unmergeable decoding segments are discarded;

[0016] For short audio not exceeding the preset time, it is extended to the specified length by repeating and splicing to meet the input requirements of the Whisper model; finally, the preprocessed audio (segment) data is used for the training and testing of the Whisper model.

[0017] As a further preferred technical solution of the above technical solution, in step S3, the loss calculation is carried out in steps to ensure that the Whisper model is punished for prediction errors in the key segments of the sequence, as shown below:

[0018] The total loss L is a weighted combination of the global loss L global and the local intermediate loss L inter , the global loss L global covers the entire sequence and is expressed as;

[0019] L global = ∑ i,j L CE (y i,j , y′ i,j );

[0020] where, L CE represents the cross-entropy loss function, y is the predicted output of the model, y′ is the true label, and y i,j and y′ i,j represent the predicted value and the true value at the position (i, j) respectively; then, assume that the position of the i-th <|transcribe|> marker is T i and the position of the i-th <|endoftext|> token is E i , the local intermediate loss L inter focuses on the subsequence between <|transcribe|> and <|endoftext|>, that is, the part of the prediction result of audio authenticity, expressed as:

[0021]

[0022] The final loss value is calculated as L = L global + α * L inter , where α is a weight factor used to increase the influence of the local intermediate loss (which means that during training, the model will not only try to minimize the errors in the entire sequence (by minimizing L global ), but also pay special attention to reducing the errors in the subsequence between the <|transcribe|> and <|endoftext|> tokens (by minimizing L inter and giving it a higher weight)).

[0023] As a further preferred technical solution of the above technical solution, during the training stage, the audio sample to be detected will first reach a random decision point, and randomly determine whether to perform speech enhancement processing on it, so as to improve the diversity of the audio, enabling the model Whisper to be trained in a more realistic environment (thereby enhancing the robustness of the model and laying a foundation for subsequent accurate detection);

[0024] Subsequently, the processed (or unprocessed) audio data is input into the model Whisper. At this time, the encoder part of the model Whisper has been frozen to ensure the stability of its pre-trained parameters during training and guarantee the robust extraction of the features of the audio data by the model Whisper; then, through the designed loss function, the loss value is calculated based on the prediction result and the true label for updating the parameters of the model Whisper. BRIEF DESCRIPTION OF THE DRAWINGS

[0025] Figure 1 is a schematic structural diagram of the model Whisper of the present invention.

[0026] Figure 2 is a schematic diagram of the transfer learning training process applied to the model Whisper of the present invention. DETAILED DESCRIPTION OF THE INVENTION

[0027] The following description is used to disclose the present invention so that those skilled in the art can implement the present invention. The preferred embodiments in the following description are only examples, and those skilled in the art can think of other obvious variations. The basic principles of the present invention defined in the following description can be applied to other embodiments, variations, improvements, equivalent solutions, and other technical solutions without departing from the spirit and scope of the present invention.

[0028] In the preferred embodiments of the present invention, those skilled in the art should note that large-scale pre-trained models such as Whisper involved in the present invention can be regarded as prior art.

[0029] The codec-based speech generation framework consists of three core modules: (1) Audio Codec: encodes speech features into discrete tokens. (2) Large Language Model (LLM): maps input text tokens to audio tokens. (3) Vocoder: reconstructs the audio waveform from audio tokens. During training, the LLM realizes speech generation by learning the relationship between text embeddings and audio tokens, while the audio codec undertakes the key task of retaining key information in the audio through compression and quantization. A general audio codec usually consists of an encoder, a codec, and a decoder, which are responsible for audio feature extraction, compression, and restoration respectively. In the inference stage, the LLM maps the text input to audio tokens, and the vocoder converts them into high-quality speech. This LLM-based speech synthesis architecture greatly improves the authenticity and naturalness of the synthesized speech.

[0030] However, it is still uncertain whether existing anti-spoofing technologies can effectively deal with deepfake audio based on codec systems. Recent research shows that anti-spoofing models trained based on vocoders often fail to detect deepfake audio generated by these advanced codec systems. The reason is that the unique properties of codec-based deepfake audio exacerbate these challenges, bringing three key obstacles to the current Audio Deepfake Detection (ADD) system:

[0031] (1) The surging number of forged audios: The rapid development of codec systems has significantly reduced the technical threshold of speech forgery, making it more convenient to generate forged audios on a large scale.

[0032] (2) Stronger spoofing ability: Codec-based deepfake audio exhibits higher authenticity and subtlety, making it more difficult to distinguish from real audio.

[0033] (3) Enhancement of audio adaptability: Emerging systems demonstrate extremely strong adaptability in different languages, environments, and audio types, further increasing the detection difficulty.

[0034] In addition, there has been no systematic research report on the identification of deepfake audio generated by decoder systems so far, which sounds the alarm for the limitations of existing anti-forgery technologies. Therefore, it is urgent and important to develop a new anti-deception system with higher generalization ability and cross-type detection ability.

[0035] To address the above issues, the present invention intends to propose an innovative detection scheme, the core of which includes:

[0036] (1) A transfer learning strategy based on the large-scale pre-trained model Whisper, enabling the model to efficiently adapt to different types of deepfake voices;

[0037] (2) Introducing additional pre-trained audio features and combining a multi-audio feature fusion mechanism to capture subtle differences in forged audio.

[0038] The Whisper model has been trained on 680,000 hours of supervised audio data and can perform tasks such as multilingual speech recognition, speech translation, and speech activity monitoring. And it has achieved excellent results on multiple public datasets. In addition, studies have found that the audio features provided by the Whisper model are highly robust and can well adapt to tasks such as audio classification. Moreover, many researchers have enabled the model to be effectively applied to tasks such as speech understanding through the method of transfer learning with Whisper. This provides the possibility to utilize the advantages of Whisper in various applications in the field of speech processing.

[0039] Transfer learning is a technique that can transfer knowledge learned in one domain to another domain and is mostly used to adapt pre-trained models to different downstream tasks. Since the dataset for the task of identifying the authenticity of audio is limited, it is very valuable to use a model trained on a large-scale dataset for other tasks and fine-tune and transfer learn it. For example, many previous studies have explored using unsupervised pre-trained models such as Wav2vec and Hubert to obtain audio features and then input these audio features into downstream models for specific tasks. However, these methods require adding an additional classifier after the pre-trained model to address the difference in the shape of audio features in different domains. In contrast, when using a large-scale model with a supervised learning method such as Whisper for transfer learning, the labels can be directly converted into text and used as the target in the fine-tuning stage, thus effectively eliminating the need for additional layers and classifiers.

[0040] Preferred embodiment.

[0041] The present invention discloses a deepfake audio detection and protection method based on the large-scale pre-trained model Whisper, comprising the following steps:

[0042] Step S1: (As Figure 1 shown, the research solution of the present invention selects to freeze the encoder layer of Whisper while fine-tuning the decoder layer of Whisper. The target of each segment is a text sequence, which contains the prefix "CodecFake Task:", followed by the labels "spoofing" and "genuine", where each label corresponds to a different single token. During the test phase, the present invention still uses "CodecFake Task:" as the decoding prefix to predict the subsequent labels. In order to effectively utilize the global speech information of the entire audio, it is necessary to extract the transcription text of the entire audio as a prompt and add the prompt information to the training and testing processes, so as to prove the effectiveness of the transcription prompt method. Although it is feasible to use only audio as the input data for the Whisper model, it essentially limits the audio within a narrow window of up to 30 seconds. To solve this limitation, this research proposes an innovative method) input the audio to be detected into the pre-trained model Whisper (hereinafter referred to as the model Whisper), and use the transcription text of the entire audio as a prompt during the fine-tuning process of each audio segment, so that the model Whisper can utilize the global text information;

[0043] Step S2: During the training and decoding processes (the model Whisper requires the audio to be padded or cropped into segments not exceeding a preset time (30 seconds). For long-form transcription tasks, Whisper provides a continuous transcription interface that utilizes the context information of the previous segment. However, during the fine-tuning process of transfer learning, the audio segments must still meet the 30-second limit), preprocess the audio data to adapt to the model Whisper;

[0044] Step S3: In terms of the design of the loss function, based on the default cross-entropy loss function, design a detection cross-entropy loss function to evaluate the performance of the model Whisper during the training process (paying special attention to the accuracy of the predicted content between the specific tokens recognized as <|transcribe|> and <|endoftext|>).

[0045] Specifically, in Step S1:

[0046] Define the audio as x and the full text transcribed by the model Whisper as transcript x , in this framework, the prediction process of the model Whisper for each segment is expressed as:

[0047] P(y|x,transcript x ) → result;

[0048] where y represents the input token sequence of the decoder in the model, and result is the classification result of the model Whisper. During the training process, the prompt of the transcribed text should be used as a reference rather than the target of prediction. Therefore, in the input and target of the decoder, from | <startofprev>|The position of the prompt indicating the start of the mark remains unchanged, and the model Whisper directly outputs the same content based on the input. When encountering| <startoftranscript>|Marker, start predicting subsequent content;

[0049] (This invention plans to use the officially provided Whisper-large model to transcribe audio to obtain its corresponding text information) During the testing phase, not only the prefix but also the complete text needs to be provided as a prompt to ensure that the Whisper model refers to the complete text information when generating the output.

[0050] More specifically, in step S2:

[0051] For long audio exceeding the preset time, use the Whisper model to split it into smaller decoding segments and record the end time of each segment (such as T 0 , T 1 , …, T k ). Subsequently, adopt the greedy merging strategy to merge the decoding segments into merged segments in sequence, making their lengths (as much as possible) close to but not exceeding the preset time, and the remaining unmergeable decoding segments are discarded;

[0052] For short audio not exceeding the preset time, extend it to the specified length by repeating and splicing to meet the input requirements of the Whisper model; finally, the preprocessed audio (segment) data is used for the training and testing of the Whisper model.

[0053] Furthermore, in step S3, the loss calculation is carried out step by step to ensure that the Whisper model is penalized for prediction errors in the key segments of the sequence, as shown below:

[0054] The total loss L is a weighted combination of the global loss L global and the local intermediate loss L inter . The global loss L global covers the entire sequence and is expressed as;

[0055] L global = ∑ i,j L CE (y i,j , y′ i,j );

[0056] where L CE represents the cross-entropy loss function, y is the predicted output of the model, y′ is the true label, and y i,j and y′ i,j represent the predicted value and the true value at the position (i, j) respectively; then, assume that the position of the i-th <|transcribe|> marker is T i , and the position of the i-th <|endoftext|> token is E i , the local intermediate loss L inter Focuses on the subsequence between <|transcribe|> and <|endoftext|>, that is, the part of the prediction result of audio authenticity, expressed as:

[0057]

[0058] The final loss value is calculated as L = L global + α * L inter , where α is a weight factor used to increase the influence of the local intermediate loss (this means that during training, the model will not only try to minimize the errors in the entire sequence (by minimizing L global ), but also pay special attention to reducing the errors in the subsequence between the <|transcribe|> and <|endoftext|> tokens (by minimizing L inter and giving it a higher weight)).

[0059] Furthermore, the complete training process of the model is as Figure 2 shown. In the training stage, the audio sample to be detected first arrives at a random decision point, randomly determining whether to perform voice enhancement processing on it to improve the diversity of the audio, enabling the model Whisper to be trained in a more realistic environment (thereby enhancing the robustness of the model and laying a foundation for subsequent accurate detection);

[0060] Subsequently, the processed (or unprocessed) audio data is input into the model Whisper. At this time, the encoder part of the model Whisper has been frozen to ensure the stability of its pre-trained parameters during training and guarantee the robust extraction of the features of the audio data by the model Whisper; then, through the designed loss function, the loss value is calculated based on the prediction result and the true label for updating the parameters of the model Whisper.

[0061] It should be noted that in this process, if the audio transcription text is used as the prompt information for the model to assist the model in better understanding the audio content and making accurate judgments, then the model will be named "Whisper-Trans-Prompt", which focuses on directly extracting key information from the audio signal itself for training and inference. On the contrary, if there is no need to rely on the prompt of the audio transcription text, the model is called "Whisper-Trans", which also focuses on directly extracting key information from the audio signal itself for training and inference. These two different settings lay the foundation for verifying the importance of introducing audio transcription text as a prompt in subsequent comparative experiments. In either case, the model can finally efficiently output the prediction results, clearly identifying the audio sample as "Spoofing" or "Genuine".

[0062] For the present invention:

[0063] The large-scale pre-trained model Whisper based on the Transformer architecture is adopted to achieve audio authenticity identification through transfer learning. Different from existing speech classification or recognition tasks (such as tasks based on Hubert, Wav2vec, etc.), the present invention not only utilizes the powerful multilingual speech recognition ability of the Whisper model, but also through an innovative fine-tuning strategy, embeds the complete transcription text of the audio as the prompt information into the decoder, overcoming the problem of single dependence on traditional acoustic features. In models such as Hubert and Wav2vec, feature extraction mainly relies on acoustic or time-frequency features, and it is difficult to fully capture the context and semantic information in speech, especially showing problems of feature sparsity and information loss in long speech processing.

[0064] In contrast, Whisper effectively aligns acoustic and semantic information by embedding the transcription text of the speech, making it show stronger robustness in long audio processing and complex spoofed audio scenarios. In addition, through the effective utilization of global speech information, the present invention breaks through the input length limit of 30-second audio segments of the Whisper model, and fine-tunes the global speech information through the sequence-to-sequence (Seq2Seq) paradigm, hoping to improve the accuracy and generalization ability of the model in the audio authenticity identification task. Compared with existing methods, this strategy performs well in a multilingual environment and ensures the capture and detection of spoofed audio details through dual alignment of semantics and acoustics in various spoofed types of scenarios.

[0065] ii. LoRA empowers the Whisper model to achieve high-efficiency and low-cost breakthroughs in spoofed audio detection.

[0066] The present invention applies the **Low-Rank Adaptation (LoRA)** method to the Whisper model to address the computational cost and resource consumption issues in forged audio detection. Different from existing transfer learning methods (such as the full-parameter fine-tuning strategies of Wav2vec and Hubert), the LoRA method injects low-rank decomposition matrices into the linear weights of the Transformer attention layer and freezes all the parameters of the Whisper model encoder, only fine-tuning a small part of the gated weights, which significantly reduces the number of parameters to be updated and remarkably decreases the computational cost and video memory requirements.

[0067] In traditional methods such as Wav2vec and Hubert, fine-tuning usually requires updating most of the model parameters, which easily leads to high video memory occupancy and low training efficiency, especially with high costs when training on large-scale audio datasets. However, with the LoRA fine-tuning strategy of the present invention, a graphics card with only 20GB of video memory can efficiently train the Whisper base version with 74 million parameters, significantly improving the resource utilization efficiency. In the pre-experiment, the LoRA fine-tuning strategy not only significantly reduced the model's training time and computational resource consumption but also demonstrated performance similar to or even better than the full fine-tuning method in the forged audio detection task. This low-cost and high-performance optimization strategy provides a new solution for audio authenticity identification and can strongly support large-scale industrial applications and cross-scenario deployments.

[0068] Compared with existing methods, the advantages of the proposed Whisper and LoRA combination solution of the present invention are as follows:

[0069] Few parameter updates and low computational cost: Compared with the traditional full-parameter fine-tuning strategy, the LoRA strategy only needs to fine-tune a small number of weight parameters, reducing the hardware requirements and computational burden.

[0070] Strong ability to fuse acoustic-semantic information: By using the speech transcription text as prompt information, the Whisper model shows stronger robustness and generalization ability in long speech and complex forgery scenarios.

[0071] Wide applicability: In the detection of multiple languages, multiple scenarios, and future unknown forgery types, it has stronger generalization ability and lower costs, laying a technical foundation for large-scale deployment.

[0072] It is worth mentioning that technical features such as the large-scale pre-trained model Whisper involved in this patent application for invention should be regarded as prior art. The specific structures, working principles, and possible control methods and spatial arrangement methods involved in these technical features can adopt conventional selections in the art and should not be regarded as the invention points of this patent for invention. This patent for invention will not be further specifically elaborated.

[0073] For those skilled in the art, it is still possible to modify the technical solutions described in the foregoing embodiments, or perform equivalent replacements for some of the technical features. Any modifications, equivalent replacements, improvements, etc. made within the spirit and principles of the present invention shall be included in the protection scope of the present invention.< / startoftranscript> < / startofprev> < / startoftranscript> < / startofprev>

Claims

1. A deep fake audio detection and protection method based on a large-scale pre-trained model Whisper, characterized in that: The following steps are involved: Step S1: Input the audio to be detected into the pre-trained Whisper model, and use the transcript of the entire audio as a prompt during the fine-tuning process of each audio segment, so that the Whisper model can utilize global text information; Step S2: During the training and decoding process, the audio data is preprocessed to adapt the model Whisper; Step S3: In terms of the design of the loss function, based on the default cross entropy loss function, a detection cross entropy loss function is designed to evaluate the performance of the model Whisper during the training process.

2. According to claim 1, a deep fake audio detection and protection method based on a large-scale pre-trained model Whisper is characterized in that: In step S1: Define the audio as x and the full text transcribed by the model Whisper as transcript x , under this framework, the prediction process of model Whisper for each segment is expressed as: P(y|x,transcript x )→result; Among them, y represents the input token sequence of the decoder in the model, and result is the classification result of the model Whisper. During the training process, the prompts of the transcribed text should be used as a reference rather than a target for prediction. Therefore, in the input and target of the decoder, from | <startofprev>The position of the prompt starting with the | mark remains unchanged, and the model Whisper directly outputs the same content based on the input. When it encounters | <startoftranscript> |Mark, start predicting the subsequent content;< / startoftranscript> < / startofprev> When conducting the testing phase, it is necessary to provide not only the prefix but also the complete text as a hint to ensure that the Whisper model refers to the complete text information when generating the output.

3. According to claim 2, a deep fake audio detection and protection method based on a large-scale pre-trained model Whisper is characterized in that: In step S2: For long audio that exceeds the preset time, the Whisper model is used to split it into smaller decoded segments and record the end time of each segment. Then, the greedy merging strategy is used to merge the decoded segments into merged segments in sequence, so that their length is close to but not longer than the preset time. The remaining decoded segments that cannot be merged are discarded. For short audio that does not exceed the preset time, it is extended to the specified length by repetition and splicing to meet the input requirements of the Whisper model; finally, the preprocessed audio data is used for training and testing of the Whisper model.

4. According to claim 3, a deep fake audio detection and protection method based on a large-scale pre-trained model Whisper is characterized in that: In step S3, the loss calculation is performed in steps to ensure that the model Whisper is penalized for prediction errors within the key segments of the sequence, as shown below: The total loss L is the global loss L global and local intermediate loss L inter The weighted combination of the global loss L global Covering the entire sequence, it is represented as; L global =∑ i,j L CE (y i,j ,y′ i,j ); Among them, L CE represents the cross entropy loss function, y is the predicted output of the model, y′ is the real label, and y i,j and y′ i,j Represent the predicted value and true value at position (i, j) respectively; then, assume that the position of the i-th <|transcribe|> tag is Ti, and the position of the i-th <|endoftext|> tag is E i , local intermediate loss L inter Focusing on the subsequence between <|transcribe|> and <|endoftext|>, that is, the prediction result of the audio authenticity, it is expressed as: The final loss value is calculated as L = L global +α*L inter , where α is a weight factor used to increase the influence of local intermediate loss.

5. According to claim 4, a deep fake audio detection and protection method based on a large-scale pre-trained model Whisper is characterized in that: During the training phase, the audio sample to be tested will first arrive at a random decision point, which will randomly determine whether to perform speech enhancement processing on it, thereby improving the diversity of the audio and allowing the Whisper model to be trained in an environment closer to reality; Subsequently, the processed audio data is input into the Whisper model. At this time, the encoder part of the Whisper model has been frozen to ensure that its pre-trained parameters remain stable during the training process, ensuring that the Whisper model can robustly extract the features of the audio data. Then, through the designed loss function, the loss value is calculated based on the prediction results and the true label, which is used to update the parameters of the Whisper model.

Citation Information

Patent Citations

  • Self-supervised voice authentic identification training method and system based on multi-center single classification

    CN117612562A

  • Counterfeit voice detection method and system based on time sequence multi-scale feature representation learning

    CN117809694A

  • Training method of forged voice detection model, forged voice detection method and device

    CN118366433A

  • False voice detection method based on self-supervised pre-training and supervised fine tuning

    CN119091912A

  • Multitask learning-based forged voice detection method and system

    CN119107979A