Audio deep pseudo detection method and device, computer equipment and storage medium

By refining and optimizing audio features and data augmentation in low-dimensional latent spaces, combining self-distillation models and classifiers, the performance degradation of existing audio deep pseudo-detection models in the face of unknown fake audio is solved, and higher detection accuracy and real-time performance are achieved.

CN120544602APending Publication Date: 2025-08-26PING AN TECH (SHENZHEN) CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510654431.8
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-05-20
Publication Date
2025-08-26

AI Technical Summary

Technical Problem

The existing audio depth forgery detection model has deteriorated detection performance when facing unknown types or post-processed deep forgery audio, making it difficult to cope with variable forgery technologies.

Method used

The pre-trained deep neural network is used to extract the latent representation of the audio, and refine and optimize and data enhancement are performed in the low-dimensional latent space, and enhance enhancement features are generated in combination with the self-distillation model, and audio deep pseudo-detection is finally performed through the classifier.

Benefits of technology

It improves the adaptability to unknown forgery methods, improves detection accuracy, reduces computing resource consumption, ensures real-time performance, and is suitable for real-time online monitoring systems.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120544602A_ABST
    Figure CN120544602A_ABST
Patent Text Reader

Abstract

The invention can be applied to the field of intelligent medical treatment and finance, and discloses an audio deep pseudo detection method and device, computer equipment and a storage medium, and the method comprises the steps: employing a pre-trained deep neural network to extract the potential representation of an audio, mapping the obtained potential representation of the audio to a low-dimensional potential space, and obtaining the potential features of the audio; refining and optimizing the potential features of the audio in the potential space to obtain optimized potential features of the audio; expanding the optimized potential features of the audio in the potential space by adopting a data enhancement strategy, and generating enhanced potential features of the audio in cooperation with a self-distillation model; and carrying out audio deep false detection on the obtained enhanced potential features of the audio by adopting a classifier. The method can improve the performance of audio deep pseudo detection.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the fields of artificial intelligence technology and speech processing technology, and in particular to a method, device, computer equipment and storage medium for detecting deepfake audio. Background Art

[0002] With the development of artificial intelligence, speech synthesis technology has emerged, enabling the creation of fake voices that resemble real people. However, these fake voices, also known as deepfakes, can be manipulated by malicious actors to commit various illegal or fraudulent activities. For example, deepfakes can be used to imitate applicants' voices to falsify records for health insurance certification, or to illegally access the bank accounts of users protected by voice protections through deepfakes. Deepfakes can also be used to impersonate users to commit telephone fraud. These activities not only seriously infringe on personal privacy and property security but also pose a serious threat to social order. To address these issues, existing technologies have proposed deepfake detection models to identify them.

[0003] Currently, existing audio deepfake detection models usually only perform well on training datasets. However, when faced with deepfake audio of unknown types or that has been post-processed, the detection performance deteriorates and it is difficult to cope with the changing forgery techniques. Summary of the Invention

[0004] The present invention provides an audio deepfake detection method, apparatus, computer equipment and medium to solve the technical problem that when faced with deepfake audio of unknown type or that has been post-processed, the detection performance is reduced and it is difficult to cope with the changing forgery techniques.

[0005] In a first aspect, a method for detecting deepfake audio is provided, comprising:

[0006] A pre-trained deep neural network is used to extract the latent representation of the audio, and the obtained latent representation of the audio is mapped to a low-dimensional latent space to obtain the latent features of the audio;

[0007] Refining and optimizing the latent features of the audio in the latent space to obtain the optimized latent features of the audio;

[0008] A data augmentation strategy is used in the latent space to expand the optimized latent features of the audio, and a self-distillation model is used to generate enhanced latent features of the audio.

[0009] A classifier is employed to perform audio deepfake detection on the obtained enhanced latent features of the audio.

[0010] In a second aspect, an audio deepfake detection device is provided, comprising:

[0011] A feature extraction module is used to extract the latent representation of the audio using a pre-trained deep neural network, map the obtained latent representation of the audio to a low-dimensional latent space, and obtain the latent features of the audio;

[0012] A feature optimization module is used to refine and optimize the latent features of the audio in the latent space to obtain the optimized latent features of the audio;

[0013] The feature enhancement module is used to expand the optimized latent features of the audio using a data enhancement strategy in the latent space, and to generate enhanced latent features of the audio in conjunction with the self-distillation model;

[0014] The feature classification module is used to perform audio deepfake detection on the obtained enhanced latent features of the audio using a classifier.

[0015] In a third aspect, a computer device is provided, comprising a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein the processor implements the steps of the above-mentioned audio deepfake detection method when executing the computer program.

[0016] In a fourth aspect, a computer-readable storage medium is provided, which stores a computer program. When the computer program is executed by a processor, the steps of the above-mentioned audio deep fake detection method are implemented.

[0017] In the scheme implemented by the above-mentioned audio deep fake detection method, device, computer equipment and storage medium, the latent representation of the audio can be extracted by using a pre-trained deep neural network, and the obtained latent representation of the audio can be mapped to a low-dimensional latent space to obtain the latent features of the audio; the latent features of the audio are refined and optimized in the latent space to obtain the optimized latent features of the audio; the optimized latent features of the audio are expanded by using a data enhancement strategy in the latent space, and enhanced latent features of the audio are generated in conjunction with a self-distillation model; a classifier is used to perform audio deep fake detection on the enhanced latent features of the obtained audio. In the present invention, for an intelligent assistant for medical insurance certification procedures under medical business, or for an intelligent customer service for voice-protected bank accounts under financial business , a solution for processing the latent representation of audio can be used. By mapping the latent representation of the extracted audio to a low-dimensional latent space, and then refining, optimizing and enhancing the feature representation in the latent space, the enhanced latent features of the audio are input into the classifier for audio deep fake detection. It can effectively filter out noise, highlight deep fake features, and make the classifier more sensitive and discriminative to forged audio, expand data distribution, improve adaptability to unknown forgery methods, improve the performance of audio deep fake detection, and improve detection accuracy. Moreover, the use of low-dimensional latent space operations can reduce the computing resource consumption of high-dimensional data processing, improve the inference speed, reduce computing overhead, and ensure real-time performance. It is suitable for deployment in real-time online monitoring systems and provides efficient protection for audio security. BRIEF DESCRIPTION OF THE DRAWINGS

[0018] In order to more clearly illustrate the technical solutions of the embodiments of the present invention, the following briefly introduces the drawings required for use in the description of the embodiments of the present invention. Obviously, the drawings described below are only some embodiments of the present invention. For ordinary technicians in this field, other drawings can be obtained based on these drawings without paying any creative labor.

[0019] Figure 1 2. It is a schematic diagram of an application environment of an audio deepfake detection method according to an embodiment of the present invention;

[0020] Figure 2 This is a flowchart of a method for detecting deepfake audio in accordance with an embodiment of the present invention;

[0021] Figure 3 yes Figure 2 A schematic flow chart of a specific implementation of step S20;

[0022] Figure 4 2 is a schematic structural diagram of an audio deepfake detection device according to an embodiment of the present invention;

[0023] Figure 5 is a structural diagram of a computer device in one embodiment of the present invention;

[0024] Figure 6 FIG. 2 is another structural diagram of a computer device according to an embodiment of the present invention. DETAILED DESCRIPTION

[0025] The following will clearly and completely describe the technical solutions in the embodiments of the present invention in conjunction with the accompanying drawings. Obviously, the described embodiments are only part of the embodiments of the present invention, not all of them. All other embodiments obtained by ordinary technicians in this field based on the embodiments of the present invention without making any creative efforts shall fall within the scope of protection of the present invention.

[0026] The audio deepfake detection method provided by the embodiment of the present invention can be applied in the following situations: Figure 1 In the application environment, it is used in intelligent assistants or intelligent customer service in application scenarios such as medical care and finance, and is usually implemented through the server, wherein the client communicates with the server through the network. The server can obtain audio through the client, use a pre-trained deep neural network to extract the latent representation of the audio, map the obtained latent representation of the audio to a low-dimensional latent space, and obtain the latent features of the audio; refine and optimize the latent features of the audio in the latent space to obtain the optimized latent features of the audio; use a data enhancement strategy in the latent space to expand the optimized latent features of the audio, and cooperate with the self-distillation model to generate enhanced latent features of the audio; use a classifier to perform audio deep fake detection on the enhanced latent features of the obtained audio, and finally feed back the audio deep fake detection results to the client. In the present invention, for the intelligent assistant of the medical insurance authentication procedure under the medical business, or for the intelligent customer service of the voice-protected bank account under the financial business, A solution for processing the latent representation of audio can be used. By mapping the latent representation of the extracted audio to a low-dimensional latent space, and then refining, optimizing and enhancing the feature representation in the latent space, the enhanced latent features of the audio are input into the classifier for audio deep fake detection. This can effectively filter noise, highlight deep fake features, and make the classifier more sensitive and discriminative to forged audio, expand data distribution, improve adaptability to unknown forgery methods, improve the performance of audio deep fake detection, and improve detection accuracy. Moreover, the use of low-dimensional latent space operations can reduce the computing resource consumption of high-dimensional data processing, improve inference speed, reduce computing overhead, ensure real-time performance, and be suitable for deployment in real-time online monitoring systems, providing efficient protection means for audio security. Among them, the client can be, but is not limited to, various personal computers, laptops, smart phones, tablets and portable wearable devices. The server can be implemented with an independent server or a server cluster consisting of multiple servers. The present invention is described in detail below through specific embodiments.

[0027] See also Figure 2As shown, Figure 2 A flowchart of a method for detecting deepfake audio data provided by an embodiment of the present invention includes the following steps:

[0028] S10: Use a pre-trained deep neural network to extract the latent representation of the audio, map the obtained latent representation of the audio to a low-dimensional latent space, and obtain the latent features of the audio.

[0029] The audio deepfake detection method provided by the present invention can be applied to intelligent customer service or intelligent assistants in various application scenarios such as medical care, finance and insurance, and is usually implemented through the server. For example, in the intelligent assistant under the medical insurance certification program in the medical application field, the applicant can verify the applicant's identity, confirm his / her application intention or understand the oral description of his / her health status through voice communication methods such as telephone authentication or remote consultation. For example, the audio can be "medical insurance application", "condition description" or "health status", etc. Or, for example, in the intelligent customer service of a voice-protected bank account in the financial application field, the user can perform information setting operations and transaction operations on the bank account through voice. For example, the audio can be: "password reset" or "transfer service", etc. After receiving the audio, the intelligent assistant or intelligent customer service can use a pre-trained deep neural network to extract the latent representation of the audio, map the obtained latent representation of the audio to a low-dimensional latent space, and obtain the latent features of the audio. The use of latent space operations not only reduces the computing resource consumption of high-dimensional data processing, but also improves the inference speed, reduces computing overhead, and ensures real-time performance.

[0030] Among them, step S10, that is, using a pre-trained deep neural network to extract the potential representation of the audio, mapping the obtained potential representation of the audio to a low-dimensional latent space, and obtaining the potential features of the audio, is specifically:

[0031] A pre-trained deep neural network is used to convert the input audio into a latent representation of the audio. The deep neural network can be an audio feature extraction network, such as Wav2Vec2.0. This network is trained on large-scale speech data to capture the temporal dynamics and spectral details of the audio, and converts the latent representation of the audio into a high-dimensional feature representation.

[0032] The obtained latent representation of the audio is mapped to a low-dimensional latent space through a dimensionality reduction module to obtain the latent features of the audio. The dimensionality reduction module uses a fully connected layer or a convolutional downsampling layer to map the latent representation of the high-dimensional audio to a low-dimensional latent space. This can accommodate the problem of mismatch between the dimensions of the pre-trained deep neural network and the current dimension, and ensure that the low-dimensional latent space can retain key information while reducing computational complexity.

[0033] In an embodiment of the invention, before step S10, that is, before extracting the latent representation of the audio using a pre-trained deep neural network, the method further includes the following steps:

[0034] An initial audio signal to be tested is received and preprocessed to obtain preprocessed audio. The preprocessing includes noise reduction and normalization. Preprocessing the initial audio signal to be tested reduces noise interference with the audio feature extraction process, improving the reliability and accuracy of the extracted potential representation of the audio.

[0035] S20: Refining and optimizing the latent features of the audio in the latent space to obtain optimized latent features of the audio.

[0036] By refining and optimizing the audio's latent features to make the optimized latent features of the obtained audio clearer, noise can be effectively suppressed and the difference between forged signals can be enhanced. That is, redundant noise is removed and forgery traces are highlighted to facilitate subsequent classification, making the classifier more sensitive and discriminative to forged audio, thereby improving detection performance.

[0037] In some embodiments of the invention, Figure 3 As shown, a specific audio latent feature refinement and optimization solution is provided. In S20, the latent features of the audio are refined and optimized in the latent space to obtain the optimized latent features of the audio, which specifically includes the following steps S21-S23:

[0038] S21: Using a batch normalization method or a layer normalization method to normalize the latent features of the audio to obtain standard latent features of the audio.

[0039] By standardizing the potential features of audio, the interference caused by uneven data distribution can be eliminated and the reliability and accuracy of the features can be improved.

[0040] S22: Capture the key information of the standard latent features of the audio through the self-attention network, and perform weighted reconstruction on the standard latent features of the audio to obtain the reconstructed latent features of the audio.

[0041] The attention mechanism is used to calculate the relevance of the input. The greater the attention, the greater the corresponding weight. The self-attention network automatically captures the key information of the standard latent features of the audio, and reconstructs the standard latent features of the audio in a weighted manner, which helps to highlight possible traces of forgery.

[0042] Audio is a typical temporal feature, so that the audio can start from the leftmost time and end at the rightmost time. In order to adapt to the temporal characteristics of the standard potential features of the audio, in step S22, that is, capturing the key information of the standard potential features of the audio through the self-attention network, the self-attention network performs a masking operation on all information of the subsequent time sequence of the current time point of the standard potential features of the input audio to erase the information of the subsequent time of the current time point, and then erases the information of the time point to the right of the current time point, so that the calculated attention is only related to the information of the current time point and all previous time points, but not to future information. Therefore, in the attention calculation process, the current input is only correlated with the previous input.

[0043] S23: Perform residual connection on the reconstructed latent features of the audio and refine them through multi-layer stacking to obtain the optimized latent features of the audio.

[0044] Residual connections are used to connect the input and output to prevent gradient vanishing and gradient explosion. Multi-layer stacking can further refine the potential features. Each layer can perform nonlinear mapping and feature fusion on the input features, gradually suppressing useless information and strengthening anomalies related to deep fakes, which is conducive to increasing the number of model parameters.

[0045] S30: Use data augmentation strategies in the latent space to expand the optimized latent features of the audio, and cooperate with the self-distillation model to generate enhanced latent features of the audio.

[0046] Data augmentation strategies can be used to simulate various forgery techniques and noise conditions, generate diverse feature representations, and improve the robustness of the model when facing unseen samples. Self-distillation technology uses the model's own output as soft labels to guide training, compares the outputs of models after different data augmentation, and minimizes the KL divergence between the outputs of models after different data augmentation to achieve consistency optimization. This allows the model to better understand the feature distribution under different enhancement methods, so that it has better adaptability to various forgery methods. Ensuring that the data distribution is fully expanded and refined in the latent space can capture a wider range of deep fake features and reduce the risk of overfitting due to insufficient training data. Through self-distillation and data augmentation technology, it can adapt to a variety of noise and disturbance conditions, maintain stable performance in cross-domain data and actual environments, reduce the risk of false positives and false negatives, and enhance generalization and robustness.

[0047] In some embodiments of the invention, in S30, that is, in the data enhancement strategy used in the latent space to expand the optimized latent features of the audio, the data enhancement strategy includes random perturbation, feature mixing and / or model pruning, and the optimized latent features of the audio are expanded in the latent space by random perturbation, feature mixing and / or model pruning.

[0048] S40: Using a classifier to perform audio deepfake detection on the obtained enhanced latent features of the audio.

[0049] The classifier can perform binary classification on the input and output a credibility score. Binary classification refers to classifying the input into real audio and deepfake audio. The classifier can also perform multi-classification on the input and output a credibility score. Multi-classification distinguishes the input based on different deepfake techniques. During training, the classifier and the front-end modules are jointly optimized for end-to-end training, jointly minimizing classification loss and regularization loss, ensuring stable operation in various environments.

[0050] Specifically, in step S40, the classifier may adopt a lightweight fully connected neural network or a convolution-based classifier.

[0051] The audio deepfake detection method provided by the present invention can be applied to intelligent customer service or intelligent assistants in various application scenarios, such as healthcare, finance, and insurance, and is typically implemented on a server. For example, in an intelligent assistant for medical insurance certification in the medical field, an applicant can verify their identity, confirm their application intention, or obtain a verbal description of their health status through voice communication such as telephone authentication or remote consultation. For example, the audio could be "medical insurance application," "condition description," or "health status." Alternatively, in a financial application, an intelligent customer service representative for a voice-protected bank account can use voice to set bank account information and conduct transactions. For example, the audio could be "password reset" or "transfer service." After receiving the audio, the intelligent assistant or intelligent customer service representative can perform deepfake detection on the audio and perform corresponding operations based on the audio deepfake detection result. If the audio deepfake detection result is authentic, the corresponding voice information is returned based on the corresponding operation steps stored on the server to guide the applicant or user to the next voice operation. If the audio deepfake detection result is a deepfake, an alarm can be issued and the corresponding applicant or user information can be recorded.

[0052] In some embodiments, after step S40, that is, after performing audio deepfake detection on the obtained enhanced latent features of the audio using a classifier, the audio deepfake detection method further includes:

[0053] The final audio deepfake judgment is made based on the output of the classifier and the preset threshold.

[0054] The output of the classifier can be a probability distribution. When the classifier performs binary classification on the input, the real audio can be represented by 1 and the deep fake audio can be represented by 0. The hidden layer data can be compressed in the last layer of the model, and then the softmax function is used to calculate the probability of the two categories of classification corresponding to the input. According to the probability distribution of the two categories of classification corresponding to the classifier output, combined with the preset threshold, the final audio deep fake judgment is made, and the final audio deep fake detection result report is output.

[0055] It can be seen that in the above scheme, the intelligent assistant for the medical insurance authentication procedure under the medical business, or the intelligent customer service for the voice-protected bank account under the financial business, can utilize the audio latent representation processing scheme, by mapping the extracted audio latent representation to a low-dimensional latent space, and then refining, optimizing and enhancing the feature representation in the latent space. The enhanced latent features of the audio obtained are input into the classifier for audio deep fake detection, which can effectively filter out noise, highlight deep fake features, and make the classifier more sensitive and discriminative to forged audio, expand data distribution, improve adaptability to unknown forgery methods, improve the performance of audio deep fake detection, and improve detection accuracy. Moreover, the use of low-dimensional latent space operations can reduce the computing resource consumption of high-dimensional data processing, improve the inference speed, reduce computing overhead, and ensure real-time performance. It is suitable for deployment in real-time online monitoring systems and provides efficient protection for audio security.

[0056] It should be understood that the size of the serial numbers of the steps in the above embodiments does not mean the order of execution. The execution order of each process should be determined by its function and internal logic, and should not constitute any limitation on the implementation process of the embodiments of the present invention.

[0057] In one embodiment, an audio deep fake detection device is provided, which corresponds to the audio deep fake detection method in the above embodiment. Figure 4 As shown, the audio deepfake detection device includes a feature extraction module 101, a feature optimization module 102, a feature enhancement module 103 and a feature classification module 104. The functional modules are described in detail as follows:

[0058] A feature extraction module 101 is configured to extract a latent representation of the audio using a pre-trained deep neural network, map the obtained latent representation of the audio to a low-dimensional latent space, and obtain latent features of the audio;

[0059] A feature optimization module 102 is used to refine and optimize the latent features of the audio in the latent space to obtain optimized latent features of the audio;

[0060] A feature enhancement module 103 is used to expand the optimized latent features of the audio using a data enhancement strategy in the latent space, and to generate enhanced latent features of the audio in conjunction with the self-distillation model;

[0061] The feature classification module 104 is used to use a classifier to perform audio deepfake detection on the obtained enhanced potential features of the audio.

[0062] In one embodiment, the feature extraction module 101 is further configured to:

[0063] An initial audio signal to be tested is received and preprocessed to obtain preprocessed audio.

[0064] In one embodiment, the feature optimization module 102 is specifically configured to:

[0065] Using batch normalization or layer normalization to normalize the potential features of the audio to obtain the standard potential features of the audio;

[0066] The self-attention network is used to capture the key information of the standard latent features of the audio, and the standard latent features of the audio are weighted reconstructed to obtain the reconstructed latent features of the audio;

[0067] The reconstructed latent features of the audio are residually connected and refined through multi-layer stacking to obtain the optimized latent features of the audio.

[0068] In one embodiment, the feature optimization module 102 is further configured to:

[0069] A mask operation is performed on all information of the subsequent time sequence of the current time point of the standard latent feature of the input audio.

[0070] In one embodiment, the feature classification module 104 is further configured to:

[0071] The final audio deepfake judgment is made based on the output of the classifier and the preset threshold.

[0072] The present invention provides an audio deep fake detection device, which maps the latent representation of the extracted audio to a low-dimensional latent space, and then refines, optimizes and enhances the feature representation in the latent space. The enhanced latent features of the obtained audio are input into a classifier for audio deep fake detection. It can effectively filter noise, highlight deep fake features, and make the classifier more sensitive and discriminative to forged audio, expand data distribution, improve adaptability to unknown forgery methods, improve the performance of audio deep fake detection, and improve detection accuracy. Moreover, the use of low-dimensional latent space operations can reduce the computing resource consumption of high-dimensional data processing, improve the inference speed, reduce computing overhead, and ensure real-time performance. It is suitable for deployment in real-time online monitoring systems, and provides an efficient protection means for audio security.

[0073] For the specific definition of the audio deep fake detection device, please refer to the definition of the audio deep fake detection method above, which will not be repeated here. The various modules in the above-mentioned audio deep fake detection device can be implemented in whole or in part by software, hardware, and a combination thereof. The above-mentioned modules can be embedded in or independent of the processor in the computer device in hardware form, or can be stored in the memory of the computer device in software form, so that the processor can call and execute the operations corresponding to the above modules.

[0074] In one embodiment, a computer device is provided. The computer device may be a server, and its internal structure diagram may be as follows: Figure 5 As shown. The computer device includes a processor, a memory, a network interface and a database connected via a system bus. The processor of the computer device is used to provide computing and control capabilities. The memory of the computer device includes a non-volatile and / or volatile storage medium and an internal memory. The non-volatile storage medium stores an operating system, a computer program and a database. The internal memory provides an environment for the operation of the operating system and the computer program in the non-volatile storage medium. The network interface of the computer device is used to communicate with an external client via a network connection. When the computer program is executed by the processor, it implements the functions or steps on the server side of an audio deep fake detection method.

[0075] In one embodiment, a computer device is provided. The computer device may be a client, and its internal structure diagram may be as follows: Figure 6 As shown. The computer device includes a processor, memory, network interface, display screen and input device connected through a system bus. Among them, the processor of the computer device is used to provide computing and control capabilities. The memory of the computer device includes a non-volatile storage medium and an internal memory. The non-volatile storage medium stores an operating system and a computer program. The internal memory provides an environment for the operation of the operating system and computer program in the non-volatile storage medium. The network interface of the computer device is used to communicate with an external server through a network connection. When the computer program is executed by the processor, it realizes the functions or steps of the client side of an audio deep fake detection method

[0076] In one embodiment, a computer device is provided, including a memory, a processor, and a computer program stored in the memory and executable on the processor. When the processor executes the computer program, the following steps are performed:

[0077] A pre-trained deep neural network is used to extract the latent representation of the audio, and the obtained latent representation of the audio is mapped to a low-dimensional latent space to obtain the latent features of the audio;

[0078] Refining and optimizing the latent features of the audio in the latent space to obtain the optimized latent features of the audio;

[0079] A data augmentation strategy is used in the latent space to expand the optimized latent features of the audio, and a self-distillation model is used to generate enhanced latent features of the audio.

[0080] A classifier is employed to perform audio deepfake detection on the obtained enhanced latent features of the audio.

[0081] In one embodiment, a computer-readable storage medium is provided, on which a computer program is stored. When the computer program is executed by a processor, the following steps are implemented:

[0082] A pre-trained deep neural network is used to extract the latent representation of the audio, and the obtained latent representation of the audio is mapped to a low-dimensional latent space to obtain the latent features of the audio;

[0083] Refining and optimizing the latent features of the audio in the latent space to obtain the optimized latent features of the audio;

[0084] A data augmentation strategy is used in the latent space to expand the optimized latent features of the audio, and a self-distillation model is used to generate enhanced latent features of the audio.

[0085] A classifier is employed to perform audio deepfake detection on the obtained enhanced latent features of the audio.

[0086] It should be noted that the above functions or steps that can be implemented by the computer-readable storage medium or computer device can be found in the relevant descriptions of the server side and the client side in the aforementioned method embodiment. To avoid repetition, they will not be described one by one here.

[0087] Those skilled in the art will appreciate that all or part of the processes in the above-mentioned embodiments can be implemented by instructing the relevant hardware through a computer program. The computer program can be stored in a non-volatile computer-readable storage medium. When the computer program is executed, it can include the processes of the embodiments of the above-mentioned methods. Among them, any reference to memory, storage, database or other media used in the embodiments provided in this application can include non-volatile and / or volatile memory. Non-volatile memory can include read-only memory (ROM), programmable ROM (PROM), electrically programmable ROM (EPROM), electrically erasable programmable ROM (EEPROM) or flash memory. Volatile memory can include random access memory (RAM) or external cache memory. By way of illustration and not limitation, RAM is available in various forms, such as static RAM (SRAM), dynamic RAM (DRAM), synchronous DRAM (SDRAM), double data rate SDRAM (DDRSDRAM), enhanced SDRAM (ESDRAM), synchronous link (Synchlink) DRAM (SLDRAM), memory bus (Rambus) direct RAM (RDRAM), direct memory bus dynamic RAM (DRDRAM), and memory bus dynamic RAM (RDRAM).

[0088] Those skilled in the art will clearly understand that for the sake of convenience and brevity of description, only the division of the above-mentioned functional units and modules is used as an example. In actual applications, the above-mentioned functions can be distributed and completed by different functional units and modules as needed, that is, the internal structure of the device can be divided into different functional units or modules to complete all or part of the functions described above.

[0089] The embodiments described above are only used to illustrate the technical solutions of the present invention, rather than to limit the same. Although the present invention has been described in detail with reference to the aforementioned embodiments, those skilled in the art should understand that they can still modify the technical solutions described in the aforementioned embodiments, or make equivalent replacements for some of the technical features therein. These modifications or replacements do not deviate the essence of the corresponding technical solutions from the spirit and scope of the technical solutions of the various embodiments of the present invention, and should all be included in the scope of protection of the present invention.

Claims

1. A method for detecting deepfake audio, characterized in that: include: A pre-trained deep neural network is used to extract the latent representation of the audio, and the obtained latent representation of the audio is mapped to a low-dimensional latent space to obtain the latent features of the audio; Refining and optimizing the latent features of the audio in the latent space to obtain the optimized latent features of the audio; A data augmentation strategy is used in the latent space to expand the optimized latent features of the audio, and a self-distillation model is used to generate enhanced latent features of the audio. A classifier is employed to perform audio deepfake detection on the obtained enhanced latent features of the audio.

2. The method for detecting deepfake audio according to claim 1, wherein: Before extracting the latent representation of the audio using a pre-trained deep neural network, the method includes: An initial audio signal to be tested is received and preprocessed to obtain preprocessed audio.

3. The audio deepfake detection method according to claim 1, wherein: The refining and optimizing the latent features of the audio in the latent space to obtain the optimized latent features of the audio includes: Using batch normalization or layer normalization to normalize the potential features of the audio to obtain the standard potential features of the audio; The self-attention network is used to capture the key information of the standard latent features of the audio, and the standard latent features of the audio are weighted reconstructed to obtain the reconstructed latent features of the audio; The reconstructed latent features of the audio are residually connected and refined through multi-layer stacking to obtain the optimized latent features of the audio.

4. The audio deepfake detection method according to claim 3, wherein: In the method of capturing key information of the standard latent features of the audio through the self-attention network, the self-attention network performs a mask operation on all information of the subsequent time sequence of the current time point of the standard latent features of the input audio.

5. The audio deepfake detection method according to claim 1, wherein: The data enhancement strategy used in the latent space to expand the optimized latent features of the audio includes random perturbation, feature mixing and / or model pruning.

6. The method for detecting deepfake audio according to claim 1, wherein: The classifier used to perform audio deepfake detection on the enhanced potential features of the obtained audio adopts a fully connected neural network or a convolution-based classifier.

7. The audio deepfake detection method according to claim 1, wherein: After performing audio deepfake detection on the obtained enhanced latent features of the audio using a classifier, the method further includes: The final audio deepfake judgment is made based on the output of the classifier and the preset threshold.

8. An audio deepfake detection device, characterized in that: include: A feature extraction module is used to extract the latent representation of the audio using a pre-trained deep neural network, map the obtained latent representation of the audio to a low-dimensional latent space, and obtain the latent features of the audio; A feature optimization module is used to refine and optimize the latent features of the audio in the latent space to obtain the optimized latent features of the audio; The feature enhancement module is used to expand the optimized latent features of the audio using a data enhancement strategy in the latent space, and to generate enhanced latent features of the audio in conjunction with the self-distillation model; The feature classification module is used to perform audio deepfake detection on the obtained enhanced latent features of the audio using a classifier.

9. A computer device comprising a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein: When the processor executes the computer program, the steps of the audio deep fake detection method according to any one of claims 1 to 7 are implemented.

10. A computer-readable storage medium storing a computer program, characterized in that: When the computer program is executed by a processor, the steps of the audio deep fake detection method according to any one of claims 1 to 7 are implemented.