Cross-domain deep counterfeit voice detection method and system

By assisting cross-domain processing and self-supervised feature extraction for source speech, combined with single-class learning and domain-invariant learning, the problem of rising detection false alarm rate caused by environmental changes in the existing technology is solved, and more stable speech detection accuracy is achieved.

CN120148549APending Publication Date: 2025-06-13TONGXIANG GENERAL ARTIFICIAL INTELLIGENCE RESEARCH INSTITUTE +1
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510350268.6
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-03-24
Publication Date
2025-06-13

AI Technical Summary

Technical Problem

The existing deep fake voice detection methods fail to fully consider the changing characteristics of voice in different environments, resulting in an increase in detection false alarm rate.

Method used

By assisting cross-domain processing of source speech, cross-domain speech (assisted speech) caused by distortion caused by different voice encoders and transmission conditions is generated. Then, frame-level feature extraction and low-dimensional feature space projection are used to perform frame-level feature extraction and low-dimensional feature space projection, combined with single-class learning and domain-invariant learning, the total loss function value is calculated for gradient update, and finally the speech detection result is determined.

Benefits of technology

Accurate modeling of real voice features under various conditions is achieved, which reduces the influence of environmental factors, improves detection accuracy, and reduces the false alarm rate.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120148549A_ABST
    Figure CN120148549A_ABST
Patent Text Reader

Abstract

The invention relates to a cross-domain deep counterfeit voice detection method and system in the technical field of voice detection. The method comprises the following steps: generating an auxiliary voice; performing frame-level feature extraction on the source voice and the auxiliary voice through a self-supervised feature extractor, and projecting the extracted features to a low-dimensional feature space to obtain a first feature vector and a second feature vector respectively; carrying out single-class learning on the first feature vector to calculate a single-class loss function value, and carrying out domain invariant learning on the second feature vector to calculate a cross entropy loss function value; calculating a total loss function value; performing gradient updating on the single classifier, performing gradient updating on the adversarial domain classifier, and performing gradient updating on parameters of the self-supervised feature extractor and the projection network; the test voice in the test set is input, the voice detection value is obtained, the voice detection result is judged based on the voice detection value, and the problem that the detection false alarm rate is increased due to the fact that environment changes are not considered in the existing forged voice detection process is solved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of voice detection, and specifically relates to a cross-domain deepfake voice detection method and system. Background Art

[0002] With the increasing popularity of artificial intelligence and machine learning, voice deepfake technology has become a key issue in industries such as media, security, and communication. Voice deepfake technology involves synthesizing or manipulating voice recordings to imitate an individual's voice, bringing threats including identity theft, misinformation, and social engineering attacks. These challenges require powerful and practical detection solutions.

[0003] Traditional voice forensics methods rely on analyzing acoustic features or voice patterns, but usually cannot cope with the complexity of modern deepfake generation techniques. With the development of neural network-based deepfake models, the voice content they generate is becoming more and more realistic and difficult to distinguish from real recordings. Especially in real-world applications, data quality and transmission variations may distort the voice signal, which poses a major challenge to existing detection systems.

[0004] Existing deepfake voice detection methods mainly focus on identifying anomalies in fake voices, such as unnatural timing or artifacts introduced during the generation process. However, since existing methods do not fully consider the variation characteristics in different voice environments during detection, for example, in Internet-based communication and VoIP / PSTN networks, factors such as voice compression, channel characteristics, and codec conversion will introduce additional distortions, and these distortions are easily misjudged as abnormal features of fake voices by existing detection methods, resulting in an increase in the false alarm rate. Summary of the Invention

[0005] Aiming at the shortcomings in the prior art, the present invention provides a cross-domain deepfake voice detection method and system, which solves the problem of the increase in the detection false alarm rate caused by the failure to consider environmental changes in the existing fake voice detection process.

[0006] To solve the above technical problems, the present invention is solved by the following technical solutions:

[0007] A cross-domain deepfake voice detection method includes the following steps:

[0008] Performing auxiliary cross-domain processing on the source voice in the training set to generate auxiliary voice;

[0009] Performing frame-level feature extraction on both the source voice and the auxiliary voice through a self-supervised feature extractor, and projecting the extracted frame-level features into a low-dimensional feature space through a projection network to obtain a first feature vector and a second feature vector respectively;

[0010] Perform one-class learning on the source speech based on the first feature vector, and calculate the one-class classification loss function value. Perform domain-invariant learning on the auxiliary speech based on the second feature vector, and calculate the cross-entropy loss function value;

[0011] Calculate the total loss function value based on the one-class classification loss function value and the cross-entropy loss function value;

[0012] Use the one-class classification loss function value to perform gradient update on the one-class classifier for one-class learning, use the cross-entropy loss function value to perform gradient update on the adversarial domain classifier for domain-invariant learning, and use the total loss function value to perform gradient update on the parameters of the self-supervised feature extractor and the projection network;

[0013] Input the test speech in the test set into the self-supervised feature extractor, the projection network, and the one-class classifier after gradient update to obtain the speech detection value, and determine the speech detection result based on the speech detection value.

[0014] Optionally, the auxiliary cross-domain processing includes the following steps:

[0015] The source speech includes the first real speech and the forged speech;

[0016] Perform transmission distortion processing on the first real speech to obtain the real speech after transmission variation;

[0017] Perform codec distortion processing on the first real speech to obtain the real speech after compression variation.

[0018] Optionally, the transmission distortion processing includes: performing jitter, delay, and packet loss processing on the first real speech;

[0019] The codec distortion processing uses a media encoder to perform codec compression simulation on the first real speech.

[0020] Optionally, the calculation formula for the one-class classification loss function value is:

[0021] ,

[0022] where, represents the batch size extracted from the source speech for training at one time; represents the scaling factor; represents the th true label corresponding to the data point; represents the th first feature vector corresponding to the source speech; represents the optimization direction of the first real speech embedding vector; and are respectively normalized to and ; When is the case, is represented as When is the case, is represented as and when > is the case, it is used to constrain and to form a boundary between the first real speech and the forged speech distribution.

[0023] Optionally, perform domain-invariant learning on the auxiliary speech based on the second feature vector and calculate the cross-entropy loss function value, including the following steps:

[0024] Set more than one adversarial domain classifier, map the second feature vector through the adversarial domain classifier to more than one label, and calculate the cross-entropy loss function value.

[0025] Optionally, when two sets of adversarial domain classifiers are set, set the two sets of adversarial classifiers as a speaker classifier and a conditional classifier, and calculate the cross-entropy loss function value, including the following steps:

[0026] Map the second feature vector through the speaker classifier to the speaker label, calculate the first cross-entropy loss function value, and the calculation formula is as follows:

[0027] ;

[0028] Map the second feature vector through the conditional classifier to the conditional label, calculate the second cross-entropy loss function value, and the calculation formula is as follows:

[0029] ;

[0030] Among them, represents the batch size extracted once from the auxiliary speech for training; represents the i-th auxiliary speech corresponding to the second feature vector in the low-dimensional feature space; represents the number of speaker domain categories; represents the number of conditional domain categories; represents corresponding to the true label of the th domain in the speaker domain; represents corresponding to the true label of the th domain in the conditional domain; represents the calculation of the gradient reversal layer; represents the speaker domain classifier; Represents a conditional domain classifier.

[0031] Optionally, the formula for calculating the total loss function value is:

[0032] , where λ 1 and λ 2 are both hyperparameters used to balance the trade-off between different objectives during training.

[0033] Optionally, determining the speech detection result based on the speech detection value includes the following steps:

[0034] Set a speech threshold and determine whether the speech detection value is greater than the speech threshold;

[0035] If the speech detection value is greater than the speech threshold, it is determined that the test speech is real speech, otherwise it is determined that the test speech is forged speech.

[0036] A cross-domain deepfake speech detection system, the deepfake speech detection system executes a cross-domain deepfake speech detection method as described above, including an auxiliary cross-domain data generation unit, a self-supervised feature extraction unit, a projection network unit, a one-class learning unit, a domain-invariant learning unit, a gradient update unit, and a test unit;

[0037] The auxiliary cross-domain data generation unit is used to perform auxiliary cross-domain processing on the source speech in the training set to generate auxiliary speech;

[0038] The self-supervised feature extraction unit is used to perform frame-level feature extraction on the source speech and the auxiliary speech respectively;

[0039] The projection network unit is used to project the source speech and the auxiliary speech after frame-level feature extraction into a low-dimensional feature space to obtain a first feature vector and a second feature vector respectively;

[0040] The one-class learning unit is used to perform one-class learning based on the first feature vector and calculate the one-class classification loss function value;

[0041] The domain-invariant learning unit is used to perform domain-invariant learning based on the second feature vector and calculate the cross-entropy loss function value;

[0042] The gradient update unit is used to calculate the total loss function value based on the one-class classification loss function value and the cross-entropy loss function value, and use the one-class classification loss function value to perform gradient update on the one-class classifier of one-class learning, use the cross-entropy loss function value to perform gradient update on the adversarial domain classifier of domain-invariant learning, and use the total loss function value to perform gradient update on the parameters of the self-supervised feature extractor and the projection network;

[0043] The test unit is used to input the test speech in the test set into the self-supervised feature extractor, projection network, and single classifier after gradient update, obtain the speech detection value, and determine the speech detection result based on the speech detection value.

[0044] A computer-readable storage medium stores a computer program, which, when executed by a processor, implements the above-mentioned cross-domain deepfake speech detection method.

[0045] Adopting the technical solution provided by the present invention, compared with the prior art, it has the following beneficial effects:

[0046] 1. By performing auxiliary cross-domain processing on the first real speech in the source speech, cross-domain speech (auxiliary speech) caused by distortions caused by different speech encoders and transmission conditions is generated; then, through self-supervised feature extraction, a self-supervised pre-training model is used to extract high-dimensional and domain-generalizable features from the speech data, ensuring accurate modeling of real speech features under various conditions, including those affected by compression or transmission.

[0047] 2. Through domain-invariant representation learning, aligning the feature distributions in different domains reduces the influence of environmental factors and achieves stable detection accuracy. Such domain-invariant features ensure strong robustness against distortions caused by codec conversion or network transmission channels.

[0048] 3. Incorporating a one-class learning method to model the distribution of real speech enables the system to effectively identify the deviations indicating synthetic speech. Description of the Drawings

[0049] To more clearly illustrate the technical solutions in the embodiments of the present invention or the prior art, the following will briefly introduce the drawings required for the description of the embodiments or the prior art. Obviously, the following drawings are only some embodiments of the present invention. For those of ordinary skill in the art, other drawings can be obtained based on these drawings without creative efforts.

[0050] Figure 1 It is a flowchart of a cross-domain deepfake speech detection method proposed in the first embodiment.

[0051] Figure 2 It is a flowchart of the generation of auxiliary speech proposed in the first embodiment.

[0052] Figure 3 It is a data processing flowchart of the self-supervised feature extractor, projection network, single classifier, and speaker / condition classifier proposed in the first embodiment. Detailed Embodiments

[0053] The present invention will be further described in detail below in conjunction with embodiments. The following embodiments are explanations of the present invention, and the present invention is not limited to the following embodiments.

[0054] Embodiment 1

[0055] As Figure 1 shown, a cross-domain deepfake voice detection method includes the following steps: performing auxiliary cross-domain processing on the source voice in the training set to generate auxiliary voice; wherein, the auxiliary cross-domain processing includes the following steps: the source voice includes the first real voice and the fake voice; performing transmission distortion processing on the first real voice to obtain the real voice after transmission variation; performing codec distortion processing on the first real voice to obtain the real voice after compression variation. The transmission distortion processing includes: performing jitter, delay, and packet loss processing on the first real voice; the codec distortion processing uses a media encoder to perform codec compression simulation on the first real voice.

[0056] Specifically, as Figure 2 shown, the auxiliary cross-domain data generation 202: During the voice transmission process, the voice signal will be affected by the distortion introduced by the transmission codecs 221, 225 and transmission-related problems (such as jitter 222, delay 223, and packet loss 224). During the voice compression process, corresponding distortions will be generated according to the media codecs 226, 227 and their specific configurations. In the above two scenarios, different transmission codecs, different transmission-related problems, and different media codecs will cause the voice to generate corresponding distortions, thereby causing domain shift in the voice data, that is, being distributed in different domains.

[0057] Therefore, this embodiment uses the following detailed methods to simulate the generation of cross-domain voice data in the voice transmission scenario. First, given the original voice signal with a length of data points, it is divided into data packets, and each data packet contains data points:

[0058] , where , , represents the th data point in the th data packet.

[0059] And jitter 222 usually manifests as fluctuations in the arrival time of data packets. To simulate jitter, we can add a random offset to the timestamp of each data packet. The original timestamp of each data packet is defined as , then the timestamp after jitter can be calculated by the following formula:

[0060] ;

[0061] Where represents the random jitter offset for each data packet. The data packets are sorted in ascending order according to , and the sorted data packets are reorganized to form the jittered voice signal , and the reorganization formula is as follows:

[0062] , where is the index of the sorted data packet.

[0063] The delay 223 refers to the movement of the entire signal relative to its original position on the time axis. We simulate the delay by shifting the original voice signal to the right by time stamps, as shown below: .

[0064] Packet loss 224 refers to the complete removal of some data packets in the signal. We define a packet loss rate , and for each original voice data packet , we define a random variable that follows a Bernoulli distribution:

[0065] ,

[0066] where indicates that the data packet is removed, indicates that the data packet is retained, and then the voice signal after packet loss processing can be obtained by the following formula:

[0067] ,

[0068] where is an indicator function that takes the value of 1 when the condition in the parentheses holds, otherwise 0.

[0069] In this embodiment, the voice data is encoded 226 and decoded 227 using one or more voice media codecs (including MP3, M4A, and OGG) to generate domain-specific data for the compression codec in a flexible combination manner, where the M4A codec uses Advanced Audio Coding (AAC), and each codec is configured with multiple bitrate settings.

[0070] Specifically, the bitrate settings for the MP3 codec cover a range from 32 kbps to 128 kbps, including but not limited to the following common values: 32 kbps, 64 kbps, and 128 kbps; the bitrate settings for the M4A codec cover a range from 16 kbps to 64 kbps; the bitrate settings for the OGG codec cover a range from 45 kbps to 96 kbps. These bitrate values cover a range from low bitrates to high bitrates, enabling the simulation of voice compression in different scenarios.

[0071] So far, this application generates an auxiliary voice 203 only from the first real voice by giving a source voice dataset 201 containing the first real voice and the forged voice, using the prior knowledge of voice transmission and compression, and simulating the voice transmission process and the codec process of voice compression. The auxiliary voice 203 constitutes the second real voice, which is obtained by performing the above simulation process on the first real voice, thus ensuring that the second real voice has the same speaker domain as the first real voice, but introducing the distortion generated by voice transmission and voice compression, thereby forming a cross-domain dataset.

[0072] After completing the conversion of the auxiliary data, both the source voice and the auxiliary voice are subjected to frame-level feature extraction by a self-supervised feature extractor, and the extracted frame-level features are projected into a low-dimensional feature space through a projection network to obtain a first feature vector and a second feature vector respectively.

[0073] Specifically, as Figure 3 shown, the feature extraction process utilizes a self-supervised pre-trained model such as Wav2Vec 2.0. The self-supervised pre-trained model is set to run inside the self-supervised feature extractor. These models are pre-trained on a large unlabeled voice corpus. In process 206, the self-supervised feature extractor consists of a convolutional neural network and multiple stacked Transformer encoder layers.

[0074] The role of the convolutional neural network is to convert the input raw waveform into a series of hidden layer features. This process can effectively extract local features in the voice signal, such as the time-frequency structure and energy distribution of the voice, providing richer information for subsequent feature processing; while the role of the Transformer encoder layer is to convert these hidden layer feature sequences into an output frame-level feature sequence. Through the self-attention mechanism, the Transformer encoder layer can capture the long-range dependencies of voice features in the time series, further enhancing the representational ability of the features, enabling the features to better reflect the semantics and context information of the voice.

[0075] As Figure 3As shown, the projection network first performs average pooling on the frame-level features to obtain representative features representing the entire speech. Then, a feed-forward neural network is used to reduce the dimension of the representative features, thereby projecting the representative features into a low-dimensional feature space.

[0076] In this application, the source speech is input as the original waveform, so that a series of frame-level features are output through the self-supervised pre-training model. Then, through the projection network, the frame-level feature sequence is projected into the low-dimensional feature space to obtain the first feature vector. At the same time, the auxiliary speech is also input as the original waveform, and the corresponding frame-level features are output through the self-supervised pre-training model. Then, through the projection network, the corresponding frame-level feature sequence is projected into the low-dimensional feature space to obtain the second feature vector.

[0077] Next, one-class learning is performed based on the first feature vector, and the one-class classification loss function value is calculated. Domain-invariant learning is performed based on the second feature vector, and the cross-entropy loss function value is calculated.

[0078] This application uses the one-class learning method to learn the compact distribution of the first true speech in the low-dimensional feature space 208, and at the same time pushes the forged speech away from the first true speech, so as to facilitate the separation of the forged speech. The construction of the one-class classifier is as Figure 3 shown. First, the input features are normalized, and the target vector is weighted and normalized. Then, the cosine similarity between the input features and the target vector is calculated. Finally, the positive and negative class scores are calculated. Define the self-supervised feature extractor mapping 206 to map the input speech into frame-level features; define the projection network 207 to map the input frame-level features into a dimensional first feature vector in the low-dimensional feature space 208 ; for the input source speech sample , the first feature vector is , and its label represent the first true speech and the forged speech respectively. Then, the one-class classification loss function value is calculated, and its calculation formula is:

[0079] ,

[0080] where, represents the batch size extracted from the source speech for training at one time; represents the scaling factor; represents the true label corresponding to the th data point; represents the th first feature vector corresponding to the source speech; represents the optimization direction of the first true speech embedding vector; and are respectively normalized to and ; , when , is expressed as , when , is expressed as , and when > is used to constrain the angle between and to form a boundary between the first real speech and the forged speech distribution.

[0081] On the other hand, when performing domain-invariant learning, domain-invariant learning is performed on the auxiliary speech based on the second feature vector, and the cross-entropy loss function value is calculated, including the following steps: setting more than one set of adversarial domain classifiers, mapping the second feature vector through the adversarial domain classifier to more than one set of labels, and calculating the cross-entropy loss function value.

[0082] In this embodiment, taking the adoption of two sets of adversarial domain classifiers as an example, the learning process of promoting the model to perform domain-invariant features is described. When two sets of adversarial domain classifiers are set, the two sets of adversarial classifiers are set as the speaker classifier and the conditional classifier, so that the model performs domain-invariant representation learning under different speakers and different conditions. These two sets of domain classifiers adopt the same structure, as shown in Figure 3 the process 211 / 212 in 206 defines a self-supervised feature extractor mapping to map the input speech into frame-level features; defining a projection network 207 to map the input frame-level features into a dimensional second feature vector in the low-dimensional feature space 208 ; then, defining a speaker domain classifier mapping 209, 211 (or a conditional domain classifier mapping 210, 212) to map the second feature vector to the speaker domain label (or the conditional domain label ). For the input auxiliary speech sample , the second feature vector is

[0083] First, map the second feature vector to the speaker label through the speaker classifier, and calculate the first cross-entropy loss function value. The calculation formula is as follows:

[0084] ;

[0085] Then, map the second feature vector to the conditional label through the conditional classifier, and calculate the value of the second cross-entropy loss function. The calculation formula is as follows:

[0086] ;

[0087] Among them, among them, represents the batch size for training extracted from the auxiliary speech at one time; represents the i-th auxiliary speech corresponding to the second feature vector in the low-dimensional feature space; represents the number of speaker domain categories; represents the number of conditional domain categories; represents corresponding to the true label of the -th domain in the speaker domain; represents corresponding to the true label of the -th domain in the conditional domain; represents the calculation of the gradient reversal layer; represents the speaker domain classifier; represents the conditional domain classifier.

[0088] This application trains and to generate robust features that can deceive the domain classifiers and , thereby increasing the cross-entropy loss. At the same time, the domain classifiers and improve their ability to distinguish specific domain features by minimizing this cross-entropy loss; To achieve end-to-end optimization, this application introduces a gradient reversal layer between and (or ), which does not interfere with the forward propagation process, but reverses the gradient by multiplying a negative coefficient during the backpropagation process.

[0089] Finally, as Figure 1 shown, in the forward propagation processes 204 and 205, the source speech data passes through the self-supervised feature extractor 206, the projection network 207, and the one-class classifier 213, and finally calculates the one-class classification loss; The auxiliary cross-domain data (auxiliary speech) passes through the self-supervised feature extractor 206, the projection network 207, the gradient reversal layers 209 and 210, and the speaker classifier 211 or the conditional classifier 212, where the cross-entropy loss is calculated. Finally, based on the one-class classification loss function value and the cross-entropy loss function value, the total loss function value is calculated to obtain: Among them and are hyperparameters used to balance the trade - off between different objectives during training.

[0090] During the backpropagation processes 216 and 217, the parameters of the one - class classifier are updated only through the gradients calculated from ; the parameters of the speaker classifier and the conditional classifier are updated only through the gradients calculated from their respective cross - entropy loss parts ( and ); the parameters of the self - supervised feature extractor and the projection network are updated through the gradients calculated from the total loss .

[0091] After sufficient training, the second true speech in different domains becomes indistinguishable in the low - dimensional feature space 208, thus making various speech features more compactly distributed in the low - dimensional feature space.

[0092] After completing the above gradient update, the test speech in the test set can be input into the self - supervised feature extractor, projection network, and one - class classifier after gradient update to obtain the speech detection value, and the speech detection result is determined based on the speech detection value. Specifically, determining the speech detection result based on the speech detection value includes the following steps: setting a speech threshold and determining whether the speech detection value is greater than the speech threshold; if the speech detection value is greater than the speech threshold, it is determined that the test speech is true speech, otherwise it is determined that the test speech is forged speech, thus completing the final forged speech detection.

[0093] This application, through the fusion of auxiliary speech generation, self - supervised learning feature extraction, one - class learning, and domain - invariant learning, enables the model for detection to effectively handle the changes of speech in different environments, and solves the problem that the false alarm rate of traditional methods often increases due to domain adaptation and environmental distortion.

[0094] Embodiment 2

[0095] A cross - domain deep - fake speech detection system includes an auxiliary cross - domain data generation unit, a self - supervised feature extraction unit, a projection network unit, a one - class learning unit, a domain - invariant learning unit, a gradient update unit, and a test unit;

[0096] The auxiliary cross - domain data generation unit is used to perform auxiliary cross - domain processing on the source speech in the training set to generate auxiliary speech;

[0097] The self - supervised feature extraction unit is used to perform frame - level feature extraction on the source speech and the auxiliary speech respectively;

[0098] A projection network unit for projecting the source speech and the auxiliary speech after frame-level feature extraction into a low-dimensional feature space respectively to obtain a first feature vector and a second feature vector;

[0099] A one-class learning unit for performing one-class learning based on the first feature vector and calculating the value of the one-class classification loss function;

[0100] A domain-invariant learning unit for performing domain-invariant learning based on the second feature vector and calculating the value of the cross-entropy loss function;

[0101] A gradient update unit for calculating the value of the total loss function based on the value of the one-class classification loss function and the value of the cross-entropy loss function, and performing gradient update on the one-class classifier of the one-class learning using the value of the one-class classification loss function, performing gradient update on the adversarial domain classifier of the domain-invariant learning using the value of the cross-entropy loss function, and performing gradient update on the parameters of the self-supervised feature extractor and the projection network using the value of the total loss function;

[0102] A testing unit for inputting the test speech in the test set into the self-supervised feature extractor, the projection network and the one-class classifier after gradient update to obtain a speech detection value, and determining the speech detection result based on the speech detection value.

[0103] Since the deepfake speech detection system of this embodiment executes a cross-domain deepfake speech detection method as described in Embodiment 1, it will not be repeated here.

[0104] Embodiment 3

[0105] A computer-readable storage medium stores a computer program, which when executed by a processor, implements a cross-domain deepfake speech detection method as described in Embodiment 1.

[0106] More specific examples of the computer-readable storage medium may include, but are not limited to: electrical connections with one or more wire segments, portable computer disks, hard disks, random access memory (RAM), read-only memory (ROM), erasable programmable read-only memory (EPROM or flash memory), optical fibers, portable compact disk read-only memory (CD-ROM), optical storage devices, magnetic storage devices, or any suitable combination of the above.

[0107] In this application, a computer-readable storage medium may be any tangible medium that contains or stores a program, which can be used by or in conjunction with an instruction execution system, apparatus, or device. And in this application, a computer-readable signal medium may include a data signal propagated in a baseband or as part of a carrier wave, which carries computer-readable program code. Such a propagated data signal may take various forms, including but not limited to electromagnetic signals, optical signals, or any suitable combination of the above. A computer-readable signal medium may also be any computer-readable medium other than a computer-readable storage medium, which can send, propagate, or transmit a program for use by or in conjunction with an instruction execution system, apparatus, or device. The program code contained on a computer-readable medium may be transmitted using any appropriate medium, including but not limited to: wireless segments, wire segments, optical cables, RF, etc., or any suitable combination of the above.

[0108] In several embodiments provided in this application, it should be understood that the disclosed apparatus and method can be implemented in other ways. For example, the apparatus embodiments described above are merely illustrative. For example, the division of the modules, units, or components is only a logical function division. In actual implementation, there may be other division methods. For example, multiple units, modules, or components can be combined or integrated into another apparatus, or some features can be ignored or not executed.

[0109] The unit may or may not be physically separated. The components shown as units may be one physical unit or multiple physical units, that is, they may be located in one place, or may be distributed to multiple different places. Some or all of the units can be selected according to actual needs to achieve the purpose of the solution of this embodiment.

[0110] In addition, in each embodiment of the present invention, the functional units may be integrated in a processing unit, or each unit may exist physically separately, or two or more units may be integrated in one unit. The above integrated units may be implemented in the form of hardware or in the form of software functional units.

[0111] Specifically, according to the embodiments disclosed in the present invention, the processes described above with reference to the flowcharts can be implemented as computer software programs. For example, the embodiments of the present disclosure include a computer program product, which includes a computer program carried on a computer-readable medium, and the computer program contains program code for executing the method shown in the flowchart. In such an embodiment, the computer program can be downloaded and installed from the network through the communication part, and / or installed from a removable medium.

[0112] When the computer program is executed by a central processing unit (CPU), the above functions defined in the method of the present application are executed. It should be noted that the above computer-readable medium of the present application can be a computer-readable signal medium, a computer-readable storage medium, or any combination of the two. The computer-readable storage medium can be, for example, but not limited to, an electrical, magnetic, optical, electromagnetic, infrared, or semiconductor system, apparatus, or device, or any combination of the above.

[0113] The flowcharts and block diagrams in the accompanying drawings illustrate the possible architectures, functions, and operations of systems, methods, and computer program products according to various embodiments of the present invention. In this regard, each block in the flowchart or block diagram may represent a module, a program segment, or a part of code that contains one or more executable instructions for implementing the specified logical function. It should also be noted that in some alternative implementations, the functions marked in the blocks may occur in a different order than marked in the accompanying drawings. For example, two consecutive blocks shown may actually be executed substantially in parallel, and they may sometimes be executed in the reverse order, depending on the functions involved. It should also be noted that each block in the block diagram and / or flowchart, and the combination of blocks in the block diagram and / or flowchart, can be implemented by a dedicated hardware-based system for performing the specified functions or operations, or can be implemented by a combination of dedicated hardware and computer instructions.

[0114] As described above, the above are only specific embodiments of the present invention, but the protection scope of the present invention is not limited thereto. Any changes or substitutions within the technical scope disclosed by the present invention should be covered by the protection scope of the present invention. Therefore, the protection scope of the present invention should be subject to the protection scope of the claims.

Claims

1. A cross-domain deep fake voice detection method, characterized in that: The following steps are involved: Perform auxiliary cross-domain processing on the source speech in the training set to generate auxiliary speech; Extracting frame-level features from both the source speech and the auxiliary speech through a self-supervised feature extractor, and projecting the extracted frame-level features into a low-dimensional feature space through a projection network to obtain a first feature vector and a second feature vector, respectively; Performing single-class learning on the source speech based on the first feature vector and calculating the single-class loss function value, performing domain-invariant learning on the auxiliary speech based on the second feature vector and calculating the cross-entropy loss function value; Calculate the total loss function value based on the single classification loss function value and the cross entropy loss function value; Use the single classification loss function value to perform gradient updates on the single classifier of single-class learning, use the cross entropy loss function value to perform gradient updates on the adversarial domain classifier of domain-invariant learning, and use the total loss function value to perform gradient updates on the parameters of the self-supervised feature extractor and projection network; The test speech in the test set is input into the self-supervised feature extractor, the projection network and the single classifier after gradient update to obtain the speech detection value, and the speech detection result is determined based on the speech detection value.

2. A cross-domain deep fake voice detection method according to claim 1, characterized in that: The auxiliary cross-domain processing includes the following steps: The source voice includes a first real voice and a forged voice; Performing transmission distortion processing on the first real speech to obtain a real speech after transmission distortion; The first real speech is subjected to encoding and decoding distortion processing to obtain a compressed and distorted real speech.

3. A cross-domain deep fake voice detection method according to claim 2, characterized in that: The transmission distortion processing includes: performing jitter, delay and packet loss processing on the first real voice; The codec distortion processing uses a media encoder to perform codec compression simulation on the first real speech.

4. A cross-domain deep fake voice detection method according to claim 1, characterized in that: The calculation formula of the single classification loss function value is: , in, Indicates the number of batches extracted from the source speech for training at one time; represents the scaling factor; Indicates The true label corresponding to the data point; Indicates The first feature vector corresponding to the source speech; represents the optimization direction of the first real speech embedding vector; and They are normalized to and ; ,when hour, Expressed as ,when hour, Expressed as , and when> When used to constrain and The angle between them forms the boundary between the first real speech and the forged speech distribution.

5. A cross-domain deep fake voice detection method according to claim 1, characterized in that: Performing domain-invariant learning on the auxiliary speech based on the second feature vector and calculating the cross entropy loss function value includes the following steps: One or more adversarial domain classifiers are set, and the second feature vector is mapped to one or more labels through the adversarial domain classifier, and the cross entropy loss function value is calculated.

6. A cross-domain deep fake voice detection method according to claim 5, characterized in that: When two groups of adversarial domain classifiers are set, the two groups of adversarial classifiers are set as speaker classifiers and conditional classifiers, and the cross entropy loss function value is calculated, including the following steps: The second feature vector is mapped to the speaker label through the speaker classifier, and the first cross entropy loss function value is calculated. The calculation formula is as follows: ; The second feature vector is mapped to the conditional label through the conditional classifier, and the second cross entropy loss function value is calculated. The calculation formula is as follows: ; in, Indicates the number of batches extracted from the auxiliary speech for training at one time; Indicates the i-th auxiliary voice The corresponding second eigenvector in the low-dimensional feature space; Indicates the number of speaker domain categories; Indicates the number of conditional domain categories; express The corresponding speaker domain The true labels of the domains; express The corresponding condition domain The true labels of the domains; Represents the gradient flip layer calculation; represents the speaker domain classifier; Represents a conditional domain classifier.

7. A cross-domain deep fake voice detection method according to claim 6, characterized in that: The formula for calculating the total loss function value is: , where λ1 and λ2 are hyperparameters used to balance the trade-offs between different objectives during training.

8. A cross-domain deep fake voice detection method according to claim 1, characterized in that: Determining a speech detection result based on the speech detection value comprises the following steps: Setting a speech threshold and determining whether the speech detection value is greater than the speech threshold; If the voice detection value is greater than the voice threshold, the test voice is determined to be a real voice, otherwise the test voice is determined to be a forged voice.

9. A cross-domain deep fake voice detection system, characterized in that: The deep fake voice detection system performs a cross-domain deep fake voice detection method as described in any one of claims 1 to 8, including an auxiliary cross-domain data generation unit, a self-supervised feature extraction unit, a projection network unit, a single-class learning unit, a domain-invariant learning unit, a gradient update unit, and a testing unit; The auxiliary cross-domain data generating unit is used to perform auxiliary cross-domain processing on the source speech in the training set to generate auxiliary speech; The self-supervisory feature extraction unit is used to perform frame-level feature extraction on the source speech and the auxiliary speech respectively; The projection network unit is used to project the source speech and the auxiliary speech after the frame-level feature extraction into the low-dimensional feature space to obtain the first feature vector and the second feature vector respectively; The single-class learning unit is used to perform single-class learning based on the first feature vector and calculate a single-class loss function value; The domain invariant learning unit is used to perform domain invariant learning based on the second feature vector and calculate a cross entropy loss function value; The gradient updating unit is used to calculate the total loss function value based on the single classification loss function value and the cross entropy loss function value, and use the single classification loss function value to perform gradient update on the single classifier of single-class learning, use the cross entropy loss function value to perform gradient update on the adversarial domain classifier of domain-invariant learning, and use the total loss function value to perform gradient update on the parameters of the self-supervised feature extractor and the projection network; The testing unit is used to input the test speech in the test set into the self-supervised feature extractor, the projection network and the single classifier after gradient update to obtain the speech detection value, and determine the speech detection result based on the speech detection value.

10. A computer-readable storage medium storing a computer program, characterized in that: When the computer program is executed by a processor, a cross-domain deep fake voice detection method as described in any one of claims 1-8 is implemented.