Audio depth forgery detection method and device

By using style alignment and Poincaré sphere model structured regularization in the deep audio forgery detection model, the problem of poor detection performance of audio data with diverse styles in the existing technology is solved, and more efficient audio forgery detection is achieved.

CN120833787APending Publication Date: 2025-10-24ALIPAY (HANGZHOU) INFORMATION TECH CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510958128.7
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-07-11
Publication Date
2025-10-24

AI Technical Summary

Technical Problem

Existing audio deepfake detection methods perform poorly when dealing with audio data with diverse styles.

Method used

A deep audio forgery detection model is adopted, including a first encoder, a second encoder, a style alignment module, and a classifier. The audio samples are style-aligned using a style library, and the features are mapped to a Poincaré sphere model. A structured regularization term is constructed to describe the hierarchical relationship, and the model parameters are iteratively updated.

Benefits of technology

It improves the generalization ability and robustness of the audio deepfake detection model, and enhances the detection accuracy of audio data of different styles.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120833787A_ABST
    Figure CN120833787A_ABST
Patent Text Reader

Abstract

The embodiment of the invention discloses an audio deep counterfeiting detection method and device. According to the method, audio styles of audio samples of a source domain are learned through a learnable style library, then the style library is used for carrying out style alignment on the audio samples, depth features of the audio samples subjected to style alignment are mapped to a Poincare sphere model, classification is carried out through a classifier, and the audio samples in the source domain are obtained. A hierarchical structure of a classification result of a classifier in a Poincare sphere model is described by constructing a structured regular term, so that the model learns an internal structure and a decision boundary of data. By means of the training method, the audio deep forgery detection model can have better generalization ability and robustness, and the accuracy of the detection result of the audio deep forgery detection model is improved. The training device of the audio depth pseudo-depth modeling detection model, the audio depth counterfeiting detection method and the audio depth counterfeiting detection device in the embodiment of the specification also have the above beneficial effects.
Need to check novelty before this filing date? Find Prior Art

Description

TECHNICAL FIELD

[0001] The present application relates to the technical field of artificial intelligence, and in particular to an audio deepfake detection method and device. BACKGROUND

[0002] Existing audio deepfake detection methods mainly fall into two categories: one is an audio deepfake detection method based on data enhancement and domain adaptation, and the other is an audio deepfake detection method based on domain-invariant feature extraction. These two methods do not work well when dealing with audio data with varying styles. SUMMARY

[0003] One or more embodiments of the present specification provide an audio deepfake detection method and device, which can at least partially solve the above technical problems.

[0004] In a first aspect, a training method of a deep audio fake detection model is provided, the deep audio fake detection model comprising a first encoder, a second encoder, a style alignment module and a classifier; the method comprising:

[0005] obtaining an audio sample of a source domain;

[0006] extracting shallow features of the audio sample using the first encoder;

[0007] training a style library using the shallow features of the audio sample, to obtain a style library capable of covering the audio style of the source domain;

[0008] aligning the shallow features of the audio sample with the style library using the style alignment module;

[0009] extracting deep features of the audio sample after style alignment using the second encoder, and mapping the extracted deep features to a Poincare ball model to obtain hyperbolic space features;

[0010] assigning the hyperbolic space features to a data prototype using the classifier, assigning the data prototype to a top-level prototype, and constructing a structured regularization term to describe the hierarchical relationship of the assignment result; the data prototype is used to represent the category of the audio sample, and the top-level prototype is a higher-level category of the data prototype;

[0011] iteratively updating the parameters of the first encoder, the second encoder and the classifier, and the positions of the data prototype and the top-level prototype, with the purpose of minimizing the structured regularization term, until the deep audio fake detection model satisfying the preset condition is obtained.

[0012] As an optional implementation of the method of the first aspect, training the style library using the audio sample specifically comprises:

[0013] initializing a set of base style vectors to constitute the style library;

[0014] extracting a style vector of the audio sample from shallow features of the audio sample;

[0015] determining similarity between the style vector of the audio sample and each base style vector in the style library; normalizing the similarity into a weight coefficient, and performing weighted summation on the base style vectors in the style library according to the weight coefficient to obtain a new style vector;

[0016] constructing an orthogonal style loss function based on the base style vectors;

[0017] constructing a style reconstruction loss function based on the new style vector and the style vector;

[0018] updating the base style vectors by using the orthogonal style loss function and the style reconstruction loss function until a style library capable of covering audio styles of the source domain is obtained.

[0019] Further, extracting a style vector of the audio sample from shallow features of the audio sample, specifically comprising:

[0020] extracting channel mean and channel standard deviation of each channel from the shallow features of the audio sample as the style vector of the audio sample.

[0021] As an optional implementation of the method of the first aspect, the style alignment module utilizes the style library to perform style alignment on the shallow features of the audio sample, specifically comprising:

[0022] extracting a style vector of the audio sample from shallow features of the audio sample;

[0023] determining similarity between the style vector of the audio sample and each base style vector in the style library; normalizing the similarity into a weight coefficient, and performing weighted summation on the base style vectors in the style library according to the weight coefficient to obtain a reconstructed new style vector;

[0024] performing style alignment processing on the style vector of the audio sample by using the new style vector.

[0025] Further, in the training method:

[0026] extracting a style vector of the audio sample from shallow features of the audio sample, specifically comprising: extracting channel mean and channel standard deviation of each channel from the shallow features of the audio sample as the style vector of the audio sample;

[0027] aligning a style vector of the audio sample with the new style vector, specifically including: decomposing the new style vector into a new channel mean and a new channel standard deviation for each channel;

[0028] aligning a corresponding channel feature in a shallow feature of the audio sample with the new channel mean and the new channel standard deviation for the channel.

[0029] As an optional implementation of the method of the first aspect, the method further includes:

[0030] initializing positions of the data prototypes and the top-level prototypes in the Poincare ball model before assigning the hyperbolic space feature of the audio sample to the data prototype using the classifier.

[0031] As an optional implementation of the method of the first aspect, an expression of the structured regularization term is:

[0032]

[0033] wherein H represents a parameter to be learned, Z represents a set of the audio samples, z i represents a hyperbolic space feature of an i-th audio sample, N represents a total number of the audio samples, P c(i) represents a data prototype to which the hyperbolic space feature of the i-th audio sample is assigned, c(i) represents an index of the data prototype to which the hyperbolic space feature of the i-th audio sample is assigned, P j represents a j-th data prototype, represents a top-level prototype to which the j-th data prototype is assigned, M P represents a number of the top-level prototypes.

[0034] As an optional implementation of the method of the first aspect, wherein in each iteration, the hyperbolic space feature of the audio sample is assigned to a data prototype closest to the audio sample using the classifier, and the data prototype is assigned to a top-level prototype closest to the data prototype using the classifier.

[0035] In a second aspect, a training device of a deep audio forgery detection model is provided, the deep audio forgery detection model including: a first encoder, a second encoder, a style alignment module, and a classifier; the device including:

[0036] a first data acquisition module configured to acquire audio samples of a source domain;

[0037] a first feature extraction module configured to extract a shallow feature of the audio sample using the first encoder;

[0038] a style library construction module configured to train a style library using shallow features of the audio samples, to obtain a style library capable of covering audio styles of the source domain;

[0039] a feature reconstruction module configured to perform style alignment on the shallow features of the audio samples by the style alignment module using the style library;

[0040] a second feature extraction module configured to perform deep feature extraction on the style-aligned audio sample features using the second encoder;

[0041] a feature mapping module configured to map the deep features extracted by the second feature extraction module into a Poincare ball model to obtain hyperbolic space features;

[0042] a training module configured to assign the hyperbolic space features of the audio samples to data prototypes, assign the data prototypes to top-level prototypes, and construct a structured regularization term to describe a hierarchical relationship of the assignment results using the classifier; the data prototypes are used to represent categories of the audio samples, and the top-level prototypes are higher-level categories of the data prototypes;

[0043] The training module is further configured to iteratively update parameters of the first encoder, the second encoder, and the classifier, and positions of the data prototypes and the top-level prototypes, with the purpose of minimizing the structured regularization term, until the deep audio forgery detection model satisfying a preset condition is obtained.

[0044] As an optional implementation of the device of the second aspect, the style library construction module is specifically used for:

[0045] initializing a set of base style vectors to constitute the style library;

[0046] extracting a style vector of the audio sample from the shallow features of the audio sample;

[0047] determining a similarity between the style vector of the audio sample and each base style vector in the style library; normalizing the similarity into a weight coefficient, and performing weighted summation on the base style vectors in the style library according to the weight coefficient to obtain a new style vector;

[0048] constructing an orthogonal style loss function based on the base style vectors;

[0049] constructing a style reconstruction loss function based on the new style vector and the style vector;

[0050] updating the base style vectors using the orthogonal style loss function and the style reconstruction loss function, until a style library capable of covering audio styles of the source domain is obtained.

[0051] Further, the style library construction module is specifically configured to:

[0052] extract a channel mean and a channel standard deviation of each channel from the shallow features of the audio sample as a style vector of the audio sample.

[0053] As an optional implementation of the device of the second aspect, the feature reconstruction module is specifically configured to:

[0054] extract a style vector of the audio sample from the shallow features of the audio sample;

[0055] determine a similarity between the style vector of the audio sample and each base style vector in the style library, normalize the similarity into a weight coefficient, and perform weighted summation on the base style vectors in the style library according to the weight coefficient to obtain a reconstructed new style vector;

[0056] perform style alignment processing on the style vector of the audio sample by using the new style vector.

[0057] Further, in the training device described above:

[0058] the style library construction module is specifically configured to extract a channel mean and a channel standard deviation of each channel from the shallow features of the audio sample as a style vector of the audio sample;

[0059] the feature reconstruction module is specifically configured to decompose the new style vector into a new channel mean and a new channel standard deviation of each channel, and perform style alignment on corresponding channel features in the shallow features of the audio sample by using the new channel mean and the new channel standard deviation of each channel.

[0060] As an optional implementation of the device of the second aspect, the training module is further configured to initialize positions of the data prototypes and the top-level prototypes in the Poincare ball model before assigning the hyperbolic space features to the data prototypes by using the classifier.

[0061] As an optional implementation of the device of the second aspect, an expression of the structured regularization term is:

[0062]

[0063] wherein H represents a parameter to be learned, Z represents a set of the audio samples, z i represents a hyperbolic space feature of an i-th audio sample, N represents a total number of the audio samples, P c(i)a data prototype to which the hyperbolic space feature of the i-th audio sample is assigned, c(i) represents an index of the data prototype to which the hyperbolic space feature of the i-th audio sample is assigned, P j represents the j-th data prototype, represents a top-level prototype to which the j-th data prototype is assigned, M P represents the number of top-level prototypes.

[0064] As an optional implementation of the device of the second aspect, the training module is specifically configured to:

[0065] In each iteration, the hyperbolic space feature of the audio sample is assigned to the data prototype closest to the audio sample data by using the classifier, and the data prototype is assigned to the top-level prototype closest to the data prototype by using the classifier.

[0066] In a third aspect, an audio deep forgery detection method is provided, comprising:

[0067] obtaining to-be-recognized audio data in a target domain;

[0068] extracting shallow features of the to-be-recognized audio data by using a first encoder of a deep audio forgery detection model; the deep audio forgery detection model is obtained by pre-training using the above-mentioned audio deep forgery detection model training method;

[0069] aligning the shallow features of the to-be-recognized audio data in style by using a style alignment module of the deep audio forgery detection model;

[0070] extracting deep features of the audio sample features after style alignment by using a second encoder of the deep audio forgery detection model, and mapping the extracted deep features to a Poincare ball model to obtain hyperbolic space features;

[0071] assigning the hyperbolic space features to data prototypes by using a classifier of the deep audio forgery detection model, assigning the data prototypes to top-level prototypes, and finally obtaining a recognition result of the to-be-recognized audio data;

[0072] determining whether the to-be-recognized audio data is a forged audio data according to the recognition result.

[0073] In a fourth aspect, an audio deep forgery detection device is provided, comprising:

[0074] a second data acquisition module configured to obtain to-be-recognized audio data in a target domain;

[0075] The detection module is configured to extract shallow features of the to-be-identified audio data by using a first encoder of a deep audio forgery detection model; perform style alignment on the shallow features of the to-be-identified audio data by using a style alignment module of the deep audio forgery detection model and a style library; perform deep feature extraction on the style-aligned audio sample features by using a second encoder of the deep audio forgery detection model, and map the extracted deep features to a Poincare ball model to obtain hyperbolic space features; assign the hyperbolic space features to data prototypes by using a classifier of the deep audio forgery detection model, assign the data prototypes to top-level prototypes, and finally obtain an identification result of the to-be-identified audio data; and the deep audio forgery detection model is obtained by pre-training according to the training method of the audio deep forgery detection model.

[0076] The determination module is configured to determine whether the to-be-identified audio data is forged audio data according to the identification result.

[0077] In a fifth aspect, a computer-readable storage medium is provided, and the computer-readable storage medium stores a computer program. When the computer program is executed on an electronic device, the electronic device is caused to execute the training method of the audio deep forgery detection model or the audio deep forgery detection method.

[0078] In a sixth aspect, an electronic device is provided, and the electronic device includes:

[0079] at least one memory configured to store a program;

[0080] at least one processor configured to execute the program stored in the memory, and when the program stored in the memory is executed, the processor is configured to execute the training method of the audio deep forgery detection model or the audio deep forgery detection method.

[0081] The training method of the audio deep forgery detection model has the beneficial effect that the method learns the audio style of the audio sample of the source domain by using a learnable style library, and then performs style alignment on the audio sample by using the style library, so that the samples from different domains or with different styles are more consistent when input to the subsequent classifier, thereby providing more robust and more generalizable input features.

[0082] In addition, the method maps the deep features of the style-aligned audio sample to a Poincare ball model and classifies them by using a classifier. The hierarchical structure of the classification results of the classifier in the Poincare ball model is described by constructing a structured regularization term, so that the model learns the internal structure and decision boundary of the data.

[0083] Through the training method, the audio deepfake detection model has better generalization ability and robustness, and the accuracy of the detection result of the audio deepfake detection model is improved.

[0084] The training device of the audio deepfake detection model, the audio deepfake detection method and device provided in the embodiments of the present specification also have the beneficial effects described above. BRIEF DESCRIPTION OF DRAWINGS

[0085] In order to more clearly illustrate the technical solutions in the embodiments of the present specification or the prior art, the drawings needed to be used in the embodiments or the prior art description will be briefly introduced as follows. Obviously, the drawings in the following description are some embodiments of the present specification, and those skilled in the art can also obtain other drawings according to these drawings without creative labor.

[0086] Figure 1 An exemplary flowchart of a training method of an audio deepfake detection model is shown.

[0087] Figure 2 An exemplary structural diagram of an audio deepfake detection model is shown.

[0088] Figure 3 An exemplary flowchart of a training method of a style library is shown.

[0089] Figure 4 An exemplary structural diagram of a training device of an audio deepfake detection model is shown.

[0090] Figure 5 An exemplary flowchart of an audio deepfake detection method is shown.

[0091] Figure 6 An exemplary structural diagram of an audio deepfake detection device is shown.

[0092] Figure 7 An exemplary structural diagram of an electronic device provided by the embodiments of the present specification is shown. DETAILED DESCRIPTION

[0093] First of all, it should be noted that the terms used in the embodiments of the present application are only for the purpose of describing specific embodiments, and are not intended to limit the present application. The singular forms "a", "said" and "the" used in the embodiments of the present application and the appended claims are also intended to include the plural forms, unless the context clearly indicates otherwise.

[0094] In order for those skilled in the technical field to better understand the technical solutions in the specification, the technical solutions in the specification will be clearly and completely described below in combination with the drawings in the specification. Obviously, the described embodiments are only part of the embodiments of the specification, not all. Therefore, those skilled in the art should realize that various changes and modifications can be made to the embodiments described herein without departing from the scope and spirit of the present application. Also, for the sake of clarity and brevity, the description below omits the description of well-known functions and structures.

[0095] It should be noted that the steps of the corresponding method are not necessarily performed in the order shown and described in the specification in other embodiments. In some other embodiments, the steps included in the method can be more or less than described in the specification. In addition, a single step described in the specification can be divided into multiple steps for description in other embodiments; and multiple steps described in the specification can be combined into a single step for description in other embodiments.

[0096] Existing audio deepfake detection methods are mainly divided into two categories: one is an audio deepfake detection method based on data enhancement and domain adaptation, and the other is an audio deepfake detection method based on domain-invariant feature extraction. With the development of generative artificial intelligence, audio data forgery schemes in the industry are emerging in an endless stream, and the styles of the audio data they forge are varied. The above two schemes are not effective in dealing with audio data with varied styles.

[0097] In order to more effectively deal with this problem, one or more embodiments of the specification propose an audio deepfake detection method and device.

[0098] The audio deepfake detection method and device described in one or more embodiments of the specification will be further described in detail below in combination with the drawings and specific embodiments of the specification, but this detailed description does not constitute a limitation on the embodiments of the specification.

[0099] Please refer to Figure 1 , Figure 1 The flowchart of a training method (hereinafter referred to as the training method) for an audio deepfake detection model proposed in one or more embodiments of the specification is shown in Figure 2 The structure of the audio deepfake detection model includes a first encoder, a second encoder, a style alignment module, and a classifier. Figure 1 As shown in

[0100] S100: Obtain audio samples of the source domain.

[0101] The audio sample data described above includes different styles of fake audio samples and real audio samples.

[0102] The style here can include but is not limited to:

[0103] Changes in recording conditions, such as different types of microphones, room acoustic characteristics, background noise, etc.

[0104] The acoustic characteristics of the speaker, such as voice line, speech rate, timbre, etc. can synthesize the specific traces of the fake through the general acoustic model.

[0105] The fake audio samples described above can also be classified according to their fake methods, such as fake audio samples using WaveNet, fake audio samples using Tacotron, fake audio samples using VALL-E, etc. Real audio samples can also be classified according to their characteristics, such as male voice audio samples, female voice audio samples, etc.

[0106] S102: Extracting shallow features of the audio sample using the first encoder.

[0107] After obtaining the audio sample, the audio sample needs to be shallowly feature extracted. It should be noted that the first encoder for feature extraction of the audio sample can be adaptively selected according to the needs, and the present embodiment does not limit this. For example, the wav2vec model can be used to shallowly feature extract the audio sample to obtain the shallow features of the audio sample.

[0108] S104: Training the style library using the shallow features of the audio sample to obtain a style library that can cover the audio styles of the source domain.

[0109] The style library described above is composed of a set of learnable base style vectors, and the parameters of these base style vectors are constantly updated during the training process, eventually enabling these base style vectors to learn various audio styles in the source domain.

[0110] Please refer to Figure 3 , Figure 3 A flowchart showing a training method of a style library is shown. As Figure 3 shown, the training method of the style library described above includes steps S300 to S310:

[0111] S300: Initialize a set of base style vectors to constitute a style library.

[0112] Specifically, suppose the shallow features of the input audio sample are C is the number of channels, T is the time dimension, and the number of base style vectors is K. The set of base style vectors described above can be represented as: B={b1,b2,…,b K}, wherein The base style vector is randomly initialized at the beginning of training.

[0113] S302: Extract the style vector of the audio sample from the shallow feature of the audio sample.

[0114] In the specific implementation, the parameters of the shallow feature can be selected to form the above-mentioned style vector according to the requirements, and the embodiment does not limit this.

[0115] In some embodiments, the channel mean and channel standard deviation of the shallow feature can be selected to construct the above-mentioned style vector.

[0116] Specifically, the expression of the channel mean is:

[0117]

[0118] wherein μ c represents the channel mean of the channel c in the shallow feature of the audio sample, F c,t represents the channel feature of the channel c in the shallow feature of the audio sample at time t.

[0119] The expression of the channel standard deviation is:

[0120]

[0121] wherein σ c represents the channel standard deviation of the channel c in the shallow feature of the audio sample, and ε is a small constant to prevent division by zero.

[0122] The style vector of the audio sample is represented as: s in =[μ1,…,μ C ,σ1,…,σ C ].

[0123] S304: Determine the similarity between the style vector of the audio sample and each base style vector in the style library; normalize the similarity into a weight coefficient, and perform weighted summation on the base style vectors in the style library according to the weight coefficient to obtain a new style vector.

[0124] In some embodiments, the similarity between the style vector of the audio sample and the base style vector in the style library can be represented by the cosine similarity. Specifically, the cosine similarity calculation formula is:

[0125]

[0126] After obtaining the cosine similarity, the Softmax function can be used to normalize the cosine similarity into a weight coefficient to ensure that the weight is positive and the sum is 1. The expression of the weight coefficient is:

[0127]

[0128] According to the weight coefficient w k The base style vectors in the style library are weighted and summed to obtain a new style vector S new is:

[0129]

[0130] S306: Construct an orthogonal style loss function based on the base style vectors.

[0131] The orthogonal style loss aims to encourage the base style vectors in the style library to be mutually orthogonal, so as to ensure that they represent different, non-redundant style dimensions.

[0132] In this step, the orthogonal style loss function used is:

[0133]

[0134] where I is the identity matrix, denotes the square of the norm, and this loss will promote B T tend to the identity matrix, so that the dot product between the base style vectors approaches 0 (orthogonal) and the modulus length approaches 1 (normalized).

[0135] S308: Construct a style reconstruction loss function based on the new style vector and the style vector.

[0136] The style reconstruction loss aims to ensure that the new style vector (s new ) generated by the weighted sum of the base style vectors can effectively reconstruct or approximate the original style vector (s in ). This guarantees that the learned style library has sufficient expressive power to cover the audio styles in the source domain audio samples.

[0137] In this step, the style reconstruction loss function used is:

[0138]

[0139] where denotes the square of the Euclidean distance, i.e. L recon is an L2 loss function.

[0140] S310: Update the base style vectors using the orthogonal style loss function and the style reconstruction loss function until a style library is obtained that can cover the audio styles of the source domain.

[0141] Specifically, the orthogonal style loss function and the style reconstruction loss function can be weighted and summed to obtain a total loss function, and the parameters of the base style vector are updated using the total loss function until the preset convergence condition is met.

[0142] S106: The style alignment module aligns the shallow features of the audio sample with the style library.

[0143] First, the style vector of the audio sample needs to be extracted from the shallow features of the audio sample. Then, the similarity between the style vector of the audio sample and each base style vector in the style library is determined, the similarity is normalized into a weight coefficient, and the base style vectors in the style library are weighted and summed according to the weight coefficient to obtain a reconstructed new style vector. The specific process of the reconstructed new style vector can refer to the specific steps in steps S302 to S304, which will not be repeated here.

[0144] Finally, the new style vector is used to perform style alignment processing on the style vector of the audio sample.

[0145] In some embodiments, the style alignment processing of the new style vector and the style vector of the audio sample can be implemented by adaptive instance normalization (Adaptive Instance Normalization, AdalN).

[0146] Specifically, the new style vector can be decomposed into a new channel mean and a new channel standard deviation corresponding to each channel according to the channels of the shallow features of the audio sample. The new channel mean can be represented as μ new = s new [1…C], and the new channel standard deviation can be represented as σ new = s new [1+C…2C].

[0147] Then, the corresponding channel features in the shallow features of the audio sample are aligned using the new channel mean and the new channel standard deviation of each channel:

[0148]

[0149] where F aligned,c represents the channel feature of the channel c after style alignment.

[0150] S108: The second encoder is used to extract deep features from the style-aligned audio sample features, and the extracted deep features are mapped to a Poincare ball model to obtain hyperbolic space features.

[0151] In this step, the style-aligned audio sample features can be subjected to deep feature extraction by a second encoder, so as to reduce the dimensionality of the features for subsequent structural analysis in the Poincare ball model.

[0152] Next, the extracted deep features are mapped from the Euclidean space to the Poincare ball model to obtain the hyperbolic space features of the audio samples.

[0153] It should be noted that the structure of the second encoder described above can be adaptively selected according to requirements, and the present embodiment does not limit this.

[0154] S110: Assigning the hyperbolic space features to data prototypes using a classifier, assigning the data prototypes to top-level prototypes, and constructing a structured regularization term to describe the hierarchical relationship of the assignment results.

[0155] Before performing this step S110, the positions of the data prototypes and the top-level prototypes need to be initialized in the Poincare ball model in advance. The data prototypes and the top-level prototypes are constructed according to the labels of the audio samples, wherein the data prototypes are used to represent fine-grained categories, and the top-level prototypes refer to higher-level categories than the data prototypes. The top-level prototypes can have multiple layers, and the highest layer is the global top-level prototype. In the spatial structure, the data prototypes can be understood as child nodes, and the top-level prototypes are parent nodes at higher levels in the hierarchical structure. The top-level prototypes organize multiple related data prototypes (or lower-level top-level prototypes) together to form a tree structure, and the root node of this tree structure is the global top-level prototype.

[0156] For example, if the data prototypes P1, P2, P3 represent different deep fake technologies (such as WaveNet, Tacotron, VALL-E), then a top-level prototype Tfake can represent the macro category of "deep fake audio", which is the parent node of P1, P2, and P3. Similarly, if P4 and P5 represent different types of real audio data (such as male voice and female voice), then another top-level prototype Treal can represent all "real audio data".

[0157] The data prototypes and the top-level prototypes can be randomly initialized in the space of the Poincare ball model, or the hyperbolic space features of the audio samples can be preliminarily clustered (such as K-means clustering), and the positions of the data prototypes and the top-level prototypes are initialized according to the results of the preliminary clustering.

[0158] The above-mentioned assignment of the hyperbolic space features of the audio samples to the data prototypes by the classifier means that the classifier is used to determine which data prototype the hyperbolic space features of the audio samples belong to.

[0159] In some embodiments, hard assignment or soft assignment can be employed to realize the assignment of the hyperbolic space feature of the audio sample to the data prototype.

[0160] Hard assignment refers to, in each iteration, assigning the hyperbolic space feature z i of the audio sample to the data prototype P j closest to it in the Poincare ball model, which can be realized by calculating the hyperbolic distance d H (z i , P j ) between the hyperbolic space feature of the audio sample and the data prototype, and selecting the minimum value of d H .

[0161] Soft assignment refers to calculating the probability distribution of the hyperbolic space feature z i of the audio sample belonging to each data prototype P j using a similarity or attention mechanism. For example, a softmax layer can be used to generate the assignment probability softmax(-similarity(z i , P i )) of z j , where similarity(z i , P j ) represents the similarity between the hyperbolic space feature z i of the audio sample and the data prototype P j , which is negatively correlated with the distance between z i and P j .

[0162] The above-mentioned assignment of the data prototype to the top-level prototype using the classifier refers to using the classifier to determine which top-level prototype the data prototype belongs to, which is the process of establishing the parent-child relationship between the data prototype and the top-level prototype.

[0163] In some embodiments, hard assignment or soft assignment can be employed to realize the assignment of the data prototype to the top-level prototype.

[0164] Hard assignment refers to, in each iteration, assigning the data prototype P j to the top-level prototype T k closest to it in the Poincare ball model, which can be realized by calculating the hyperbolic distance d H (P j , T k ) between the data prototype P j and the top-level prototype T k , and selecting the minimum value of d H .

[0165] Soft assignment means using similarity or attention mechanism to assign the data prototype P j Count the top-level prototypes it belongs to k For example, a softmax layer can be used to generate the data prototype P j The probability of assignment softmax(-similarity(P j ,T k )), where similarity(P j ,T k ) represents the data prototype P j With the top prototype T k The similarity between j and T k The distance between them is negatively correlated.

[0166] In some more specific implementations, some high-level relationships can also be pre-defined. For example, if it is known that certain data prototypes (such as various counterfeit technology prototypes) belong to the general category of "counterfeiting", then these data prototypes can be assigned to the top-level prototype representing "counterfeiting" at the beginning.

[0167] After the above allocation, the hyperbolic space features of the audio samples are allocated to the data prototypes, the data prototypes are allocated to the top prototypes, and the top prototypes are allocated to the top prototypes at a higher level than them. Finally, a tree-like hierarchical allocation result can be obtained.

[0168] It should be noted that the model structure of the above-mentioned classifier can be adaptively selected according to needs, and this embodiment does not impose any limitation on this.

[0169] After obtaining the classification results of the classifier, a structured regularization term can be constructed to describe the hierarchical relationship of the allocation results, that is, to describe the hierarchical relationship between the nodes in the allocation results of the tree hierarchy.

[0170] Specifically, the above structured regularization term can be described by the following expression:

[0171]

[0172] Among them, R structure (H, Z) represents the structured regularization term, H represents the parameters to be learned, Z represents the set of audio samples, and z i represents the hyperbolic space feature of the i-th audio sample, N represents the total number of audio samples, P c(i) represents the data prototype to which the hyperbolic spatial feature of the i-th audio sample is assigned, c(i) represents the index of the data prototype to which the hyperbolic spatial feature of the i-th audio sample is assigned, P j represents the jth data prototype, Indicates the top prototype to which the j-th data prototype is assigned, M P Indicates the number of top-level prototypes.

[0173] S112: Iteratively update the parameters of the first encoder, the second encoder, and the classifier, as well as the positions of the data prototype and the top prototype, with the goal of minimizing the structured regularization term, until a deep audio forgery detection model that meets preset conditions is obtained.

[0174] In this step, the above-mentioned structured regularization term is used as the loss function, and the model parameters are updated with the goal of minimizing the structured empirical risk. That is, the parameters of the first and second encoders, the parameters of the classifier, and the positions of the data prototypes and top prototypes are iteratively updated with the goal of minimizing the structured regularization term. During the iterative process, the structured regularization term forces the hyperbolic space feature z of each audio sample to be i Approaching the data prototype P to which it belongs j , so that each data prototype P j Approaching the top prototype T to which it belongs k .

[0175] Because the Poincare sphere model is better able to represent hierarchical structures (in hyperbolic space, greater distances indicate deeper / less correlated hierarchies), during iterative training, the hyperbolic space features of the aforementioned tree-like hierarchy, from the root node to each top-level prototype, then to each data prototype, and finally to each specific audio sample data, gradually converge. Accordingly, the parameters of the aforementioned audio forgery detection model (including encoder parameters, classifier parameters, and positional parameters of all data prototypes and top-level prototypes) are updated and adjusted according to the overall loss function until the aforementioned tree-like hierarchy converges to a hierarchy that meets the preset convergence criteria, resulting in the aforementioned audio deepfake detection model.

[0176] It should be noted that the convergence condition of the above-mentioned iterative training can be adaptively set according to needs, and this embodiment does not impose any limitation on this.

[0177] The above is a training method for an audio deep fake detection model proposed in one or more embodiments of this specification. This method learns the audio style of audio samples in the source domain through a learnable style library, and then uses the style library to align the audio samples in style, so that samples from different fields or with different styles are more consistent when sent to subsequent classifiers, thereby providing more robust and more generalizable input features.

[0178] Further, the method maps the style-aligned audio sample deep features into a Poincare ball model and classifies them through a classifier. The hierarchical structure of the classifier's classification results in the Poincare ball model is described by constructing a structured regularization term, so that the model learns the intrinsic structure and decision boundary of the data.

[0179] Through the above training method, the audio deep fake detection model can have better generalization ability and robustness, and improve the accuracy of the detection result of the audio deep fake detection model.

[0180] In general, the main goal of style alignment is to handle domain differences and style changes at the feature level. It "standardizes" the features by mapping the style of the input audio to a unified style space, so that samples from different domains or with different styles are more consistent when fed into the subsequent classifier. This is equivalent to data preprocessing and feature enhancement at the front end of the classification task, with the goal of providing more robust and more generalizable input features.

[0181] The main goal of structured empirical risk minimization is to optimize the deep embedding space learned by the audio deep fake detection model and the generalization ability of the classifier. It constructs the hierarchical structure of the data in the Poincare ball model and minimizes the structured empirical risk, so that the model focuses on how to learn the intrinsic structure and decision boundary of the data, so as to ensure that the model not only correctly classifies known data, but also better generalizes to unknown domain data.

[0182] Corresponding to the above training method, one or more embodiments of the present specification propose a training device (hereinafter referred to as training device) of an audio deep fake detection model. The audio deep fake detection model includes a first encoder, a second encoder, a style alignment module and a classifier. Please refer to Figure 4 , Figure 4 The structure of the training device of the audio deep fake detection model proposed in one or more embodiments of the present specification is schematically shown. It should be noted that the above training method can be implemented by relying on Figure 4 the training device shown, but is not limited to this device.

[0183] As Figure 4 shown, the above training device includes:

[0184] The first data acquisition module 401 is configured to acquire audio samples of a source domain.

[0185] The first feature extraction module 402 is configured to extract shallow features of the audio samples using the first encoder.

[0186] The style library construction module 403 is configured to train a style library using the shallow features of the audio samples to obtain a style library capable of covering the audio styles of the source domain.

[0187] The feature reconstruction module 404 is configured to perform style alignment on the shallow features of the audio samples by the style alignment module using the style library.

[0188] The second feature extraction module 405 is configured to perform deep feature extraction on the style-aligned audio sample features using a second encoder.

[0189] The feature mapping module 406 is configured to map the deep features extracted by the second feature extraction module into a Poincare ball model to obtain hyperbolic space features.

[0190] The training module 407 is configured to assign the hyperbolic space features of the audio samples to data prototypes, assign the data prototypes to top-level prototypes, and construct a structured regularization term to describe the hierarchical relationship of the assignment results using a classifier.

[0191] The training module 407 is further configured to iteratively update the parameters of the first encoder, the second encoder, and the classifier, and the positions of the data prototypes and the top-level prototypes, with the purpose of minimizing the structured regularization term, until an audio deep fake detection model that meets the preset conditions is obtained.

[0192] For the above-mentioned data acquisition module 401, the audio samples acquired by the module include fake audio samples and real audio samples of different styles. The style here can include but is not limited to:

[0193] Changes in recording conditions, such as different types of microphones, room acoustic characteristics, background noise, etc.

[0194] Acoustic characteristics of the speaker, such as voice line, speech rate, timbre, etc., can synthesize specific traces of fake through general acoustic patterns.

[0195] The above-mentioned fake audio samples can also be classified according to their fake methods, such as fake audio samples using WaveNet for fake, fake audio samples using Tacotron for fake, fake audio samples using VALL-E for fake, etc. Real audio samples can also be classified according to their characteristics, such as male voice audio samples, female voice audio samples, etc.

[0196] For the above-mentioned first feature extraction module 402, the module is used to perform shallow feature extraction on the audio samples using a first encoder after acquiring the audio samples.

[0197] It should be noted that the first encoder used for feature extraction of the audio samples can be adaptively selected according to requirements, and the present embodiment does not limit this. For example, a wav2vec model can be used to perform shallow feature extraction on the above-mentioned audio samples to obtain shallow features of the audio samples.

[0198] For the above style library construction module 403, the module is configured to train the style library using the shallow features of the audio samples to obtain a style library capable of covering the audio styles of the source domain. The specific training process of the style library can adopt the specific steps in step S104 of the above training method, which will not be repeated here.

[0199] For the above feature reconstruction module 404, the module is specifically configured to:

[0200] First, the style vector of the audio sample is extracted from the shallow feature of the audio sample. Then, the similarity between the style vector of the audio sample and each base style vector in the style library is determined, the similarity is normalized into a weight coefficient, and the base style vectors in the style library are weighted and summed according to the weight coefficient to obtain a reconstructed new style vector. The specific process of the above reconstructed new style vector can refer to the specific steps in steps S302 to S304, which will not be repeated here.

[0201] Finally, the feature reconstruction module 404 performs style alignment processing on the style vector of the audio sample using the new style vector.

[0202] In some embodiments, the feature reconstruction module 404 can realize the style alignment processing of the new style vector and the style vector of the audio sample by adaptive instance normalization (Adaptive Instance Normalization, AdalN).

[0203] Specifically, the feature reconstruction module 404 can decompose the new style vector into a new channel mean and a new channel standard deviation corresponding to each channel according to the channels of the shallow feature of the audio sample. The new channel mean can be represented as μ new = s new [1…C], and the new channel standard deviation can be represented as σ new = s new [1+C…2C].

[0204] Then, the feature reconstruction module 404 uses the new channel mean and the new channel standard deviation of each channel to perform style alignment on the corresponding channel feature in the shallow feature of the audio sample:

[0205]

[0206] Wherein, F aligned,c represents the channel feature of the channel c after style alignment.

[0207] For the second feature extraction module 405 described above, the module can use a second encoder to perform deep feature extraction on the style-aligned audio sample features, thereby reducing the dimensionality of the features for subsequent structural analysis in the Poincaré ball model.

[0208] It should be noted that the structure of the second encoder described above can be adaptively selected according to requirements, and the present embodiment does not limit this.

[0209] For the feature mapping module 406 described above, the module is used to map the extracted deep features in Euclidean space to the Poincaré ball model to obtain the hyperbolic space features of the audio sample.

[0210] For the training module 407 described above, the module uses a classifier to assign hyperbolic space features to data prototypes and assign data prototypes to top-level prototypes, and constructs a structured regularization term to describe the hierarchical relationship of the assignment results.

[0211] Before the training module 407 assigns hyperbolic space features to data prototypes using a classifier, the training module 407 needs to initialize the positions of data prototypes and top-level prototypes in the Poincaré ball model in advance. Data prototypes and top-level prototypes are constructed according to the labels of audio samples, where data prototypes are used to represent fine-grained categories, and top-level prototypes refer to higher-level categories than data prototypes. There can be multiple layers of top-level prototypes, and the highest layer is the global top-level prototype. In the spatial structure, data prototypes can be understood as child nodes, while top-level prototypes are parent nodes at higher levels in the hierarchical structure. Top-level prototypes organize multiple related data prototypes (or lower-level top-level prototypes) together to form a tree structure, and the root node of this tree structure is the global top-level prototype.

[0212] For example, if data prototypes P1, P2, and P3 represent different deep fake technologies (such as WaveNet, Tacotron, and VALL-E), a top-level prototype Tfake can represent the macro category of "deep fake audio", which is the parent node of P1, P2, and P3. Similarly, if P4 and P5 represent different types of real audio data (such as male voice and female voice), another top-level prototype Treal can represent all "real audio data".

[0213] Data prototypes and top-level prototypes can be randomly initialized in the space of the Poincaré ball model, or the hyperbolic space features of the audio samples can be preliminarily clustered (such as K-means clustering), and the positions of the data prototypes and top-level prototypes are initialized according to the results of the preliminary clustering.

[0214] The aforementioned use of a classifier to assign the hyperbolic space features of an audio sample to a data prototype refers to using a classifier to determine to which data prototype the hyperbolic space features of an audio sample belong.

[0215] In some implementations, the training module 407 may implement allocation of the hyperbolic spatial features of the audio samples to the data prototypes using a hard allocation or a soft allocation approach.

[0216] Hard allocation means that in each iteration, the training module 407 assigns the hyperbolic spatial features z of the audio samples to i Assigned to the data prototype P closest to it in the Poincare sphere model j , which can be achieved by calculating the hyperbolic distance d between the hyperbolic space features of the audio sample and the data prototype H (z i ,P j ) and select d H to achieve the minimum value.

[0217] Soft assignment means that the training module 407 uses similarity or attention mechanism to assign hyperbolic space features z to the audio samples. i Calculate which data prototype P it belongs to j For example, the training module 407 can use a softmax layer to generate z i The distribution probability softmax(-similarity(z i ,P j )), where similarity(z i ,P j ) represents the hyperbolic space feature z of the audio sample i With data prototype P j The similarity between i and P j The distance between them is negatively correlated.

[0218] The above-mentioned use of a classifier to assign a data prototype to a top-level prototype refers to using a classifier to determine to which top-level prototype the data prototype belongs. This process is a process of establishing a parent-child relationship between the data prototype and the top-level prototype.

[0219] In some implementations, the training module 407 may implement the allocation of data prototypes to top-level prototypes using hard allocation or soft allocation.

[0220] Hard allocation means that in each iteration, the training module 407 assigns the data prototype P j Assigned to the top prototype T closest to it in the Poincare sphere model k , which can be achieved by calculating the data prototype P j With the top prototype T kThe hyperbolic distance d between H (P j ,T k ) and select d H to achieve the minimum value.

[0221] Soft assignment means that the training module 407 uses similarity or attention mechanism to assign the data prototype P j Count the top-level prototypes it belongs to k For example, the training module 407 can use a softmax layer to generate the data prototype P j The probability of assignment softmax(-similarity(P j ,T k )), where similarity(P j ,T k ) represents the data prototype P j With the top prototype T k The similarity between j and T k The distance between them is negatively correlated.

[0222] In some more specific embodiments, the training module 407 can also pre-define some high-level relationships. For example, if it is known that certain data prototypes (such as various counterfeit technology prototypes) belong to the general category of "counterfeiting", then these data prototypes can be assigned to the top-level prototype representing "counterfeiting" at the beginning.

[0223] After the above allocation, the hyperbolic space features of the audio samples are allocated to the data prototypes, the data prototypes are allocated to the top prototypes, and the top prototypes are allocated to the top prototypes at a higher level than them. Finally, a tree-like hierarchical allocation result can be obtained.

[0224] It should be noted that the model structure of the above-mentioned classifier can be adaptively selected according to needs, and this embodiment does not impose any limitation on this.

[0225] After obtaining the classification result of the classifier, the training module 407 can construct a structured regularization term to describe the hierarchical relationship of the allocation result, that is, to describe the hierarchical relationship between the nodes in the allocation result of the tree-like hierarchical structure.

[0226] Specifically, the above structured regularization term can be described by the following expression:

[0227]

[0228] Among them, R structure (H, Z) represents the structured regularization term, H represents the parameters to be learned, Z represents the set of audio samples, and z ihyperbolic space feature of the i-th audio sample, N represents the total number of audio samples, P c(i) data prototype to which the hyperbolic space feature of the i-th audio sample is assigned, c(i) represents the index of the data prototype to which the hyperbolic space feature of the i-th audio sample is assigned, P j represents the j-th data prototype, represents the top prototype to which the j-th data prototype is assigned, M P represents the number of top prototypes.

[0229] For the training module 407 described above, the module is further configured to iteratively update the parameters of the first encoder, the second encoder and the classifier, and the positions of the data prototypes and the top prototypes, with the purpose of minimizing the structured regularization term, until a deep audio forgery detection model satisfying a preset condition is obtained.

[0230] Specifically, the training module 407 uses the structured regularization term described above as a loss function, and updates the model parameters with the purpose of structured empirical risk minimization. That is, the training module 407 iteratively updates the parameters of the first and second encoders, the parameters of the classifier, and the positions of the data prototypes and the top prototypes, with the purpose of minimizing the structured regularization term. During the iteration process, the structured regularization term forces the hyperbolic space feature z i of each audio sample to approach the data prototype P j to which it belongs, so that each data prototype P j approaches the top prototype T k to which it belongs.

[0231] Since the Poincare ball model can better represent the hierarchical structure (in hyperbolic space, the farther the distance, the deeper / less relevant the hierarchy), in the iterative training process, the tree-like hierarchical structure described above gradually converges from the root node to each top prototype, to each data prototype, and finally to the hyperbolic space feature of each specific audio sample data. Correspondingly, the parameters of the audio forgery detection model described above (including the encoder parameters, the classifier parameters, and the position parameters of all data prototypes and top prototypes) will be updated and adjusted according to the overall loss function, until the tree-like hierarchical structure converges to a hierarchical structure that meets the preset convergence condition. At this point, the audio deep forgery detection model described above is obtained.

[0232] It should be noted that the convergence condition of the above iterative training can be adaptively set according to requirements, and the present embodiment does not limit this.

[0233] For the training device of the above-mentioned audio deepfake detection model, taking a module as an example of a software functional unit, the first data acquisition module 401 can include code running on a computing instance. The computing instance can include at least one of a physical host (computing device), a virtual machine, and a container. Further, the above-mentioned computing instance can be one or more. For example, the first data acquisition module 401 can include code running on multiple hosts / virtual machines / containers. The multiple hosts / virtual machines / containers for running the code can be distributed in the same region, or distributed in different regions. Further, the multiple hosts / virtual machines / containers for running the code can be distributed in the same availability zone (AZ), or distributed in different AZs, each AZ including a data center or multiple data centers with similar geographical locations. Generally, one region can include multiple AZs.

[0234] Similarly, the multiple hosts / virtual machines / containers for running the code can be distributed in the same virtual private cloud (VPC), or distributed in multiple VPCs. Generally, one VPC is set in one region, and communication between two VPCs in the same region and between VPCs in different regions needs to be set in each VPC to set a communication gateway, and the interconnection between VPCs is realized through the communication gateway.

[0235] Taking a module as an example of a hardware functional unit, the first data acquisition module 401 can include at least one computing device, such as a server, etc. Alternatively, the first data acquisition module 401 can also be a device implemented by an application-specific integrated circuit (ASIC) or a programmable logic device (PLD), etc. The above-mentioned PLD can be implemented by a complex programmable logical device (CPLD), a field-programmable gate array (FPGA), a generic array logic (GAL), or any combination thereof.

[0236] The multiple computing devices included in the first data acquisition module 401 can be distributed in the same region or in different regions. The multiple computing devices included in the first data acquisition module 401 can be distributed in the same AZ or in different AZs. Similarly, the multiple computing devices included in the first data acquisition module 401 can be distributed in the same VPC or in multiple VPCs. The multiple computing devices can be any combination of servers, ASICs, PLDs, CPLDs, FPGAs, and GALs.

[0237] In other embodiments, the first data acquisition module 401 can be used to perform any step of the training method of the audio deepfake detection model described above, the first feature extraction module 402 can be used to perform any step of the training method of the audio deepfake detection model described above, the style library construction module 403 can be used to perform any step of the training method of the audio deepfake detection model described above, the feature reconstruction module 404 can be used to perform any step of the training method of the audio deepfake detection model described above, the second feature extraction module 405 can be used to perform any step of the training method of the audio deepfake detection model described above, the feature mapping module 406 can be used to perform any step of the training method of the audio deepfake detection model described above, and the training module 407 can be used to perform any step of the training method of the audio deepfake detection model described above. The steps responsible for the implementation of the first data acquisition module 401, the first feature extraction module 402, the style library construction module 403, the feature reconstruction module 404, the second feature extraction module 405, the feature mapping module 406, and the training module 407 can be specified as needed, and the functions of the audio deepfake detection model training device described above are implemented by the first data acquisition module 401, the first feature extraction module 402, the style library construction module 403, the feature reconstruction module 404, the second feature extraction module 405, the feature mapping module 406, and the training module 407 respectively implementing different steps in the training method of the audio deepfake detection model described above.

[0238] In this implementation, the training device of the audio deepfake detection model can also be applied to computing devices such as computers and servers, or to computing device clusters including at least one computing device, to realize the specific functions of the audio deepfake detection model training device.

[0239] Corresponding to the training method described above, one or more embodiments of the present specification also propose an audio deepfake detection method. Please refer to Figure 5 , Figure 5A flowchart of an audio deepfake detection method (hereinafter referred to as the detection method) proposed in one or more embodiments of the present specification, the detection method comprising steps S500-S510.

[0240] S500: Obtain to-be-recognized audio data in a target domain.

[0241] S502: Extract shallow features of the to-be-recognized audio data using a first encoder of a deep audio fake detection model.

[0242] S504: Perform style alignment on the shallow features of the to-be-recognized audio data using a style alignment module of the deep audio fake detection model using a style library.

[0243] S506: Perform deep feature extraction on the style-aligned audio sample features using a second encoder of the deep audio fake detection model, and map the extracted deep features to a Poincare ball model to obtain hyperbolic space features.

[0244] S508: Assign the hyperbolic space features to data prototypes using a classifier of the deep audio fake detection model, assign the data prototypes to top-level prototypes, and finally obtain a recognition result of the to-be-recognized audio data.

[0245] S510: Determine whether the to-be-recognized audio data is fake audio data according to the recognition result.

[0246] In the above detection method, the audio deepfake detection model is obtained by pre-training using the training method of the audio deepfake detection model described above. When the audio fake detection model is applied to the target domain, the audio data of the target domain is first aligned in style using the pre-trained style library, and then sent to the classifier for classification, so that the parameters of the audio fake detection model do not need to be fine-tuned. Then, the classifier in the audio deepfake detection model assigns the hyperbolic space features of the to-be-recognized audio data to the pre-set categories (such as real category and fake category), thereby obtaining the recognition result of the to-be-recognized audio data, and further determining whether the to-be-recognized audio data is fake audio data according to the recognition result.

[0247] Corresponding to the above detection method, an audio deepfake detection device is also proposed in one or more embodiments of the present specification. Please refer to Figure 6 , Figure 6 A structural diagram of an audio deepfake detection device proposed in one or more embodiments of the present specification. It should be noted that the above detection method can be implemented by relying on Figure 6 the detection device, but is not limited to this device.

[0248] As Figure 6 shown, the detection device comprises:

[0249] The second data acquisition module 601 is configured to acquire to-be-recognized audio data in a target domain.

[0250] The detection module 602 is configured to extract shallow features of the to-be-recognized audio data by using a first encoder of a deep audio forgery detection model; perform style alignment on the shallow features of the to-be-recognized audio data by using a style library through a style alignment module of the deep audio forgery detection model; perform deep feature extraction on the audio sample features after the style alignment by using a second encoder of the deep audio forgery detection model, and map the extracted deep features to a Poincare ball model to obtain hyperbolic space features; and assign the hyperbolic space features to data prototypes by using a classifier of the deep audio forgery detection model, assign the data prototypes to top-level prototypes, and finally obtain a recognition result of the to-be-recognized audio data.

[0251] The determination module 603 is configured to determine whether the to-be-recognized audio data is forged audio data according to the recognition result.

[0252] In the detection device, the audio deep forgery detection model is obtained by using the training method of the audio deep forgery detection model. When the audio deep forgery detection model is applied to the target domain, the audio data in the target domain is first aligned in style by using the pre-trained style library, and then is sent to the classifier for classification, so that the parameters of the audio deep forgery detection model do not need to be fine-tuned. Then, the classifier in the audio deep forgery detection model is used to assign the hyperbolic space features of the to-be-recognized audio data to a preset category (for example, a real category and a forged category), so as to obtain a recognition result of the to-be-recognized audio data, and then determine whether the to-be-recognized audio data is forged audio data according to the recognition result.

[0253] For the audio deep forgery detection device, as an example of a module as a software functional unit, the second data acquisition module 601 can include code running on a computing instance. The computing instance can include at least one of a physical host (computing device), a virtual machine, and a container. Further, the computing instance can be one or more. For example, the second data acquisition module 601 can include code running on multiple hosts / virtual machines / containers. The multiple hosts / virtual machines / containers for running the code can be distributed in the same region (region), or can be distributed in different regions. Further, the multiple hosts / virtual machines / containers for running the code can be distributed in the same availability zone (availability zone, AZ), or can be distributed in different AZs, each of which includes a data center or multiple data centers with similar geographical locations. Generally, one region can include multiple AZs.

[0254] Likewise, the multiple hosts / virtual machines / containers for running the code can be distributed in the same virtual private cloud (VPC) or in multiple VPCs. Among them, usually one VPC is set in one region, and a communication gateway needs to be set in each VPC for cross-region communication between two VPCs in the same region and between VPCs in different regions, and the interconnection between VPCs is realized through the communication gateway.

[0255] As an example of a hardware functional unit, the second data acquisition module 601 can include at least one computing device, such as a server, etc. Alternatively, the second data acquisition module 601 can also be a device implemented by an application-specific integrated circuit (ASIC) or a programmable logic device (PLD), etc. Among them, the above-mentioned PLD can be implemented by a complex programmable logic device (CPLD), a field-programmable gate array (FPGA), a generic array logic (GAL), or any combination thereof.

[0256] The multiple computing devices included in the second data acquisition module 601 can be distributed in the same region or in different regions. The multiple computing devices included in the second data acquisition module 601 can be distributed in the same AZ or in different AZs. Likewise, the multiple computing devices included in the fine-tuning module 501 can be distributed in the same VPC or in multiple VPCs. Among them, the multiple computing devices can be any combination of servers, ASICs, PLDs, CPLDs, FPGAs, and GALs, etc.

[0257] In other embodiments, the second data acquisition module 601 can be used to perform any step of the above-mentioned audio deepfake detection method, the detection module 602 can be used to perform any step of the above-mentioned audio deepfake detection method, and the determination module 603 can be used to perform any step of the above-mentioned audio deepfake detection method. The steps responsible for the second data acquisition module 601, the detection module 602, and the determination module 603 can be specified as needed, and the above-mentioned audio deepfake detection device realizes the entire function of the above-mentioned audio deepfake detection method by realizing different steps in the above-mentioned audio deepfake detection method through the second data acquisition module 601, the detection module 602, and the determination module 603 respectively.

[0258] In the present implementation, the audio deepfake detection apparatus can also be applied to a computing device such as a computer or a server, or a computing device cluster comprising at least one computing device, to implement an audio deepfake detection function.

[0259] In some embodiments, an electronic device is also provided. Please refer to Figure 7 The electronic device comprises a bus 701, a processor 702, a memory 703 and a communication interface 704. The processor 702, the memory 703 and the communication interface 704 communicate with each other through the bus 701. The electronic device can be a server or a terminal device. It should be understood that the number of processors and memories in the electronic device is not limited in the present application.

[0260] The bus 701 can be a peripheral component interconnect (PCI) bus, an extended industry standard architecture (EISA) bus or the like. The bus can be divided into an address bus, a data bus, a control bus and the like. For ease of representation, Figure 7 only one line is used in the figure, but it does not mean that there is only one bus or only one type of bus. The bus 701 can include a path for transmitting information between various components (for example, the processor 702, the memory 703 and the communication interface 704) of the electronic device.

[0261] The processor 702 can include any one or more of a processor CPU, a graphics processing unit (GPU), a microprocessor (MP) or a digital signal processor (DSP).

[0262] The memory 703 can include a volatile memory such as a random access memory (RAM). The memory 703 can also include a non-volatile memory such as a read-only memory (ROM), a flash memory, a hard disk drive (HDD) or a solid state drive (SSD).

[0263] The memory 703 stores executable program codes, and the processor 702 executes the executable program codes to implement the training method of the audio deepfake detection model as described above, or to implement the audio deepfake detection method as described above.

[0264] The communication interface 704 uses a transceiving module such as, but not limited to, a network interface card, a transceiver, and the like to enable communication between the electronic device and other devices or communication networks.

[0265] One or more embodiments in the specification provide a computer-readable storage medium storing a computer program, when the computer program is executed on an electronic device, causes the electronic device to perform the training method of the audio deepfake detection model described above, or perform the audio deepfake detection method described above.

[0266] The computer-readable storage medium can be any available medium or data center and the like data storage device containing one or more available media that the electronic device can store. The available media can be a magnetic medium, (for example, a floppy disk, a hard disk, a magnetic tape), an optical medium (for example, a DVD), or a semiconductor medium (for example, a solid state disk) and the like. The computer-readable storage medium includes instructions instructing the electronic device to perform the training method of the audio deepfake detection model described above, or perform the audio deepfake detection method described above.

[0267] It can be understood that the structure illustrated by the embodiments in the specification does not constitute a specific limitation on the system of the embodiments in the specification. In other embodiments of the specification, the above system can include more or fewer components than the illustration, or combine certain components, or split certain components, or different component arrangements. The illustrated components can be implemented in hardware, software, or a combination of software and hardware.

[0268] Each of the embodiments in the specification is described in a progressive manner, and the same or similar parts between each embodiment can be referred to each other, and each embodiment focuses on the difference from other embodiments. Especially, for the device embodiments, since they are basically similar to the method embodiments, the description is relatively simple, and the relevant parts can be referred to the part of the method embodiment.

[0269] The above describes specific embodiments of the specification. Other embodiments are within the scope of the appended claims. In some cases, the actions or steps recited in the claims can be performed in an order different than the order in the embodiments and still achieve the desired result. In addition, the processes depicted in the figures do not necessarily require the particular order shown or sequential order to achieve the desired results. In certain implementations, multitasking and parallel processing are also possible or advantageous.

[0270] It should be noted that the above-mentioned are only specific embodiments of the present application, and obviously the present application is not limited to the above-mentioned embodiments, and there are many similar changes. All the changes directly derived or thought from the disclosure of the present application by those skilled in the art should belong to the protection scope of the present application.

Claims

1. A method of training a deep audio forgery detection model, the deep audio forgery detection model comprising: a first encoder, a second encoder, a style alignment module and a classifier; the method comprises: obtaining an audio sample of a source domain; extracting shallow features of the audio sample by using the first encoder; training a style library by using the shallow features of the audio sample, to obtain a style library capable of covering audio styles of the source domain; aligning the shallow features of the audio sample with the style library by using the style alignment module; extracting deep features of the audio sample after style alignment by using the second encoder, and mapping the extracted deep features into a Poincare ball model to obtain hyperbolic space features; assigning the hyperbolic space features to data prototypes by using the classifier, assigning the data prototypes to top-level prototypes, and constructing a structured regularization term to describe the hierarchical relationship of the assignment results; the data prototypes are used to represent the categories of the audio sample, and the top-level prototypes are higher-level categories of the data prototypes; iteratively updating parameters of the first encoder, the second encoder and the classifier, and positions of the data prototypes and the top-level prototypes, with the purpose of minimizing the structured regularization term, until the deep audio forgery detection model satisfying the preset condition is obtained.

2. The method of claim 1, wherein training the style library by using the audio sample comprises: initializing a set of base style vectors to constitute the style library; extracting a style vector of the audio sample from the shallow features of the audio sample; determining a similarity between the style vector of the audio sample and each base style vector in the style library; normalizing the similarity into a weight coefficient, and performing weighted summation on the base style vectors in the style library according to the weight coefficient to obtain a new style vector; constructing an orthogonal style loss function based on the base style vectors; constructing a style reconstruction loss function based on the new style vector and the style vector; updating the base style vectors by using the orthogonal style loss function and the style reconstruction loss function, until the style library capable of covering the audio styles of the source domain is obtained.

3. The method of claim 2, wherein extracting the style vector of the audio sample from the shallow features of the audio sample comprises: extracting a channel mean and a channel standard deviation of each channel from the shallow features of the audio sample as the style vector of the audio sample.

4. The method of claim 1, wherein aligning the shallow features of the audio sample with the style library by using the style alignment module comprises: extracting a style vector of the audio sample from the shallow features of the audio sample; determining a similarity between the style vector of the audio sample and each base style vector in the style library; normalizing the similarity into a weight coefficient, and performing weighted summation on the base style vectors in the style library according to the weight coefficient to obtain a reconstructed new style vector; aligning the style vector of the audio sample with the new style vector by using the new style vector.

5. The method of claim 4, wherein: extracting a style vector of the audio sample from the shallow features of the audio sample, specifically comprising: extracting a channel mean and a channel standard deviation of each channel from the shallow features of the audio sample as the style vector of the audio sample; performing style alignment on the style vector of the audio sample by using the new style vector, specifically comprising: decomposing the new style vector into a new channel mean and a new channel standard deviation of the each channel; performing style alignment on the corresponding channel feature in the shallow features of the audio sample by using the new channel mean and the new channel standard deviation of the each channel.

6. The method of claim 1, further comprising: initializing positions of the data prototypes and the top prototypes in the Poincare ball model before assigning the hyperbolic space features to the data prototypes by using the classifier.

7. The method of claim 1, wherein the structured regularizer is expressed as: wherein H represents a parameter to be learned, Z represents a set of the audio samples, z i represents a hyperbolic space feature of an i-th audio sample, N represents a total number of the audio samples, P c(i) represents a data prototype to which the hyperbolic space feature of the i-th audio sample is assigned, c(i) represents an index of the data prototype to which the hyperbolic space feature of the i-th audio sample is assigned, P j represents a j-th data prototype, represents a top-level prototype to which the j-th data prototype is assigned, M P represents a number of the top-level prototypes.

8. The method of claim 1, wherein, in each iteration, assigning the hyperbolic space features of the audio sample to the data prototype closest to the audio sample data by using the classifier, and assigning the data prototype to the top prototype closest to the data prototype by using the classifier.

9. An apparatus for training a deep audio forgery detection model, the deep audio forgery detection model comprising: a first encoder, a second encoder, a style alignment module, and a classifier; the apparatus comprises: a first data acquisition module configured to acquire audio samples of a source domain; a first feature extraction module configured to extract shallow features of the audio samples by using the first encoder; a style library construction module configured to train a style library by using the shallow features of the audio samples, to obtain a style library capable of covering audio styles of the source domain; a feature reconstruction module configured to perform style alignment on the shallow features of the audio samples by using the style library through the style alignment module; a second feature extraction module configured to perform deep feature extraction on the audio sample features after style alignment by using the second encoder; a feature mapping module configured to map the deep features extracted by the second feature extraction module into a Poincare ball model to obtain hyperbolic space features; a training module configured to assign the hyperbolic space features of the audio samples to data prototypes and assign the data prototypes to top prototypes by using the classifier, and to construct a structured regularizer to describe a hierarchical relationship of the assignment results; the data prototypes are used to represent categories of the audio samples, and the top prototypes are higher-level categories of the data prototypes; the training module is further configured to iteratively update parameters of the first encoder, the second encoder, and the classifier, and positions of the data prototypes and the top prototypes, with a purpose of minimizing the structured regularizer, until the deep audio forgery detection model satisfying a preset condition is obtained.

10. The apparatus of claim 9, wherein the style library construction module is specifically configured to: initialize a set of base style vectors to constitute the style library; extract a style vector of the audio sample from the shallow features of the audio sample; determining similarities between the style vector of the audio sample and each base style vector in the style library; normalizing the similarities into weight coefficients, and performing a weighted summation of the base style vectors in the style library according to the weight coefficients to obtain a new style vector; constructing an orthogonal style loss function based on the base style vectors; constructing a style reconstruction loss function based on the new style vector and the style vector; updating the base style vectors using the orthogonal style loss function and the style reconstruction loss function until a style library capable of covering the audio styles of the source domain is obtained.

11. The apparatus of claim 10, wherein the style library construction module is specifically configured to: extract a channel mean and a channel standard deviation of each channel from the shallow features of the audio sample as the style vector of the audio sample.

12. The apparatus of claim 9, wherein the feature reconstruction module is specifically configured to: extract the style vector of the audio sample from the shallow features of the audio sample; determine similarities between the style vector of the audio sample and each base style vector in the style library; normalize the similarities into weight coefficients, and perform a weighted summation of the base style vectors in the style library according to the weight coefficients to obtain a reconstructed new style vector; perform style alignment processing on the style vector of the audio sample using the new style vector.

13. The apparatus of claim 12, wherein: the style library construction module is specifically configured to extract a channel mean and a channel standard deviation of each channel from the shallow features of the audio sample as the style vector of the audio sample; the feature reconstruction module is specifically configured to decompose the new style vector into a new channel mean and a new channel standard deviation of each channel; and perform style alignment on the corresponding channel features in the shallow features of the audio sample using the new channel mean and the new channel standard deviation of each channel.

14. The apparatus of claim 9, wherein the training module is further configured to initialize positions of the data prototypes and the top-level prototypes in the Poincare ball model before assigning the hyperbolic space features of the audio samples to the data prototypes using the classifier.

15. The apparatus of claim 9, wherein an expression of the structured regularization term is: wherein, H represents a parameter to be learned, Z represents a set of the audio samples, z i represents a hyperbolic space feature of an i-th audio sample, N represents a total number of the audio samples, P c(i) represents a data prototype to which the hyperbolic space feature of the i-th audio sample is assigned, c(i) represents an index of the data prototype to which the hyperbolic space feature of the i-th audio sample is assigned, P j represents a j-th data prototype, represents a top-level prototype to which the j-th data prototype is assigned, M P represents a number of the top-level prototypes.

16. The apparatus of claim 9, wherein the training module is specifically configured to: in each iteration, assign the hyperbolic space features of the audio samples to the data prototype closest to the audio sample data using the classifier, and assign the data prototype to the top-level prototype closest to the data prototype using the classifier.

17. An audio deep forgery detection method, comprising: obtaining to-be-recognized audio data in a target domain; extracting shallow features of the to-be-recognized audio data using a first encoder of a deep audio forgery detection model; the deep audio forgery detection model is pre-trained using the method of any one of claims 1 to 8; performing style alignment on the shallow features of the to-be-recognized audio data using the style library by a style alignment module of the deep audio forgery detection model. The second encoder of the deep audio forgery detection model performs deep feature extraction on the style-aligned audio sample features, and maps the extracted deep features to a Poincare ball model to obtain hyperbolic space features; The classifier of the deep audio forgery detection model assigns the hyperbolic space features to data prototypes, and assigns the data prototypes to top-level prototypes, and finally obtains the recognition result of the to-be-recognized audio data; According to the recognition result, it is determined whether the to-be-recognized audio data is a fake audio data.

18. An audio deep forgery detection device, comprising: a second data acquisition module configured to acquire to-be-recognized audio data in a target domain; a detection module configured to extract shallow features of the to-be-recognized audio data using a first encoder of a deep audio forgery detection model; perform style alignment on the shallow features of the to-be-recognized audio data using a style alignment module of the deep audio forgery detection model and a style library; perform deep feature extraction on the style-aligned audio sample features using a second encoder of the deep audio forgery detection model, and map the extracted deep features to a Poincare ball model to obtain hyperbolic space features; assign the hyperbolic space features to data prototypes using a classifier of the deep audio forgery detection model, assign the data prototypes to top-level prototypes, and finally obtain the recognition result of the to-be-recognized audio data; the deep audio forgery detection model is pre-trained using the method of any one of claims 1-8; a determination module configured to determine whether the to-be-recognized audio data is a fake audio data according to the recognition result.

19. A computer readable storage medium, the computer readable storage medium storing a computer program, when the computer program is executed on an electronic device, the electronic device is caused to execute the method of any one of claims 1-8, or execute the method of claim 17.

20. An electronic device, comprising: at least one memory for storing programs; at least one processor for executing the programs stored in the memory, when the programs stored in the memory are executed, the processor is used to execute the method of any one of claims 1-8, or execute the method of claim 17.