Unsupervised cross-domain forged speech detection method based on fusion loss constraint
By adopting an unsupervised cross-domain detection method based on fusion loss constraints in the speech depth forgery detection technology, combining the loss calculation of first-order and second-order statistical information, the problem of poor detection performance of the existing technology under different data distributions is solved, and the high accuracy and low error rates of cross-domain speech detection are achieved.
Patent Information
- Application Number
- CN202510027827.X
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-01-08
- Publication Date
- 2025-05-13
- Estimated Expiration
- 2045-01-08
AI Technical Summary
The existing speech depth forgery detection technology faces speech with different data distributions, poor detection performance and poor generalization capabilities, especially when the speech of the source and target domains is not independently distributed, the model performance will drop sharply.
Unsupervised cross-domain forged speech detection method based on fusion loss constraints is adopted. By combining first-order statistical information and second-order statistical information at the feature level, the losses of shallow detail characteristics and deep global semantic features between the source domain and the target domain are calculated, network parameters are updated, and the model parameters of the deep forged detection model are optimized to achieve domain adaptation effect.
In the unsupervised situation, the generalization of unseen fake speech is improved, the generalization ability of the detection network in cross-domain situations is improved, and the high accuracy and low error rates of the forged speech detection method in the face of domain drift caused by different languages are ensured.
Smart Images

Figure CN119559952B_ABST
Abstract
Description
Technical Field
[0001] The invention relates to the field of multimedia security technology, and in particular to an unsupervised cross-domain forged speech detection method based on fusion loss constraints. Background Art
[0002] Traditional methods of generating fake voices are mainly through text-to-speech and speech conversion. Among them, text-to-speech refers to converting a given text into natural speech; speech conversion refers to changing only the identity of the speaker in the voice. In recent years, with the continuous development of voice deep fake technology, the content of fake voices has become richer and more diversified, which has led to the gradual increase in the potential security risks brought by fake voices. For example: Meta's universal large-scale voice text-guided generation model Voicebox (Le M, Vyas A, Shi B, et al. Voicebox: Text-guided multilingual universal speech generation at scale [J]. Advances in neural information processing systems, 2023: 14005-14034.) can efficiently complete multilingual speech synthesis, content editing, style conversion and other tasks. This kind of fake voice with different data distribution increases the chances of deceiving victims and automatic speaker verification systems.
[0003] However, the training data types of existing voice deep fake detection technologies are relatively simple and have the same distribution. This leads to the problems of poor detection performance and poor generalization ability when facing data with different distributions, such as different voice types, different forgery methods, etc.
[0004] At present, there are some generalized voice deep fake detection methods that target different forgery methods (Ren Y, Peng H, Li L, et al. Generalized voice spoofing detection via integral knowledge amalgamation [J]. IEEE / ACM Transactions on Audio, Speech, and Language Processing, 2023, 31: 2461-2475.) and different data sets (Xie Y, Cheng H, Wang Y, et al. Domain generalization via aggregation and separation for audio deepfake detection [J]. IEEE Transactions on Information Forensics and Security, 2023. 344-358). However, different languages are also one of the problems that cause the failure of deep fake detection models. That is, when the forged voice detection model trained on the source domain voice is directly used for the test of the target domain forged voice, the model performance will drop sharply. Because the success of the usual voice deep fake detection model depends on the fact that the training set and the test set are independent and identically distributed, but in real application scenarios, there are many types of voice languages to be detected, which will lead to domain mismatch problems. At the same time, due to the current situation of insufficient and unbalanced voice deep fake detection data, the target domain data to be detected may not always have sufficient labels. To solve this problem, this method proposes to use fusion loss to constrain the training of the deep fake detection network. Specifically, it combines the first-order statistical information with the second-order statistical information at the feature level to avoid the problem of poor effect caused by the overly single and insufficient feature distance calculation. Based on this, the distance between the source domain and the target domain is minimized as much as possible, and then the mixed regularization loss is used to calculate the output result and label value of the source domain classification layer to obtain the loss value, and then optimize the model parameters of the deep fake detection model to achieve domain adaptation effect. During the training process, it is ensured that the target domain voice labels are inaccessible, and only the source domain voice labels are accessible. In this case, the high accuracy and low error rate of the cross-domain deep fake detection model are maintained. Summary of the invention
[0005] The purpose of the present invention is to provide an unsupervised cross-domain forged speech detection method based on fusion loss constraints, when the source domain and target domain speech are not independent and identically distributed, the detection network trained with source domain data can also be used to detect the target domain speech.
[0006] The present invention adopts the following technical solution:
[0007] An unsupervised cross-domain forged speech detection method based on fusion loss constraints, comprising:
[0008] The detection network construction step includes a shallow feature pre-extraction module, a composite residual backbone network based on a compression-excitation module, and a classification module; the acoustic features obtained after speech preprocessing are input into the shallow feature pre-extraction module, the shallow feature pre-extraction module extracts shallow artifact features and outputs them to the composite residual backbone network, the composite residual backbone network extracts deep artifact features and outputs them to the classification module; the classification module outputs a binary classification result for distinguishing the authenticity of speech;
[0009] The detection network training step includes inputting the pre-processed labeled source domain training set speech samples and the unlabeled target domain training set speech samples in batches into the detection network for training; during the training process, the global loss constraint is used to constrain the detection network and update the network parameters; the global loss constraint is obtained by weighted summation of the hybrid regularization loss and the fusion loss; wherein the hybrid regularization loss is used as a generalization booster; the fusion loss is the loss of the shallow detail features and deep global semantic features between the fusion source domain speech and the target domain speech;
[0010] In the speech detection step, a target domain speech sample to be detected is obtained, and after preprocessing, the acoustic features of the target domain speech sample are input into the trained detection network for detection to generate a discrimination result.
[0011] Preferably, the preprocessing includes: performing length unification, CQT pre-transformation and pre-emphasis processing on each speech sample to convert it into acoustic features with the same dimension.
[0012] Preferably, the shallow feature pre-extraction module includes a two-dimensional convolution, batch normalization, ReLU activation function and maximum pooling layer connected in sequence, takes the acoustic features of the speech sample as input, and extracts shallow artifact features.
[0013] Preferably, the composite residual backbone network based on the compression-excitation module includes four residual blocks connected in sequence, and after the shallow artifact features pass through the four residual blocks, deep artifact features, namely residual features, are obtained.
[0014] Preferably, the classification module includes a global average pooling layer, a dimension compression function, multiple fully connected layers and multiple ReLU activation functions connected in sequence; the feature vector obtained by the deep artifact feature through a global average pooling layer is then reduced in dimension through a dimension compression function, and then connected to multiple fully connected layers and ReLU activation functions to output a binary classification result to determine the authenticity of the speech.
[0015] Preferably, the hybrid regularization randomly mixes the source domain speech samples and their labels, and weights the cross entropy loss of the two sets of labels to obtain the hybrid regularization loss, thereby learning the continuous relationship of multiple speech samples. The process is expressed as follows:
[0016]
[0017] in, represents two sets of randomly selected mixed training samples, represents the labels of two randomly selected mixed training samples, x a is a training sample in the source domain, x b is another training sample in the source domain; a For x a The label, y b For x b The label of , λ is the mixing coefficient randomly sampled from the Beta distribution, λ~Beta(α,α), α∈(0,∞) is a hyperparameter;
[0018] The implementation of hybrid regularization is through the following equivalent loss function The implementation process is formalized as:
[0019]
[0020] in, The softmax probability of including mixed samples is determined by the detection network Output: CE(·,·) is the standard cross entropy loss.
[0021] Preferably, the fusion loss is a loss function based on first-order statistical information and a loss function based on second-order statistical information, which are expressed as follows:
[0022]
[0023] in, represents the fusion loss; represents the loss function based on first-order statistical information; represents the loss function based on second-order statistical information;
[0024] Loss function based on first-order statistics It is expressed as follows:
[0025]
[0026] Among them, δ∈(0,∞) is a hyperparameter that determines the degree of shortening the distance between the two domains; Represents the MMD distance between the source domain and target domain speech samples;
[0027]
[0028] Among them, φ(·) is a specific feature representation that operates on the source domain and target domain speech samples; represents the source domain speech set; represents the target domain speech set, Indicates a specific piece of voice data belonging to the source domain; Indicates a specific piece of voice data belonging to the target domain;
[0029] Loss function based on second-order statistics It is expressed as follows:
[0030]
[0031] in, The label representing the source domain speech; and Represent the covariance of the source domain and the target domain respectively; represents the cross entropy loss function; represents the squared log Euclidean distance; β is a hyperparameter;
[0032] and It is calculated by the centralization matrix J and is expressed as follows:
[0033]
[0034] in, and represent the activation values of the source domain and the target domain respectively.
[0035] Preferably, the global loss constraint It is expressed as follows:
[0036]
[0037] Among them, σ and ξ are empirical parameters; represents fusion loss; CE mixup represents the hybrid regularization loss.
[0038] Compared with the prior art, the present invention has the following beneficial effects:
[0039] The present invention designs an unsupervised cross-domain forged speech detection method based on fusion loss in an unsupervised situation, uses a hybrid regularization loss as a generalization booster to improve the generalization of the detection network for unseen forged speech, and to a certain extent improves the generalization ability of the detection network itself in a cross-domain situation, and then jointly constrains the detection network with the fusion loss that fuses the deep detail features and global semantic features between the source domain and the target domain to construct a cross-domain forged speech detection network; while learning to extract forged information, the detection network eliminates the model's dependence on the language type, ensuring the high accuracy and low error rate of the forged speech detection method in dealing with domain drift caused by language differences. BRIEF DESCRIPTION OF THE DRAWINGS
[0040] Figure 1 A schematic diagram of a network model of a method according to an embodiment of the present invention;
[0041] Figure 2 It is a flowchart of an unsupervised cross-domain forged speech detection method based on fusion loss constraint according to an embodiment of the present invention;
[0042] Figure 3 Schematic diagram of the detection results of an embodiment of the present invention. DETAILED DESCRIPTION
[0043] In order to make the purpose, technical solutions and advantages of the embodiments of the present invention clearer, the technical solutions in the embodiments of the present invention will be clearly and completely described below in conjunction with the drawings in the embodiments of the present invention. Obviously, the described embodiments are part of the embodiments of the present invention, not all of the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by ordinary technicians in this field without making creative work belong to the scope of protection of the present invention. Therefore, the following detailed description of the embodiments of the present invention provided in the drawings is not intended to limit the scope of the invention claimed for protection, but merely represents the selected embodiments of the present invention. Based on the embodiments of the present invention, all other embodiments obtained by ordinary technicians in this field without making creative work belong to the scope of protection of the present invention.
[0044] See also Figure 1 and Figure 2 As shown, the present invention provides an unsupervised cross-domain forged speech detection method based on fusion loss constraints, comprising the following steps.
[0045] S101, detection network construction step, the detection network includes a shallow feature pre-extraction module, a composite residual backbone network based on a compression-excitation module and a classification module; the acoustic features obtained after speech preprocessing are input into the shallow feature pre-extraction module, the shallow feature pre-extraction module extracts shallow artifact features and outputs them to the composite residual backbone network, the composite residual backbone network extracts deep artifact features and outputs them to the classification module; the classification module outputs a binary classification result for distinguishing the authenticity of the speech.
[0046] Specifically, the preprocessing includes: performing length unification, CQT pre-transformation and pre-emphasis processing on each speech sample to convert it into acoustic features with the same dimension. In this embodiment, a specific preprocessing process is as follows.
[0047] (1) Assume that the speech sample is X, and there are X i =[x1,x2,…,x m ], 1≤i≤m, m is the number of speech samples. Preprocessing operations are performed on each speech sample in different language datasets of the source domain (English) and the target domain (Chinese), and their speech length is uniformly controlled to 6 seconds. Speech longer than 6 seconds is trimmed and retained to 6 seconds, and speech shorter than 6 seconds is repeated to 6 seconds to unify the speech length.
[0048] (2) Perform CQT transformation on the speech sample X after the above speech length processing, assuming that the highest and lowest sounds processed are f max and f min .f k represents the frequency of the kth component, k=0,1,…,K-1, b is the number of spectrum lines contained in one octave, and its calculation method is: K is the total number of frequency segments, indicating how many intervals the frequency range is divided into. The constant Q is only related to b and is calculated as follows: N k is the frequency-dependent window length, where f s Represents the sampling frequency. After CQT transformation, its expression is: ω(n) is a window function of length ; k is the frequency index of the CQT spectrum. CQT is similar to Fourier transform, but the horizontal axis frequency of its spectrum is not linear, but based on log2.
[0049] (3) Then, the speech after CQT transformation is pre-emphasized, which is formally expressed as: where f E is the pre-emphasis function, It is the acoustic characteristic after pre-emphasis processing.
[0050] In this embodiment, the proportional coefficient of the pre-emphasis processing is α=0.97. This operation can enhance the high-frequency components and suppress the low-frequency components to a certain extent. In order to reduce spectrum leakage, smooth speech, and reduce noise interference, a Hanning window is used here. Each audio is resampled to 16000 Hz. In order to ensure that the feature scale of each batch input is the same, the padding function provided by numpy is used to fill the speech sampling length to 96000, that is, 6s for each speech.
[0051] Furthermore, the construction process of the detection network is specifically as follows.
[0052] S1011, the function of the shallow feature pre-extraction module is to extract shallow artifact features, which is mainly composed of a 7×7 two-dimensional convolution, batch normalization, ReLU activation function and maximum pooling layer. The acoustic features of the speech obtained after preprocessing As the input of the preprocessing convolution block, it undergoes 7×7 two-dimensional convolution, batch normalization, and ReLU activation function to generate shallow features X S , the process can be formalized as: Among them, f relu is the ReLU activation function, f bn is batch normalization, f cnt7 is a 7×7 two-dimensional convolution, θ cnt7 is the parameter of 7×7 two-dimensional convolution, θ bn is the parameter of batch normalization, f maxpool is the maximum pooling layer.
[0053] S1012, see Figure 1 In this embodiment, the total number of residual blocks based on compression-excitation is 4, and the source domain and target domain data actually share a set of network parameters. The function of the residual block based on compression-excitation is to extract deep artifact features. The input features of the residual block are The workflow of each compression-excitation based residual block is as follows.
[0054] Taking input feature X s As the input of the separable residual block, one side undergoes a 1×1 two-dimensional convolution to obtain the feature X sxx , X sxx =f cnt1 (X s ,θ cnt1 ). Among them, X s is the input feature, f cnt1 is a 1×1 two-dimensional convolution, θ cnt1 is the parameter of the 1×1 2D convolution.
[0055] At the same time, the other side first undergoes 3×3 two-dimensional convolution, batch normalization, and ReLU activation function in sequence. After repeating the above steps, it undergoes another 3×3 two-dimensional convolution to obtain feature X syy , X syy =f cnt3 (f relu (f bn (f cnt3 (f relu (f bn (f cnt3 (X s ,θ cnt3 )),θ bn ),θ cnt3 ),θ bn )),θ cnt3 ).
[0056] In addition, the features obtained by three 3×3 convolutions of the residual block based on compression-excitation A compression operation is performed, where C1 represents the number of channels, H1 represents the height of the feature map, and W1 represents the width of the feature map, thereby obtaining the global characteristics of the channel level This step is implemented by a global average pooling layer; then the feature is excitation operated to learn the relationship between different channels, and the weights of different channels are multiplied by the original feature map to get the final feature. This step is implemented by two fully connected (FC) layers, a ReLU layer and a sigmoid layer. The first FC layer plays a dimensionality reduction role and converts the feature dimension into After that, ReLU activation is used, and the final FC layer restores the features to their original dimensions. Then multiply the learned activation values of each channel (after sigmoid activation) by the original feature X syy , assign weights to each channel to get Such an attention mechanism allows the model to pay more attention to the channel features with the most information, while suppressing unimportant channel features. Its formal expression is: syy ′=f relu (f fc2 (f relu (f fc1 (f global (X syy )),θ fc1 )),θ fc2 ).
[0057] In order to avoid network function degradation caused by information loss, skip connections are used, that is, summing X sxx With X syy ′, and then undergo batch normalization to obtain feature X D, the process can be formalized as: D =f bn (f relu (X sxx +X syy ′),θ bn ). Among them, f relu is the ReLU activation function, f bn is batch normalization, θ bn is the parameter of batch normalization. Among them, f cnt3 is a 3×3 two-dimensional convolution, θ cnt3 is the parameter of 3×3 two-dimensional convolution, f bn is batch normalization, θ bn is the parameter of batch normalization, f relu is the ReLU activation function.
[0058] The composite residual backbone network based on the compression-excitation module consists of four residual blocks mentioned above. After the above preprocessing convolution block, the shallow feature representation X is obtained s , X s After four more residual blocks, the residual feature X is obtained. RES , the process can be formalized as: RES =(M RSB2D (M RSB2D (M RSB2D (M RSB2D (X s )))))) Among them, M RSB2D It is a residual block. The backbone network contains four residual blocks of this structure to extract deep artifact features for subsequent classification tasks.
[0059] S1013,X RES After a global average pooling layer, the feature vector is obtained, and then after a dimension reduction function, it is connected to the classification layer f composed of multiple fully connected layers and ReLU activation functions. cls , output the binary classification result, so as to achieve the purpose of distinguishing the true and false situations of the target language speech. The process can be formalized as follows:
[0060] y=f fc5 ((f relu (f fc4 (f relu (f fc3 (f squeeze (f avg (X RES ))),θ fc3 )),θ fc4 )),θ fc5 );
[0061] where favg is the global average pooling layer, f squeeze is the dimension compression function, f fc3 、f fc4 、f fc5 All are fully connected layers, θ fc3 ,θ fc4 and θ fc5 are the parameters of the fully connected layer, f relu is the ReLU activation function, and y is the predicted value.
[0062] S102, detection network training step, inputting the pre-processed labeled source domain training set speech samples and unlabeled target domain training set speech samples in batches into the detection network for training; during the training process, using the global loss constraint detection network, to update the network parameters; the global loss constraint is obtained by weighted sum of the hybrid regularization loss and the fusion loss; wherein the hybrid regularization loss is used as a generalization booster; the fusion loss is the loss of shallow detail features and deep global semantic features between the fusion source domain speech and the target domain speech.
[0063] In this step, the hybrid regularization loss is first used as a generalization booster to improve the generalization of the detection network for unseen forged speech. It improves the generalization ability of the detection network itself in cross-domain situations to a certain extent. Then, the fusion loss L that integrates the shallow detail features and deep global semantic features between the source domain and the target domain speech is used to generate the generalization loss. fusion Combined with the constraint detection network, a cross-domain forged speech detection network is constructed, as follows.
[0064] S1021, inputting the acoustic feature group of the labeled source domain speech and the unlabeled target domain speech to be trained into the forged speech detection network for training in an unsupervised manner to generate the forged speech detection network.
[0065] S1022, using the mixed regularization loss to calculate the classification layer output results and label values, obtain the loss value, and then optimize the model parameters. Mixed regularization randomly mixes different source domain speech samples and their labels, and weights the cross entropy loss of the two sets of labels to obtain a mixed regularization loss, thereby learning the continuous relationship between multiple samples, which helps to improve the classification performance of the model on unseen data. The process can be formalized as: in, represents two sets of randomly selected mixed training samples, represents the labels of two randomly selected mixed training samples, x a is a training sample in the source domain, x b is another training sample in the source domain; a For x a The label, y b For xb , λ is the mixing coefficient randomly sampled from the Beta distribution, λ~Beta(α,α), α∈(0,∞) is a hyperparameter.
[0066] The implementation of hybrid regularization is carried out through the following equivalent loss function, which can be formalized as: in, The softmax probability of including mixed samples is determined by the detection network Output: CE(·,·) is the standard cross entropy loss.
[0067] S1023, the fusion loss adopts a measurement method that combines the alignment of first-order statistical information and second-order statistical information to jointly minimize the distance between the source domain and the target domain. The feature vectors output by the source domain and the target domain are measured at the fully connected layer of the network to construct a loss function, and finally the loss function is minimized during the training process to align the distribution of the source domain and the target domain as much as possible and minimize the domain difference, thereby achieving the effect of domain adaptation under unsupervised conditions. First, the loss function calculation method based on the first-order statistical information is specifically to consider the standard distribution distance metric, the maximum mean difference MMD (Maximum Mean Discrepency), hoping to minimize the distance between domains, which is formally expressed as: Among them, φ(·) is a specific feature representation, φ(·): X→H, that is, the feature space mapping from X to H, H is the reproducing kernel Hilbert space, and X is a non-empty set, which is the source domain and target domain speech samples. φ(·) operates on the source domain and target domain data. Then the MMD distance between them is calculated. denote the source domain and the target domain respectively, Represents a specific piece of speech data belonging to the source domain or the target domain respectively. In order to minimize the distance between domains and train a classifier at the same time, the following loss function is minimized, that is, the loss function based on first-order statistical information: Among them, the hyperparameter δ determines the extent to which the distance between the two domains is shortened.
[0068] Secondly, the calculation method of the loss function based on the second-order statistical information is as follows: Consider the problem of classifying speech X in the true and false binary classification problem. Assume that there is a classifier, which is formally expressed as: Among them, To represent a common neural network, the performance of the classifier f depends on the parameter θ, which is optimized by minimizing the cross entropy loss function. The cross entropy loss function is formally expressed as follows: Among them, X, Z are the sets of input speech and its labels. For each speech x i∈X, and its label z i ∈Z, use the inner product <·,·> to calculate the classifier result With the corresponding audio tag z i The similarity measure between i is a two-dimensional one-hot encoded vector. Specifically, if x i Belongs to the first category, then z i =1, otherwise z i =0.
[0069] In the standard supervised case, the optimal parameter θ should be obtained by minimizing the cross entropy loss function H(X,Z); however, in the unsupervised case, the choice of θ should promote the source domain To the target domain Good migration effect. Due to the unsupervised situation, that is, the target domain There are no labels to use, so domain adaptation should be performed at the feature level. In the case of relevance alignment, the problem of minimizing the above cross entropy function is replaced by: only calculate the speech of the source domain and its tags The cross entropy loss between them is used to align the covariance feature representation to obtain the optimal parameter θ. S , Mathematically, it is a symmetric and positive definite matrix of a Riemann manifold with non-zero curvature. Using the ordinary Euclidean metric to measure the correlation is not optimal for it, so it is modified to the squared logarithmic Euclidean distance: To solve the problem of domain adaptation, this step is calculated on the activation value calculated at a certain layer of the network f(·,θ), and its overall formal expression is as follows: The activation values of the source domain and the target domain are and The activation values of d dimensions are stacked in columns, and the covariance of the source domain and the target domain are represented by C S , It is calculated by the centering matrix J (centering matrix), which is formally expressed as follows: and They represent the set of source domain speech and the label of source domain speech respectively, and β is a hyperparameter. Finally, the fusion loss of the present invention is
[0070] S1024, global loss constraint It is expressed as follows:
[0071]
[0072] Among them, σ and ξ are empirical parameters; represents fusion loss; CE mixup represents the hybrid regularization loss.
[0073] In the speech detection step S103, a target domain speech sample to be detected is obtained, and after preprocessing, the acoustic features of the target domain speech sample are input into the trained detection network for detection to generate a discrimination result.
[0074] In this embodiment, it is assumed that the speech sample to be detected Then, the detected speech sample X is processed according to the preprocessing operation in S101. * Preprocessing is performed to obtain the acoustic features X' corresponding to the detected speech samples. * , and then input the acoustic features into the trained forged speech detection network for detection to generate the discrimination results.
[0075] The acoustic feature X' * The sample is input into the trained forged speech detection network to obtain the predicted probability p of the speech sample to be tested. When p≥0.5, the speech to be tested is natural speech, otherwise it is forged speech.
[0076] See also Figure 3 As shown, in this embodiment, in order to evaluate the unsupervised cross-domain forged speech detection method based on fusion loss constraints, the source domain and target domain languages contain training sets, validation sets and test sets. The source domain speech training set has 22,541 speech samples, the validation set has 13,552 speech samples, and the test set has 19,190 speech samples; the target domain speech training set has 26,850 speech samples, the validation set has 18,124 speech samples, and the test set has 18,124 speech samples. The applicant uses an unsupervised method to train the detection model on the cross-domain data set, evaluates the performance of the model on the test set, and compares it with the currently common domain-adapted forged speech detection algorithm. In order to be able to intuitively evaluate the performance of the model, the applicant uses the mainstream evaluation indicators: Equal Error Rate (EER) and Accuracy (ACC). Equal Error Rate is the error rate when the false acceptance rate and the false rejection rate are equal. It reflects the performance of the forged speech detection algorithm. The smaller this value is, the better the performance. The accuracy rate reflects the classification accuracy of the detection model. The higher the accuracy rate, the smaller the number of misclassified speech. Figure 3 The experimental results in show that the unsupervised cross-domain forged speech detection method based on fusion loss constraints has great advantages in detection performance, surpassing the existing detection methods.
[0077] Although the specific embodiments of the present invention are described above, those skilled in the art should understand that the specific embodiments described in this patent are only illustrative and not intended to limit the scope of the present invention. Equivalent modifications and changes made by those skilled in the art in accordance with the spirit of the present invention should be included in the scope of protection of the claims of the present invention.
Claims
1. An unsupervised cross-domain forged speech detection method based on fusion loss constraints, characterized in that: include: The detection network construction step includes a shallow feature pre-extraction module, a composite residual backbone network based on a compression-excitation module, and a classification module; the acoustic features obtained after speech preprocessing are input into the shallow feature pre-extraction module, the shallow feature pre-extraction module extracts shallow artifact features and outputs them to the composite residual backbone network, the composite residual backbone network extracts deep artifact features and outputs them to the classification module; the classification module outputs a binary classification result for distinguishing the authenticity of speech; The detection network training step is to input the pre-processed labeled source domain training set speech samples and the unlabeled target domain training set speech samples into the detection network in batches for training; During the training process, a global loss constraint is used to detect the network and update the network parameters; the global loss constraint is obtained by weighted summing the hybrid regularization loss and the fusion loss; wherein the hybrid regularization loss is used as a generalization booster; the fusion loss is the loss of shallow detail features and deep global semantic features between the fusion source domain speech and the target domain speech; The speech detection step is to obtain the target domain speech sample to be detected, and after preprocessing, the acoustic features of the target domain speech sample are input into the trained detection network for detection to generate a discrimination result; The fusion loss is a loss function based on first-order statistical information and a loss function based on second-order statistical information, which is expressed as follows: in, represents the fusion loss; represents the loss function based on first-order statistical information; represents the loss function based on second-order statistical information; Loss function based on first-order statistics It is expressed as follows: Among them, δ∈(0,∞) is a hyperparameter that determines the degree of shortening the distance between the two domains; Represents the MMD distance between the source domain and target domain speech samples; Among them, φ(·) is a specific feature representation that operates on the source domain and target domain speech samples; represents the source domain speech set; represents the target domain speech set, Indicates a specific piece of voice data belonging to the source domain; Indicates a specific piece of voice data belonging to the target domain; Loss function based on second-order statistics It is expressed as follows: in, The label representing the source domain speech; and Represent the covariance of the source domain and the target domain respectively; represents the cross entropy loss function; represents the squared log Euclidean distance; β is a hyperparameter; and It is calculated by the centralization matrix J and is expressed as follows: in, and represent the activation values of the source domain and the target domain respectively.
2. The unsupervised cross-domain forged speech detection method based on fusion loss constraint according to claim 1 is characterized in that: The preprocessing includes: performing length unification, CQT pre-transformation and pre-emphasis processing on each speech sample, and converting it into acoustic features with the same dimension.
3. The unsupervised cross-domain forged speech detection method based on fusion loss constraint according to claim 1 is characterized in that: The shallow feature pre-extraction module includes a two-dimensional convolution, batch normalization, ReLU activation function and maximum pooling layer connected in sequence, takes the acoustic features of the speech sample as input, and extracts shallow artifact features.
4. The unsupervised cross-domain forged speech detection method based on fusion loss constraint according to claim 1 is characterized in that: The composite residual backbone network based on the compression-excitation module includes four residual blocks connected in sequence. After the shallow artifact features pass through the four residual blocks, the deep artifact features, namely the residual features, are obtained.
5. The unsupervised cross-domain forged speech detection method based on fusion loss constraint according to claim 1 is characterized in that: The classification module includes a global average pooling layer, a dimension compression function, multiple fully connected layers and multiple ReLU activation functions connected in sequence; the feature vector obtained by the deep artifact feature through a global average pooling layer is then reduced in dimension through a dimension compression function, and then connected to multiple fully connected layers and ReLU activation functions to output a binary classification result to determine the authenticity of the speech.
6. The unsupervised cross-domain forged speech detection method based on fusion loss constraint according to claim 1 is characterized in that: Mixed regularization randomly mixes two sets of source domain speech samples and their labels, and weights the cross entropy loss of the two sets of labels to obtain a mixed regularization loss, thereby learning the continuous relationship between multiple speech samples. The process is expressed as follows: in, represents two sets of mixed training samples randomly selected from the source domain, represents the labels corresponding to two sets of randomly selected mixed training samples, x a is a training sample in the source domain, x b is another training sample in the source domain; a For x a The label, y b For x b The label of , λ is the mixing coefficient randomly sampled from the Beta distribution, λ~Beta(α,α), α∈(0,∞) is a hyperparameter; The implementation of hybrid regularization is through the following equivalent loss function The implementation process is formalized as: in, The softmax probability of including mixed samples is determined by the detection network Output: CE(·,·) is the standard cross entropy loss.
7. The unsupervised cross-domain forged speech detection method based on fusion loss constraint according to claim 1 is characterized in that: Global loss constraint It is expressed as follows: Among them, σ and ξ are empirical parameters; represents fusion loss; CE mixup represents the hybrid regularization loss.
Citation Information
Patent Citations
Synthetic speech detection method based on partially connected microarchitecture search and residual auto-encoder
CN117373427A