A method and device for voice forgery identification based on dynamic multi-view fusion

Through dynamic multi-view fusion and clustering integration algorithms, combined with pseudo-label enhancement and confidence calibration, K nearest neighbor classifiers are trained, and the problem of insufficient overfitting and generalization capabilities in the existing speech depth forgery detection technology is solved, achieving higher detection accuracy and robustness.

CN119864055BActive Publication Date: 2025-06-10SOUTH CHINA UNIV OF TECH
View PDF 4 Cites 0 Cited by

Patent Information

Application Number
CN202510355281.0
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-03-25
Publication Date
2025-06-10
Estimated Expiration
2045-03-25

AI Technical Summary

Technical Problem

The existing speech depth forgery detection technology has problems such as overfitting, insufficient generalization ability and insufficient feature utilization.

Method used

A speech forgery identification method of dynamic multi-view fusion is proposed. By constructing a speech multi-view dataset, the dynamic weight clustering ensemble algorithm is used for clustering, and the K nearest neighbor classifier is trained in combination with pseudo-label enhancement, dynamic K value selection and confidence calibration.

Benefits of technology

It improves the robustness and accuracy of the detection model, reduces the error rate, and significantly improves the ability to generalize unseen data points.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119864055B_ABST
    Figure CN119864055B_ABST
Patent Text Reader

Abstract

The present invention discloses a method and device for voice forgery identification based on dynamic multi-view fusion. The method includes the following steps: obtaining multi-view dynamic feature data according to the original voice signal to construct a multi-view data set; clustering the data in the multi-view data set by using a dynamic weight clustering ensemble algorithm to generate pseudo-labels; integrating the pseudo-labels with the multi-view dynamic feature data to obtain a first enhanced data set and training a K-nearest neighbor classifier; using the trained K-nearest neighbor classifier to predict unseen data points other than the original voice signal in the voice signal to be identified. The present invention uses unsupervised learning to mine multi-perspective voice information, solving the deficiencies of existing methods in generalization, robustness, and feature utilization. The present invention has excellent performance on multiple data sets, with significant improvement in key indicators, strong generalization and robustness against different forged data points, providing a new path for voice deep forgery detection.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of voice deepfake detection, and particularly relates to a method and device for voice forgery identification by dynamic multi-view fusion. Background Art

[0002] With the rapid development of artificial intelligence-generated content (AIGC) technology, voice synthesis technology has made remarkable progress in anthropomorphism, authenticity, and naturalness. In particular, the application of large models has made the boundary between forged voice and real voice increasingly blurred. Although this technological progress has promoted the development of the voice synthesis field, it has also brought serious security risks, such as identity fraud and false information dissemination. Therefore, it is of great practical significance to develop efficient and robust voice deepfake detection technology.

[0003] Existing voice deepfake detection technologies mainly rely on supervised learning methods, the core of which includes acoustic feature extraction and classification algorithms. For example, Chinese patent document CN119170052A discloses a voice forgery detection method based on a multi-scale GMM-ResNet model, which uses a feature fusion method to train in the ResNet model; Chinese patent document CN117809694A discloses a forged voice detection method and system based on temporal multi-scale feature representation learning, which extracts preliminary features through wav2vec2.0 and then inputs them into a convolutional network based on multi-scale time series to extract a feature matrix. The feature matrix is input into the SCG-Res2Net50 model for training. The above technologies are all implemented based on self-supervised learning models, and moreover, for the extraction of acoustic features, it is still concentrated on the deep mining method. Although the existing methods have achieved certain results on specific datasets, there are still the following limitations:

[0004] 1. Overfitting problem: Existing methods mainly rely on supervised learning and are prone to overfitting in the case of limited data, resulting in insufficient generalization ability of the model in unknown forged audio scenarios.

[0005] 2. Limitations of feature mining: Current research mainly focuses on the deep mining and fusion of acoustic features (i.e., "deep analysis"), while the multi-angle research of voice data (i.e., "width analysis") has not been fully explored, restricting the further improvement of detection performance.

[0006] 3. Neglect of unsupervised learning: Existing methods mostly rely on labeled data, while unsupervised learning methods have significant advantages in utilizing a large amount of unlabeled data and adapting to different data distributions, but their application in the field of voice deepfake detection is still relatively limited. Summary of the Invention

[0007] The present invention aims to solve the problems of overfitting, insufficient generalization ability, and insufficient feature utilization existing in the existing voice deep forgery detection technology, and proposes a voice forgery discrimination method and device based on dynamic multi-view fusion. By constructing a voice multi-view data set, a dynamic weight clustering integration algorithm is used for clustering and combined with pseudo-label enhancement, dynamic K value selection, and confidence calibration to train a K-nearest neighbor classifier, so as to improve the robustness and accuracy of the detection K-nearest neighbor classifier.

[0008] The purpose of the present invention is achieved by at least one of the following technical solutions.

[0009] A voice forgery discrimination method based on dynamic multi-view fusion includes the following steps:

[0010] Including the following steps:

[0011] S1. According to the original voice signal, obtain multi-view feature data including a reconstruction error view, an acoustic feature view, and a quality assessment view, and construct a multi-view data set;

[0012] S2. Use a dynamic weight clustering integration algorithm to cluster the multi-view feature data in the multi-view data set to generate pseudo-labels;

[0013] S3. Integrate the pseudo-labels with the multi-view feature data to obtain a first enhanced data set for training a K-nearest neighbor classifier;

[0014] S4. Use the trained K-nearest neighbor classifier to predict the unseen data points other than the original voice signal in the voice signal to be discriminated.

[0015] Further, in step S1, a voice generation model HiFiGAN is used to calculate the reconstruction error view, and time-frequency domain residual analysis is introduced to quantify local anomalies in the voice generation process, specifically as follows:

[0016]

[0017] Among them, represents the reconstruction error; is a variable representing the original voice signal; represents the original voice signal at time value; represents the total duration of the voice signal (the length of the time series), that is, the total number of time points from the start time to the end time; represents the decoder in the voice generation model HiFiGAN; represents the encoder in the voice generation model HiFiGAN; is a hyperparameter used to balance the weights of the two terms in the formula and is a manually set constant; represents the original speech signal At time the total variation, which is an index used to measure the degree of local variation of the signal;

[0018] The audio coding model Encodec is used to obtain the acoustic feature view, extract multi-scale acoustic coding, combine with a convolutional network for feature dimensionality reduction, and then calculate the inner product and normalize the dimensionality-reduced vector, as follows:

[0019] The original speech signal (such as in wav or MP3 format) is input into the audio coding model Encodec to generate multi-scale acoustic coding; these codings are an abstract representation of the audio data, which contain important information about the audio, such as the pitch, timbre, rhythm, and other features at different scales. The multi-scale acoustic coding is first passed through a convolutional layer to extract local features, then the global average pooling layer is used to calculate the mean of each channel to compress it into a one-dimensional vector, and finally the inner product of the one-dimensional vector itself is calculated to obtain the acoustic feature view;

[0020] The UTMOS model specifically for speech quality assessment is used to obtain the quality assessment view, as follows:

[0021] The original speech signal (such as in wav or MP3 format) is input into the UTMOS model, and the overall speech quality assessment score is output and the data is normalized to obtain the quality assessment view;

[0022] The structure of the multi-view dataset is: {reconstruction error view, acoustic feature view, quality assessment view}.

[0023] Furthermore, step S2 includes the following steps:

[0024] S2.1. Adjust the weights of each view in the multi-view dataset based on information entropy;

[0025] S2.2. Construct a hypergraph that fuses the similarity information of each view;

[0026] S2.3. Based on the hypergraph in step S2.2, use spectral clustering to generate meta-cluster groups, as follows:

[0027] First, construct the Laplacian matrix of the hypergraph and perform eigenvalue decomposition on the Laplacian matrix to obtain the eigenvalues; then automatically determine the optimal number of clusters through the eigenvalue gap method , select the first matrix composed of the eigenvectors corresponding to the smallest eigenvalues, Each row is regarded as a new data point, and meta-cluster groups are obtained by clustering with traditional clustering algorithms such as K-means;

[0028] S2.4. Calculate the scores of the data points in the multi-view dataset to the meta-cluster groups, and determine the pseudo-labels, specifically as follows:

[0029] By the distance metric method, calculate the average Euclidean distance from the data point to all data points in the meta-cluster group to obtain the score of the data point to the meta-cluster group, and then according to the minimum distance method, take the label of the meta-cluster group with the smallest distance to determine the pseudo-label of the data point.

[0030] Furthermore, in step S2.1, the weights of each view in the multi-view dataset are dynamically adjusted according to the information entropy difference between views;

[0031] In the multi-view dataset, the th view's information entropy is calculated as follows: Take the probability of each data point appearing in the th view, multiply it by the logarithm of this probability, add up the results calculated for all data points, and finally take the negative value;

[0032] After obtaining the information entropy of each view, first take the exponential of the information entropy of the th view as the numerator; then add up the exponentials of the information entropies of the three views as the denominator; divide the numerator by the denominator to obtain the weight of the th view, specifically as follows:

[0033]

[0034]

[0035] Among them, represents the weight of the th view in the multi-view dataset; represents the th view's information entropy, and the larger its value, the richer the information and the higher the uncertainty contained in this view; represents the exponential function; represents the th view's number of data points, and the number of data points in all views is the same; represents the th rd data point appearing probability in the

[0036] Furthermore, in step S2.2, construct a hypergraph that fuses the similarity information of each view, specifically as follows:

[0037] First, perform multi-view clustering. The multi-view dataset is , represents the th view; After clustering the data in the multi-view dataset times respectively, a set of clustering partitions is obtained, where represents the th clustering member (i.e., the base cluster), and each clustering member consists of multiple clustering subsets, denoted as , where is the total number of clustering subsets in the th clustering partition, represents the rd clustering subset in the th clustering member;

[0038] The data points in the multi-view dataset are uniformly numbered to calculate the similarity between data points without distinguishing specific views. Then, the multi-view similarity is fused. First, the relevant similarities under each view are calculated and summed, and then the cross-view Figure 1 consistency coefficient is added to the product of the cosine similarity of the feature vectors of the th data point and the th data point in the multi-view dataset, obtaining the similarity between the th data point and the th data point in the multi-view dataset Specifically as follows:

[0039]

[0040] Among them, represents the similarity between the th data point and the th data point in the multi-view dataset; represents the weight of the th view in the multi-view dataset; During the clustering process of the th view, the th data point will be assigned to a clustering subset, and all the features in this clustering subset together form the feature set of the th data point under the th view ; represents the th view, the th data point belongs to the feature set and the th data point belongs to the feature set between the similarity; represents the cross-view Figure 1The consistency coefficient is a constant parameter set artificially; denotes the feature vector of the th data point and the cosine similarity between the feature vectors of the

[0041] th data point.

[0042] S3.1. Concatenate the pseudo-label and multi-view feature data to obtain the first augmented dataset. The structure of the first augmented dataset is: {reconstruction error view, acoustic feature view, quality assessment view, pseudo-label}. Divide the first augmented dataset into a validation set and a training set;

[0043] S3.2. Adjust the value of K of the K-nearest neighbor classifier according to the data point density in the training set;

[0044] S3.3. Input the training set into the K-nearest neighbor classifier. The K-nearest neighbor classifier calculates the distances between the data points in the training set using the weighted Euclidean distance according to the value of K in step S3.2 to find the K nearest neighbors of each data point, and determines the data point category through the voting method according to the pseudo-label (neighbor label) to obtain the K-nearest neighbor classifier with preliminary training completed. Evaluate the performance (accuracy, recall rate, F1 value) of the K-nearest neighbor classifier with preliminary training completed using the validation set. If the evaluation result meets the requirements or reaches the set maximum number of iterations, execute step S3.4; otherwise, the parameter in the data point density calculation will be increased or decreased by the set step size, the number of iterations will be incremented by 1, and return to step S3.2;

[0045] S3.4. Calibrate the output confidence of the K-nearest neighbor classifier with preliminary training completed using the Bayesian model to obtain the trained K-nearest neighbor classifier.

[0046] Furthermore, in step S3.2, adaptively adjust the value of K according to the data point density in the training set; for the th data point in the training set, traverse all the data points in the training set. For the th data point , calculate the distance between the th data point and the th data point ; if the distance is less than the pre-set range , then the value of the indicator function is 1; otherwise, it is 0. Add up the values of the indicator functions corresponding to all data points to obtain the th data point The number of local data points , which reflects the number of the remaining data points within the set range centered on , and embodies the density of the data points around this data point;

[0047] After obtaining the number of local data points of the th data point , by taking the square root of and rounding up, the adaptively adjusted K value for the th data point is obtained as follows: The adaptively adjusted K value is specifically as follows:

[0048]

[0049]

[0050] Among them, represents the adaptively adjusted K value for the th data point in the training set; represents the number of local data points of the th data point , that is, the number of the remaining data points within the set range centered on , which reflects the density of the data points around the data point ; represents rounding up; represents the total number of data points in the training set; represents the indicator function, when holds, the value of is 1, otherwise it is 0; represents the distance between the th data point and the th data point and the th data point .

[0051] Furthermore, in step S3.3, if the evaluation result shows that the K value of the K-nearest neighbor classifier is too small, it may mean that the calculation range of the local density of the data points is too small, then increase the set step size of the parameter ; if the evaluation result shows that the K value of the K-nearest neighbor classifier is too large, it may mean that the calculation range of the local density of the data points is too large, then decrease the set step size of the parameter .

[0052] Furthermore, in step S3.4, a Bayesian probability model is adopted to calibrate the confidence of the output of the preliminarily trained K-nearest neighbor classifier;

[0053] For the feature vector of the data point input into the K-nearest neighbor classifier , it is necessary to determine the probability that the class of this data point is equal to a class of the data point under the condition of the given feature vector ;

[0054] Judge whether the class of the -th neighbor among the K nearest neighbors is the same as the class through the indicator function ; when holds, has a value of 1, otherwise 0;

[0055] Sum the weights of the neighbors with class among the K nearest neighbors, that is , and then divide by the total weight of the K nearest neighbors to obtain the probability part that the data point belongs to class preliminarily estimated based on the information of the K nearest neighbors;

[0056] Introduce the information entropy of the feature vector , which reflects the uncertainty of the information contained in the feature vector . Adjust the influence degree of this information entropy term in the whole probability calculation through a hyperparameter , that is .

[0057] Add the probability part that the data point belongs to class preliminarily estimated based on the information of the K nearest neighbors to to obtain the final , is the probability that the data point belongs to class under the condition of the given feature vector of the data point . Considering both the neighbor information and the uncertainty of the feature vector, the confidence calibration of the output of the K-nearest neighbor classifier is realized, making the probability estimation of the classification result more accurate and contributing to more reliable voice forgery identification, as follows:

[0058]

[0059] Among them, represents the probability that the class of the data point is under the condition of the given feature vector of the data point​ is a category of data points; denotes an indicator function, when holds, the value of is 1, otherwise 0; represents the feature vector of the data points input to the K-nearest neighbor classifier, which contains various feature information of the data points in the multi-view feature data and is the basis for classification and confidence calculation; represents the number of the nearest neighbors in the K-nearest neighbor classifier; represents the weight of the th nearest neighbor; represents the category of the th nearest neighbor; represents the sum of the weights of the nearest neighbors with category among the nearest neighbors, reflecting the contribution degree of the nearest neighbors with category = 1 among the nearest neighbors to the category = 1 of the feature vector is a manually set hyperparameter used to adjust the information entropy of the feature vector ; represents the information entropy of the feature vector , and the information entropy is an index used to measure information uncertainty. Here, the information entropy reflects the uncertainty or information richness of the features of the feature vector itself. The higher the information entropy of the feature vector , the greater the uncertainty of its features. In confidence calibration, by introducing the information entropy, the influence of the characteristics of the data points themselves on the confidence of the classification result can be comprehensively considered.

[0060] Furthermore, in step S4, the speech signal to be identified is subjected to steps S1 to S3 to obtain a second enhanced data set corresponding to the speech signal to be identified; the first enhanced data set and the second enhanced data set have the same structure;

[0061] The second enhanced data set is input into the trained K-nearest neighbor classifier to predict the unseen data points other than the original speech signal in the speech signal to be identified.

[0062] The present invention also provides a speech forgery identification device with dynamic multi-view fusion for implementing the above method, including the following modules:

[0063] Data acquisition and preprocessing module: used to obtain the original speech signal and the speech signal to be identified;

[0064] Multi-view Feature Extraction Module: It is used to obtain multi-view feature data based on the original speech signal and construct a multi-view data set;

[0065] Integrated Clustering Module: It is used to cluster the data in the multi-view data set by using the dynamic weight clustering integration algorithm to generate pseudo-labels;

[0066] Classifier Training Module: It is used to integrate the pseudo-labels with the multi-view feature data to obtain the first enhanced data set and train the K-nearest neighbor classifier;

[0067] Prediction Module: It is used to use the trained K-nearest neighbor classifier to predict the unseen data points other than the original speech signal in the speech signal to be authenticated.

[0068] Compared with the prior art, the advantages of the present invention are as follows:

[0069] Construction of multi-view feature data: The present invention extracts multi-perspective features from three perspectives: speech acoustic features, speech overall quality evaluation, and speech generation reconstruction error, ensuring the complementarity and consensus of the data.

[0070] Dynamic multi-view weight allocation: Traditional multi-view methods usually use fixed weights to fuse information from different views, which may lead to overemphasis on the information of some views and neglect of the information of other views, resulting in the problem of feature redundancy. The present invention dynamically adjusts the view weights through information entropy, can adaptively allocate weights according to the information richness of different views, avoids the problem of feature redundancy, and improves the effectiveness of feature fusion.

[0071] Hypergraph-driven clustering integration: Hypergraph can better model the high-order associations between multi-views. The present invention introduces hypergraph theory into speech forgery detection, can more comprehensively consider the relationships between different views, and thus improves the stability and accuracy of clustering.

[0072] Adaptive hybrid classifier: The performance of the traditional K-nearest neighbor classifier may be affected when facing complex data distributions. The present invention trains the K-nearest neighbor classifier by combining pseudo-label enhancement, dynamic K value selection, and confidence calibration, can better adapt to different data distributions, improves the accuracy and robustness of the classifier, and reduces the error rate by more than 15% compared with the traditional K-nearest neighbor classifier. Description of the Drawings

[0073] Figure 1 It is a schematic structural diagram of a voice forgery identification device with dynamic multi-view fusion in an embodiment of the present invention.

[0074] Figure 2 It is a schematic structural diagram of the audio coding model Encodec in an embodiment of the present invention.

[0075] Figure 3 This is the flowchart of the steps of a method for voice forgery identification with dynamic multi-view fusion in an embodiment of the present invention. Detailed implementation manners

[0076] To make the objectives, technical solutions and advantages of the present invention clearer and more understandable, the following takes examples in combination with the accompanying drawings to elaborate on the specific implementation of the present invention in detail.

[0077] In one embodiment, a method for voice forgery identification with dynamic multi-view fusion, as Figure 3 shown, includes the following steps:

[0078] S1. According to the original voice signal, obtain multi-view feature data and construct a multi-view data set;

[0079] The multi-view feature data includes a reconstruction error view, an acoustic feature view, and a quality assessment view;

[0080] The HiFiGAN voice generation model (see Jungil Kong, Jaehyeon Kim, and Jaekyoung Bae. 2020. HiFi-GAN: generative adversarial networks for efficient and high fidelity speech synthesis. In Proceedings of the 34th International Conference on Neural Information Processing Systems (NIPS '20). Curran Associates Inc., Red Hook, NY, USA, Article 1428, 17022–17033.) is used to calculate the reconstruction error view, introducing time-frequency domain residual analysis to quantify local anomalies in the voice generation process, specifically as follows:

[0081]

[0082] Among them, represents the reconstruction error; is a variable representing the original voice signal; represents the original voice signal at time value; represents the total duration of the voice signal (the length of the time series), that is, the total number of time points from the start time to the end time; Represents the decoder in the voice generation model HiFiGAN; Represents the encoder in the voice generation model HiFiGAN; Is a hyperparameter used to balance the weights of the two terms in the formula and is a manually set constant; Represents the original speech signal At time The total variation, which is a metric for measuring the degree of local variation of a signal; in one embodiment, .

[0083] Such as Figure 2 As shown, the audio coding model Encodec is used to obtain the acoustic feature view, extract multi-scale acoustic coding, combine with a convolutional network for feature dimensionality reduction, and then take the inner product and normalize the dimensionality-reduced vector, as follows:

[0084] The original speech signal (such as wav, MP3 format) is input into the audio coding model Encodec (see D'efossez, A., Copet, J., Synnaeve, G., & Adi, Y. (2022). High Fidelity Neural Audio Compression. ArXiv, abs / 2210.13438 .), generating multi-scale acoustic coding; these codings are an abstract representation of audio data, which contain important information of the audio, such as the manifestation of features such as pitch, timbre, and rhythm at different scales. The multi-scale acoustic coding is first passed through a convolutional layer to extract local features, then the global average pooling layer is used to calculate the mean of each channel to compress it into a one-dimensional vector, and finally the inner product of the one-dimensional vector itself is taken to obtain the acoustic feature view;

[0085] The UTMOS model dedicated to speech quality assessment (see Saeki, T., Xin, D., Nakata, W., Koriyama, T., Takamichi, S., Saruwatari, H. (2022) UTMOS: UTokyo-SaruLab System for VoiceMOS Challenge 2022. Proc. Interspeech 2022, 4521 - 4525, doi:10.21437 / Interspeech.2022 - 439) is used to obtain the quality assessment view, as follows:

[0086] The original speech signal (such as wav, MP3 format) is input into the UTMOS model, and the overall speech quality assessment score is output and the data is normalized to obtain the quality assessment view;

[0087] The structure of the multi-view dataset is: {reconstruction error view, acoustic feature view, quality assessment view}.

[0088] S2. Cluster the data in the multi-view dataset using the dynamic weight clustering ensemble algorithm to generate pseudo-labels, including the following steps:

[0089] S2.1. Adjust the weights of each view in the multi-view dataset based on information entropy;

[0090] Dynamically adjust the weights of each view in the multi-view dataset according to the information entropy difference between views;

[0091] In the multi-view dataset, the th view's information entropy is calculated as follows: Take the probability of each data point in the th view, take the logarithm and then multiply by this probability, add up the results calculated for all data points, and finally take the negative;

[0092] After obtaining the information entropy of each view, first take the exponent of the information entropy of the th view as the numerator; then add up the exponents of the information entropy of the three views as the denominator; divide the numerator by the denominator to get the weight of the th view, specifically as follows:

[0093]

[0094]

[0095] Among them, represents the weight of the th view in the multi-view dataset; represents the th view's information entropy, and the larger its value, the richer the information and the higher the uncertainty contained in this view; represents the exponential function; represents the th view's number of data points, and the number of data points in all views is the same; represents the th view's th data point appearance probability.

[0096] S2.2. Construct a hypergraph that fuses the similarity information of each view , is the vertex set, including all data points in the multi-view dataset, is the edge set. When the similarity of multiple data points in different views meets the set threshold, these data points are used as a hyperedge, specifically as follows:

[0097] First, perform multi-view clustering. The multi-view dataset is , representing the th view. After performing times of clustering on the data in the multi-view dataset, a set of clustering partitions is obtained, where represents the th clustering member (i.e., the base cluster). Each clustering member consists of multiple clustering subsets, denoted as , where is the total number of clustering subsets in the th clustering partition, represents the th clustering subset in the th clustering member;

[0098] Uniformly number the data points in the multi-view dataset for calculating the similarity between data points without distinguishing specific views, and then fuse the multi-view similarity as follows:

[0099]

[0100] Among them, represents the similarity between the th data point and the th data point in the multi-view dataset; represents the weight of the th view in the multi-view dataset; During the clustering process of the th view, the th data point will be assigned to a clustering subset, and all the features in this clustering subset together constitute the feature set of the th data point under the th view; represents the th view, the th data point belongs to the feature set and the th data point belongs to the feature set between the similarity; represents the cross-view Figure 1 consistency coefficient, which is a constant parameter set by humans. In this embodiment, ; represents the feature vector of the th data point and the th data point's feature vector The cosine similarity between

[0101] S2.3. Based on the hypergraph in step S2.2, use the normalized cut (Ncut) algorithm combined with the spectral clustering idea to partition the hypergraph, and use spectral clustering to generate meta-cluster groups (see Bernhard Schölkopf; John Platt; Thomas Hofmann, "Learning with Hypergraphs: Clustering, Classification, and Embedding," in Advances in Neural Information Processing Systems 19: Proceedings of the 2006 Conference, MIT Press, 2007, pp. 1601-1608.), specifically as follows:

[0102] First, construct the Laplacian matrix of the hypergraph And perform eigenvalue decomposition on the Laplacian matrix to obtain eigenvalues; then automatically determine the optimal number of clusters by the eigenvalue gap method , this method avoids artificially presetting the number of clusters and improves the adaptive ability of clustering; select the first non-zero smallest eigenvalue corresponding eigenvectors to form a matrix , these eigenvectors contain the key information about data point clustering in the hypergraph; regard each row of the matrix as a new data point. In one embodiment, use the k-means algorithm to cluster the matrix to obtain meta-cluster groups . Among them, represents the total number of meta-clusters, and each meta-cluster contains a group of similar data points.

[0103] S2.4. Calculate the scores of data points in the multi-view dataset to the meta-cluster groups, and determine the pseudo-labels, specifically as follows:

[0104] In one embodiment, through the distance metric method, calculate the average Euclidean distance from the data point to all data points in the meta-cluster group to obtain the score of the data point to the meta-cluster group, and then according to the minimum distance method, take the label of the meta-cluster group with the smallest distance to determine the pseudo-label of the data point.

[0105] S3. Integrate the pseudo-labels with the multi-view feature data to obtain the first enhanced dataset, and train the K-nearest neighbor classifier, including the following steps:

[0106] S3.1. Concatenate the pseudo-labels and multi-view feature data to obtain the first augmented dataset, which can increase the dimension and information content of the features, provide more information for the classifier, and improve the performance of the classifier. The structure of the first augmented dataset is: {reconstruction error view, acoustic feature view, quality assessment view, pseudo-labels}. Divide the first augmented dataset into a validation set and a training set;

[0107] S3.2. Adjust the value of K in the K-nearest neighbor classifier according to the data point density in the training set, specifically as follows:

[0108]

[0109]

[0110] where, represents the adaptively adjusted value of K for the th data point in the training set; ; represents the th data point , that is, the number of local data points of the th data point, which is the number of the remaining data points within a set range centered on . It reflects the data point density around the data point ; represents rounding up; represents the total number of data points in the training set; represents the indicator function. When holds, has a value of 1, otherwise 0; represents the th data point and the th data point ;

[0111] S3.3. Input the training set into the K-nearest neighbor classifier. The K-nearest neighbor classifier calculates the distances between the data points in the training set using the weighted Euclidean distance according to the value of K in step S3.2 to find the K nearest neighbors of each data point, and determines the class of the data point through the voting method based on the pseudo-labels (neighbor labels) to obtain the K-nearest neighbor classifier that is initially trained. Evaluate the performance of the initially trained K-nearest neighbor classifier using the validation set. In one embodiment, the accuracy rate, recall rate, and F1 value are used to evaluate the performance of the initially trained K-nearest neighbor classifier. If the evaluation result meets the requirements or reaches the set maximum number of iterations, execute step S3.4. Otherwise, the parameter in the data point density calculation will be increased or decreased by a set step size, the number of iterations is incremented by 1, and return to step S3.2;

[0112] If the evaluation result shows that the K value of the K-nearest neighbor classifier is small, it may mean that the local density calculation range of the data points is too small, then the parameter Increase the set step size; if the evaluation result shows that the K value of the K-nearest neighbor classifier is large, it may mean that the local density calculation range of the data points is too large, then the parameter Decrease the set step size.

[0113] If the distribution of the data points is relatively uniform, a smaller step size can be selected, such as 0.1. If the distribution of the data points varies greatly, a relatively larger step size is selected, such as 0.5.

[0114] S3.4. Calibrate the output confidence of the K-nearest neighbor classifier that has been preliminarily trained using the Bayesian model to obtain the trained K-nearest neighbor classifier.

[0115] In one embodiment, a Bayesian probability model is used to calibrate the confidence of the output of the K-nearest neighbor classifier that has been preliminarily trained, as follows:

[0116]

[0117] Among them, assuming that the category is binary classification, = 1 represents one of the categories; represents the probability that the category of the data point is under the condition of the given feature vector of the data point; represents the indicator function, when holds, the value of is 1, otherwise it is 0; represents the feature vector of the data point input to the K-nearest neighbor classifier, which contains various feature information of the data point in the multi-view feature data and is the basis for classification and confidence calculation; represents the number of the nearest neighbors in the K-nearest neighbor classifier; represents the weight of the th nearest neighbor; represents the category of the th nearest neighbor; represents the sum of the weights of the nearest neighbors whose category is among the nearest neighbors, reflecting the contribution degree of the nearest neighbors whose category is = 1 to the category of the feature vector = 1 considering the weights of the nearest neighbors; is a hyperparameter set by humans to adjust the information entropy of the feature vector , and in one embodiment, ; Represents the feature vector of the information entropy. Information entropy is a metric used to measure the uncertainty of information. Here, the information entropy reflects the uncertainty or information richness of the features of the feature vector itself. The higher the information entropy of the feature vector , the greater the uncertainty of its features. In confidence calibration, by introducing information entropy, the influence of the characteristics of the data points themselves on the confidence of the classification result can be comprehensively considered.

[0118] S4. Use the trained K-nearest neighbor classifier to predict the unseen data points other than the original speech signal in the speech signal to be authenticated;

[0119] Perform steps S1 - S3 on the speech signal to be authenticated to obtain a second enhanced data set corresponding to the speech signal to be authenticated; the first enhanced data set and the second enhanced data set have the same structure;

[0120] Input the second enhanced data set into the trained K-nearest neighbor classifier to predict the unseen data points other than the original speech signal in the speech signal to be authenticated.

[0121] In one embodiment, experiments are conducted on the ASVspoof2019 LA (known attacks) and In-the-Wild 2021 (unknown attacks) data sets.

[0122] Comparison methods: LCNN - BiLSTM (supervised learning), AST (semi-supervised). These comparison methods are selected to comprehensively evaluate the performance of the present invention. LCNN - BiLSTM is a supervised learning method based on deep learning, and AST is a semi-supervised learning method. By comparing with these methods, the advantages of the present invention can be more clearly demonstrated. The experimental results are shown in Table 1.

[0123] Table 1 Experimental results table

[0124]

[0125] Table 1 shows that the method of the present invention performs excellently in terms of the equal error rate (EER) index, achieving the best balance between performance and robustness.

[0126] In one embodiment, a voice forgery discrimination device for dynamic multi-view fusion, as Figure 1 shown, includes the following modules:

[0127] Data acquisition and preprocessing module: used to obtain the original speech signal and the speech signal to be authenticated;

[0128] Multi-view Feature Extraction Module: It is used to obtain multi-view feature data according to the original speech signal and construct a multi-view data set;

[0129] Integrated Clustering Module: It is used to cluster the data in the multi-view data set by using the dynamic weight clustering integration algorithm to generate pseudo-labels;

[0130] Classifier Training Module: It is used to integrate the pseudo-labels with the multi-view feature data to obtain the first enhanced data set and train the K-nearest neighbor classifier;

[0131] Prediction Module: It is used to use the trained K-nearest neighbor classifier to predict the unseen data points other than the original speech signal in the speech signal to be authenticated.

[0132] A voice forgery authentication method and device based on dynamic multi-view fusion proposed by the present invention uses unsupervised learning to mine multi-perspective voice information, and solves the deficiencies of existing methods in generalization, robustness, and feature utilization. In terms of method, multi-view data construction accurately extracts features, and dynamic weight clustering integration improves the clustering effect. Combining pseudo-label enhancement, dynamic K-value selection, and confidence calibration to train the K-nearest neighbor classifier enhances the classification performance. The device ensures the high efficiency and real-time performance of actual applications. Experiments show that the present invention has excellent performance on multiple data sets, significantly improves key indicators, has strong generalization and robustness against different forged data points, and provides a new path for voice deep forgery detection.

[0133] Through the above embodiments, it is detailedly shown how to use the technical method of the present invention for voice forgery authentication. From multi-view data construction to integrated clustering, then to classifier training and prediction, each step is closely combined with the core technology, giving full play to the advantages of dynamic multi-view fusion, adaptive clustering, and hybrid classifiers, and is expected to achieve good detection effects in actual applications.

Claims

1. A dynamic multi-view fusion voice forgery identification method, characterized in that: The following steps are involved: S1. According to the original speech signal, obtain multi-view dynamic feature data and construct a multi-view dataset; S2. Clustering the data in the multi-view dataset using a dynamic weight clustering ensemble algorithm to generate pseudo labels; including the following steps: S2.1, adjusting the weight of each view in a multi-view dataset based on information entropy; S2.

2. Constructing a hypergraph integrating similarity information of each view; Constructing a hypergraph integrating similarity information of each view, specifically as follows: First, multi-view clustering is performed. The multi-view dataset is , Indicates Views; The data in the multi-view dataset are After clustering, we get a set of cluster divisions ,in Representative cluster members, each cluster member consists of multiple cluster subsets, denoted as ,in To indicate the The total number of cluster subsets in the sub-clustering partition, Indicates The first of the cluster members Cluster subsets; The data points in the multi-view dataset are uniformly numbered and the multi-view similarities are fused as follows: in, Represents the first data points and The similarity between data points; Represents the first The weight of the view; In the clustering process of views, data points will be classified into a cluster subset, and all the features in this cluster subset together constitute the first The data point in Feature set under view ; Indicated in In the view, The feature set to which the data point belongs With The feature set to which the data point belongs Between Similarity; represents the cross-view consistency coefficient; Indicates The feature vector of the data point and The feature vector of the data point The cosine similarity between ; S2.

3. Based on the hypergraph in step S2.2, use spectral clustering to generate meta-clustering groups: first construct the Laplacian matrix of the hypergraph And the Laplace matrix Perform eigendecomposition to obtain eigenvalues; then automatically determine the optimal number of clusters using the eigenvalue gap method , before selecting The eigenvectors corresponding to the smallest eigenvalues ​​form a matrix , the matrix Each row is considered as a new data point and clustered using the K-means clustering algorithm to obtain meta-cluster groups; S2.

4. Calculate the scores of data points in the multi-view dataset to meta-cluster groups and determine pseudo labels as follows: The average Euclidean distance from the data point to all data points in the meta-clustering group is calculated by the distance measurement method to obtain the score of the data point to the meta-clustering group. Then, according to the minimum distance method, the meta-clustering group label with the smallest distance is taken to determine the pseudo label of the data point. S3, integrating the pseudo labels with the multi-view dynamic feature data to obtain a first enhanced data set, and training a K nearest neighbor classifier; S4. Use the trained K nearest neighbor classifier to predict the unseen data points other than the original speech signal in the speech signal to be identified.

2. According to the method of claim 1, the method is characterized in that: In step S1, the multi-view dynamic feature data includes a reconstruction error view, an acoustic feature view, and a quality assessment view; The speech generation model HiFiGAN is used to calculate the reconstruction error view, as follows: in, represents the reconstruction error; represents the original speech signal; Represents the original speech signal At the moment The value of Indicates the total duration of the speech signal; Represents the decoder in the speech generation model HiFiGAN; Represents the encoder in the speech generation model HiFiGAN; is a hyperparameter; Represents the original speech signal At the moment The total variation of The audio encoding model Encodec is used to obtain the acoustic feature view, as follows: The original speech signal is input into the audio coding model Encodec to generate a multi-scale acoustic code. The multi-scale acoustic code is first subjected to a convolutional layer to extract local features, and then a global average pooling layer is used to calculate the mean of each channel and compress it into a one-dimensional vector. Finally, the inner product of the one-dimensional vector itself is calculated to obtain an acoustic feature view. The UTMOS model is used to obtain the quality assessment view, as follows: Input the original speech signal into the UTMOS model, output the overall speech quality assessment score, and normalize the data to obtain the quality assessment view; The structure of the multi-view dataset is: {sequence number of the original speech signal, reconstruction error view, acoustic feature view, quality assessment view}.

3. According to the method of claim 1, the method is characterized in that: In step S2.1, the weight of each view in the multi-view dataset is dynamically adjusted according to the information entropy difference between the views, as follows: in, Represents the first The weight of each view; Indicates The information entropy of each view; represents the exponential function; Indicates The number of data points in each view is the same for all views; Indicates In the view Data points Probability of occurrence.

4. According to the method of claim 1, the method is characterized in that: Step S3 includes the following steps: S3.

1. Concatenate the pseudo-labels and the multi-view dynamic feature data to obtain a first enhanced data set. The structure of the first enhanced data set is: {sequence number of the original speech signal, reconstruction error view, acoustic feature view, quality assessment view, pseudo-label}. Divide the first enhanced data set into a validation set and a training set. S3.2, adjust the K value of the K nearest neighbor classifier according to the density of data points in the training set; S3.3, input the training set into the K nearest neighbor classifier. The K nearest neighbor classifier uses the weighted Euclidean distance to calculate the distance between the data points in the training set according to the K value in step S3.2 to find the K nearest neighbors of each data point. The data point category is determined according to the pseudo-label by voting method to obtain the K nearest neighbor classifier that has been preliminarily trained. The performance of the K nearest neighbor classifier that has been preliminarily trained is evaluated using the validation set. If the evaluation result meets the requirements or reaches the set maximum number of iterations, execute step S3.

4. Otherwise, the parameters in the data point density calculation are set to zero. The set step size will be increased or decreased, the number of iterations will be increased by 1, and the process will return to step S3.2; S3.

4. Use the Bayesian model to calibrate the output confidence of the initially trained K nearest neighbor classifier to obtain a trained K nearest neighbor classifier.

5. According to the method of claim 4, the method is characterized in that: In step S3.2, the K value is adaptively adjusted according to the density of the data points in the training set, as follows: in, Indicates that for the first Data points , K value after adaptive adjustment; Indicates Data points The number of local data points, that is, The range of the setting is centered The number of remaining data points in ; Indicates rounding up; represents the total number of data points in the training set; represents the indicator function, when When established, The value of is 1, otherwise it is 0; Indicates Data points and Data points The distance between.

6. The method for identifying voice forgery based on dynamic multi-view fusion according to claim 4, characterized in that: In step S3.4, the Bayesian probability model is used to calibrate the confidence of the output of the K nearest neighbor classifier that has been initially trained, as follows: in, Represents the feature vector at a given data point Under the condition that The probability of is a category of data points; represents the indicator function, when When established, The value of is 1, otherwise it is 0; The feature vector representing the data point input into the K nearest neighbor classifier; Represents the number of nearest neighbors in the K nearest neighbor classifier; Indicates The weight of the nearest neighbor; Indicates The category of the nearest neighbors; Express The category in the nearest neighbor Sum the weights of the nearest neighbors; is a hyperparameter; Represents the feature vector Information entropy.

7. The method for identifying voice forgery based on dynamic multi-view fusion according to claim 1, characterized in that: In step S4, steps S1 to S3 are performed on the speech signal to be identified to obtain a second enhanced data set corresponding to the speech signal to be identified; The first augmented dataset and the second augmented dataset have the same structure; The second enhanced data set is input into the trained K nearest neighbor classifier to predict the unseen data points other than the original speech signal in the speech signal to be identified.

8. A device for implementing the method for identifying voice forgery by dynamic multi-view fusion as described in any one of claims 1 to 7, characterized in that: Includes the following modules: Data acquisition and preprocessing module: used to obtain the original voice signal and the voice signal that needs to be identified; Multi-view feature extraction module: used to obtain multi-view dynamic feature data based on the original speech signal and construct a multi-view dataset; Integrated clustering module: used to cluster the data in the multi-view dataset using a dynamic weight clustering integrated algorithm to generate pseudo labels; Classifier training module: used to integrate pseudo labels with multi-view dynamic feature data to obtain a first enhanced data set and train a K-nearest neighbor classifier; Prediction module: used to use the trained K nearest neighbor classifier to predict the unseen data points other than the original speech signal in the speech signal to be identified.

Citation Information

Patent Citations

  • Counterfeit voice detection method and system based on time sequence multi-scale feature representation learning

    CN117809694A

  • Voice forgery detection method based on multi-scale GMM-ResNet model

    CN119170052A

  • Tracing method and device for robust counterfeit speech algorithm

    CN116959425A

  • Multi-view data label-free clustering method based on trusted neighbor information aggregation

    CN117909778A