Method for constructing source tracing model of original speaker of forged voice
By constructing a traceability model for fake voice original speakers, using multi-scale voiceprint feature extraction and residual correction networks, the voiceprint features of the original speakers are reversely restored, and the problem of the inability to trace the origin of the speech original speakers in the existing technology is solved, and high-precision identity recognition of the original speakers is achieved.
Patent Information
- Application Number
- CN202510436216.0
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-04-08
- Publication Date
- 2025-07-04
AI Technical Summary
The existing forged voice detection technology cannot trace the voiceprint information of the original speaker in the voice in the voice, and cannot identify the identity of the original speaker, resulting in the inability to effectively supervise the abuse of the forged voice technology.
A tracing model for fake voice original speakers was constructed, and the voiceprint characteristics of the original speakers were reversely restored through multi-scale vocalprint feature extraction, Transformer-CLAP layered purification module and 3-level RCB residual correction module, and a three-stage training strategy was used to train the model.
Effectively distinguish the residual original speaker characteristics in forged voice from the target speaker characteristics generated by forged voice, improve the traceability accuracy of original speakers, and provide efficient and reliable traceability solutions for fake voice original speakers.
Smart Images

Figure CN120260579A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical fields of artificial intelligence and computer security, and more particularly to a method for constructing a forgery speech original speaker tracing model. Background Art
[0002] With the rapid development of voice deep forgery technology, "voice cloning" can imitate the voice of a specific person through an artificial intelligence model, and only need a few seconds of voice samples of the original speaker to generate a high-quality forged voice of the target speaker. The use threshold of forgery voice-related tools is gradually decreasing, which will expand the scope of abuse of forgery voice technology and bring multiple social security risks such as voice fraud, false news, and espionage.
[0003] The existing forgery voice detection technology still focuses on the binary classification forgery determination problem, that is, to identify whether a piece of voice is a real sample or a forged one. However, the above forgery voice detection scheme cannot trace the voiceprint information of the original speaker in the forged voice, so as to further identify the identity of the original speaker. Therefore, it is impossible to form sufficient supervision and deterrence against the original speakers who abuse the forgery voice technology and cause social harm. In addition, the current forgery voice tracing mainly faces three major technical bottlenecks: 1. In the process of deep voice forgery, the voice generator will introduce voiceprint perturbations of the target speaker (such as fundamental frequency perturbation, formant shift), and this perturbation will occupy the dominant part of the forged voiceprint, resulting in mis-matching of traditional voiceprint recognition systems; 2. The forged voice generated through adversarial training will deliberately erase the voiceprint information of the original speaker, resulting in very sparse and difficult-to-capture voiceprint features of the original speaker remaining in the forged voice; 3. The voiceprint of the target speaker and the voiceprint of the original speaker in the forged voice are highly coupled, and the existing technology lacks the ability of disentanglement and reverse recovery when focusing on voiceprint features, and cannot reconstruct the voiceprint information of the original speaker from the tampered acoustic features. Therefore, how to establish a forgery voice original speaker tracing model with reverse mapping ability is an urgent problem to be solved by those skilled in the art. Summary of the Invention
[0004] In view of this, the present invention provides a method for constructing a forgery voice original speaker tracing model, which reversely recovers the voiceprint of the original speaker in the forged voice, so as to trace the identity information of the original speaker and form a deterrence against the abuse of voice forgery technology.
[0005] In order to achieve the above object, the present invention provides the following technical solutions:
[0006] A method for constructing a forgery voice original speaker tracing model, comprising the following steps:
[0007] S1, multi-scale voiceprint feature extraction: converting the forged audio sample X A→B into a coarse-grained voiceprint feature sample E A→B, construct the voiceprint feature data pair (E A , E B , E A→B );
[0008] S2. Target voiceprint hierarchical purification: Construct a Transformer-CLAP hierarchical purification module, introduce the multi-head self-attention mechanism to model long-range dependencies, and focus on the key voiceprint feature dimensions; adopt a hierarchical fusion mechanism to fuse the outputs of different attention heads at multiple scales; combine contrast learning to enhance the voiceprint feature decoupling ability to obtain the purified target speaker's voiceprint feature E'. B ;
[0009] S3. Construct a residual correction network: Construct a 3-level RCB residual correction module, model the feature differences through the residual learning mechanism, reverse-derive and restore the original speaker identity label to obtain the restored original speaker's voiceprint feature E'. A ;
[0010] S4. Train the spoofed speech original speaker tracing model: Based on the Transformer-CLAP hierarchical purification module and the 3-level RCB residual correction module, adopt a three-stage training strategy to train the spoofed speech original speaker tracing model.
[0011] Optionally, S1 is specifically;
[0012] S11. Construct a language dataset: For each original speaker A, target speaker B, and their original audio samples X A and X B , use different speech spoofing methods F to generate spoofed audio samples X A→B = F(X A , X B ), to obtain the audio sample data pair (X A , X B , X A→B );
[0013] S12. Frequency domain feature preprocessing: Use the Mel filter bank to convert the audio sample data pair (X A , X B , X A→B ) into a feature spectrum with dimensions of f×T, where f is the frequency domain feature dimension and T is the number of time frames;
[0014] S13. Voiceprint encoding and feature extraction: Use the pre-trained ECAPA-TDNN architecture as the core structure of the voiceprint encoder, input the feature spectrum, capture the frequency domain and temporal features of the speech, and obtain the voiceprint feature data pair (E A , E B , E A→B ), where E A is the original speaker's voiceprint feature, EB is the voiceprint feature of the target speaker, E A→B is the voiceprint feature after forgery for imitating the target speaker B.
[0015] Optionally, S2 is specifically:
[0016] S21, Multi-Head Self-Attention Feature Focusing: Generate a multi-head self-attention matrix through a dynamic masking mechanism, and combine orthogonal loss to constrain the independent subspaces of different attention heads to focus on the voiceprint features, and output the multi-head attention output matrix H';
[0017] S22, Multi-Head Self-Attention Hierarchical Fusion: Use the FPN structure to hierarchically extract local details, context associations, and global statistical features, and cross-layer fuse multi-scale features through a gated attention mechanism to enhance the representational ability of the forged voiceprint features, and obtain the forged voiceprint feature E' A→B ;
[0018] S23, Contrastive Learning Feature Decoupling: Based on the contrastive learning framework, optimize the feature decoupling boundary by constructing positive and negative sample pairs and a loss function with an adaptive temperature coefficient, and separate the coupling of the target speaker's voiceprint features and forged features.
[0019] Optionally, S21 is specifically:
[0020] Calculate the multi-head self-attention matrix: Project the forged voiceprint feature E A→B into the multi-head self-attention module and project it into the query Q and key K spaces:
[0021] Q = E A→B ·W q
[0022] K = E A→B ·W k
[0023] Generate the cosine similarity matrix cos sim , and calculate after normalizing with the L2 norm:
[0024] cos sim = (Q·K T ) / (||Q||2 * ||K||2 + ε)
[0025] Use the dynamic masking mechanism to binarize the cosine similarity matrix cos sim to obtain the output matrix H of a single attention head:
[0026]
[0027] Cross-Head Feature Orthogonal Constraint: For the output matrix H of multiple attention heads, construct an orthogonal loss function L orth, force different attention heads to focus on independent subspaces of voiceprint features:
[0028]
[0029] Obtain the multi-head attention output matrix H' after feature focusing;
[0030] In the formula, W q , W k represent learnable parameter matrices in the query space and the key space respectively, T represents the matrix transpose operation, ε is a positive floating-point number used to prevent division by zero errors, θ is a learnable threshold parameter, and ||·|| F represents the Frobenius norm, which is used to measure the size of the matrix.
[0031] Optionally, S22 is specifically:
[0032] Horizontal connection: For the multi-head attention output matrix H' after feature focusing, a 3-level FPN structure is adopted; among them, the underlying structure E low extracts local detail features through one-dimensional convolution; the middle-level structure E mid captures context-related features through dilated convolution; the high-level structure E high obtains statistical characteristics through global average pooling;
[0033] Cross-layer aggregation: Adopt a gated attention mechanism to obtain the gated signal g:
[0034]
[0035] In the formula, represents concatenating the feature matrices along the channel dimension, and W g is a learnable weight matrix used to map the concatenated features to a gated signal, and σ(·) represents the Sigmoid activation function;
[0036] Use the gated signal g to fuse the feature outputs of different attention heads at multiple scales, and finally obtain the forged voiceprint features E' after feature focusing and feature fusion A→B ,
[0037] E' A→B = g⊙E low +(1 - g)⊙E mid + g⊙E high
[0038] In the formula, ⊙ represents element-wise multiplication.
[0039] Optionally, S23 is specifically:
[0040] Construct sample pairs: For the voiceprint feature data pair (E A , EB , E A→B ), and its characteristic variant E' A→B , respectively construct 1 group of positive sample pairs and n groups of negative sample pairs Among them, E k (k ∈ [1, n]) is a randomly sampled voiceprint feature vector with the top 30% similarity to E B in the voice data set of S1;
[0041] Enhanced feature separability: Adopt a temperature coefficient adaptive mechanism to construct a contrast loss function L CLAp , automatically learn the feature decoupling boundary, and optimize to obtain the target speaker voiceprint feature E' B enhanced by feature decoupling,
[0042]
[0043] In the formula, τ is a learnable temperature parameter; sim(·) represents the cosine similarity function, and exp(·) represents the exponential function.
[0044] Optionally, S3 is specifically:
[0045] S31. Feature difference calculation: Perform feature difference operation on the forged voiceprint feature E A→B and the purified target speaker voiceprint feature E' B to obtain the preliminary residual feature ΔE0 = E A→B - E' B , providing input for subsequent residual modeling;
[0046] S32. Construction of residual feature modeling network: Construct a 3-level RCB residual correction module, combine fully connected mapping, GeLU non-linear activation, lightweight multi-head attention and layer normalization module, and gradually model the non-linear difference of voiceprint features and stabilize the training;
[0047] S33. Residual prediction and correction: Refine the residual feature through 3-level RCB iteration, use the predicted residual ΔE B to inversely correct the forged feature, and combine the learnable parameter α to restore the original speaker voiceprint feature E' A ;
[0048] S34. Bimodal joint optimization: Combine the feature reconstruction loss and the identity cross-entropy loss to achieve bimodal optimization of voiceprint restoration.
[0049] 8. A method for constructing a forged speech original speaker traceability model according to claim 7, wherein S33 is specifically:
[0050] Residual Iterative Prediction: Input the preliminary residual feature ΔE0 into the 3-level RCB residual correction module, and obtain the accurate residual feature ΔE through gradual refinement. B ,
[0051] ΔE1 = RCB1(ΔE0, E′ B )
[0052] ΔE2 = RCB2(ΔE1, E′ B )
[0053] ΔE B = RCB3(ΔE2, E′ B )
[0054] where RCB i (i ∈ [1, 3]) corresponds to the 3-level RCB residual correction module;
[0055] Original Speaker Voiceprint Feature Restoration: Use the accurate residual feature ΔE B to perform reverse correction on the forged voiceprint feature E A→B to obtain the restored original speaker voiceprint feature E' A = E A→B - αΔE B ; where α is a learnable residual scaling factor that is automatically optimized through backpropagation.
[0056] Optionally, S34 is specifically:
[0057] Construct Feature Reconstruction Loss: Calculate the mean square error between the restored original speaker voiceprint feature E′ A and the original voiceprint feature E A to obtain the reconstruction loss L recon = (||E' A - E A ||2) 2 , which constrains the consistency of the feature space;
[0058] Construct Identity Consistency Loss: Adopt a pre-trained original speaker classifier to map the voiceprint feature vector to a probability distribution vector p = Classifer(E' A ), and then calculate the cross-entropy loss to ensure that the restored feature can be accurately traced back to the original identity ID A ; where y i is the true one-hot label of the original speaker A i , and p i is the probability distribution vector predicted by the original speaker classifier for E' A .
[0059] Optionally, S4 is specifically:
[0060] S41. Pre - training the Transformer - CLAP hierarchical purification module: Freeze the parameters of the speaker encoder in S1, and use the pair of speaker feature data (E A , E B , E A→B ) to train the Transformer - CLAP hierarchical purification module alone, so that the E' purified by the Transformer - CLAP hierarchical purification module B is close to the true target speaker's voiceprint feature E B ;
[0061] S42. Pre - training the 3 - level RCB residual correction module: Freeze the parameters of the speaker encoder in S1 and the Transformer - CLAP hierarchical purification module in S2, and use the data pair (E A , E′ B , E A→B ) to train the 3 - level RCB residual correction module alone, so that the E′ restored by the 3 - level RCB residual correction module A is close to the true original speaker's voiceprint feature E A ;
[0062] S43. Joint fine - tuning: Unfreeze all the parameters of the Transformer - CLAP hierarchical purification module and the 3 - level RCB residual correction module, and conduct end - to - end overall training on the forged speech original speaker traceability model, using cosine annealing learning rate scheduling.
[0063] As can be seen from the above - mentioned technical solutions, compared with the prior art, the present invention provides a method for constructing a forged speech original speaker traceability model, having the following beneficial effects:
[0064] 1. The multi - scale speaker feature hierarchical purification technology proposed by the present invention can break through the problem of mixed forged features. Compared with traditional methods that rely on single - scale speaker features (such as MFCC or shallow networks), the present invention can effectively distinguish the remaining original speaker features in forged speech from the target speaker features generated by forgery, significantly reducing misjudgments caused by feature confusion;
[0065] 2. The residual correction network proposed by the present invention for reversely restoring the original speaker's voiceprint can solve the problem of non - linear distortion modeling. Compared with traditional residual networks that cannot model non - linear spectral distortions introduced by forgery tools (such as the frequency domain distortion of voice cloning models), the present invention can model non - linear feature differences step by step, effectively improving the accuracy of original speaker traceability;
[0066] 3. The present invention uses a three - stage training strategy and a lightweight architecture, which can achieve efficient deployment and strong robustness, providing an efficient and reliable solution for voiceprint forgery traceability. Description of the Drawings
[0067] To more clearly illustrate the technical solutions in the embodiments of the present invention or the prior art, the following will briefly introduce the accompanying drawings required in the description of the embodiments or the prior art. Obviously, the accompanying drawings in the following description are only embodiments of the present invention. For those of ordinary skill in the art, without creative efforts, other accompanying drawings can also be obtained based on the provided drawings.
[0068] Figure 1 It is a flowchart of the method for constructing a forged voice original speaker tracing model of the present invention;
[0069] Figure 2 It is a schematic diagram of the Transformer-CLAP hierarchical purification module of the present invention;
[0070] Figure 3 It is a schematic diagram of the 3-level RCB residual correction module of the present invention. Specific embodiments
[0071] The following will clearly and completely describe the technical solutions in the embodiments of the present invention with reference to the accompanying drawings in the embodiments of the present invention. Obviously, the described embodiments are only a part of the embodiments of the present invention, rather than all of the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those of ordinary skill in the art without creative efforts belong to the scope of protection of the present invention.
[0072] Based on artificial intelligence voice forgery technology, it is possible to make the voice of the original speaker A sound like the target speaker B while keeping the speech content unchanged. In other words, the biometric feature obtained by the voiceprint recognition system for the forged voice will change from E A (the identity identifier pointing to the original speaker A) to E A→B (similar to E B , the identity identifier pointing to the target speaker B). Therefore, based on the above technology, the original speaker A can forge the voice of the target speaker B and carry out harmful acts that damage personal or social interests, such as voice fraud. However, the existing forged voice detection technology mainly focuses on binary classification of forgery determination, rather than tracing the identity of the forger, and it is difficult to form a strong deterrent to those who abuse voice forgery technology.
[0073] The embodiments of the present invention disclose a method for constructing a forged voice original speaker tracing model, as Figure 1 shown, including the following steps:
[0074] S1. Multi-scale voiceprint feature extraction: Convert the forged audio sample X A→B into a coarse-grained voiceprint feature sample E A→B , and construct a voiceprint feature data pair (E A , EB , E A→B );
[0075] S2. Target voiceprint hierarchical purification: Construct a Transformer-CLAP hierarchical purification module, introduce a multi-head self-attention mechanism to model long-range dependencies, and focus on key voiceprint feature dimensions; adopt a hierarchical fusion mechanism to fuse the outputs of different attention heads at multiple scales; combine contrastive learning to enhance the voiceprint feature decoupling ability to obtain the purified target speaker voiceprint feature E'. B ; The Transformer-CLAP hierarchical purification module is as Figure 2 shown;
[0076] S3. Construct a residual correction network: Construct a 3-level RCB residual correction module, model feature differences through a residual learning mechanism, reverse-derive and restore the original speaker identity identifier to obtain the restored original speaker voiceprint feature E'. A ; The 3-level RCB residual correction module is as Figure 3 shown;
[0077] S4. Train the spoofed speech original speaker traceability model: Based on the Transformer-CLAP hierarchical purification module and the 3-level RCB residual correction module, adopt a three-stage training strategy to train the spoofed speech original speaker traceability model.
[0078] Furthermore, S1 is specifically;
[0079] S11. Construct a language dataset: For each original speaker A, target speaker B, and their original audio samples X A and X B , use different speech spoofing methods F to generate spoofed audio samples X A→B = F(X A , X B ), to obtain audio sample data pairs (X A , X B , X A→B );
[0080] In an embodiment of the present invention, for a original speakers A = {A1, A2,..., A a} and b target speakers B = {B1, B2,..., B b}, each original speaker or target speaker respectively has c original audio samples X A = {x A1 , x A2 ,..., x Ac}, X B = {x B1 , x B2 ,..., x Bc}, using d different voice forgery methods F = {f1, f2, …, f d}, respectively generate forged audio samples X A→B = F(X A , X B ), and a total of a × b × c × d groups of (X A , X B , X A→B ) audio sample data pairs are obtained;
[0081] S12. Frequency domain feature preprocessing: Use a Mel filter bank to convert the audio sample data pair (X A , X B , X A→B ) into a feature spectrogram with dimensions f × T, where f is the frequency domain feature dimension and T is the number of time frames. In the embodiment of the present invention, the frequency domain feature dimension f is set to 80, and the number of time frames T is set to 256;
[0082] S13. Speaker verification coding and feature extraction: Use a pre-trained ECAPA-TDNN architecture as the core structure of the speaker verification encoder. The ECAPA-TDNN architecture includes 3 layers of aggregated SE-Res2Net modules (each layer includes a Res2Net multi-scale feature extraction network and an SE channel attention network) and a dynamic self-attention statistical pooling module. Input the feature spectrogram to capture the frequency domain and temporal features of the speech, and obtain the speaker verification feature data pair (E A , E B , E A→B ), where E A is the original speaker's voiceprint feature, E B is the target speaker's voiceprint feature, and E A→B is the forged voiceprint feature used to imitate the target speaker B. In the embodiment of the present invention, the voiceprint feature is an embedding vector E = (e1, e2,...., e 128 ) with a dimension of 256.
[0083] Furthermore, S2 is specifically as follows;
[0084] S21. Multi-head self-attention feature focusing; Generate a multi-head self-attention matrix through a dynamic masking mechanism, and combine the orthogonal loss to constrain the independent subspaces that different attention heads focus on the voiceprint features, and output the multi-head attention output matrix H';
[0085] S22. Multi-head self-attention hierarchical fusion; Use the FPN structure to hierarchically extract local details, context associations, and global statistical features, and cross-layer fuse multi-scale features through a gated attention mechanism to enhance the representation ability of the forged voiceprint features, and obtain the forged voiceprint feature E' A→B ;
[0086] S23. Contrastive learning feature decoupling: Based on the contrastive learning framework, by constructing positive and negative sample pairs and a loss function with an adaptive temperature coefficient, optimize the feature decoupling boundary to separate the coupling between the target speaker's voiceprint features and forged features.
[0087] Further, S21 is specifically as follows:
[0088] Calculate the multi-head self-attention matrix: Project the forged voiceprint feature E A→B into the multi-head self-attention module and project it into the query Q (Query) and key K (Key) spaces:
[0089] Q = E A→B ·W q
[0090] K = E A→B ·W k
[0091] Generate the cosine similarity matrix cos sim , and calculate it after L2 norm normalization:
[0092] cos sim =(Q·K T ) / (||Q||2 * ||K||2 + ε)
[0093] Adopt a dynamic masking mechanism to binarize the cosine similarity matrix cos sim to obtain the output matrix H of a single attention head:
[0094]
[0095] Cross-head feature orthogonal constraint: For the output matrix H of multiple attention heads, construct an orthogonal loss function L orth to force different attention heads to focus on independent subspaces of voiceprint features:
[0096]
[0097] Obtain the output matrix H' of the multi-head attention after feature focusing;
[0098] In the formula, W q , W k respectively represent learnable parameter matrices in the query space and the key space, T represents the matrix transpose operation, ε is a positive floating-point number used to prevent division by zero errors, θ is a learnable threshold parameter, and ||·|| FDenoted as the Frobenius norm, it is used to measure the size of a matrix. In the embodiments of the present invention, the number of attention mechanism heads is set to 8. For the output matrices {H1, H2, …, H8} of the 8 attention heads, the focused multi-head attention output matrix is {H'1, H'2, …, H'8}, and the initial value of θ is set to 0.5.
[0099] Further, S22 is specifically:
[0100] Horizontal connection: For the multi-head attention output matrix H' after feature focusing, a 3-level FPN structure is adopted; among them, the underlying structure E low extracts local detail features through one-dimensional convolution; the middle structure E mid captures context correlation features through dilated convolution; the upper structure E high obtains statistical characteristics through global average pooling; in the embodiments of the present invention, the underlying structure E low ={H'1, H'2}, and the convolution kernel size is set to 5; the middle structure E mid ={H′3, H′4, H′5}, and the dilation rate is set to 2. The upper structure E high ={H′6, H′7, H′8} obtains statistical characteristics through global average pooling;
[0101] Cross-layer aggregation: Adopt a gated attention mechanism to obtain a gating signal g:
[0102]
[0103] In the formula, represents concatenating the feature matrices in the channel dimension, and W g is a learnable weight matrix used to map the concatenated features to a gating signal, and σ(·) represents the Sigmoid activation function;
[0104] Use the gating signal g to fuse the feature outputs of different attention heads in multiple scales, and finally obtain the forged voiceprint feature E' after feature focusing and feature fusion g→B ,
[0105] E' A→B = g⊙E low +(1 - g)⊙E mid + g⊙E high
[0106] In the formula, ⊙ represents element-wise multiplication.
[0107] Further, S23 is specifically:
[0108] Construct sample pairs: For the voiceprint feature data pairs (E A , E B , E A→B) and its characteristic variant E′ A→B , construct 1 group of positive sample pairs respectively and n groups of negative sample pairs wherein, E k (k∈[1,n]) is a random voiceprint feature vector randomly sampled from the voice dataset of S1 and having a similarity of the top 30% with E B . In the embodiment of the present invention, the feature similarity is calculated based on MFCC features;
[0109] Enhanced feature separability: Adopt a temperature coefficient adaptive mechanism to construct a contrastive loss function L CLAP , automatically learn the feature decoupling boundary, and optimize to obtain the target speaker voiceprint feature E′ after feature decoupling enhancement B ,
[0110]
[0111] In the formula, τ is a learnable temperature parameter; sim(·) represents the cosine similarity function, and exp(·) represents the exponential function; in the embodiment of the present invention, the initial value of τ is set to 0.07, and the purified target speaker voiceprint feature is an embedding vector with a dimension of 256.
[0112] Furthermore, S3 is specifically:
[0113] S31. Feature difference calculation: Perform feature difference operation on the forged voiceprint feature E A→B and the purified target speaker voiceprint feature E′ B to obtain a preliminary residual feature ΔE0 = E A→B - E′ B , providing input for subsequent residual modeling;
[0114] S32. Construction of residual feature modeling network: Construct a 3-level RCB residual correction module, combine fully connected mapping, GeLU non-linear activation, lightweight multi-head attention and layer normalization module, and model the non-linear difference of voiceprint features step by step and train stably;
[0115] In the embodiment of the present invention, S32 is specifically:
[0116] Construct a fully connected layer: The fully connected layer is used to map the residual feature ΔE i and the target speaker voiceprint feature E′ B to the hidden layer space, and their input and output dimensions are set to [256→512], [512→256], [256→128] respectively;
[0117] Construct a GeLU non-linear activation function: Adopt the GeLU activation function to introduce non-linear transformation ability for modeling the non-linear difference in voiceprint features;
[0118] Construct a lightweight multi-head attention module: Adopt a lightweight attention mechanism and focus on key residual dimensions:
[0119] Q′ = ΔE i ·W′ q
[0120] K′ = E′ B ·W′ k
[0121]
[0122] Wherein, W′ q 、W′ k are learnable parameter matrices, d k is a scaling factor, V′ is the projection of the value vector, and MHFA is the output of this multi-head feature attention module;
[0123] Construct a layer normalization module: Connect the layer normalization (LayerNorm) module to stabilize the training process.
[0124] S33. Residual prediction and correction: Refine the residual features through 3-level RCB iteration, and use the predicted residual ΔE B to inversely correct the forged features, and combine the learnable parameter α to restore the speaker's voiceprint feature E′ A ;
[0125] S34. Bimodal joint optimization: Combine the feature reconstruction loss and the identity cross-entropy loss to achieve bimodal optimization of voiceprint restoration.
[0126] Furthermore, S33 is specifically:
[0127] Residual iterative prediction: Input the preliminary residual feature ΔE0 into the 3-level RCB residual correction module, and obtain the accurate residual feature ΔE B ,
[0128] ΔE1 = RCB1(ΔE0, E′ B )
[0129] ΔE2 = RCB2(ΔE1, E′ B )
[0130] ΔE B = RCB3(ΔE2, E′ B )
[0131] In the formula, RCB i (i ∈ [1, 3]) corresponds to the 3-level RCB residual correction module;
[0132] Original speaker voiceprint feature restoration: Use the precise residual feature ΔE B Perform reverse correction on the forged voiceprint feature E A→B To obtain the restored original speaker voiceprint feature E' A = E A→B - αΔE B ; where α is a learnable residual scaling factor, automatically optimized through backpropagation, and set to 1.0 as the initial value in the embodiments of the present invention.
[0133] Furthermore, S34 is specifically as follows:
[0134] Construct the feature reconstruction loss: Calculate the mean square error between the restored original speaker voiceprint feature E' A And the original voiceprint feature E A To obtain the reconstruction loss L recon = (||E' A - E A ||2) 2 , to constrain the consistency of the feature space;
[0135] Construct the identity consistency loss: Adopt a pre-trained original speaker classifier. In the embodiments of the present invention, the configuration of the original speaker classifier is 2 fully connected layers and a Softmax layer, used to map the voiceprint feature vector to a probability distribution vector p = Classifer(E' A ), and then calculate the cross-entropy loss To ensure that the restored feature can be accurately traced back to the original identity ID A ; where y i Is the true one-hot label of the original speaker A i , p i Is the probability distribution vector predicted by the original speaker classifier for E' A .
[0136] Furthermore, S4 is specifically as follows:
[0137] S41. Pre-train the Transformer-CLAP hierarchical purification module: Freeze the parameters of the voiceprint encoder in S1, and use the voiceprint feature data pairs (E A , E B , E A→B ) to train the Transformer-CLAP hierarchical purification module separately, so that the E' B Purified by the Transformer-CLAP hierarchical purification module is close to the true target speaker voiceprint feature E B ; In the embodiments of the present invention, the AdamW optimizer is adopted, and the initial learning rate is set to 3e-4;
[0138] S42, pre-trained 3-level RCB residual correction module: freeze the parameters of the voiceprint encoder in S1 and the Transformer-CLAP layered purification module in S2, and use the data pair (E A ,E′ B ,E A→B ) Train the 3-level RCB residual correction module separately, so that the 3-level RCB residual correction module can recover the obtained E' A Close to the real speaker's voiceprint feature E A ; In the embodiment of the present invention, the AdamW optimizer is used, and the initial learning rate is set to 2e-4;
[0139] S43, joint fine-tuning: unfreeze all parameters of the Transformer-CLAP hierarchical purification module and the 3-level RCB residual correction module, perform end-to-end overall training on the forged speech original speaker tracing model, and use cosine annealing learning rate scheduling. In this embodiment of the present invention, the initial learning rate is set to 1e-4 and the minimum learning rate is set to 1e-6.
[0140] The various embodiments in this specification are described in a progressive manner, and each embodiment focuses on the differences from other embodiments. The same or similar parts between the various embodiments can be referenced to each other.
[0141] The above description of the disclosed embodiments enables one skilled in the art to implement or use the present invention. Various modifications to these embodiments will be apparent to one skilled in the art, and the general principles defined herein may be implemented in other embodiments without departing from the spirit or scope of the present invention. Therefore, the present invention will not be limited to the embodiments shown herein, but rather to the widest scope consistent with the principles and novel features disclosed herein.
Claims
1. A method for constructing a tracing model of the original speaker of forged speech, characterized in that, It includes the following steps: S1. Multi-scale voiceprint feature extraction: Convert the forged audio sample X through a voiceprint encoder A→B into a voiceprint feature sample E with coarse granularity A→B , and construct a voiceprint feature data pair (E A , E B , E A→B ); S2. Target voiceprint hierarchical purification: Construct a Transformer-CLAP hierarchical purification module, introduce a multi-head self-attention mechanism to model long-range dependencies, and focus on key voiceprint feature dimensions; Adopt a hierarchical fusion mechanism to fuse the outputs of different attention heads at multiple scales; combine contrastive learning to enhance the ability to decouple voiceprint features, and obtain the purified target speaker voiceprint feature E' B ; S3. Construct a residual correction network: Construct a 3-level RCB residual correction module, model the feature differences through the residual learning mechanism, reverse-derive and restore the identity identifier of the original speaker, and obtain the restored voiceprint feature E' of the original speaker A ; S4. Train the forger voice originator tracing model: Based on the Transformer-CLAP hierarchical purification module and the 3-level RCB residual correction module, adopt a three-stage training strategy to train the forger voice originator tracing model.
2. The method for constructing a tracing model of the original speaker of forged speech according to claim 1, wherein S1 is specifically as follows; S11. Construct a language dataset: For each original speaker A, target speaker B, and their original audio sample X A and X B , use different voice forgery methods F to generate forged audio samples X A→B = F(X A , X B ), obtaining the audio sample data pairs (X A , X B , X A→B ); S12. Frequency-domain feature preprocessing: Use a Mel filter bank to convert the audio sample data pairs (X A , X B , X A→B ) into a feature spectrum with dimensions f×T, where f is the frequency-domain feature dimension and T is the number of time frames; S13. Voiceprint Encoding and Feature Extraction: Use the pre-trained ECAPA-TDNN architecture as the core structure of the voiceprint encoder. Input the feature spectrum to capture the frequency-domain and temporal features of the speech, and obtain the voiceprint feature data pairs (E A , E B , E A→B ), where E A is the voiceprint feature of the original speaker, E B is the voiceprint feature of the target speaker, and E A→B is the forged voiceprint feature used to imitate the target speaker B.
3. A method for constructing a tracing model of the original speaker of forged speech according to claim 1, characterized in that, S2 is specifically as follows; S21. Multi-head self-attention feature focusing: Generate a multi-head self-attention matrix through a dynamic masking mechanism, and combine the orthogonal loss to constrain the independent subspaces of different attention heads to focus on voiceprint features, and output the multi-head attention output matrix H'; S22. Multi-head self-attention hierarchical fusion: Use the FPN structure to hierarchically extract local details, context associations, and global statistical features, and cross-layer fuse multi-scale features through the gated attention mechanism to enhance the representation ability of forged voiceprint features, obtaining the forged voiceprint feature E'. A→B ; S23. Contrastive learning feature decoupling; Based on the contrastive learning framework, by constructing positive and negative sample pairs and a loss function with an adaptive temperature coefficient, optimize the feature decoupling boundary and separate the coupling between the target speaker's voiceprint features and forged features.
4. A method for constructing a tracing model of the original speaker of forged speech according to claim 3, characterized in that, S21 is specifically as follows: Calculate the multi-head self-attention matrix: project the forged voiceprint feature E A→B into the multi-head self-attention module and project it into the query Q and key K spaces: Q = E A→B ·W q K = E A→B ·W k Generate the cosine similarity matrix cos sim , and calculate it after normalizing with the L2 norm: cos sim = (Q·K T ) / (‖Q‖2 * ‖K‖2 + ε) Adopt a dynamic masking mechanism to binarize the cosine similarity matrix cos sim to obtain the output matrix H of a single attention head: Orthogonal Constraint on Cross-Head Feature: For the output matrix H of multiple attention heads, construct an orthogonal loss function L orth , forcing different attention heads to focus on independent subspaces of voiceprint features: Obtain the multi-head attention output matrix H' after feature focusing; where, W q , W k respectively represent learnable parameter matrices in the query space and the key space, T represents the matrix transpose operation, ε is a positive floating-point number used to prevent division-by-zero errors, θ is a learnable threshold parameter, ||·|| F represents the Frobenius norm, which is used to measure the size of a matrix.
5. A method for constructing a source tracing model of a forged voice original speaker, as claimed in claim 3, wherein S22 is specifically as follows: Horizontal connection: For the multi-head attention output matrix H' after feature focusing, a 3-level FPN structure is adopted; among them, the underlying structure E low Extracts local detail features through one-dimensional convolution; the middle-level structure E mid Captures context correlation features through dilated convolution; the high-level structure E high Obtains statistical characteristics through global average pooling; Cross-layer aggregation: Adopt a gated attention mechanism to obtain the gated signal g: g = σ(W g [E low ⊕ E mid ⊕ E high ) wherein, ⊕ represents concatenating the feature matrices in the channel dimension, and W g is a learnable weight matrix for mapping the concatenated features to a gating signal, and σ(·) represents the Sigmoid activation function; The gating signal g is used to fuse the feature outputs of different attention heads at multiple scales, and finally the forged voiceprint feature E' after feature focusing and feature fusion is obtained. A→B , E′ A→B = g ⊙ E low + (1 - g) ⊙ E mid + g ⊙ E high In the formula, ⊙ represents element-wise multiplication.
6. A method for constructing a tracing model of the original speaker of forged speech, according to claim 3, characterized in that S23 is specifically as follows: Construct sample pairs: For the voiceprint feature data pairs (E A , E B , E A→B ) and their feature variants E′ A→B , construct 1 set of positive sample pairs and n sets of negative sample pairs respectively. Among them, E k (k ∈ [1, n]) is a random voiceprint feature vector randomly sampled from the voice dataset of S1 and having a similarity of the top 30% with E B ; Enhanced feature separability: By adopting a temperature coefficient adaptive mechanism, a contrastive loss function L is constructed CLAP to automatically learn the feature decoupling boundary and optimize to obtain the target speaker's voiceprint feature E' enhanced by feature decoupling B , In the formula, τ is a learnable temperature parameter; sim(·) represents the cosine similarity function, and exp(·) represents the exponential function.
7. A method for constructing a tracing model of the original speaker of forged speech, according to claim 1, characterized in that S3 is specifically as follows: S31. Feature difference calculation: Take the forged voiceprint feature E A→B and the purified target speaker voiceprint feature E' B to perform feature difference operation, and obtain the preliminary residual feature ΔE0 = E A→B - E' B , which provides input for subsequent residual modeling; S32. Residual feature modeling network construction: Construct a 3-level RCB residual correction module, combine fully connected mapping, GeLU non-linear activation, lightweight multi-head attention and layer normalization modules, and gradually model the non-linear differences of voiceprint features and stabilize the training; S33. Residual prediction and correction: The residual features are refined through three-level RCB iteration, and the predicted residual ΔE is used to B perform reverse correction on the forged features, and the speaker voiceprint features E' of the original speaker are restored by combining the learnable parameter α A ; S34. Bimodal joint optimization: Combine the feature reconstruction loss and the identity cross-entropy loss to achieve bimodal optimization of voiceprint restoration.
8. A method for constructing a tracing model of the original speaker of forged speech, according to claim 7, characterized in that S33 is specifically as follows: Residual iterative prediction: The preliminary residual feature ΔE0 is input into the 3-level RCB residual correction module, and the accurate residual feature ΔE is obtained through step-by-step refinement B , ΔE1 = RCB1(ΔE0, E′ B ) ΔE2 = RCB2(ΔE1, E′ B ) ΔE B = RCB3(ΔE2, E′ B ) where RCB i a 3-level RCB residual correction module corresponding to (i ∈ [1, 3]); Original speaker voiceprint feature restoration: Using the precise residual feature ΔE B Perform reverse correction on the forged voiceprint feature E A→B to obtain the restored original speaker voiceprint feature E' A = E A→B - αΔE B ; where α is a learnable residual scaling factor and is automatically optimized through backpropagation.
9. A method for constructing a tracing model of the original speaker of forged speech according to claim 7, characterized in that, S34 is specifically as follows: Construct the feature reconstruction loss: Calculate the recovered speaker's voiceprint feature E' A and the original voiceprint feature E A to obtain the reconstruction loss L recon =(‖E' A -E A ‖2) 2 , to constrain the consistency of the feature space; Constructing identity consistency loss: Adopt a pre-trained speaker classifier to map the voiceprint feature vector to a probability distribution vector P = Classifer(E' A ), and then calculate the cross-entropy loss Ensure that the restored features can be accurately traced back to the original identity ID A ; where y i is the true one-hot label of speaker A i , and p i is the probability distribution vector predicted by the speaker classifier for E' A .
10. A method for constructing a tracing model of the original speaker of forged speech according to claim 1, characterized in that, S4 is specifically as follows: S41. Pre-train the Transformer-CLAP hierarchical purification module: Freeze the parameters of the speaker encoder in S1, and use the speaker feature data pairs (E A , E B , E A→B ) to train the Transformer-CLAP hierarchical purification module separately, so that the E' B purified by the Transformer-CLAP hierarchical purification module is close to the true target speaker feature E B ; S42. Pre-train the 3-level RCB residual correction module: Freeze the parameters of the voiceprint encoder in S1 and the Transformer-CLAP hierarchical purification module in S2, and use the data pairs (E A , E' B , E A→B ) to train the 3-level RCB residual correction module separately, so that the E' A restored by the 3-level RCB residual correction module is close to the true voiceprint feature E A ; S43. Joint fine-tuning: Unfreeze all the parameters of the Transformer-CLAP hierarchical purification module and the 3-level RCB residual correction module, conduct end-to-end overall training on the forger voice originator tracing model, and adopt a cosine annealing learning rate schedule.
Citation Information
Cited By
Voiceprint recognition processing method for dispatching telephone
CN122177124A
A voiceprint recognition processing method for dispatching telephones
CN122177124B