Diffusion forged voice tracing method and system

By constructing and training a diffusion forged speech traceability model, combining voiceprint feature purification and refining technology, the problem that the existing technology cannot accurately trace the diffusion forged speech is solved, and high-precision forged speech traceability and reliable judicial evidence collection support is achieved.

CN120126503APending Publication Date: 2025-06-10ZHEJIANG UNIV +1
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510415768.3
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-04-02
Publication Date
2025-06-10

AI Technical Summary

Technical Problem

The existing technology cannot accurately trace the origin and spread the forgers of forged voice, resulting in the inability to protect the rights and interests of victims, which poses potential hidden dangers to the safety of citizens' personal and property.

Method used

By collecting clean speech data sets, calculating the average spectrum corresponding to each phoneme, constructing a diffusion forged speech data set, and training a spectrum encoder and a diffusion forged speech traceability model, combining voiceprint feature purification and refining technology to achieve accurate traceability of diffusion forged speech.

Benefits of technology

It realizes high-precision forged speech traceability, can accurately identify the voiceprint characteristics of diffuse forged speech, provide reliable judicial evidence for evidence, and improves the robustness of the model and data processing efficiency.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120126503A_ABST
    Figure CN120126503A_ABST
Patent Text Reader

Abstract

The invention discloses a diffusion forged voice tracing method and system, and relates to the technical field of artificial intelligence and voice safety. The method comprises the following steps: collecting a clean voice data set, calculating an average frequency spectrum corresponding to each phoneme, and constructing a diffusion forged voice data set; training a frequency spectrum encoder, predicting an average frequency spectrum by using the frequency spectrum encoder, calculating a loss value, updating model parameters of the frequency spectrum encoder, and performing multi-round iteration; constructing and training a diffusion forged voice traceability model, selecting a forged voice and a reference voice, enabling the forged voice and the reference voice to sequentially pass through each layer of the diffusion forged voice traceability model, calculating a loss value, updating parameters of the diffusion forged voice traceability model, and carrying out multi-round iteration; tracing and diffusing the forged voice, and completing a forged voice tracing task. The method can accurately trace and diffuse the voiceprint of a counterfeiter of the counterfeited voice, and plays an auxiliary role in judicial evidence collection.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical fields of artificial intelligence and voice security, and particularly to a method and system for tracing the source of spoofed voice generated by diffusion models. Background Art

[0002] Nowadays, with the progress of artificial intelligence, significant breakthroughs have been made in voice conversion technology, especially in the field of voice generation based on diffusion models, whose fluency and fidelity have reached the leading level in the industry. In the complex voice conversion requirements of multiple scenarios, multiple languages, and multiple voices, diffusion models have become one of the core technologies due to their high-quality generation capabilities and robustness.

[0003] However, cases of using voice conversion technology to create spoofed voice generated by diffusion models for fraud occur frequently. Existing technologies are unable to accurately trace the forgers of spoofed voice generated by diffusion models, which fails to guarantee the rights and interests of victims and poses potential risks to the personal and property safety of citizens. Therefore, a method and system for tracing the source of spoofed voice generated by diffusion models are particularly important.

[0004] Therefore, it is an urgent problem for those skilled in the art to propose a method and system for tracing the source of spoofed voice generated by diffusion models to solve the difficulties existing in the prior art. Summary of the Invention

[0005] In view of this, the present invention provides a method and system for tracing the source of spoofed voice generated by diffusion models, which can accurately trace the voiceprint of the forger of spoofed voice generated by diffusion models and play an auxiliary role in judicial evidence collection.

[0006] To achieve the above object, the present invention adopts the following technical solutions:

[0007] A method for tracing the source of spoofed voice generated by diffusion models includes:

[0008] S1. Collect a clean voice dataset, calculate the average spectrum corresponding to each phoneme, and construct a spoofed voice dataset generated by diffusion models;

[0009] S2. Train a spectrum encoder, use the spectrum encoder to predict the average spectrum, calculate the loss value, update the model parameters of the spectrum encoder, and perform multiple rounds of iteration to obtain a trained spectrum encoder;

[0010] S3. Combine the trained spectrum encoder, construct and train a spoofed voice source tracing model, select spoofed voice and reference voice, make them pass through each layer of the spoofed voice source tracing model in turn, calculate the loss value, update the model parameters of the spoofed voice source tracing model, and perform multiple rounds of iteration;

[0011] S4. Trace the source of spoofed voice generated by diffusion models to complete the task of tracing the source of spoofed voice.

[0012] For the above method, optionally, constructing the diffusion forged speech dataset in S1 specifically includes:

[0013] S101. Collect a clean speech dataset A = {a 1 , a 2 , …, a m} containing n speakers and m pieces of data, where a i = (b i , d i ) is an audio-label data pair, b i is audio data, and d i is a label representing the speaker of the audio data;

[0014] S102. Resample all the audio data in the clean speech dataset A to a unified sampling rate of rt, and obtain the audio data Save the resampled audio data in a unified format, where Resample(·) is the resampling function;

[0015] S103. Calculate the average spectrum corresponding to each phoneme according to the clean speech dataset;

[0016] S104. Based on the speech conversion method of the diffusion model, generate audio or audio spectrum, and use an open-source pre-trained model or self-train to reproduce the speech conversion effect to obtain the diffusion speech conversion model VC(·);

[0017] S105. Construct the diffusion forged speech dataset: Generate a diffusion forged speech dataset C = {c 1 , c 2 , …, c g} containing g pieces of data through the diffusion speech conversion model VC(·), is a forged audio-reference audio-forged speech speaker label data pair;

[0018] where the diffusion forged speech data is the reference audio, d j represents the forged speech speaker, d k represents the speaker of the forged speech, and at the same time satisfies d j ≠ d k .

[0019] For the above method, optionally, calculating the average spectrum corresponding to each phoneme in S103 specifically includes:

[0020] Speech spectrum calculation: Calculate the spectra of all the audio data in the clean speech dataset A to obtain the speech spectra;

[0021] Phoneme-spectrum alignment: Use the phoneme-spectrum alignment algorithm to align the spectrum frames in all the speech spectra with the phonemes;

[0022] Average spectrum calculation: For a specific phoneme, average the spectrum frames corresponding to this phoneme over the entire dataset as the average spectrum of this phoneme.

[0023] In the above method, optionally, the training of the spectrum encoder in S2 specifically includes:

[0024] S201. Randomly select data a from the clean speech dataset A i As training data, calculate the speech spectrum to obtain the speech spectrum where P(·) is the spectrum extraction function;

[0025] S202. Use the phoneme-spectrum alignment algorithm to align the spectrum frames in the obtained speech spectrum s i with the phonemes;

[0026] S203. Use the spectrum encoder to predict the average spectrum of the selected training data to obtain the predicted average spectrum where E(·) is the spectrum encoder;

[0027] S204. Replace the spectrum frames in the speech spectrum s i with the average spectrum of the phonemes calculated in step S103 to obtain the true average spectrum of this data

[0028] S205. Calculate the loss value to obtain Calculate the gradient and use the optimization model to backpropagate and update the model parameters of the spectrum encoder, where Loss 1 (·) is the first loss calculation function;

[0029] S206. Repeat steps S201 to S205 until the loss value is less than the threshold δ or the specified number of iterations MaxIter is reached;

[0030] S207. Denote the spectrum encoder after iteration as E * (·), and in subsequent steps, the model parameters of E * (·) are fixed and no longer changed.

[0031] In the above method, optionally, the construction and training of the diffusion forged speech traceability model in S3 specifically includes:

[0032] S301. Randomly select data from the clean speech dataset A and the diffusion forged speech dataset C as training data;

[0033] If the selected data is clean speech data a i , then the forged audio reference audio X r is of the same length as Consistent 0-matrix audio, label Y = d i ;

[0034] If the selected data is the diffusion forged speech data c i , then the forged audio Reference audio Label Y = d j , where Pad(·) is an audio splicing function used to unify the lengths of the forged audio X vc and the reference audio X r ;

[0035] S302. Calculate the speech spectrum to obtain the forged spectrum S vc = P(X vc ), and the reference spectrum S r = P(X r );

[0036] S303. Use the spectrum encoder E * (·) to predict the average spectrum of the data and obtain the average spectrum S avg = E * (S vc );

[0037] S304. Construct a UNet with α layers as the forged feature purification layer to purify the voiceprint features of the forged speech speaker from S vc , S avg , S r and obtain the processed feature tmp 1 = UNet(S vc , S avg , S r );

[0038] S305. Construct a ResNet with β layers as the residual network layer to refine the obtained voiceprint features and obtain the processed feature tmp 2 = ResNet(tmp 1 );

[0039] S306. Construct an Attentive Pooling with γ channels as the attention pooling layer and a BatchNormalization as the batch normalization layer to pool the voiceprint features and obtain the processed feature tmp 3 = BN(AP(tmp 2 ));

[0040] S307. Construct a fully connected layer with a hidden layer size of σ as the fully connected layer, and construct a Batch Normalization as the batch normalization layer to normalize the voiceprint features to obtain the forged voiceprint V vc = BN(FC(tmp 3 ));

[0041] S308. Update model parameters: Calculate the loss value to obtain L 2 = Loss 2 (V vc , Y), calculate the gradient, and use the optimization model to backpropagate and update the model parameters of the forged feature purification layer UNet, residual network layer ResNet, attention pooling layer Attentive Pooling + batch normalization layer Batch Normalization, and fully connected layer Fully Connected + batch normalization layer Batch Normalization of the diffusion forged voice traceability model. Among them, Loss 2 (·) is the second loss calculation function;

[0042] S309. Repeat the iterative steps S301 to S308 until the loss value is less than the threshold δ or the specified number of iterations MaxIter is reached.

[0043] In the above method, optionally, the tracing of the diffusion forged voice in S4 specifically includes:

[0044] S401. Conduct investigation and evidence collection to obtain the diffusion forged voice x with a sampling rate of rt vc , the voice of the speaker of the forged voice x r , and the voice of the suspected speaker of the forged voice x s ;

[0045] S402. Calculate the forged voiceprint, and repeat the iterative steps S302 to S307. Among them, the forged audio X vc = x vc , and the reference audio X r = x r , to obtain the forged voiceprint V vc ;

[0046] S403. Calculate the voiceprint of the suspected voice, and repeat the iterative steps S302 to S307. Among them, the forged audio X vc = x s , and the reference audio X r is a zero matrix audio with the same length as , to obtain the suspected voiceprint V s ;

[0047] S404. Calculate the similarity between the forged voiceprint and the suspected voiceprint to obtain a similarity score Score = Similarity(V vc , V s ), where Similarity(·) is a similarity calculation function;

[0048] S405. Trace the source of the forged voice. If the similarity score Score is greater than or equal to the threshold Δ, the speaker of the suspected forged voice is the forger of the forged audio; if the similarity score Score is less than the threshold Δ, the speaker of the suspected forged voice is not the forger of the forged audio.

[0049] A diffusion forged voice source tracing system that executes a diffusion forged voice source tracing method described in any one of the above, including: a voice data preprocessing module, a spectrum encoder training module, a diffusion forged language source tracing model construction module, a diffusion forged language source tracing model training module, and a diffusion forged voice source tracing module;

[0050] The voice data preprocessing module, the spectrum encoder training module, the diffusion forged language source tracing model construction module, the diffusion forged language source tracing model training module, and the diffusion forged voice source tracing module are connected in sequence;

[0051] Voice data preprocessing module: Collect a clean voice dataset, calculate the average spectrum corresponding to each phoneme, and construct a diffusion forged voice dataset;

[0052] Spectrum encoder training module: Randomly select data from the clean voice dataset as training data, and iteratively update the spectrum encoder;

[0053] Diffusion forged language source tracing model construction module: Establish a diffusion forged language source tracing model, and the model includes: a forged feature purification layer UNet, a residual network layer ResNet, an attention pooling layer Attentive Pooling + a batch normalization layer BatchNormalization, and a fully connected layer Fully Connected + a batch normalization layer BatchNormalization;

[0054] Diffusion forged language source tracing model training module: Use the optimized model to backpropagate and update the model parameters of the diffusion forged voice source tracing model, and iteratively update the diffusion forged voice source tracing model;

[0055] Diffusion forged voice source tracing module: Trace the source of the diffusion forged voice and complete the task of forging voice source tracing.

[0056] As can be seen from the above technical solutions, compared with the prior art, the present invention provides a method and system for tracing the source of diffused forged speech, which has the following beneficial effects: 1) The present invention has high-precision tracing ability. By constructing a tracing model for diffused forged speech and combining voiceprint feature purification and refinement technologies, it can accurately identify the voiceprint features of diffused forged speech, achieve accurate tracing of forgers, and provide a reliable basis for judicial evidence collection; 2) By adopting multi-level model architectures such as spectral encoders, UNets, and ResNets, and combining attention pooling and voiceprint normalization technologies, the extraction accuracy of voiceprint features and the robustness of the model are significantly improved, ensuring the reliability of the tracing results; 3) It has an efficient data processing process. By steps such as preprocessing of voice data and calculation of average phoneme spectra, the quality of data input is optimized. At the same time, a forged speech dataset is generated using a diffused speech conversion model, providing high-quality data support for model training; 4) Lightweight design and easy to deploy. The model adopts the Attentive Pooling technology, effectively reducing the computational complexity, realizing the lightweight design of the model, and making it easier to be deployed and applied in resource-constrained environments; 5) It has flexibility and scalability, supports unified processing of various voice formats and sampling rates, and can adapt to different scenario requirements by adjusting model parameters (such as the number of UNet layers, the number of ResNet layers, etc.), with strong scalability and adaptability. BRIEF DESCRIPTION OF THE DRAWINGS

[0057] In order to more clearly illustrate the technical solutions in the embodiments of the present invention or the prior art, the following will briefly introduce the drawings required for the description of the embodiments or the prior art. Obviously, the drawings in the following description are only the embodiments of the present invention. For those of ordinary skill in the art, other drawings can be obtained according to the provided drawings without creative efforts.

[0058] Figure 1 It is a flowchart of a method for tracing the source of diffused forged speech provided by the present invention;

[0059] Figure 2 It is a schematic diagram of a tracing model for diffused forged speech provided by the present invention;

[0060] Figure 3 It is a structural block diagram of a system for tracing the source of diffused forged speech provided by the present invention. DETAILED DESCRIPTION OF THE EMBODIMENTS

[0061] Next, the technical solutions in the embodiments of the present invention will be clearly and completely described in conjunction with the accompanying drawings in the embodiments of the present invention. Obviously, the described embodiments are only a part of the embodiments of the present invention, rather than all the embodiments. All other embodiments obtained by those of ordinary skill in the art based on the embodiments of the present invention without creative efforts shall fall within the protection scope of the present invention.

[0062] In the present application, relational terms such as first and second are only used to distinguish one entity or operation from another entity or operation, and do not necessarily require or imply any actual relationship or order between these entities or operations. The terms "comprising", "including" or any other variant thereof are intended to cover non-exclusive inclusion, so that a process, method, article or device comprising a series of elements includes not only those elements but also other elements not expressly listed, or elements inherent to such process, method, article or device. Without further limitation, an element defined by the statement "comprising an..." does not exclude the presence of additional identical elements in the process, method, article or device comprising the element.

[0063] Referring to Figure 1 as shown, the present invention discloses a method for tracing the source of spoofed speech, including:

[0064] S1. Collect a clean speech data set, calculate the average spectrum corresponding to each phoneme, and construct a spoofed speech data set;

[0065] S2. Train a spectrum encoder, use the spectrum encoder to predict the average spectrum, calculate the loss value, update the model parameters of the spectrum encoder, and perform multiple rounds of iteration to obtain a trained spectrum encoder;

[0066] S3. Combine the trained spectrum encoder, construct and train a spoofed speech source tracing model, select spoofed speech and reference speech, make them pass through each layer of the spoofed speech source tracing model in turn, calculate the loss value, update the model parameters of the spoofed speech source tracing model, and perform multiple rounds of iteration;

[0067] S4. Trace the source of the spoofed speech and complete the spoofed speech source tracing task.

[0068] Further, the specific steps of constructing the spoofed speech data set in S1 include:

[0069] S101. Collect a clean speech data set A = {a 1 , a 2 , …, a m} containing n speakers and m pieces of data, where a i = (b i , d i) is an audio-label data pair, b i is audio data, d i is the label representing the speaker of the audio data;

[0070] S102. Resample all the audio data in the clean speech dataset A to a unified sampling rate of rt to obtain the audio data Save the resampled audio data in a unified format, where Resample(·) is the resampling function;

[0071] S103. Calculate the average spectrum corresponding to each phoneme according to the clean speech dataset;

[0072] S104. Based on the voice conversion method of the diffusion model, generate audio or audio spectrum, and use an open-source pre-trained model or reproduce the voice conversion effect by self-training. For example, the DiffVC model, which has open-sourced both the model and parameters and can be directly used, to obtain the diffusion voice conversion model VC(·);

[0073] S105. Construct a diffusion forged speech dataset: Generate a diffusion forged speech dataset C = {c 1 , c 2 , …, c g} containing g pieces of data through the diffusion voice conversion model VC(·), which is a forged audio-reference audio-forged speech speaker label data pair;

[0074] where the diffusion forged speech data is the reference audio, d j represents the forged speech speaker, d k represents the speaker of the speech to be forged, and at the same time satisfies d j ≠ d k .

[0075] Furthermore, the specific calculation of the average spectrum corresponding to each phoneme in S103 includes:

[0076] Speech spectrum calculation: Calculate the spectrum of all the audio data in the clean speech dataset A to obtain the speech spectrum;

[0077] Phoneme-spectrum alignment: Use the phoneme-spectrum alignment algorithm to align the spectrum frames in all the speech spectra with the phonemes;

[0078] Average spectrum calculation: Divide the phonemes using the International Phonetic Alphabet IPA, and for a specific phoneme, average the spectrum frames corresponding to the phoneme over the entire dataset as the average spectrum of the phoneme.

[0079] Further, the specific training of the spectrum encoder in S2 includes:

[0080] S201. Randomly select data a from the clean speech dataset A i As training data, calculate the speech spectrum to obtain the speech spectrum where P(·) is the spectrum extraction function;

[0081] S202. Use the phoneme-spectrum alignment algorithm to align the spectrum frames in the obtained speech spectrum s i with phonemes;

[0082] S203. Use the spectrum encoder to predict the average spectrum of the selected training data to obtain the predicted average spectrum where E(·) is the spectrum encoder;

[0083] S204. Replace the spectrum frames in the speech spectrum s i with the average spectrum of the phonemes calculated in step S103 to obtain the true average spectrum of this data

[0084] S205. Calculate the loss value to obtain Calculate the gradient and use the optimization model to backpropagate and update the model parameters of the spectrum encoder. Among them, Loss 1 (·) is the first loss calculation function;

[0085] S206. Repeat steps S201 to S205 until the loss value is less than the threshold δ or reach the specified number of iterations MaxIter;

[0086] S207. Denote the spectrum encoder after iteration as E * (·). In subsequent steps, the model parameters of E * (·) are fixed and no longer change.

[0087] Furthermore, as shown in Figure 2 S3. Specifically, constructing and training the diffusion forgery speech traceability model includes:

[0088] S301. Randomly select data from the clean speech dataset A and the diffusion forgery speech dataset C as training data;

[0089] If the selected data is clean speech data a i , then the forged audio The reference audio X r is a zero matrix audio with the same length as , and the label Y = d i ;

[0090] If the selected data is diffusion forgery speech data c i , then the forged audio The reference audio Label Y = d j , where Pad(·) is an audio splicing function for uniformly forging audio X vc and the reference audio X r length;

[0091] S302. Calculate the speech spectrum to obtain the forged spectrum S vc = P(X vc ), the reference spectrum S r = P(X r );

[0092] S303. Use the spectrum encoder E * (·) to predict the average spectrum of the data and obtain the average spectrum S avg = E * (S vc );

[0093] S304. Construct a UNet with α layers as the forged feature purification layer, and purify the voiceprint features of the forged speech speaker from S vc , S avg , S r to obtain the processed feature tmp 1 = UNet(S vc , S avg , S r );

[0094] S305. Construct a ResNet with β layers as the residual network layer, refine the obtained voiceprint features, and obtain the processed feature tmp 2 = ResNet(tmp 1 );

[0095] S306. Construct an Attentive Pooling (AP) with γ channels as the attention pooling layer, and construct a BatchNormalization (BN) as the batch normalization layer to pool the voiceprint features and obtain the processed feature tmp 3 = BN(AP(tmp 2 ));

[0096] S307. Construct a Fully Connected (FC) with a hidden layer size of σ as the fully connected layer, and construct a BatchNormalization (BN) as the batch normalization layer to normalize the voiceprint features and obtain the forged voiceprint V vc = BN(FC(tmp 3 ));

[0097] S308. Model parameter update: Calculate the loss value to obtain L 2= Loss 2 (V vc , Y), calculate the gradient, and use the optimized model to backpropagate and update the model parameters of the forgery feature purification layer UNet, residual network layer ResNet, attention pooling layer AP + batch normalization layer BN, and fully connected layer FC + batch normalization layer BN of the diffusion forgery voice traceability model, where Loss 2 (·) is the second loss calculation function;

[0098] S309. Repeatedly iterate steps S301 to S308 until the loss value is less than the threshold δ or reach the specified number of iterations MaxIter.

[0099] Furthermore, the traceability of the diffusion forgery voice in S4 specifically includes:

[0100] S401. Investigate and obtain evidence, and acquire the diffusion forgery voice x with a sampling rate of rt vc , the voice x of the speaker of the forged voice r , and the voice x of the suspected speaker of the forged voice s ;

[0101] S402. Calculate the voiceprint of the forged voice, and repeatedly iterate steps S302 to S307. Among them, the forged audio X vc = x vc , and the reference audio X r = x r , to obtain the forged voiceprint V vc ;

[0102] S403. Calculate the voiceprint of the suspected voice, and repeatedly iterate steps S302 to S307. Among them, the forged audio X vc = x s , and the reference audio X r is a zero matrix audio with the same length as , to obtain the suspected voiceprint V s ;

[0103] S404. Calculate the similarity between the forged voiceprint and the suspected voiceprint, and obtain the similarity score Score = Similarity(V vc , V s ), where Similarity(·) is the similarity calculation function;

[0104] S405. Trace the forged voice. If the similarity score Score is greater than or equal to the threshold Δ, then the suspected speaker of the forged voice is the forger of the forged audio; if the similarity score Score is less than the threshold Δ, then the suspected speaker of the forged voice is not the forger of the forged audio.

[0105] Refer to Figure 3As shown in the figure, a diffusion forgery voice traceability system, which executes a diffusion forgery voice traceability method described in any one of the above, includes: a voice data preprocessing module, a spectrum encoder training module, a diffusion forgery language traceability model construction module, a diffusion forgery language traceability model training module, and a diffusion forgery voice traceability module;

[0106] The voice data preprocessing module, the spectrum encoder training module, the diffusion forgery language traceability model construction module, the diffusion forgery language traceability model training module, and the diffusion forgery voice traceability module are connected in sequence;

[0107] Voice data preprocessing module: Collect a clean voice dataset, calculate the average spectrum corresponding to each phoneme, and construct a diffusion forgery voice dataset;

[0108] Spectrum encoder training module: Randomly select data from the clean voice dataset as training data, and iteratively update the spectrum encoder;

[0109] Diffusion forgery language traceability model construction module: Establish a diffusion forgery language traceability model, and the model includes: a forgery feature purification layer UNet, a residual network layer ResNet, an attention pooling layer Attentive Pooling + a batch normalization layer BatchNormalization, and a fully connected layer Fully Connected + a batch normalization layer BatchNormalization;

[0110] Diffusion forgery language traceability model training module: Use the optimized model to backpropagate and update the model parameters of the diffusion forgery voice traceability model, and iteratively update the diffusion forgery voice traceability model;

[0111] Diffusion forgery voice traceability module: Trace the diffusion forgery voice and complete the forgery voice traceability task.

[0112] In a specific embodiment, the clean speech dataset uses the VCTK dataset, which contains 109 speakers and 40,000 pieces of data; the resampling function in FFmpeg is used as Resample(·), the unified sampling rate rt is set to 16,000 Hz, the Mel Frequency Cepstral Coefficients (MFCC) are used as the spectral extraction function P(·) to calculate the speech spectrum, and the Montreal forced alignment algorithm is used as the phoneme-spectrum alignment algorithm to align the spectral frames in all speech spectra with phonemes; the open-source project DiffVC is used as the diffusion voice conversion model VC(·) to generate a diffusion forged speech dataset containing 4,000 pieces of data, a 3-layer Transformer structure is used as the spectral encoder E(·), and the model parameters of the spectral encoder E(·) are updated by backpropagation using the optimization model, repeating the iteration until the loss value is less than the threshold δ or reaching the specified number of iterations MaxIter. Take δ as 1e-8 and MaxIter as 1,000,000; the iterated spectral encoder is denoted as E * (·), and in the subsequent steps, the model parameters of E * (·) will no longer change; Diffusion forged speech traceability model construction and training stage: Construct a UNet with 10 layers as the forged feature purification layer processing layer, construct a ResNet with 10 layers as the forged feature purification layer processing layer, construct an Attentive Pooling (AP) with 1024 channels as the attention pooling layer, construct a BatchNormalization (BN) as the batch normalization layer for pooling the voiceprint features, construct a Fully Connected (FC) with a hidden layer size of 1024 as the fully connected layer, and construct a BatchNormalization (BN) as the batch normalization layer for normalizing the voiceprint features; use the optimization model to backpropagate and update the model parameters of the diffusion forged speech traceability model, repeating the iteration until the loss value is less than the threshold δ or reaching the specified number of iterations MaxIter. Take δ as 1e-8 and MaxIter as 1,000,000; Diffusion forged speech traceability stage, investigation and evidence collection, obtain the diffusion forged speech x with a sampling rate of 1600 vc , the speech x of the forged speech speaker r , the speech x of the suspected forged speech speaker s , calculate the forged speech voiceprint, where the forged audio X vc = x vc , the reference audio X r = x r , and obtain the forged voiceprint V vc ; calculate the voiceprint of the suspected speech, where the forged audio X vc = x s , and the reference audio X r is the length of Consistent 0 matrix audio to obtain the suspected voiceprint V s ; Calculate the voiceprint similarity to obtain the similarity score Score = Similarity(V vc , V s ). The cosine similarity function (Cosine Similarity) is used as Similarity(·). If the similarity score Score is greater than or equal to the threshold Δ, then the speaker of the suspected forged voice is the forger of the forged audio; if the similarity score Score is less than the threshold Δ, then the speaker of the suspected forged voice is not the forger of the forged audio. Here, Δ is taken as 0.2.

[0113] Each embodiment in this specification is described in a progressive manner. The same or similar parts among the embodiments can be referred to each other, and the key points of each embodiment are the differences from other embodiments. In particular, for the system or system embodiments, since they are basically similar to the method embodiments, the description is relatively simple, and the relevant parts can be referred to the description of the method embodiments. The systems and system embodiments described above are only illustrative. The units described as separate components may or may not be physically separated, and the components shown as units may or may not be physical units, that is, they may be located in one place or distributed to multiple network units. Some or all of the modules can be selected according to actual needs to achieve the purpose of the solution of this embodiment. Those of ordinary skill in the art can understand and implement it without creative work.

[0114] The above description of the disclosed embodiments enables those skilled in the art to implement or use the present invention. Various modifications to these embodiments will be obvious to those skilled in the art, and the general principles defined herein can be implemented in other embodiments without departing from the spirit or scope of the present invention. Therefore, the present invention will not be limited to these embodiments shown herein, but will be accorded the widest scope consistent with the principles and novel features disclosed herein.

Claims

1. A method for tracing the source of a spread forged voice, characterized in that: include: S1. Collect clean speech data sets, calculate the average spectrum corresponding to each phoneme and construct a diffusion forged speech data set; S2, training the spectrum encoder, using the spectrum encoder to predict the average spectrum, calculating the loss value, updating the model parameters of the spectrum encoder and performing multiple rounds of iterations to obtain a trained spectrum encoder; S3. Combine the trained spectrum encoder to build and train the diffusion forged speech source tracing model, select the forged speech and reference speech, make them pass through each layer of the diffusion forged speech source tracing model in turn, calculate the loss value, update the parameters of the diffusion forged speech source tracing model and perform multiple rounds of iterations; S4. Trace and spread the forged voice and complete the task of tracing the forged voice.

2. A method for tracing the source of a spread forged voice according to claim 1, characterized in that: The construction of the diffusion forged voice dataset in S1 specifically includes: S101. Collect a clean speech dataset A containing n speakers and m pieces of data. m }, where a i =(b i ,d i ) is an audio-label data pair, b i is the audio data, d i is a label representing the speaker of the audio data; S102, resample all audio data in the clean speech data set A to a uniform sampling rate of rt to obtain audio data Save the resampled audio data into a unified format, where Resample(·) is the resampling function; S103, calculating the average spectrum corresponding to each phoneme according to the clean speech data set; S104, a diffusion model-based voice conversion method generates audio or audio spectrum, uses an open source pre-trained model or self-training to reproduce the voice conversion effect, and obtains a diffusion voice conversion model VC(·); S105, constructing a diffusion forged speech data set: Generate a diffusion forged speech data set C containing g data items through a diffusion speech conversion model VC(·) = {c1, c2, …, c g }, It is a pair of fake audio-reference audio-fake speech speaker label data; Among them, spreading fake voice data It is the reference audio. is the reference audio of the speaker whose voice is forged, d j represents the fake voice speaker, d k represents the speaker of the forged voice, and satisfies d j ≠d k .

3. A method for tracing the source of a spread forged voice according to claim 2, characterized in that: Calculating the average spectrum corresponding to each phoneme in S103 specifically includes: Speech spectrum calculation: Calculate the spectrum of all audio data in the clean speech data set A to obtain the speech spectrum; Phoneme-spectrum alignment: Using the phoneme-spectrum alignment algorithm, all spectrum frames in the speech spectrum are aligned with the phonemes; Average spectrum calculation: For a specific phoneme, the spectrum frames corresponding to the phoneme in the entire dataset are averaged to obtain the average spectrum of the phoneme.

4. A method for tracing the source of a spread forged voice according to claim 3, characterized in that: Training the spectrum encoder in S2 specifically includes: S201, randomly select data a from the clean speech data set A i As training data, calculate the speech spectrum and get the speech spectrum Where, P(·) is the spectrum extraction function; S202, using the phoneme-spectrum alignment algorithm, the obtained speech spectrum s i The spectral frames in are aligned with the phonemes; S203: Use the spectrum encoder to predict the average spectrum of the selected training data to obtain the predicted average spectrum. Where, E(·) is the spectrum encoder; S204: The speech spectrum s i The spectrum frame in step S100 is replaced with the average spectrum of the phoneme calculated in step S103 to obtain the true average spectrum of the data. S205, calculate the loss value, and obtain Calculate the gradient and use the optimization model to update the model parameters of the spectrum encoder, where Loss1(·) is the first loss calculation function; S206, repeating steps S201 to S205 until the loss value is less than a threshold value δ or reaches a specified number of iterations MaxIter; S207, the spectrum encoder after the iteration is recorded as E * (·), in the following step E * The model parameters of (·) are fixed and do not change.

5. A method for tracing the source of a spread forged voice according to claim 4, characterized in that: The construction and training of the diffusion forged voice tracing model in S3 specifically includes: S301, randomly selecting data from the clean speech data set A and the diffused forged speech data set C as training data; If the selected data is clean voice data a i , then fake audio Reference Audio X r For length and Consistent 0 matrix audio, label Y=d i ; If the selected data is the diffusion forged voice data c i , then fake audio Reference Audio Label Y = d j , where Pad(·) is the audio patching function used to unify the forged audio X vc With Reference Audio X r length; S302, calculate the speech spectrum and obtain the forged spectrum S vc =P(X vc ), reference spectrum S r =P(X r ); S303, using spectrum encoder E * (·) Predict the average spectrum of the data and obtain the average spectrum S avg =E * (S vc ); S304, construct a UNet with a layer number of α as a forged feature purification layer, from S vc ,S avg ,S r Purify the voiceprint features of the forged voice speaker and obtain the processed features tmp1 = UNet (S vc ,S avg ,S r ); S305, construct a ResNet with a layer number of β as a residual network layer, refine the obtained voiceprint features, and obtain the processed features tmp2 = ResNet (tmp1); S306, construct an Attentive Pooling with a channel number of γ as an attention pooling layer, construct a BatchNormalization as a batch normalization layer, pool the voiceprint features, and obtain the processed feature tmp3 = BN (AP (tmp2)); S307, construct a Fully Connected layer with a hidden layer size of σ as a fully connected layer, construct a BatchNormalization layer as a batch normalization layer, normalize the voiceprint features, and obtain the forged voiceprint V vc =BN(FC(tmp3)); S308, model parameter update: calculate the loss value, and obtain L2=Loss2(V vc ,Y), calculate the gradient, and use the optimization model to backpropagate and update the model parameters of the forged feature purification layer UNet, the residual network layer ResNet, the attention pooling layer Attentive Pooling + batch normalization layer BatchNormalization, and the fully connected layer Fully Connected + batch normalization layer BatchNormalization of the diffusion forged speech tracing model, where Loss2(·) is the second loss calculation function; S309, repeat iterative steps S301 to S308 until the loss value is less than the threshold δ or reaches the specified number of iterations MaxIter.

6. A method for tracing the source of a spread forged voice according to claim 5, characterized in that: The specific source-tracing and diffusion of forged voice in S4 include: S401, investigate and collect evidence, obtain the diffuse forged voice x with a sampling rate of rt vc , the forged voice speaker voice x r , suspected fake voice speaker voice x s ; S402, calculate the forged voiceprint, and repeat the iterative steps S302 to S307, wherein the forged audio X vc =x vc , Reference Audio X r =x r , get the fake voiceprint V vc ; S403, calculate the suspected voiceprint, and repeat the steps S302 to S307, wherein the forged audio X vc =x s , Reference Audio X r For length and Consistent 0 matrix audio, get the suspected voiceprint V s ; S404, calculate the similarity between the forged voice print and the suspected voice print, and obtain a similarity score Score = Similarity (V vc ,V s ), where Similarity(·) is the similarity calculation function; S405, tracing the source of the forged voice, if the similarity score Score is greater than or equal to the threshold Δ, then the speaker of the suspected forged voice is the forger of the forged audio; if the similarity score Score is less than the threshold Δ, then the speaker of the suspected forged voice is not the forger of the forged audio.

7. A diffusion forged voice tracing system, executing a diffusion forged voice tracing method according to any one of claims 1 to 6, comprising: Speech data preprocessing module, spectrum encoder training module, diffusion forged language tracing model construction module, diffusion forged language tracing model training module and diffusion forged speech tracing module; The speech data preprocessing module, the spectrum encoder training module, the diffusion forged language tracing model construction module, the diffusion forged language tracing model training module and the diffusion forged speech tracing module are connected in sequence; Speech data preprocessing module: collect clean speech data sets, calculate the average spectrum corresponding to each phoneme and construct a diffusion-forged speech data set; Spectral encoder training module: randomly selects data from the clean speech dataset as training data and iteratively updates the spectral encoder; Diffusion forged language traceability model construction module: Establish a diffusion forged language traceability model, which includes: forged feature purification layer UNet, residual network layer ResNet, attention pooling layer Attentive Pooling + batch normalization layer BatchNormalization and fully connected layer Fully Connected + batch normalization layer BatchNormalization; Diffusion forged language tracing model training module: Use the optimization model feedback to update the model parameters of the diffusion forged speech tracing model, and iteratively update the diffusion forged speech tracing model; Diffusion forged voice tracing module: trace the diffusion forged voice and complete the forged voice tracing task.