Speech anti-cloning method and device based on non-learnable speech samples
By generating unlearning speech samples and adding perturbations using the minimization of the objective function, the problem of inability to prevent speech cloning in the prior art is solved, and the effect of preventing speech cloning during speech synthesis is achieved.
Patent Information
- Application Number
- CN202510436285.1
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-04-09
- Publication Date
- 2025-07-08
- Estimated Expiration
- 2045-04-09
AI Technical Summary
Existing detection-based speech protection methods cannot effectively prevent speech cloning and cannot prevent cloning of cloned speech before speech synthesis.
By generating non-learning speech samples, add perturbations to the speech samples to be protected using the objective function minimization method, and generate non-learning speech samples to prevent the target speech synthesis model from learning the speech characteristics of the speaker.
Effectively prevent speech cloning, ensure that the target speech synthesis model cannot learn the voice characteristics of the speaker, and improve the security of speech prevention cloning.
Smart Images

Figure CN119943063B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of voice anti-cloning, and more particularly, to a voice anti-cloning method and apparatus based on unlearnable voice samples. Background Art
[0002] Voice cloning is a technology that replicates human voice characteristics through artificial intelligence technology to generate a voice highly similar to that of a real person. To prevent unauthorized users from using cloned voices, related technicians have proposed voice protection methods based on detection; voice protection based on detection aims to use advanced detection algorithms to determine whether an audio is real or synthetic. The protection methods for voice detection mainly focus on liveness detection and acoustic signal analysis, based on physical characteristics such as vocal cord vibration and pronunciation gestures of the human voice system, as well as low-dimensional signal features such as Mel Frequency Cepstral Coefficient (MFCC) and signal power linearity. There are also some voice detection protections through an auxiliary Speaker Verification (SV) system that learns the differences between fake and real audio to determine the authenticity of a given audio segment.
[0003] However, these voice protection methods based on detection are auxiliary judgment measures provided after the cloning voice synthesis and cannot achieve voice anti-cloning. Summary of the Invention
[0004] In view of this, the purpose of this application is to provide a voice anti-cloning method and apparatus based on unlearnable voice samples, which can generate unlearnable voice samples through the updated target perturbation, so that the target voice synthesis model cannot learn the voice characteristics of the speaker, and thus prevent voice cloning during the voice synthesis process.
[0005] In a first aspect, an embodiment of this application provides a voice anti-cloning method based on unlearnable voice samples, and the method includes:
[0006] Obtain a voice sample to be protected and the text information corresponding to the voice sample to be protected;
[0007] Update and iterate the perturbation added to the voice sample to be protected with the minimum of the objective function as the optimization goal to obtain the target perturbation; where the objective function includes a first loss for measuring the distance between the voice sample to be protected and the synthesized voice, a second loss for measuring the voice feature hiding characteristic of the perturbation, and a third loss for measuring the auditory inaudible characteristic of the perturbation; the synthesized voice is synthesized based on the linear spectrogram of the voice sample to be protected with the added perturbation and the text information;
[0008] Add the target perturbation to the voice sample to be protected to obtain an unlearnable voice sample.
[0009] In a possible implementation manner, the perturbation added to the speech sample to be protected is updated and iterated with the minimum of the objective function as the optimization objective to obtain the target perturbation, including:
[0010] For the current round of optimization process, the linear spectrogram and text information of the speech sample to be protected added with the perturbation of the previous round are input into the alternative speech synthesis model to obtain the synthesized speech of the current round;
[0011] According to the mel spectrogram of the speech sample to be protected and the mel spectrogram of the synthesized speech of the current round, calculate the first loss of the current round;
[0012] According to the mel spectrogram of the synthesized speech of the current round and the mel spectrogram of the random Gaussian distribution noise, calculate the second loss of the current round;
[0013] If the third loss is enabled, then according to the speech sample to be protected added with the perturbation of the previous round and the speech sample to be protected, calculate the third loss of the current round; and update the perturbation of the previous round based on the first loss, the second loss, and the third loss of the current round to obtain the perturbation of the current round.
[0014] In a possible implementation manner, the method further includes:
[0015] If the third loss is not enabled, then update the perturbation of the previous round based on the first loss and the second loss of the current round to obtain the perturbation of the current round.
[0016] In a possible implementation manner, according to the mel spectrogram of the speech sample to be protected and the mel spectrogram of the synthesized speech of the current round, calculating the first loss of the current round includes:
[0017] Substitute the mel spectrogram of the speech sample to be protected and the mel spectrogram of the synthesized speech of the current round into the following first loss function to obtain the first loss of the current round;
[0018] ;
[0019] Wherein, is the first loss of the current round, is the mel spectrogram of the speech sample to be protected , is the mel spectrogram of the synthesized speech of the current round ; is the speech sample to be protected and the synthesized speech of the current round between distance.
[0020] In a possible implementation manner, according to the mel spectrogram of the synthesized speech of the current round and the mel spectrogram of the random Gaussian distribution noise, calculating the second loss of the current round includes:
[0021] Substitute the Mel spectrogram of the synthesized speech in this round and the Mel spectrogram of the random Gaussian noise into the following second loss function to obtain the second loss in this round;
[0022] ;
[0023] where, is the second loss in this round, is the Mel spectrogram of the synthesized speech in this round , is the Mel spectrogram of the random Gaussian noise , is the KL divergence between the synthesized speech in this round and the random Gaussian noise , is the distance between the synthesized speech in this round and the random Gaussian noise .
[0024] In a possible implementation, calculate the third loss in this round according to the protected speech sample added with the perturbation in the previous round and the protected speech sample, including:
[0025] Substitute the protected speech sample added with the perturbation in the previous round and the protected speech sample into the following third loss function to obtain the third loss in this round;
[0026] ;
[0027] ;
[0028] where, is the third loss in this round, is the short-term intelligibility index loss between the protected speech sample added with the perturbation in the previous round and the protected speech sample , is the short-time Fourier transform loss function, is the protected speech sample, is the perturbation in the previous round, is the value obtained after the short-time Fourier transform of the protected speech sample added with the perturbation in the previous round, is the value obtained after the short-time Fourier transform of the protected speech sample, is the distance between the value obtained after the short-time Fourier transform of the protected speech sample added with the perturbation in the previous round and the value obtained after the short-time Fourier transform of the protected speech sample .
[0029] Second aspect, an embodiment of the present application further provides a voice anti-cloning device based on unlearnable voice samples, and the device includes:
[0030] An acquisition module, configured to acquire a voice sample to be protected and text information corresponding to the voice sample to be protected;
[0031] A perturbation update module, configured to update and iterate the perturbation added to the voice sample to be protected with the minimum of the objective function as the optimization goal to obtain a target perturbation; wherein, the objective function includes a first loss for measuring the distance between the voice sample to be protected and the synthesized voice, a second loss for measuring the voice feature hiding characteristic of the perturbation, and a third loss for measuring the auditory inaudible characteristic of the perturbation; the synthesized voice is synthesized based on the linear spectrogram and text information of the voice sample to be protected added with the perturbation;
[0032] A sample generation module, configured to add the target perturbation to the voice sample to be protected to obtain an unlearnable voice sample.
[0033] In a possible implementation manner, the perturbation update module is specifically configured to, for the current round of optimization process, input the Mel spectrogram and text information of the voice sample to be protected added with the perturbation of the previous round into a voice synthesis model to obtain the synthesized voice of the current round; calculate the first loss of the current round according to the Mel spectrogram of the voice sample to be protected and the Mel spectrogram of the synthesized voice of the current round; calculate the second loss of the current round according to the Mel spectrogram of the synthesized voice of the current round and the Mel spectrogram of the random Gaussian distribution noise; if the third loss is enabled, calculate the third loss of the current round according to the voice sample to be protected added with the perturbation of the previous round and the voice sample to be protected; and update the perturbation of the previous round based on the first loss, the second loss, and the third loss of the current round to obtain the perturbation of the current round.
[0034] Third aspect, an embodiment of the present application further provides an electronic device, including: a processor, a storage medium, and a bus. The storage medium stores machine-readable instructions executable by the processor. When the electronic device runs, the processor communicates with the storage medium through the bus, and the processor executes the machine-readable instructions to perform the steps of the voice anti-cloning method based on unlearnable voice samples according to any one of the first aspect.
[0035] Fourth aspect, an embodiment of the present application further provides a computer-readable storage medium, on which a computer program is stored. When the computer program is run by a processor, it performs the steps of the voice anti-cloning method based on unlearnable voice samples according to any one of the first aspect.
[0036] The embodiments of the present application provide a voice anti-cloning method and device based on non-learnable voice samples. The method includes: obtaining a voice sample to be protected and the corresponding text information of the voice sample to be protected; updating and iterating the perturbation added to the voice sample to be protected with the minimum of the objective function as the optimization goal to obtain the target perturbation; where the objective function includes a first loss for measuring the distance between the voice sample to be protected and the synthesized voice, a second loss for measuring the voice feature hiding property of the perturbation, and a third loss for measuring the auditory inaudible property of the perturbation; the synthesized voice is synthesized based on the linear spectrogram of the voice sample to be protected with the perturbation added and the text information; adding the target perturbation to the voice sample to be protected to obtain a non-learnable voice sample. Through the present application, a non-learnable voice sample can be generated through the updated target perturbation, so that the target voice synthesis model cannot learn the voice features of the speaker, and thus prevent voice cloning during the voice synthesis process. BRIEF DESCRIPTION OF THE DRAWINGS
[0037] In order to more clearly illustrate the technical solutions of the embodiments of the present application, the drawings required for the embodiments will be briefly introduced below. It should be understood that the following drawings only show some embodiments of the present application and should not be regarded as limiting the scope. For those of ordinary skill in the art, other related drawings can be obtained based on these drawings without creative efforts.
[0038] Figure 1 The flowchart of the voice anti-cloning method based on non-learnable voice samples provided by the embodiments of the present application is shown;
[0039] Figure 2 The flowchart of updating the perturbation provided by the embodiments of the present application is shown;
[0040] Figure 3 The structural schematic diagram of the voice anti-cloning device based on non-learnable voice samples provided by the embodiments of the present application is shown;
[0041] Figure 4 The structural schematic diagram of an electronic device provided by the embodiments of the present application is shown. DETAILED DESCRIPTION
[0042] To make the objectives, technical solutions, and advantages of the embodiments of this application clearer, the technical solutions in the embodiments of this application will be clearly and completely described below with reference to the accompanying drawings in the embodiments of this application. It should be understood that the accompanying drawings in this application are only for the purposes of illustration and description, and are not used to limit the protection scope of this application. Additionally, it should be understood that the schematic drawings are not drawn to scale. The flowcharts used in this application illustrate the operations implemented according to some embodiments of this application. It should be understood that the operations in the flowchart may not be implemented in sequence, and steps without a logical context relationship may be reversed or implemented simultaneously. Moreover, those skilled in the art can add one or more other operations to the flowchart or remove one or more operations from the flowchart under the guidance of the content of this application.
[0043] In addition, the described embodiments are only some embodiments of this application, rather than all embodiments. The components of the embodiments of this application usually described and illustrated in the accompanying drawings here can be arranged and designed in various different configurations. Therefore, the following detailed description of the embodiments of this application provided in the accompanying drawings is not intended to limit the scope of the claimed application, but merely represents the selected embodiments of this application. Based on the embodiments of this application, all other embodiments obtained by those skilled in the art without creative efforts belong to the protection scope of this application.
[0044] To enable those skilled in the art to use the content of this application, the following implementation manners are given in combination with a specific application scenario, the "voice anti-cloning field". For those skilled in the art, without departing from the spirit and scope of this application, the general principles defined here can be applied to other embodiments and application scenarios. Although this application is mainly described around the "voice anti-cloning field", it should be understood that this is only an exemplary embodiment.
[0045] It should be noted that the term "including" will be used in the embodiments of this application to indicate the existence of the features stated thereafter, but does not exclude adding other features.
[0046] The voice anti-cloning method based on non-learnable voice samples provided in the embodiments of this application will be described in detail below.
[0047] Refer to Figure 1 As shown, it is a schematic flowchart of the voice anti-cloning method based on non-learnable voice samples provided in the embodiments of this application. The following will explain each step of the embodiments of this application by way of example:
[0048] S101. Obtain the voice sample to be protected and the text information corresponding to the voice sample to be protected.
[0049] In the embodiments of the present application, the voice sample to be protected refers to the voice data that needs to be protected against cloning and can be extracted from the voice files published by users. The voice files can be files containing only voices or videos with voices, etc. The text information corresponding to the voice sample to be protected corresponds one-to-one with the voice sample to be protected in the original dataset and can be directly obtained from the original dataset.
[0050] S102. Update and iterate the perturbation added to the voice sample to be protected with the minimum of the objective function as the optimization goal to obtain the target perturbation.
[0051] In the embodiments of the present application, the objective function refers to the objective function used to optimize the perturbation. The objective function includes a first loss for measuring the distance between the voice sample to be protected and the synthesized voice, a second loss for measuring the voice feature hiding property of the perturbation, and a third loss for measuring the auditory inaudible property of the perturbation. The synthesized voice is synthesized based on the linear spectrogram and text information of the voice sample to be protected with the perturbation added.
[0052] Among them, the voice feature hiding property refers to the degree of hiding of the voice features (such as timbre) of the speaker. The auditory inaudible property refers to the degree of perception of the perturbation by human hearing during the voice feature hiding process.
[0053] Here, the embodiments of the present application introduce the Error-Minimizing Noise (EM) technology. By updating the perturbation, the error of the alternative voice synthesis model is reduced to replace the normal learning of the model, so that there is no learnable content in the "voice sample with the target perturbation added (i.e., the unlearnable voice sample)", so that the target voice synthesis model cannot learn the privacy-sensitive voice information of the user in the unlearnable voice sample, preventing voice cloning. The alternative voice synthesis model refers to the voice synthesis model that assists in optimizing the perturbation. Its input can include the linear spectrogram, text information, etc. of the voice sample to be protected, and its output is the synthesized voice corresponding to the text information. The target voice synthesis model refers to the model for voice synthesis in actual applications trained with unlearnable voice samples. Its input is text information, and its output is the synthesized voice corresponding to the text information.
[0054] In addition, since the embodiments of the present application, as the defense party, can only modify the voice sample to be protected and need to keep the text input accurate, the perturbation in the present application can only be added to the voice sample.
[0055] Furthermore, minimizing the objective function can protect the speech features in the speech sample to be protected from being learned by the target speech synthesis model. Therefore, in the embodiments of the present application, a TTS (Text-to-Speech) model can be selected as the alternative speech synthesis model, and the objective function of the alternative speech synthesis model is used as the perturbed objective function to generate protected speech (i.e., the speech sample to be protected after adding the target perturbation). Specifically, a generative TTS model with an encoder-decoder structure (such as a variational autoencoder) is adopted to construct the alternative speech synthesis model to learn the distribution of the input speech samples. The requirements for the input of this alternative speech synthesis model are relatively complex (which can include one or more of text information, speaker identity, the speech sample to be protected, and the linear spectrogram of the speech sample to be protected). In view of the differences in different TTS model architectures, the present application focuses on the generator g of the alternative speech synthesis model. By optimizing the proposed objective function, the error of the alternative speech synthesis model is significantly reduced, making the speech sample after adding the perturbation "unlearnable".
[0056] Among them, the selection of the objective function needs to meet the following three criteria: (1) The objective function can be optimized by perturbation; (2) The objective function should be general enough to face various TTS models and does not rely on prior knowledge; (3) The optimization complexity of the objective function is less than the preset complexity, such as having the fastest convergence rate or the largest information entropy; the optimization complexity refers to the difficulty of optimizing the objective function.
[0057] Refer to Figure 2 As shown, it is the flowchart of updating the perturbation provided by the embodiments of the present application. Taking the minimum of the objective function as the optimization goal, the perturbation added to the speech sample to be protected is updated iteratively to obtain the target perturbation, including:
[0058] S201. For the current round of optimization process, the linear spectrogram and text information of the speech sample to be protected with the previous round of perturbation added are input into the alternative speech synthesis model to obtain the synthesized speech of this round.
[0059] Here, in addition to the "linear spectrogram of the speech sample to be protected with the previous round of perturbation added" and "text information", other data can also be input into the alternative speech synthesis model according to the actual situation to improve the speech anti-cloning effect.
[0060] S202. According to the Mel spectrogram of the speech sample to be protected and the Mel spectrogram of the synthesized speech of this round, calculate the first loss of this round.
[0061] In the embodiments of the present application, since the structures and objective functions of different TTS models are different, in order to improve the transferability of unlearnable speech samples among different TTS models, a general first loss function is proposed, and most mainstream generative speech synthesis models will select this as one of the optimization objectives.
[0062] Specifically, the mel spectrogram of the speech sample to be protected and the mel spectrogram of the synthesized speech in this round are substituted into the following first loss function to obtain the first loss in this round;
[0063] ;
[0064] Wherein, is the first loss in this round, is the mel spectrogram of the speech sample to be protected and is the mel spectrogram of the synthesized speech in this round and is the distance between the speech sample to be protected and the synthesized speech in this round .
[0065] S203. Calculate the second loss in this round according to the mel spectrogram of the synthesized speech in this round and the mel spectrogram of random Gaussian distributed noise.
[0066] Wherein, the random Gaussian distributed noise refers to random noise whose probability conforms to the Gaussian distribution.
[0067] Fowl et al. found through research that using targeted adversarial attacks can achieve a better poisoning effect, that is, making the model unable to learn the training data. In order to further improve the speech feature protection effect in the present application, a speech perturbation concealment technology (SPEC) based on the Kullback-Leibler Divergence is proposed. The SPEC technology expects the synthesized speech to be random noise, not containing the speech features of the speaker and the speech content cannot be clearly heard. The SPEC technology calculates the KL divergence value between the synthesized speech and the random Gaussian distributed noise, and supplements it with distance as a part to improve the poisoning effect of the perturbation.
[0068] Specifically, the mel spectrogram of the synthesized speech in this round and the mel spectrogram of random Gaussian distributed noise are substituted into the following second loss function to obtain the second loss in this round;
[0069] ;
[0070] Wherein, is the second loss in this round, For the mel spectrogram of the synthesized speech in this round and for the mel spectrogram of the random Gaussian distributed noise and for the KL divergence between the synthesized speech in this round and the random Gaussian distributed noise and for the distance between the synthesized speech in this round and the random Gaussian distributed noise
[0071] S204. If the third loss is enabled, calculate the third loss in this round based on the protected speech sample with the previous round of perturbation added and the protected speech sample
[0072] Here, protecting the protected speech sample based on the SPEC technology can effectively interfere with the training of the TTS model, and the generated perturbation should not affect the normal use of the synthesized speech. Zhang et al. found that the speech enhancement technology can effectively reduce the human ear's perception of such adversarial noise. This application proposes a noise perception optimization technology based on the speech enhancement technology
[0073] In the speech enhancement technology, the short-time objective intelligibility (STOI) and the short-time Fourier transform (STFT) loss function are two commonly used speech enhancement technologies. As the core of the perception module, STOI objectively quantifies the intelligibility of the synthesized speech by calculating the short-time envelope correlation (the score ranges from 0 to 1, and the higher the score, the better the sound quality) between the protected speech sample and the protected speech sample with perturbation added (i.e., the protected speech sample). This index is highly correlated with human auditory perception. Optimizing the STOI loss function can enhance the naturalness of the sound, and the specific calculation process follows the existing literature
[0074] Meanwhile, this application combines the STFT transform in the time domain and frequency domain dimensions within the radius (referring to the space measured by the norm) to optimize human auditory perception. And STFT has excellent feature extraction ability. This application uses the distance to construct the short-time Fourier transform loss function
[0075] Specifically, substitute the protected speech sample with the previous round of perturbation added and the protected speech sample into the following third loss function to obtain the third loss in this round
[0076] ;
[0077] ;
[0078] Among them, is the third loss in this round, is the short-term intelligibility index loss between the protected speech sample with the previous round's perturbation added and the protected speech sample ; is the short-time Fourier transform loss function, is the protected speech sample, is the previous round's perturbation, is the value obtained after the short-time Fourier transform of the protected speech sample with the previous round's perturbation added, is the value obtained after the short-time Fourier transform of the protected speech sample, is the distance between the value obtained after the short-time Fourier transform of the protected speech sample with the previous round's perturbation added and the value obtained after the short-time Fourier transform of the protected speech sample ;
[0079] S205. Update the previous round's perturbation based on the first loss, the second loss, and the third loss in this round to obtain the perturbation in this round.
[0080] Step 1. Substitute the first loss, the second loss, and the third loss in this round into the following formula to obtain the objective function value in this round;
[0081] ;
[0082] Among them, is the objective function value in this round, is the hyperparameter of the preset second loss function, is the hyperparameter of the preset third loss function.
[0083] Step 2. Update the previous round's perturbation based on the objective function value to obtain the perturbation in this round;
[0084] ;
[0085] Among them, is the perturbation in this round obtained after updating the previous round's perturbation, is the gradient vector of the objective function value in this round with respect to the protected speech sample, is the sign function representing the direction of the gradient of the objective function value in this round with respect to the protected speech sample, is the minimum value of the perturbation, is the maximum value of the perturbation, To limit the value of the perturbation within .
[0086] S206. If the third loss is not enabled, update the previous-round perturbation based on the current-round first loss and the current-round second loss to obtain the current-round perturbation.
[0087] Step 1. Substitute the current-round first loss and the current-round second loss into the following formula to obtain the current-round objective function value;
[0088] ;
[0089] where, is the current-round objective function value, is the hyperparameter of the preset second loss function.
[0090] Step 2. Update the previous-round perturbation based on the objective function value to obtain the current-round perturbation;
[0091] ;
[0092] where, is the current-round perturbation obtained after updating the previous-round perturbation, is the gradient vector of the current-round objective function value with respect to the speech sample to be protected, is the sign function used to represent the direction of the gradient of the current-round objective function value with respect to the speech sample to be protected, is the minimum value of the perturbation, is the maximum value of the perturbation, To limit the value of the perturbation within .
[0093] S103. Add the target perturbation to the speech sample to be protected to obtain an unlearnable speech sample.
[0094] Furthermore, use the text information as the sample and the unlearnable speech sample as the label to train the target speech synthesis model. Make the target speech synthesis model unable to learn the speech features of the speaker from the unlearnable speech sample to prevent voice cloning.
[0095] Here, the Word Error Rate (WER) and Speaker Similarity (SIM) can be used to measure the quality of the speech synthesis of the trained target speech synthesis model. The higher the WER, the lower the quality of the synthesized speech, and the lower the SIM, the lower the similarity between the synthesized speech and the speech of the original speaker.
[0096] The embodiments of the present application provide a voice anti-cloning method based on non-learnable voice samples. The method includes: obtaining a voice sample to be protected and the text information corresponding to the voice sample to be protected; updating and iterating the perturbation added to the voice sample to be protected with the minimum of the objective function as the optimization target to obtain the target perturbation; wherein the objective function includes a first loss for measuring the distance between the voice sample to be protected and the synthesized voice, a second loss for measuring the voice feature hiding property of the perturbation, and a third loss for measuring the auditory inaudible property of the perturbation; the synthesized voice is synthesized based on the Mel spectrogram of the voice sample to be protected with the added perturbation and the text information; adding the target perturbation to the voice sample to be protected to obtain a non-learnable voice sample. Through the present application, a non-learnable voice sample can be generated through the updated target perturbation, so that the target voice synthesis model cannot learn the voice features of the speaker, thereby preventing voice cloning in the voice synthesis process.
[0097] Based on the same inventive concept, the embodiments of the present application also provide a voice anti-cloning device based on non-learnable voice samples corresponding to the voice anti-cloning method based on non-learnable voice samples. Since the principle of solving problems by the device in the embodiments of the present application is similar to that of the above-mentioned voice anti-cloning method based on non-learnable voice samples in the embodiments of the present application, the implementation of the device can refer to the implementation of the method, and the repeated parts will not be elaborated.
[0098] Refer to Figure 3 As shown, it is a schematic diagram of a voice anti-cloning device based on non-learnable voice samples provided by the embodiments of the present application. The device includes:
[0099] An acquisition module 301, configured to obtain a voice sample to be protected and the text information corresponding to the voice sample to be protected;
[0100] A perturbation update module 302, configured to update and iterate the perturbation added to the voice sample to be protected with the minimum of the objective function as the optimization target to obtain the target perturbation; wherein the objective function includes a first loss for measuring the distance between the voice sample to be protected and the synthesized voice, a second loss for measuring the voice feature hiding property of the perturbation, and a third loss for measuring the auditory inaudible property of the perturbation; the synthesized voice is synthesized based on the linear spectrogram of the voice sample to be protected with the added perturbation and the text information;
[0101] A sample generation module 303, configured to add the target perturbation to the voice sample to be protected to obtain a non-learnable voice sample.
[0102] In a possible implementation, the perturbation update module 302 is specifically configured to, for the current round of optimization process, input the mel spectrogram and text information of the protected speech sample with the perturbation added in the previous round into the speech synthesis model to obtain the synthesized speech of the current round; calculate the first loss of the current round according to the mel spectrogram of the protected speech sample and the mel spectrogram of the synthesized speech of the current round; calculate the second loss of the current round according to the mel spectrogram of the synthesized speech of the current round and the mel spectrogram of the random Gaussian distribution noise; if the third loss is enabled, calculate the third loss of the current round according to the protected speech sample with the perturbation added in the previous round and the protected speech sample; and update the perturbation of the previous round based on the first loss, the second loss, and the third loss of the current round to obtain the perturbation of the current round.
[0103] The embodiment of the present application provides a voice anti-cloning device based on an unlearnable speech sample. The device includes: an acquisition module 301, configured to acquire a protected speech sample and text information corresponding to the protected speech sample; a perturbation update module 302, configured to update and iterate the perturbation added to the protected speech sample with the minimum of the objective function as the optimization target to obtain the target perturbation; where the objective function includes a first loss for measuring the distance between the protected speech sample and the synthesized speech, a second loss for measuring the speech feature hiding characteristic of the perturbation, and a third loss for measuring the auditory inaudible characteristic of the perturbation; the synthesized speech is synthesized based on the linear spectrogram and text information of the protected speech sample with the perturbation added; a sample generation module 303, configured to add the target perturbation to the protected speech sample to obtain an unlearnable speech sample. Through the present application, an unlearnable speech sample can be generated through the updated target perturbation, so that the target speech synthesis model cannot learn the speech features of the speaker, thereby preventing voice cloning in the speech synthesis process.
[0104] As Figure 4 shown, an electronic device 400 provided by the embodiment of the present application includes: a processor 401, a memory 402, and a bus. The memory 402 stores machine-readable instructions executable by the processor 401. When the electronic device runs, the processor 401 communicates with the memory 402 through the bus, and the processor 401 executes the machine-readable instructions to perform the steps of the voice anti-cloning method based on the unlearnable speech sample as described above.
[0105] Specifically, the above-mentioned memory 402 and processor 401 can be general-purpose memory and processor, which are not specifically limited here. When the processor 401 runs the computer program stored in the memory 402, it can execute the voice anti-cloning method based on the unlearnable speech sample as described above.
[0106] Corresponding to the above voice anti-cloning method based on non-learnable voice samples, an embodiment of the present application further provides a computer-readable storage medium, on which a computer program is stored. When the computer program is run by a processor, it executes the steps of the above voice anti-cloning method based on non-learnable voice samples.
[0107] Those skilled in the art can clearly understand that for the convenience and brevity of description, the specific working processes of the systems and devices described above can refer to the corresponding processes in the method embodiments, which will not be elaborated herein. In the several embodiments provided in the present application, it should be understood that the disclosed systems, devices, and methods can be implemented in other ways. The device embodiments described above are only illustrative. For example, the division of the modules is only a logical function division, and there can be other division methods in actual implementation. For another example, multiple modules or components can be combined or integrated into another system, or some features can be ignored or not executed. Another point is that the displayed or discussed couplings or direct couplings or communication connections to each other can be through some communication interfaces. The indirect couplings or communication connections of the devices or modules can be electrical, mechanical, or other forms.
[0108] The modules described as separate components may or may not be physically separated. The components shown as modules may or may not be physical units, that is, they can be located in one place or distributed to multiple network units. Some or all of the units can be selected according to actual needs to achieve the purpose of the solution of this embodiment.
[0109] In addition, in each embodiment of the present application, the functional units can be integrated into one processing unit, or each unit can exist physically alone, or two or more units can be integrated into one unit.
[0110] If the function is implemented in the form of a software functional unit and sold or used as an independent product, it can be stored in a non-volatile computer-readable storage medium executable by a processor. Based on this understanding, the technical solution of the present application, in essence, or the part that contributes to the prior art, or a part of this technical solution, can be embodied in the form of a software product. The computer software product is stored in a storage medium and includes several instructions for causing a computer device (which can be a personal computer, a server, or a network device, etc.) to execute all or part of the steps of the information processing method described in each embodiment of the present application. The aforementioned storage medium includes: various media such as USB flash drives, mobile hard disks, ROM, RAM, magnetic disks, or optical discs that can store program codes.
[0111] The above are only specific embodiments of the present application, but the protection scope of the present application is not limited thereto. Any person skilled in the art within the technical scope disclosed by the present application can easily think of changes or substitutions, which should all be covered within the protection scope of the present application. Therefore, the protection scope of the present application shall be subject to the protection scope of the claims.
Claims
1. A voice anti-cloning method based on non-learnable voice samples, characterized in that, The method includes: Obtaining a speech sample to be protected and text information corresponding to the speech sample to be protected; Updating and iterating the perturbation added to the speech sample to be protected with the minimum of the objective function as the optimization goal to obtain the target perturbation, including: for the current round of optimization process, inputting the linear spectrogram of the speech sample to be protected with the perturbation of the previous round added and the text information into the alternative speech synthesis model to obtain the synthesized speech of the current round; calculating the first loss of the current round according to the Mel spectrogram of the speech sample to be protected and the Mel spectrogram of the synthesized speech of the current round; calculating the second loss of the current round according to the Mel spectrogram of the synthesized speech of the current round and the Mel spectrogram of the random Gaussian distribution noise; if the third loss is enabled, calculating the third loss of the current round according to the speech sample to be protected with the perturbation of the previous round added and the speech sample to be protected; and updating the perturbation of the previous round based on the first loss of the current round, the second loss of the current round, and the third loss of the current round to obtain the perturbation of the current round; wherein, the objective function includes a first loss for measuring the distance between the speech sample to be protected and the synthesized speech, a second loss for measuring the speech feature hiding property of the perturbation, and a third loss for measuring the auditory inaudible property of the perturbation; the synthesized speech is synthesized based on the linear spectrogram of the speech sample to be protected with the perturbation added and the text information; Adding the target perturbation to the speech sample to be protected to obtain an unlearnable speech sample.
2. The voice anti-cloning method based on non-learnable voice samples according to claim 1, characterized in that, The method further includes: If the third loss is not enabled, updating the perturbation of the previous round based on the first loss of the current round and the second loss of the current round to obtain the perturbation of the current round.
3. The voice anti-cloning method based on non-learnable voice samples according to claim 1, characterized in that, The calculating the first loss of the current round according to the Mel spectrogram of the speech sample to be protected and the Mel spectrogram of the synthesized speech of the current round includes: Substituting the Mel spectrogram of the speech sample to be protected and the Mel spectrogram of the synthesized speech of the current round into the following first loss function to obtain the first loss of the current round; ; Among them, is the first loss in this round, is the Mel spectrogram of the voice sample to be protected , is the Mel spectrogram of the synthesized voice in this round , is the voice sample to be protected and the synthesized voice in this round the distance between.
4. The voice anti-cloning method based on non-learnable voice samples according to claim 1, wherein The calculating the second loss of the current round according to the Mel spectrogram of the synthesized speech of the current round and the Mel spectrogram of the random Gaussian distribution noise includes: Substituting the Mel spectrogram of the synthesized speech of the current round and the Mel spectrogram of the random Gaussian distribution noise into the following second loss function to obtain the second loss of the current round; ; Among them, is the second loss in this round, is the Mel spectrogram of the synthesized speech in this round , is the Mel spectrogram of the random Gaussian distribution noise , is the KL divergence between the synthesized speech in this round and the random Gaussian distribution noise , is the distance between the synthesized speech in this round and the random Gaussian distribution noise .
5. The voice anti-cloning method based on non-learnable voice samples according to claim 1, characterized in that, The calculating the third loss of the current round according to the speech sample to be protected with the perturbation of the previous round added and the speech sample to be protected includes: Substituting the speech sample to be protected with the perturbation of the previous round added and the speech sample to be protected into the following third loss function to obtain the third loss of the current round; ; ; Among them, is the third loss in this round, is the short-term intelligibility index loss between the protected speech sample with the previous round of perturbation added and the protected speech sample ; is the short-time Fourier transform loss function, is the protected speech sample, is the previous round of perturbation, is the value obtained after the short-time Fourier transform of the protected speech sample with the previous round of perturbation added, is the value obtained after the short-time Fourier transform of the protected speech sample, is the value obtained after the short-time Fourier transform of the protected speech sample with the previous round of perturbation added and the protected speech sample after the short-time Fourier transform, distance.
6. A voice anti-cloning device based on non-learnable voice samples, characterized in that, The apparatus includes: An obtaining module, configured to obtain a speech sample to be protected and text information corresponding to the speech sample to be protected; A perturbation update module, which is used to update and iterate the perturbation added to the protected speech sample with the minimum of the objective function as the optimization goal to obtain the target perturbation, including: for the current round of optimization process, input the linear spectrogram of the protected speech sample with the previous round of perturbation added and the text information into the alternative speech synthesis model to obtain the synthesized speech of this round; calculate the first loss of this round according to the Mel spectrogram of the protected speech sample and the Mel spectrogram of the synthesized speech of this round; calculate the second loss of this round according to the Mel spectrogram of the synthesized speech of this round and the Mel spectrogram of the random Gaussian distribution noise; if the third loss is enabled, calculate the third loss of this round according to the protected speech sample with the previous round of perturbation added and the protected speech sample; and update the previous round of perturbation based on the first loss, the second loss and the third loss of this round to obtain the perturbation of this round; wherein, the objective function includes a first loss for measuring the distance between the protected speech sample and the synthesized speech, a second loss for measuring the speech feature hiding property of the perturbation, and a third loss for measuring the auditory inaudible property of the perturbation; the synthesized speech is synthesized based on the linear spectrogram of the protected speech sample with perturbation added and the text information. A sample generation module, which is used to add the target perturbation to the protected speech sample to obtain an unlearnable speech sample.
7. An electronic device, characterized in that, Including: A processor, a storage medium and a bus, the storage medium stores machine-readable instructions executable by the processor. When the electronic device runs, the processor communicates with the storage medium through the bus, and the processor executes the machine-readable instructions to perform the steps of the speech anti-cloning method based on unlearnable speech samples according to any one of claims 1 to 5.
8. A computer-readable storage medium, characterized in that, A computer program is stored on the computer-readable storage medium, and when the computer program is run by the processor, it performs the steps of the speech anti-cloning method based on unlearnable speech samples according to any one of claims 1 to 5.
Citation Information
Patent Citations
Black-box intelligent speech recognition system confrontation sample generation method and related device
CN116343759A