Anti-voice conversion watermark embedding method based on diffusion model and program product
By using a diffusion model-based anti-speech-switching watermarking embedding method, log-Mel spectrum and diffusion prior are generated, and model parameters are optimized. This solves the balance problem between imperceptibility and robustness in existing anti-speech-switching methods, and achieves efficient watermarking embedding and improved speech-switching capabilities.
Patent Information
- Application Number
- CN202511341537.9
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-09-19
- Publication Date
- 2025-10-28
- Estimated Expiration
- 2045-09-19
AI Technical Summary
Existing anti-speech conversion methods struggle to balance anti-speech conversion capability and robustness while ensuring imperceptibility. Traditional watermark embedding methods ignore the overall structure of the speech signal, resulting in watermark information being concentrated in the easily attenuated high-frequency part, and excessive watermark strength affecting imperceptibility.
A diffusion-based anti-speech conversion watermarking embedding method is adopted. By generating log-Mel spectrum and diffusion prior, and combining forward diffusion module, feature tensor generation module, residual feature extraction module and reverse diffusion module, the model parameters are optimized using the total loss function to generate watermarked speech. A frequency-weighted loss function under multi-scale short-time Fourier transform is designed to optimize the watermark embedding position.
It improves the anti-speech conversion performance and robustness of watermarked speech, ensuring the watermark's imperceptibility, while also enhancing the watermark's strength and anti-speech conversion capability, thus strengthening its robustness in complex transmission processes.
Smart Images

Figure CN120853589A_ABST
Abstract
Description
Technical Field
[0001] This invention belongs to the field of digital audio watermarking technology, specifically relating to a diffusion-based anti-speech conversion watermarking embedding method and program product. Background Technology
[0002] Deep speech generation technology has already achieved the ability to produce lifelike speech. While enriching people's entertainment and lives, it is also easily abused by criminals for voice spoofing, posing a significant threat to personal privacy and social security. Therefore, resisting and disrupting voice spoofing has become a field of practical significance. This technology, through special processing of the voice to be protected, destroys the conversion result of the target voice while ensuring imperceptibility, thus preventing harm at its source. In recent years, while this field has made rapid progress, it also faces many challenges and bottlenecks.
[0003] In 2021, Huang et al. proposed an iterative optimization method to embed watermarks into the frequency domain of the speech to be protected (Huang C, Lin YY, Lee H, et al. Defending your voice: Adversarial attack on voice conversion[C] / / Proceedings of the 2021 IEEE Spoken Language Technology Workshop (SLT). Shenzhen, China: IEEE, 2021: 552-559.), pioneering a new direction of resisting speech conversion by disrupting the output. In 2023, Li et al. proposed embedding imperceptible adversarial watermarking in the time domain of speech (Li J, Ye D, Tang L, et al. VoiceGuard: Protecting voice privacy with strong and imperceptible adversarial perturbation in the time domain[C] / / Proceedingsof the 32nd International Joint Conference on Artificial Intelligence (IJCAI-23). Macau, China: IJCAI, 2023: 4812-4820.), avoiding the reconstruction loss caused by time-frequency conversion of the speech to be protected, and achieving effective destruction of the conversion result while improving imperceptibility. Both of the above methods generate watermarks iteratively, which suffers from slow inference speed. In 2024, Dong et al. proposed a method to directly generate adversarial watermarks using a GAN generator and then embed them into the speech to be protected (Dong S, Chen B, Ma K, et al. Active defense against voice conversion through generative adversarial network[J]. IEEE Signal Processing Letters, 2024, 31: 706-710.). They also designed a SWCSM network to reconstruct waveforms from spectrograms, which accelerated the inference speed while ensuring the ability to destroy the conversion results and the imperceptibility of the speech to be protected.However, all three methods mentioned above embed watermarks in the speech to be protected. This method has two problems: (1) Existing methods are difficult to simultaneously take into account the anti-speech conversion capability and the imperceptibility of speech samples. In order to improve the practicality of adversarial samples, compromises are often needed on the anti-speech conversion capability; (2) Watermark information is easily erased in the complex propagation process, which makes the robustness of this method need to be improved when facing attacks such as compression and filtering.
[0004] Based on the above analysis, the existing anti-speech conversion methods are still difficult to guarantee the success rate and robustness of resisting speech conversion while ensuring imperceptibility. The reasons are as follows: (1) Traditional watermark embedding methods only make local modifications to the waveform or spectrum, ignoring the overall structure of the speech signal, and causing the watermark information to be concentrated in the high-frequency part that is easy to attenuate; (2) In the above methods, the watermark is an additional part, and its excessive watermark strength will lead to a decrease in imperceptibility, thereby reducing the imperceptibility of the protected speech, and it is difficult to achieve a balance between anti-speech conversion ability and imperceptibility. Summary of the Invention
[0005] To address the aforementioned issues, this invention proposes a diffusion-based method and program for embedding watermarks against speech conversion, which can improve the anti-speech conversion performance and robustness of watermarked speech while ensuring the watermark's imperceptibility.
[0006] To achieve the above-mentioned technical objectives and effects, the present invention is implemented through the following technical solution:
[0007] In a first aspect, the present invention provides a diffusion-based method for embedding anti-speech conversion watermarks, comprising:
[0008] The acquired real-time carrier speech is input into a trained diffusion model-based anti-speech-transformation watermarking embedding network to generate watermarked speech.
[0009] The training method for the diffusion-based anti-speech-transformation watermark embedding network includes:
[0010] The received historical carrier speech is preprocessed to generate a log-Mel spectrum and a diffusion prior;
[0011] The historical carrier speech, log-Mel spectrum, and diffusion prior are fed into a diffusion-based anti-speech-transformation watermarking embedding network to generate watermarked speech.
[0012] With the goal of minimizing the total loss function, the model parameters of the anti-speech-switching watermarking embedding network are continuously updated to complete the training of the anti-speech-switching watermarking embedding network. The total loss function includes the adversarial loss function between the speech conversion result of the carrier speech and the conversion result of the watermarked speech, the diffusion loss function between the predicted noise and the actual added noise, and the frequency-weighted loss function of the watermarked speech and the carrier speech under multi-scale short-time Fourier transform.
[0013] In conjunction with the first aspect, optionally, the method for generating the diffusion prior includes:
[0014] The carrier speech is segmented and windowed, and the short-time Fourier transform is performed on each frame of carrier speech after segmentation and windowing to obtain the amplitude spectrum. The amplitude spectrum is mapped through the Mel filter bank and the logarithm is taken to generate the corresponding log-Mel spectrum. The log-Mel spectrum is also used as conditional information.
[0015] Frame-level energy is calculated frame by frame from the log-Mel spectrum to obtain a frame-level energy sequence;
[0016] Based on the frame-level energy sequence, a diagonal covariance matrix is constructed in chronological order.
[0017] The spectral flatness of the carrier speech is calculated, and a suitable Gaussian noise is determined based on the relationship between the spectral flatness and a preset threshold. The diffusion prior is then obtained by combining the covariance matrix.
[0018] In conjunction with the first aspect, optionally, the formula for calculating the spectral flatness is:
[0019] ,
[0020] In the formula, For spectral flatness, Total number of frequency bands For the first Spectral values of each frequency band A positive number less than a set threshold. Represents the logarithmic function with base e;
[0021] when In this case, super-Gaussian noise should be selected;
[0022] when In this case, standard Gaussian noise should be selected;
[0023] when In this case, sub-Gaussian noise should be selected.
[0024] In conjunction with the first aspect, optionally, the anti-speech conversion watermarking embedding network includes:
[0025] The forward diffusion module is used to randomly select diffusion time steps and inject Gaussian noise into the carrier speech according to the noise schedule to generate noisy speech;
[0026] The feature tensor generation module is used to perform scale unification operations on noisy speech, log-Mel spectrum and diffusion time step, and add them in the feature channels to form a uniformly modulated feature tensor.
[0027] The residual feature extraction module is used to perform residual feature extraction processing on the feature tensor to obtain prediction noise;
[0028] The reverse diffusion module is used to perform reverse diffusion updates on noisy speech based on predicted noise and diffusion priors to generate watermarked speech.
[0029] In conjunction with the first aspect, optionally, the residual feature extraction module includes several sequentially arranged dilated residual blocks, with the residual components of the previous dilated residual block being output to the next dilated residual block; the Skip components of all dilated residual blocks are added together and averaged, then dimensionality is reduced by 1×1 convolution and ReLU activation is performed to obtain the prediction noise.
[0030] In conjunction with the first aspect, optionally, the mathematical expression of the total loss function is:
[0031] ,
[0032] In the formula, For the total loss function, Let be the adversarial loss function between the speech conversion result of the carrier speech and the conversion result of the watermarked speech. To predict the diffusion loss function between the noise and the actual added noise, Let be the frequency-weighted loss function of the watermarked speech and the carrier speech under multi-scale short-time Fourier transform. , , These are the weights for the adversarial loss function, the diffusion loss function, and the frequency-weighted loss function, respectively.
[0033] In conjunction with the first aspect, optionally, the diffusion loss function between the predicted noise and the actual added noise... The mathematical expression is:
[0034]
[0035] In the formula, and These represent the actual injected Gaussian noise and the predicted noise, respectively. It is the inverse of the covariance matrix. Indicates time step , This represents conditional information, which is a log-Mel spectrum. Represents the carrier's speech. This indicates noisy speech. This indicates the diffusion time step t and the carrier speech. Actual injected Gaussian noise Take the expected value together. This represents a weighted L2 norm.
[0036] In conjunction with the first aspect, optionally, the difference in waveform between the carrier speech and the watermarked speech can be denoted as the perturbation. The waveform difference is divided into multiple sub-bands by short-time Fourier transform and transformed to the frequency domain to obtain the amplitude spectrum of the corresponding sub-bands. ;
[0037] Based on the sensing weights of each frequency band and the amplitude spectrum of the sub-band Generate a frequency-weighted loss function under single-scale short-time Fourier transform. ;
[0038] Based on the frequency-weighted loss functions under all single-scale short-time Fourier transforms, frequency-weighted loss functions for watermarked speech and carrier speech under multi-scale short-time Fourier transforms are generated. The frequency-weighted loss function under the single-scale short-time Fourier transform Frequency-weighted loss function under multi-scale short-time Fourier transform The mathematical expressions are as follows:
[0039] ,
[0040] ,
[0041] In the formula, For the first The center frequency of each frequency band The total number of frequency bands. ∈[0,1] is the first The perceptual weight of each frequency band is assigned; a higher weight value indicates that the human ear is more sensitive to that frequency band. Indicates the first The weighted weights corresponding to each scale Indicates the first Each scale.
[0042] In conjunction with the first aspect, optionally, the adversarial loss function between the speech conversion result of the carrier speech and the conversion result of the watermarked speech... The mathematical expression is:
[0043] ,
[0044] In the formula, This represents the Mel spectrum generated after the carrier speech and source speech are input into the speech conversion model for speech conversion. This represents the Mel spectrum generated after the watermarked speech and the source speech are input into the speech conversion model for speech conversion. express Norm.
[0045] In a second aspect, the present invention provides a computer program product, including a computer program / instruction that, when executed by a processor, implements the diffusion-based anti-speech conversion watermarking embedding method described in any one aspect.
[0046] Compared with the prior art, the beneficial effects of the present invention are as follows:
[0047] This invention provides a diffusion-based method and program product for resisting speech conversion watermarking, which can improve the resistance to speech conversion and robustness of watermarked speech while ensuring the imperceptibility of the watermark.
[0048] Furthermore, this invention leverages the powerful generative capabilities of the diffusion model and designs a diffusion prior determined by the input data, unlike traditional diffusion models. Compared to existing methods, the watermarked speech generated by this invention has higher quality and can improve the watermark strength while maintaining the imperceptibility of the watermarked speech. At the same time, a frequency-weighted loss function under multi-scale short-time Fourier transform is designed. By focusing the watermark on frequency bands that are not sensitive to the human ear, the imperceptibility of the watermarked speech is effectively improved, indirectly enhancing the anti-speech conversion ability of the watermarked speech. Attached Figure Description
[0049] To more clearly illustrate the technical solutions in the embodiments of the present invention or the prior art, the drawings used in the embodiments will be briefly described below. Obviously, the drawings described below are only some embodiments of the present invention. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort, wherein:
[0050] Figure 1 This is a flowchart illustrating an embodiment of the anti-speech conversion watermarking embedding method based on a diffusion model according to the present invention.
[0051] Figure 2 This is a training framework diagram of an anti-speech conversion watermarking embedding network according to an embodiment of the present invention;
[0052] Figure 3 This is an application framework diagram of an anti-speech conversion watermarking embedding network according to an embodiment of the present invention. Detailed Implementation
[0053] The technical solutions of the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only a part of the embodiments of the present invention, and not all of them. Based on the embodiments of the present invention, all other embodiments obtained by those of ordinary skill in the art without creative effort are within the scope of protection of the present invention.
[0054] Furthermore, if the embodiments of this invention involve descriptions such as "first" or "second," these descriptions are for descriptive purposes only and should not be construed as indicating or implying their relative importance or implicitly specifying the number of technical features indicated. Therefore, a feature defined with "first" or "second" may explicitly or implicitly include at least one of those features. Additionally, the technical solutions of the various embodiments can be combined with each other, but this must be based on the ability of those skilled in the art to implement them. If the combination of technical solutions is contradictory or impossible to implement, it should be considered that such a combination of technical solutions does not exist and is not within the scope of protection claimed by this invention.
[0055] Example 1
[0056] This invention provides a diffusion-based method for embedding anti-speech conversion watermarks, comprising the following steps:
[0057] The acquired real-time carrier speech is input into a trained diffusion model-based anti-speech-transformation watermarking embedding network to generate watermarked speech.
[0058] The training method for the diffusion-based anti-speech-transformation watermark embedding network includes:
[0059] The received historical carrier speech is preprocessed to generate a log-Mel spectrum and diffusion prior; in the specific implementation process, such as... Figure 2 As shown, the preprocessing module preprocesses the received historical carrier speech to generate a log-Mel spectrum (i.e., conditional information). ) and diffusion prior P(0, );
[0060] The historical carrier speech, log-Mel spectrum, and diffusion prior are fed into a diffusion-based anti-speech-transformation watermarking embedding network to generate watermarked speech.
[0061] With the goal of minimizing the total loss function, the model parameters of the anti-speech-switching watermarking embedding network are continuously updated to complete the training of the anti-speech-switching watermarking embedding network. The total loss function includes the adversarial loss function between the speech conversion result of the carrier speech and the conversion result of the watermarked speech, the diffusion loss function between the predicted noise and the actual added noise, and the frequency-weighted loss function of the watermarked speech and the carrier speech under multi-scale short-time Fourier transform.
[0062] Based on the above scheme, it is possible to improve the anti-speech conversion performance and robustness of watermarked speech while ensuring the imperceptibility of the watermark.
[0063] In one specific embodiment of the present invention, the method for generating the diffusion prior includes:
[0064] The carrier speech is segmented and windowed, and the short-time Fourier transform is performed on each frame of carrier speech after segmentation and windowing to obtain the amplitude spectrum. The amplitude spectrum is mapped through the Mel filter bank and the logarithm is taken to generate the corresponding log-Mel spectrum. The log-Mel spectrum is also used as conditional information.
[0065] Frame-level energy is calculated frame by frame from the log-Mel spectrum to obtain a frame-level energy sequence;
[0066] Based on the frame-level energy sequence, a diagonal covariance matrix is constructed in chronological order. ;
[0067] The spectral flatness of the carrier speech is calculated, and a suitable Gaussian noise is determined based on the relationship between spectral flatness and a preset threshold. This noise is then combined with the covariance matrix. The diffusion prior P(0, ).
[0068] In one specific embodiment of the present invention, the formula for calculating the spectral flatness is:
[0069] ,
[0070] In the formula, For spectral flatness, This represents the total number of frequency bands. For the first Spectral values of each frequency band A positive number less than a set threshold. Represents the logarithmic function with base e;
[0071] when In this case, super-Gaussian noise should be selected;
[0072] when In this case, standard Gaussian noise should be selected;
[0073] when In this case, sub-Gaussian noise should be selected.
[0074] Unlike traditional diffusion models that uniformly use standard Gaussian noise as the diffusion prior, based on the above scheme, a generalized diffusion prior P(0, This design allows the forward diffusion process to more closely resemble the natural distribution of the input data, laying the foundation for subsequent fast and accurate noise reduction.
[0075] In one specific embodiment of the present invention, the anti-speech conversion watermarking embedding network includes:
[0076] The forward diffusion module is used to randomly select diffusion time steps and inject Gaussian noise into the carrier speech according to the noise schedule to generate noisy speech;
[0077] The feature tensor generation module is used to perform scale unification operations on noisy speech, log-Mel spectrum, and diffusion time step, and add them in the feature channels to form a unified modulated feature tensor. In specific implementation, the feature tensor generation module includes a two-layer spectrum upsampler and a diffusion time step embedder. The two-layer spectrum upsampler is used to upsample the received log-Mel spectrum, and the diffusion time step embedder is used to map the diffusion time step to the embedding vector of the time step. The outputs of the two are added in the channel dimension to the residual features obtained by 1×1 convolution and ReLU processing of the noisy speech corresponding to the current diffusion time step, so as to achieve unified modulation of the diffusion time step, Mel spectrum, and noisy speech.
[0078] The residual feature extraction module is used to perform residual feature extraction processing on the feature tensor to obtain prediction noise. ;
[0079] Backdiffusion module, used for prediction noise-based and diffusion prior P(0, The noisy speech is updated by reverse diffusion to generate watermarked speech.
[0080] In the above scheme, the formula used for reverse diffusion is:
[0081] = [ ]+ , ~ P( ),
[0082] In the formula, Indicates diffusion time step The corresponding noisy speech, Indicates diffusion time step The corresponding noisy speech, and Represents the noise scheduling parameters during the diffusion process. Indicates noisy speech diffusion time step The predicted noise obtained from the conditional information c, P is the Gaussian noise P that is followed during the random sampling process, used to ensure diversity and the random diffusion characteristics of the model.
[0083] In one specific embodiment of the present invention, the residual feature extraction module includes several dilated residual blocks arranged in sequence. The residual components of the previous dilated residual block are output to the next dilated residual block. The Skip components of all dilated residual blocks are added together and averaged. The dimensionality is reduced by 1×1 convolution and ReLU activation is performed to obtain the prediction noise.
[0084] In the above scheme, the number of dilated residual blocks can be set to 30 or other values during implementation, depending on actual needs. Each dilated residual block includes a 3×1 dilated convolution and a gated activation unit, and then the residual feature is mapped into residual components and Skip components through 1×1 channel projection.
[0085] In one specific embodiment of the present invention, the mathematical expression of the total loss function is:
[0086] ,
[0087] In the formula, For the total loss function, Let be the adversarial loss function between the speech conversion result of the carrier speech and the conversion result of the watermarked speech. To predict the diffusion loss function between the noise and the actual added noise, Let be the frequency-weighted loss function of the watermarked speech and the carrier speech under multi-scale short-time Fourier transform. , , These are the weights for the adversarial loss function, the diffusion loss function, and the frequency-weighted loss function, respectively.
[0088] In one specific embodiment of the present invention, the diffusion loss function between the predicted noise and the actual added noise The mathematical expression is:
[0089]
[0090] In the formula, and These represent the actual injected Gaussian noise and the predicted noise, respectively. It is the inverse of the covariance matrix. Indicates time step , This represents conditional information, which is a log-Mel spectrum. Represents the carrier's speech. This indicates noisy speech. This indicates the diffusion time step t and the carrier speech. Actual injected Gaussian noise Take the expected value together. Represents a weighted norm 2;
[0091] In one specific embodiment of the present invention, the difference in waveform between the carrier speech and the watermarked speech is denoted as the disturbance. The waveform difference is divided into multiple sub-bands by short-time Fourier transform and transformed to the frequency domain to obtain the amplitude spectrum of the corresponding sub-bands. ;
[0092] Based on the sensing weights of each frequency band and the amplitude spectrum of the sub-band Generate a frequency-weighted loss function under single-scale short-time Fourier transform. ;
[0093] Based on the frequency-weighted loss functions under all single-scale short-time Fourier transforms, frequency-weighted loss functions for watermarked speech and carrier speech under multi-scale short-time Fourier transforms are generated. The frequency-weighted loss function under the single-scale short-time Fourier transform Frequency-weighted loss function under multi-scale short-time Fourier transform The mathematical expressions are as follows:
[0094] ,
[0095] ,
[0096] In the formula, For the first The center frequency of each frequency band The total number of frequency bands. ∈[0,1] is the first The perceptual weight of each frequency band is assigned; a higher weight value indicates that the human ear is more sensitive to that frequency band. Indicates the first The weighted weights corresponding to each scale Indicates the first Each scale.
[0097] In one specific embodiment of the present invention, the adversarial loss function between the speech conversion result of the carrier speech and the conversion result of the watermarked speech is... The mathematical expression is:
[0098] ,
[0099] In the formula, This represents the Mel spectrum generated after the carrier speech and source speech are input into the speech conversion model for speech conversion. This represents the Mel spectrum generated after the watermarked speech and the source speech are input into the speech conversion model for speech conversion. express Norm.
[0100] In the above scheme, during implementation, one or more of four speech conversion models can be selected: Adain-VC, AgainVC, TriAAN, and Vqvc+. All four models have an output sampling rate of 22kHz, and the front end uniformly uses a 1024-point STFT (Short Time Fourier Transform) to map the received speech waveform into a log-Mel spectrum of a predetermined size (e.g., 80×128). All models use officially pre-trained weights. During training, the carrier speech, the watermarked speech, and the corresponding source speech are input into the four models respectively to complete the speech conversion, thus simulating a real speech conversion attack scenario.
[0101] The following describes in detail the diffusion-based anti-speech conversion watermarking embedding method in this invention embodiment with reference to a specific implementation method.
[0102] like Figure 1 As shown, the diffusion model-based anti-speech conversion watermark embedding method includes the following steps:
[0103] Step 1: Read the audio from the historical recording;
[0104] Step 2: Initialize the speech-resistant watermarking embedding network, the speech-resistant watermarking embedding network (i.e. Figure 2 During training, the watermark embedding network needs to be integrated with the attack layer (i.e., Figure 2 The model parameters of the anti-speech conversion watermarking embedding network are adjusted in conjunction with the VC model in the anti-speech conversion watermarking embedding network. Specifically, this refers to the adjustment of the model parameters of the noise prediction network (i.e., the back diffusion network) in the anti-speech conversion watermarking embedding network.
[0105] Step 3: Set the weights of the adversarial loss function between the speech conversion result of the carrier speech and the conversion result of the watermarked speech in the anti-speech conversion watermarking embedding network, the diffusion loss function between the predicted noise and the actual added noise, and the frequency weighted loss function of the watermarked speech and the carrier speech under multi-scale short-time Fourier transform.
[0106] Step 4: Read the source speech that matches the carrier speech and input it into the attack layer for training the anti-speech-switching watermarking embedding network. Use historical carrier speech and corresponding source speech as training data, dividing them into training speech datasets and test speech datasets. The training speech dataset is used for training the anti-speech-switching watermarking embedding network, and the test speech dataset is used for testing the anti-speech-switching watermarking embedding network. The specific training process of the anti-speech-switching watermarking embedding network is as follows: Figure 2 As shown;
[0107] Step 5: Calculate the total loss function, which includes the adversarial loss function between the speech conversion result of the carrier speech and the conversion result of the watermarked speech, the diffusion loss function between the predicted noise and the actual added noise, and the frequency weighted loss function of the watermarked speech and the carrier speech under multi-scale short-time Fourier transform. The network parameters of the anti-speech conversion watermarking embedding network are adjusted and optimized with the goal of minimizing the total loss function.
[0108] Step 6: Read the real-time carrier speech, input it into the pre-trained anti-speech conversion watermarking embedding network, calculate and output the watermarked speech.
[0109] Further explanation of step 2, details of the anti-speech-switching watermark embedding network and attack layer:
[0110] 2-1: Details of the anti-speech-transfer watermarking embedding network are shown in Table 1 below, where b is the batch size:
[0111] Table 1. Detailed configuration of network parameters for the anti-speech conversion watermarking embedding network.
[0112]
[0113] 2-2: The attack layer contains four speech conversion models (i.e. Figure 2 The four speech conversion models (Adain-VC, AgainVC, TriAAN, and Vqvc+) all use a 22kHz sampling rate for their conversion outputs. A 1024-point STFT is used in the front end to map the received speech signal to an 80×128 log-Mel spectrum. During the prediction phase, the input data is clipped into 128-frame windows, and the amplitude is normalized to the [-1,1] interval before being fed into their respective encoders. All models use officially pre-trained weights, and the parameters are frozen during the prediction phase. During training, the carrier speech, the watermarked speech, and the corresponding source speech are input into the four models to complete speaker conversion, thus simulating real speech conversion attack scenarios. This allows the anti-speech conversion watermarking embedding network to adapt to these attacks during training and generate watermarked speech that resists them.
[0114] Specifically, the differentiable surrogate model Again-VC was used to simulate the conversion process of the speech conversion model, and the specific network parameter configuration is shown in Table 2. Each encoder convolutional layer contains two 3×1 Convs, and each decoder convolutional layer contains four 3×1 Convs. After the decoder output, the features are first passed through a two-layer GRU to capture long-range temporal dependencies, and then reduced to the target number of channels by a 1×1 linear projection layer. The output of the differentiable surrogate model Again-VC is a converted Mel spectrum of 1×80×128. After inverse MelGAN transformation, the speech waveform is obtained with a size of 1×32768 and a sampling rate of 22kHz.
[0115] Table 2. Detailed Network Parameter Configuration for the Differentiable Agent Model Again-VC
[0116]
[0117] Further explanation of the specific training process in step 4:
[0118] 4-1: First, read 1×32768 single-channel carrier speech from the file list. For carrier speech Frame segmentation and windowing are performed, and short-time Fourier transforms are applied to the speech in each frame after frame segmentation and windowing to obtain the amplitude spectrum. The amplitude spectrum is then mapped through a Mel filter bank and its logarithm is taken to generate an 80×128 log-Mel spectrum, which is used as conditional information c. Frame-level energy is calculated frame by frame for the log-Mel spectrum to obtain a frame-level energy sequence. Based on the frame-level energy sequence, a diagonal covariance matrix is constructed in chronological order. .
[0119] Furthermore, computing the speech of the carrier. spectral flatness :
[0120] ,
[0121] In the formula, For spectral flatness, This represents the total number of frequency bands. For the first Spectral values of each frequency band For positive numbers less than a set threshold, in the specific implementation process, It is a very small positive number.
[0122] Will Divided into three levels, when When using super-Gaussian noise as the diffusion prior, when When standard Gaussian noise is used as the diffusion prior, when In this case, sub-Gaussian noise is used as the diffusion prior. This yields a generalized diffusion prior P(0, ..., depending on the input data.) This design makes the forward diffusion process more closely resemble the natural distribution of the input data, laying the foundation for subsequent fast and accurate noise reduction.
[0123] Subsequently, time step t is randomly selected, and the speech is transmitted to the carrier according to the noise schedule. Injecting Gaussian noise (i.e., performing forward noise addition) yields noisy speech. To ensure conditional alignment, the log-Mel spectrum is upsampled to 80×32768 via two stages of transposed convolution. Frame-level energy is repeatedly expanded into a 1×32768 prior vector by sampling 256 times per step. A randomly selected diffusion time step t is first obtained by sine and cosine positional encoding to obtain a 128-dimensional vector, then sequentially mapped to a 512-dimensional embedding vector for the time step through two 1×1 fully connected layers (equivalent to 1×1 Conv), and then copied along the time domain to 512×32768. The above two channels of information are then superimposed in the channel dimension and combined with the noisy speech expanded to 64×32768 via 1×1 convolution. The components are summed to form a unified modulation tensor, which is then fed into a 30-layer dilated residual block. The Skip components (i.e., skip connections) output by each dilated residual block are summed by channel and averaged, then reduced to 1 × 32768 by 1 × 1 convolution and ReLU activation, yielding the prediction noise (i.e., noise estimation) for this step. Using this predicted noise and diffusion prior P(0, ), for noisy speech Perform sequential reverse diffusion update, which reconstructs the denoised speech point by point in the time domain, and directly outputs the final watermarked speech with a size of 1×32768. Finally, regarding the watermarked audio... The waveform amplitude is clipped to [-1, 1].
[0124] 4-2: During the attack phase, the watermarked audio will be... With the corresponding carrier speech Speech conversion models (Adain-VC, AgainVC, TriAAN, Vqvc+) were pre-trained using the target speech and the same paired source speech as inputs. Both input speech pairs consisted of speech waveforms with a size of 1×32768 and a sampling rate of 22kHz. Ideally, the converted speech output from the source speech after conversion between the carrier speech and the source speech would have a Mel-spectrum size of 80×128. However, the converted speech outputting the watermarked speech after conversion with the source speech, but with a corrupted size, also had a Mel-spectrum size of 80×128.
[0125] Further explanation of the calculation of the total loss function in step 5:
[0126] 5-1 To improve the accuracy of the anti-speech-switching watermark embedding network in predicting noise during the denoising process, and thus improve the imperceptibility of the watermarked speech, this embodiment introduces a diffusion loss function between the predicted noise and the actual added noise. :
[0127]
[0128] In the formula, and These represent the actual injected Gaussian noise and the predicted noise, respectively. It is the inverse of the covariance matrix. Indicates time step , This represents conditional information, which is a log-Mel spectrum. Represents the carrier's speech. This indicates noisy speech. This indicates the diffusion time step t and the carrier speech. Actual injected Gaussian noise Take the expected value together. This represents the weighted L2 norm. The diffusion loss function improves the quality of generated samples by training the diffusion model to accurately predict the noise to be removed in each denoising step by minimizing the difference between the predicted noise and the actual injected Gaussian noise (i.e., the real noise).
[0129] 5-2 In order to concentrate the watermark information in a region where the human ear is less sensitive, thereby improving the imperceptibility of the watermarked speech, the present invention first introduces a frequency-weighted loss function under single-scale short-time Fourier transform. :
[0130]
[0131] In the formula, For the first The center frequency of each frequency band ∈[0,1] is the first The perceptual weights for each frequency band are assigned, with higher weight values indicating less sensitivity of the human ear to that band. The anti-speech-switching watermark embedding network will preferentially embed the watermark information into that frequency band. To avoid the problem of singular focus caused by single-scale short-time Fourier transform, this implementation proposes calculating frequency-weighted losses at three different scales (corresponding to large, medium, and small scales, respectively). This is achieved by introducing a lightweight ScaleAttention network with soft attention weights. Adaptive fusion of different scales yields a frequency-weighted loss function under multi-scale short-time Fourier transform. :
[0132] ,
[0133] In the formula, Indicates the first The weighted weights corresponding to each scale Indicates the first Each scale.
[0134] 5-3 To enhance the anti-speech conversion capability of watermarked speech, this embodiment introduces an adversarial loss function between the speech conversion result of the carrier speech and the conversion result of the watermarked speech. :
[0135] ,
[0136] In the formula, in the formula, This represents the Mel spectrum generated after the carrier speech and source speech are input into the speech conversion model for speech conversion. This represents the Mel spectrum generated after the watermarked speech and the source speech are input into the speech conversion model for speech conversion. express The norm is used in this process, where the source speech for the paired carrier speech and watermarked speech is the same. This loss maximizes the difference between the conversion results, thereby improving the anti-speech conversion capability of the watermarked speech generated by this method.
[0137] 5-4 In summary, the total loss function This includes an adversarial loss function between the speech conversion result of the carrier speech and the conversion result of the watermarked speech. The diffusion loss function between predicted noise and actual added noise Frequency-weighted loss function of watermarked speech and carrier speech under multi-scale short-time Fourier transform To balance the imperceptibility of the watermark and the speech-to-speech resistance of the watermarked speech, different weights are assigned to the three elements, as follows:
[0138]
[0139] In the formula, For the total loss function, Let be the adversarial loss function between the speech conversion result of the carrier speech and the conversion result of the watermarked speech. To predict the diffusion loss function between the noise and the actual added noise, Let be the frequency-weighted loss function of the watermarked speech and the carrier speech under multi-scale short-time Fourier transform. , , These are the weights for the adversarial loss function, the diffusion loss function, and the frequency-weighted loss function, respectively.
[0140] Ultimately, the Adam optimizer was used to adjust only the relevant parameters of the speech-resistant watermark embedding network.
[0141] Step 6 will be further explained in detail, and analyzed from the perspectives of invisibility results and attack robustness results:
[0142] 6-1: The trained anti-speech-to-text watermark embedding network is tested on a pre-defined dataset, using SVA. resist As an indicator of a watermark's ability to resist speech conversion, this indicator represents the proportion of speech pairs in the test dataset where the corrupted converted speech and the carrier speech are identified by the system as belonging to the same speaker, out of all speech pairs. A smaller indicator indicates better resistance to speech conversion. SVA quality As an indicator of watermark imperceptibility, this indicator represents the proportion of speech pairs in the test dataset where the watermarked speech and the carrier speech are judged to be from the same speaker to the total number of speech pairs. The larger the indicator, the better the imperceptibility. Table 3 below shows the test results of the watermarked speech generated by this method in terms of resistance to speech conversion and imperceptibility under different proxy models as the attack layer. It can be seen that the method in the embodiment of the present invention performs well in both of these aspects. This is mainly because: (1) In the embodiment of the present invention, the diffusion model is used to directly generate the watermarked speech after embedding the watermark. As a generator with excellent generation effect, the diffusion model's reconstruction result is closer to the natural distribution of the original speech, which alleviates the decrease in imperceptibility caused by excessive watermark information energy. Moreover, this process is performed directly on the waveform diagram, avoiding the quality loss that occurs during the reconstruction of the waveform diagram from the Mel spectrogram. (2) The frequency weighted loss function used in the embodiment of the present invention guides the watermark to be added to a frequency band that is not sensitive to the human ear, thereby reducing the impact of the addition of the watermark on the imperceptibility of the watermarked speech.
[0143] Table 3. Resistance to speech conversion and imperceptibility of the generated watermarked speech.
[0144]
[0145] 6-2: To test the robustness of the watermarked speech generated by the method in this embodiment, the anti-speech conversion capability of the watermarked speech was tested under MP3 compression and Gaussian filtering scenarios. The MP3 compression was divided into two scenarios: from 352kbps to 128kbps and from 352kbps to 64kbps. See Table 4 below:
[0146] Table 4. Robustness of the methods in the embodiments of the present invention under different attack scenarios
[0147]
[0148] The robustness of the method in the embodiments of the present invention under different attack scenarios is shown in Table 4. It can be seen that the watermarked speech still maintains good anti-speech conversion ability and imperceptibility when facing attacks. This is mainly because most of these attack scenarios are disrupted in the high-frequency part, while the method reconstructs the watermarked speech through a diffusion model. This process ensures that the watermark appears in various frequency bands of the speech sample, and the anti-speech conversion ability will not be greatly erased.
[0149] Example 2
[0150] Based on the same inventive concept as in Embodiment 1, this embodiment of the invention provides a computer program product, including a computer program / instruction, which, when executed by a processor, implements the anti-speech conversion watermarking embedding method based on the diffusion model as described in Embodiment 1.
[0151] Those skilled in the art will understand that embodiments of this application can be provided as methods, systems, or computer program products. Therefore, this application can take the form of a completely hardware embodiment, a completely software embodiment, or an embodiment combining software and hardware aspects. Furthermore, this application can take the form of a computer program product embodied on one or more computer-usable storage media (including but not limited to disk storage, CD-ROM, optical storage, etc.) containing computer-usable program code.
[0152] This application is described with reference to flowchart illustrations and / or block diagrams of methods, apparatus (systems), and computer program products according to embodiments of this application. It will be understood that each block of the flowchart illustrations and / or block diagrams, and combinations of blocks in the flowchart illustrations and / or block diagrams, can be implemented by computer program instructions. These computer program instructions can be provided to a processor of a general-purpose computer, special-purpose computer, embedded processor, or other programmable data processing apparatus to produce a machine, such that the instructions, which execute via the processor of the computer or other programmable data processing apparatus, generate instructions for implementing the flowchart... Figure 1 a process or multiple processes and / or boxes Figure 1 A device that provides the functions specified in one or more boxes.
[0153] These computer program instructions may also be stored in a computer readable memory that can direct a computer or other programmable data processing device to work in a specific manner, so that the instructions stored in the computer readable memory produce an article of manufacture comprising an instruction device, which implements the process Figure 1 a process or multiple processes and / or boxes Figure 1 The function specified in one or more boxes.
[0154] These computer program instructions can also be loaded onto a computer or other programmable data processing device so that a series of operational steps are executed on the computer or other programmable device to produce a computer-implemented process, thereby providing the instructions executed on the computer or other programmable device for implementing the process. Figure 1 a process or multiple processes and / or boxes Figure 1 A step that specifies a function in one or more boxes.
[0155] The embodiments of the present invention have been described above with reference to the accompanying drawings. However, the present invention is not limited to the specific embodiments described above. The specific embodiments described above are merely illustrative and not restrictive. Those skilled in the art can make many other forms under the guidance of the present invention without departing from the spirit and scope of the claims. All of these forms are within the protection scope of the present invention.
[0156] The basic principles, main features, and advantages of the present invention are shown and described above. Those skilled in the art should understand that the present invention is not limited to the above embodiments. The above embodiments and descriptions are merely illustrative of the principles of the present invention. Various changes and modifications may be made to the present invention without departing from the spirit and scope of the present invention. Such changes and modifications are intended to fall within the scope of the present invention. The scope of protection claimed in the present invention is defined by the appended claims and their equivalents.
Claims
1. A diffusion-based method for embedding watermarks to resist speech conversion, characterized in that, include: The acquired real-time carrier speech is input into a trained diffusion model-based anti-speech-transformation watermarking embedding network to generate watermarked speech. The training method for the diffusion-based anti-speech-transformation watermark embedding network includes: The received historical carrier speech is preprocessed to generate a log-Mel spectrum and a diffusion prior; The historical carrier speech, log-Mel spectrum, and diffusion prior are fed into a diffusion-based anti-speech-transformation watermarking embedding network to generate watermarked speech. With the goal of minimizing the total loss function, the model parameters of the anti-speech-switching watermarking embedding network are continuously updated to complete the training of the anti-speech-switching watermarking embedding network. The total loss function includes the adversarial loss function between the speech conversion result of the carrier speech and the conversion result of the watermarked speech, the diffusion loss function between the predicted noise and the actual added noise, and the frequency-weighted loss function of the watermarked speech and the carrier speech under multi-scale short-time Fourier transform.
2. The diffusion-based anti-speech-transformation watermark embedding method according to claim 1, characterized in that, The method for generating the diffusion prior includes: The carrier speech is segmented and windowed, and the short-time Fourier transform is performed on each frame of carrier speech after segmentation and windowing to obtain the amplitude spectrum. The amplitude spectrum is mapped through the Mel filter bank and the logarithm is taken to generate the corresponding log-Mel spectrum. The log-Mel spectrum is also used as conditional information. Frame-level energy is calculated frame by frame from the log-Mel spectrum to obtain a frame-level energy sequence; Based on the frame-level energy sequence, a diagonal covariance matrix is constructed in chronological order. The spectral flatness of the carrier speech is calculated, and a suitable Gaussian noise is determined based on the relationship between the spectral flatness and a preset threshold. The diffusion prior is then obtained by combining the covariance matrix.
3. The diffusion-based anti-speech-transformation watermark embedding method according to claim 2, characterized in that, The formula for calculating the spectral flatness is: , In the formula, For spectral flatness, This represents the total number of frequency bands. For the first Spectral values of each frequency band A positive number less than a set threshold. Represents the logarithmic function with base e; when In this case, super-Gaussian noise should be selected; when In this case, standard Gaussian noise should be selected; when In this case, sub-Gaussian noise should be selected.
4. The diffusion-based anti-speech-transformation watermark embedding method according to claim 1, characterized in that, The anti-speech conversion watermarking embedding network includes: The forward diffusion module is used to randomly select diffusion time steps and inject Gaussian noise into the carrier speech according to the noise schedule to generate noisy speech; The feature tensor generation module is used to perform scale unification operations on noisy speech, log-Mel spectrum and diffusion time step, and add them in the feature channels to form a uniformly modulated feature tensor. The residual feature extraction module is used to perform residual feature extraction processing on the feature tensor to obtain prediction noise; The reverse diffusion module is used to perform reverse diffusion updates on noisy speech based on predicted noise and diffusion priors to generate watermarked speech.
5. The anti-speech conversion watermarking embedding method based on a diffusion model according to claim 4, characterized in that: The residual feature extraction module includes several sequentially arranged dilated residual blocks. The residual components of the previous dilated residual block are output to the next dilated residual block. The Skip components of all dilated residual blocks are added together and averaged. The dimensionality is reduced by 1×1 convolution and ReLU activation is performed to obtain the prediction noise.
6. The anti-speech conversion watermarking embedding method based on a diffusion model according to claim 1, characterized in that: The mathematical expression for the total loss function is: , In the formula, For the total loss function, Let be the adversarial loss function between the speech conversion result of the carrier speech and the conversion result of the watermarked speech. To predict the diffusion loss function between the noise and the actual added noise, Let be the frequency-weighted loss function of the watermarked speech and the carrier speech under multi-scale short-time Fourier transform. , , These are the weights for the adversarial loss function, the diffusion loss function, and the frequency-weighted loss function, respectively.
7. The anti-speech conversion watermarking embedding method based on a diffusion model according to claim 6, characterized in that: The diffusion loss function between the predicted noise and the actual added noise The mathematical expression is: , In the formula, and These represent the actual injected Gaussian noise and the predicted noise, respectively. It is the inverse of the covariance matrix. Indicates time step , This represents conditional information, which is a log-Mel spectrum. Represents the carrier's speech. This indicates noisy speech. This indicates the diffusion time step t and the carrier speech. Actual injected Gaussian noise Take the expected value together. This represents a weighted L2 norm.
8. The anti-speech conversion watermarking embedding method based on a diffusion model according to claim 6, characterized in that: The difference in waveform between the carrier speech and the watermarked speech is denoted as the perturbation. The waveform difference is divided into multiple sub-bands by short-time Fourier transform and transformed to the frequency domain to obtain the amplitude spectrum of the corresponding sub-bands. ; Based on the sensing weights of each frequency band and the amplitude spectrum of the sub-band Generate a frequency-weighted loss function under single-scale short-time Fourier transform. ; Based on the frequency-weighted loss functions under all single-scale short-time Fourier transforms, frequency-weighted loss functions for watermarked speech and carrier speech under multi-scale short-time Fourier transforms are generated. The frequency-weighted loss function under the single-scale short-time Fourier transform Frequency-weighted loss function under multi-scale short-time Fourier transform The mathematical expressions are as follows: , , In the formula, For the first The center frequency of each frequency band The total number of frequency bands. ∈[0,1] is the first The perceptual weight of each frequency band is assigned; a higher weight value indicates that the human ear is more sensitive to that frequency band. Indicates the first The weighted weights corresponding to each scale Indicates the first Each scale.
9. A diffusion-based anti-speech conversion watermarking embedding method according to claim 6, characterized in that: The adversarial loss function between the speech conversion result of the carrier speech and the conversion result of the watermarked speech. The mathematical expression is: , In the formula, This represents the Mel spectrum generated after the carrier speech and source speech are input into the speech conversion model for speech conversion. This represents the Mel spectrum generated after the watermarked speech and the source speech are input into the speech conversion model for speech conversion. express Norm.
10. A computer program product, characterized in that, It includes a computer program / instruction that, when executed by a processor, implements the diffusion-based anti-speech conversion watermarking embedding method according to any one of claims 1 to 9.
Citation Information
Patent Citations
End-to-end real-time speech synthesis method
CN113409759A
Voice watermark injection and right confirmation method and system based on diffusion model
CN117894323A
Voice conversion confrontation audio generation method and device based on conditional diffusion model
CN119132309A
Model training method, audio generation method, watermark detection method and device
CN120236605A
Bandwidth extension and speech enhancement of audio
US20230326476A1