Text-to-speech synthesis model watermarking method based on VITS
By introducing a group residual vector quantization module and a watermark encoder into the VITS model, embedding watermark information, and designing a watermark extractor based on ResNet and timing attention mechanism, the problem of active traceability and copyright protection of synthetic voice in an open environment is solved, and efficient and robust watermark embedding and extraction effects are achieved.
Patent Information
- Application Number
- CN202510193427.6
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-02-21
- Publication Date
- 2025-05-27
- Estimated Expiration
- Not applicable · inactive patent
AI Technical Summary
It is difficult for the existing technology to realize the active traceability and copyright protection of synthetic speech in an open environment, especially when facing the continuous evolution of generative models, the traditional passive evidence method has inherent defects such as insufficient generalization ability and lag in detection.
A text-to-speech synthesis model watermarking method based on VITS is designed. By introducing a group residual vector quantization module after the posterior encoder trained by the VITS model, the continuous hidden variables output by the posterior encoder are quantized into discrete data, and watermark information data of the same latitude as the quantized output is generated through the watermark encoder, and finally superimposed with the hidden variables, and the input decoder generates watermarked speech. At the same time, a trainable watermark extractor based on ResNet and timing attention mechanism was designed to extract watermark information features through multi-layer convolution and residual connections, and finally restore the watermark information through multi-layer linear layers and softmax activation function.
It realizes the invisibility and robustness of watermarks while ensuring the quality of synthetic speech, enhances the deep correlation between the characteristics of generated speech content and the watermark, and provides innovative solutions in the field of watermarks for speech synthesis models.
Smart Images

Figure CN120048270A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of audio watermarking technology, and in particular to a watermarking method based on a VITS text-to-speech synthesis model. Background Art
[0002] In recent years, speech synthesis systems driven by deep learning technology have made breakthrough progress in the field of text-to-speech (TTS). End-to-end models represented by VALLE, SPEAR-TTS, Glow-TTS and VITS abandon the paradigm of traditional TTS system training in modules, and achieve joint modeling of acoustic features and speech waveforms through a unified deep neural network architecture. This architectural innovation significantly improves the perceived naturalness and rhythmic coherence of synthesized speech. In particular, the VITS model introduces variational reasoning and adversarial training mechanisms, and shows excellent performance in complex scenarios such as speech cloning and multi-language synthesis. With the widespread application of advanced models such as VITS on Internet platforms, synthesized speech is breaking through the perceptual boundaries of "machine voice". Its high-fidelity characteristics not only promote the development of voice interaction technology, but also raise severe AI security challenges. Therefore, it is crucial for regulators to regulate synthetic speech through traceability methods. Passive forensics is one of the most common traceability options. Traditional passive forensics technology relies on forgery trace detection, but in the face of continuously evolving generative models, such methods have inherent defects such as insufficient generalization ability and detection lag. How to achieve active traceability and copyright protection of synthesized speech in an open environment has become a core issue that needs to be urgently addressed in the field of artificial intelligence security.
[0003] As a key technology of the active defense system, digital watermark technology provides reliable traceability credentials for multimedia content by embedding imperceptible identification information into carrier data. By adding a unique identification to the TTS model through digital watermark technology to achieve active tracking, it can better trace the source of the synthesized speech and maintain the intellectual property rights of the owner, better manage and supervise its use, and reduce the risk of content being abused or pirated. Summary of the invention
[0004] The present invention aims to provide a watermarking method for a text-to-speech synthesis model based on VITS. This method introduces a group residual vector quantization module after the posterior encoder trained by the VITS model, quantizes the continuous latent variables output by the posterior encoder into discrete data, and generates watermark information data with the same latitude as the quantized output through the watermark encoder, which is finally superimposed with the latent variables and input into the decoder to generate watermarked speech. In the watermark extraction stage, a trainable watermark extractor based on ResNet and temporal attention mechanism is designed, which extracts watermark information features through multi-layer convolution and residual connection, and finally restores the watermark information through multi-layer linear layers and softmax activation function. The specific scheme is as follows:
[0005] 1. Speech Encoder Module: Generate intermediate latent variables through the VITS speech encoder, and then group and residual quantize them into discrete data z.
[0006] 2. Watermark Encoder: Encode binary watermark information into a latent variable w with the same dimension as z.
[0007] 3. Final Latent Variable Generate: Superimpose z and w
[0008] to generate the final latent variable.
[0009] 4. Watermarked Speech Generate: After the final latent variable is quantized and restored, it is input into the VITS decoder to generate watermarked speech.
[0010] 5. Attack Simulation: Process the synthesized speech through the attack simulation module and use it as the input for the subsequent watermark extractor.
[0011] 6. Message Decoder: Extract features from the spectrogram corresponding to the speech after simulated attack, and decode and restore the watermark information.
[0012] 7. Loss Function
[0013] In addition to the loss function for the original VITS model training, a quantization loss is introduced
[0014]
[0015] i and c represent the i-th group and the c-th residual quantizer, and the cross-entropy watermark loss used to measure the accuracy of extracting watermark information from the watermarked speech, where l ij and p ij are the one-hot encoding and prediction probability of the i-th digit watermark for the j digits respectively:
[0016]
[0017] The watermark method for the text-to-speech synthesis model based on VITS embeds watermark information in the speech encoding stage during the training process of the VITS model through the method of endogenous watermarking, and jointly optimizes with the decoder, watermark extractor, and discriminator in the model to achieve a stable end-to-end embedding process. On the one hand, it enhances the deep association between the features of the generated speech content and the watermark, and on the other hand, it ensures the invisibility and robustness of the watermark. The present invention proposes a watermark method for the text-to-speech synthesis model based on VITS, which improves the invisibility and robustness of the watermark on the premise of ensuring the quality of the synthesized speech by designing an efficient embedding strategy and extraction mechanism, providing an innovative solution for the field of watermarks in speech synthesis models. BRIEF DESCRIPTION OF THE DRAWINGS
[0018] To more clearly illustrate the technical solutions of the embodiments of the present invention, the drawings used in the description of the related technologies will be introduced below. The drawings described below are only the embodiments of the present invention.
[0019] Figure 1 Structural schematic diagram of the VITS watermark model provided by the present invention
[0020] Figure 2 Schematic diagram of the watermark embedding and extraction framework provided by the present invention
[0021] Figure 3 Structural schematic diagram of the watermark extractor provided by the present invention SPECIFIC IMPLEMENTATION METHODS
[0022] The technical solutions in the embodiments of the present invention will be clearly and completely described below with reference to the drawings in the embodiments of the present invention. It should be noted that the following detailed description is exemplary and is intended to provide further explanation of the present invention. The present invention provides a watermark method for the text-to-speech synthesis model based on VITS, as Figure 1 shown, including a watermark encoding module, a speech encoding quantization module, an analog attack module, and a watermark extraction module. The main steps include:
[0023] S101. The speech is encoded by a speech encoder (WaveNet residual structure) to obtain an intermediate latent variable z, and then passes through a grouped residual vector quantization module, where the number of codebooks is set to 8 (each layer contains 4 codebooks, and two residual layers are used), to obtain the final output z_q of the speech module.
[0024] S102. The watermark encoder inputs an m-bit binary watermark message M, first maps it to a low-dimensional vector representation through an embedding layer, and then sends it to multiple fully connected layers (FC layers), each followed by a ReLU activation function, so as to extract useful latent representations, and finally outputs a latent vector w with the same dimension as z_q.
[0025] S103. Implement watermark embedding by linearly superimposing \(z_q\) and \(w\), then obtain the input of the decoder (Hifigan v1 generator) through quantization restoration, and finally generate watermarked speech through the decoder.
[0026] S104. Add an attack simulation module before the watermark extraction module. First, perform feature extraction on the corresponding spectrogram of the watermarked speech after specific attack operations through a two-dimensional convolutional layer, then extract the \(r\) vector through 4 ResNet modules and connect it to \(m\) groups of two-layer linear layers, introduce the self-attention mechanism, calculate the probability distribution of each numerical value, and finally obtain the predicted value through softmax.
[0027] S105. All experiments were carried out in the PyTorch 1.13.0 framework and ran on a single RTX 3090 GPU. The specific training process is as follows:
[0028]
[0029] The experiment used the Adam optimizer with a learning rate of 0.0002 to train the model and jointly optimize the model loss.
[0030] To evaluate the quality of the synthesized watermarked speech, in the experimental stage, objective metrics including Perceptual Evaluation of Speech Quality (PESQ) and Short-Time Objective Intelligibility (STOI), and the subjective metric Mean Opinion Score (MOS) were used to evaluate the quality of the synthesized watermarked speech. Among them, the higher the PESQ value, the better the auditory quality of the speech. STOI quantifies the word intelligibility on a scale of 0 - 1 by measuring speech intelligibility, and MOS uses a grading judgment method.
[0031] We used the state-of-the-art deep learning-based audio watermarking framework WavMark as a baseline for comparison experiments. Since this framework only implements the function of embedding watermarks in audio, we used WavMark to embed watermarks into the speech waveforms generated by VITS, and compared the obtained watermarked speech with the watermarked speech generated by our proposed model.
[0032]
[0033] The synthesized speech was respectively subjected to the following seven attack simulations: normal extraction without attack (Normal); resampling at 90% of the original sampling rate and recovery rate (RS-90); adding white noise with a signal-to-noise ratio of 35 db (Noise-W35); randomly removing 0.1% of the sample points (SD-01); reducing the amplitude to 90% of the original value (AR-90); attenuating the obtained speech by 0.3 times, delaying the volume by 15%, and then covering the original speech (EA-0315). The watermark extraction accuracy under different attacks obtained through experiments is shown in the table, indicating that the speech generated by the model proposed in this paper maintains a higher extraction accuracy than the baseline when facing attacks.
[0034]
[0035] As described above, only the preferred specific embodiments of the present invention are given, but the protection scope of the present invention is not limited thereto. Any changes or substitutions that can be easily thought of by those familiar with the technology within the technical scope disclosed by the present invention should be covered within the protection scope of the present invention. Therefore, the protection scope of the present invention should be subject to the protection scope of the claims.
Claims
1. A VITS-based text-to-speech synthesis model watermarking method, characterized in that: include: A. The input speech X is encoded into an intermediate continuous latent variable z by the speech encoder, and then the group residual vector quantization module is introduced to quantize z into discrete data z_q. B. Designed watermark encoder E w The m-bit binary bit stream watermark information M is converted into a latent vector w with the same latitude as z_q, and the watermark extractor accurately extracts the binary watermark information from the watermarked speech after attack simulation. C. The embedding module linearly superimposes z_q and w, and obtains the final decoder input z' after quantization restoration. D. z' is input into the decoder to generate watermarked speech X'.
2. The VITS-based text-to-speech synthesis model watermarking method according to claim 1, characterized in that: Step A further comprises the following steps: A1. The speech is encoded by the speech encoder (WaveNet residual structure) to obtain the intermediate latent variable z, and then passes through the group residual vector quantization module, where the number of codebooks is set to 8 (each layer contains 4 codebooks and two residual layers are used) to obtain the final output z_q of the speech module.
3. The VITS-based text-to-speech synthesis model watermarking method according to claim 1, characterized in that: Step B further comprises the following steps: B1. The watermark encoder structure mainly consists of an embedding layer and two fully connected layers. First, the binary watermark information is mapped into a low-dimensional vector representation through an embedding layer, and then sent to multiple fully connected layers (FC layers). Each layer is followed by a ReLU activation function to extract useful potential representations. B2. The watermark extractor takes the spectrum corresponding to the watermarked speech after attack simulation as the original input, uses a two-dimensional convolutional layer to extract its features, then extracts the r vector through four ResNet modules and connects it with m groups of two-layer linear layers, introduces a self-attention mechanism, calculates the probability distribution of each bit value, and finally obtains the predicted value through softmax, outputting the embedded watermark information M'.
4. The VITS-based text-to-speech synthesis model watermarking method according to claim 1, characterized in that: Step C further comprises the following steps: C1, linear superposition, based on the core idea of adaptive watermark embedding strategy, dynamically adjusts the weighting coefficient of the watermark vector according to the speech characteristics of each frame. Specifically, the square sum of the hidden variables z_q is used to define the energy value of each frame, and the adaptive adjustment coefficient α is calculated based on the energy value. i , used to control the strength of watermark embedding.
5. The VITS-based text-to-speech synthesis model watermarking method according to claim 1, characterized in that: Step D further comprises the following steps: D1, the final target hidden variable generates watermarked speech through the decoder of the HifiGan v1 generator.
Citation Information
Patent Citations
Audio watermark embedding and extracting method based on deep learning
CN115831131A
Voice communication system and method
CN117544603A
Digital audio-frequency water-print inlaying and detecting method based on auditory characteristic and integer lift ripple
CN1529246A
Encoding method and decoding method
US20240395265A1