Voice synthesis method, device and equipment embedding watermark information and storage medium

By embedding watermark information during the audio generation process, and utilizing the latent watermark encoder and convolutional network of the audio generation model, the method solves the problems of low watermark embedding efficiency and poor concealment in the existing technology, and achieves high-quality watermark information embedding and audio generation.

CN120164453BActive Publication Date: 2025-11-18PING AN TECH (SHENZHEN) CO LTD
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202510429064.1
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-04-07
Publication Date
2025-11-18
Estimated Expiration
2045-04-07

AI Technical Summary

Technical Problem

Existing audio watermarking technologies suffer from low embedding efficiency and concealment due to their post-generation embedding method, and are susceptible to noise interference during the generation process, affecting audio quality. In particular, the watermark information embedding effect is poor in financial and medical scenarios.

Method used

The target watermark information is encoded by a latent watermark encoder in the audio generation model and embedded during the audio generation process. The initial speech information is generated by combining the original audio features. The watermark information is extracted and reconstructed using a convolutional network and a decoder. Similarity verification is performed to ensure the embedding quality.

Benefits of technology

It improves the concealment and efficiency of watermark embedding, reduces the impact on audio quality, enhances the robustness and reliability of watermark information, supports multiple watermark formats and embedding methods, and adapts to multiple audio generation models.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120164453B_ABST
    Figure CN120164453B_ABST
Patent Text Reader

Abstract

The application belongs to the technical field of artificial intelligence, can be applied to financial or medical scenes, and discloses a speech synthesis method and device embedding watermark information, equipment and a storage medium, comprising: obtaining target watermark information; encoding the target watermark information through a latent watermark encoder to obtain watermark information coding; taking an intermediate latent space of an audio generation model as a watermark embedding position and embedding the watermark information coding, combining original audio features to obtain target latent representation of the target watermark information; processing the target latent representation through a generator to generate initial speech information; extracting watermark information features of the initial speech information through a convolutional network, and reconstructing the watermark information features through a decoder to obtain decoded watermark information; comparing the decoded watermark information with the target watermark information to obtain a similarity value, and when the similarity value is greater than a preset similarity threshold, outputting the initial speech information as target speech information.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the field of artificial intelligence and is applied to online processing scenarios such as finance and digital healthcare. In particular, it relates to a speech synthesis method, device, equipment, and storage medium that embeds watermark information. Background Technology

[0002] With the development of audio generation models, such as text-to-speech (TTS) and voice conversion (VC) technologies, and their widespread application in the commercial field, the protection of audio content copyright and the verification of its authenticity have become critical issues that the industry urgently needs to address.

[0003] However, existing audio watermarking technologies have shown obvious limitations in addressing these challenges. Traditional audio watermarking technologies generally adopt the "post-generation embedding" method, that is, the watermark is added after the audio is generated. However, this method is disconnected from the audio generation process, resulting in low watermark embedding efficiency and concealment, and it is easily affected by weak noise during the generation process, which affects the quality of the generated audio.

[0004] For example, in financial or medical scenarios, financial institutions (such as banks) need to embed watermark information for investment strategy guidance, medical equipment operation instructions, medical training courses, etc. However, traditional watermark embedding generally adopts the "embedding after generation" method, resulting in poor quality in the above audio, and its watermark embedding efficiency and concealment are also relatively low.

[0005] Therefore, there is an urgent need for a method that can embed watermark information during the audio generation process. Summary of the Invention

[0006] The purpose of this application is to provide a speech synthesis method, apparatus, device and storage medium for embedding watermark information. Its main purpose is to embed watermark information during the audio generation process, so as to improve the concealment and embedding efficiency of watermark embedding, and also reduce the impact of watermark information on audio quality during the embedding process, thereby improving the audio quality.

[0007] Firstly, in order to solve the above-mentioned technical problems, embodiments of this application provide a speech synthesis method for embedding watermark information, which adopts the following technical solution:

[0008] Obtain the target watermark information to be embedded;

[0009] The target watermark information is encoded by a latent watermark encoder of an audio generation model to obtain the watermark information encoding.

[0010] The intermediate latent space of the audio generation model is used as the watermark embedding position. The watermark information is encoded and embedded according to the watermark embedding position, and combined with the original audio features, the target latent representation of the target watermark information is obtained.

[0011] The generator of the audio generation model processes the latent representation of the target to generate initial speech information.

[0012] The watermark information features of the initial speech information are extracted by the convolutional network of the audio generation model, and the watermark information features are reconstructed by the decoder of the audio generation model to obtain the decoded watermark information.

[0013] The decoded watermark information is compared with the target watermark information to obtain a similarity value. When the similarity value is greater than a preset similarity threshold, the initial speech information is output as the target speech information.

[0014] Secondly, in order to solve the above-mentioned technical problems, this application also provides a speech synthesis device with embedded watermark information, which adopts the following technical solution:

[0015] The acquisition module is used to acquire the target watermark information that needs to be embedded.

[0016] The encoding module is used to encode the target watermark information through the latent watermark encoder of the audio generation model to obtain the watermark information encoding;

[0017] The embedding module is used to take the intermediate latent space of the audio generation model as the watermark embedding position, encode and embed the watermark information according to the watermark embedding position, and combine it with the original audio features to obtain the target latent representation of the target watermark information.

[0018] The generation module is used to process the target latent representation through the generator of the audio generation model to generate initial speech information;

[0019] The reconstruction module is used to extract watermark information features from the initial speech information through the convolutional network of the audio generation model, and reconstruct the watermark information features through the decoder of the audio generation model to obtain decoded watermark information.

[0020] The comparison module is used to compare the decoded watermark information with the target watermark information to obtain a similarity value. When the similarity value is greater than a preset similarity threshold, the initial speech information is output as the target speech information.

[0021] Thirdly, in order to solve the above-mentioned technical problems, embodiments of this application also provide a computer device, including: at least one processor; and a memory communicatively connected to the at least one processor; wherein the memory stores instructions executable by the at least one processor, the instructions being executed by the at least one processor to enable the at least one processor to perform the speech synthesis method for embedding watermark information as described above.

[0022] Fourthly, in order to solve the above-mentioned technical problems, embodiments of this application also provide a computer-readable storage medium storing a computer program, which, when executed by a processor, implements the speech synthesis method for embedding watermark information as described above.

[0023] Compared with the prior art, the embodiments of this application have the following main advantages:

[0024] By obtaining the target watermark information that needs to be embedded, the specific content that needs to be embedded can be clearly identified, which is beneficial for subsequent watermark processing, embedding and verification work.

[0025] Encoding the target watermark information through the latent watermark encoder of the audio generation model can improve the robustness of the target watermark information, prevent the target watermark information from being tampered with and attacked, and improve the reliability of the target watermark information.

[0026] By embedding the target watermark information into the intermediate latent space of the audio generation model, the watermark information does not need to be embedded after the audio is generated, which improves the concealment and embedding efficiency of the watermark. When embedding the watermark information, the impact on the audio quality can also be adjusted to improve the audio quality.

[0027] By processing the target's latent representation through a generator to generate initial speech information, the impact of watermark embedding on the quality of the initial speech signal can be reduced, meeting the requirements of high-fidelity speech synthesis. Furthermore, the audio generation model with embedded watermark carries a unique identifier each time it generates initial speech information, making it easy to trace the source and preventing the content from being forged or tampered with.

[0028] The initial speech information is decoded by a convolutional network and a decoder, and the decoded watermark information of the initial speech information is reconstructed. This ensures that the watermark information in the initial speech information can be completely extracted, and improves the robustness of watermark information embedding.

[0029] By embedding watermark information when generating initial speech information and verifying the embedded watermark information, this method can support multiple watermark formats and embedding methods, thus improving its versatility and compatibility. Attached Figure Description

[0030] To more clearly illustrate the solutions in this application, the accompanying drawings used in the description of the embodiments of this application will be briefly introduced below. Obviously, the accompanying drawings described below are some embodiments of this application. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.

[0031] Figure 1 This is an exemplary system architecture diagram to which this application can be applied;

[0032] Figure 2 A flowchart of an embodiment of the speech synthesis method for embedding watermark information according to this application;

[0033] Figure 3 This is a schematic diagram of a speech synthesis device for embedding watermark information according to an embodiment of the present application;

[0034] Figure 4 This is a schematic diagram of the structure of one embodiment of the device according to this application. Detailed Implementation

[0035] Unless otherwise defined, all technical and scientific terms used herein have the same meaning as commonly understood by one of ordinary skill in the art to which this application pertains; the terminology used herein in the specification of the application is for the purpose of describing particular embodiments only and is not intended to be limiting of the application; the terms "comprising" and "having," and any variations thereof, in the specification, claims, and foregoing drawings of this application, are intended to cover non-exclusive inclusion. The terms "first," "second," etc., in the specification, claims, or foregoing drawings of this application are used to distinguish different objects, not to describe a particular order.

[0036] In this document, the term "embodiment" means that a particular feature, structure, or characteristic described in connection with an embodiment may be included in at least one embodiment of this application. The appearance of this phrase in various places throughout the specification does not necessarily refer to the same embodiment, nor is it a separate or alternative embodiment mutually exclusive with other embodiments. It will be explicitly and implicitly understood by those skilled in the art that the embodiments described herein can be combined with other embodiments.

[0037] To enable those skilled in the art to better understand the present application, the technical solutions in the embodiments of the present application will be clearly and completely described below with reference to the accompanying drawings.

[0038] like Figure 1As shown, system architecture 100 may include terminal device 101, network 102, and server 103. Terminal device 101 may be a laptop 1011, tablet 1012, or mobile phone 1013. Network 102 is used as a medium to provide a communication link between terminal device 101 and server 103. Network 102 may include various connection types, such as wired, wireless communication links, or fiber optic cables.

[0039] Users can use terminal device 101 to interact with server 103 via network 102 to receive or send messages, etc. Various communication client applications can be installed on terminal device 101, such as web browser applications, shopping applications, search applications, instant messaging tools, email clients, social media platform software, etc.

[0040] Terminal device 101 can be various electronic devices with a display screen and support web browsing. In addition to laptops 1011, tablets 1012, or mobile phones 1013, terminal device 101 can also be e-book readers, MP3 players (Moving Picture Experts Group Audio Layer III), MP4 players (Moving Picture Experts Group Audio Layer IV), laptops, and desktop computers.

[0041] Server 103 can be a server that provides various services, such as a backend server that provides support for the pages displayed on terminal device 101.

[0042] It should be noted that the speech synthesis method for embedding watermark information provided in this application embodiment is generally executed by a server / terminal device, and correspondingly, the speech synthesis device for embedding watermark information is generally set in the server / terminal device.

[0043] It should be understood that Figure 1 The number of terminal devices, networks, and servers shown is merely illustrative. Depending on implementation needs, any number of terminal devices, networks, and servers can be included.

[0044] Continue to refer to Figure 2This document illustrates a flowchart of an embodiment of the speech synthesis method for embedding watermark information according to this application. The order of steps in the flowchart can be changed, and some steps can be omitted, depending on different requirements. The speech synthesis method for embedding watermark information provided in this application can be applied to any scenario requiring speech synthesis for embedding watermark information, and thus can be applied to products in these scenarios. The speech synthesis method for embedding watermark information includes the following steps:

[0045] Step S201: Obtain the target watermark information to be embedded.

[0046] In this embodiment, the speech synthesis method embedding watermark information runs on an electronic device (e.g., Figure 1 The server / terminal device shown can obtain the target watermark information via wired or wireless connection. It should be noted that the aforementioned wireless connection methods may include, but are not limited to, 3G / 4G / 5G connections, WiFi connections, Bluetooth connections, WiMAX connections, Zigbee connections, UWB (ultra wideband) connections, and other currently known or future wireless connection methods.

[0047] In this embodiment, the target watermark information refers to the watermark information that needs to be embedded in the audio. The target watermark information includes, but is not limited to, text information, image information, and audio information, such as copyright statements, author identifiers, and product numbers in text information, company logos and product icons in image information, and specific audio segments or rhythm patterns in audio information, depending on the specific situation.

[0048] Specifically, once the target watermark information to be embedded in the audio is determined, the target watermark information is format-converted into a suitable embedding format. For example, text information can be converted into a binary code stream format. Then, preprocessing operations are performed on the target watermark information to obtain preprocessed target watermark information. These preprocessing operations include, but are not limited to, text length adjustment, scrambling, and intensity adjustment. Text length adjustment is to prevent the information from becoming too long by truncating or compressing the target watermark information. Scrambling is to break the periodicity and regularity of the target watermark information, such as using Arnold transform. Intensity adjustment ensures that the target watermark information can be accurately extracted after embedding without affecting the quality of the audio signal. Through the above preprocessing methods, such as scrambling, the target watermark information becomes difficult to directly observe or identify, improving its concealment. Encryption ensures the security of the target watermark information and prevents tampering.

[0049] In one embodiment, after obtaining the target watermark information to be embedded, the method further includes:

[0050] The validity of the target watermark information is detected, wherein the validity includes integrity check, readability check and robustness test;

[0051] When the integrity check, readability check, and robustness test all pass, the target watermark information is recorded and stored in a preset watermark information database.

[0052] In this embodiment, the target watermark information undergoes validity testing, which includes integrity checks, readability checks, and robustness tests. Specifically, it checks whether the target watermark information is complete and without omissions or errors; ensures that the watermark information can be correctly read or recognized; and sets a robustness threshold to determine whether the stability and durability of the target watermark information are greater than the parameter set by the robustness threshold. If it is greater than the robustness threshold, the robustness test is passed; if it is less than or equal to the robustness threshold, the robustness test is failed. After all the above validity tests are passed, the target watermark information, such as the content or identifier of the target watermark information and the embedding time of the watermark information, is recorded. Subsequently, the target watermark information is stored in a preset watermark information database. The watermark information database refers to a database used to store all target watermark information that has passed the validity test, including but not limited to relational databases, to improve the efficiency of subsequent target watermark information verification steps.

[0053] In this embodiment, after obtaining the target watermark information, the validity of the target watermark information is checked to ensure the accuracy and reliability of the target watermark information and prevent the subsequent verification steps from failing to verify it, thus avoiding the waste of computing resources. Furthermore, storing the target watermark information in a dedicated database for storing watermark information allows for convenient querying and retrieval, improving the management efficiency of the target watermark information.

[0054] In the financial sector, for example, in a feasible case A, a bank or financial institution needs to protect the copyright of its voice advertisements. The institution first determines the target watermark information to be embedded, "Copyright owned by XX Company", and converts "Copyright owned by XX Company" into a binary code stream. Then, it checks the validity of "Copyright owned by XX Company". If the validity is verified, "Copyright owned by XX Company" is stored in a preset watermark information database.

[0055] In this embodiment, by obtaining the target watermark information to be embedded, the specific content to be embedded can be clearly defined, which is beneficial for subsequent watermark processing, embedding and verification.

[0056] Step S202: Encode the target watermark information using the latent watermark encoder of the audio generation model to obtain the watermark information encoding.

[0057] In this embodiment, the audio generation model is a composite model, including a potential watermark encoder, a generator, a convolutional network, and a decoder, to achieve the generation of audio signals and the embedding of watermark information.

[0058] In this embodiment, the target watermark information from step S201 is input into the audio generation model, and the target watermark information is encoded by the latent watermark encoder of the audio generation model to encode the watermark information into an encoding form suitable for subsequent embedding in the audio generation model, thereby obtaining the watermark information encoding. The watermark information encoding is suitable for embedding in the latent space of the subsequent audio generation model.

[0059] Following the above feasible example A, the target watermark information "Copyright owned by XX Company" is input into the audio generation model. At this time, the potential watermark encoder of the model will be called to encode "Copyright owned by XX Company" into a form suitable for subsequent embedding in the audio generation model, thus obtaining the watermark information encoding corresponding to "Copyright owned by XX Company".

[0060] In one embodiment, the latent watermark encoder using the audio generation model encodes the target watermark information to obtain the watermark information encoding, including:

[0061] The target watermark information is padded or truncated based on the number of neurons in the input layer of the potential watermark encoder to obtain at least one watermark sequence.

[0062] At least one of the watermark sequences is input into the latent watermark encoder through the input layer of the latent watermark encoder;

[0063] The hidden layer of the latent watermark encoder converts at least one of the watermark sequences into a vector corresponding to a high-dimensional latent representation.

[0064] The high-dimensional latent representation is output through the output layer of the latent watermark encoder to obtain the watermark information encoding.

[0065] In this embodiment, the aforementioned latent watermark encoder refers to a multilayer perceptron (MLP) model, including an input layer, multiple hidden layers, and an output layer. Each hidden layer includes multiple neurons. Since the number of neurons in the input layer must match the sequence length of the target watermark information, the target watermark information needs to be padded or truncated based on the number of neurons in the input layer to obtain at least one watermark sequence, ensuring that the sequence length of the target watermark information matches the number of neurons in the input layer. Subsequently, at least one watermark sequence is input into the latent watermark encoder through its input layer. Each hidden layer of the latent watermark encoder converts the watermark sequence (watermark information) into a vector corresponding to a high-dimensional latent representation. Specifically, each hidden layer receives input from the previous hidden layer or the input layer, and performs weighted summation, bias adjustment, and activation functions, including but not limited to ReLU, sigmoid, and tanh functions. Finally, the output layer of the latent watermark encoder outputs the vector corresponding to the high-dimensional latent representation, thus obtaining the watermark information encoding.

[0066] In this embodiment, the target watermark information is mapped to a vector corresponding to a high-dimensional latent representation through the MLP model, which can flexibly adapt to various audio generation models and improve the flexibility and adaptability of watermark embedding. The target watermark information is converted into watermark information encoding through the MLP model, which can generate watermark information encoding that is difficult to tamper with or forge, thereby improving the security of watermarking technology and making the embedded target watermark information more reliable and difficult to remove.

[0067] In one embodiment, after the target watermark information is encoded by the latent watermark encoder of the audio generation model to obtain the watermark information encoding, the method further includes:

[0068] Obtain random noise of a preset type, and preset the noise intensity of the random noise;

[0069] The random noise is added to the watermark information encoding according to the preset type of random noise and the corresponding noise intensity.

[0070] In this embodiment, the aforementioned preset type of random noise includes random noise and specific distribution noise. Random noise refers to the generation of random numbers and pseudo-random number sequences, while specific distribution noise refers to Gaussian noise, uniform noise, etc. The aforementioned noise intensity refers to the pre-evaluated noise intensity. The type of random noise and the corresponding noise intensity are selected, and the random noise is added to the watermark information encoding by direct addition or weighted addition. After adding the random noise to the watermark information encoding, the watermark information encoding also needs to be normalized to ensure that the numerical range is within a suitable range, thus maintaining the quality and stability of the subsequently generated audio signal.

[0071] In this embodiment, by adding random noise to the watermark information encoding, the concealment and robustness of the watermark information are enhanced, making the watermark information difficult to detect and resistant to various attacks, thereby improving the security of the watermark information.

[0072] In this embodiment, by encoding the target watermark information through the latent watermark encoder of the audio generation model, the robustness of the target watermark information can be improved, preventing the target watermark information from being tampered with and attacked, and improving the reliability of the target watermark information.

[0073] Step S203: Use the intermediate latent space of the audio generation model as the watermark embedding position, encode and embed the watermark information according to the watermark embedding position, and combine it with the original audio features to obtain the target latent representation of the target watermark information.

[0074] In this embodiment, the model architecture of the audio generation model is analyzed, and multiple candidate intermediate latent spaces of the audio generation model are analyzed from the model architecture. According to actual needs, one intermediate latent space is selected from the multiple candidate intermediate latent spaces as the watermark embedding position. According to the selected watermark embedding position, the watermark information obtained in step S202 is encoded and embedded, and combined with the original audio features, the target latent representation of the target watermark information is obtained.

[0075] Specifically, step S202 describes the main components of the audio generation model, including a latent watermark encoder, a generator, a convolutional network, and a decoder. Multiple candidate intermediate latent spaces are identified from all components. These candidate intermediate latent spaces are typically located at the output of each component or within its internal processing layer. The stealth, robustness, and potential impact on audio quality of each candidate intermediate latent space are evaluated to obtain the evaluation result for each candidate intermediate latent space. Robustness refers to the ability to accurately detect the audio signal even after various processing methods (such as compression, cropping, and format conversion). Stealth refers to the difficulty in detecting the watermark information after embedding. The impact refers to whether watermark embedding will lead to a significant decrease in audio quality. Based on the requirements of watermark embedding, such as concealment, robustness, and impact on audio quality, a candidate intermediate latent space is selected from multiple candidate intermediate latent spaces as the watermark embedding position. The watermark information obtained in step S202 is encoded and embedded into the watermark embedding position using a preset embedding algorithm. Combined with the features of the original audio, the target latent representation of the target watermark information is obtained. The embedding algorithm includes, but is not limited to, least significant bit substitution, phase coding, discrete wavelet transform (DWT), etc. The features of the original audio include, but are not limited to, spectral features, temporal features, prosodic features, speech features, etc.

[0076] Following the feasible example A above, the embedding positions of each component in the audio generation model are identified, resulting in multiple candidate intermediate latent spaces. The robustness, concealment, and impact on audio quality of these candidate intermediate latent spaces are evaluated, yielding evaluation results. The embedding requirements for watermark information are defined as robustness and impact on audio quality. Based on these requirements, a candidate intermediate latent space with high overall robustness and low impact on audio quality is selected from the multiple evaluation results. This candidate intermediate latent space is then used as the watermark embedding position. The watermark information obtained in step S202 is encoded and embedded into this position using a preset embedding algorithm (such as least significant bit substitution, phase coding, Discrete Wavelet Transform (DWT), etc.). After embedding, combined with the features of the original audio, a target latent representation containing the watermark information is obtained.

[0077] In this embodiment, the target watermark information is embedded into the intermediate latent space of the audio generation model, eliminating the need to embed the target watermark information after audio generation, thus improving the concealment and embedding efficiency of the watermark. Furthermore, the impact on audio quality can be adjusted during the embedding of the watermark information to improve audio quality.

[0078] In one embodiment, encoding and embedding the watermark information according to the watermark embedding position, and combining it with the original audio features to obtain the target potential representation of the target watermark information, includes:

[0079] Acquire the original audio signal and extract the original audio features from the original audio signal;

[0080] The watermark information is encoded and embedded into the watermark embedding position;

[0081] The watermark information encoding embedded at the watermark embedding location is combined with the original audio features to generate the target potential representation of the target watermark information.

[0082] In this embodiment, an unprocessed raw audio signal is obtained from an audio source. The raw audio signal refers to the standard electrical signal that subsequently generates the initial speech information. A preset feature extraction method is used to extract features from the raw audio signal to obtain raw audio features. The preset feature extraction method includes, but is not limited to, Mel-frequency cepstral coefficients (MFCC) and linear predictive cepstral coefficients (LPCC). The watermark information is encoded and embedded into the determined watermark embedding position. The watermark information encoding after watermark embedding is combined with the raw audio features using weighted superposition, feature fusion, and other methods to finally generate the target potential representation of the target watermark information. The target potential representation includes the features of the raw audio signal and the watermark information.

[0083] In this embodiment, combining the watermarked information with the original audio features ensures that the watermark information can still be reliably detected and extracted after the audio signal has undergone various processing steps. Encoding and combining the original audio features with the watermark information minimizes the impact on the original audio features, thereby maintaining the sound quality and audibility of the subsequently generated initial speech signal.

[0084] Step S204: The target latent representation is processed by the generator of the audio generation model to generate initial speech information.

[0085] In this embodiment, the generator includes an input layer, a transform layer, a decoding layer, and an output layer. The transform layer includes a convolutional sub-layer, a fully connected sub-layer, and a pooling sub-layer. The latent representation of the target is input into the generator of the audio generation model. The latent representation of the target is processed collaboratively by the various layers of the generator to finally generate the initial speech information. It is worth mentioning that at this time, it is uncertain whether the initial speech information contains watermark information or whether the watermark information can be accurately extracted.

[0086] Following the above feasible example A, the latent representation of the target with "Copyright owned by XX Company" is input into the generator of the audio generation model through the input layer of the generator. The intermediate features of the latent representation of the target with "Copyright owned by XX Company" are extracted through the transform layer of the generator, and the intermediate features are decoded through the decoding layer of the generator to obtain the initial speech information with the watermark information "Copyright owned by XX Company". The initial speech information with the watermark information "Copyright owned by XX Company" is output through the output layer of the generator.

[0087] In this embodiment, the generator processes the potential representation of the target to generate initial speech information, which can reduce the impact of watermark embedding on the quality of the initial speech signal and meet the requirements of high-fidelity speech synthesis. Moreover, the audio generation model with embedded watermark carries a unique identifier each time it generates initial speech information, which is convenient for tracing the source and preventing the content from being forged or tampered with.

[0088] In one embodiment, the process of processing the target latent representation through the generator of the audio generation model to generate initial speech information includes:

[0089] The potential representation of the target is input into the generator through the input layer of the generator;

[0090] The generator's transform layer extracts multiple features of the target's latent representation, and a feature fusion operation is performed on these multiple features to obtain intermediate features;

[0091] The intermediate features are converted using the decoding layer of the generator to generate initial speech information;

[0092] The initial speech information is output using the output layer of the generator.

[0093] In this embodiment, the target latent representation is input into the generator of the audio generation model through the generator's input layer. The generator's transformation layer performs feature extraction on the input target latent representation. Specifically, the feature extraction includes: performing a linear transformation on the target latent representation through a fully connected sublayer in the transformation layer to convert the target latent representation into a new representation in the feature space, such as multiplying the target latent representation by a weight matrix and adding a bias vector to obtain an output vector. This output vector represents the target latent representation in the new feature space. Then, a preset nonlinear activation function (such as ReLU or Sigmoid) is applied. The generator uses functions such as d to map the output vector to a nonlinear space. After nonlinear activation, multiple convolutional kernels in the convolutional sub-layer extract features from the output vector, resulting in multiple local feature maps. A pooling sub-layer then performs dimensionality reduction on these feature maps. These dimensionality-reduced feature maps are then fused (e.g., concatenated, added, dot product) to obtain intermediate features. The generator's decoding layer then uses specific waveform generation techniques (e.g., pulse code modulation, parametric synthesis) to decode these intermediate features, ultimately generating the initial speech information. Finally, the generator's output layer outputs this initial speech information.

[0094] In this embodiment, by processing the target potential representation through a generator, initial speech information including watermark information can be generated, avoiding the need to embed watermark information into the speech information after subsequent speech information generation, thus improving the efficiency of watermark embedding; at the same time, it also avoids the problem of audio quality degradation caused by watermark information when embedding speech information, thus improving audio quality.

[0095] In one embodiment, before the generator of the audio generation model processes the target latent representation to generate initial speech information, the method further includes:

[0096] Step S2041: Obtain the audio training dataset with watermark information, and use the audio training dataset as the training set;

[0097] Step S2042: Obtain the perceptual loss function and define a preset loss function, and combine the preset loss function with the perceptual loss function to obtain a combined loss function;

[0098] Step S2043: Construct an initial generator, train the initial generator using the training set, and generate initial training speech information;

[0099] Step S2044: Calculate the combined loss function based on the initial training speech information to obtain the calculation result of the combined loss function;

[0100] Step S2045: Update the parameters of the initial generator based on the calculation result of the combined loss function;

[0101] Repeat steps S2042 to S2045 until the calculated result of the combined loss function is less than the preset loss threshold, and obtain the generator that has been trained.

[0102] In this embodiment, a large amount of audio training dataset with watermark information is acquired and used as the training set. The training set is preprocessed, including denoising and normalization, to improve its data quality. A perceptual loss function and a preset loss function are obtained. The perceptual loss function measures the perceptual difference between the initial training speech information and the target training speech information, where the target training speech information is a predefined set of parameters. The preset loss function includes adversarial loss functions. The perceptual loss function and the preset loss function are combined to obtain a combined loss function. An initial generator is constructed, with the main construction being the same as described above. The initial generator is trained using the training set to generate initial training speech information. The combined loss function is calculated to obtain its result, which represents the difference between the initial training speech information and the target training speech information. Based on the calculated difference, the parameters of the initial generator are updated using a backpropagation algorithm. The above steps are repeated until the combined loss function result is less than a preset loss threshold, resulting in a trained generator.

[0103] In this embodiment, by adding a perceptual loss function to train the initial generator, the impact of watermark embedding on the naturalness of speech can be reduced, thereby improving the audio quality of the speech information generated by the generator.

[0104] Step S205: Extract the watermark information features of the initial speech information through the convolutional network of the audio generation model, and reconstruct the watermark information features through the decoder of the audio generation model to obtain the decoded watermark information.

[0105] In this embodiment, after generating the initial voice information, it is also necessary to verify the initial voice information to prevent the watermark information in the initial voice information from being unable to be extracted or being extracted incompletely.

[0106] In this embodiment, the watermark information features of the initial speech information are extracted by the convolutional network in the audio generation model, and then the watermark information features are reconstructed by the decoder in the deep audio model to obtain the decoded watermark information. The convolutional network includes an input layer, a convolutional layer, a pooling layer and an output layer; the decoder includes an input layer, a transposed convolutional layer and an output layer.

[0107] Following the above feasible example A, the initial speech information with the watermark "Copyright, XX Company" generated by the generator is input into the convolutional network through the input layer of the convolutional network. The watermark information features of the initial speech information with the watermark "Copyright, XX Company" are extracted through multiple convolutional layers and pooling layers of the convolutional network. The extracted watermark information features are then input into the decoder in the audio generation model. The watermark information features are deconvolved through the transposed convolutional layer of the decoder to reconstruct the decoded watermark information of "Copyright, XX Company".

[0108] It is worth mentioning that the above-mentioned examples are also applicable to the medical field. For example, a hospital's watermark information can be embedded into a model, and initial voice information with the hospital's watermark information can be generated. The system can then verify whether the watermark information in the initial voice information can be extracted normally and whether the extracted watermark information is the same as the target watermark information.

[0109] In this embodiment, the initial speech information is decoded by a convolutional network and a decoder to reconstruct the decoded watermark information of the initial speech information, ensuring that the watermark information in the initial speech information can be completely extracted and improving the robustness of watermark information embedding.

[0110] In one embodiment, the step of extracting watermark information features from the initial speech information through the convolutional network of the audio generation model, and reconstructing the watermark information features through the decoder of the audio generation model to obtain decoded watermark information, includes:

[0111] The initial speech information is convolved through multiple convolutional layers of the convolutional network to obtain intermediate watermark features.

[0112] The intermediate watermark features are reduced in dimensionality by the pooling layer of the convolutional network to obtain the watermark information features.

[0113] The watermark information features are deconvolutionally processed by the transposed convolutional layer of the decoder to obtain the initial watermark information.

[0114] The resolution of the initial watermark information is increased by the upsampling layer of the decoder to obtain the decoded watermark information.

[0115] In this embodiment, initial speech information is input into the convolutional network through its input layer. Multiple convolutional layers of the network perform convolution operations on the initial speech information to extract intermediate watermark features. Each convolutional layer uses a set of convolutional kernels to perform convolution operations on the initial speech information, thereby extracting the intermediate watermark features. Then, pooling layers perform dimensionality reduction on these intermediate watermark features to obtain the watermark information features. The main function of the pooling layer is to reduce the dimensionality of the intermediate watermark features while retaining important features. To reduce computation and avoid overfitting, the intermediate watermark features are refined into watermark information features after pooling. The watermark information features are then fed into the transposed convolutional layer of the decoder through the input layer. The transposed convolutional layer performs deconvolution on the watermark information features to obtain the initial watermark information. Finally, the resolution of the initial watermark information is improved through the upsampling layer of the decoder to obtain the decoded watermark information. The main function of the upsampling layer is to increase the resolution of the feature map, making it closer to the size of the original watermark information. After processing by the upsampling layer, the decoded watermark information is finally reconstructed.

[0116] In this embodiment, by working together with a convolutional network and a decoder, the watermark features in the initial speech information can be extracted efficiently and accurately, and then reconstructed into the decoded watermark information, thus improving the efficiency of watermark information reconstruction.

[0117] Step S206: Compare the decoded watermark information with the target watermark information to obtain a similarity value. When the similarity value is greater than a preset similarity threshold, output the initial speech information as the target speech information.

[0118] In this embodiment, the decoded watermark information reconstructed in step S205 is compared with the target watermark information stored in the preset watermark information database in step S201. The similarity value between the decoded watermark information and the target watermark information is calculated. Then, the similarity value is compared with a preset similarity threshold. When the similarity value is greater than the preset similarity threshold, the initial speech information is output as the target speech information. The similarity calculation method includes, but is not limited to, structural similarity calculation methods.

[0119] Following the above feasible example A, the decoded watermark information "Copyright owned by XX Company" is compared with the target watermark information "Copyright owned by XX Company" stored in the preset watermark information database. The similarity value between the two is calculated. When the similarity value is greater than the preset similarity threshold, the initial voice information is output as the target voice information.

[0120] In this embodiment, by embedding watermark information when generating initial voice information and verifying the embedded watermark information, the method can support multiple watermark formats and embedding methods, thus improving its versatility and compatibility.

[0121] It should be emphasized that, in order to further ensure the privacy and security of the aforementioned target watermark information, the aforementioned target watermark information can also be stored in a blockchain node.

[0122] The blockchain referred to in this application is a novel application model of computer technologies such as distributed data storage, peer-to-peer transmission, consensus mechanisms, and encryption algorithms. Essentially, a blockchain is a decentralized database, a chain of data blocks linked together using cryptographic methods. Each data block contains information about a batch of network transactions, used to verify the validity of the information (anti-counterfeiting) and generate the next block. A blockchain can include an underlying blockchain platform, a platform product service layer, and an application service layer.

[0123] The embodiments of this application can acquire and process relevant data based on artificial intelligence technology. Artificial intelligence (AI) refers to the theories, methods, technologies, and application systems that use digital computers or machines controlled by digital computers to simulate, extend, and expand human intelligence, perceive the environment, acquire knowledge, and use that knowledge to obtain optimal results.

[0124] Foundational technologies for artificial intelligence generally include sensors, dedicated AI chips, cloud computing, distributed storage, big data processing, operating / interactive systems, and mechatronics. AI software technologies mainly encompass computer vision, robotics, biometrics, speech processing, natural language processing, and machine learning / deep learning.

[0125] Those skilled in the art will understand that all or part of the processes in the methods of the above embodiments can be implemented by instructing related hardware with computer-readable instructions. These computer-readable instructions can be stored in a computer-readable storage medium. When executed, the program can include the processes of the embodiments of the above methods. The aforementioned storage medium can be a non-volatile storage medium such as a magnetic disk, optical disk, or read-only memory (ROM), or random access memory (RAM).

[0126] It should be understood that although the steps in the flowcharts of the accompanying figures are shown sequentially as indicated by the arrows, these steps are not necessarily executed in the order indicated by the arrows. Unless explicitly stated herein, there is no strict order restriction on the execution of these steps, and they can be executed in other orders. Moreover, at least some steps in the flowcharts of the accompanying figures may include multiple sub-steps or multiple stages. These sub-steps or stages are not necessarily completed at the same time, but can be executed at different times, and their execution order is not necessarily sequential, but can be performed alternately or in turn with other steps or at least some of the sub-steps or stages of other steps.

[0127] Further reference Figure 3 As a response to the above Figure 2 To implement the method shown, this application provides an embodiment of a speech synthesis device that embeds watermark information. This device embodiment is similar to... Figure 2 Corresponding to the method embodiments shown, the device can be specifically applied to various computer devices.

[0128] like Figure 3 As shown, the speech synthesis device 300 with embedded watermark information described in this embodiment includes: an acquisition module 301, an encoding module 302, an embedding module 303, a generation module 304, a reconstruction module 305, and a comparison module 306. Wherein:

[0129] The acquisition module 301 is used to acquire the target watermark information that needs to be embedded.

[0130] Encoding module 302 is used to encode the target watermark information through the latent watermark encoder of the audio generation model to obtain watermark information encoding;

[0131] In one embodiment, the encoding module further includes:

[0132] The alignment submodule is used to fill or truncate the target watermark information according to the number of neurons in the input layer of the potential watermark encoder to obtain at least one watermark sequence.

[0133] The first input submodule is used to input at least one of the watermark sequences into the potential watermark encoder through the input layer of the potential watermark encoder;

[0134] The transformation submodule is used to transform at least one of the watermark sequences into a vector corresponding to a high-dimensional latent representation through the hidden layer of the latent watermark encoder.

[0135] The first output submodule is used to output the vector corresponding to the high-dimensional latent representation through the output layer of the latent watermark encoder to obtain the watermark information encoding.

[0136] In one embodiment, the apparatus further includes:

[0137] The noise acquisition module is used to acquire random noise of a preset type and to preset the noise intensity of the random noise;

[0138] The noise addition module is used to add the random noise into the watermark information encoding according to the preset type of random noise and the corresponding noise intensity.

[0139] The embedding module 303 is used to use the intermediate latent space of the audio generation model as the watermark embedding position, encode and embed the watermark information according to the watermark embedding position, and combine it with the original audio features to obtain the target latent representation of the target watermark information.

[0140] In one embodiment, the embedding module further includes:

[0141] An audio signal acquisition submodule is used to acquire the original audio signal and extract the original audio features from the original audio signal;

[0142] An encoding embedding submodule is used to encode and embed the watermark information into the watermark embedding position;

[0143] The combined submodule is used to combine the watermark information encoding embedded at the watermark embedding location with the original audio features to generate the target potential representation of the target watermark information.

[0144] The generation module 304 is used to process the target latent representation through the generator of the audio generation model to generate initial speech information.

[0145] In one embodiment, the generation module further includes:

[0146] The second input submodule is used to input the target latent representation into the generator through the input layer of the generator;

[0147] The feature fusion submodule is used to extract multiple features of the latent representation of the target using the transform layer of the generator, and to perform feature fusion operation on the multiple features to obtain intermediate features;

[0148] The conversion submodule uses the decoding layer of the generator to convert the intermediate features and generate initial speech information;

[0149] The second output submodule is used to output the initial speech information using the output layer of the generator.

[0150] In one embodiment, the apparatus further includes:

[0151] The training set acquisition module is used in step S2041 to acquire the audio training dataset with watermark information and use the audio training dataset as the training set.

[0152] The loss function combination module is used in step S2042 to obtain the perceived loss function and define a preset loss function, and to combine the preset loss function with the perceived loss function to obtain a combined loss function.

[0153] An initial generator construction module is used in step S2043 to construct an initial generator, train the initial generator using the training set, and generate initial training speech information.

[0154] The loss function calculation module is used in step S2044 to calculate the combined loss function based on the initial training speech information and obtain the combined loss function calculation result.

[0155] The parameter update module is used in step S2045 to update the parameters of the initial generator based on the calculation result of the combined loss function.

[0156] The iterative module is used to repeatedly execute steps S2042 to S2045 until the calculated result of the combined loss function is less than a preset loss threshold, thereby obtaining the trained generator.

[0157] The reconstruction module 305 is used to extract the watermark information features of the initial speech information through the convolutional network of the audio generation model, and reconstruct the watermark information features through the decoder of the audio generation model to obtain the decoded watermark information.

[0158] In one embodiment, the reconstruction module further includes:

[0159] The convolutional submodule is used to perform convolution operations on the initial speech information through multiple convolutional layers of the convolutional network to obtain intermediate watermark features;

[0160] The dimensionality reduction submodule is used to perform dimensionality reduction on the intermediate watermark features through the pooling layer of the convolutional network to obtain the watermark information features;

[0161] The deconvolution submodule is used to perform a deconvolution operation on the watermark information features through the transposed convolutional layer of the decoder to obtain the initial watermark information.

[0162] The upsampling submodule is used to improve the resolution of the initial watermark information through the upsampling layer of the decoder to obtain the decoded watermark information.

[0163] The comparison module 306 is used to compare the decoded watermark information with the target watermark information to obtain a similarity value. When the similarity value is greater than a preset similarity threshold, the initial speech information is output as the target speech information.

[0164] In this embodiment, by obtaining the target watermark information to be embedded, the specific content to be embedded can be clearly defined, which is beneficial for subsequent watermark processing, embedding and verification work.

[0165] Encoding the target watermark information through the latent watermark encoder of the audio generation model can improve the robustness of the target watermark information, prevent the target watermark information from being tampered with and attacked, and improve the reliability of the target watermark information.

[0166] By embedding the target watermark information into the intermediate latent space of the audio generation model, the watermark information does not need to be embedded after the audio is generated, which improves the concealment and embedding efficiency of the watermark. When embedding the watermark information, the impact on the audio quality can also be adjusted to improve the audio quality.

[0167] By processing the target's latent representation through a generator to generate initial speech information, the impact of watermark embedding on the quality of the initial speech signal can be reduced, meeting the requirements of high-fidelity speech synthesis. Furthermore, the audio generation model with embedded watermark carries a unique identifier each time it generates initial speech information, making it easy to trace the source and preventing the content from being forged or tampered with.

[0168] The initial speech information is decoded by a convolutional network and a decoder, and the decoded watermark information of the initial speech information is reconstructed. This ensures that the watermark information in the initial speech information can be completely extracted, and improves the robustness of watermark information embedding.

[0169] By embedding watermark information when generating initial speech information and verifying the embedded watermark information, this method can support multiple watermark formats and embedding methods, thus improving its versatility and compatibility.

[0170] To address the aforementioned technical problems, embodiments of this application also provide a device (computer device). Please refer to [link / reference] for details. Figure 4 , Figure 4 This is a basic structural block diagram of the computer device in this embodiment.

[0171] The computer device 4 includes a memory 41, a processor 42, and a network interface 43 that are interconnected via a system bus. It should be noted that only the computer device 4 with memory 41, processor 42, and network interface 43 is shown in the figure; however, it should be understood that it is not required to implement all the components shown, and more or fewer components can be implemented alternatively. Those skilled in the art will understand that the computer device described here is a device capable of automatically performing numerical calculations and / or information processing according to pre-set or stored instructions, and its hardware includes, but is not limited to, microprocessors, application-specific integrated circuits (ASICs), field-programmable gate arrays (FPGAs), digital signal processors (DSPs), embedded devices, etc.

[0172] The computer device can be a desktop computer, laptop, handheld computer, or cloud server, etc. The computer device can interact with the user via a keyboard, mouse, remote control, touchpad, or voice control.

[0173] The memory 41 includes at least one type of readable storage medium, including flash memory, hard disk, multimedia card, card-type memory (e.g., SD or DX memory), random access memory (RAM), static random access memory (SRAM), read-only memory (ROM), electrically erasable programmable read-only memory (EEPROM), programmable read-only memory (PROM), magnetic memory, magnetic disk, optical disk, etc. In some embodiments, the memory 41 may be an internal storage unit of the computer device 4, such as the hard disk or memory of the computer device 4. In other embodiments, the memory 41 may also be an external storage device of the computer device 4, such as a plug-in hard disk, smart media card (SMC), secure digital (SD) card, flash card, etc., equipped on the computer device 4. Of course, the memory 41 may also include both the internal storage unit and its external storage device of the computer device 4. In this embodiment, the memory 41 is typically used to store the operating system and various application software installed on the computer device 4, such as computer-readable instructions for speech synthesis methods embedding watermark information. In addition, the memory 41 can also be used to temporarily store various types of data that have been output or will be output.

[0174] In some embodiments, the processor 42 may be a central processing unit (CPU), controller, microcontroller, microprocessor, or other data processing chip. The processor 42 is typically used to control the overall operation of the computer device 4. In this embodiment, the processor 42 is used to execute computer-readable instructions stored in the memory 41 or to process data, such as executing computer-readable instructions for the speech synthesis method embedding watermark information.

[0175] The network interface 43 may include a wireless network interface or a wired network interface, which is typically used to establish communication connections between the computer device 4 and other electronic devices.

[0176] In the implementation of the electronic device of this application, by obtaining the target watermark information to be embedded, the specific content to be embedded can be clearly identified, which is beneficial to subsequent watermark processing, embedding and verification work;

[0177] Encoding the target watermark information through the latent watermark encoder of the audio generation model can improve the robustness of the target watermark information, prevent the target watermark information from being tampered with and attacked, and improve the reliability of the target watermark information.

[0178] By embedding the target watermark information into the intermediate latent space of the audio generation model, the watermark information does not need to be embedded after the audio is generated, which improves the concealment and embedding efficiency of the watermark. When embedding the watermark information, the impact on the audio quality can also be adjusted to improve the audio quality.

[0179] By processing the target's latent representation through a generator to generate initial speech information, the impact of watermark embedding on the quality of the initial speech signal can be reduced, meeting the requirements of high-fidelity speech synthesis. Furthermore, the audio generation model with embedded watermark carries a unique identifier each time it generates initial speech information, making it easy to trace the source and preventing the content from being forged or tampered with.

[0180] The initial speech information is decoded by a convolutional network and a decoder, and the decoded watermark information of the initial speech information is reconstructed. This ensures that the watermark information in the initial speech information can be completely extracted, and improves the robustness of watermark information embedding.

[0181] By embedding watermark information when generating initial speech information and verifying the embedded watermark information, this method can support multiple watermark formats and embedding methods, thus improving its versatility and compatibility.

[0182] This application also provides another embodiment, namely, providing a storage medium (computer-readable storage medium) storing computer-readable instructions that can be executed by at least one processor to cause the at least one processor to perform the steps of the speech synthesis method for embedding watermark information as described above.

[0183] In the implementation of the computer-readable storage medium of this application, by obtaining the target watermark information to be embedded, the specific content to be embedded can be clearly identified, which is beneficial to subsequent watermark processing, embedding and verification work.

[0184] Encoding the target watermark information through the latent watermark encoder of the audio generation model can improve the robustness of the target watermark information, prevent the target watermark information from being tampered with and attacked, and improve the reliability of the target watermark information.

[0185] By embedding the target watermark information into the intermediate latent space of the audio generation model, the watermark information does not need to be embedded after the audio is generated, which improves the concealment and embedding efficiency of the watermark. When embedding the watermark information, the impact on the audio quality can also be adjusted to improve the audio quality.

[0186] By processing the target's latent representation through a generator to generate initial speech information, the impact of watermark embedding on the quality of the initial speech signal can be reduced, meeting the requirements of high-fidelity speech synthesis. Furthermore, the audio generation model with embedded watermark carries a unique identifier each time it generates initial speech information, making it easy to trace the source and preventing the content from being forged or tampered with.

[0187] The initial speech information is decoded by a convolutional network and a decoder, and the decoded watermark information of the initial speech information is reconstructed. This ensures that the watermark information in the initial speech information can be completely extracted, and improves the robustness of watermark information embedding.

[0188] By embedding watermark information when generating initial speech information and verifying the embedded watermark information, this method can support multiple watermark formats and embedding methods, thus improving its versatility and compatibility.

[0189] The software tools or components not belonging to our company that appear in the embodiments of this application are merely examples and do not represent actual use.

[0190] Through the above description of the embodiments, those skilled in the art can clearly understand that the methods of the above embodiments can be implemented by means of software plus necessary general-purpose hardware platforms. Of course, they can also be implemented by hardware, but in many cases the former is a better implementation method. Based on this understanding, the technical solution of this application, in essence, or the part that contributes to the prior art, can be embodied in the form of a software product. This computer software product is stored in a storage medium (such as ROM / RAM, magnetic disk, optical disk), and includes several instructions to cause a terminal device (which may be a mobile phone, computer, server, air conditioner, or network device, etc.) to execute the methods described in the various embodiments of this application.

[0191] Obviously, the embodiments described above are only some embodiments of this application, not all embodiments. The accompanying drawings show preferred embodiments of this application, but do not limit the patent scope of this application. This application can be implemented in many different forms; rather, the purpose of providing these embodiments is to provide a more thorough and comprehensive understanding of the disclosure of this application. Although this application has been described in detail with reference to the foregoing embodiments, those skilled in the art can still modify the technical solutions described in the foregoing specific embodiments, or make equivalent substitutions for some of the technical features. Any equivalent structures made using the content of this application's specification and drawings, directly or indirectly applied to other related technical fields, are similarly within the scope of patent protection of this application.

Claims

1. A speech synthesis method embedding watermark information, characterized in that, The method includes: Obtain the target watermark information to be embedded; The target watermark information is encoded by a latent watermark encoder of an audio generation model to obtain the watermark information encoding. The intermediate latent space of the audio generation model is used as the watermark embedding position. The watermark information is encoded and embedded according to the watermark embedding position. Combined with the original audio features, the target latent representation of the target watermark information is obtained. The model architecture of the audio generation model is analyzed, and multiple candidate intermediate latent spaces of the audio generation model are analyzed from the model architecture. According to actual needs, one intermediate latent space is selected from multiple candidate intermediate latent spaces as the watermark embedding position. The candidate intermediate latent space is at the output end or internal processing layer of each component of the audio generation model. The generator of the audio generation model processes the latent representation of the target to generate initial speech information. The watermark information features of the initial speech information are extracted by the convolutional network of the audio generation model, and the watermark information features are reconstructed by the decoder of the audio generation model to obtain the decoded watermark information. The decoded watermark information is compared with the target watermark information to obtain a similarity value. When the similarity value is greater than a preset similarity threshold, the initial speech information is output as the target speech information.

2. The speech synthesis method for embedding watermark information as described in claim 1, characterized in that, The latent watermark encoder using the audio generation model encodes the target watermark information to obtain the watermark information encoding, including: The target watermark information is padded or truncated based on the number of neurons in the input layer of the potential watermark encoder to obtain at least one watermark sequence. At least one of the watermark sequences is input into the latent watermark encoder through the input layer of the latent watermark encoder; The hidden layer of the latent watermark encoder converts at least one of the watermark sequences into a vector corresponding to a high-dimensional latent representation. The high-dimensional latent representation is output through the output layer of the latent watermark encoder to obtain the watermark information encoding.

3. The speech synthesis method embedding watermark information as described in claim 1 or 2, characterized in that, After the target watermark information is encoded by the latent watermark encoder of the audio generation model to obtain the watermark information encoding, the method further includes: Obtain random noise of a preset type, and preset the noise intensity of the random noise; The random noise is added to the watermark information encoding according to the preset type of random noise and the corresponding noise intensity.

4. The speech synthesis method for embedding watermark information as described in claim 1, characterized in that, The step of encoding and embedding the watermark information according to the watermark embedding position, and combining it with the original audio features to obtain the target potential representation of the target watermark information, includes: Acquire the original audio signal and extract the original audio features from the original audio signal; The watermark information is encoded and embedded into the watermark embedding position; The watermark information encoding embedded at the watermark embedding location is combined with the original audio features to generate the target potential representation of the target watermark information.

5. The speech synthesis method for embedding watermark information as described in claim 1, characterized in that, The step of processing the target latent representation through the generator of the audio generation model to generate initial speech information includes: The potential representation of the target is input into the generator through the input layer of the generator; The generator's transform layer extracts multiple features of the target's latent representation, and a feature fusion operation is performed on these multiple features to obtain intermediate features; The intermediate features are converted using the decoding layer of the generator to generate initial speech information; The initial speech information is output using the output layer of the generator.

6. The speech synthesis method for embedding watermark information as described in claim 1 or 5, characterized in that, Before processing the target latent representation through the generator of the audio generation model to generate initial speech information, the method further includes: Step S2041: Obtain the audio training dataset with watermark information, and use the audio training dataset as the training set; Step S2042: Obtain the perceptual loss function and define a preset loss function, and combine the preset loss function with the perceptual loss function to obtain a combined loss function; Step S2043: Construct an initial generator, train the initial generator using the training set, and generate initial training speech information; Step S2044: Calculate the combined loss function based on the initial training speech information to obtain the calculation result of the combined loss function; Step S2045: Update the parameters of the initial generator based on the calculation result of the combined loss function; Repeat steps S2042 to S2045 until the calculated result of the combined loss function is less than the preset loss threshold, and obtain the generator that has been trained.

7. The speech synthesis method for embedding watermark information as described in claim 1, characterized in that, The step of extracting watermark information features from the initial speech information using the convolutional network of the audio generation model, and reconstructing the watermark information features using the decoder of the audio generation model to obtain decoded watermark information, includes: The initial speech information is convolved through multiple convolutional layers of the convolutional network to obtain intermediate watermark features. The intermediate watermark features are reduced in dimensionality by the pooling layer of the convolutional network to obtain the watermark information features. The watermark information features are deconvolutionally processed by the transposed convolutional layer of the decoder to obtain the initial watermark information. The resolution of the initial watermark information is increased by the upsampling layer of the decoder to obtain the decoded watermark information.

8. A speech synthesis device with embedded watermark information, characterized in that, The device includes: The acquisition module is used to acquire the target watermark information that needs to be embedded. The encoding module is used to encode the target watermark information through the latent watermark encoder of the audio generation model to obtain the watermark information encoding; An embedding module is used to use the intermediate latent space of the audio generation model as the watermark embedding position, encode and embed the watermark information according to the watermark embedding position, and combine it with the original audio features to obtain the target latent representation of the target watermark information. In this module, the model architecture of the audio generation model is analyzed, and multiple candidate intermediate latent spaces of the audio generation model are analyzed from the model architecture. According to actual needs, one intermediate latent space is selected from the multiple candidate intermediate latent spaces as the watermark embedding position. The candidate intermediate latent space is located at the output end or internal processing layer of each component of the audio generation model. The generation module is used to process the target latent representation through the generator of the audio generation model to generate initial speech information; The reconstruction module is used to extract watermark information features from the initial speech information through the convolutional network of the audio generation model, and reconstruct the watermark information features through the decoder of the audio generation model to obtain decoded watermark information. The comparison module is used to compare the decoded watermark information with the target watermark information to obtain a similarity value. When the similarity value is greater than a preset similarity threshold, the initial speech information is output as the target speech information.

9. A computer device, characterized in that, The computer device includes: At least one processor; and, A memory communicatively connected to the at least one processor; wherein, The memory stores instructions that can be executed by the at least one processor to enable the at least one processor to perform the speech synthesis method for embedding watermark information as described in any one of claims 1 to 7.

10. A computer-readable storage medium storing a computer program, characterized in that, When the computer program is executed by the processor, it implements the speech synthesis method with embedded watermark information as described in any one of claims 1 to 7.

Citation Information

Patent Citations

  • Inaudible watermark enabled text-to-speech framework

    CN112767913A

  • Model watermark embedding method and device, computer equipment and storage medium

    CN116881871A