Watermark embedding method and device, electronic equipment and storage medium

By building a preset watermark embedding model, using watermark samples to conditionally generate and extract audio data, the problem of interference between watermarks and audio data in the prior art is solved, and the auditory quality and fidelity of watermark audio data are improved.

CN120089146APending Publication Date: 2025-06-03PING AN TECH (SHENZHEN) CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510271925.8
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-03-07
Publication Date
2025-06-03

AI Technical Summary

Technical Problem

In the prior art, audio watermarks and audio data interfere with each other, resulting in a decrease in the auditory quality of watermark audio data and the fidelity of watermark audio data cannot be effectively improved.

Method used

By constructing a preset watermark embedding model including a first decoder, a watermark extractor and a second decoder, the audio sample vector representation is conditionally generated using the watermark samples, the watermark is extracted, and the model parameters are adjusted through data reconstruction technology to ensure that the watermark data does not interfere with the feature extraction of the audio data.

Benefits of technology

The auditory quality and fidelity of watermark audio data are improved, so that the watermark data can be effectively embedded in the audio without affecting the original characteristics of the audio.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120089146A_ABST
    Figure CN120089146A_ABST
Patent Text Reader

Abstract

The embodiment of the invention provides a watermark embedding method and device, electronic equipment and a storage medium, which can be applied to the fields of financial science and technology and medical science and technology, and the watermark embedding method comprises the following steps: obtaining a preset watermark embedding model comprising a first decoder, a watermark extractor and a second decoder; performing condition generation on the watermark sample and the audio sample vector representation through a first decoder to obtain watermark audio data; performing watermark extraction on the watermark audio data through a watermark extractor to obtain a reference watermark; performing data reconstruction on the audio sample vector representation through a second decoder to obtain audio reconstruction data; and based on the reference watermark, the watermark sample, the audio reconstruction data and the training sample, adjusting parameters of a preset watermark embedding model to obtain a target watermark embedding model so as to carry out watermark embedding processing. According to the embodiment of the invention, the auditory quality of the watermark audio data can be improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present application relates to the technical field of speech processing, and can be applied to the fields of fintech and medical technology. In particular, it relates to a watermark embedding method and device, an electronic device, and a storage medium. Background Art

[0002] Audio watermarking technology refers to embedding imperceptible watermark data into audio data, and the quality of the audio after embedding the watermark should not decrease significantly. The use of audio watermarks enables audio to be traced and authenticated while still maintaining audibility, and is widely applied in the fields of fintech and medical technology. For example, in scenarios such as telephone banking, conference recording, and identity verification, users' voice commands involve sensitive information (such as transaction information, account information). By embedding watermark information (user identity information, transaction identifier), it can be detected whether the voice has been tampered with (such as clipped, inserted with false commands), and the source of voice data leakage can be traced. Another example is in the remote diagnosis scenario, where the patient's voice is an important basis for diagnosis. The remote diagnosis system judges the patient's health status based on the patient's voice. The remote diagnosis system often transmits the patient's voice to the doctor's end through a public network, which has network security risks. By embedding watermark information (patient identity information, doctor identity information, electronic medical record information), it can be detected whether the voice is complete and whether false medical information has been injected.

[0003] In related technologies, the audio data and the watermark data are directly encoded into a watermark audio vector by an encoder, and the watermark audio vector is decoded into watermark audio data by a decoder. The watermark data and the audio data interfere with each other, reducing the auditory quality of the watermark audio data. Therefore, how to improve the auditory quality of the watermark audio data has become an urgent technical problem to be solved. Summary of the Invention

[0004] The main purpose of the embodiments of the present application is to propose a watermark embedding method and device, an electronic device, and a storage medium, aiming to improve the auditory quality of the watermark audio data.

[0005] To achieve the above object, a first aspect of the embodiments of the present application proposes a watermark embedding method, and the method includes:

[0006] Obtain a watermark sample and obtain a target watermark;

[0007] Obtain the audio sample vector representation of the training sample and obtain the target audio vector representation of the target input data; the training sample is audio or text, and the target input data is audio or text;

[0008] Obtain a preset watermark embedding model, and the preset watermark embedding model includes a first decoder, a watermark extractor, and a second decoder;

[0009] Generate the watermarked audio data by conditional generation of the watermark sample and the audio sample vector representation through the first decoder; wherein, the watermark sample serves as the control condition of the first decoder;

[0010] Extract the watermark from the watermarked audio data through the watermark extractor to obtain the reference watermark;

[0011] Reconstruct the data of the audio sample vector representation through the second decoder to obtain the audio reconstruction data;

[0012] Adjust the parameters of the preset watermark embedding model based on the reference watermark, the watermark sample, the audio reconstruction data, and the training sample to obtain the target watermark embedding model;

[0013] Perform watermark embedding processing on the target watermark and the target audio vector representation according to the target watermark embedding model.

[0014] To achieve the above object, a second aspect of the embodiments of the present application proposes a watermark embedding device, the device includes:

[0015] A watermark acquisition module, configured to acquire a watermark sample and a target watermark;

[0016] An audio vector representation acquisition module, configured to acquire the audio sample vector representation of the training sample and the target audio vector representation of the target input data; the training sample is audio or text, and the target input data is audio or text;

[0017] A model acquisition module, configured to acquire a preset watermark embedding model, and the preset watermark embedding model includes a first decoder, a watermark extractor, and a second decoder;

[0018] A conditional generation module, configured to generate the watermarked audio data by conditional generation of the watermark sample and the audio sample vector representation through the first decoder; wherein, the watermark sample serves as the control condition of the first decoder;

[0019] A watermark extraction module, configured to extract the watermark from the watermarked audio data through the watermark extractor to obtain the reference watermark;

[0020] A data reconstruction module, configured to reconstruct the data of the audio sample vector representation through the second decoder to obtain the audio reconstruction data;

[0021] A parameter adjustment module, configured to adjust the parameters of the preset watermark embedding model based on the reference watermark, the watermark sample, the audio reconstruction data, and the training sample to obtain the target watermark embedding model;

[0022] A watermark embedding module, configured to perform watermark embedding processing on the target watermark and the target audio vector representation according to the target watermark embedding model.

[0023] To achieve the above object, a third aspect of the embodiments of the present application proposes an electronic device, which includes a memory and a processor. The memory stores a computer program, and when the processor executes the computer program, the watermark embedding method of the first aspect is implemented.

[0024] To achieve the above object, a fourth aspect of the embodiments of the present application proposes a computer-readable storage medium, which stores a computer program, and when the computer program is executed by a processor, the watermark embedding method of the first aspect is implemented.

[0025] The watermark embedding method, device, electronic device and storage medium proposed by the present application construct a preset watermark embedding model including a first decoder, a watermark extractor and a second decoder; obtain the vector representation of the audio sample, and use the watermark sample to perform conditional generation on the vector representation of the audio sample through the first decoder to obtain watermarked audio data; extract the watermark from the watermarked audio data through the watermark extractor to obtain a reference watermark; reconstruct the data of the audio sample vector representation through the second decoder to obtain audio reconstruction data; based on the reference watermark, watermark sample, audio reconstruction data and training samples, train the preset watermark embedding model to enable the model to learn the mapping relationship between the vector representation of the audio sample and the watermarked audio under a given condition (i.e., the watermark sample), and the mapping relationship between the vector representation of the audio sample and the audio. The watermark data will not interfere with the feature extraction of the audio data (i.e., the watermark sample will not interfere with the generation of the vector representation of the audio sample), thereby improving the auditory quality of the watermarked audio data and improving the watermark audio fidelity. Description of the Drawings

[0026] Figure 1 is a flowchart of the watermark embedding method provided by the embodiments of the present application;

[0027] Figure 2 is Figure 1 a flowchart of step S102 in

[0028] Figure 3 is Figure 1 another flowchart of step S102 in

[0029] Figure 4 is Figure 1 a flowchart of step S104 in

[0030] Figure 5 is Figure 1 a flowchart of step S107 in

[0031] Figure 6 is another flowchart of the watermark embedding method provided by the embodiments of the present application;

[0032] Figure 7It is a schematic structural diagram of the watermark embedding device provided by an embodiment of the present application;

[0033] Figure 8 It is a schematic hardware structure diagram of the electronic device provided by an embodiment of the present application. Detailed implementation manners

[0034] In order to make the objectives, technical solutions and advantages of the present application clearer, the present application will be further described in detail below with reference to the accompanying drawings and embodiments. It should be understood that the specific embodiments described herein are only used to explain the present application and are not used to limit the present application.

[0035] It should be noted that although functional module division is performed in the device schematic diagram and the logical order is shown in the flowchart, in some cases, the steps shown or described may be executed in a different module division in the device or a different order in the flowchart. Terms such as "first" and "second" in the description, claims and the above-mentioned drawings are used to distinguish similar objects and do not necessarily need to be used to describe a specific order or sequence.

[0036] Unless otherwise defined, all technical and scientific terms used herein have the same meaning as commonly understood by those skilled in the technical field to which the present application belongs. The terms used herein are only for the purpose of describing the embodiments of the present application and are not intended to limit the present application.

[0037] First, several nouns involved in the present application are analyzed:

[0038] Artificial intelligence (AI): It is a new technical science that studies, develops theories, methods, technologies and application systems for simulating, extending and expanding human intelligence; artificial intelligence is a branch of computer science. Artificial intelligence attempts to understand the essence of intelligence and produce a new intelligent machine that can respond in a way similar to human intelligence. The research in this field includes robots, speech recognition, image recognition, natural language processing and expert systems, etc. Artificial intelligence can simulate the information process of human consciousness and thinking. Artificial intelligence also uses digital computers or machines controlled by digital computers to simulate, extend and expand human intelligence, perceive the environment, acquire knowledge and use knowledge to obtain the best results of theories, methods, technologies and application systems.

[0039] Audio processing (AP): Audio processing refers to the process of analyzing, modifying and optimizing audio signals through a series of technical means to meet specific requirements. Audio processing technology is widely used in fields such as audio compression, audio enhancement, audio generation and audio reconstruction, and is one of the core technologies of modern audio applications.

[0040] Audio Watermarking (AW): It is a technology that embeds watermark information into an audio signal without disturbing the quality or perceptibility of the original audio. The use of audio watermarks enables audio content to be traced and authenticated while still maintaining its listenability, providing an effective means for the management of digital content. Audio watermarking technology is widely used in scenarios such as audio copyright protection, broadcast monitoring, and multimedia data hiding.

[0041] Watermark Extraction (WE): It refers to the process of extracting the embedded watermark information from digital media (such as images, audio, video, etc.). The purpose is to recover the watermark information from the watermarked media through specific algorithms. Watermark extraction is an important part of digital watermarking technology and has a wide range of applications in multiple fields. In terms of copyright protection, by embedding copyright information in digital media, illegal copying and piracy can be effectively prevented; in terms of content authentication, digital watermarking technology can be used to verify the authenticity and integrity of digital media content and prevent tampering and forgery.

[0042] Decoder: It is a device or software program used to restore encoded data or signals to the original information. In the field of audio processing, the decoder decodes the encoded and compressed audio data into playable audio signals.

[0043] Data Reconstruction (DR): Data reconstruction refers to the process of restoring processed, compressed, or encoded data to its original or near-original state through specific algorithms or models. It is an important part of data processing and machine learning, especially widely used in data dimensionality reduction, feature extraction, anomaly detection, and generation tasks. In the field of artificial intelligence, data reconstruction is a key step in training machine learning models. Through data reconstruction, the original data can be transformed into more valuable information assets, thereby enhancing the usability and value of the data.

[0044] Conditional Generation (CG): Unconditional generation and conditional generation are a pair of relative concepts. Unconditional generation means that the model can spontaneously generate new data without any input or input pattern. Conditional generation requires the model to generate data under the control of specific conditions or labels. The conditions can be text (such as a description statement), an image (such as a sketch), or a label (such as a category). For example, in image generation, the condition can be a specific category label, and the goal is to generate an image belonging to that category; in natural language processing, conditional generation can be used in text generation tasks, such as generating an article according to a given theme or style.

[0045] Based on this, the embodiments of the present application provide a watermark embedding method and apparatus, an electronic device, and a storage medium, aiming to improve the auditory quality of watermarked audio data.

[0046] The watermark embedding method and apparatus, electronic device, and storage medium provided by the embodiments of the present application are specifically described through the following embodiments. First, a watermark embedding method in the embodiments of the present application is described.

[0047] The embodiments of the present application can acquire and process relevant data based on artificial intelligence technology. Among them, artificial intelligence (AI) is a theory, watermark embedding method, technology, and application system that uses a digital computer or a machine controlled by a digital computer to simulate, extend, and expand human intelligence, perceive the environment, acquire knowledge, and use knowledge to obtain the best results.

[0048] Artificial intelligence basic technologies generally include technologies such as sensors, dedicated artificial intelligence chips, cloud computing, distributed storage, big data processing technology, operation / interaction systems, and mechatronics. Artificial intelligence software technologies mainly include several major directions such as computer vision technology, robotics, biometric technology, speech processing technology, natural language processing technology, and machine learning / deep learning.

[0049] A watermark embedding method provided by the embodiments of the present application relates to the field of artificial intelligence technology. A watermark embedding method provided by the embodiments of the present application can be applied to a terminal, can also be applied to a server side, or can also be software running on a terminal or a server side. In some embodiments, the terminal can be a smart phone, a tablet computer, a notebook computer, a desktop computer, etc.; the server side can be configured as an independent physical server, can also be configured as a server cluster or a distributed system composed of multiple physical servers, or can also be configured as a cloud server providing basic cloud computing services such as cloud services, cloud databases, cloud computing, cloud functions, cloud storage, network services, cloud communications, middleware services, domain name services, security services, CDN, and big data and artificial intelligence platforms; the software can be an application that implements a watermark embedding method, etc., but is not limited to the above forms.

[0050] This application can be used in numerous general or specific computer system environments or configurations. For example: personal computers, server computers, handheld or portable devices, tablet devices, multi-processor systems, microprocessor-based systems, set-top boxes, programmable consumer electronic devices, network PCs, minicomputers, mainframe computers, distributed computing environments including any of the above systems or devices, and so on. This application can be described in the general context of computer-executable instructions executed by a computer, such as program modules. Generally, program modules include routines, programs, objects, components, data structures, etc. that perform specific tasks or implement specific abstract data types. This application can also be practiced in a distributed computing environment where tasks are performed by remote processing devices connected through a communication network. In a distributed computing environment, program modules can be located in local and remote computer storage media including storage devices.

[0051] It should be noted that in each specific embodiment of this application, when it comes to relevant processing based on data related to the user's identity or characteristics, such as user information, user behavior data, user historical data, and user location information, the user's permission or consent will be obtained first. Moreover, the collection, use, and processing of these data will comply with relevant laws, regulations, and standards. In addition, when this application embodiment needs to obtain the user's sensitive personal information, it will obtain the user's separate permission or separate consent through methods such as pop-up windows or redirecting to a confirmation page. After clearly obtaining the user's separate permission or separate consent, the necessary user-related data for the normal operation of this application embodiment will be obtained. The non-company software tools or components that appear in this application embodiment are only for illustrative purposes and do not represent actual use.

[0052] Figure 1 is an optional flowchart of the watermark embedding method provided by the embodiment of this application. Figure 1 The watermark embedding method in may include but is not limited to steps S101 to S108.

[0053] Step S101, obtain a watermark sample and obtain a target watermark;

[0054] Step S102, obtain the audio sample vector representation of the training sample and obtain the target audio vector representation of the target input data; the training sample is audio or text, and the target input data is audio or text;

[0055] Step S103, obtain a preset watermark embedding model, and the preset watermark embedding model includes a first decoder, a watermark extractor, and a second decoder;

[0056] Step S104: Conditionally generate watermarked audio data by using a first decoder for the watermark sample and the audio sample vector representation, where the watermark sample serves as the control condition for the first decoder.

[0057] Step S105: Extract the watermark from the watermarked audio data by using a watermark extractor to obtain a reference watermark.

[0058] Step S106: Reconstruct the data of the audio sample vector representation by using a second decoder to obtain audio reconstruction data.

[0059] Step S107: Adjust the parameters of a preset watermark embedding model based on the reference watermark, the watermark sample, the audio reconstruction data, and the training samples to obtain a target watermark embedding model.

[0060] Step S108: Perform watermark embedding processing on the target watermark and the target audio vector representation according to the target watermark embedding model.

[0061] Steps S101 to S108 illustrated in the embodiments of the present application construct a preset watermark embedding model including a first decoder, a watermark extractor, and a second decoder; obtain an audio sample vector representation, and conditionally generate watermarked audio data for the audio sample vector representation by using the watermark sample through the first decoder; extract the watermark from the watermarked audio data by using the watermark extractor to obtain a reference watermark; reconstruct the data of the audio sample vector representation by using the second decoder to obtain audio reconstruction data; train the preset watermark embedding model based on the reference watermark, the watermark sample, the audio reconstruction data, and the training samples, so that the model learns the mapping relationship between the audio sample vector representation and the watermarked audio under a given condition (i.e., the watermark sample), the mapping relationship between the audio sample vector representation and the audio, and the watermark data does not interfere with the feature extraction of the audio data (i.e., the watermark sample does not interfere with the generation of the audio sample vector representation), thereby improving the auditory quality of the watermarked audio data and improving the watermark audio fidelity.

[0062] It is easy to understand that in the embodiments of the present application, the audio sample vector representation and the target audio vector representation can be generated based on audio or based on text. For example, a text preprocessing module, a phoneme processing module, and an encoder in a text-to-speech model can be used to perform text preprocessing, phoneme conversion, and encoding on the text to obtain the above-mentioned audio sample vector representation and target audio vector representation. The text-to-speech model is used to convert the input text information into an audible speech signal. For another example, an autoencoder neural network can be used to unsupervised extract the audio sample vector representation and the target audio vector representation from the audio. The autoencoder neural network includes an encoder and a decoder. The encoder learns an effective representation of the input data, compresses the data into a low-dimensional latent space, and the decoder reconstructs the original data from it.

[0063] In the related art, usually a text-to-speech model is used to convert text into audio, and then a watermark embedding model is used to embed a watermark into the audio. This method consumes a large amount of computing resources and has poor real-time performance.

[0064] In summary, it should be particularly emphasized that in the embodiments of the present application, the target watermark embedding model takes the target watermark and the target audio vector representation as inputs for watermark embedding. The target watermark embedding model can be compatible with multiple types of decoders. For example, converting text into an audio sample vector representation can directly generate watermarked audio from the text without having to embed the watermark after converting the text into audio, reducing the computational amount and improving the real-time performance.

[0065] In step S101 of some embodiments, the watermark sample and the target watermark can be identification information, and data traceability can be performed through the identification information; the watermark sample and the target watermark can also be audio features (such as Mel spectrogram, Hilbert spectrogram), and it can be verified whether the data has been tampered with through the audio features; it is not limited thereto.

[0066] In step S102 of some embodiments, the training sample and the target input data can be user speech, music, synthetic speech, financial document text, etc., and it is not limited thereto.

[0067] In a medical scenario, the target input data is medical audio (such as a diagnostic recording) or text (such as a patient's oral description of the condition) to be embedded with a watermark, and the target watermark is medical identification information to be hidden (such as a patient ID, an institutional copyright label). The target watermark embedding model is used to achieve covert embedding and traceability protection.

[0068] In a financial scenario, the target input data is financial audio (such as a transaction recording) or text (such as a contract / transaction record) to be protected, and the target watermark is financial identification information covertly embedded (such as a transaction serial number, a customer identity identifier). The target watermark embedding model is used to achieve covert embedding and traceability protection.

[0069] Please refer to Figure 2 , in step S102 of some embodiments, when the target input data is audio, obtaining the target audio vector representation of the target input data may include, but is not limited to, steps S201 to S203:

[0070] Step S201, obtaining the target input data;

[0071] Step S202, mapping the target input data to a continuous latent space through an encoder to obtain a continuous vector representation;

[0072] Step S203, discretizing the continuous vector representation through a vector quantizer to obtain an audio vector representation.

[0073] Among them, the vector quantizer is used to map the continuous vector output by the encoder to the nearest codebook vector, realizing the discretization of the continuous vector representation. Specifically, a vector quantization variational autoencoder can be used to perform the above steps to compress the audio into a discrete audio vector representation.

[0074] Understandably, the training samples are audio. Obtaining the audio sample vector representation of the training samples may include, but is not limited to: obtaining the training samples; mapping the training samples to a continuous latent space through an encoder to obtain a continuous vector representation; and discretizing the continuous vector representation through a vector quantizer to obtain the audio sample vector representation.

[0075] After obtaining the audio sample vector representation through the vector quantizer, the embodiments of the present application can also calculate the codebook loss and the commitment loss based on the distance calculation between the encoder output and the codebook vector, and add the codebook loss and the commitment loss as regularization terms to the objective function of the above-mentioned preset watermark embedding model to optimize the total loss. Among them, the codebook loss is to make the codebook vector close to the encoder output, and the commitment loss is to make the encoder output close to the codebook vector. The two-way constraint helps to train a more stable codebook.

[0076] In steps S201 to S203 illustrated by the embodiments of the present application, the audio is mapped to a continuous vector representation through an encoder, and then the continuous vector representation is discretized through a vector quantizer. The obtained audio vector representation is a discrete representation, which is closer to human language symbols, such as phonemes in speech or notes in music, and is easier to capture the structural features in the audio, which is beneficial to subsequent conditional generation tasks.

[0077] Please refer to Figure 3 , in step S102 of some embodiments, the target input data is text. Obtaining the target audio vector representation of the target input data may include, but is not limited to, steps S301 to S303:

[0078] Step S301, obtaining the target input data;

[0079] Step S302, based on the mapping relationship between the text and the phonemes, converting the target input data into a phoneme sequence to obtain the target phoneme sequence;

[0080] Step S303, mapping the target phoneme sequence to a continuous latent space through an encoder to obtain the target audio vector representation.

[0081] Among them, the phoneme is the smallest unit in speech, and the phoneme can be determined according to the pronunciation actions in the syllable. One action corresponds to one phoneme.

[0082] Understandably, the training samples are texts. Obtaining the target audio vector representation of the training samples may include, but is not limited to: obtaining the training samples; converting the training samples into a phoneme sequence based on the mapping relationship between text and phonemes; and mapping the phoneme sequence to a continuous latent space through an encoder to obtain the audio vector representation.

[0083] In step S302 of some embodiments, the target input data can be converted into a phoneme sequence through a phoneme dictionary; or the target input data can be converted into a phoneme sequence through a pre-trained language model; and the like.

[0084] Steps S301 to S303 illustrated in the embodiments of the present application convert the text into a phoneme sequence, and then map the target phoneme sequence to an audio vector representation through an encoder, which can more accurately capture the correspondence between the text and the audio, and is beneficial to subsequent conditional generation tasks.

[0085] In step S103 of some embodiments, the preset watermark embedding model may further include a decoder, and the audio sample vector representation of the training samples is extracted through the decoder; the decoder participates in the training together with the first decoder, the watermark extractor, and the second decoder, and the parameters of the preset watermark embedding model can be adjusted through methods such as hard parameter sharing and soft parameter sharing in the future.

[0086] In step S104 of some embodiments, the first decoder can adopt architectures such as an autoregressive model, a sequence-to-sequence model, a conditional variational autoencoder, etc., and is not limited thereto.

[0087] Please refer to Figure 4 , in some embodiments, step S104 may include, but is not limited to, steps S401 to S403:

[0088] Step S401, vectorize the watermark sample to obtain the watermark sample vector representation;

[0089] Step S402, perform vector concatenation on the watermark sample vector representation and the audio sample vector representation to obtain a concatenated vector;

[0090] Step S403, perform transposed convolution and anti-pooling operations on the concatenated vector through the first decoder to obtain the watermark audio data.

[0091] In step S401 of some embodiments, the watermark sample vector representation can be obtained by performing one-hot encoding on the watermark sample; or the watermark sample vector representation can be obtained by performing label encoding on the watermark sample; and the like.

[0092] In step S403 of some embodiments, the first decoder includes a number of transposed convolutional layers and a number of anti-pooling layers, and may also include an upsampling layer, a gated recurrent unit (GRU), an activation layer (such as ReLU), etc., which is not limited thereto. For example, the audio sample vector representation can be upsampled first, and then the upsampled audio sample vector is concatenated with the watermark sample vector representation to obtain a concatenated vector; the hidden state of the concatenated vector is calculated through a gated recurrent unit and input to subsequent convolutional layers, anti-pooling layers, and activation layers to obtain watermarked audio data.

[0093] Steps S401 to S403 illustrated in the embodiments of the present application obtain a concatenated vector by concatenating the watermark sample vector representation and the audio sample vector representation, and perform transposed convolution and anti-pooling operations on the concatenated vector to obtain watermarked audio data

[0094] In step S105 of some embodiments, a watermark extractor based on a deep learning classifier can be used to directly classify and output watermark tags from the watermarked audio, which is suitable for extracting robust watermarks in adversarial attack scenarios; a watermark extractor based on frequency domain feature decoding can also be used to perform STFT / CQT transformation on the watermarked audio and decode the embedded frequency domain watermark from a specific frequency band, which is suitable for extracting covert watermarks in a fidelity-first scenario; this is not limited thereto.

[0095] In step S106 of some embodiments, data reconstruction can be performed through a second decoder based on an autoregressive model, which is suitable for high-fidelity scenarios; data reconstruction can also be performed through a second decoder based on a diffusion model, which is suitable for small-sample learning scenarios; this is not limited thereto.

[0096] Please refer to Figure 5 , in some embodiments, step S107 may include but is not limited to steps S501 to S504:

[0097] Step S501, calculate the watermark reconstruction loss based on the reference watermark and the watermark sample;

[0098] Step S502, calculate the audio reconstruction loss based on the audio reconstruction data and the training samples;

[0099] Step S503, calculate the target loss according to the watermark reconstruction loss and the audio reconstruction loss;

[0100] Step S504, adjust the parameters of the preset watermark embedding model based on the target loss to obtain the target watermark embedding model.

[0101] Among them, the watermark reconstruction loss refers to the difference between the reference watermark and the original watermark, which is used to measure the accuracy of watermark embedding; the audio reconstruction loss refers to the difference between the audio reconstruction data and the original audio or its corresponding label, which is used to measure the quality of audio reconstruction; the objective loss refers to the final loss value obtained after integrating multiple loss functions (such as watermark reconstruction loss, audio reconstruction loss, etc.) during model training, which is used to guide the adjustment of model parameters; it is the core objective of model training. By minimizing the objective loss, the model can achieve a balance among multiple tasks and optimize the performance of each subtask simultaneously. Understandably, when the training sample is audio, the audio reconstruction loss is directly calculated based on the audio and the audio reconstruction data; when the training sample is text, the training sample is provided with a sample label, and the sample label is the audio corresponding to the text, and the audio reconstruction loss is calculated based on the label and the audio reconstruction data.

[0102] In step S501 of some embodiments, the watermark reconstruction loss can be calculated by the mean square error (MSE) loss function, which is applicable to the scenario where the watermark is a numerical identifier. For example, when embedding a customer ID (binary code) in a financial transaction recording, the digital label needs to be accurately extracted; the watermark reconstruction loss can also be calculated by the structural similarity (SSIM) loss function, which is applicable to the scenario where the watermark is a spectral feature; or the watermark reconstruction loss can be calculated by the cross entropy loss, which is applicable to the scenario where the watermark is a discrete label; not limited to this.

[0103] In step S502 of some embodiments, the audio reconstruction loss can be calculated by the mean square error (MSE) loss function, which is applicable to the scenario where audio fidelity is prioritized. For example, in a telephone recording, the original waveform details need to be retained; the audio reconstruction loss can also be calculated by the scale-invariant signal-to-noise ratio (SI-SNR) loss function, which is applicable to the scenario where audio intelligibility is prioritized. For example, in a customer service call recording; not limited to this.

[0104] In step S503 of some embodiments, the watermark reconstruction loss and the audio reconstruction loss can be weighted and summed to obtain the objective loss.

[0105] Steps S501 to S504 shown in the embodiments of the present application calculate the objective loss based on the watermark reconstruction loss and the audio reconstruction loss, aiming to minimize the differences between the reference watermark and the watermark sample, and between the audio reconstruction data and the training sample, optimize the performance of the preset watermark embedding model, and improve the auditory quality of the watermarked audio data while maintaining the watermark quality.

[0106] Please refer to Figure 6 , in some embodiments, after step S503, the watermark embedding method further includes steps S601 to S602:

[0107] Step S601, calculate the mutual information between the reference watermark and the audio reconstruction data to obtain the correlation loss;

[0108] Step S602: Update the target loss based on the correlation loss.

[0109] In step S601 of some embodiments, the mutual information between the reference watermark and the audio reconstruction data can be calculated by the Mutual Information Neural Estimation (MINE), or by the frequency-domain cross-correlation method. This is not limited to these methods.

[0110] In step S602 of some embodiments, hyperparameters can be set for the correlation loss. The hyperparameters are used to control the intensity of the mutual information constraint. The correlation loss is added as a regularization term to the target loss to update the target loss. Alternatively, the correlation loss and the target loss can be directly weighted and summed, and the weighted sum value is used to update the target loss. This is not limited to these methods.

[0111] Steps S601 to S602 illustrated in the embodiments of the present application calculate the mutual information between the reference watermark and the audio reconstruction data. The mutual information represents the reduction in uncertainty of the audio reconstruction data given the watermark data, that is, the dependence between the reference watermark and the audio reconstruction data. It is used as a loss term to train the watermark embedding model. During the training process, minimizing the watermark reconstruction loss, the audio reconstruction loss, and the correlation loss can further reduce the coupling between the watermark and the audio, and effectively improve the auditory quality of the watermarked audio data.

[0112] In some embodiments, the target watermark embedding model includes a target first decoder; step S108 may include but is not limited to step S701:

[0113] Step S701: Conditionally generate target watermarked audio data through the target first decoder for the target watermark and the target audio vector representation, where the target watermark serves as the control condition for the target first decoder.

[0114] It is easy to understand that the target watermark embedding model is the target watermark embedding model after parameter adjustment, and the target first decoder is the first decoder after parameter adjustment.

[0115] In step S701 of some embodiments, the target watermark embedding model may further include a decoder. Input the target input data and the target watermark into the target watermark embedding model. The decoder extracts the target audio vector representation of the target input data, and the target first decoder conditionally generates the target watermarked audio data for the target watermark and the target audio vector representation.

[0116] In step S701 of some embodiments, the target watermark embedding model may further include a target watermark extractor; the target watermark extractor is a watermark extractor with adjusted parameters. The target watermark extractor extracts the watermark from the target watermark audio data to obtain a watermark extraction value. The target watermark is used to verify the watermark extraction value; or the watermark extraction value is used for data traceability.

[0117] Taking the fintech scenario as an example, financial institutions such as banks use voice robots to communicate with customers to complete tasks such as standardized consultation, business guidance, and question answering, such as account query and transfer operation guidance. Since such conversations usually involve a lot of sensitive information, such as financial product descriptions and investment advice, the voice robot identifier or customer identifier can be embedded in the voice of the voice robot to reduce the risk of data tampering and facilitate subsequent data traceability.

[0118] Specifically, the voice robot receives an investment analysis question input by the customer and generates an investment advice text based on the investment analysis question. The investment advice text is used as the above-mentioned target input data, and based on the mapping relationship between the text and the phoneme, the investment advice text is converted into a corresponding investment advice phoneme sequence. The investment advice phoneme sequence is mapped to a continuous latent space through an encoder to obtain an investment advice audio vector representation. The customer identifier is obtained as the target watermark, and the customer identifier is vectorized to obtain a customer identifier vector representation. The customer identifier vector representation and the investment advice audio vector representation are vector-concatenated to obtain a concatenated vector; the concatenated vector is subjected to transposed convolution and anti-pooling operations through a first decoder to obtain investment advice watermark audio data. The voice robot outputs the investment advice watermark audio data. By extracting and verifying the watermark in the watermark audio data, it can be confirmed whether the audio has been tampered with and data traceability can be performed.

[0119] Taking the medical technology scenario as an example, a remote medical consultation system transmits voice data between the doctor side and the patient side. Such voice data usually involves sensitive information such as medical condition descriptions, patient physiological data, and medical plans. The doctor identifier and patient identifier can be embedded in such voice data to reduce the risk of data tampering and facilitate subsequent data traceability.

[0120] Specifically, an electronic stethoscope collects heart sound data of a patient, uses the heart sound data as the above-mentioned target input data, maps the heart sound data to a continuous latent space through an encoder to obtain a continuous vector representation; discretizes the continuous vector representation through a vector quantizer to obtain a heart sound vector representation. The patient identifier is obtained as the target watermark, and the patient identifier is vectorized to obtain a patient identifier vector representation. The patient identifier vector representation and the heart sound vector representation are vector-concatenated to obtain a concatenated vector; the concatenated vector is subjected to transposed convolution and de-pooling operations through a first decoder to obtain watermarked heart sound data. The electronic stethoscope outputs and transmits the watermarked heart sound data to the doctor side. By extracting and verifying the watermark in the watermarked heart sound data, it can be confirmed whether the heart sound has been tampered with.

[0121] Please refer to Figure 7 , the embodiment of the present application further provides a watermark embedding device, which can implement the above watermark embedding method. The device includes:

[0122] A watermark acquisition module, configured to acquire a watermark sample and a target watermark;

[0123] An audio vector representation acquisition module, configured to acquire an audio sample vector representation of a training sample and a target audio vector representation of target input data; the training sample is audio or text, and the target input data is audio or text;

[0124] A model acquisition module, configured to acquire a preset watermark embedding model, where the preset watermark embedding model includes a first decoder, a watermark extractor, and a second decoder;

[0125] A conditional generation module, configured to perform conditional generation on the watermark sample and the audio sample vector representation through the first decoder to obtain watermarked audio data; wherein, the watermark sample is used as a control condition of the first decoder;

[0126] A watermark extraction module, configured to extract a watermark from the watermarked audio data through the watermark extractor to obtain a reference watermark;

[0127] A data reconstruction module, configured to reconstruct the audio sample vector representation through the second decoder to obtain audio reconstruction data;

[0128] A parameter adjustment module, configured to adjust the parameters of the preset watermark embedding model based on the reference watermark, the watermark sample, the audio reconstruction data, and the training sample to obtain a target watermark embedding model;

[0129] A watermark embedding module, configured to perform watermark embedding processing on the target watermark and the target audio vector representation according to the target watermark embedding model.

[0130] The specific implementation manner of this watermark embedding device is basically the same as the specific embodiment of the above watermark embedding method, and will not be elaborated here.

[0131] An embodiment of the present application further provides an electronic device, which includes a memory and a processor. The memory stores a computer program, and when the processor executes the computer program, the above-mentioned watermark embedding method is implemented. The electronic device can be any intelligent terminal including a tablet computer, an in-vehicle computer, etc.

[0132] Please refer to Figure 8 , Figure 8 which illustrates the hardware structure of an electronic device according to another embodiment. The electronic device includes:

[0133] A processor 901, which can be implemented in ways such as a general-purpose CPU (Central Processing Unit), a microprocessor, an Application Specific Integrated Circuit (ASIC), or one or more integrated circuits, and is used to execute relevant programs to implement the technical solutions provided by the embodiments of the present application;

[0134] A memory 902, which can be implemented in forms such as a Read Only Memory (ROM), a static storage device, a dynamic storage device, or a Random Access Memory (RAM). The memory 902 can store an operating system and other application programs. When implementing the technical solutions provided by the embodiments of this specification through software or firmware, the relevant program codes are stored in the memory 902, and the processor 901 is called to execute the watermark embedding method of the embodiments of the present application;

[0135] An input / output interface 903, which is used to implement information input and output;

[0136] A communication interface 904, which is used to implement communication and interaction between this device and other devices. It can implement communication through a wired method (such as USB, network cable, etc.) or through a wireless method (such as a mobile network, WIFI, Bluetooth, etc.);

[0137] A bus 905, which transmits information between various components of the device (such as the processor 901, the memory 902, the input / output interface 903, and the communication interface 904);

[0138] Among them, the processor 901, the memory 902, the input / output interface 903, and the communication interface 904 are communicatively connected to each other inside the device through the bus 905.

[0139] An embodiment of the present application further provides a computer-readable storage medium, which stores a computer program, and when the computer program is executed by a processor, the above-mentioned watermark embedding method is implemented.

[0140] As a non-transitory computer-readable storage medium, the memory can be used to store non-transitory software programs and non-transitory computer-executable programs. In addition, the memory may include high-speed random access memory, and may also include non-transitory memory, such as at least one magnetic disk storage device, a flash memory device, or other non-transitory solid-state storage devices. In some embodiments, the memory may optionally include a memory remotely disposed relative to the processor, and these remote memories can be connected to the processor through a network. Examples of the above network include but are not limited to the Internet, an intranet, a local area network, a mobile communication network, and combinations thereof.

[0141] The watermark embedding method, watermark embedding device, electronic device, and storage medium provided in the embodiments of the present application construct a preset watermark embedding model including a first decoder, a watermark extractor, and a second decoder; obtain an audio sample vector representation, and use the watermark sample to perform conditional generation on the audio sample vector representation through the first decoder to obtain watermarked audio data; extract the watermark from the watermarked audio data through the watermark extractor to obtain a reference watermark; reconstruct the data of the audio sample vector representation through the second decoder to obtain audio reconstruction data; train the preset watermark embedding model based on the reference watermark, the watermark sample, the audio reconstruction data, and the training sample, so that the model learns the mapping relationship between the audio sample vector representation and the watermarked audio under a given condition (i.e., the watermark sample), and the mapping relationship between the audio sample vector representation and the audio. The watermark data will not interfere with the feature extraction of the audio data (i.e., the watermark sample will not interfere with the generation of the audio sample vector representation), thereby improving the auditory quality of the watermarked audio data and improving the watermark audio fidelity.

[0142] The embodiments described in the embodiments of the present application are for more clearly illustrating the technical solutions of the embodiments of the present application, and do not constitute a limitation on the technical solutions provided by the embodiments of the present application. Those skilled in the art can know that with the evolution of technology and the emergence of new application scenarios, the technical solutions provided by the embodiments of the present application are equally applicable to similar technical problems.

[0143] Those skilled in the art can understand that the technical solutions shown in the figures do not constitute a limitation on the embodiments of the present application, and may include more or fewer steps than shown in the figures, or combine certain steps, or different steps.

[0144] The device embodiments described above are merely illustrative. The units described as separate components may or may not be physically separated, that is, they may be located in one place, or may be distributed to multiple network units. Some or all of the modules can be selected according to actual needs to achieve the purpose of the solution of this embodiment.

[0145] Those of ordinary skill in the art can understand that all or some of the steps in the watermark embedding method disclosed above, and the functional modules / units in the system and device can be implemented as software, firmware, hardware, and their appropriate combinations.

[0146] As used in the specification of this application and the above-mentioned drawings, the terms "first", "second", "third", "fourth", etc. (if any) are used to distinguish similar objects and are not necessarily used to describe a specific order or sequence. It should be understood that such data can be interchanged under appropriate circumstances so that the embodiments of this application described herein can be implemented in an order different from those illustrated or described herein. In addition, the terms "comprising" and "having" and any variations thereof are intended to cover non-exclusive inclusion. For example, a process, a watermark embedding method, a system, a product, or a device that includes a series of steps or units is not necessarily limited to those steps or units clearly listed, but may include other steps or units not clearly listed or inherent to these processes, watermark embedding methods, products, or devices.

[0147] It should be understood that in this application, "at least one (item)" means one or more, and "a plurality" means two or more. "And / or" is used to describe the association relationship of associated objects and indicates that there can be three relationships. For example, "A and / or B" can mean: only A exists, only B exists, and both A and B exist at the same time. Among them, A and B can be singular or plural. The character " / " generally means that the associated objects before and after are in an "or" relationship. "At least one (one) of the following" or its similar expression refers to any combination of these items, including any combination of single item (one) or plural items (ones). For example, at least one (one) of a, b, or c can mean: a, b, c, "a and b", "a and c", "b and c", or "a and b and c", where a, b, and c can be single or multiple.

[0148] In several embodiments provided in this application, it should be understood that the disclosed device and watermark embedding method can be implemented in other ways. For example, the device embodiments described above are merely illustrative. For example, the above-mentioned unit division is only a logical function division. In actual implementation, there may be other division methods. For example, multiple units or components can be combined or integrated into another system, or some features can be ignored or not executed. Another point is that the displayed or discussed coupling or direct coupling or communication connection to each other can be through some interfaces. The indirect coupling or communication connection of the device or unit can be in an electrical, mechanical, or other form.

[0149] The units described above as separate components may or may not be physically separated. The components shown as units may or may not be physical units, that is, they may be located in one place or distributed across multiple network units. Some or all of the units can be selected according to actual needs to achieve the purpose of the solution of this embodiment.

[0150] In addition, each functional unit in various embodiments of the present application may be integrated in a processing unit, may exist separately as individual physical units, or two or more units may be integrated in one unit. The above-mentioned integrated units can be implemented in the form of hardware or in the form of software functional units.

[0151] If the integrated unit is implemented in the form of a software functional unit and sold or used as an independent product, it can be stored in a computer-readable storage medium. Based on this understanding, the technical solution of the present application, in essence, or the part that contributes to the prior art, or all or part of this technical solution, can be embodied in the form of a software product. This computer software product is stored in a storage medium and includes multiple instructions for causing a computer device (which can be a personal computer, a server, or a network device, etc.) to execute all or part of the steps of the watermark embedding method in various embodiments of the present application. The aforementioned storage medium includes: USB flash drives, mobile hard disks, read-only memory (ROM), random access memory (RAM), magnetic disks, or optical discs and other various media that can store programs.

[0152] The preferred embodiments of the embodiments of the present application have been described above with reference to the accompanying drawings. However, this does not limit the scope of the rights of the embodiments of the present application. Any modifications, equivalent replacements, and improvements made by those skilled in the art without departing from the scope and essence of the embodiments of the present application shall fall within the scope of the rights of the embodiments of the present application.

Claims

1. A watermark embedding method, characterized in that: The method comprises: Get watermark sample and target watermark; Obtaining an audio sample vector representation of a training sample, and obtaining a target audio vector representation of target input data; the training sample is audio or text, and the target input data is audio or text; Acquire a preset watermark embedding model, wherein the preset watermark embedding model includes a first decoder, a watermark extractor, and a second decoder; Conditionally generating the watermark sample and the audio sample vector representation through the first decoder to obtain watermarked audio data; wherein the watermark sample serves as a control condition of the first decoder; Extracting watermarks from the watermarked audio data by the watermark extractor to obtain a reference watermark; Reconstructing the audio sample vector representation by the second decoder to obtain audio reconstructed data; Based on the reference watermark, the watermark sample, the audio reconstruction data and the training sample, adjusting the parameters of the preset watermark embedding model to obtain a target watermark embedding model; The target watermark and the target audio vector representation are subjected to watermark embedding processing according to the target watermark embedding model.

2. The method according to claim 1, characterized in that The conditionally generating the watermark sample and the audio sample vector representation by the first decoder to obtain watermarked audio data includes: Vectorizing the watermark sample to obtain a watermark sample vector representation; Performing vector concatenation on the watermark sample vector representation and the audio sample vector representation to obtain a concatenated vector; The first decoder performs transposed convolution and unpooling operations on the splicing vector to obtain the watermarked audio data.

3. The method according to claim 1, characterized in that The adjusting the parameters of the preset watermark embedding model based on the reference watermark, the watermark sample, the audio reconstruction data and the training sample to obtain a target watermark embedding model includes: Calculating a watermark reconstruction loss based on the reference watermark and the watermark sample; Calculating an audio reconstruction loss based on the audio reconstruction data and the training samples; Calculating a target loss according to the watermark reconstruction loss and the audio reconstruction loss; Based on the target loss, the parameters of the preset watermark embedding model are adjusted to obtain the target watermark embedding model.

4. The method according to claim 3, characterized in that The method further includes: updating the target watermark embedding model, specifically including: Calculating mutual information between the reference watermark and the audio reconstruction data to obtain correlation loss; Based on the correlation loss, the target loss is updated.

5. The method according to any one of claims 1 to 4, characterized in that: The target input data is audio, and obtaining a target audio vector representation of the target input data includes: Acquiring the target input data; Mapping the target input data to a continuous latent space through an encoder to obtain a continuous vector representation; The continuous vector representation is discretized by a vector quantizer to obtain the audio vector representation.

6. The method according to any one of claims 1 to 4, characterized in that: The target input data is text, and obtaining a target audio vector representation of the target input data includes: Acquiring the target input data; Based on the mapping relationship between text and phonemes, the target input data is converted into a phoneme sequence to obtain a target phoneme sequence; The target phoneme sequence is mapped to a continuous latent space through an encoder to obtain the target audio vector representation.

7. The method according to any one of claims 1 to 4, characterized in that: The target watermark embedding model includes a target first decoder; and performing watermark embedding processing on the target watermark and the target audio vector representation according to the target watermark embedding model includes: The target watermark and the target audio vector representation are conditionally generated by the target first decoder to obtain target watermark audio data; wherein the target watermark serves as a control condition for the target first decoder.

8. A watermark embedding device, characterized in that: The device comprises: A watermark acquisition module is used to obtain watermark samples and target watermarks; An audio vector representation acquisition module, used to acquire an audio sample vector representation of a training sample and a target audio vector representation of target input data; the training sample is audio or text, and the target input data is audio or text; A model acquisition module, used to acquire a preset watermark embedding model, wherein the preset watermark embedding model includes a first decoder, a watermark extractor, and a second decoder; a conditional generation module, configured to conditionally generate the watermark sample and the audio sample vector representation through the first decoder to obtain watermarked audio data; wherein the watermark sample serves as a control condition of the first decoder; A watermark extraction module, used to extract the watermark from the watermarked audio data through the watermark extractor to obtain a reference watermark; A data reconstruction module, used for reconstructing the audio sample vector representation through the second decoder to obtain audio reconstruction data; A parameter adjustment module, configured to adjust the parameters of the preset watermark embedding model based on the reference watermark, the watermark sample, the audio reconstruction data and the training sample to obtain a target watermark embedding model; The watermark embedding module is used to perform watermark embedding processing on the target watermark and the target audio vector representation according to the target watermark embedding model.

9. An electronic device, characterized in that: The electronic device comprises a memory and a processor, the memory stores a computer program, and the processor implements the watermark embedding method according to any one of claims 1 to 7 when executing the computer program.

10. A computer-readable storage medium storing a computer program, characterized in that: When the computer program is executed by a processor, the watermark embedding method according to any one of claims 1 to 7 is implemented.