Audio watermark generation method and related device
By using acoustic feature segmentation and latent space coding to encode audio, the noise problem caused by frame-by-frame watermark embedding is solved, achieving imperceptible watermark embedding and improving the listening experience and naturalness of the audio.
Patent Information
- Application Number
- CN202511457826.5
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-10-13
- Publication Date
- 2025-11-11
- Estimated Expiration
- 2045-10-13
AI Technical Summary
When existing technologies embed watermark information frame by frame into audio signals, they can easily introduce perceptible noise, leading to a decline in the user's listening experience.
By performing acoustic feature segmentation and latent space encoding on the input audio, latent space representation vectors of multiple acoustic feature blocks are obtained and fused with watermark information encoding vectors. The audio containing watermark information is then decoded and reconstructed. Latent space encoding is used to remove redundant audio information, and watermark information is embedded only on key features.
It reduces the impact of watermarks on the overall characteristics of the audio, making the watermarks less noticeable and improving the listening experience and naturalness of the audio.
Smart Images

Figure CN120932656A_ABST
Abstract
Description
Technical Field
[0001] This application relates to the field of artificial intelligence technology, and in particular to an audio watermark generation method and related apparatus. Background Technology
[0002] The neural network-based audio watermarking embedding method involves embedding watermark information frame by frame into the original audio signal. Specifically, first, a short-time Fourier transform (STFT) is performed on the original audio signal to separate the amplitude spectrum features and phase spectrum features of each audio frame. Then, the watermark information encoding is concatenated with the amplitude spectrum features of each audio frame to construct a joint feature tensor for each audio frame, which is used as the input of a convolutional neural network. The watermark embedding weight matrix corresponding to each frequency point of the amplitude spectrum is then learned. After multiplying the weight matrix of each audio frame element by element with the amplitude spectrum features of the corresponding audio frame, and combining it with the phase spectrum features of the corresponding audio frame, the system is reconstructed through an inverse short-time Fourier transform (iSTFT) to generate watermarked audio with watermark information.
[0003] However, embedding watermark information frame by frame introduces perceptible noise into the original audio signal, reducing the user's listening experience. Summary of the Invention
[0004] In view of the above problems, this application provides an audio watermark generation method and related apparatus to achieve the purpose of seamlessly embedding watermarks into audio. The specific solution is as follows:
[0005] The first aspect of this application provides a method for generating an audio watermark, including:
[0006] Get the input audio;
[0007] The input audio is subjected to acoustic feature block division and latent space encoding to obtain latent space representation vectors corresponding to multiple acoustic feature blocks of the input audio.
[0008] The latent space representation vectors corresponding to the multiple acoustic feature blocks are fused with the watermark information encoding vectors to obtain the watermark embedding vectors corresponding to the multiple acoustic feature blocks.
[0009] The watermarked audio containing watermark information is reconstructed by decoding the watermark embedding vectors corresponding to the multiple acoustic feature blocks.
[0010] In one possible implementation, the step of decoding and reconstructing the watermark-embedded audio containing watermark information based on the watermark embedding vectors corresponding to the plurality of acoustic feature blocks includes:
[0011] The watermark embedding vectors corresponding to the plurality of acoustic feature blocks are decoded respectively to obtain the reconstructed feature blocks corresponding to the plurality of acoustic feature blocks respectively;
[0012] Determine whether the reconstruction quality of the reconstruction feature blocks corresponding to the plurality of acoustic feature blocks meets the preset quality requirements;
[0013] The reconstructed feature blocks that do not meet the quality requirements are taken as target reconstructed feature blocks. The target reconstructed feature blocks are then replaced with feature compensation based on the acoustic feature blocks corresponding to them, resulting in updated reconstructed feature blocks.
[0014] The watermark-embedded audio is generated based on the reconstructed feature blocks or updated reconstructed feature blocks corresponding to the plurality of acoustic feature blocks.
[0015] In one possible implementation, determining whether the reconstruction quality of the reconstructed feature blocks corresponding to the plurality of acoustic feature blocks meets a preset quality requirement includes:
[0016] Calculate the energy of the reconstructed feature block corresponding to each of the plurality of acoustic feature blocks, and use it as the reconstructed energy corresponding to each of the plurality of acoustic feature blocks;
[0017] Based on whether the reconstruction energy corresponding to each of the plurality of acoustic feature blocks is greater than or equal to a preset energy threshold, it is determined whether the reconstruction quality of the reconstruction feature blocks corresponding to each of the plurality of acoustic feature blocks meets the quality requirements.
[0018] And / or,
[0019] For each of the plurality of acoustic feature blocks, the difference between the acoustic feature block and the corresponding reconstructed feature block is calculated, and this difference is taken as the reconstruction difference corresponding to the acoustic feature block.
[0020] Based on whether the reconstruction gap corresponding to the plurality of acoustic feature blocks is greater than or equal to a preset gap threshold, it is determined whether the reconstruction quality of the reconstruction feature blocks corresponding to the plurality of acoustic feature blocks meets the quality requirements.
[0021] In one possible implementation, calculating the difference between the acoustic feature block and the corresponding reconstructed feature block includes:
[0022] Calculate the mean square error of the acoustic feature block and the corresponding reconstructed feature block.
[0023] In one possible implementation, the step of performing feature compensation replacement on the target reconstructed feature block based on the acoustic feature block corresponding to the target reconstructed feature block includes:
[0024] For each of the plurality of acoustic feature blocks:
[0025] If the reconstruction gap corresponding to the acoustic feature block is less than the gap threshold, then feature compensation is performed on the reconstruction feature block corresponding to the acoustic feature block based on the acoustic feature block.
[0026] If the reconstructed energy corresponding to the acoustic feature block is less than the energy threshold, then the reconstructed feature block corresponding to the acoustic feature block is replaced with the acoustic feature block.
[0027] In one possible implementation, the step of performing feature compensation on the reconstructed feature block corresponding to the acoustic feature block based on the acoustic feature block includes:
[0028] The acoustic feature block and its corresponding reconstructed feature block are weighted and summed to obtain the fused feature block, which is then used as the updated reconstructed feature block.
[0029] In one possible implementation, the input audio is subjected to acoustic feature segmentation and latent space encoding to obtain latent space representation vectors corresponding to multiple acoustic feature blocks of the input audio. These latent space representation vectors are then fused with watermark information encoding vectors to obtain watermark embedding vectors corresponding to the multiple acoustic feature blocks. Finally, the watermark-embedded audio containing watermark information is decoded and reconstructed based on these watermark embedding vectors, including:
[0030] The input audio is fed into a pre-trained audio watermark embedding model to obtain the watermark-embedded audio output by the model.
[0031] In one possible implementation, the audio watermark embedding model and the watermark extraction model are jointly trained;
[0032] The training process of the joint training includes:
[0033] Obtain the original audio sample;
[0034] The original audio sample is fed into the audio watermark embedding model to obtain a watermarked audio sample containing watermark information.
[0035] Calculate the first loss between the acoustic features of the original audio sample and the watermarked audio sample;
[0036] The watermarked audio sample is sent to the simulation attack layer to simulate the attack and obtain the attacked audio sample.
[0037] The watermark extraction model is used to extract the watermark from the attacked audio sample to obtain the watermark codebook recovery vector.
[0038] Calculate the second loss between the watermark codebook recovery vector and the watermark codebook vector, wherein the watermark codebook vector is used to encode the watermark information encoding vector.
[0039] The audio watermark embedding model and the watermark extraction model are trained based on the first loss and the second loss.
[0040] A second aspect of this application provides a computer program product including computer-readable instructions that, when executed on an electronic device, cause the electronic device to implement the audio watermark generation method of the first aspect or any implementation thereof.
[0041] A third aspect of this application provides an electronic device, comprising at least one processor and a memory connected to the processor, wherein:
[0042] The memory is used to store computer programs;
[0043] The processor is used to execute the computer program so that the electronic device can implement the audio watermark generation method of the first aspect or any implementation thereof.
[0044] A fourth aspect of this application provides a computer storage medium carrying one or more computer programs, which, when executed by an electronic device, enable the electronic device to implement the audio watermark generation method of the first aspect or any implementation thereof.
[0045] Using the above technical solution, the audio watermark generation method provided in this application considers that embedding watermarks frame by frame on the amplitude spectrum involves many frequency points corresponding to redundant audio information. Embedding watermarks on these frequency points may interfere with the normal frequency distribution of the audio, making the watermark easily detectable. To avoid embedding watermarks on redundant audio information as much as possible, this application acquires input audio, performs acoustic feature segmentation and latent space encoding on the input audio, obtaining latent space representation vectors corresponding to multiple acoustic feature blocks of the input audio. These latent space representation vectors are then fused with the watermark information encoding vectors to obtain watermark embedding vectors corresponding to multiple acoustic feature blocks. Based on these watermark embedding vectors, the watermark-embedded audio containing watermark information is decoded and reconstructed. This application uses latent space encoding to remove a large amount of redundant audio information from the acoustic feature blocks, retaining only the more critical and representative features of the audio essence. Embedding watermark information encoding vectors on these features minimizes the impact of the watermark on the overall audio characteristics, making the watermark less noticeable and improving the listening experience.
[0046] Furthermore, the latent space representation vector is composed of more core features with strong correlation and integrity. Embedding watermarks in these features with strong correlation and integrity allows the watermark information to be better integrated into the overall characteristics of the input audio due to the mutual constraints and coordination between features. This avoids excessive damage to individual features, thus maintaining the naturalness and coherence of the watermark embedded in the audio and improving the listening experience of the watermark embedded in the audio. Attached Figure Description
[0047] The above and other features, advantages, and aspects of the embodiments of this disclosure will become more apparent from the accompanying drawings and the following detailed description. Throughout the drawings, the same or similar reference numerals denote the same or similar elements. It should be understood that the drawings are schematic, and the originals and elements are not necessarily drawn to scale.
[0048] Figure 1 A schematic diagram of a system architecture provided in this application;
[0049] Figure 2 A flowchart illustrating an audio watermark generation method provided in this application;
[0050] Figure 3 This application provides a schematic diagram of the structure of an audio watermark embedding model.
[0051] Figure 4 A schematic diagram illustrating the calculation process of model loss provided in this application;
[0052] Figure 5 A schematic diagram of an audio watermark generation device provided in this application;
[0053] Figure 6 This is a schematic diagram of the structure of an electronic device provided in this application. Detailed Implementation
[0054] The embodiments of this application are described below with reference to the accompanying drawings. The terminology used in the implementation section of this application is for explaining specific embodiments only and is not intended to limit the scope of this application.
[0055] The embodiments of this application will now be described with reference to the accompanying drawings. Those skilled in the art will recognize that, with technological advancements and the emergence of new scenarios, the technical solutions provided in the embodiments of this application are equally applicable to similar technical problems.
[0056] The terms "first," "second," etc., used in the specification, claims, and accompanying drawings of this application are used to distinguish similar objects and are not necessarily used to describe a specific order or sequence. It should be understood that such terms are interchangeable where appropriate; this is merely a way of distinguishing objects with the same attributes in the embodiments of this application. Furthermore, the terms "comprising" and "having," and any variations thereof, are intended to cover non-exclusive inclusion, so that a process, method, system, product, or apparatus that comprises a series of elements is not necessarily limited to those elements, but may include other elements not explicitly listed or inherent to those processes, methods, products, or apparatuses.
[0057] This application provides an audio watermark generation method and related apparatus, which can be applied to scenarios where watermarks need to be embedded in audio. For example, in the copyright protection of musical and film works, watermarks can be embedded in the musical and film works to be protected to clearly identify the copyright ownership information of the music; in important audio recording scenarios such as meeting recordings and interview recordings, watermarks can be embedded in the recorded audio to detect whether the recorded audio has been tampered with or to verify the authenticity of the recorded audio; in advertising scenarios, watermarks can be embedded in the audio of the advertisement to be placed to track the playback status of the advertising audio in real time; in covert communication scenarios, confidential information can be embedded as a watermark in ordinary audio data and transmitted through public audio transmission channels to achieve covert communication.
[0058] It should be noted that the above scenarios are merely examples and are not intended to limit this application.
[0059] Optionally, the audio watermark generation method provided in this application can be applied to, for example... Figure 1 The system architecture shown includes a terminal 100 and a server 200. The server 200 may include one or more servers (…). Figure 1 (This example uses a server as an illustration).
[0060] Either terminal 100 or server 200 can be used independently to execute the audio watermark generation method provided in the embodiments of this application. Alternatively, terminal 100 and server 200 can also be used collaboratively to execute the audio watermark generation method provided in the embodiments of this application.
[0061] The following description Figure 1 The product form of the mid-terminal 100;
[0062] The terminal 100 in this application embodiment can be a mobile phone, tablet computer, wearable device, vehicle device, augmented reality (AR) / virtual reality (VR) device, laptop computer, ultra-mobile personal computer (UMPC), netbook, personal digital assistant (PDA), etc., and this application embodiment does not impose any restrictions on it.
[0063] To enable those skilled in the art to better understand this application, the audio watermark generation method of the embodiments of this application will be described in detail below with reference to the accompanying drawings.
[0064] Reference Figure 2 , Figure 2 This is a flowchart illustrating an audio watermark generation method provided in an embodiment of this application, as shown below. Figure 2 As shown, the audio watermark generation method may include:
[0065] Step S201: Obtain the input audio.
[0066] Here, input audio refers to the audio data to which the watermark is to be embedded.
[0067] For example, in the context of copyright protection for musical and film works, input audio refers to the audio data in the musical or film works; in important audio recording scenarios such as meeting recordings and interview recordings, input audio refers to the recorded audio; in advertising scenarios, input audio refers to the advertising audio; and in covert communication scenarios, input audio refers to ordinary audio data.
[0068] Step S202: Perform acoustic feature segmentation and latent space encoding on the input audio to obtain the latent space representation vectors corresponding to the multiple acoustic feature blocks of the input audio.
[0069] In this embodiment, acoustic features are first extracted from the input audio to obtain a feature spectrogram of the input audio. Then, the feature spectrogram is divided into blocks to obtain multiple acoustic feature blocks. Finally, latent space encoding is performed on each acoustic feature block to obtain the latent space representation vector corresponding to each acoustic feature block.
[0070] Optionally, the process of "extracting acoustic features from the input audio" may include: windowing the input audio before extracting the acoustic feature FilterBank. Here, FilterBank represents a filter bank, which is an acoustic feature that simulates the characteristics of human hearing.
[0071] This embodiment extracts acoustic features by windowing the input audio. The window function can smooth the signal boundaries of the input audio, avoiding high-frequency noise caused by the discontinuity of the input audio. At the same time, dividing the long audio data into several short audio data can achieve approximately stable audio in a local range, which is conducive to more accurate extraction of acoustic features.
[0072] Meanwhile, FilterBank is used as the extracted acoustic feature. Since FilterBank is an acoustic feature that simulates the hearing characteristics of the human ear, latent space coding and subsequent watermark embedding on the basis of FilterBank can make the watermark embedding only affect specific frequency bands with low auditory sensitivity, thus significantly improving the concealment of the watermark.
[0073] The size of the feature spectrogram obtained after the above processing can be denoted as f×d, where f represents the number of feature frames and d represents the feature dimension. The feature spectrogram is then divided into blocks, resulting in N acoustic feature blocks, where the nth acoustic feature block is denoted as... , n=1,2,…,N.
[0074] Optionally, the size of each acoustic feature block can be a fixed 16×16.
[0075] Of course, other block sizes can be set in this application, or an appropriate block size can be selected according to the audio characteristics of the input audio, etc., without specific limitations here.
[0076] Optionally, the encoder of a pre-trained model can be used to perform latent space encoding on each acoustic feature block to obtain the latent space representation vector corresponding to each acoustic feature block.
[0077] Preferably, the pre-trained model can be the AudioMAE model, which is an audio self-supervised pre-trained model based on MaskedAutoencoder technology. Its core goal is to learn a general representation of audio by masking part of the audio data and reconstructing the missing parts, thereby improving the model's ability to understand audio features and supporting a variety of audio processing tasks (such as classification, generation, recognition, etc.).
[0078] Of course, the pre-trained model can also be used for other models that can achieve latent space encoding, such as the GPT series of large models, etc., without making specific limitations here.
[0079] Step S203: Fuse the latent space representation vectors corresponding to the multiple acoustic feature blocks with the watermark information encoding vectors respectively to obtain the watermark embedding vectors corresponding to the multiple acoustic feature blocks respectively.
[0080] Here, the watermark information encoding vector can be a watermark codebook vector selected from the watermark codebook library. The vector obtained by encoding.
[0081] Optionally, a pre-trained watermark information encoder can be used to encode the watermark codebook vector. Here, the watermark information encoder is obtained by training the watermark codebook vector labeled with watermark information encoding vectors.
[0082] Optionally, the watermark information encoder can be a network architecture consisting of a 3-layer DNN (Deep Neural Network). Of course, the watermark information encoder can also adopt other network architectures, which are not limited in this application.
[0083] Preferably, the watermark information encoding vector can have the same dimension as the latent space representation vector, that is, if the dimension of the latent space representation vector is... Therefore, the label dimension used when training the watermark information encoder is also... .
[0084] When the watermark information encoding vector and the latent space representation vector have the same dimension, one possible fusion method is to add the latent space representation vector corresponding to each acoustic feature block to the watermark information encoding vector to obtain the watermark embedding vector corresponding to each acoustic feature block, i.e. ,in, The latent space represents a vector. This represents the watermark information encoding vector. This represents the watermark embedding vector.
[0085] In addition to the above-mentioned addition method, there are other fusion methods, such as direct splicing, weighted addition, etc. The appropriate fusion method can be selected according to the actual application scenario, and this application does not limit it.
[0086] It should also be noted that, for different acoustic feature blocks, the watermark information encoding vector fused with their corresponding latent space representation vectors can be the same or different.
[0087] Step S204: Decode and reconstruct the watermarked embedded audio containing watermark information based on the watermark embedding vectors corresponding to the multiple acoustic feature blocks.
[0088] In this embodiment, the watermark embedding vector corresponding to each acoustic feature block can be decoded to obtain the reconstructed feature block corresponding to each acoustic feature block. For example, the nth reconstructed feature block is denoted as... Next, based on all the reconstructed feature blocks, a watermark embedding vector containing the watermark information can be obtained.
[0089] Optionally, corresponding to the above, the decoder of the pre-trained model mentioned above can be used to decode the watermark embedding vector corresponding to each acoustic feature block to obtain the reconstructed feature block corresponding to each acoustic feature block.
[0090] The audio watermarking method provided in this application considers that embedding watermarks frame-by-frame on the amplitude spectrum involves many frequency points corresponding to redundant audio information. Embedding watermarks on these frequency points may interfere with the normal frequency distribution of the audio, making the watermark easily detectable. To avoid embedding watermarks on redundant audio information as much as possible, this application acquires the input audio, performs acoustic feature segmentation and latent space encoding on the input audio, obtaining latent space representation vectors corresponding to multiple acoustic feature blocks of the input audio. These latent space representation vectors are then fused with the watermark information encoding vectors to obtain watermark embedding vectors corresponding to multiple acoustic feature blocks. Based on these watermark embedding vectors, the watermark-embedded audio containing watermark information is decoded and reconstructed. This application uses latent space encoding to remove a large amount of redundant audio information from the acoustic feature blocks, retaining only the more critical and representative features of the audio essence. Embedding watermark information encoding vectors on these features minimizes the impact of the watermark on the overall audio characteristics, making the watermark less noticeable and improving the listening experience.
[0091] Furthermore, the latent space representation vector is composed of more core features with strong correlation and integrity. Embedding watermarks in these features with strong correlation and integrity allows the watermark information to be better integrated into the overall characteristics of the input audio due to the mutual constraints and coordination between features. This avoids excessive damage to individual features, thus maintaining the naturalness and coherence of the watermark embedded in the audio and improving the listening experience of the watermark embedded in the audio.
[0092] In some embodiments of this application, the process of step S204, "decoding and reconstructing watermark-embedded audio containing watermark information based on the watermark embedding vectors corresponding to multiple acoustic feature blocks," is described.
[0093] In one possible implementation, this embodiment can decode the watermark embedding vectors corresponding to multiple acoustic feature blocks respectively to obtain the reconstructed feature blocks corresponding to the multiple acoustic feature blocks respectively, and then directly generate the watermark embedded audio based on the reconstructed feature blocks corresponding to the multiple acoustic feature blocks respectively.
[0094] Specifically, the reconstructed feature blocks corresponding to multiple acoustic feature blocks can be spliced together to form a reconstructed feature spectrogram, and then audio restoration can be performed based on the reconstructed feature spectrogram to obtain the watermarked audio.
[0095] In another possible implementation, considering that the watermark may be slightly perceived due to factors such as low energy in the reconstructed feature blocks, in order to avoid these reconstructed feature blocks that may have slight watermark perception affecting the quality of the final watermark embedded audio, this embodiment can calculate the reconstruction quality of the reconstructed feature blocks corresponding to multiple acoustic feature blocks respectively, and then determine whether the reconstruction quality of the reconstructed feature blocks corresponding to multiple acoustic feature blocks meets the preset quality requirements.
[0096] Here, reconstruction quality can be used to quantitatively evaluate the perceptibility of watermarks embedded in reconstructed feature blocks. If the reconstruction quality of a reconstructed feature block meets the quality requirements, it means that the watermark embedded in the reconstructed feature block will hardly be noticed, and audio restoration can be performed directly according to the reconstructed feature block. Conversely, if the reconstruction quality of a reconstructed feature block does not meet the quality requirements, it means that the watermark embedded in the reconstructed feature block may be slightly perceptible, and the original acoustic feature block can be used to perform feature compensation replacement on the reconstructed feature block.
[0097] In other words, this embodiment identifies reconstructed feature blocks that do not meet quality requirements as target reconstructed feature blocks. Based on the acoustic feature blocks corresponding to these target reconstructed feature blocks, feature compensation and replacement are performed on the target reconstructed feature blocks to obtain updated reconstructed feature blocks. Then, after feature compensation and replacement, watermark-embedded audio is generated based on the reconstructed feature blocks corresponding to the multiple acoustic feature blocks or the updated reconstructed feature blocks. Specifically, watermark-embedded audio is generated based on the feature compensation and replacement results (updated reconstructed feature blocks) of the reconstructed feature blocks that meet quality requirements and those that do not (the updated reconstructed feature blocks).
[0098] There are multiple ways to implement the above-mentioned "determining whether the reconstruction quality of the reconstruction feature blocks corresponding to multiple acoustic feature blocks meets the preset quality requirements". Here, we provide, but are not limited to, the following two methods.
[0099] The first method involves calculating the energy of the reconstructed feature blocks corresponding to multiple acoustic feature blocks, using this energy as the reconstructed energy for each acoustic feature block, and determining whether the reconstruction quality of the reconstructed feature blocks corresponding to multiple acoustic feature blocks meets the quality requirements based on whether the reconstructed energy is greater than or equal to a preset energy threshold.
[0100] Specifically, if the reconstructed energy corresponding to an acoustic feature block is greater than or equal to the energy threshold, then the reconstruction quality of the reconstructed feature block corresponding to that acoustic feature block is determined to meet the quality requirements; if the reconstructed energy corresponding to an acoustic feature block is less than the energy threshold, then the reconstruction quality of the reconstructed feature block corresponding to that acoustic feature block is determined to not meet the quality requirements.
[0101] Optionally, the formula for calculating the reconstruction energy corresponding to the acoustic feature block is as follows: ,in, This represents the reconstructed feature block corresponding to the nth acoustic feature block. The i-th feature amplitude value in the matrix, i=1,2,…,256; This represents the energy of the reconstructed feature block corresponding to the nth acoustic feature block, which is also the reconstruction energy corresponding to the nth acoustic feature block.
[0102] Optionally, for each of the multiple acoustic feature blocks, if the reconstruction energy corresponding to the acoustic feature block is less than the energy threshold, then "performing feature compensation replacement of the target reconstruction feature block according to the acoustic feature block corresponding to the target reconstruction feature block" may include: replacing the reconstruction feature block corresponding to the acoustic feature block with the acoustic feature block.
[0103] Specifically, if the reconstructed energy corresponding to the nth acoustic feature block Below the energy threshold This indicates that the reconstructed feature block corresponding to the nth acoustic feature block has low energy. Therefore, using this low-energy reconstructed feature block to recover the audio may introduce noise interference into the recovered audio, affecting the listening experience. Thus, the nth acoustic feature block can be used instead. Replace the reconstructed feature block corresponding to the nth acoustic feature block The newly reconstructed feature block after compensation, obtained by feature block replacement, is denoted as... .
[0104] It is understandable that the low energy of the reconstructed feature block may be caused by the embedded watermark. In order to avoid poor audio restoration due to the watermark, this application provides a method of replacing it with the original acoustic feature block, which can avoid the audio listening experience degradation caused by the watermark to a certain extent and improve the listening experience of the watermark embedded audio to a certain extent.
[0105] The second method is to calculate the difference between each acoustic feature block and the corresponding reconstructed feature block for each of the multiple acoustic feature blocks. This difference is used as the reconstruction difference for that acoustic feature block. Based on whether the reconstruction differences corresponding to the multiple acoustic feature blocks are greater than or equal to a preset difference threshold, it is determined whether the reconstruction quality of the reconstructed feature blocks corresponding to the multiple acoustic feature blocks meets the quality requirements.
[0106] Optionally, the above-mentioned "calculating the difference between the acoustic feature block and the reconstructed feature block corresponding to the acoustic feature block" may include: calculating the mean square error between the acoustic feature block and the reconstructed feature block corresponding to the acoustic feature block, i.e. ,in, This represents the mean square error between the nth acoustic feature block and its corresponding reconstructed feature block.
[0107] If the reconstruction gap (such as the mean square error mentioned above) corresponding to an acoustic feature block is greater than or equal to the gap threshold, then the reconstruction quality of the reconstructed feature block corresponding to that acoustic feature block is determined to meet the quality requirements; if the reconstruction gap corresponding to an acoustic feature block is less than the gap threshold, then the reconstruction quality of the reconstructed feature block corresponding to that acoustic feature block is determined to not meet the quality requirements.
[0108] Of course, the above-mentioned gap can also be other indicator gaps, such as Frobenius norm, Chebyshev distance, similarity, etc. This embodiment does not make specific limitations.
[0109] Optionally, for each of the multiple acoustic feature blocks, if the reconstruction gap corresponding to the acoustic feature block is less than the gap threshold, then "performing feature compensation and replacement of the target reconstruction feature block according to the acoustic feature block corresponding to the target reconstruction feature block" may include: performing feature compensation on the reconstruction feature block corresponding to the acoustic feature block according to the acoustic feature block.
[0110] Specifically, if the reconstruction gap corresponding to the nth acoustic feature block is as follows: Below the gap threshold This indicates that the reconstruction effect of the reconstructed feature block corresponding to the nth acoustic feature block is poor. This means that the original audio has been significantly modified due to factors such as watermark embedding. Therefore, using this poorly reconstructed feature block to recover the audio will likely result in a poor listening experience. To avoid this problem, we can use the nth acoustic feature block... For the reconstructed feature block corresponding to the nth acoustic feature block Feature compensation is performed, and the newly reconstructed feature blocks obtained by feature block compensation are denoted as follows: .
[0111] Optionally, the process of "performing feature compensation for the reconstructed feature block corresponding to the acoustic feature block" may include: performing a weighted summation of the acoustic feature block and the reconstructed feature block corresponding to the acoustic feature block, for example, The fused feature block corresponding to the acoustic feature block is obtained and used as the updated reconstructed feature block corresponding to the acoustic feature block. Here, 'a' represents... The corresponding weight, b represents The corresponding weight is a+b=1.
[0112] The values of a and b can be determined based on the actual scenario. For example, in one possible scenario, a = 0.15 and b = 0.85.
[0113] It is understandable that the difference between the reconstructed feature block and the original acoustic feature block is likely caused by the embedded watermark. In order to avoid poor audio restoration due to the watermark, this application provides a method of compensation using the original acoustic feature block, which can avoid the audio listening experience degradation caused by the watermark to a certain extent and improve the listening experience of the watermark-embedded audio to a certain extent.
[0114] It should also be noted that the above feature compensation method can also be replaced by feature replacement method, which is not limited in this application.
[0115] There may be instances where the reconstructed energy corresponding to an acoustic feature block is less than the energy threshold and the reconstruction gap is less than the gap threshold. In such cases, the feature block can be replaced entirely using the aforementioned feature replacement method or compensated using the aforementioned feature compensation method. This application does not limit the choice.
[0116] After the feature compensation and replacement performed in the previous section, the updated reconstructed feature blocks are used according to the corresponding acoustic feature blocks. Or rebuild feature blocks (If there is no updated reconstructed feature block, select the reconstructed feature block; if there is an updated reconstructed feature block, select the updated reconstructed feature block.) Generate a watermark embedded in the audio.
[0117] In summary, this embodiment can correct the reconstructed feature block by feature compensation replacement when the reconstruction quality of the reconstructed feature block does not meet the requirements, so that the updated reconstructed feature block obtained after compensation has higher reconstruction quality, thereby restoring a better listening experience of the watermark embedded audio.
[0118] In some embodiments of this application, an audio watermark embedding model can be pre-trained to implement the processes of steps S201 to S204 mentioned above. In this embodiment, the input audio can be sent to the pre-trained audio watermark embedding model to obtain the watermark-embedded audio output by the model.
[0119] See Figure 3 The diagram shown is a structural schematic of an audio watermark embedding model provided in this application.
[0120] like Figure 3 Optionally, the audio watermark embedding model may include: an acoustic feature segmentation module 310, a latent space encoder 320, a watermark information encoder 330, a vector fusion module 340, a decoder 350, and an audio recovery module 360. Optionally, the latent space encoder 320 may be the encoder in the pre-trained model provided above, and the decoder 350 may be the decoder in the pre-trained model.
[0121] The acoustic feature segmentation module 310 segments the input audio into acoustic feature blocks, obtaining multiple acoustic feature blocks corresponding to the input audio; the latent space encoder 320 performs latent space encoding on each of the multiple acoustic feature blocks, obtaining latent space representation vectors corresponding to each of the multiple acoustic feature blocks; the watermark information encoder 330 encodes the watermark codebook vector obtained from the watermark codebook vector library into a watermark information encoding vector; the vector fusion module 340 fuses the latent space representation vectors corresponding to each of the multiple acoustic feature blocks with the watermark information encoding vectors, obtaining watermark embedding vectors corresponding to each of the multiple acoustic feature blocks; the decoder 350 decodes the watermark embedding vectors corresponding to each of the multiple acoustic feature blocks, obtaining reconstructed feature blocks corresponding to each of the multiple acoustic feature blocks; and the audio restoration module 360 generates watermark-embedded audio based on the reconstructed feature blocks corresponding to each of the multiple acoustic feature blocks.
[0122] In the case where feature compensation and replacement can be implemented as described above, the audio watermark embedding model may optionally include: a feature compensation and replacement module 370. This module 370 determines whether the reconstruction quality of the reconstructed feature blocks corresponding to the multiple acoustic feature blocks meets preset quality requirements. It then uses the reconstructed feature blocks that do not meet the quality requirements as target reconstructed feature blocks, and performs feature compensation and replacement on the target reconstructed feature blocks based on the acoustic feature blocks corresponding to them, obtaining updated reconstructed feature blocks. Correspondingly, the audio restoration module 360 generates watermark-embedded audio based on the reconstructed feature blocks corresponding to the multiple acoustic feature blocks or the updated reconstructed feature blocks.
[0123] In one possible implementation, this embodiment also provides a watermark extraction model, which is jointly trained with an audio watermark embedding model, as described below. Figure 4 The training process of the joint training is introduced.
[0124] like Figure 4 In this embodiment, the original audio sample can be obtained and fed into the audio watermark embedding model to obtain a watermarked audio sample containing watermark information.
[0125] Optionally, the processing of the original audio samples in the audio watermark embedding model can be the same as described above. For details, please refer to the previous introduction, which will not be repeated here.
[0126] To further train the model's watermark-free embedding capability, optionally, for the original audio sample, before sending each acoustic feature block of the original audio sample into the latent space encoder 320, each acoustic feature block of the original audio sample can be masked separately, so that the audio watermark embedding model can achieve watermark-free embedding even when some audio information is missing.
[0127] Alternatively, to avoid overfitting during the training of the audio watermark embedding model, a certain amount of random noise can be added after summing the latent space representation vector and the watermark information encoding vector. , This represents Gaussian noise with a mean of 0 and a variance of 1.
[0128] Of course, the above-mentioned noise forms can also be other, and are not limited here.
[0129] After obtaining the watermarked audio sample containing the watermark information, this embodiment can calculate the first loss between the acoustic features of the original audio sample and the watermarked audio sample.
[0130] Taking the mean square error loss, where the acoustic features are the short-time Fourier transform (STFT) features of both the original audio sample and the watermarked audio sample, as an example, the first loss can be: ,in, Indicates the first loss. To find the norm, The STFT features representing the original audio sample, This represents the STFT feature of the watermarked audio sample.
[0131] In order to simulate the various processing and damage that audio may encounter in practical applications, this application can send the watermarked audio sample into the simulation attack layer to simulate the attack and obtain the attacked audio sample.
[0132] For example, one attack method can be randomly selected from the attack methods to simulate an attack on all watermarked audio samples in the same training batch, thus obtaining the attacked audio samples corresponding to each watermarked audio sample.
[0133] Of course, different attack methods can be used to simulate attacks on different watermarked audio samples in the same training batch; this is not limited here.
[0134] Optionally, the attack methods may include, but are not limited to, various time-domain signal attacks such as environmental noise and packet loss.
[0135] Furthermore, watermark extraction is performed on the attacked audio samples to obtain the watermark codebook recovery vector.
[0136] like Figure 4 Optionally, a watermark extraction model 410 can be preset, and the watermark codebook recovery vector can be extracted from the attacked audio sample through the watermark extraction model 410.
[0137] Optionally, the watermark extraction model 410 can consist of a deep residual network with a network depth of 34 and an average pooling layer. The attacked audio sample can then be fed into the deep residual network, followed by the average pooling layer, ultimately recovering the watermark codebook vector. Extraction.
[0138] Next, the watermark codebook recovery vector is calculated. and watermark codebook vector The difference loss between them is used as the second loss. Here, the watermark codebook vector refers to the vector extracted from the watermark codebook vector library above for encoding the watermark information encoding vector.
[0139] Optionally, the formula for calculating the second loss is: ,in, Represents the watermark codebook recovery vector and watermark codebook vector The cosine distance between them This indicates the second loss.
[0140] Finally, the audio watermark embedding model and the watermark extraction model 410 are trained based on the first loss and the second loss. For example, the first loss and the second loss are added together, and the sum is used as the global loss, i.e. , If we represent global loss, then we can utilize global loss. Complete the network gradient update for the data in the training batch.
[0141] By repeating the above process several times, the trained audio watermark embedding model and the optimized watermark extraction model can be obtained.
[0142] Through the above training process, not only can an audio watermark embedding model that can seamlessly embed watermarks be trained, but also a corresponding watermark extraction model can be trained. This enables the present application to not only seamlessly embed watermarks into audio, but also accurately extract watermarks from audio, thus increasing its practical value.
[0143] The above describes an audio watermark generation method provided by the embodiments of this application. The following will describe the apparatus for performing the above audio watermark generation method.
[0144] Please see Figure 5 , Figure 5 This is a schematic diagram of an audio watermark generation device provided in an embodiment of this application. Figure 5 As shown, the audio watermark generation device may include:
[0145] Data input unit 501 is used to acquire input audio;
[0146] The data processing unit 502 is used to perform acoustic feature segmentation and latent space encoding on the input audio to obtain latent space representation vectors corresponding to multiple acoustic feature blocks of the input audio. The latent space representation vectors corresponding to multiple acoustic feature blocks are fused with the watermark information encoding vector to obtain watermark embedding vectors corresponding to multiple acoustic feature blocks. The watermark embedded audio containing watermark information is decoded and reconstructed based on the watermark embedding vectors corresponding to multiple acoustic feature blocks.
[0147] Each module in the aforementioned audio watermark generation device can be implemented entirely or partially through software, hardware, or a combination thereof. These modules can be embedded in or independent of the processor in a computer device, or stored in the memory of a computer device as software, so that the processor can call and execute the corresponding operations of each module.
[0148] This application also provides an electronic device, which may include at least one processor and a memory connected to the processor, wherein:
[0149] Memory is used to store computer programs;
[0150] The processor is used to execute computer programs to enable electronic devices to implement any of the audio watermark generation methods provided in the embodiments of this application.
[0151] refer to Figure 6 The diagram illustrates a structural schematic suitable for implementing the electronic device in the embodiments of this application. The electronic device in the embodiments of this application may include, but is not limited to, fixed terminals such as mobile phones, laptops, PDAs (personal digital assistants), PADs (tablet computers), desktop computers, etc. Figure 6 The electronic device shown is merely an example and should not impose any limitation on the functionality and scope of use of the embodiments of this application.
[0152] like Figure 6 As shown, the electronic device may include a processing unit (e.g., a central processing unit, a graphics processing unit, etc.) 601, which can perform various appropriate actions and processes according to a program stored in a read-only memory (ROM) 602 or a program loaded from a storage device 608 into a random access memory (RAM) 603. When the electronic device is powered on, the RAM 603 also stores various programs and data required for the operation of the electronic device. The processing unit 601, ROM 602, and RAM 603 are interconnected via a bus 604. An input / output (I / O) interface 605 is also connected to the bus 604.
[0153] Typically, the following devices can be connected to I / O interface 605: input devices 606 including, for example, touchscreens, touchpads, keyboards, mice, cameras, microphones, accelerometers, gyroscopes, etc.; output devices 607 including, for example, liquid crystal displays (LCDs), speakers, vibrators, etc.; storage devices 608 including, for example, memory cards, hard drives, etc.; and communication devices 609. Communication device 609 allows electronic devices to communicate wirelessly or wiredly with other devices to exchange data. Although Figure 6 Electronic devices with various devices are shown, but it should be understood that it is not required to implement or have all of the devices shown. More or fewer devices may be implemented or have instead.
[0154] This application also provides a computer program product including computer-readable instructions, which, when executed on an electronic device, cause the electronic device to implement any of the audio watermark generation methods provided in this application.
[0155] This application also provides a computer-readable storage medium that carries one or more computer programs. When the one or more computer programs are executed by an electronic device, the electronic device can implement any of the audio watermark generation methods provided in this application.
[0156] It should also be noted that the device embodiments described above are merely illustrative. The units described as separate components may or may not be physically separate, and the components shown as units may or may not be physical units; that is, they may be located in one place or distributed across multiple network units. Some or all of the modules can be selected to achieve the purpose of this embodiment according to actual needs. In addition, in the device embodiment drawings provided in this application, the connection relationship between modules indicates that they have a communication connection, which can be implemented as one or more communication buses or signal lines.
[0157] Through the above description of the embodiments, those skilled in the art can clearly understand that this application can be implemented by means of software plus necessary general-purpose hardware, or it can be implemented by special-purpose hardware including application-specific integrated circuits, special-purpose CPUs, special-purpose memory, special-purpose components, etc. Generally, any function performed by a computer program can be easily implemented by corresponding hardware, and the specific hardware structure used to implement the same function can also be diverse, such as analog circuits, digital circuits, or special-purpose circuits. However, for this application, software program implementation is more often the preferred implementation method. Based on this understanding, the technical solution of this application, in essence, or the part that contributes to the prior art, can be embodied in the form of a software product. This computer software product is stored in a readable storage medium, such as a computer floppy disk, USB flash drive, mobile hard disk, ROM, RAM, magnetic disk, or optical disk, etc., and includes several instructions to cause a computer device (which may be a personal computer, training equipment, or network device, etc.) to execute the methods described in the various embodiments of this application.
[0158] In the above embodiments, implementation can be achieved, in whole or in part, through software, hardware, firmware, or any combination thereof. When implemented in software, it can be implemented, in whole or in part, as a computer program product.
[0159] The computer program product includes one or more computer instructions. When the computer program instructions are loaded and executed on a computer, all or part of the processes or functions described in the embodiments of this application are generated. The computer may be a general-purpose computer, a special-purpose computer, a computer network, or other programmable device. The computer instructions may be stored in a computer-readable storage medium or transmitted from one computer-readable storage medium to another. For example, the computer instructions may be transmitted from one website, computer, training device, or data center to another website, computer, training device, or data center via wired (e.g., coaxial cable, fiber optic, digital subscriber line (DSL)) or wireless (e.g., infrared, wireless, microwave, etc.) means. The computer-readable storage medium may be any available medium that a computer can store or a data storage device such as a training device or data center that integrates one or more available media. The available media may be magnetic media (e.g., floppy disks, hard disks, magnetic tapes), optical media (e.g., DVDs), or semiconductor media (e.g., solid-state drives (SSDs)).
Claims
1. A method for generating audio watermarks, characterized in that, include: Get the input audio; The input audio is subjected to acoustic feature block division and latent space encoding to obtain latent space representation vectors corresponding to multiple acoustic feature blocks of the input audio. The latent space representation vectors corresponding to the multiple acoustic feature blocks are fused with the watermark information encoding vectors to obtain the watermark embedding vectors corresponding to the multiple acoustic feature blocks. The watermarked audio containing watermark information is reconstructed by decoding the watermark embedding vectors corresponding to the multiple acoustic feature blocks.
2. The audio watermark generation method according to claim 1, characterized in that, The step of decoding and reconstructing the watermark-embedded audio containing watermark information based on the watermark embedding vectors corresponding to the plurality of acoustic feature blocks includes: The watermark embedding vectors corresponding to the plurality of acoustic feature blocks are decoded respectively to obtain the reconstructed feature blocks corresponding to the plurality of acoustic feature blocks respectively; Determine whether the reconstruction quality of the reconstruction feature blocks corresponding to the plurality of acoustic feature blocks meets the preset quality requirements; The reconstructed feature blocks that do not meet the quality requirements are taken as target reconstructed feature blocks. The target reconstructed feature blocks are then replaced with feature compensation based on the acoustic feature blocks corresponding to them, resulting in updated reconstructed feature blocks. The watermark-embedded audio is generated based on the reconstructed feature blocks or updated reconstructed feature blocks corresponding to the plurality of acoustic feature blocks.
3. The audio watermark generation method according to claim 2, characterized in that, The step of determining whether the reconstruction quality of the reconstructed feature blocks corresponding to the plurality of acoustic feature blocks meets the preset quality requirements includes: Calculate the energy of the reconstructed feature block corresponding to each of the plurality of acoustic feature blocks, and use it as the reconstructed energy corresponding to each of the plurality of acoustic feature blocks; Based on whether the reconstruction energy corresponding to each of the plurality of acoustic feature blocks is greater than or equal to a preset energy threshold, it is determined whether the reconstruction quality of the reconstruction feature blocks corresponding to each of the plurality of acoustic feature blocks meets the quality requirements. And / or, For each of the plurality of acoustic feature blocks, the difference between the acoustic feature block and the corresponding reconstructed feature block is calculated, and this difference is taken as the reconstruction difference corresponding to the acoustic feature block. Based on whether the reconstruction gap corresponding to the plurality of acoustic feature blocks is greater than or equal to a preset gap threshold, it is determined whether the reconstruction quality of the reconstruction feature blocks corresponding to the plurality of acoustic feature blocks meets the quality requirements.
4. The audio watermark generation method according to claim 3, characterized in that, The calculation of the difference between the acoustic feature block and the corresponding reconstructed feature block includes: Calculate the mean square error of the acoustic feature block and the corresponding reconstructed feature block.
5. The audio watermark generation method according to claim 3, characterized in that, The step of performing feature compensation replacement on the target reconstructed feature block based on the acoustic feature block corresponding to the target reconstructed feature block includes: For each of the plurality of acoustic feature blocks: If the reconstruction gap corresponding to the acoustic feature block is less than the gap threshold, then feature compensation is performed on the reconstruction feature block corresponding to the acoustic feature block based on the acoustic feature block. If the reconstructed energy corresponding to the acoustic feature block is less than the energy threshold, then the reconstructed feature block corresponding to the acoustic feature block is replaced with the acoustic feature block.
6. The audio watermark generation method according to claim 5, characterized in that, The step of performing feature compensation on the reconstructed feature block corresponding to the acoustic feature block based on the acoustic feature block includes: The acoustic feature block and its corresponding reconstructed feature block are weighted and summed to obtain the fused feature block, which is then used as the updated reconstructed feature block.
7. The audio watermark generation method according to any one of claims 1-6, characterized in that, The input audio is subjected to acoustic feature segmentation and latent space encoding to obtain latent space representation vectors corresponding to multiple acoustic feature blocks of the input audio. These latent space representation vectors are then fused with watermark information encoding vectors to obtain watermark embedding vectors corresponding to the multiple acoustic feature blocks. Watermark-embedded audio containing watermark information is then decoded and reconstructed based on these watermark embedding vectors, including: The input audio is fed into a pre-trained audio watermark embedding model to obtain the watermark-embedded audio output by the model.
8. The audio watermark generation method according to claim 7, characterized in that, The audio watermark embedding model and the watermark extraction model are jointly trained; The training process of the joint training includes: Obtain the original audio sample; The original audio sample is fed into the audio watermark embedding model to obtain a watermarked audio sample containing watermark information. Calculate the first loss between the acoustic features of the original audio sample and the watermarked audio sample; The watermarked audio sample is sent to the simulation attack layer to simulate the attack and obtain the attacked audio sample. The watermark extraction model is used to extract the watermark from the attacked audio sample to obtain the watermark codebook recovery vector. Calculate the second loss between the watermark codebook recovery vector and the watermark codebook vector, wherein the watermark codebook vector is used to encode the watermark information encoding vector. The audio watermark embedding model and the watermark extraction model are trained based on the first loss and the second loss.
9. A computer program product, characterized in that, It includes computer-readable instructions that, when executed on an electronic device, cause the electronic device to implement the audio watermark generation method as described in any one of claims 1 to 8.
10. An electronic device, characterized in that, It includes at least one processor and a memory connected to the processor, wherein: The memory is used to store computer programs; The processor is used to execute the computer program to enable the electronic device to implement the audio watermark generation method as described in any one of claims 1 to 8.
11. A computer storage medium, characterized in that, The storage medium carries one or more computer programs that, when executed by an electronic device, enable the electronic device to implement the audio watermark generation method as described in any one of claims 1 to 8.
Citation Information
Patent Citations
Time-domain digital audio watermarking method based on audio breakpoint
CN104810022A
Audio watermark processing method and device, equipment and medium
CN120636419A
Audio system and method for acoustic echo cancellation
EP1832104A1
Face-speech bridging by cycle video / audio reconstruction
US10931976B1
Inserting audio channels into descriptions of soundfields
US20150271621A1