Audio watermark generation method, device and equipment and computer storage medium

By using potential vectors to characterize consistency in the audio watermark generation network, optimizing the network, building triple losses and updating network parameters, the problem that audio watermark embedded networks cannot effectively characterize audio differences in the prior art, and improving the effect of audio watermark generation.

CN120236593AActive Publication Date: 2025-07-01IFLYTEK CO LTD

Patent Information

Application Number
CN202510707347.8
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-05-29
Publication Date
2025-07-01
Estimated Expiration
2045-05-29

AI Technical Summary

Technical Problem

Existing audio watermark embedding networks based on neural networks cannot effectively characterize the differences between audio before and after adding watermarks, resulting in the loss function being unable to fully capture audio changes.

Method used

An audio watermark generation method is proposed, by obtaining the audio dataset, extracting potential vector representations of target audio and watermark audio, constructing triple losses, and updating the audio watermark generation network to align the audio features before and after adding watermarks in the latent space.

Benefits of technology

Through the latent vector characterization consistency optimization network, the differences in audio before and after adding watermarks can be better characterized, and the parameters of audio watermark embedded in the network can be updated, thereby improving the effect of audio watermark generation.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120236593A_ABST
    Figure CN120236593A_ABST
Patent Text Reader

Abstract

The invention provides an audio watermark generation method, an audio watermark generation device, audio watermark generation equipment and a computer storage medium. The audio watermark generation method comprises the following steps: extracting a first potential vector representation of target audio data; inputting the target audio data into an audio watermark embedding network of an audio watermark generation network, adding watermark data, and obtaining watermark audio data; extracting a second potential vector representation of the watermark audio data; extracting third potential vector representation of other audio data; using the first potential vector representation, the second potential vector representation and the third potential vector representation to construct a triple loss; and the audio watermark generation network is updated by using the triple loss, and the audio watermark embedding network is used for generating the audio watermark after being updated. According to the audio watermark generation method, the audio features before and after the watermark is added are aligned in the potential space, so that the network parameters of the audio watermark embedding network are updated, and the audio watermark generation effect is improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present application relates to the technical field of digital watermarking, and particularly to an audio watermark generation method, an audio watermark generation device, an audio watermark generation equipment, and a computer storage medium. Background Art

[0002] For the audio watermark embedding and extraction method based on neural network, first, perform short-time Fourier transform (STFT, Short-Time Fourier Transform) on the original audio signal to separate the amplitude spectrum and phase spectrum features. By splicing the watermark information encoding and the amplitude spectrum features in the feature dimension, a joint feature tensor is constructed as the input of the convolutional neural network, and then the watermark embedding weight matrix corresponding to each frequency point of the amplitude spectrum is learned. After multiplying the weight matrix element by element with the original amplitude spectrum and combining the original phase spectrum, the watermark-embedded audio containing the watermark information is reconstructed through inverse short-time Fourier transform (iSTFT).

[0003] The existing watermark embedding network based on neural network only relies on calculating the mean square error between the STFT features of the original audio and the audio after adding the watermark to update the network parameters. Summary of the Invention

[0004] To solve the above technical problems, the present application proposes an audio watermark generation method, an audio watermark generation device, an audio watermark generation equipment, and a computer storage medium.

[0005] To solve the above technical problems, the present application proposes an audio watermark generation method, and the audio watermark generation method includes: Obtain the audio data set of the current batch, where the audio data set includes target audio data and other audio data; Extract the first latent vector representation of the target audio data; Input the target audio data into the audio watermark embedding network of the audio watermark generation network to add watermark data, and obtain watermarked audio data; Extract the second latent vector representation of the watermarked audio data; Extract the third latent vector representation of the other audio data; Construct a triplet loss by using the first latent vector representation, the second latent vector representation, and the third latent vector representation; Update the audio watermark generation network by using the triplet loss, where the audio watermark embedding network is used to generate audio watermarks after being updated.

[0006] Wherein, the audio watermark generation method further includes: Extract the first short-time Fourier transform feature of the target audio data; Extract the second short-time Fourier transform feature of the watermarked audio data; Construct an error loss by using the first short-time Fourier transform feature and the second short-time Fourier transform feature; The updating of the audio watermark generation network by using the triplet loss includes: Update the audio watermark generation network by using the triplet loss and the error loss.

[0007] Among them, the adding of watermark data to the audio watermark embedding network of the audio watermark generation network by inputting the target audio data to obtain watermarked audio data includes: Input the target audio data into the audio watermark embedding network to add several different watermark data respectively to obtain several watermarked audio data; The constructing of the triplet loss by using the first latent vector representation, the second latent vector representation, and the third latent vector representation includes: Construct intermediate triplet losses respectively by using the second latent vector representation of each watermarked audio data, the first latent vector representation, and the third latent vector representation; Construct the triplet loss by using the intermediate triplet losses corresponding to the several watermarked audio data.

[0008] Among them, the extracting of the third latent vector representation of the other audio data includes: Extract the third latent vector representation of the other audio data before adding watermark data, or the third latent vector representation of the other audio data after adding watermark data.

[0009] Among them, the audio watermark generation method further includes: Input the watermark data into the audio watermark extraction network of the audio watermark generation network to obtain the first watermark information encoding; Select the simulation attack method for the audio data set of the current batch; Perform attack simulation on the watermarked audio data according to the simulation attack method to generate attack audio data; Input the attack audio data into the audio watermark extraction network to extract the second watermark information encoding; Construct a difference loss by using the first watermark information encoding and the second watermark information encoding; The updating of the audio watermark generation network by using the triplet loss includes: Form a global loss by using the triplet loss and the difference loss; Update the audio watermark generation network by using the global loss.

[0010] Among them, the simulation attack method of the audio data set of other batches is different from that of the audio data set of the current batch.

[0011] Among them, inputting the target audio data into the audio watermark embedding network of the audio watermark generation network to add watermark data to obtain watermarked audio data includes: Inputting the target audio data into the audio watermark embedding network to extract the amplitude spectrum feature and phase spectrum feature of the target audio data; Selecting one of the watermark information encodings from the watermark codebook library, and performing frame-level replication according to the number of frames of the amplitude spectrum feature to obtain watermark data; Performing frame-level splicing on the amplitude spectrum feature and the watermark data to obtain a watermarked amplitude spectrum feature; Extracting a watermark modulation weight matrix based on the watermarked amplitude spectrum feature; Fusing the amplitude spectrum feature and the watermark modulation weight matrix to obtain a modulated amplitude spectrum feature; Combining and performing inverse Fourier transform on the modulated amplitude spectrum feature and the phase spectrum feature to obtain the watermarked audio data.

[0012] To solve the above technical problems, the present application also proposes an audio watermark generation device, which includes: an acquisition module, an extraction module, a generation module, and an update module; among them, The acquisition module is used to acquire the audio data set of the current batch, where the audio data set includes target audio data and other audio data; The extraction module is used to extract the first latent vector representation of the target audio data; The generation module is used to input the target audio data into the audio watermark embedding network of the audio watermark generation network to add watermark data to obtain watermarked audio data; The extraction module is used to extract the second latent vector representation of the watermarked audio data; The extraction module is used to extract the third latent vector representation of the other audio data; The update module is used to construct a triplet loss by using the first latent vector representation, the second latent vector representation, and the third latent vector representation; The update module is used to update the audio watermark generation network by using the triplet loss, where the audio watermark embedding network is updated to generate audio watermarks after the update.

[0013] To solve the above technical problems, the present application also proposes an audio watermark generation device, which includes a memory and a processor coupled to the memory; wherein, the memory is used to store program data, and the processor is used to execute the program data to implement the audio watermark generation method as described above.

[0014] To solve the above technical problems, the present application also proposes a computer storage medium, which is used to store program data, and when the program data is executed by a computer, it is used to implement the above audio watermark generation method.

[0015] Compared with the prior art, the beneficial effects of the present application are as follows: The audio watermark generation device obtains the current batch of audio data sets, where the audio data sets include target audio data and other audio data; extracts the first latent vector representation of the target audio data; inputs the target audio data into the audio watermark embedding network of the audio watermark generation network to add watermark data to obtain watermarked audio data; extracts the second latent vector representation of the watermarked audio data; extracts the third latent vector representation of the other audio data; constructs a triplet loss using the first latent vector representation, the second latent vector representation, and the third latent vector representation; and updates the audio watermark generation network using the triplet loss, where the audio watermark embedding network is updated to generate audio watermarks after the update. Through the above audio watermark generation method, the audio features before and after adding watermarks are aligned in the latent space to update the network parameters of the audio watermark embedding network, thereby improving the audio watermark generation effect. Description of the Drawings

[0016] To more clearly illustrate the technical solutions in the embodiments of the present application, the following will briefly introduce the drawings required for the description of the embodiments. Obviously, the following drawings are only some embodiments of the present application. For those of ordinary skill in the art, without creative efforts, other drawings can also be obtained based on these drawings. Among them: Figure 1 is a schematic flowchart of the first embodiment of the audio watermark generation method provided by the present application; Figure 2 is a schematic architecture diagram of the audio watermark based on neural network and random alternating simulation attack provided by the present application; Figure 3 is Figure 1 a specific schematic flowchart of step S13 of the audio watermark generation method shown; Figure 4 is a schematic flowchart of the second embodiment of the audio watermark generation method provided by the present application; Figure 5It is a schematic structural diagram of an embodiment of the audio watermark generation device provided by the present application; Figure 6 It is a schematic structural diagram of an embodiment of the audio watermark generation device provided by the present application; Figure 7 It is a schematic structural diagram of an embodiment of the computer storage medium provided by the present application. Specific embodiments

[0017] Next, the technical solutions in the embodiments of the present application will be clearly and completely described in conjunction with the accompanying drawings in the embodiments of the present application. Obviously, the described embodiments are only a part of the embodiments of the present application, rather than all the embodiments. Based on the embodiments in the present application, all other embodiments obtained by those of ordinary skill in the art without creative efforts shall fall within the protection scope of the present application.

[0018] The terms "first", "second", "third", "fourth", etc. (if any) in the specification and claims of the present application and the above accompanying drawings are used to distinguish similar objects, and do not necessarily need to be used to describe a specific order or sequence. It should be understood that such data can be interchanged under appropriate circumstances so that the embodiments of the present application described herein, for example, can be implemented in an order other than those illustrated or described herein. In addition, the terms "comprising" and "having" and any variations thereof are intended to cover non-exclusive inclusion. For example, a process, method, system, product, or device that includes a series of steps or units does not necessarily need to be limited to those steps or units clearly listed, but may include other steps or units not clearly listed or inherent to these processes, methods, products, or devices.

[0019] With the increasing popularity of digital multimedia (including but not limited to audio, images, and videos, etc.) and the profound transformation of the dissemination methods, the copyright protection and authenticity verification of these contents have become increasingly important. Because some low-cost tampering, forgery, and splicing methods can batch modify or generate information carriers transmitted in the mobile Internet, posing a cognitive risk to ordinary audiences. Audio, as one of the widely used information carriers in digital media, also brings corresponding security challenges due to its characteristics of rapid dissemination and easy replication and modification. Therefore, ensuring the integrity protection and copyright guarantee of audio has become a focus area of research.

[0020] Traditional protection measures for digital multimedia information, similar to the method of encoding first and then decoding in a communication system, that is, encrypting the multimedia information carrier to make the multimedia information carrier in a chaotic state. It is invisible to the outside world during information transmission. Even if intercepted midway, without the corresponding decoding algorithm, unauthorized users cannot obtain the valid information carried in the multimedia information carrier, thus achieving the purpose of protecting digital media and copyright. This protection method assumes that unauthorized persons cannot obtain the corresponding decoding algorithm. However, with the progress of technology, the difficulty of breaking the encoding and decoding algorithms has been significantly reduced. Unauthorized persons can illegally decode the information transmitted in ciphertext from the mobile Internet, implant harmful information in it, and then use the same encoding algorithm to release it back to the mobile Internet for transmission. Due to the lack of effective authentication, it is impossible to distinguish the illegally tampered multimedia carrier, causing cognitive disruption to relevant users of the mobile Internet and resulting in the risk of public transmission vulnerabilities.

[0021] Digital watermarking technology provides a hidden and effective means to indicate the copyright ownership or confirm the authenticity of the content by embedding imperceptible information into digital works. The key challenge of this technology is to ensure that the embedded watermark is not only transparent to the end user but also resistant to various signal processing operations. Digital watermarking technology is a crucial research branch in the field of information hiding. It can effectively protect the security of information without affecting its visibility. Different from traditional encryption technologies, digital watermarks protect the content by leaving unique copyright marks on digital media without restricting its normal use. The watermarks embedded in digital media maintain the readability of the content and enhance its security. Especially in the field of digital audio, using audio digital watermarks can achieve multiple functions such as copyright protection, content verification, and copy control, which has significant research significance for enhancing the security of audio.

[0022] However, the development of audio watermarking faces more complex challenges than image and video watermarking. In addition to balancing imperceptibility and robustness, it must also take into account the high sensitivity of the human auditory system, which makes the requirements for audio quality more stringent. Therefore, in-depth research on audio digital watermarking technology is not only a technical requirement but also an inevitable trend in the pursuit of high-quality audio experience.

[0023] Existing watermark embedding networks based on neural networks, due to the problems of fixed time-frequency resolution and spectral leakage in STFT, the loss functions they provide cannot fully represent the differences between audio before and after adding watermarks.

[0024] Therefore, the audio watermark generation method provided by this application designs a potential vector representation consistency optimization network to align the audio before and after adding the watermark in the latent space, which can well represent the difference between the audio before and after adding the watermark.

[0025] For details, please refer to Figure 1 and Figure 2 , Figure 1 which is a schematic flowchart of the first embodiment of the audio watermark generation method provided by this application. Figure 2 And is a schematic architecture diagram of the audio watermark based on neural network and random alternating simulation attack provided by this application.

[0026] The audio watermark generation method of this application is applied to an audio watermark generation device. Among them, the audio watermark generation device of this application can be a server, a terminal device, or a system composed of a server and a terminal device cooperating with each other. Correspondingly, each part included in the audio watermark generation device, such as each unit, subunit, module, and submodule, can be all set in the server, all set in the terminal device, or respectively set in the server and the terminal device.

[0027] Furthermore, the above server can be hardware or software. When the server is hardware, it can be implemented as a distributed server cluster composed of multiple servers or as a single server. When the server is software, it can be implemented as multiple software or software modules, such as software or software modules used to provide a distributed server, or as a single software or software module, which is not specifically limited here.

[0028] As Figure 1 shown, the specific steps are as follows: Step S11: Obtain the audio data set of the current batch, where the audio data set includes target audio data and other audio data.

[0029] In the embodiment of this application, the audio data set of the current batch is used as the training data set for updating the network parameters of the Figure 2 shown audio watermark generation network or audio watermark generation architecture, and the audio data set of the current batch is used for network update in the current iteration stage.

[0030] Among them, the target audio data is the audio data currently used for audio watermark generation, and the other audio data is any audio data in the current batch of audio data set except the target audio data.

[0031] Step S12: Extract the first potential vector representation of the target audio data.

[0032] In an embodiment of the present application, the audio watermark generation device extracts the first potential vector representation of the target audio data before adding the watermark. Among them, extracting the potential vector representation of the audio is to convert the original audio signal into a low-dimensional and dense vector for use in machine learning tasks, such as the watermark generation task of the present application.

[0033] Specifically, as Figure 2 shown, the audio watermark generation device inputs the target audio data before adding the watermark into the potential vector representation consistency optimization network to extract the relevant potential vector representation.

[0034] In a specific embodiment, the main network of the potential vector representation consistency optimization network provided by the present application can be a ViT basic network (Vision Transformer) for directly extracting the first potential vector representation of the target audio data .

[0035] For the watermarked audio data after adding the watermark, the potential vector representation consistency optimization network extracts the relevant second potential vector representation . For other audio data, the potential vector representation consistency optimization network extracts the relevant third potential vector representation .

[0036] Step S13: Input the target audio data into the audio watermark embedding network of the audio watermark generation network to add watermark data, and obtain watermarked audio data.

[0037] In an embodiment of the present application, the audio watermark generation device inputs the target audio data into the audio watermark embedding network as Figure 2 shown, which is used to encode and fuse the watermark information in the watermark codebook library into the audio data to add the relevant watermark data and generate watermarked audio data.

[0038] As Figure 2 shown, the audio watermark generation network of the present application specifically includes an audio watermark embedding network, a potential vector representation consistency optimization network, and an audio watermark extraction network.

[0039] Regarding the watermark addition principle of the audio watermark embedding network provided by the present application, please continue to refer to Figure 2 and Figure 3 , Figure 3 which is Figure 1 the specific process schematic diagram of step S13 of the audio watermark generation method shown.

[0040] As Figure 3 shown, the specific steps are as follows: Step S131: Input the target audio data into the audio watermark embedding network to extract the amplitude spectrum feature and phase spectrum feature of the target audio data.

[0041] In the embodiment of the present application, after the audio watermark embedding network performs windowing and short-time Fourier transform on the audio data, the amplitude spectrum feature and phase spectrum feature of the STFT are obtained, and their feature sizes are both , where represents the number of feature frames, represents the feature dimension.

[0042] Step S132: Select one of the watermark information encodings from the watermark codebook library, and perform frame-level replication according to the number of frames of the amplitude spectrum feature to obtain watermark data.

[0043] In the embodiment of the present application, the audio watermark embedding network randomly selects a watermark information encoding from the watermark codebook library and performs frame-level replication according to the number of frames of the amplitude spectrum feature. The size of the replicated watermark encoding is , where represents the number of feature frames, represents the number of bits of the watermark information encoding.

[0044] Step S133: Perform frame-level splicing on the amplitude spectrum feature and the watermark data to obtain the watermark amplitude spectrum feature.

[0045] In the embodiment of the present application, the audio watermark embedding network performs frame-level splicing on the amplitude spectrum feature extracted in step S131 and the watermark data extracted in step S132 to obtain the amplitude spectrum feature incorporating the watermark information, that is, the watermark amplitude spectrum feature, and its feature size is .

[0046] Step S134: Extract the watermark modulation weight matrix based on the watermark amplitude spectrum feature.

[0047] In the embodiment of the present application, the watermark modulation weight matrix is an adaptive mask in deep watermark embedding, and is generated by a neural network (such as Figure 2 the deep residual network (Deep Residual Network) with a network depth of 50 shown), and is mainly used to slightly adjust the amplitude spectrum, maintain inaudible loss; resist signal processing attacks; dynamically optimize the embedding position and intensity according to the audio content.

[0048] Compared with the traditional method, the watermark modulation weight matrix combined with deep learning can better balance inaudibility and robustness.

[0049] The watermark modulation weight matrix is a weight matrix with the same size as the original amplitude spectrum, and each of its elements represents: How to adjust the original amplitude spectrum to embed the watermark while minimizing the damage to the audio quality.

[0050] The value range is usually where is the modulation intensity, such as 0.05, to ensure fine-tuning rather than complete replacement.

[0051] Step S135: Fuse the amplitude spectrum feature with the watermark modulation weight matrix to obtain the modulated amplitude spectrum feature.

[0052] In the embodiment of the present application, the audio watermark embedding network multiplies the original amplitude spectrum feature with the watermark modulation weight matrix to obtain the amplitude spectrum feature modulated by the watermark information, that is, the modulated amplitude spectrum feature.

[0053] Step S136: Combine the modulated amplitude spectrum feature with the phase spectrum feature and perform inverse Fourier transform to obtain the watermarked audio data.

[0054] In the embodiment of the present application, the audio watermark embedding network combines the phase spectrum feature with the modulated amplitude spectrum feature generated in step S135 and performs inverse Fourier transform to obtain the audio sample with the watermark information added, that is, the watermarked audio data.

[0055] Furthermore, in the above step S132, the audio watermark embedding network needs to select multiple watermark information from the watermark codebook library, so as to add watermarks to the same audio sample with different watermark information.

[0056] Specifically, in the training process of watermark embedding, the current watermark embedding technology only randomly adds one watermark information to each training audio sample each time, which is likely to cause the problem of watermark information dependence. That is, for the same audio sample, different watermark information is selected for embedding during the training of the watermark embedding network, and different network parameter update gradients will be obtained.

[0057] Therefore, the audio watermark embedding network provided by the present application selects to adopt the method of adding multiple groups of watermarks together, and adds multiple groups of watermark information to the same training audio sample at the same time to solve the information dependence problem brought by a single watermark.

[0058] Specifically, the audio watermark generation device inputs the target audio data into the audio watermark embedding network to add several different watermark data respectively, and obtains several watermarked audio data for calculating the triple loss corresponding to each watermarked audio data respectively.

[0059] Step S14: Extract the second latent vector representation of the watermarked audio data.

[0060] In the embodiment of the present application, the latent vector representation consistency optimization network extracts the second latent vector representation related to the watermarked audio data .

[0061] Step S15: Extract the third latent vector representation of other audio data.

[0062] In the embodiment of the present application, the latent vector representation consistency optimization network extracts the third latent vector representation related to other audio data .

[0063] Among them, the third latent vector representation can be the latent vector representation before adding watermark data to other audio data, or the latent vector representation after adding watermark data. And the other audio data can be any audio data in the audio dataset except the target audio data.

[0064] To sum up, in order to ensure that the audio samples before and after adding watermark are basically the same, the present application sends the audio samples before adding watermark and the audio samples after adding watermark into the ViT base network (Vision Transformer), and can extract their corresponding latent vector representations respectively and .

[0065] Step S16: Construct a triplet loss using the first latent vector representation, the second latent vector representation, and the third latent vector representation.

[0066] In the embodiment of the present application, after the training of the current batch, for each audio sample, the audio watermark generation device constructs a latent vector representation triplet ( , , ), where and respectively represent the latent vector representations obtained by extracting through the ViT-base model before and after adding watermark information to the current audio sample; represents the latent vector representation obtained before or after adding watermark to other audio samples in this training batch.

[0067] For each audio sample, calculate the triplet loss function of the representation : , where represents the cosine distance between two speaker models, is the cosine distance boundary (usually set to 0.2), , respectively represent the latent vector representations of the current audio sample and the latent vector representation of another randomly selected audio sample after the first addition of watermark information to the sample audio in the current training batch.

[0068] Further, as shown in the above step S136, in this application, multiple groups of watermark information are simultaneously added to the same training audio sample. Therefore, the audio watermark generation device can generate multiple watermarked audio data for the target audio data, and each watermarked audio data corresponds to a second latent vector representation .

[0069] Specifically, for the current audio data, the audio watermark generation device selects N - 1 watermark information encodings from the watermark codebook library, adds them to the audio data, respectively sends them into the ViT basic network to extract the corresponding latent vector representations, and calculates their corresponding triplet loss functions 、...、 .

[0070] Further, the audio watermark generation device can also calculate the mean square error between the STFT features of the audio sample before adding the watermark and the STFT features of the audio sample after adding the watermark: , where represents the STFT features of the audio obtained after adding the watermark information to the audio sample for the first time. Therefore, the audio watermark generation device also needs to calculate the mean square error loss of N - 1 STFT features 、...、 .

[0071] Step S17: Update the audio watermark generation network using the triplet loss, where the audio watermark embedding network is updated for generating audio watermarks after the update

[0072] In the embodiment of this application, the audio watermark generation device accumulates and averages all the triplet loss functions in step S16 to obtain the overall loss function of the latent vector representation consistency optimization network and the mean square error loss function of the audio watermark embedding network .

[0073] In this application, by simultaneously adding N watermark information embedding encodings to the above one sample and performing averaging processing on the loss function, the dependence on a single watermark information embedding encoding is weakened

[0074] In this application, an audio watermark generation device obtains an audio data set of the current batch, where the audio data set includes target audio data and other audio data; extracts a first latent vector representation of the target audio data; inputs the target audio data into an audio watermark embedding network of an audio watermark generation network to add watermark data to obtain watermarked audio data; extracts a second latent vector representation of the watermarked audio data; extracts a third latent vector representation of the other audio data; constructs a triplet loss by using the first latent vector representation, the second latent vector representation, and the third latent vector representation; and updates the audio watermark generation network by using the triplet loss, where the audio watermark embedding network is updated for generating audio watermarks after the update. Through the above audio watermark generation method, the audio features before and after adding watermarks are aligned in the latent space to update the network parameters of the audio watermark embedding network, thereby improving the audio watermark generation effect.

[0075] Furthermore, in the attack simulation of the existing system, multiple groups of attack methods are randomly mixed together to attack the training audio samples, and the effects generated between multiple groups of attack methods may cancel each other out. Therefore, this application adopts an alternating training method for attack methods. For one training data batch, the same attack method is used, and for the next training data batch, another attack method is randomly used. In this way, it can be ensured that in one update of network parameters, only one simulation attack is affected.

[0076] For details, please refer to Figure 4 , Figure 4 which is a schematic flowchart of the second embodiment of the audio watermark generation method provided by this application.

[0077] As Figure 4 shown, the specific steps are as follows: Step S21: Input the watermark data into the audio watermark extraction network of the audio watermark generation network to obtain a first watermark information code.

[0078] Step S22: Select a simulation attack method for the audio data set of the current batch.

[0079] Step S23: Perform an attack simulation on the watermarked audio data according to the simulation attack method to generate attack audio data.

[0080] In the embodiment of this application, the audio watermark generation device sends each audio sample after adding a watermark into the simulation attack simulation layer, randomly selects an attack method from the attack methods to perform an attack simulation on all audio samples in the training batch, and obtains the attacked audio samples with watermark information, that is, the attack audio data.

[0081] Step S24: Input the attack audio data into the audio watermark extraction network to extract the second watermark information encoding.

[0082] In the embodiment of the present application, the audio watermark generation device inputs the watermark data added during the audio watermark generation process and the attack audio data generated in step S23 into Figure 2 the audio watermark extraction network shown in the figure to extract the first watermark information encoding and the second watermark information encoding.

[0083] Specifically, the audio watermark generation device first sends the above audio data into a deep residual network with a network depth of 34, and then passes through an average pooling layer (Average Pooling), and finally realizes the extraction of the embedded watermark information.

[0084] Step S25: Use the first watermark information encoding and the second watermark information encoding to construct a differential loss.

[0085] In the embodiment of the present application, the audio watermark generation device measures the cosine distance between the extracted watermark information encoding and the embedded watermark information encoding to obtain a differential loss function between the two watermark information encodings:

[0086] where, is the embedded watermark information encoding, is the extracted watermark information encoding.

[0087] Step S26: Use the triplet loss and the differential loss to form a global loss.

[0088] Step S27: Update the audio watermark generation network using the global loss.

[0089] In the embodiment of the present application, the audio watermark generation device adds all the constructed loss functions to obtain a global loss function of the audio watermark generation network: . The audio watermark generation device completes the network gradient update of the current training stage according to the global loss function constructed from the audio data set of the current batch.

[0090] For the audio data set of the next batch and the training stage, the audio watermark generation device repeats the above steps. When selecting the simulation attack method in step S22, it is necessary to randomly switch to another attack method to simulate the attack on all audio samples in the training batch, so as to achieve the same attack simulation on the sample audio in the current training batch and different attack simulations on the sample audio in different training batches. Finally, the optimization of the audio watermark information embedding and extraction network is realized.

[0091] The audio watermark generation method of this application designs a potential vector representation consistency optimization network for aligning the audio before and after adding watermarks in the latent space.

[0092] The audio watermark generation method of this application designs to adopt the method of adding multiple groups of watermarks together. For the same training audio sample, multiple groups of random watermark information are added simultaneously to solve the dependence problem caused by adding a single watermark information to the current sample.

[0093] The audio watermark generation method of this application proposes to adopt the method of alternating training with attack methods. For one training data batch (batch), the same attack method is used, and for the next training data batch, another attack method is randomly adopted. Through this method of alternating training with a single attack method, the interference of the attack method is uniformly restricted in the current network space.

[0094] To implement the above audio watermark generation method, this application also proposes an audio watermark generation device. For details, please refer to Figure 5 , Figure 5 is the structural schematic diagram of an embodiment of the audio watermark generation device provided by this application.

[0095] The audio watermark generation device 500 of this embodiment includes: an acquisition module 51, an extraction module 52, a generation module 53, and an update module 54.

[0096] Among them, the acquisition module 51 is used to acquire the audio data set of the current batch, where the audio data set includes target audio data and other audio data.

[0097] The extraction module 52 is used to extract the first potential vector representation of the target audio data.

[0098] The generation module 53 is used to input the target audio data into the audio watermark embedding network of the audio watermark generation network to add watermark data to obtain watermarked audio data.

[0099] The extraction module 52 is used to extract the second potential vector representation of the watermarked audio data.

[0100] The extraction module 52 is used to extract the third potential vector representation of the other audio data.

[0101] The update module 54 is used to construct a triplet loss by using the first potential vector representation, the second potential vector representation, and the third potential vector representation.

[0102] The update module 54 is used to update the audio watermark generation network by using the triplet loss, where the audio watermark embedding network is updated to generate audio watermarks after the update.

[0103] To implement the above audio watermark generation method, the present application also proposes an audio watermark generation device. For details, please refer to Figure 6 , Figure 6 which is a schematic structural diagram of an embodiment of the audio watermark generation device provided by the present application.

[0104] The audio watermark generation device 400 in this embodiment includes a processor 41, a memory 42, an input / output device 43, and a bus 44.

[0105] The processor 41, the memory 42, and the input / output device 43 are respectively connected to the bus 44. Program data is stored in the memory 42, and the processor 41 is configured to execute the program data to implement the audio watermark generation method described in the above embodiment.

[0106] In the embodiment of the present application, the processor 41 may also be referred to as a CPU (Central Processing Unit). The processor 41 may be an integrated circuit chip with signal processing capabilities. The processor 41 may also be a general-purpose processor, a digital signal processor (DSP, Digital Signal Process), an application-specific integrated circuit (ASIC, Application Specific Integrated Circuit), a field-programmable gate array (FPGA, Field Programmable Gate Array), or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components. The general-purpose processor may be a microprocessor, or the processor 41 may also be any conventional processor, etc.

[0107] The present application also provides a computer storage medium. Please continue to refer to Figure 7 , Figure 7 which is a schematic structural diagram of an embodiment of the computer storage medium provided by the present application. A computer program 61 is stored in the computer storage medium 600. When the computer program 61 is executed by a processor, it is used to implement the audio watermark generation method described in the above embodiment.

[0108] When the embodiments of the present application are implemented in the form of software functional units and sold or used as independent products, they can be stored in a computer-readable storage medium. Based on this understanding, the technical solution of the present application, in essence, or the part that contributes to the prior art, or all or part of the technical solution, can be embodied in the form of a software product. This computer software product is stored in a storage medium and includes several instructions for causing a computer device (which may be a personal computer, a server, or a network device, etc.) or a processor to execute all or part of the steps of the methods described in various embodiments of the present application. The aforementioned storage medium includes: various media that can store program codes, such as USB flash drives, mobile hard disks, read-only memories (ROM, Read-Only Memory), random access memories (RAM, Random Access Memory), magnetic disks, or optical discs.

[0109] The above are only the embodiments of the present application, and do not limit the patent scope of the present application. Any equivalent structure or equivalent process transformation made by using the content of the specification and drawings of the present application, or directly or indirectly applied in other related technical fields, shall be equally included in the patent protection scope of the present application.

Claims

1. An audio watermark generation method, characterized in that The audio watermark generation method includes: Obtain the audio data set of the current batch, where the audio data set includes target audio data and other audio data; Extract the first latent vector representation of the target audio data; Input the target audio data into the audio watermark embedding network of the audio watermark generation network to add watermark data, and obtain watermarked audio data; Extract the second latent vector representation of the watermarked audio data; Extract the third latent vector representation of the other audio data; Construct a triplet loss using the first latent vector representation, the second latent vector representation, and the third latent vector representation; Update the audio watermark generation network using the triplet loss, where the audio watermark embedding network is updated to generate audio watermarks after the update.

2. The audio watermark generation method according to claim 1, characterized in that The audio watermark generation method further includes: Extract the first short-time Fourier transform feature of the target audio data; Extract the second short-time Fourier transform feature of the watermarked audio data; Construct an error loss using the first short-time Fourier transform feature and the second short-time Fourier transform feature; The updating the audio watermark generation network using the triplet loss includes: Update the audio watermark generation network using the triplet loss and the error loss.

3. The audio watermark generation method according to claim 1 or 2, characterized in that The inputting the target audio data into the audio watermark embedding network of the audio watermark generation network to add watermark data and obtaining watermarked audio data includes: Input the target audio data into the audio watermark embedding network to add several different watermark data respectively, and obtain several watermarked audio data; The constructing a triplet loss using the first latent vector representation, the second latent vector representation, and the third latent vector representation includes: Construct intermediate triplet losses respectively using the second latent vector representation of each watermarked audio data, the first latent vector representation, and the third latent vector representation; Construct the triplet loss using the intermediate triplet losses corresponding to the several watermarked audio data.

4. The audio watermark generation method according to claim 1, characterized in that The extracting the third latent vector representation of the other audio data includes: Extract the third latent vector representation of the other audio data before adding watermark data, or the third latent vector representation of the other audio data after adding watermark data.

5. The audio watermark generation method according to claim 1, characterized in that The audio watermark generation method further includes: Input the watermark data into the audio watermark extraction network of the audio watermark generation network to obtain the first watermark information encoding; Select the simulation attack method for the audio data set of the current batch; Perform attack simulation on the watermarked audio data according to the simulation attack method to generate attack audio data; Input the attack audio data into the audio watermark extraction network to extract the second watermark information encoding; Construct a difference loss using the first watermark information encoding and the second watermark information encoding; Updating the audio watermark generation network by using the triplet loss includes: Combining the triplet loss and the difference loss to form a global loss; Updating the audio watermark generation network by using the global loss.

6. The audio watermark generation method according to claim 5, wherein The simulation attack modes of the audio data sets of other batches are different from those of the audio data set of the current batch.

7. The audio watermark generation method according to claim 1, wherein Adding watermark data to the audio watermark embedding network of the audio watermark generation network by inputting the target audio data to obtain watermarked audio data, includes: Inputting the target audio data into the audio watermark embedding network to extract the amplitude spectrum feature and phase spectrum feature of the target audio data; Selecting one of the watermark information encodings from the watermark codebook library, and performing frame-level replication according to the number of frames of the amplitude spectrum feature to obtain watermark data; Performing frame-level splicing on the amplitude spectrum feature and the watermark data to obtain a watermarked amplitude spectrum feature; Extracting a watermark modulation weight matrix based on the watermarked amplitude spectrum feature; Fusing the amplitude spectrum feature and the watermark modulation weight matrix to obtain a modulated amplitude spectrum feature; Combining and performing inverse Fourier transform on the modulated amplitude spectrum feature and the phase spectrum feature to obtain the watermarked audio data.

8. An audio watermark generation device, characterized in that, The audio watermark generation device includes: an acquisition module, an extraction module, a generation module, and an update module; wherein, The acquisition module is configured to acquire the audio data set of the current batch, wherein the audio data set includes target audio data and other audio data; The extraction module is configured to extract the first latent vector representation of the target audio data; The generation module is configured to add watermark data to the audio watermark embedding network of the audio watermark generation network by inputting the target audio data to obtain watermarked audio data; The extraction module is configured to extract the second latent vector representation of the watermarked audio data; The extraction module is configured to extract the third latent vector representation of the other audio data; The update module is configured to construct a triplet loss by using the first latent vector representation, the second latent vector representation, and the third latent vector representation; The update module is configured to update the audio watermark generation network by using the triplet loss, wherein the audio watermark embedding network is updated to generate audio watermarks after the update.

9. An audio watermark generation device, characterized in that, The audio watermark generation device includes a memory and a processor coupled to the memory; Wherein, the memory is used to store program data, and the processor is used to execute the program data to implement the audio watermark generation method according to any one of claims 1 to 7.

10. A computer storage medium, characterized in that, The computer storage medium is used to store program data, and when the program data is executed by a computer, it is used to implement the audio watermark generation method according to any one of claims 1 to 7.

Citation Information

Patent Citations

  • Audio watermark adding and detecting methods and devices

    CN105741845A

  • Detection method and device and road side unit

    CN112861570A

  • Zero sample learning method based on combination of aligned variational auto-encoder and triple

    CN114022739A

  • Communication signal modulation mode open set identification method and system based on deep learning

    CN114567528A

  • Media data de-duplication method and target model training method and device

    CN115292541A

Cited By

  • Image watermark generation method, image watermark generation device and computer storage medium

    CN120410831A