Audio watermark generation method and device and computer storage medium

By designing the frame selection logic and auditory masking effect post-processing matrix in the audio watermark generation method, the problems of audio watermark redundancy and hearing difference in the prior art are solved, and more efficient and lossless audio watermark generation is achieved.

CN120236594AInactive Publication Date: 2025-07-01IFLYTEK CO LTD

Patent Information

Application Number
CN202510707348.2
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-05-29
Publication Date
2025-07-01
Estimated Expiration
Not applicable · inactive patent

AI Technical Summary

Technical Problem

The existing neural network-based audio watermark embedding network not only ensures high detection rate, but also causes watermarks to be added to each frame of the audio, resulting in redundancy and hearing differences.

Method used

Design an audio watermark generation method, select the number of frames and frame numbers that need to be added to the watermark through frame selection logic, combine the auditory masking effect post-processing matrix, adjust the number of frames and embedding weights added to the watermark to reduce the redundancy of watermark information and reduce the difference in hearing.

Benefits of technology

By reducing the redundant information of audio watermarks, the difference in audio listening experience between the watermark before and after adding the watermark is reduced, and the effectiveness and hearing losslessness of the audio watermark are improved.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120236594A_ABST
    Figure CN120236594A_ABST
Patent Text Reader

Abstract

The invention provides an audio watermark generation method, audio watermark generation equipment and a computer storage medium. The audio watermark generation method comprises the following steps: acquiring target audio data; extracting amplitude spectrum features and phase spectrum features of the target audio data; acquiring watermark adding frame information of the amplitude spectrum features; selecting one watermark information code from a watermark codebook library, and performing frame-level copying according to the frame number of the amplitude spectrum characteristics to obtain first watermark data; performing frame selection on the first watermark data according to the watermark adding frame information to obtain second watermark data; fusing the amplitude spectrum feature with the second watermark data to obtain a watermark amplitude spectrum feature; and fusing the watermark amplitude spectrum feature and the phase spectrum feature to obtain watermark audio data. According to the audio watermark generation method, the watermark adding selector is designed, the frame number and the frame number needing to be added with the watermark are selected through the frame selection logic, and it is guaranteed that after the audio watermark is generated, the audio listening feeling is not affected.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present application relates to the field of digital watermarking technology, and particularly to an audio watermark generation method, an audio watermark generation device, and a computer storage medium. Background Art

[0002] For the audio watermark embedding and extraction method based on neural network, first, the original audio signal is subjected to short-time Fourier transform (STFT, Short-Time Fourier Transform) to separate the amplitude spectrum and phase spectrum features. By splicing the watermark information coding and the amplitude spectrum features in the feature dimension, a joint feature tensor is constructed as the input of the convolutional neural network, and then the watermark embedding weight matrix corresponding to each frequency point of the amplitude spectrum is learned. After multiplying the weight matrix element by element with the original amplitude spectrum and combining it with the original phase spectrum, the watermark-embedded audio is reconstructed through inverse short-time Fourier transform (iSTFT).

[0003] The existing watermark embedding network based on neural network needs to prioritize ensuring the detection rate of the watermark. Therefore, watermarks are added to each frame of the audio. Excessive watermark information causes a certain redundancy and increases the perceptual difference of the audio before and after adding. Summary of the Invention

[0004] To solve the above technical problems, the present application proposes an audio watermark generation method, an audio watermark generation device, and a computer storage medium.

[0005] To solve the above technical problems, the present application proposes an audio watermark generation method, and the audio watermark generation method includes: Obtain target audio data; Input the target audio data into the audio watermark embedding network of the audio watermark generation network to extract the amplitude spectrum feature and phase spectrum feature of the target audio data; Obtain the watermark-adding frame information of the amplitude spectrum feature, where the watermark-adding frame information is used to indicate the audio frames to which watermarks are added; Select one of the watermark information codings from the watermark codebook library, and perform frame-level replication according to the number of frames of the amplitude spectrum feature to obtain first watermark data; Perform frame selection on the first watermark data according to the watermark-adding frame information to obtain second watermark data; Fuse the amplitude spectrum feature and the second watermark data to obtain a watermark amplitude spectrum feature; Fuse the watermark amplitude spectrum feature and the phase spectrum feature to obtain watermark audio data.

[0006] Wherein, the obtaining the watermark-adding frame information of the amplitude spectrum feature includes: Input the amplitude spectrum feature into the frame selection module to determine the watermark-added frame information of the amplitude spectrum feature; Among them, the frame selection module includes a neural network.

[0007] Among them, the watermark-added frame information includes an auditory masking effect post-processing matrix; among them, the parameters corresponding to the audio frames with watermark added in the auditory masking effect post-processing matrix are set to a first preset value, and the parameters corresponding to the audio frames without watermark added are set to a second preset value.

[0008] Among them, obtaining the watermark-added frame information of the amplitude spectrum feature includes: Determine the high-energy audio frames of the target audio data based on the amplitude spectrum feature; Determine the high-energy audio frames as the audio frames with watermark added and set the corresponding frame weight to a first preset value; Determine the remaining audio frames other than the audio frames determined as the audio frames with watermark added in the target audio data as the audio frames without watermark added and set the corresponding frame weight to a second preset value.

[0009] Among them, after determining the high-energy audio frames as the audio frames with watermark added and setting the corresponding frame weight to a first preset value, the audio watermark generation method further includes: Select the adjacent audio frames of the high-energy audio frames as the audio frames with watermark added and set the corresponding frame weight to a first preset value.

[0010] Among them, after fusing the amplitude spectrum feature with the second watermark data to obtain the watermark amplitude spectrum feature, the audio watermark generation method further includes: Extract the watermark modulation weight matrix based on the watermark amplitude spectrum feature; Modulate the watermark amplitude spectrum feature with the watermark modulation weight matrix to obtain the modulated amplitude spectrum feature.

[0011] Among them, after fusing the watermark amplitude spectrum feature with the phase spectrum feature to obtain the watermark audio data, the audio watermark generation method further includes: Input the watermark audio data into the audio quality evaluation module to obtain the audio quality score; Determine the audio quality loss using the audio quality score; Update the audio watermark generation network using the audio quality loss; among them, after the audio watermark embedding network is updated, it is used to generate audio watermarks.

[0012] Among them, after fusing the watermark amplitude spectrum feature with the phase spectrum feature to obtain the watermark audio data, the audio watermark generation method further includes: Input the first watermark data and / or the second watermark data into the audio watermark extraction network of the audio watermark generation network to obtain the first watermark information encoding; Select the simulation attack method for the target audio data; Perform attack simulation on the watermarked audio data according to the simulation attack method to generate attack audio data; Input the attack audio data into the audio watermark extraction network to extract the second watermark information encoding; Construct a differential loss by using the first watermark information encoding and the second watermark information encoding; The updating of the audio watermark generation network by using the audio quality loss includes: Update the audio watermark generation network by using the differential loss and the audio quality loss.

[0013] To solve the above technical problems, the present application also proposes an audio watermark generation device, which includes a memory and a processor coupled to the memory; wherein, the memory is used to store program data, and the processor is used to execute the program data to implement the audio watermark generation method as described above.

[0014] To solve the above technical problems, the present application also proposes a computer storage medium, which is used to store program data, and when the program data is executed by a computer, it is used to implement the above audio watermark generation method.

[0015] Compared with the prior art, the beneficial effect of the present application is that: the audio watermark generation device obtains target audio data; inputs the target audio data into the audio watermark embedding network of the audio watermark generation network to extract the amplitude spectrum feature and phase spectrum feature of the target audio data; obtains the watermark-added frame information of the amplitude spectrum feature, where the watermark-added frame information is used to indicate the audio frame to which the watermark is added; selects one of the watermark information encodings from the watermark codebook library, performs frame-level replication according to the number of frames of the amplitude spectrum feature to obtain the first watermark data; performs frame selection on the first watermark data according to the watermark-added frame information to obtain the second watermark data; fuses the amplitude spectrum feature with the second watermark data to obtain the watermarked amplitude spectrum feature; fuses the watermarked amplitude spectrum feature with the phase spectrum feature to obtain the watermarked audio data. Through the above audio watermark generation method, a watermark addition selector is designed, and the number of frames and frame numbers that need to add watermarks are selected through the frame selection logic, ensuring that the audio watermark has no impact on the audio listening feeling after generation. Description of the Drawings

[0016] To more clearly illustrate the technical solutions in the embodiments of the present application, the following will briefly introduce the accompanying drawings required for the description of the embodiments. Obviously, the accompanying drawings in the following description are only some embodiments of the present application. For those of ordinary skill in the art, without creative efforts, other accompanying drawings can be obtained based on these drawings. Among them: Figure 1 FIG. is a schematic flowchart of an embodiment of the audio watermark generation method provided by the present application; Figure 2 FIG. is a schematic architecture diagram of the audio watermark generation architecture provided by the present application; Figure 3 FIG. is a schematic flowchart of a second embodiment of the audio watermark generation method provided by the present application; Figure 4 FIG. is a schematic structural diagram of an embodiment of the audio watermark generation device provided by the present application; Figure 5 FIG. is a schematic structural diagram of an embodiment of the computer storage medium provided by the present application. Specific Embodiments

[0017] The following will clearly and completely describe the technical solutions in the embodiments of the present application in conjunction with the accompanying drawings in the embodiments of the present application. Obviously, the described embodiments are only a part of the embodiments of the present application, rather than all the embodiments. Based on the embodiments in the present application, all other embodiments obtained by those of ordinary skill in the art without creative efforts fall within the scope of protection of the present application.

[0018] The terms "first", "second", "third", "fourth", etc. (if any) in the specification and claims of the present application and the above accompanying drawings are used to distinguish similar objects, and do not necessarily describe a specific order or sequence. It should be understood that such data can be interchanged under appropriate circumstances, so that the embodiments of the present application described here, for example, can be implemented in an order other than those illustrated or described here. In addition, the terms "include" and "have" and any variations thereof are intended to cover non-exclusive inclusion. For example, a process, method, system, product, or device that includes a series of steps or units does not necessarily have to be limited to those steps or units clearly listed, but may include other steps or units not clearly listed or inherent to these processes, methods, products, or devices.

[0019] In today's digital age, with the rapid development of multimedia technology and the Internet, the creation, dissemination, and sharing of audio content have become unprecedentedly convenient. Online music platforms make a vast amount of music works easily accessible, and audio forms such as Internet radio and audiobooks are also enriching people's lives day by day. However, this convenience has also brought many serious problems. The illegal copying, dissemination, and tampering of digital audio products are becoming increasingly rampant, seriously damaging the legitimate rights and interests of creators, distributors, and copyright owners. Taking the music industry as an example, the widespread circulation of pirated music has led to a significant reduction in the income of creators, dampened their creative enthusiasm, and hindered the healthy development of the music industry. Against this background, audio watermarking technology, as an effective means of copyright protection and content authentication, has emerged and received extensive attention.

[0020] Audio watermarking technology utilizes the redundancy of audio signals and the characteristics of the human auditory system to secretly embed watermark information into audio data without affecting the original auditory quality of the audio. The embedded watermark has a certain degree of robustness and can resist common audio processing operations such as compression, filtering, resampling, etc. At the same time, it also has detectability, and the watermark information can be extracted from the audio through specific algorithms to verify the copyright ownership, content integrity of the audio, or perform other relevant authentication operations.

[0021] However, with the booming development of AIGC (Artificial Intelligence Generated Content) technology, audio watermarking technology is facing many serious challenges. With its powerful generation ability, AIGC technology can quickly and efficiently synthesize various types of audio content, which not only enriches audio resources but also greatly disrupts the ecological order of audio content, posing an impact on the application and function of traditional audio watermarking technology. Especially at present, with the application of deep learning in audio watermarking systems, audio watermarks can still be effectively detected under various attack situations such as editing, encoding / decoding, and rerecording. On the other hand, the way of adding watermarks by deep learning often has a certain impact on the listening perception of the audio while ensuring a high detection rate.

[0022] Existing watermark embedding networks based on neural networks need to prioritize ensuring the detection rate of watermarks, so watermarks are added to each frame of the audio. Excessive watermark information causes a certain degree of redundancy and increases the listening perception difference of the audio before and after adding.

[0023] Therefore, the audio watermark generation method provided in this application designs a watermark addition selector to determine which audio frames in the audio data need to add watermarks, thereby reducing the watermark information that needs to be added to the audio and minimizing the listening perception difference of the audio before and after adding watermarks.

[0024] For details, please refer to Figure 1 andFigure 2 , Figure 1 is a schematic flowchart of an embodiment of the audio watermark generation method provided by this application. Figure 2 is a schematic architecture diagram of the audio watermark generation architecture provided by this application.

[0025] The audio watermark generation method of this application is applied to an audio watermark generation device. Among them, the audio watermark generation device of this application can be a server, a terminal device, or a system in which the server and the terminal device cooperate with each other. Correspondingly, each part included in the audio watermark generation device, such as each unit, subunit, module, and submodule, can be all set in the server, all set in the terminal device, or respectively set in the server and the terminal device.

[0026] Furthermore, the above-mentioned server can be hardware or software. When the server is hardware, it can be implemented as a distributed server cluster composed of multiple servers or as a single server. When the server is software, it can be implemented as multiple software or software modules, such as software or software modules used to provide a distributed server, or as a single software or software module, which is not specifically limited here.

[0027] As Figure 1 shown, the specific steps are as follows: Step S11: Obtain target audio data.

[0028] In the embodiment of this application, the target audio data can be used on the one hand to input Figure 2 the audio watermark embedding network shown to add an audio watermark, and on the other hand, it can also be used to Figure 2 update the network parameters of the audio watermark embedding network shown and the audio watermark generation architecture where it is located.

[0029] Step S12: Input the target audio data into the audio watermark embedding network of the audio watermark generation network to extract the amplitude spectrum feature and phase spectrum feature of the target audio data.

[0030] In the embodiment of this application, after the audio watermark embedding network performs windowing and short-time Fourier transform on the audio data, the amplitude spectrum feature and phase spectrum feature of the STFT (short-time Fourier transform) are obtained, and their feature sizes are both , where represents the number of feature frames, represents the feature dimension.

[0031] Step S13: Obtain the watermark-adding frame information of the amplitude spectrum feature, where the watermark-adding frame information is used to indicate the audio frame to which the watermark is added.

[0032] In the embodiments of the present application, as Figure 2 shown, the audio watermark generation device inputs the amplitude spectrum features of the target audio data into the frame selection module, which is a watermark addition selector, and its function is to determine which audio frames need to add watermark information.

[0033] In a specific embodiment, Figure 2 the frame selection module shown may be in the form of a neural network or a frame selection model constructed based on this. Specifically, the neural network selected by the frame selection module includes, but is not limited to, CNN (Convolutional Neural Network), RNN (Recurrent Neural Network), LSTM (Long Short-Term Memory), Transformer (self-attention model), etc.

[0034] Among them, the output result of the frame selection module can be the number of frames or the frame numbers in the target audio data that need to add watermark information. When the output of the frame selection module is the frame number, the audio watermark embedding network can directly add watermark information according to the above frame number; when the output of the frame selection module is the number of frames, the audio watermark embedding network can randomly select the corresponding number of audio frames of the target audio data to add watermark information, and can also introduce other networks or models to complete the selection of the audio frames that need to add watermark information corresponding to the number of frames.

[0035] In another specific embodiment, since the current audio watermark addition technology does not consider the human auditory masking effect, for example, although some watermarks cause spectral differences, if added in the auditory masking effect area, it will basically have no impact on the human auditory perception. Therefore, a post-processing method based on the human auditory masking effect is proposed to adjust the number of frames for watermark addition and the embedding weight, so that the added watermark is basically the same as before in terms of auditory perception.

[0036] Specifically, the audio watermark embedding network inputs the amplitude spectrum features of the target audio data into the auditory masking effect post-processing module to implement the frame selection logic of the audio frames that need to add watermarks. Among them, the human auditory masking effect mainly includes two parts. First, for the same moment, the audio with higher energy can mask the audio with lower energy; second, for different moments, the audio with higher energy can mask the audio with lower energy in the short time before and after.

[0037] According to the above principle, the auditory masking effect post-processing module can set the following rules to execute the frame selection logic: Rule 1: For the amplitude spectra of all frames of the original audio, select the first 40% (the threshold here can be adjusted according to the scenario) of the amplitude spectra as the number of frames allowed for watermark addition, and set the frame weights of all these frames to 1; set the frame weights of the remaining frames to 0, indicating that watermark addition is not allowed.

[0038] Rule 2: For the audio frames selected in Rule 1, select the adjacent frames before and after them as the number of frames allowed for watermark addition, set the frame weights of all these frames to 1, and set the frame weights of the remaining frames to 0, indicating that watermark addition is not allowed.

[0039] It should be noted that the auditory masking effect post-processing module can determine the auditory masking effect post-processing matrix according to the above Rule 1 and Rule 2, or can also determine the auditory masking effect post-processing matrix only according to the above Rule 1.

[0040] In yet another specific implementation, the audio watermark embedding network can further input the output result of the frame selection module into the auditory masking effect post-processing module to achieve modulation of the frame selection result.

[0041] Specifically, the auditory masking effect post-processing module can set the frame weights corresponding to the first 40% of the audio frames with the highest energy among the amplitude spectra selected by the frame selection module to 1, and set the frame weights of the remaining frames to 0. Among them, the energy level of the amplitude spectrum is mainly reflected in the amplitude value of the amplitude spectrum.

[0042] Step S14: Select one of the watermark information encodings from the watermark codebook library, and perform frame-level replication according to the number of frames of the amplitude spectrum feature to obtain the first watermark data.

[0043] In the embodiment of the present application, the audio watermark embedding network randomly selects a watermark information encoding from the watermark codebook library and performs frame-level replication according to the number of frames of the amplitude spectrum feature. The size of the replicated watermark encoding is , where represents the number of feature frames, represents the number of bits of the watermark information encoding. Then, the watermark information encoding is transformed to dimensions through the projection layer, which is consistent with the audio features of the target audio data.

[0044] Step S15: Perform frame selection on the first watermark data according to the watermark addition frame information to obtain the second watermark data.

[0045] In the embodiment of the present application, since the number of frames of the watermark information encoding is the same as the number of frames of the target audio data after frame-level replication of the watermark information encoding, at this time, the watermark addition frame information determined in step S13 can directly act on the watermark information encoding, that is, perform frame screening on the watermark information encoding to be added to the audio data.

[0046] In this application, by adding watermark frame information to screen the watermark information encoding, the watermark information finally added to the target audio data is reduced, the redundancy of the watermark information is reduced, and the perceptual difference of the audio before and after adding the watermark is minimized as much as possible.

[0047] Specifically, the added watermark frame information in this application can be represented in a matrix form. The number of matrix parameters is the same as the number of frames of the target audio data. When the matrix parameter value is 1, it means that the weight of the corresponding frame is 1, indicating that watermark information needs to be added; when the matrix parameter value is 0, it means that the weight of the corresponding frame is 0, indicating that watermark information does not need to be added. Therefore, the method of screening the watermark information encoding can be to perform a dot product of the watermark information encoding and the weight matrix of the added watermark frame information to obtain the watermark information to be added to the target audio data.

[0048] Step S16: Fuse the amplitude spectrum feature with the second watermark data to obtain the watermark amplitude spectrum feature.

[0049] In the embodiment of this application, the audio watermark embedding network fuses the amplitude spectrum feature with the second watermark data screened in step S15, so as to obtain the watermark amplitude spectrum feature after adding the watermark information.

[0050] Furthermore, the audio watermark embedding network can also use the watermark modulation weight matrix to adjust the watermark amplitude spectrum feature. Among them, the watermark modulation weight matrix is an adaptive mask in deep watermark embedding, generated by a neural network, mainly used for slightly adjusting the amplitude spectrum to maintain perceptual lossless, counteracting signal processing attacks, and dynamically optimizing the embedding position and strength according to the audio content.

[0051] Compared with the traditional method, the watermark modulation weight matrix combined with deep learning can better balance inaudibility and robustness.

[0052] The watermark modulation weight matrix is a weight matrix with the same size as the original amplitude spectrum, and each of its elements represents: How to adjust the original amplitude spectrum to embed the watermark while minimizing the damage to the audio quality.

[0053] The value range is usually Among them, is the modulation intensity, such as 0.05, to ensure fine-tuning the amplitude rather than completely replacing it.

[0054] Finally, the audio watermark embedding network performs a dot product of the original amplitude spectrum feature and the watermark modulation weight matrix to obtain the amplitude spectrum feature modulated by the watermark information, that is, the modulated amplitude spectrum feature.

[0055] It should be noted that, as Figure 2As shown, the functions of the auditory masking effect post - processing module also include adjusting the embedding weights of the watermark modulation weight matrix so that the added watermark is basically indistinguishable in auditory perception from before it was added.

[0056] Step S17: Fuse the watermark amplitude spectrum feature and the phase spectrum feature to obtain watermarked audio data.

[0057] In the embodiment of the present application, the audio watermark embedding network combines the phase spectrum feature with the watermark amplitude spectrum feature generated in step S16, performs an inverse Fourier transform, and obtains the audio sample after adding the watermark information, that is, the watermarked audio data.

[0058] In the present application, the audio watermark generation device obtains target audio data; inputs the target audio data into the audio watermark embedding network of the audio watermark generation network to extract the amplitude spectrum feature and the phase spectrum feature of the target audio data; obtains the watermark - adding frame information of the amplitude spectrum feature, where the watermark - adding frame information is used to indicate the audio frames to which the watermark is added; selects one of the watermark information encodings from the watermark codebook library, performs frame - level replication according to the number of frames of the amplitude spectrum feature to obtain the first watermark data; performs frame selection on the first watermark data according to the watermark - adding frame information to obtain the second watermark data; fuses the amplitude spectrum feature and the second watermark data to obtain the watermark amplitude spectrum feature; fuses the watermark amplitude spectrum feature and the phase spectrum feature to obtain the watermarked audio data. Through the above - mentioned audio watermark generation method, a watermark - adding selector is designed, and the number of frames and frame numbers that need to add the watermark are selected through the frame - selection logic, ensuring that the audio watermark has no impact on the auditory perception of the audio after generation.

[0059] Furthermore, existing neural - network - based watermark embedding networks only rely on calculating the mean squared error between the STFT features of the original audio and the audio after adding the watermark to update the network parameters. Since the MSE (Mean Squared Error Loss) of the spectral features is used as the loss function, the impact of the differences in different spectra on the auditory perception is not considered. For example, the same difference has significantly different auditory perception effects when acting on different amplitude spectra.

[0060] In response to this, the present application trains an audio quality assessment model MosNET as a reward function, so that the audio after adding the watermark can obtain better audio quality, rather than treating all frame watermarks equally in order to pursue the minimum mean squared error loss.

[0061] Specifically, for the audio sample after adding the watermark, this application sends it into an audio quality evaluation module, namely the MosNET module. Among them, the MosNET module can adopt a neural network, including but not limited to CNN, RNN, LSTM, Transformer, etc. The MosNET module can score the audio quality fed into it. It is obtained through training with a large number of supervised audios. The higher the quality, the higher the output score. Its loss function is denoted as .

[0062] In addition to the mean square error loss obtained by calculating the mean square error between the STFT features of the original audio and the audio after adding the watermark, this application can also introduce the above-mentioned audio quality loss and jointly train and update Figure 2 the audio watermark generation network shown.

[0063] For details, please continue to refer to Figure 3 . Figure 3 which is the flowchart of the second embodiment of the audio watermark generation method provided by this application.

[0064] As Figure 3 shown, the specific steps are as follows: Step S21: Input the first watermark data and / or the second watermark data into the audio watermark extraction network of the audio watermark generation network to obtain the first watermark information encoding.

[0065] Step S22: Select the simulation attack method for the target audio data.

[0066] Step S23: Perform an attack simulation on the watermarked audio data according to the simulation attack method to generate attack audio data.

[0067] In the embodiment of this application, the audio watermark generation device sends the N audios after adding the watermark to each audio sample into the simulation attack simulation layer, randomly selects an attack method from the attack methods to perform an attack simulation on all the audio samples in the training batch, and obtains the watermarked audio samples after being attacked, that is, the attack audio data.

[0068] Among them, the N audios can be generated by adding different watermarks to each audio sample.

[0069] Step S24: Input the attack audio data into the audio watermark extraction network to extract the second watermark information encoding.

[0070] In the embodiment of this application, the audio watermark generation device inputs the watermark data added during the audio watermark generation process and the attack audio data generated in step S23 into Figure 2 the audio watermark extraction network shown to extract the first watermark information encoding and the second watermark information encoding.

[0071] Specifically, the audio watermark generation device first sends the above audio data into a deep residual network with a network depth of 34, and then passes through an average pooling layer (Average Pooling), and finally realizes the extraction of the embedded watermark information.

[0072] Step S25: Construct a differential loss using the first watermark information encoding and the second watermark information encoding.

[0073] In the embodiment of the present application, the audio watermark generation device performs cosine distance measurement on the extracted watermark information encoding and the embedded watermark information encoding to obtain a differential loss function between the two watermark information encodings:

[0074] Step S26: Update the audio watermark generation network using the differential loss and the audio quality loss.

[0075] In the embodiment of the present application, the audio watermark generation device adds up all the constructed loss functions to obtain a global loss function of the audio watermark generation network: . The audio watermark generation device completes the network gradient update of the current training stage according to the global loss function constructed based on the audio data set of the current batch.

[0076] Among them, the audio watermark generation device calculates the mean square error between the STFT features of the audio samples before adding the watermark and the audio samples after adding the watermark:

[0077] The audio watermark generation method provided by the present application first designs a watermark addition selector, and uses a neural network to judge which frames need to add watermarks; secondly, trains an audio quality evaluation model MosNET as a reward function, so that the audio after adding watermarks can obtain better audio quality, rather than treating all frame watermarks equally in order to pursue the minimum mean square error loss; finally, proposes to adopt a post-processing method based on the human auditory masking effect to adjust the number of frames with watermarks added and the embedding weights, so that the added watermark is basically indistinguishable from the one before addition in terms of auditory perception.

[0078] To implement the above audio watermark generation method, the present application also proposes an audio watermark generation device, specifically please refer to Figure 4 , Figure 4 is a schematic structural diagram of an embodiment of the audio watermark generation device provided by the present application.

[0079] The audio watermark generation device 400 in this embodiment includes a processor 41, a memory 42, an input / output device 43, and a bus 44.

[0080] The processor 41, the memory 42, and the input / output device 43 are respectively connected to the bus 44. Program data is stored in the memory 42, and the processor 41 is configured to execute the program data to implement the audio watermark generation method described in the above embodiments.

[0081] In the embodiments of the present application, the processor 41 may also be referred to as a CPU (Central Processing Unit). The processor 41 may be an integrated circuit chip with signal processing capabilities. The processor 41 may also be a general-purpose processor, a digital signal processor (DSP, Digital Signal Process), an application-specific integrated circuit (ASIC, Application Specific Integrated Circuit), a field-programmable gate array (FPGA, Field Programmable Gate Array), or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components. The general-purpose processor may be a microprocessor or the processor 41 may also be any conventional processor, etc.

[0082] The present application also provides a computer storage medium. Please continue to refer to Figure 5 , Figure 5 FIG. is a schematic structural diagram of an embodiment of the computer storage medium provided by the present application. A computer program 61 is stored in the computer storage medium 600. When the computer program 61 is executed by a processor, it is used to implement the audio watermark generation method described in the above embodiments.

[0083] When the embodiments of the present application are implemented in the form of software functional units and sold or used as independent products, they can be stored in a computer-readable storage medium. Based on such an understanding, the technical solution of the present application, in essence, or the part that contributes to the prior art, or all or part of the technical solution, can be embodied in the form of a software product. The computer software product is stored in a storage medium and includes several instructions for causing a computer device (which may be a personal computer, a server, or a network device, etc.) or a processor to execute all or part of the steps of the methods described in various embodiments of the present application. The foregoing storage medium includes: various media such as a USB flash drive, a mobile hard disk, a read-only memory (ROM, Read-Only Memory), a random access memory (RAM, Random Access Memory), a magnetic disk, or an optical disc that can store program codes.

[0084] The above are only the embodiments of the present application, and do not thus limit the patent scope of the present application. Any equivalent structure or equivalent process transformation made by using the content of the specification and drawings of the present application, or directly or indirectly applied in other related technical fields, shall be similarly included in the patent protection scope of the present application.

Claims

1. An audio watermark generation method, characterized in that, The audio watermark generation method includes: Obtain target audio data; Input the target audio data into the audio watermark embedding network of the audio watermark generation network, and extract the amplitude spectrum feature and phase spectrum feature of the target audio data; Obtain the watermark-adding frame information of the amplitude spectrum feature, where the watermark-adding frame information is used to indicate the audio frames to which watermarks are added; Select one of the watermark information encodings from the watermark codebook library, and perform frame-level replication according to the number of frames of the amplitude spectrum feature to obtain the first watermark data; Perform frame selection on the first watermark data according to the watermark-adding frame information to obtain the second watermark data; Fuse the amplitude spectrum feature and the second watermark data to obtain the watermark amplitude spectrum feature; Fuse the watermark amplitude spectrum feature and the phase spectrum feature to obtain the watermark audio data.

2. The audio watermark generation method according to claim 1, wherein the obtaining of the watermark-adding frame information of the amplitude spectrum feature includes: Input the amplitude spectrum feature into the frame selection module to determine the watermark-adding frame information of the amplitude spectrum feature; wherein, the frame selection module includes a neural network.

3. The audio watermark generation method according to claim 1 or 2, wherein the watermark-adding frame information includes an auditory masking effect post-processing matrix; wherein, the parameters corresponding to the audio frames with watermarks added in the auditory masking effect post-processing matrix are set to a first preset value, and the parameters corresponding to the audio frames without watermarks added are set to a second preset value.

4. The audio watermark generation method according to claim 3, wherein the obtaining of the watermark-adding frame information of the amplitude spectrum feature includes: Based on the amplitude spectrum feature, determine the high-energy audio frames of the target audio data; Determine the high-energy audio frames as the audio frames with watermarks added, and set the corresponding frame weight to a first preset value; Determine the remaining audio frames in the target audio data other than the audio frames determined to have watermarks added as the audio frames without watermarks added, and set the corresponding frame weight to a second preset value.

5. The audio watermark generation method according to claim 4, wherein after determining the high-energy audio frames as the audio frames with watermarks added and setting the corresponding frame weight to a first preset value, the audio watermark generation method further includes: Select the adjacent audio frames of the high-energy audio frames and determine them as the audio frames with watermarks added, and set the corresponding frame weight to a first preset value.

6. The audio watermark generation method according to claim 1, wherein after fusing the amplitude spectrum feature and the second watermark data to obtain the watermark amplitude spectrum feature, the audio watermark generation method further includes: Extract a watermark modulation weight matrix based on the watermark amplitude spectrum feature; Modulate the watermark amplitude spectrum feature with the watermark modulation weight matrix to obtain a modulated amplitude spectrum feature.

7. The audio watermark generation method according to claim 1, wherein after fusing the watermark amplitude spectrum feature and the phase spectrum feature to obtain the watermark audio data, the audio watermark generation method further includes: Input the watermarked audio data into an audio quality assessment module to obtain an audio quality score; Determine the audio quality loss using the audio quality score; Update the audio watermark generation network using the audio quality loss; wherein, after the update, the audio watermark embedding network is used to generate an audio watermark.

8. The audio watermark generation method according to claim 7, wherein After fusing the watermark amplitude spectrum feature and the phase spectrum feature to obtain watermarked audio data, the audio watermark generation method further includes: Input the first watermark data, and / or the second watermark data into the audio watermark extraction network of the audio watermark generation network to obtain a first watermark information code; Select a simulation attack mode for the target audio data; Perform an attack simulation on the watermarked audio data according to the simulation attack mode to generate attack audio data; Input the attack audio data into the audio watermark extraction network to extract a second watermark information code; Construct a difference loss using the first watermark information code and the second watermark information code; The updating the audio watermark generation network using the audio quality loss includes: Updating the audio watermark generation network using the difference loss and the audio quality loss.

9. An audio watermark generation device, characterized in that, The audio watermark generation device includes a memory and a processor coupled to the memory; wherein, the memory is used to store program data, and the processor is used to execute the program data to implement the audio watermark generation method according to any one of claims 1 to 8.

10. A computer storage medium, characterized in that, The computer storage medium is used to store program data, and when the program data is executed by a computer, it is used to implement the audio watermark generation method according to any one of claims 1 to 8.

Citation Information

Patent Citations

  • Watermark embedding and detecting method based on audio content classification

    CN104700841A

  • Method for realizing watermark encryption and decryption based on audio signal

    CN104851428A

  • DWT-SVD-ICA-based digital audio watermarking algorithm

    CN105741844A

  • Audio watermark embedding, extracting and television program interaction method and device

    CN109584890A

  • Echo delay determination method and device, equipment and storage medium

    CN113707160A

Cited By

  • Voice watermark encoding and decoding method, device, equipment and medium

    CN120708628A

  • A speech watermarking encoding and decoding method, apparatus, device and medium

    CN120708628B

  • Audio watermark embedding method and device, electronic equipment and storage medium

    CN121331148A

  • Audio watermark embedding method and device, electronic equipment and storage medium

    CN121331148B

  • Audio watermark model training method and device and audio watermark embedding method and device

    CN122177138A