High-precision dynamic watermarking method and device for streaming audio
By embedding sampling point-level watermark signals in streaming audio, and using neural networks to achieve high-precision dynamic watermarks, the problem of limited dynamic and flexibility of audio watermarks in the prior art is solved, and the carrying capacity and adaptability of watermark information are improved.
Patent Information
- Application Number
- CN202510984996.2
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-07-17
- Publication Date
- 2025-08-12
- Estimated Expiration
- 2045-07-17
Smart Images

Figure CN120472915A_ABST
Abstract
Description
Technical Field
[0001] The present application relates to the field of digital watermark technology, and in particular to a high-precision dynamic watermark method and device for streaming audio. Background Art
[0002] Digital audio content is experiencing groundbreaking growth, penetrating deeply into entertainment, telecommunications, and information distribution. In particular, breakthroughs in AI-generated content technology are enabling the industrialized production of high-fidelity audio, such as speech synthesis and music creation. This technological wave, while reshaping the content production paradigm, is also triggering profound changes in intellectual property rights, digital content traceability, and authenticity verification systems. Faced with an increasingly complex digital content ecosystem, building robust and intelligent security mechanisms has become a core issue in digital governance.
[0003] Existing technologies often rely on digital audio watermarking to enhance the security of audio data. This technology embeds imperceptible digital information within the audio signal. Current mainstream solutions typically use least significant bit (LSB) replacement, which permanently replaces the last one or two bits of the sampled value. This approach cannot dynamically adjust the embedding position or capacity, nor can it flexibly adjust the watermark content. This makes it difficult to meet the diverse and demanding requirements of complex application scenarios such as live streaming, real-time communication, and processing of long audio files.
[0004] There is currently no effective solution to the problem of limited dynamics and flexibility of audio watermarking in related technologies. Summary of the Invention
[0005] In this embodiment, a high-precision dynamic watermarking method and apparatus for streaming audio are provided to solve the problem of limited dynamism and flexibility of audio watermarking in related technologies.
[0006] In a first aspect, this embodiment provides a high-precision dynamic watermarking method for streaming audio, the method comprising:
[0007] Acquire a streaming audio input signal and a corresponding watermark information segment; the watermark information segment is composed of bit information segments with different states and variable lengths;
[0008] Determining a sampling point-level watermark signal according to the audio input signal and the watermark information segment;
[0009] Embedding the sampling point-level watermark signal into the audio input signal to obtain an audio output signal, and streaming the audio output signal;
[0010] Perform watermark detection on the sampling points of the received audio signal to be tested to obtain a watermark confidence vector for the sampling point; perform threshold judgment on the watermark confidence vector to restore the corresponding watermark information segment.
[0011] In some embodiments, determining a sampling point-level watermark signal according to the audio input signal and the watermark information segment includes:
[0012] Mapping the watermark information segment into a watermark feature representation that is point-by-point aligned with the audio input signal;
[0013] splicing the watermark feature representation and the representation of the audio input signal along the channel dimension to obtain a sampling point-level fusion feature representation;
[0014] The sampling point level fusion feature representation is input into a preset first neural network, and the disturbance information on each of the sampling points is calculated to obtain a sampling point level watermark signal.
[0015] In some embodiments, mapping the watermark information segment into a watermark feature representation that is point-by-point aligned with the audio input signal comprises:
[0016] Get the learnable embedding layer and the corresponding mapping relationship;
[0017] Based on the mapping relationship, extracting a set of target row vectors corresponding to each of the bit information segments in the watermark information segment from the learnable embedding layer; and obtaining a feature vector corresponding to the bit information segment based on the set of target row vectors;
[0018] The feature vectors are combined in time sequence into a watermark feature representation that is aligned point by point with the audio input signal.
[0019] In some embodiments, performing watermark detection on sampling points of a received audio signal to be tested to obtain a watermark confidence vector for the sampling points includes:
[0020] The received audio signal to be tested is input into a preset second neural network, and watermark detection is performed on the sampling points to obtain a multi-dimensional watermark confidence vector for the sampling points; the multi-dimensional watermark confidence vector includes watermark existence information and watermark content information.
[0021] In some embodiments, performing a threshold decision on the watermark confidence vector to recover the corresponding watermark information segment includes:
[0022] Smoothing the watermark confidence vector to obtain a smoothed result;
[0023] Based on a preset decision threshold, binarizing the smoothed result to obtain a binary decision sequence;
[0024] In a case where the watermark existence information indicates that a watermark exists, obtaining the bit information segment based on the corresponding binary determination sequence;
[0025] The watermark information segment is restored based on the bit information segment corresponding to each of the sampling points.
[0026] In some embodiments, the first neural network adopts an encoder-decoder structure, the first neural network adopts causal convolution operation, and has a state transfer mechanism across audio processing segments;
[0027] The second neural network includes an encoder component, which includes at least a convolutional layer and a channel attention mechanism layer, and the convolutional layer adopts a depth-separable convolution operation.
[0028] In some embodiments, the first neural network and the second neural network are obtained based on end-to-end joint training of an initial generation network and an initial detection network; the forward data propagation process of the end-to-end joint training includes:
[0029] Based on the audio test signal, a watermark test segment with a preset number of bits is randomly generated within a preset time length range;
[0030] Converting the watermark test segment into a test feature vector at the sampling point level, concatenating the test feature vector with the audio test signal and inputting the resultant vector into an initial generation network, outputting a one-dimensional test watermark signal of the same length as the audio test signal; and embedding the one-dimensional test watermark signal into the audio test signal to obtain a watermarked test signal.
[0031] Perform data enhancement processing on the watermarked test signal, input the enhanced watermarked test signal into an initial detection network, and output watermark detection confidence at a sampling point level, including watermark existence test information and watermark content test information, wherein the data enhancement processing includes: at least one of: channel distortion enhancement, watermark mask enhancement, and tampering simulation enhancement.
[0032] In some embodiments, the end-to-end joint training is optimized using a multi-task loss function; the multi-task loss function includes at least one of a perceptual loss function and a watermark loss function;
[0033] The perceptual loss function includes at least one of the following: a depth feature loss function, a loudness domain loss function, a frequency domain loss function, and a time domain loss function;
[0034] The watermark loss function includes at least one of the following: a watermark existence detection loss function, a multi-bit message detection loss function, and a boundary consistency loss function.
[0035] In some embodiments, the end-to-end joint training is optimized using a curriculum training strategy; the curriculum training strategy includes at least one of the following: dynamic adjustment of the weight of the perceptual loss function, and progressive increase in the number of bits of the watermark information segment.
[0036] In a second aspect, this embodiment provides a high-precision dynamic watermarking device for streaming audio, the device comprising:
[0037] A data acquisition module is used to acquire a streaming audio input signal and a corresponding watermark information segment; the watermark information segment is composed of bit information segments with different states and variable lengths;
[0038] a sampling point level watermark generation module, configured to determine a sampling point level watermark signal according to the audio input signal and the watermark information segment;
[0039] a watermark embedding module, configured to embed the sampling point-level watermark signal into the audio input signal to obtain an audio output signal, and stream output the audio output signal;
[0040] The watermark extraction module is used to perform watermark detection on the sampling points of the received audio signal to be tested, obtain the watermark confidence vector for the sampling point, perform threshold judgment on the watermark confidence vector, and restore the watermark information segment.
[0041] In a third aspect, the present application further provides a computer device comprising a memory and a processor, wherein the memory stores a computer program, and the processor implements the high-precision dynamic watermarking method for streaming audio according to the first aspect when executing the computer program.
[0042] In a fourth aspect, the present application further provides a computer-readable storage medium having a computer program stored thereon, which, when executed by a processor, implements the high-precision dynamic watermarking method for streaming audio described in the first aspect.
[0043] Compared with the related art, the high-precision dynamic watermarking method and device for streaming audio provided in this embodiment obtains a streaming audio input signal and a corresponding watermark information segment; the watermark information segment is composed of bit information segments with different states and variable lengths; a sampling point-level watermark signal is determined based on the audio input signal and the watermark information segment; the sampling point-level watermark signal is embedded in the audio input signal to obtain an audio output signal, and the audio output signal is streamed out; watermark detection is performed on the sampling points of the received audio signal to be tested to obtain a watermark confidence vector for the sampling point; a threshold judgment is performed on the watermark confidence vector to restore the corresponding watermark information segment, which solves the problem of poor dynamics and flexibility in watermark processing. Based on sampling point-level operations, the length and content of the watermark can be dynamically changed according to the sampling points of the audio input signal, supporting highly flexible dynamic watermark information embedding and greatly enhancing the watermark information carrying capacity. It is suitable for scenarios that require real-time updating of identification, tracking of dynamic states, or embedding of complex metadata.
[0044] The details of one or more embodiments of the present application are set forth in the following drawings and description to make other features, objects, and advantages of the present application more readily apparent. BRIEF DESCRIPTION OF THE DRAWINGS
[0045] The drawings described herein are used to provide a further understanding of the present application and constitute a part of the present application. The illustrative embodiments of the present application and their descriptions are used to explain the present application and do not constitute an improper limitation on the present application. In the drawings:
[0046] Figure 1 This is a hardware structure block diagram of the application environment of the high-precision dynamic watermarking method for streaming audio in an embodiment of the present application;
[0047] Figure 2 Schematic diagram of the process of a high-precision dynamic watermarking method for streaming audio in an embodiment of the present application;
[0048] Figure 3 This is a flow chart of a method for determining a sampling point-level watermark signal in an embodiment of the present application;
[0049] Figure 4 This is a flow chart of a method for recovering a watermark information segment based on post-processing in an embodiment of the present application;
[0050] Figure 5 This is a schematic diagram of the structure of the first neural network in the embodiment of the present application;
[0051] Figure 6 This is a schematic diagram of the structure of the second neural network in the embodiment of this application;
[0052] Figure 7Schematic diagram of the process of a high-precision dynamic watermark training method for streaming audio in an embodiment of the present application;
[0053] Figure 8 Schematic diagram of the flow of a high-precision dynamic watermarking method for streaming audio in a preferred embodiment of the present application;
[0054] Figure 9 This is a structural block diagram of a high-precision dynamic watermarking device for streaming audio in an embodiment of the present application.
[0055] Reference numerals: 102, terminal; 104, cloud server; 106, communication network; 108, other terminals; 91, data acquisition module; 92, sampling point-level watermark generation module; 93, watermark embedding module; 94, watermark extraction module. DETAILED DESCRIPTION
[0056] In order to more clearly understand the purpose, technical solutions and advantages of the present application, the present application is described and illustrated below in conjunction with the accompanying drawings and embodiments.
[0057] Unless otherwise defined, technical or scientific terms used in this application shall have the ordinary meanings as understood by persons of ordinary skill in the art to which this application belongs. The terms "a," "an," "the," "these," and similar expressions in this application do not denote limitations on quantity and may be singular or plural. The terms "comprise," "include," "have," and any variations thereof, as used in this application, are intended to cover non-exclusive inclusions. For example, a process, method, system, product, or device comprising a series of steps or modules (units) is not limited to the listed steps or modules (units) but may include unlisted steps or modules (units) or other steps or modules (units) inherent to the process, method, product, or device. The terms "connected," "connected," "coupled," and similar expressions used in this application are not limited to physical or mechanical connections but may include electrical connections, whether direct or indirect. As used in this application, "plurality" means two or more. "And / or" describes an association between associated objects, indicating that three possible relationships exist. For example, "A and / or B" can mean: A exists alone; A and B exist simultaneously; or B exists alone. Generally, the character " / " indicates that the objects in the preceding and following relationship are in an "or" relationship. The terms "first", "second", "third", etc. involved in this application are only used to distinguish similar objects and do not represent a specific ordering of the objects.
[0058] The high-precision dynamic watermarking method embodiment for streaming audio provided in this embodiment can be applied to a variety of application environments. For example, it can be applied to Figure 1The server-terminal application environment shown in FIG. This environment typically involves a terminal 102, a cloud server 104, and a communication network 106 connecting the two. Terminal 102 communicates with cloud server 104 via communication network 106. In this environment, for example, in applications such as copyright protection or content tracking, a content provider may first upload original audio (i.e., an audio input signal) to cloud server 104. Cloud server 104 is deployed with an embedding processing unit that implements the method of this embodiment. This embedding processing unit generates a dynamic watermark message containing a copyright identifier, content ID, and other information based on preset rules or specific requirements. The dynamic watermark message may have a time-varying state and / or variable length to improve tracking accuracy and security. Subsequently, the embedding processing unit on cloud server 104 invokes the high-precision embedding algorithm of this embodiment to embed this dynamic watermark message into the original audio with high precision (e.g., at or near the sampling point level), generating watermarked audio (i.e., the audio input signal). This watermarked audio may be stored in a data storage system associated with cloud server 104 and distributed to terminal 102 via communication network 106. When copyright verification or source tracing is required for audio on terminal 102, the audio to be tested or its features can be uploaded to cloud server 104. A detection processing unit deployed on cloud server 104 performs high-precision watermark detection and dynamic message extraction (e.g., sampling point-level detection and extraction) on the received data, and returns the verification results to the requesting party via communication network 106 or for platform management. Terminal 102 can be a variety of devices, such as personal computers, laptops, smartphones, tablets, smart home devices (e.g., smart speakers, smart TVs), in-vehicle infotainment systems, or wearable devices (e.g., smart watches).
[0059] Furthermore, the high-precision dynamic watermarking method for streaming audio provided in this embodiment can also be applied in end-to-end application scenarios. This scenario typically involves a sending end (end 102) and one or more receiving ends (other ends 108) communicating via a communications network 106. For example, in real-time voice calls, secure instant messaging, or online collaboration applications, when sending an audio stream, end 102 can utilize its locally deployed embedding processing unit implementing the method of this embodiment. This processing unit can embed a dynamic watermark message containing specific information (such as the sender's identity authentication fingerprint, session synchronization signal, message integrity check code, or low-rate control instructions) into the transmitted audio stream in real time with high precision (e.g., sample point accuracy). The watermarked audio stream is then transmitted via the communications network 106 to the other end 108. The receiving device uses its locally deployed detection processing unit implementing the method of this embodiment to detect the presence of the watermark from the received audio stream and extract the embedded dynamic watermark information. The extracted information can be used by applications on the receiving device to perform corresponding operations, such as performing authentication, ensuring that data has not been tampered with, synchronizing application status, or responding to control instructions. It is understood that in terminal-to-terminal scenarios, the communication involved can also be one-to-many, such as in live broadcast scenarios, or many-to-many, such as in video conferencing scenarios.
[0060] In this embodiment, a high-precision dynamic watermarking method for streaming audio is provided. Figure 2 FIG. 1 is a flow chart of a high-precision dynamic watermarking method for streaming audio according to an embodiment of the present invention. Figure 2 As shown, the process includes the following steps:
[0061] Step S210: obtaining a streamed audio input signal and a corresponding watermark information segment; the watermark information segment is composed of bit information segments with different states and variable lengths.
[0062] Specifically, streaming audio refers to a technology that transmits audio content to a receiving end in real time through a network or other digital transmission media. Unlike traditional download-and-play, streaming audio allows users to use specific transmission protocols and encoding methods in audio files to ensure that audio data can be continuously and stably transmitted to the audience. This implementation receives streaming audio input and obtains a corresponding audio input signal. The audio input signal can be pre-processed data, and the pre-processing means include but are not limited to normalization, noise reduction and other pre-processing operations, thereby improving the quality of the audio signal.
[0063] This embodiment supports the embedding of dynamic watermarks, which means that the embedded watermark information is not a static sequence that remains unchanged, but rather dynamic information that can be flexibly set. The dynamism is reflected in two aspects: first, the "state" of the watermark can change, that is, the specific bit content of the bit information segment at each time point can change, where the capacity of the bit information segment is B (B is an integer greater than or equal to 1), which determines that B bits of information can be embedded at a certain moment; second, the "length" of each state can change, that is, the time span of each state or the number of sample points included can change, where the length change range (Lmin, Lmax) can be set to improve the reliability of the change. The specific "state" content and "length" change rules are defined and provided by the upper-layer application based on its business logic and other information, making the application of watermarks extremely flexible and able to adapt to a wide range of scenarios that require dynamic information marking. The following are some typical situations:
[0064] The first type is human-triggered or externally-triggered switching: In this case, the watermark state change is triggered by a manual user switch command or a switch command sent by an external control system. For example, in a live broadcast scenario, the host can manually switch the embedded watermark content to identify different live broadcast segments or promotional information; or the broadcast monitoring system can instruct the embedding of different channel identifiers or time codes.
[0065] The second type involves automatic switching based on context or events: changes in the watermark state are tied to the current business scenario or a specific event. For example, in an e-commerce livestream, the embedded watermark can be associated with the link or ID of the currently displayed product. When the host switches between products, the corresponding watermark information is automatically updated. Another example is a multi-speaker video conference or voice chat scenario where the embedded watermark can be automatically switched based on the microphone currently speaking (via voice activity detection (VAD) or sound source localization), using the watermark to identify the current speaker's ID.
[0066] The third type is automatic generation and labeling based on content or time: in this case, the watermark state is automatically generated and updated by the system based on the audio content itself or time information. For example, a B-bit timestamp representing the current precise time (such as year, month, day, hour, minute, and second) is periodically generated and embedded as the watermark state for subsequent accurate tracing. For another example, an automatic speaker recognition function is integrated, and when different speakers are identified, a unique B-bit identifier representing the speaker's identity is automatically embedded. For another example, some form of hash value or acoustic fingerprint can be calculated for a fragment of audio content, and its B-bit representation can be embedded as a watermark for fine-grained content integrity verification or identification.
[0067] It should be noted that the above typical cases are used to explain different watermark embedding scenarios and are not used to limit them.
[0068] Step S220: Determine a sampling point-level watermark signal according to the audio input signal and the watermark information segment.
[0069] Specifically, the watermark information segment is processed in conjunction with the audio input signal to generate a sampling-point-level watermark signal that is compatible with the audio input signal at the sampling point precision. This minimizes interference with the audio input signal and is imperceptible to the human ear. In some implementations, the process of determining the sampling-point-level watermark signal can be implemented using a neural network. The neural network can automatically learn the content of the audio input signal and the watermark information segment at the sampling point level to generate a sample-level watermark signal that meets the requirements, greatly reducing the difficulty of designing the watermarking method. In other embodiments, this step can also be achieved through time-frequency analysis methods, such as performing time-frequency analysis on the audio input signal to extract its spectral energy distribution and masking threshold and other characteristics, and then dynamically adjusting the amplitude and frequency band position of the watermark information segment based on the psychoacoustic model - enhancing the watermark strength in the high-energy frequency band and reducing the watermark amplitude in the low-energy frequency band to maintain concealment; at the same time, a pseudo-random sequence can be generated in combination with the audio content (such as using the audio spectrum hash value as the seed), and the watermark information segment can be mapped into a broadband noise signal related to the statistical characteristics of the audio through spread spectrum modulation, or the intermediate frequency DCT coefficients or wavelet subbands can be adaptively selected in the transform domain, and the quantization step size (QIM) can be designed according to the local energy difference of the host audio, so that the quantization error of the watermark signal matches the characteristics of the audio signal to achieve adaptability of the watermark signal to the host audio, taking into account both robustness and imperceptibility.
[0070] Step S230: embed the sampling point-level watermark signal into the audio input signal to obtain an audio output signal, and stream the audio output signal.
[0071] Specifically, the sampling point level watermark signal is superimposed on the audio input signal in the form of direct addition of time domain signals. Specifically, for each sample point n in the audio input signal, the corresponding generated sampling point level watermark signal δ[n] is added to the original audio input signal value x[n]. In order to more accurately control the embedding strength to balance robustness and audio fidelity, an intensity control factor α can be used to scale the watermark signal before addition. Therefore, the audio output signal x is calculated. wm [n] A more complete operation can be expressed as: x wm [n] = x[n] + αδ[n]. A larger α enhances the robustness of the watermark but may introduce audible distortion; a smaller α helps maintain high fidelity but may reduce the watermark's ability to survive noise or compression. The α value can be designed as a fixed value or dynamically adjusted based on the specific application. Because this superposition process is performed sample by sample, it ensures that the watermark is added to the audio with the highest temporal accuracy.
[0072] After superposition, the generated watermarked audio output signal x wm [n], stored in the system output cache. The data in the cache is then continuously output to the next link according to the correct timing. For example, it is directly sent to the network transmission module for real-time broadcast or written to a local audio file stream. Through this buffering and timing management, this embodiment can achieve low-latency, smooth and continuous streaming audio output, meeting the requirements of real-time applications. The output audio stream not only carries the original content, but also secretly carries the embedded dynamic watermark information.
[0073] Step S240: Perform watermark detection on the sampling points of the received audio signal to be tested to obtain a watermark confidence vector for the sampling point; perform threshold judgment on the watermark confidence vector to restore the corresponding watermark information segment.
[0074] Specifically, this step aims to initially extract potential watermark information from the received audio signal x'[n] under test, which may have been distorted by channel transmission. A watermark confidence vector can be output for each or a portion of representative sample points in the audio signal. The specific bits of the watermark can be recovered based on the watermark confidence vector, and then, combined with time information, the individual bits can be integrated into segmented watermark information segments. This step can be implemented using the neural network or time-frequency analysis method corresponding to step S220. The specific implementation method is not limited in this embodiment.
[0075] In this embodiment, a streaming audio input signal and a corresponding watermark information segment are obtained; the watermark information segment consists of bit information segments with different states and variable lengths; a sampling point-level watermark signal is determined according to the audio input signal and the watermark information segment; the sampling point-level watermark signal is embedded in the audio input signal to obtain an audio output signal, and the audio output signal is streamed; watermark detection is performed on the sampling points of the received audio signal to be tested to obtain a watermark confidence vector for the sampling point; a threshold judgment is performed on the watermark confidence vector to restore the corresponding watermark information segment, thereby solving the problem of poor dynamics and flexibility in watermark processing. Based on sampling point-level operations, the length and content of the watermark can be dynamically changed according to the sampling points of the audio input signal, supporting highly flexible dynamic watermark information embedding, and greatly enhancing the watermark information carrying capacity. It is suitable for scenarios that require real-time updating of identification, tracking of dynamic states, or embedding of complex metadata.
[0076] In some embodiments, based on step S210, obtaining a streaming audio input signal includes:
[0077] Step S211 : obtaining an initial audio signal by dividing a continuous audio stream into short speech segments.
[0078] Specifically, the initial audio signal can consist of a single short speech segment or multiple short speech segments, depending on actual needs. Receiving streaming input as short speech segments effectively processes continuous audio data streams, particularly in real-time applications such as live streaming, real-time communication, or processing long audio files. Unlike loading an entire audio file all at once, streaming input means that audio data is delivered to the processing system in successive small chunks (called "frames" or "segments"). Short speech segments can be of fixed or variable length. Each frame (i.e., each short speech segment) contains a certain number of audio samples. The frame length (or window size) is determined based on application requirements and the input requirements of processing units (such as downstream neural networks). For example, it can be set to hundreds to thousands of samples, corresponding to a duration of tens to hundreds of milliseconds. This streaming processing approach enables the system to handle audio input of any length with low memory usage and processing latency, while also enabling simultaneous watermarking during speech without impacting the original real-time application experience.
[0079] Step S212: pre-process the initial audio signal to obtain an audio input signal.
[0080] Specifically, after receiving each audio frame, the system performs necessary preprocessing. Normalization aims to eliminate volume differences that may arise from different audio sources or during transmission, adjusting the audio amplitude to a standard range (e.g., [-1, 1]). This helps improve the stability and consistency of subsequent neural network processing. Dynamic normalization methods are often used in streaming, such as root mean square (RMS) normalization based on a sliding window or exponential moving average (EMA) normalization based on historical data. In other embodiments, if the input audio source contains significant background noise, an appropriate noise reduction algorithm (such as spectral subtraction, Wiener filtering, or a deep learning-based noise reduction model) is used to improve the clarity of the original audio. This may help enhance the stealth of watermark embedding or the robustness of extraction. The audio data frames, after framing and preprocessing, are sequentially fed into the subsequent watermark embedding process (e.g., steps S220 and S230).
[0081] Furthermore, when acquiring the audio signal x'[n] to be tested, a streaming framing method compatible with the embedded end can also be used for processing. Necessary preprocessing, such as normalization and noise reduction, can also be performed on the audio signal x'[n] to eliminate volume differences that may be introduced by different audio sources or during transmission, thereby improving audio clarity.
[0082] In this embodiment, streaming processing is used to achieve lower memory usage and shorter processing delay to cope with audio input of any length, and pre-processing is used to eliminate volume differences that may be introduced by different audio sources or during the transmission process, thereby improving audio clarity and increasing high-quality input time for subsequent watermark processing steps.
[0083] In some embodiments, based on step S220, a sampling point level watermark signal is determined according to the audio input signal and the watermark information segment, see Figure 3 , specifically including the following steps:
[0084] Step S221: Map the watermark information segment into a watermark feature representation that is aligned point by point with the audio input signal.
[0085] Specifically, because raw binary bit information may not be effective for neural networks to directly learn embedding strategies, it needs to be projected into a higher-dimensional continuous feature space. The information-to-feature mapping aims to map the B-bit information segment m[n] at each time point n into a rich and learnable C-dimensional feature vector e[n]. The feature vectors e[n] calculated at all time points n are sequentially combined to form a watermark feature representation M (dimension N×C) that has the same length N as the audio input signal x[n] and is point-by-point aligned.
[0086] Step S222: concatenate the watermark feature representation and the representation of the audio input signal along the channel dimension to obtain a sampling point-level fusion feature representation.
[0087] Specifically, after obtaining the point-by-point aligned watermark feature representation M, it needs to be fused with the corresponding audio input signal representation x (typically considered to be N×1 dimensional) to produce an input data that contains information from both. In this embodiment, fusion is performed using concatenation along the channel dimension. This means that the representation of the audio input signal x (1 channel) and the watermark feature representation M (C channels) are placed side by side along the feature dimension, forming a point-by-point fused feature representation F of N×(1+C) dimensions. The advantage of this fusion approach is that it fully preserves the information of the original audio signal while incorporating the learned features representing the watermark intent as additional channel information, with the two being strictly aligned in time. This enables the subsequent first neural network to simultaneously perceive the audio content and the corresponding watermark instructions in its convolutional or other processing layers, thereby learning to generate an accurate, content-relevant, point-by-point watermark signal.
[0088] In step S223, the sampling point-level fusion feature representation is input into a preset first neural network, and the disturbance information at each sampling point is calculated to obtain a sampling point-level watermark signal.
[0089] Specifically, the first neural network can be configured to support streaming processing, which means that the first neural network can adopt mechanisms such as causal convolution and cross-segment state management, so that it can process continuously input audio streams and their corresponding feature representations with low latency, frame by frame or block by block, without waiting for the entire audio file or long periods of data. This streaming processing capability is crucial for real-time watermark embedding applications (such as live broadcast and real-time communication).
[0090] The first neural network, through its multi-layered structure (e.g., a U-Net-based encoder-decoder architecture), deeply processes the input sample-level fused feature representation F. The first neural network must not only understand the information to be embedded (the watermark information segment m[n]) contained in the watermark feature channel but also fully consider the content of the audio input signal x[n] in the audio feature channel. This is because the generated sample-level watermark signal δ[n] must not only meet the information embedding requirements but also minimize the impact on auditory quality when superimposed on the original audio input signal x[n] (i.e., maintain high fidelity). Therefore, during training, the first neural network learns a balancing strategy: it adjusts the shape and amplitude of the generated sample-level watermark signal δ[n] based on the characteristics of the audio content (such as energy, frequency content, and temporal structure). This ensures that the generated watermark signal δ[n] embeds a stronger signal in areas that are less noticeable to the human ear (taking advantage of the auditory masking effect) to improve robustness, while embedding a weaker signal in sensitive areas to ensure imperceptibility.
[0091] Ultimately, for each input frame of sample-level fused feature representation F (length N), the first neural network outputs a sample-level watermark signal δ[n] of exactly the same length and dimension N×1. δ[n] is a real-valued sequence representing the specific perturbation value to be added to each corresponding sample point in the original audio input signal. This "point-by-point" nature ensures the highest possible temporal accuracy for watermark embedding. This generated δ[n] is then used to modify the audio input signal x[n] in step S230.
[0092] In this embodiment, by collaboratively applying a neural network architecture designed for sample point processing, high-precision embedding of audio watermarks at the single sampling point level is achieved, making the temporal positioning accuracy far superior to traditional audio frame-based methods, providing a more advantageous technical foundation for applications that require fine control and analysis (such as precise tampering positioning).
[0093] In some embodiments, the step S221 of mapping the watermark information segment into a watermark feature representation that is point-by-point aligned with the audio input signal includes:
[0094] Step S310: Obtain a learnable embedding layer and a corresponding mapping relationship.
[0095] Specifically, in this embodiment, the mapping is achieved through a learnable embedding layer E. The learnable embedding layer E can be regarded as a parameterized lookup table that stores 2B different C-dimensional vectors, each of which corresponds to one of the two possible states (0 or 1) of a bit in the B bits. By obtaining the preset mapping relationship at the same time, the required information can be found in the learnable embedding layer E to represent the watermark information segment, thereby achieving mapping. In one embodiment, for the bit information m[n] at time point n, the i-th bit m of the bit information m[n] is i [n], the mapping relationship indicates that the corresponding C-dimensional vector is accurately selected from the learnable embedding layer E according to its value (0 or 1). For example, the index in the learnable embedding layer E is 2i+m i Since the parameters of the learnable embedding layer E are learnable, the model can automatically optimize the mapping relationship from bits to features during training.
[0096] Step S320: extracting a set of target row vectors corresponding to each bit information segment in the watermark information segment from the learnable embedding layer based on the mapping relationship; and obtaining a feature vector corresponding to the bit information segment based on the set of target row vectors.
[0097] Specifically, the B C-dimensional vectors selected for each B-bit information segment are combined to generate a single C-dimensional feature vector e[n] representing the overall watermark information at time point n. For example, the combination operation can be implemented by summing, and the calculation formula is: , where e[n] represents the feature vector, B represents the total number of bits in the bit information segment, and i represents the bit index of the bit information segment. represents the 2i+mth layer in the learnable embedding layer E i A row vector of [n] rows.
[0098] Step S330: Combine the feature vectors in time sequence to form a watermark feature representation that is aligned point by point with the audio input signal.
[0099] Specifically, the feature vectors e[n] calculated at all time points n are combined in sequence to form a watermark feature representation M (dimension N×C) that has the same length N as the audio input signal x[n] and is point-by-point aligned.
[0100] In this embodiment, binary discrete data is projected into a higher-dimensional continuous feature space based on a learnable embedding layer, and the mapping relationship from bits to features is automatically optimized through the learnable embedding layer.
[0101] In some embodiments, in step S240, watermark detection is performed on the sampling points of the received audio signal to be tested to obtain a watermark confidence vector for the sampling points, which specifically includes the following steps:
[0102] In step S410, the received audio signal to be tested is input into a preset second neural network, and watermark detection is performed on the sampling points to obtain a multi-dimensional watermark confidence vector for the sampling points; the multi-dimensional watermark confidence vector includes watermark existence information and watermark content information.
[0103] Specifically, the second neural network is configured to receive the audio signal under test and output a sample-level multidimensional confidence vector, i.e., a watermark confidence vector. The structure of the second neural network includes an encoder component for feature extraction, which includes at least a convolutional layer and a channel attention mechanism. For example, an architecture that includes depthwise separable convolution and a channel attention mechanism, such as an architecture based on the SeaNet encoder, can be used.
[0104] Among them, one element of the watermark confidence vector c[n] is the watermark existence information c0[n], which is used to indicate the confidence or probability estimate of the presence of the watermark signal at the current sample point n. This value is usually between 0 and 1. The closer the value is to 1, the higher the possibility that the network believes that the watermark exists at this point. The remaining B elements in the watermark confidence vector c[n] are watermark content information. , the B elements of the watermark content information correspond to the confidence that each bit in the embedded B-bit information segment is "1", and these values are usually between 0 and 1, c i The closer [n] is to 1, the more likely the network believes that the i-th bit is "1" at that moment. Preferably, the watermark existence information c0[n] is located in the first position, and the watermark content information It is located at the subsequent B bits of data, so the dimension of the watermark confidence vector c[n] is (B+1).
[0105] It's understandable that the output confidence vector c[n] is a continuous-valued, potentially noisy raw network output. It represents the neural network's "initial judgment" or "belief strength" regarding the watermark's existence and content at each sample point, rather than the final, discrete detection result. This sequence of sample-level confidence vectors forms the basis for subsequent refinement.
[0106] In this embodiment, based on the second neural network, the sampling point level operation is further delayed to the watermark extraction stage, overcoming the limitation of most existing technologies that can only embed and extract static or fixed pattern information.
[0107] In some embodiments, in step S240, a threshold value is determined for the watermark confidence vector to recover the corresponding watermark information segment. Figure 4, specifically including the following steps:
[0108] Step S510: Smoothing the watermark confidence vector to obtain a smoothed result.
[0109] Specifically, the goal of calculating the sampling-point-level multidimensional watermark confidence vector c[n] is to transform this raw, potentially unstable confidence data into a final, clear and reliable discrete watermark. However, the watermark confidence vector c[n] is a raw confidence sequence, which may contain jitter and errors caused by noise, channel distortion, or network prediction uncertainty. Directly applying a threshold to it may result in a large number of bit errors and fragmented results. Therefore, post-processing operations are required, mainly including smoothing and binarization.
[0110] Smoothing, specifically temporal smoothing, is used to eliminate transient noise and isolated mispredictions in the confidence sequence, ensuring that the recovered watermark message exhibits a stable paragraph structure with a certain duration, similar to that of the original embedding. A common approach to temporal smoothing is to apply a sliding window technique. For each confidence vector c[n] of a sample point n, all confidence vectors within a temporally adjacent window are examined. The confidence of the central point n is then modified based on the dominance of the states within the window (watermark presence / absence, and the 0 / 1 state of each bit) (e.g., average confidence calculated by the mean, median, or weighted average, or directly by majority voting). For example, if a point's confidence contradicts the majority of its preceding and following neighbors and its own confidence is low, it may be adjusted to align with its neighbors. Low-confidence points, whose confidences are close to a decision threshold (e.g., 0.5), can also be treated specifically by assigning them lower weights or employing a more conservative update strategy during the smoothing process. Through time smoothing, noise can be effectively suppressed and broken segments can be connected, making the confidence sequence more stable and better reflecting the structure of the original embedded segmented watermark message.
[0111] Step S520 : Binarize the smoothed result based on a preset decision threshold to obtain a binary decision sequence.
[0112] Specifically, since the smoothed result after smoothing is still a continuous confidence vector, the watermark bit information cannot be directly extracted. It is necessary to use binarization processing to convert the smoothed result into a discrete binary (0 or 1) judgment result. Specifically, the smoothed result includes watermark existence information c0[n] and watermark content information c i[n], a preset decision threshold (usually 0.5) is applied to the smoothed result. If a confidence value in the smoothed result is greater than the decision threshold, the corresponding state is judged to be 1 (watermark exists, or the bit is 1); if it is less than or equal to the threshold, it is judged to be 0 (watermark does not exist, or the bit is 0). After binarization, the discrete watermark information sequence m'[n] is obtained, that is, the binary decision sequence, including the watermark existence flag bit (recovered from the watermark existence information c0[n]) and B message bits (recovered from the watermark content information c i [n] recovered).
[0113] Step S530: When the watermark existence information indicates that a watermark exists, a bit information segment is obtained based on a corresponding binary determination sequence.
[0114] Specifically, the binary decision sequence is discrete, relatively stable, and as close as possible to the original embedded watermark bit information segment. At the same time, based on the recovered existence flag bit sequence, it is possible to determine which time periods contain watermarks.
[0115] Step S540: Restore the watermark information segment based on the bit information segment corresponding to each sampling point.
[0116] Specifically, based on the existence judgment of the watermark, each B message bit is combined into a final detection result, that is, a watermark information segment.
[0117] In this embodiment, through post-processing steps based on smoothing and binarization, jitter and errors caused by noise, channel distortion or network prediction uncertainty are eliminated, and the accuracy of watermark information segment recovery is improved. This overcomes the limitation of most existing technologies that can only embed static or fixed pattern information, greatly enhances the flexibility and information carrying capacity of the watermark, and is suitable for scenarios such as real-time updating of identification, tracking of dynamic status or embedding of complex metadata.
[0118] In some embodiments, the first neural network adopts an encoder-decoder structure, uses causal convolution operations, and has a state transfer mechanism across audio processing segments.
[0119] Specifically, the first neural network is configured to receive sample-level fused feature representations and generate sample-level watermark signals. For example, an encoder-decoder structure, such as a network based on the U-Net architecture or the Demucs architecture, can be used. The convolution operation in the network is transformed into a causal convolution to ensure that the output calculated at the current time point does not depend on any future input information. Based on the causal operation, a state transfer mechanism can be implemented across audio processing segments to maintain the necessary temporal context dependencies in streaming processing.
[0120] In this embodiment, by causally transforming the first neural network architecture and adopting a state transfer mechanism across audio processing segments, low-latency, continuous streaming processing of audio signals is supported, which can be seamlessly integrated into systems with high real-time requirements such as live broadcast and real-time communication, meeting the key needs of modern audio applications.
[0121] In some of these embodiments, see Figure 5 The first deep neural network includes an encoder, a bottleneck layer, and a decoder connected in sequence. The encoder includes multiple stages of first residual downsampling modules connected in sequence; each stage of the first residual downsampling module includes a connected residual convolution block and a downsampling layer; the decoder includes a first residual upsampling module that is symmetrical to the first residual downsampling module and connected to the encoder to introduce skip connections and restore signal details; each stage of the first residual upsampling module includes a connected residual convolution block and an upsampling layer. The bottleneck layer includes a first long short-term neural network to capture temporal dependencies.
[0122] The first deep neural network may also include two independent first one-dimensional convolutional layers, each connected to an encoder to receive the original audio (i.e., the audio input signal) and the sample-level information to be embedded (i.e., the watermark feature representation). Feature extraction is performed on each of these layers through independent first one-dimensional convolutional layers, and then fused to form a sample-level fused feature representation. This sample-level fused feature representation is fed into the encoder of the first deep neural network and subsequently processed through a bottleneck layer and a decoder to generate a sample-level watermark signal. Finally, the sample-level watermark signal is embedded into the original audio through a point-by-point superposition process, resulting in the watermarked audio (i.e., the audio output signal).
[0123] In this embodiment, the watermark signal can be adaptively adjusted according to the original audio content to balance robustness and imperceptibility.
[0124] In some embodiments, the second neural network includes an encoder component, which includes at least a convolutional layer and a channel attention mechanism layer, where the convolutional layer uses a depthwise separable convolution operation. Specifically, the second neural network is configured to receive a test audio signal and output a sample-level watermark confidence vector. For example, an architecture including depthwise separable convolution and a channel attention mechanism can be used, such as an architecture based on the SeaNet encoder.
[0125] In some of these embodiments, see Figure 6 The second neural network also includes a one-dimensional transposed convolutional layer connected to the encoder component. The encoder component includes a second one-dimensional convolutional layer, a multi-stage second residual downsampling module, a second long short-term neural network, and a third one-dimensional convolutional layer connected in sequence; each stage of the second residual downsampling module includes a connected residual convolution block and a downsampling layer.
[0126] Specifically, the second one-dimensional convolutional layer performs preliminary feature extraction on the audio signal under test. The multi-stage second residual downsampling module integrates depthwise separable convolution and a channel attention mechanism to enhance feature representation. The second long-short-time neural network processes sequence information. The third one-dimensional convolutional layer refines features. Finally, a one-dimensional transposed convolutional layer outputs a multidimensional confidence vector at the sampling point level as the watermark detection result (i.e., the watermark confidence vector). This vector indicates the likelihood of detecting the predetermined watermark information at each time position in the audio.
[0127] In this embodiment, by causally transforming the first neural network architecture and adopting a state transfer mechanism across audio processing segments, the second neural network is designed to adopt an architecture that includes depthwise separable convolution and channel attention mechanism, thereby supporting low-latency, continuous streaming processing of audio signals, and can be seamlessly integrated into systems with high real-time requirements such as live broadcast and real-time communication, meeting the key needs of modern audio applications.
[0128] In some embodiments, the first neural network and the second neural network are obtained by performing end-to-end joint training on the initial generation network and the initial detection network. Figure 7 , the forward data propagation process of end-to-end joint training includes:
[0129] Step S610: Based on the audio test signal, a watermark test segment with a preset number of bits is randomly generated within a preset time length range.
[0130] Specifically, the audio test signal can be audio data obtained from a database, or can be streaming input audio data received in a frame-by-frame manner. Furthermore, the audio data can be pre-processed to improve data quality. Figure 7 The original information end in is the watermark test segment, and the original information end is aligned with the audio (ie, the audio test signal) in the time domain.
[0131] Step S620: convert the watermark test segment into a test feature vector at the sampling point level, concatenate the test feature vector with the audio test signal, and input the resultant vector into the initial generation network, outputting a one-dimensional test watermark signal of the same length as the audio test signal; and embedding the one-dimensional test watermark signal into the audio test signal to obtain a watermarked test signal.
[0132] Specifically, Figure 7 The sampling point-level information in is the test feature vector. Based on the learnable embedding layer E, bits are mapped to features to obtain the sampling point-level test feature vector. The sampling point-level test feature vector is aligned with the audio test signal in terms of length and sampling points. Figure 7 The watermarked audio shown in FIG is the watermarked test signal.
[0133] Step S630: perform data enhancement processing on the watermarked test signal, input the enhanced watermarked test signal into the initial detection network, and output the watermark detection confidence at the sampling point level, including watermark 1-bit existence test information and watermark content test information of a preset number of bits.
[0134] Specifically, to further improve the performance of the neural network, especially in terms of the positioning accuracy of the watermark signal and the robustness to channel distortion and malicious attacks (such as audio clipping, splicing and tampering) that may be encountered in practical applications, one or more data augmentation techniques can be further used to apply simulated transformations or perturbations to the training data, so that the neural network model can improve its ability to operate stably under diverse conditions. Figure 7 The watermarked test signal can be enhanced through data enhancement methods such as active masking, channel simulation, and tampering simulation to achieve accurate positioning of the model's sampling points, resistance to channel interference, and resistance to active attack destruction. This embodiment does not limit the data enhancement method.
[0135] Figure 7 The sampling point level logic value shown in is the watermark detection confidence. In the watermark embedding stage, the sampling point level information contains the preset B-bit information. In the watermark extraction stage, the corresponding watermark content test information (i.e. Figure 7 The sampling point level information in the test result is a logical value) and there is also a B bit. In addition, there is one more bit of test information in the test result (i.e. Figure 7 The sampling point level watermark exists in the logical value of bit).
[0136] After the end-to-end joint training is completed, the optimized neural network parameters are loaded into the modules running the first neural network and the second neural network for the actual watermark embedding and extraction tasks.
[0137] In this embodiment, although end-to-end joint training does not directly participate in the watermark embedding and extraction at runtime, it is the key support for the effective operation of the entire system. With the powerful fitting ability of deep learning and combined with optimized training methods such as data enhancement technology, it can resist common audio distortions (such as compression and noise) while maximally maintaining the auditory quality of the original audio, achieving a good balance between robustness and imperceptibility, and improving the efficiency and accuracy of watermark processing.
[0138] In some embodiments, the data enhancement processing includes at least one of channel distortion enhancement, watermark mask enhancement, and tampering simulation enhancement.
[0139] Channel distortion enhancement aims to simulate various signal degradations that audio signals may experience during transmission, storage, or playback, thereby improving the model's robustness to common channel distortions. For example, noise of varying types and intensities, such as additive white Gaussian noise and specific environmental noise, can be added to training samples (i.e., watermarked test signals) to simulate varying signal-to-noise ratio (SNR) conditions. Another example is the ability to modify the sampling rate of training samples, such as by downsampling and then upsampling, to simulate distortion introduced by sample rate conversion or mismatches between devices. Another example is the application of lossy compression, where training samples are encoded and decoded at different bit rates using common audio codecs (such as MP3, AAC, and Opus) to simulate information loss caused by compression. Optionally, simulation of other linear or nonlinear distortions, such as equalization, clipping, echo, and reverberation, can also be included.
[0140] Watermark mask enhancement is primarily used to train the detection model's ability to accurately identify the boundaries of watermark-embedded regions, which is crucial for resisting content cropping attacks and accurately extracting watermark information. Specifically, a random binary mask can be generated during training. This mask defines alternating regions of 1 (indicating the presence of a watermark) and 0 (indicating the absence of a watermark) along the time dimension. This random binary mask can be applied to training samples, for example, by multiplying it with an ideal watermark perturbation signal, or by selectively retaining or removing watermark components from watermarked audio, thereby creating training samples containing alternating watermark and non-watermark regions. The training objective is then adjusted accordingly to not only detect the presence of a watermark, but also accurately predict the boundaries defined by the mask. Parameters such as the mask's pattern and paragraph length distribution can be randomly varied during training.
[0141] Tampering simulation enhancement aims to improve the model's ability to detect malicious tampering behaviors such as splicing and replacement of audio content. Such tampering usually destroys the continuity or intrinsic association between the original audio and the embedded watermark. For example, one or more time segments can be randomly selected in the watermarked test signal. The selected segments are replaced with other audio content, such as audio segments from different sources, corresponding original unwatermarked segments, or segments embedded with different watermark messages / keys. This operation simulates editing tampering behaviors such as cutting and pasting, effectively destroying the temporal continuity or consistency of the audio and / or watermark signals. By training the neural network model to identify such discontinuities or inconsistencies introduced by tampering (for example, by monitoring decoding errors, feature mutations, etc.), its ability to detect splicing-type tampering can be improved.
[0142] It should be understood that the three data augmentation techniques described above can be implemented individually or in any combination on the same training sample (i.e., the watermarked test signal). The specific parameters used in the augmentation operation (such as noise level, compression bit rate, mask parameters, replacement segment position and length, etc.) are typically randomized within a preset range to maximize the diversity of the training data.
[0143] In this embodiment, by selecting and combining a variety of data enhancement methods, the positioning performance and anti-channel robustness can be further enhanced, or the robustness against malicious attacks such as cropping can be improved, so as to adapt to different application scenarios.
[0144] In some embodiments, end-to-end joint training is optimized using a multi-task loss function; the multi-task loss function includes a perceptual loss function and a watermark loss function. The perceptual loss function includes at least one of the following: a deep feature loss function, a loudness domain loss function, a frequency domain loss function, and a time domain loss function. The watermark loss function includes at least one of the following: a watermark presence detection loss function, a multi-bit message detection loss function, and a boundary consistency loss function.
[0145] Specifically, a multi-task loss function L is used to optimize the first and second neural networks in the end-to-end joint training. The goal of the multi-task loss function L is to simultaneously guide the learning of the two neural networks so that they can achieve optimal performance in mutual collaboration. Specifically, it is configured to jointly optimize two seemingly contradictory but critical goals for the watermarking system: one is to maintain high perceptual quality of the watermarked audio, that is, to be as imperceptible to the human ear as possible; the other is to ensure that the embedded watermark information can be accurately and robustly detected and extracted, even after the audio has been subjected to common distortions such as compression and noise.
[0146] The multi-task loss function L can be constructed by weighted summation of multiple specific loss terms, and its overall form can be expressed as:
[0147] ;
[0148] Among them, w wav 、w mel 、w loud 、w deep 、w det 、w msg and w bound Represents the weight coefficient of each loss, which is used to adjust the relative importance of different targets in the total loss; L wav represents the time domain loss function; L mel represents the frequency domain loss function; L loud represents the loudness domain loss function; L deep represents the deep feature loss function; Ldet Watermark existence detection loss function; L msg Multi-bit message detection loss function; represents L bound represents the boundary consistency loss function.
[0149] Time domain loss function (L wav ), which can directly constrain the difference at the time domain waveform level, by calculating the L2 distance between the original audio test signal and the watermarked test signal in the time domain waveform, thereby limiting the basic acoustic distortion.
[0150] Frequency domain loss function (L mel ), which can measure the perceptual difference from the frequency domain perspective. It is obtained by calculating the L2 distance between the multi-scale Mel spectrum representation of the original audio test signal and the watermarked test signal, achieving an effect that is closer to the frequency perception characteristics of the human ear.
[0151] Loudness domain loss function (L loud ), which can use the auditory masking principle to perform more refined perception optimization, and its calculation method is:
[0152] ;
[0153] ;
[0154] in, Represents a one-dimensional test watermark signal for embedding in a specific frequency band b and time window w With the original audio test signal The loudness difference between This method allows for greater loudness variations to be tolerated in time-frequency regions where the human ear is insensitive, while requiring smaller differences in sensitive regions.
[0155] Deep feature loss function (L deep ), which aims to capture higher-level and more abstract perceptual feature differences and is calculated as follows:
[0156] ;
[0157] in, represents the L2 norm, represents the features extracted from the first three layers of a pre-trained audio model (such as Wav2Vec2.0), w i The layer feature weights are preset, and the representation ability of the pre-trained model is used to constrain the perceptual artifacts introduced by the watermark.
[0158] The four loss functions listed above (deep feature loss, loudness domain loss, frequency domain loss, and time domain loss) are perceptual loss functions. Perceptual loss functions focus on minimizing the impact of watermark embedding on the listening experience of the original audio, aiming to maintain high audio fidelity. Depending on the actual application scenario, you can choose to use at least one of these perceptual loss functions.
[0159] Watermark existence detection loss function (L det ), which is used to train the detector to accurately determine whether a watermark exists at each sample point. It is obtained by calculating the binary cross entropy (BCE) between the predicted watermark existence probability p[n] and the true watermark mask mask[n], specifically:
[0160] ;
[0161] Where N is the audio length.
[0162] Multi-bit message detection loss function (L msg ), which is used to train the detector to accurately recover the embedded B-bit message content, and is obtained by calculating the binary cross entropy (BCE) between the predicted confidence of each watermark bit (c[n][b]) and the true watermark bit (m[n][b]), specifically:
[0163] ;
[0164] Where B is the number of watermark bits.
[0165] Boundary consistency loss function (L bound ), which aims to enhance the temporal continuity and stability of the detection results and reduce isolated errors. It is obtained by calculating the binary cross entropy (BCE) between the predicted confidence difference (△c[n][b]) of adjacent sample points and the true bit difference (△m[n][b]), specifically:
[0166] ;
[0167] ;
[0168] .
[0169] The three loss functions listed above (watermark presence detection loss, multi-bit message detection loss, and boundary consistency loss) are watermark loss functions. Watermark loss functions focus on improving watermark detection and extraction performance, ensuring accurate information transmission and system robustness. Depending on the actual application, you can choose to use at least one of these watermark loss functions.
[0170] In this embodiment, by adjusting and weighting these perceptual loss terms and watermark loss terms, the neural network can be effectively guided to learn how to achieve high-precision and high-robust dynamic watermark embedding and detection while maintaining audio quality.
[0171] In some embodiments, the end-to-end joint training is optimized using a curriculum training strategy; the curriculum training strategy includes at least one of the following: dynamic adjustment of the weight of the perceptual loss function, and progressive increase in the number of bits of the watermark information segment.
[0172] Specifically, to further optimize the stability of the training process and improve the overall performance of the final model (for example, achieving a good balance between embedding capacity, robustness, and audio fidelity), a curriculum training strategy can be further adopted. The core idea of this strategy is to present training tasks or data to the model in a certain order or in increasing difficulty, simulating the learning process of humans or animals, from easy to difficult.
[0173] The strategy of dynamically adjusting the weight of the perceptual loss function aims to guide the model's learning focus by dynamically adjusting the weight of the perceptual loss term in the overall loss function. In the early stages of training, the perceptual loss term is assigned a relatively low weight. This allows the model to initially focus on learning the basic patterns and mapping relationships for watermark embedding and extraction, without being overly constrained by audio quality (measured by the perceptual loss). This reduces the risk of early training crashing or falling into a local optimum due to overly strong constraints. As training progresses (for example, with increasing training steps or epochs), the weight of the perceptual loss term is gradually, smoothly, or staged. This gradual increase in weight allows the model, after mastering basic embedding capabilities, to prioritize maintaining the perceptual quality of the original audio. Ultimately, by the end of training, it reaches a balance between effectively embedding and detecting watermarks and producing high-quality watermarked audio. This strategy helps achieve stable convergence and optimizes final audio fidelity. In a specific implementation, the perceptual loss term can be omitted or used with very low weight at the beginning of training (for example, the first 10,000 steps). Data augmentation can be omitted at this stage to allow the model to first learn the basic task. Afterwards, the perceptual loss term can be gradually introduced and weighted, and the aforementioned data augmentation techniques can be introduced at the same time (or later).
[0174] Among them, the strategy of gradually increasing the number of bits of the watermark information segment aims to gradually adapt the model to more difficult embedding and extraction tasks by gradually increasing the number of bits B (i.e., embedding capacity) of the watermark message to be embedded. The model first learns robust embedding and detection at a lower bit number, and after mastering the basic capabilities, it challenges a higher information embedding rate. Training starts with a preset initial bit number B0 (for example, B0=1 or a smaller value). As training progresses (for example, according to the number of training steps t), the number of bits B of the watermark message increases. t is gradually increased until it reaches the preset maximum number of bits B max The increase in the number of bits can follow a predetermined strategy. For example, the following progressive expansion strategy can be used to calculate the number of bits B used in the training step t: t :
[0175] ;
[0176] Among them, B t is the number of watermark message bits used in training step t; B max is the target maximum watermark bit number; B0 is the initial watermark bit number; △B is the number of bits added each time (for example, △B=1); t is the current training step number; T is the bit number update period (that is, the number of bits is increased every T steps); Indicates a floor operation.
[0177] This approach of incrementally increasing the difficulty of tasks helps prevent training instability when the model directly handles high-complexity tasks (high bitrate embeddings), allowing the model to be trained to more effectively learn feature representations that are robust under different embedding capacities.
[0178] It should be understood that the two curriculum learning strategies listed above can be used alone or in combination. For example, one can start training at a lower number of bits and a lower perceptual loss weight, and then gradually increase the perceptual loss weight and the number of watermark bits simultaneously or in an interleaved manner.
[0179] In this embodiment, by implementing a training method that includes one or more curriculum strategies, it helps to achieve a more stable and efficient training process, and enables the final trained model to achieve better balance and optimization in multiple performance dimensions such as audio quality preservation, watermark embedding capacity and robustness, thereby being suitable for efficient and reliable audio watermark embedding and detection tasks in actual deployment.
[0180] The present embodiment is described and illustrated below through preferred embodiments. Figure 8 This is a flowchart of the high-precision dynamic watermarking method for streaming audio according to the preferred embodiment of the present invention.
[0181] S1, watermark embedding stage: embed watermark information with sampling point level positioning accuracy into the streaming speech input in real time.
[0182] S1.1. Obtaining an audio input signal and a watermark information segment: Receive an audio signal streaming input by dividing a continuous audio stream into short speech segments, and perform necessary preprocessing operations such as normalization and noise reduction on the obtained audio signal to obtain a final audio input signal; simultaneously input a watermark information segment of the same length and aligned with the audio input signal, wherein the watermark information segment consists of B-bit information segments with different states and variable lengths, where B is an integer greater than or equal to 1.
[0183] S1.2. Sampling point-level watermark signal generation: Map the dynamic watermark message into a watermark feature representation that is point-by-point aligned with the audio input signal, and concatenate the watermark feature representation with the representation of the audio input signal along the channel dimension to obtain a sampling point-level fused feature representation; configure the sampling point-level fused feature representation input to a first neural network that supports streaming processing to generate a sampling point-level watermark signal corresponding to the audio input signal.
[0184] S1.3. Watermark signal embedding and output: The sampling point-level watermark signal is superimposed on the audio input signal in the form of direct addition of time domain signals to generate a watermarked audio signal, which is then stored in the system output cache to achieve voice streaming output.
[0185] S2, watermark extraction stage: detect whether there is a watermark in the speech and extract the corresponding sampling point level watermark information.
[0186] S2.1. Sample-level confidence detection: The received audio signal is input into the second neural network (detector network). The second neural network outputs a multidimensional confidence vector for each or part of the sample points in the audio signal. The vector contains the confidence of the watermark existence and the confidence of each watermark message bit;
[0187] S2.2. Post-processing of watermark information segments: Post-process the multi-dimensional confidence vector, including temporal smoothing and binarization, to restore the embedded segmented watermark message.
[0188] In this preferred embodiment, a high-precision dynamic audio digital watermarking method is provided to address the limitations of existing audio watermarking technologies in terms of embedding and detection accuracy, information dynamics, and streaming processing capabilities. This method utilizes a streaming deep neural network architecture to accurately embed and extract dynamic watermark messages consisting of bit segments with different states and variable lengths by operating at the sample point level of the audio signal. Specifically, this preferred embodiment utilizes a first neural network to generate a sample point-level watermark signal and combines it with the original audio. At the same time, a second neural network is used to output a sample point-level multidimensional confidence vector and restore the watermark information through post-processing. This effectively overcomes the shortcomings of traditional watermarking methods in terms of accuracy, flexibility, and real-time performance. It aims to provide a more accurate, flexible, and efficient copyright protection, content tracking, and management solution for digital audio content (especially for scenarios such as real-time streaming and AI-generated content).
[0189] It should be noted that the steps shown in the above process or the flowchart in the accompanying drawings can be executed in a computer system such as a set of computer-executable instructions, and although a logical order is shown in the flowchart, in some cases, the steps shown or described can be executed in an order different from that shown here.
[0190] It should be noted that the user-related information (including but not limited to the audio content itself, possible embedded identity identification, device information, etc.) and data (including but not limited to parameters used for watermark generation / detection, feature data, storage data, display data, etc.) involved in this application are all information and data authorized or licensed by the user and fully authorized by all parties, and the collection, use and processing of relevant data must comply with relevant laws, regulations and standards, and corresponding operation entrances must be provided for users to choose to authorize or refuse.
[0191] This embodiment also provides a high-precision dynamic watermark device for streaming audio. The device is used to implement the above-mentioned embodiments and preferred implementations. Details that have already been described will not be repeated. The terms "module," "unit," "subunit," etc. used below may refer to a combination of software and / or hardware that implements a predetermined function. Although the devices described in the following embodiments are preferably implemented in software, implementation using hardware, or a combination of software and hardware, is also possible and contemplated.
[0192] Figure 9 This is a structural block diagram of the high-precision dynamic watermarking device for streaming audio according to this embodiment. Figure 9 As shown, the device includes: a data acquisition module 91, a sampling point level watermark generation module 92, a watermark embedding module 93 and a watermark extraction module 94.
[0193] The data acquisition module 91 is used to acquire the streaming audio input signal and the corresponding watermark information segment; the watermark information segment is composed of bit information segments with different states and variable lengths;
[0194] A sampling point level watermark generation module 92 is used to determine a sampling point level watermark signal according to the audio input signal and the watermark information segment;
[0195] The watermark embedding module 93 is used to embed the sampling point level watermark signal into the audio input signal to obtain the audio output signal and stream the audio output signal;
[0196] The watermark extraction module 94 is configured to perform watermark detection on the sampling points of the received audio signal to be tested, obtain a watermark confidence vector for the sampling point, perform threshold judgment on the watermark confidence vector, and restore the watermark information segment.
[0197] In some embodiments, the sampling point-level watermark generation module 92 is further configured to map the watermark information segment into a watermark feature representation that is point-aligned with the audio input signal; concatenate the watermark feature representation with the representation of the audio input signal along the channel dimension to obtain a sampling point-level fused feature representation; and input the sampling point-level fused feature representation into a preset first neural network to calculate the disturbance information at each sampling point to obtain a sampling point-level watermark signal.
[0198] In some of the embodiments, the sampling point-level watermark generation module 92 is further used to obtain a learnable embedding layer and a corresponding mapping relationship; based on the mapping relationship, extract a set of target row vectors corresponding to each bit information segment in the watermark information segment from the learnable embedding layer; based on the set of target row vectors, obtain a feature vector corresponding to the bit information segment; and combine the individual feature vectors in time sequence into a watermark feature representation that is point-by-point aligned with the audio input signal.
[0199] In some embodiments, the watermark extraction module 94 is further configured to input the received audio signal to be tested into a preset second neural network, perform watermark detection on the sampling points, and obtain a multi-dimensional watermark confidence vector for the sampling points; the multi-dimensional watermark confidence vector includes watermark existence information and watermark content information.
[0200] In some of the embodiments, the watermark extraction module 94 is further configured to smooth the watermark confidence vector to obtain a smoothed result; binarize the smoothed result based on a preset decision threshold to obtain a binary determination sequence; obtain a bit information segment based on the corresponding binary determination sequence when watermark existence information indicates the existence of a watermark; and restore the watermark information segment based on the bit information segment corresponding to each sampling point.
[0201] In some embodiments, the first neural network adopts an encoder-decoder structure, adopts causal convolution operations, and has a state transfer mechanism across audio processing segments; the second neural network includes an encoder component, the encoder component includes at least a convolution layer and a channel attention mechanism layer, and the convolution layer adopts depth-separable convolution operations.
[0202] In some embodiments, the first neural network and the second neural network are obtained based on end-to-end joint training of an initial generation network and an initial detection network; the forward data propagation process of the end-to-end joint training includes: based on an audio test signal, randomly generating a watermark test segment with a preset number of bits within a preset time length range; converting the watermark test segment into a test feature vector at the sampling point level, splicing the test feature vector with the audio test signal and inputting the resultant into the initial generation network, and outputting a one-dimensional test watermark signal of the same length as the audio test signal; embedding the one-dimensional test watermark signal into the audio test signal to obtain a watermarked test signal; performing data enhancement processing on the watermarked test signal, inputting the enhanced watermarked test signal into the initial detection network, and outputting a watermark detection confidence at the sampling point level, including watermark existence test information and watermark content test information.
[0203] In some embodiments, the data enhancement processing includes at least one of channel distortion enhancement, watermark mask enhancement, and tampering simulation enhancement.
[0204] In some embodiments, end-to-end joint training is optimized using a multi-task loss function; the multi-task loss function includes at least one of a perceptual loss function and a watermark loss function; the perceptual loss function includes at least one of the following: a deep feature loss function, a loudness domain loss function, a frequency domain loss function, and a time domain loss function; the watermark loss function includes at least one of the following: a watermark presence detection loss function, a multi-bit message detection loss function, and a boundary consistency loss function.
[0205] In some embodiments, the end-to-end joint training is optimized using a curriculum training strategy; the curriculum training strategy includes at least one of the following: dynamic adjustment of the weight of the perceptual loss function, and progressive increase in the number of bits of the watermark information segment.
[0206] It should be noted that the above modules can be functional modules or program modules, and can be implemented through software or hardware. For modules implemented through hardware, the above modules can be located in the same processor; or the above modules can be located in different processors in any combination.
[0207] This embodiment further provides a computer device, including a memory and a processor, wherein the memory stores a computer program, and the processor is configured to run the computer program to execute the steps in any one of the above method embodiments.
[0208] Optionally, the computer device may further include a transmission device and an input / output device, wherein the transmission device is connected to the processor, and the input / output device is connected to the processor.
[0209] It should be noted that, for specific examples in this embodiment, reference may be made to the examples described in the above embodiments and optional implementation modes, and will not be repeated in this embodiment.
[0210] In addition, in conjunction with the high-precision dynamic watermarking method for streaming audio provided in the above embodiments, this embodiment may also provide a computer storage medium for implementation. The computer storage medium stores a computer program; when the computer program is executed by a processor, it implements any of the high-precision dynamic watermarking methods for streaming audio provided in the above embodiments.
[0211] It should be understood that the specific embodiments described herein are only used to explain this application and are not used to limit it. Based on the embodiments provided in this application, all other embodiments obtained by ordinary technicians in this field without creative work are within the scope of protection of this application.
[0212] Obviously, the accompanying drawings are merely examples or embodiments of the present application. A person skilled in the art can also apply the present application to other similar situations based on these drawings without inventive effort. Furthermore, it is understandable that, although the work involved in this development process may be complex and lengthy, certain design, manufacturing, or production changes based on the technical content disclosed in this application are merely routine technical means for a person skilled in the art and should not be considered to constitute a deficiency in the disclosure of the present application.
[0213] The term "embodiment" as used in this application refers to specific features, structures, or characteristics described in conjunction with the embodiment that can be included in at least one embodiment of this application. The appearance of this phrase in various places in the specification does not necessarily mean that the embodiment is the same, nor does it mean that it is mutually exclusive with other embodiments and is independent or optional. It is understood, either explicitly or implicitly, by those skilled in the art that the embodiments described in this application can be combined with other embodiments when there is no conflict.
[0214] The above-described embodiments merely represent several implementation methods of the present application. While the descriptions are relatively specific and detailed, they should not be construed as limiting the scope of patent protection. It should be noted that a person of ordinary skill in the art may make various modifications and improvements without departing from the spirit of the present application, and these modifications and improvements fall within the scope of protection of the present application. Therefore, the scope of protection of the present application shall be determined by the appended claims.
Claims
1. A high-precision dynamic watermarking method for streaming audio, characterized in that: The method comprises: Acquire a streaming audio input signal and a corresponding watermark information segment; the watermark information segment is composed of bit information segments with different states and variable lengths; Determining a sampling point-level watermark signal according to the audio input signal and the watermark information segment; Embedding the sampling point-level watermark signal into the audio input signal to obtain an audio output signal, and streaming the audio output signal; Perform watermark detection on the sampling points of the received audio signal to be tested to obtain a watermark confidence vector for the sampling point; perform threshold judgment on the watermark confidence vector to restore the corresponding watermark information segment.
2. The high-precision dynamic watermarking method for streaming audio according to claim 1, characterized in that: Determining a sampling point-level watermark signal according to the audio input signal and the watermark information segment includes: Mapping the watermark information segment into a watermark feature representation that is point-by-point aligned with the audio input signal; splicing the watermark feature representation and the representation of the audio input signal along the channel dimension to obtain a sampling point-level fusion feature representation; The sampling point level fusion feature representation is input into a preset first neural network, and the disturbance information on each of the sampling points is calculated to obtain a sampling point level watermark signal.
3. The high-precision dynamic watermarking method for streaming audio according to claim 2, characterized in that: Mapping the watermark information segment into a watermark feature representation that is point-by-point aligned with the audio input signal, comprising: Get the learnable embedding layer and the corresponding mapping relationship; Based on the mapping relationship, extracting a set of target row vectors corresponding to each of the bit information segments in the watermark information segment from the learnable embedding layer; and obtaining a feature vector corresponding to the bit information segment based on the set of target row vectors; The feature vectors are combined in time sequence into a watermark feature representation that is aligned point by point with the audio input signal.
4. The high-precision dynamic watermarking method for streaming audio according to claim 2, characterized in that: Perform watermark detection on the sampling points of the received audio signal to be tested, and obtain the watermark confidence vector for the sampling point, including: The received audio signal to be tested is input into a preset second neural network, and watermark detection is performed on the sampling points to obtain a multi-dimensional watermark confidence vector for the sampling points; the multi-dimensional watermark confidence vector includes watermark existence information and watermark content information.
5. The high-precision dynamic watermarking method for streaming audio according to claim 4, characterized in that: Performing a threshold judgment on the watermark confidence vector to restore the corresponding watermark information segment includes: Smoothing the watermark confidence vector to obtain a smoothed result; Based on a preset decision threshold, binarizing the smoothed result to obtain a binary decision sequence; When the watermark existence information indicates that a watermark exists, obtaining the bit information segment based on the corresponding binary determination sequence; The watermark information segment is restored based on the bit information segment corresponding to each of the sampling points.
6. The high-precision dynamic watermarking method for streaming audio according to claim 4, characterized in that: The first neural network adopts an encoder-decoder structure, adopts causal convolution operation, and has a state transfer mechanism across audio processing segments; The second neural network includes an encoder component, which includes at least a convolutional layer and a channel attention mechanism layer, and the convolutional layer adopts a depth-separable convolution operation.
7. The high-precision dynamic watermarking method for streaming audio according to claim 4, characterized in that: The first neural network and the second neural network are obtained based on end-to-end joint training of an initial generation network and an initial detection network; The forward data propagation process of the end-to-end joint training includes: Based on the audio test signal, a watermark test segment with a preset number of bits is randomly generated within a preset time length range; Converting the watermark test segment into a test feature vector at the sampling point level, concatenating the test feature vector with the audio test signal and inputting the resultant vector into an initial generation network, outputting a one-dimensional test watermark signal of the same length as the audio test signal; and embedding the one-dimensional test watermark signal into the audio test signal to obtain a watermarked test signal. Data enhancement processing is performed on the watermarked test signal, the enhanced watermarked test signal is input into an initial detection network, and watermark detection confidence at a sampling point level is output, including watermark presence test information and watermark content test information; wherein the data enhancement processing includes: at least one of: channel distortion enhancement, watermark mask enhancement, and tampering simulation enhancement.
8. The high-precision dynamic watermarking method for streaming audio according to claim 7, characterized in that: The end-to-end joint training is optimized using a multi-task loss function; the multi-task loss function includes at least one of a perceptual loss function and a watermark loss function; The perceptual loss function includes at least one of the following: a depth feature loss function, a loudness domain loss function, a frequency domain loss function, and a time domain loss function; The watermark loss function includes at least one of the following: a watermark existence detection loss function, a multi-bit message detection loss function, and a boundary consistency loss function.
9. The high-precision dynamic watermarking method for streaming audio according to claim 8, characterized in that: The end-to-end joint training is optimized using a curriculum training strategy; the curriculum training strategy includes at least one of the following: dynamically adjusting the weight of the perceptual loss function, and gradually increasing the number of bits of the watermark information segment.
10. A high-precision dynamic watermarking device for streaming audio, characterized in that: The device comprises: A data acquisition module is used to acquire a streaming audio input signal and a corresponding watermark information segment; the watermark information segment is composed of bit information segments with different states and variable lengths; a sampling point level watermark generation module, configured to determine a sampling point level watermark signal according to the audio input signal and the watermark information segment; a watermark embedding module, configured to embed the sampling point-level watermark signal into the audio input signal to obtain an audio output signal, and stream output the audio output signal; The watermark extraction module is used to perform watermark detection on the sampling points of the received audio signal to be tested, obtain the watermark confidence vector for the sampling point, perform threshold judgment on the watermark confidence vector, and restore the watermark information segment.
11. A computer-readable storage medium having a computer program stored thereon, characterized in that: When the computer program is executed by a processor, the steps of the method according to any one of claims 1 to 9 are implemented.
Citation Information
Patent Citations
Blind audio watermark embedding and watermark extraction processing method
CN104795071A
Identity authentication audio watermarking algorithm based on deep learning
CN111091841A
Data processing method and device, terminal, cloud service and storage medium
CN114840824A
Homomorphic encryption audio watermark system in cloud computing environment
CN116935866A
Audio traceability device and method based on semi-fragile watermark
CN118887963A