Wavelet domain audio zero watermark method and system based on VMD-DCT

Through the DTCWT-VMD-DCT audio zero watermark algorithm, the problem of insufficient robustness of the existing technology under complex attacks is solved, and more efficient feature extraction and noise resistance is achieved, which improves the robustness of the audio zero watermark technology, especially in conventional and synchronous attacks.

CN120259064APending Publication Date: 2025-07-04QIQIHAR UNIVERSITY
View PDF 0 Cites 2 Cited by

Patent Information

Application Number
CN202510454099.0
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-04-11
Publication Date
2025-07-04

AI Technical Summary

Technical Problem

The existing audio zero watermarking technology is not robust enough in the face of complex and malicious attacks, especially under synchronous attacks, and traditional transform domain methods have limitations in feature extraction and noise immunity.

Method used

The audio zero watermark algorithm based on DTCWT-VMD-DCT is adopted to measure robustness by performing dual-tree complex wavelet transformation, variational mode decomposition and discrete cosine transformation on the audio signal, combining normalized correlation and bit error rate, and using DTCWT to capture signal characteristics in more directions and scales, VMD's adaptability and noise immunity are enhanced to enhance algorithm stability.

Benefits of technology

Improves the robustness of audio zero watermarking technology under conventional attacks, especially under noise, filtering and compression attacks, and improves in synchronous attacks.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120259064A_ABST
    Figure CN120259064A_ABST
Patent Text Reader

Abstract

The invention discloses a wavelet domain audio zero watermark algorithm based on dual-tree complex wavelet transform (DTCWT), variational mode decomposition (VMD) and discrete cosine transform (DCT). Firstly, encryption protection is carried out on a watermark image, then framing is carried out on an original audio signal, dual-tree complex wavelet transform (DTCWT) and variational mode decomposition (VMD) are carried out on each frame, a first IMF is taken to carry out DCT, and a feature vector is obtained. And finally, carrying out XOR operation on the polarity feature vector and the scrambled watermark to obtain a zero watermark key. Experiments show that the method can effectively resist conventional attacks, and especially has good robustness to noise attacks.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention belongs to the technical field of digital watermarking, and particularly relates to a zero-watermark algorithm for audio in the wavelet domain based on dual-tree complex wavelet transform (DTCWT), variational mode decomposition (VMD), and discrete cosine transform (DCT). Background Art

[0002] With the spread of multimedia information, such as audio, images, videos, etc., becoming more extensive on the Internet. Although this trend has promoted information sharing and dissemination, it has also brought problems in aspects such as information security, intellectual property protection, and identity authentication. To address these issues, researchers and industry experts have conducted extensive exploration and innovation in the field of digital watermarking technology, aiming to develop an effective way to covertly embed copyright information or identity verification data into multimedia files without disturbing the original content. Among them, zero-watermark technology has been widely adopted because it is better imperceptible.

[0003] Compared with traditional watermarking technology that realizes copyright protection by embedding hidden information in digital media, zero-watermark technology realizes content identification and authentication by examining the characteristics of the digital media content itself. Therefore, for zero-watermark, its construction area and the selection of appropriate features are particularly important, which are also important factors affecting zero-watermark algorithms. Since zero-watermark does not change the original audio, high robustness has become the key point of zero-watermark schemes.

[0004] Dragomiretskiy et al. proposed the variational mode decomposition (VMD) algorithm. Through in-depth analysis of the VMD algorithm and comparison with the EMD algorithm, it is concluded that the VMD algorithm has better mathematical theory support, feature extraction ability, and anti-noise ability than the EMD algorithm.

[0005] Nowadays, due to the simplicity and efficiency of the transform domain, most zero-watermark schemes are studied by transforming to the transform domain. The transform domain provides a simple and efficient way to capture the potential features of signals. Especially based on classical transforms such as discrete wavelet transform (DWT), singular value decomposition (SVD), and fast Fourier transform (FFT), by performing mathematical transforms on signals, unique features can be extracted in the frequency domain or wavelet domain. However, in the face of some more complex and malicious attacks (such as synchronization attacks), these traditional transform domain methods may show certain limitations. Therefore, in recent years, researchers have begun to invest more energy in finding more efficient and stable feature extraction algorithms to enhance the anti-attack ability and robustness of watermarking technology. The audio zero-watermark algorithm based on the EMD algorithm still has a large research space. Summary of the Invention

[0006] The invention provides an audio zero-watermark algorithm based on DTCWT-VMD-DCT, and the robustness is measured by the normalized correlation (NC) and the bit error rate (BER).

[0007] The object of the present invention is achieved by the following technical solutions:

[0008] A wavelet domain audio zero-watermark algorithm based on DTCWT-VMD-DCT includes the following steps:

[0009] (1) Watermark image preprocessing:

[0010] Perform Arnold scrambling on the N×N (N = 32) watermark binary image M. Then reduce the dimension of the scrambled binary watermark image. Wherein, each pixel is represented by a one-dimensional signal vector:

[0011]

[0012]

[0013] Wherein, x and y respectively represent the horizontal and vertical coordinates of the pixel point, c(i) represents each pixel point, and S represents the length of the watermark sequence, which is also the total number of pixels of the binary image.

[0014] (2) Frame division:

[0015] Divide the audio signal y evenly into S non-overlapping frames (S is the length of the watermark sequence). The length of each frame is represented by flength. Therefore, we have:

[0016]

[0017] Wherein, length(y) represents the length of the audio signal y.

[0018] (3) Embedding of the watermark:

[0019] The embedding process of the watermark is as Figure 1 shown, and the specific embedding steps are as follows:

[0020] Step 1: Perform the dual-tree complex wavelet transform (DTCWT) on the audio signal frame by frame.

[0021] Step 2: Perform variational mode decomposition (VMD) on each frame after the dual-tree complex wavelet transform to obtain all modal components IMF and the residue I r .

[0022] Step 3: For each frame, select the first modal component and perform discrete cosine transform (DCT) on it to obtain the transformed vector X i (1≤i≤S).

[0023] Step 4: Select the maximum coefficient η in X for each frame i (1 ≤ i ≤ S), that is, the highest amplitude of each frame of the signal.

[0024] Step 5: Take the highest amplitude of all frames of the audio signal and find its average value, that is, Mean(η i ).

[0025]

[0026] Step 6: Binarize according to the magnitude relationship between η of each frame i and Mean(η i ). Obtain the polarity vector E i .

[0027]

[0028] Step 7: Perform an exclusive OR operation on the watermark sequence C and the polarity vector E i to obtain the key Key. Key is expressed as:

[0029]

[0030] (4) Extraction of the watermark:

[0031] The extraction process of the watermark is as shown in Figure 1 and the specific extraction steps are as follows:

[0032] Step 1: Frame the watermarked audio signal y1 that has been attacked.

[0033] Step 2: Perform DTCWT and VMD on each frame, and perform DCT on the first modal component of each frame to obtain the transformed vector X i ´ (1 ≤ i ≤ S).

[0034] Step 3: Select the maximum amplitude η in X´ for each frame i ´ (1 ≤ i ≤ S), and then find its average value Mean(η i ´) according to the maximum amplitudes of all frames. Then form the polarity vector E i ´ through the magnitude relationship between η i ´ and Mean(η i ´). Among them, Figure 2 shows the maximum absolute amplitude coefficient values of the first 128 frames among 1024 frames under different attacks.

[0035] Step 4: Perform an exclusive OR operation on the key Key and the polarity vector E i ´ to obtain the watermark sequence C´. C´ is expressed as:

[0036]

[0037] Step 5: Perform inverse Arnold scrambling on C´ for decryption to obtain the binary watermark image.

[0038]

[0039] (5) Watermark system:

[0040] Step 1: Form a watermark image with the information for authenticating the audio ownership.

[0041] Step 2: Input the watermark image and the audio to be protected into the system, and the system performs the above operations (1)-(4).

[0042] Step 3: Save the zero-watermark key to a trusted third-party library and return the key number to the user for copyright protection.

[0043] Compared with the current technology, the present invention has the following advantages:

[0044] 1. The present invention is a zero-watermark algorithm based on DTCWT and VMD. The dual-tree complex wavelet transform is applied, which has good frequency resolution in more directions and scales and can capture the signal features more accurately.

[0045] 2. The present invention applies the VMD algorithm. VMD has self-adaptability. This self-adaptive characteristic enables VMD to effectively avoid the problems of mode mixing and over-decomposition when processing non-stationary signals, ensuring the accuracy and stability of the signal decomposition result. VMD can effectively extract the signal components in the presence of noise, so it has good anti-noise performance. Description of the Drawings

[0046] Figure 1 is the flowchart of the present invention;

[0047] Figure 2 is the schematic diagram of the maximum amplitude coefficient under different attacks;

[0048] Figure 3 is the schematic diagram of the audio directory of the Musdb18 dataset;

[0049] Figure 4 is the schematic diagram of the binary watermark image;

[0050] Figure 5 is the schematic diagram of the watermark image after Arnold transformation;

[0051] Figure 6 is the watermark image extracted after being attacked: Figure 6 (a) is without attack,Figure 6 (b) is Gaussian noise (20 db), Figure 6 (c) is Gaussian noise (10 db), Figure 6 (d) is Gaussian low-pass filtering, Figure 6 (e) is resampling, Figure 6 (f) is requantization, Figure 6 (g) is MP3 compression, Figure 6 (h) is TSM time shifting, Figure 6 (i) is cropping attack;

[0052] Figure 7 are the changing trends of four groups of audio signals with respect to the watermark based on the noise signal-to-noise ratio: Figure 7 (a) is the changing trend of the NC of the extracted watermark, Figure 7 (b) is the changing trend of the BER of the extracted watermark;

[0053] Figure 8 are the changing trends of four groups of audio signals with respect to the watermark based on the MP3 compression ratio: Figure 8 (a) is the changing trend of the NC of the extracted watermark, Figure 8 (b) is the changing trend of the BER of the extracted watermark; Detailed implementation manners

[0054] The technical solutions of the present invention will be further described below in conjunction with the accompanying drawings, but are not limited thereto. Any modification or equivalent replacement of the technical solutions of the present invention without departing from the spirit and scope of the technical solutions of the present invention shall be covered by the protection scope of the present invention.

[0055] The present invention proposes a wavelet domain audio zero watermark algorithm and system based on DTCWT-VMD-DCT. The system includes a memory, a processor, and a computer program stored on the memory. The computer program is executed by the processor to achieve copyright protection. Before watermark embedding, the binary watermark image is encrypted by Arnold to ensure the security of the watermark. Then it is converted into a one-dimensional vector so that the watermark image can adapt to the subsequent feature extraction and embedding processes. The audio signal is framed, and DTCWT-VMD is performed frame by frame. The first intrinsic mode component is taken for DCT transformation, and its maximum value is selected. The average value is calculated according to the maximum value, and a key is generated through the relationship between the two. Finally, the key is stored in a reliable third-party library.

[0056] To implement the above algorithm, the present invention is divided into the following four steps:

[0057] (1) Watermark image preprocessing:

[0058] Perform Arnold scrambling on the N×N (N = 32) binary watermark image M. Then, reduce the dimension of the scrambled binary watermark image. Among them, each pixel is represented by a one-dimensional signal vector:

[0059]

[0060]

[0061] Among them, x and y respectively represent the horizontal and vertical coordinates of the pixel point, c(i) represents each pixel point, S represents the length of the watermark sequence, which is also the total number of pixels in the binary image.

[0062] (2) Frame division:

[0063] Evenly divide the audio signal y into S non-overlapping frames (S is the length of the watermark sequence). The length of each frame is represented by flength. Therefore, we have

[0064] .

[0065] Among them, length(y) represents the length of the audio signal y.

[0066] (3) Embedding of the watermark:

[0067] Step 1: Perform the dual-tree complex wavelet transform (DTCWT) on the original signal frame by frame.

[0068] Step 2: Perform variational mode decomposition (VMD) on each frame after the dual-tree complex wavelet transform, set the penalty factor α = 2000, and the number of decomposed modes k = 10 to obtain all the mode components IMF and the residue I of each frame r .

[0069] Step 3: For each frame, select the first mode component and perform discrete cosine transform (DCT) on it to obtain the transformed vector X i (1 ≤ i ≤ S).

[0070] Step 4: Select the maximum coefficient η in X of each frame i (1 ≤ i ≤ S), that is, the highest amplitude of each frame of the signal.

[0071] Step 5: Take the highest amplitudes of all frames of the audio signal and find their average value, that is, Mean(η i ).

[0072]

[0073] Step 6: According to η of each frame i and Mean(η i) Binaryize based on the size relationship between the two to obtain the polarity vector E i :

[0074]

[0075] Step Seven: XOR the watermark sequence C with the polarity vector E i to obtain the key Key. Key is expressed as:

[0076]

[0077] (4) Watermark extraction:

[0078] Step One: Frame the watermarked audio signal y1 that has been attacked

[0079] Step Two: Perform DTCWT and VMD on each frame, and perform DCT on the first modal component of each frame to obtain the transformed vector X i ´ (1 ≤ i ≤ S).

[0080] Step Three: Select the maximum amplitude η i ´ (1 ≤ i ≤ S) in X´ of each frame, and then calculate its average value Mean(η i ´) based on the maximum amplitudes of all frames. Then, form the polarity vector E i ´ through the size relationship between η i ´ and Mean(η i ´).

[0081] Step Four: XOR the key Key with the polarity vector E i ´ to obtain the watermark sequence C´. C´ is expressed as:

[0082]

[0083] Step Five: Decrypt C´ by performing inverse Arnold scrambling to obtain the binary watermark image

[0084]

[0085] (5) Watermark system:

[0086] Step One: Form a watermark image with the information that can authenticate the audio ownership

[0087] Step Two: Input the watermark image and the audio to be protected into the system, and the system performs the above operations (1)-(4).

[0088] Step Three: Save the zero-watermark key to a trusted third-party library and return the key number to the user for copyright protection

[0089] The present invention utilizes the DTCWT transform, which can not only capture the local features of a signal in more directions and scales, but also has higher computational efficiency and can better remove redundant information in the signal. The VMD algorithm not only has high adaptability but also has stronger anti-noise ability. The wavelet domain audio zero-watermarking algorithm based on DTCWT, VMD, and DCT has a mathematical optimization framework, has high practical value in the field of information security, and this algorithm can be applied to other fields and has wide applicability.

[0090] The following will be described from the theoretical basis and experimental data:

[0091] 1. Dual-Tree Complex Wavelet Transform (DTCWT)

[0092] The Dual-Tree Complex Wavelet Transform (DTCWT) is a complex wavelet transform method. Since the sampling frequencies of the filters of the two trees are the same and the delay between them is exactly one sampling interval, it has good translational invariance while reducing the computational complexity and making it easier to implement. Compared with the traditional wavelet transform, DTCWT has more directions to select subbands. Because DTCWT adopts an even-odd decomposition structure, it has translational invariance and also reduces the loss caused by the decomposition and reconstruction of the image, making it have better robustness.

[0093] The one-dimensional DTCWT can be expressed as:

[0094]

[0095] where and represent the wavelets corresponding to Tree A and Tree B.

[0096] 2. Variational Mode Decomposition (VMD)

[0097] Variational Mode Decomposition (VMD) is a signal processing method widely used in the fields of signal processing, image processing, etc. Its core idea is to find a set of basis functions through an optimization problem, so that the signal can be decomposed into a series of frequency-localized modes (IMFs). These modes reflect the components of the signal in different frequency bands, which helps to analyze the time-frequency characteristics of the signal more accurately. Compared with Empirical Mode Decomposition (EMD), the latter is prone to over-decomposition problems during the decomposition process, that is, the number of modes is too large, resulting in overlapping between modes, thereby reducing the decomposition accuracy. While VMD effectively suppresses this mode overlapping phenomenon by determining the number of decomposed modes and controlling the bandwidth. In addition, compared with Empirical Mode Decomposition, VMD has a more solid theoretical basis and can be better applied to non-stationary signals.

[0098] Assume that the original signal f is decomposed into k modal components. To minimize the sum of the estimated bandwidths of each modal component, a corresponding constraint condition is determined, and its expression is:

[0099]

[0100]

[0101] where: K is the number of decomposed modes, and the Kth modal component and the center frequency after decomposition correspond to {u k} and {ω k} respectively.

[0102] To better solve the above constraint condition, the Lagrange multiplier operator λ is introduced to make the constraint condition become part of the objective function, and the optimal solution is obtained by minimizing the augmented Lagrangian function. The augmented Lagrangian expression is:

[0103]

[0104] where α is the quadratic penalty factor.

[0105] Finally, the Alternating Direction Method of Multipliers (ADMM) is used for iterative optimization.

[0106] 1. Discrete Cosine Transform (DCT)

[0107] Discrete Cosine Transform (DCT) is a signal processing technology mainly used for time-frequency conversion. DCT decomposes the input signal into a linear combination of a set of cosine functions, and each cosine function corresponds to different frequency components of the signal. Different from the Fourier transform, the frequency components generated by DCT are real values instead of complex values.

[0108] The definitions of the one-dimensional discrete cosine transform and its inverse transform are as follows:

[0109]

[0110]

[0111] Among them, (k) is:

[0112]

[0113] Example:

[0114] In this experiment, audio segments of several styles were selected from 150 pieces of music in the MUSDB18 dataset. The music sampling rate was 44.1 kHz, and the duration was 2 - 3 minutes. The music directory of the dataset is as Figure 3 shown. A 32×32 watermark image was used in this experiment. Since the zero-watermark scheme does not modify the original data, it has good invisibility. Therefore, the experiment focused on evaluating the robustness of the watermark scheme proposed in the present invention. The experiment tested the robustness of the scheme by applying various attack means, such as noise attack, filtering attack, and compression attack. We used two metrics, the bit error rate (BER) and the normalized cross-correlation coefficient (NC), to measure the robustness performance of the watermark scheme. Next, the performance of the scheme of the present invention in terms of robustness was further evaluated through specific experimental data.

[0115] Next, the robustness in the present invention was judged through specific experimental data.

[0116] By performing different attacks on the watermarked audio signal, the watermark image extracted by the watermark scheme of the present invention was compared with its original watermark image, and the robustness of the algorithm was verified through the bit error rate (BER) and the normalized cross-correlation coefficient (NC). During this process, 8 different types of attacks were applied, and the specific parameters of each attack are shown in Table 1.

[0117] Table 1 Explanation of attack types and parameters

[0118]

[0119] As Figure 4 shown is the original watermark image. The original watermark was scrambled by the Arnold transformation k times to obtain the watermark image after Arnold scrambling as shown in Figure 5 shown. If the non-k times inverse Arnold transformation scrambling is used, the watermark information cannot be correctly restored.

[0120] Table 2 Performance of audio signals under attacks

[0121]

[0122] Table 2 shows the NC values and BER values obtained from randomly selecting 2 audio signals from the MUSDB18 dataset and subjecting them to the above 8 attacks. Figure 6 This is the result graph of the watermark image extracted by Mixture59 of this scheme. The specific results are shown in Table 2.

[0123] It can be seen from the experimental data shown in Table 2 that when the audio signal with the watermark image is subjected to 8 attacks, it has good ability to resist conventional attacks. However, its performance in resisting synchronization attacks needs to be improved. Figure 7 (a) and Figure 7 (b) show the changing trends of the NC values and BER values of the watermark extracted from four groups of audio signals with the change of the noise signal-to-noise ratio. Figure 8 (a) and Figure 8 (b) show the changing trends of the NC values and BER values of the watermark extracted from four groups of audio signals with the change of the MP3 compression ratio. The experimental results show that when facing resampling, requantization, Gaussian noise, and MP3 compression attacks, this scheme has high robustness.

[0124] Next, the robustness of the algorithm mentioned in the present invention is further tested, and conclusions are drawn through comparative experiments. The NC values and BER values obtained after subjecting the algorithm of the present invention and the existing algorithms to the above 8 attacks are compared. As shown in Table 3.

[0125] It can be seen from the comparison results in Table 3 that this scheme is superior to the other three schemes in resisting conventional attacks. And when resisting synchronization attacks, the robustness of this scheme also achieves good results compared with the other three schemes.

[0126] Table 3 Comparative experiments

[0127]

Claims

1. A wavelet-domain audio zero-watermarking algorithm based on DTCWT-VMD-DCT, comprising the following steps: (1) Preprocessing of the watermark image: Perform Arnold scrambling on the N × N (N = 32) binary watermark image M. Then reduce the dimension of the scrambled binary watermark image. Among them, Each pixel is represented by a one-dimensional signal vector: ; ; Where x and y respectively represent the horizontal and vertical coordinates of the pixel point, c(i) represents each pixel point, S represents the length of the watermark sequence and also the total number of pixels of the binary image; (2) Frame division: The audio signal y is evenly divided into S non-overlapping frames (S is the length of the watermark sequence); The length of each frame is represented by flength. Therefore, we have: ; Where length(y) represents the length of the audio signal y; (3) Embedding of the watermark: Step 1: Perform the dual-tree complex wavelet transform (DTCWT) on the audio signal frame by frame; Step 2: Perform variational mode decomposition (VMD) on each frame after double-tree complex wavelet transform, set the penalty factor α = 2000, and the number of decomposed modes k = 10 to obtain all mode components IMF and the residue I of each frame r; Step 3: For each frame, select the first modal component and perform discrete cosine transform (DCT) on it to obtain the transformed vector X i where (1 ≤ i ≤ S); Step 4: Select the maximum coefficient η in X for each frame i (1 ≤ i ≤ S), which is the highest amplitude of each frame of signal; Step 5: Take the maximum amplitude of all frames of the audio signal and find its average value, i.e., Mean(η i ); ; Step 6: According to the magnitude relationship between η of each frame i and Mean(η i ), perform binarization to obtain the polarity vector E i; ; Step 7: Exclusive-OR the watermark sequence C with the polarity vector E i to obtain the key Key. Key is expressed as: ; (4) Extraction of the watermark: Step 1: Divide the watermarked audio signal y1 that has been attacked into frames; Step 2: Perform DTCWT and VMD on each frame, and perform DCT on the first modal component of each frame to obtain the transformed vector X i ´ (1 ≤ i ≤ S); Step 3: Select the maximum amplitude η in X´ for each frame i ´ (1 ≤ i ≤ S), and then calculate the average value Mean(η i ´) based on the maximum amplitudes of all frames. Then, form a polarity vector E i ´ according to the magnitude relationship between η i ´ and Mean(η i ´); Step 4: Perform an exclusive OR operation on the key Key and the polarity vector E i ´ to obtain the watermark sequence C´. C´ is expressed as: ; Step 5: Perform inverse Arnold scrambling on C´ for decryption to obtain the binary watermark image; ; (5) Watermark system: Step 1: Form a watermark image with the information that can authenticate the audio ownership; Step 2: Input the watermark image and the audio to be protected into the system, and the system performs operations from Step 1 to Step 4; Step 3: Save the zero-watermark key to a trusted third-party library and return the key number to the user for copyright protection.

2. The wavelet domain audio zero-watermarking algorithm based on DTCWT-VMD-DCT according to claim 1. Its main feature lies in that during the watermark embedding process in step three, the penalty factor α of variational mode decomposition is set to 2000, and the number of decomposed modes k is 10, obtaining all mode components IMF and the residue I of each frame r .

Citation Information

Cited By

  • Self-adaptive digital watermark embedding method and device and readable storage medium

    CN122115186A

  • An adaptive digital watermark embedding method, device and readable storage medium

    CN122115186B