Watermark embedding method based on wavelet domain statistical features
By combining wavelet domain statistical features and deep neural networks, mid-to-high frequency sub-bands are selected as embedding regions, and the covariance matrix is calculated to generate weights. This achieves a balance between the invisibility and robustness of the watermark, solves the problems of contradictory embedding region selection and insufficient anti-attack capability in existing watermarking technologies, and improves the concealment and anti-attack capability of the watermark.
Patent Information
- Application Number
- CN202510787705.0
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-06-13
- Publication Date
- 2025-12-16
- Estimated Expiration
- 2045-06-13
AI Technical Summary
Existing watermarking technologies in data transactions suffer from several drawbacks: an imbalance between invisibility and robustness due to contradictions in embedding region selection; insufficient resistance to attacks caused by fixed embedding strength; and the defect that detection relies on the original carrier data.
An invisible watermark embedding method based on wavelet domain statistical features is adopted. The mid-to-high frequency sub-bands are selected as the embedding region through wavelet multi-level decomposition, the covariance matrix is calculated to generate weights, and a deep neural network is used for adaptive embedding of the invisible watermark. The robustness and concealment are improved by combining the chaotic encryption algorithm.
While ensuring invisibility, it improves the robustness of watermarks against common attacks such as compression and noise, dynamically adjusts the embedding intensity of different texture regions, enhances robustness in smooth regions, suppresses distortion in edge regions, and has strong resistance to reverse analysis.
Smart Images

Figure CN120672553B_ABST
Abstract
Description
TECHNICAL FIELD
[0001] The present application relates to the field of watermark embedding, in particular to a kind of invisible watermark embedding method based on wavelet domain statistical characteristics. BACKGROUND
[0002] In the new economic form of data elements becoming core assets, digital authentication is the key link to protect the value and transaction order of data circulation. As an implicit identification means, digital watermark technology can embed copyright information imperceptibly into data content, enabling ownership tracking and infringement evidence in cross-platform, multi-link circulation. However, with the expansion of data transaction scale and the upgrading of attack means, existing watermark technology faces serious challenges: data undergoes format conversion, re-compression, local editing and other operations in multiple transactions, which easily leads to loss or damage of watermark information; attackers use public transaction data to reverse analyze watermark rules and then tamper with or forge copyright marks in bulk, seriously threatening the credibility of data transactions; At the same time, privacy protection regulations require transaction platforms to desensitize user identity, timestamp and other metadata, and traditional watermark technology cannot dynamically separate copyright marks from sensitive information, increasing compliance risks.
[0003] The current mainstream technology generally has the problem of imbalance between robustness, concealment and functionality. Spatial domain methods directly modify pixel values, although the algorithm is simple, but the resistance to regular signal processing attacks (such as compression, filtering) is weak; frequency domain methods disperse watermark energy by adjusting transform domain coefficients, but high-frequency embedding is easily disturbed and destroyed by noise, low-frequency modification leads to visual artifacts, and the vulnerability to geometric deformation makes it difficult to adapt to multi-platform transmission scenarios. More importantly, existing frequency domain methods usually rely on mathematical tools such as Fourier transform and discrete cosine transform to map data from spatial domain to frequency domain space, and achieve watermark embedding by adjusting specific frequency band coefficients. This strategy of dispersing watermark energy can theoretically improve the anti-detection ability of the watermark, but in actual application it exposes significant drawbacks. High-frequency coefficients, as an area with low sensitivity of the human sensory system, can achieve high concealment, but this frequency band is easily disturbed by channel noise, lossy compression and other regular signal processing operations. SUMMARY
[0004] In view of the deficiencies of the prior art, the present application proposes a kind of invisible watermark embedding method based on wavelet domain statistical characteristics, to solve the imbalance between invisibility and robustness caused by the contradiction of embedding area selection in existing watermark technology, the lack of anti-attack ability caused by fixed embedding strength and the defect of detection dependence on original carrier data.
[0005] In order to achieve the above purpose, the technical scheme adopted by the present application is as follows:
[0006] The invisible watermark embedding method based on wavelet domain statistical characteristics comprises the following steps:
[0007] S1, wavelet multi-level decomposition is performed on an input image, and a middle-high frequency subband is selected as a watermark embedding region;
[0008] S2, the watermark embedding region is evenly divided into a plurality of subband blocks with a size of , a covariance matrix of each subband block is calculated, and a weight is generated according to an eigenvalue of the covariance matrix;
[0009] S3, using a deep neural network, watermark information is encoded into a frequency domain disturbance, and the weight obtained in step S2 is combined to realize adaptive embedding of an invisible watermark.
[0010] Further, the wavelet multi-level decomposition of the input image in step S1 is specifically: the original image is decomposed into a low frequency subband (LL3), a middle-high frequency subband (HL3 and LH3), and a high frequency subband (HH3) through three-level discrete wavelet transform (DWT). The low frequency subband contains the main information of the image, and modification is easy to cause visual distortion; the high frequency subband is sensitive to noise and has poor robustness. Therefore, the middle-high frequency subband (HL3 and LH3) is selected as the watermark embedding region, which can balance invisibility and robustness.
[0011] Further, the size of the subband block is , which meets the requirement of human eye insensitivity to high frequency noise in the size area and meets the requirement of invisibility.
[0012] Further, the covariance matrix of each subband block in step S2 is specifically:
[0013]
[0014]
[0015] wherein, is the i-th subband block, and are the indexes of the pixels in the width and height of the subband block, and represent the mean and covariance matrix of the i-th subband block, respectively. The mean measures the average value of the pixels in the block, which is used for data centralization, and the covariance describes the correlation between the pixels in the block, which reflects the local texture features. Further, the weight generated according to the eigenvalue of the covariance matrix in step S2 is specifically:
[0016]
[0017]
[0018] wherein, is the weight of the th subband block, and is the eigenvalue of . If , it represents a strong edge region, and the weight is small to suppress distortion; if , it represents a smooth region, and the weight is large to enhance robustness.
[0019] Further, the step S3 of encoding the watermark information into the frequency domain perturbation using the deep neural network comprises:
[0020] determining the input of the deep neural network:
[0021]
[0022] wherein, represents the concatenation operation; is the feature after concatenating the th subband and the th subband in the mid-high frequency subband, and has a size of ; is the feature after up-sampling the original watermark, and has a size of ;
[0023] the deep neural network is a U-shaped network comprising an input layer, 5 layers of encoders, 4 layers of decoders, and an output layer;
[0024] the output of the deep neural network is the frequency domain perturbation , which is limited to by a TanH function to prevent the perturbation from being too large.
[0025] Further, for and , concatenation is performed along the channel dimension, and the calculation of the is as follows:
[0026]
[0027] wherein, is the channel index, , .
[0028] Further, the calculation of the is as follows:
[0029] using a chaotic encryption algorithm, the original watermark Convert to a sequence of real numbers Then, two-dimensional interpolation was used to... Perform upsampling to obtain the upsampled features. ;
[0030] in, .
[0031] Furthermore, in step S3, the invisible watermark adaptive embedding is achieved by combining the weights obtained in step S2, specifically as follows:
[0032]
[0033] in, It's a sub-band with a watermark added. It is the global embedding strength, which controls the overall strength of the watermark signal and avoids differences in optimal strength between different images. It is optimized through backpropagation. It is the first The weight of each block.
[0034] Furthermore, the loss function used for training the deep neural network is:
[0035]
[0036] in, , and It is the coefficient of loss;
[0037] As for the invisibility loss, to ensure the invisibility of the watermark, the original image is calculated. With watermarked images The mean squared error (MSE) constrains visual differences.
[0038]
[0039] in, It is the original image. It is a watermarked image. This represents the square of the L2 norm, which is the sum of squares of the pixel-by-pixel differences;
[0040] For robustness loss, calculate and The negative cosine similarity is used to minimize the similarity between them.
[0041]
[0042] in, for and cosine similarity, is the subband after embedding the watermark, is the original image Add Gaussian noise attack after frequency domain decomposition of the subband;
[0043] Statistical alignment loss is for statistical alignment loss, to ensure the rationality of the disturbance embedded, the disturbance intensity The matching between the subband block weight Ensure that the high weight area (smooth area) embeds stronger disturbance.
[0044]
[0045] Wherein, Indicates the L2 norm of the frequency domain disturbance .
[0046] Compared with the prior art, the beneficial effects of the present application are: the present application selects the high-frequency subband in the wavelet domain as the embedding area, which improves the robustness of the watermark to conventional attacks such as compression and noise while ensuring invisibility. The adaptive embedding weight mechanism based on the eigenvalue of the covariance matrix can dynamically adjust the embedding strength of different texture regions, enhance the robustness of smooth regions, and suppress distortion in edge regions. The nonlinear disturbance generated by the deep neural network and combined with the chaotic encryption preprocessing breaks the predictability of the linear statistical rule. BRIEF DESCRIPTION OF DRAWINGS
[0047] Figure 1 is the flowchart of the invisible watermark embedding method based on wavelet domain statistical characteristics of the present application. DETAILED DESCRIPTION
[0048] The technical solutions of the present application will be further described in detail below in combination with the drawings.
[0049] As Figure 1 shown, the embodiment of the present application provides an invisible watermark embedding method based on wavelet domain statistical characteristics, which includes the following steps:
[0050] S1, after inputting a 512x512 pixel RGB image, perform three-level discrete wavelet transform (DWT). Specifically, first, use a 3-level DWT operator to decompose the input image:
[0051]
[0052] Wherein is a three-level DWT operator, and Daubechies 8 wavelet basis is used, is the input image, and the size is . For the image , and These are the width and height of the image. express It is a three-channel color image.
[0053] Generating low-frequency subbands ( ), mid-to-high frequency sub-band , (each ), high-frequency subband ( ).
[0054] Then select and The sub-band serves as the watermark embedding area, avoiding low-frequency distortion and high-frequency noise sensitivity issues.
[0055] S2, Covariance weight calculation, firstly... and Subbands are divided into Pixel blocks, each sub-band generates a total of There are several blocks. Then, the mean vector of each block is calculated. The three-channel values of the 64 pixels within the block are averaged. The covariance matrix is then calculated. Based on the centered pixel value, the covariance is calculated per channel to reflect the texture correlation within the block. Finally, the covariance matrix is decomposed into eigenvalues to obtain the principal eigenvalues. and secondary eigenvalues The weights are generated according to the following formula w:
[0056]
[0057] S3. The watermark is embedded using a deep neural network. The structure of the deep neural network (DNN) is shown in the table below:
[0058] Table 1 Deep Neural Network Architecture
[0059]
[0060] Before inputting into the DNN, the 256-bit binary watermark sequence is encrypted using a Logistic chaotic mapping to generate a real number sequence. The watermark sequence is generated randomly. Then... Perform bilinear interpolation upsampling to generate and and Frequency domain watermarking features with subband size matching Then the original subband features were compared with... The input network is spliced together to generate nonlinear frequency domain perturbations. And embed the watermark using the following formula:
[0061] where global intensity It is initialized to 0.1 by backpropagation optimization.
[0062] The DNN model training is constrained by three losses, invisibility loss The original image is calculated The mean square error of the watermarked image is constrained to constrain the pixel-level difference. Robustness loss The cosine similarity of The Gaussian noise attack with a mean of 0.1 is applied, and the negative value of the cosine similarity of the subband before and after the attack is minimized. Alignment loss The L2 norm of the disturbance is forced to match the weight , ensuring the effectiveness of the weight mechanism. The total loss is as follows:
[0063]
[0064] The coefficients of different losses , , are set to 1.0, 0.5, and 0.2, respectively.
[0065] During training, the AdamW optimizer is used in the model training stage, the initial learning rate is set to 1e-4, and the weight decay is set to 1e-5 for parameter optimization. The batch size is 16, and the training is performed for 2000 epochs. The method is evaluated on the COCO dataset.
[0066] In terms of robustness, the embodiment is constrained by robustness loss, and after applying Gaussian noise attack to the image embedded with watermark, the cosine similarity of the embedded subband and the attacked subband can be kept at a high level (about 0.9), while the cosine similarity of the traditional frequency domain method is often lower than 0.7 after Gaussian noise attack.
[0067] In terms of invisibility, the embodiment constrains the mean square error of the original image and the watermarked image by invisibility loss, so that the pixel-level difference between the two is small, and the corresponding peak signal-to-noise ratio is greater than 30dB, close to the quality of the original image, while the traditional frequency domain method (such as DCT low-frequency embedding) has obvious visual artifacts.
[0068] In terms of adaptive ability, the embodiment generates weights based on the eigenvalues of the covariance matrix, with larger weights in smooth areas to allow stronger disturbance to enhance robustness, and smaller weights in strong edge areas to suppress disturbance to avoid distortion.
[0069] In terms of anti-reverse analysis, the embodiment converts the binary watermark into a real number sequence through chaotic encryption, and generates a nonlinear disturbance using a DNN, breaking the statistical rules of traditional linear embedding, so that the success rate of Gaussian noise attack is less than 30%, while the success rate of traditional linear method attack is greater than 50%.
[0070] Finally, it should be noted that: the above embodiments are intended to illustrate the technical solutions of the present application and do not constitute any form of limitation on the present application. Those skilled in the art should fully understand that it is entirely feasible to modify the technical solutions described in the foregoing embodiments or to equivalently replace any part or all of the technical features thereof. These modifications or replacements should be considered as reasonable extensions of the present application, as long as they do not deviate from the protection scope determined by the claims of the present application.
Claims
1. An invisible watermark embedding method based on wavelet domain statistical features, characterized in that, Includes the following steps: S1, perform wavelet multi-level decomposition on the input image, and select the mid-to-high frequency sub-band as the watermark embedding region; S2, divide the watermark embedding area into multiple equal parts of size. For each sub-band, calculate the covariance matrix of each sub-band, and then generate weights based on the eigenvalues of the covariance matrix. The calculation of the covariance matrix for each sub-band block is specifically as follows: in, It is the first He has a block on his body. and It is the pixel index of the sub-band block in width and height. and They represent the first The mean and covariance matrix of each sub-band block; The process of generating weights based on the eigenvalues of the covariance matrix is as follows: in, It is the first The weight of each block, and yes eigenvalues; S3. Use a deep neural network to encode the watermark information as a frequency domain perturbation, and combine it with the weights obtained in step S2 to achieve adaptive embedding of the invisible watermark.
2. The method according to claim 1, characterized in that, Step S1 describes performing wavelet multi-level decomposition on the input image, specifically by decomposing the original image into low-frequency sub-band, mid-to-high-frequency sub-band, and high-frequency sub-band through a 3-level discrete wavelet transform.
3. The method according to claim 1, characterized in that, The size of the sub-band block is It meets the requirement that the human eye is insensitive to high-frequency noise within this size area, thus satisfying the requirement of invisibility.
4. The method according to claim 1, characterized in that, Step S3, which involves encoding the watermark information into a frequency domain perturbation using a deep neural network, includes: Determine the input of a deep neural network : in, Indicates a splicing operation; To make the mid-to-high frequency sub-band Sub-band and The feature after sub-band splicing, its size is ; The feature is a sampled feature of the original watermark, and its size is [size missing]. ; The deep neural network is a U-shaped network containing an input layer, 5 encoder layers, 4 decoder layers, and an output layer; The output of the deep neural network is a frequency domain perturbation. .
5. The method according to claim 4, characterized in that, The The calculation is as follows: in, It is a channel index. , .
6. The method according to claim 4, characterized in that, The The calculation is as follows: Using a chaotic encryption algorithm, the original watermark, defined as a binary sequence, is... Convert to a sequence of real numbers Then, two-dimensional interpolation was used to... Perform upsampling to obtain the upsampled features. ; in, .
7. The method according to claim 4, characterized in that, Step S3, the invisible watermark adaptive embedding is achieved by combining the weights obtained in step S2, specifically as follows: in, It is a sub-band block with a watermark added. It is the global embedding strength. It is the first The weight of each block.
8. The method according to claim 4, characterized in that, The loss function used for training the deep neural network is: in, , and It is the coefficient of loss; in, It is the original image. It is a watermarked image. Represents the square of the L2 norm; in, for and cosine similarity, It is a sub-band after the watermark is embedded. This is the original image. Subbands after frequency domain decomposition following the addition of Gaussian noise attack; in, Indicates frequency domain perturbation The L2 norm.
Citation Information
Patent Citations
Adaptive robust watermark embedding method and system based on deep neural network
CN114549273A
Image high-capacity robust watermarking method based on wavelet neural network
CN116029887A