Screen-shooting-resistant robust image watermark soft fusion network method and system based on UNet architecture
By introducing discrete wavelet transformation and self-attention mechanisms into the UNet architecture, the joint embedding of the airspace and frequency domain is realized, and the model robustness is improved through non-micronoise simulation training, which solves the shortcomings of the existing technology in anti-screen shooting and noise interference, and significantly improves the detectability and noise resistance of watermarks.
Patent Information
- Application Number
- CN202510625134.0
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-05-15
- Publication Date
- 2025-06-13
- Estimated Expiration
- 2045-05-15
AI Technical Summary
The prior art has shortcomings in terms of anti-screen camera robustness and noise interference, especially when dealing with complex textures and high-frequency components, the watermark embedding position deviates from the significant area of the image after screen camera, affecting robustness.
The anti-screen-resistant robust image watermark soft fusion network method based on UNet architecture is adopted, and the joint embedding of the airspace and frequency domain is realized through discrete wavelet transformation downsampling, cover self-attention CSA and channel compression attention SE modules, and a non-micronoise simulation training process is designed to improve the robustness of the model.
It significantly improves the detectability and noise resistance of watermarks under screen shooting conditions, maintains a low bit error rate and a high peak signal-to-noise ratio PSNR and structural similarity SSIM, demonstrating excellent screen shooting resistance.
Smart Images

Figure CN120147099A_ABST
Abstract
Description
Technical Field
[0001] The present invention belongs to the field of information security, and particularly relates to a robust image watermark soft fusion network method and system based on the UNet architecture. Background Art
[0002] With the acceleration of the global enterprise digital transformation process, the frequency of cross-domain display of business secrets through digital conference systems has increased significantly, and the problem of screen capture (screen shooting) leakage has become increasingly severe. According to Verizon's "2023 Data Breach Investigations Report", 11% of physical security incidents stem from illegal secret stealing behavior of taking pictures of the screen with mobile devices. Against this background, robust watermark technology, with its powerful tracking and tracing capabilities, has become a key technical means to combat screen capture leakage and protect copyrights. By embedding encrypted identification information, this technology enables the watermark information to remain completely extractable even after the presentation document has gone through complex processes such as screen capture and social media dissemination, thereby building an active defense system of "once leaked, traceable".
[0003] Many advances have been made in the research on robust watermarks. Traditional methods mainly rely on spatial domain and transform domain techniques. Among spatial domain methods, the least significant bit (LSB) modification algorithm has received wide attention due to its high visual imperceptibility, but it has obvious shortcomings in resisting screen capture attacks and noise interference. For transform domain-based methods, such as singular value decomposition (SVD), discrete wavelet transform (DWT), and discrete cosine transform (DCT), the statistical characteristics of transform coefficients are adjusted to improve the concealment and anti-interference ability of watermarks. However, these methods rely on manual feature extraction and are difficult to adaptively capture the main energy distribution of images. Especially when dealing with complex textures and high-frequency components, their feature representation capabilities have significant limitations, resulting in a deviation between the watermark embedding position and the significant regions of the image after screen capture, thus affecting robustness.
[0004] With the development of deep learning, the success of convolutional neural networks (CNNs) in the field of computer vision has promoted the transformation of watermark technology. Deep learning methods can, through an autonomous feature learning mechanism, explore the internal representation laws of data and construct multi-scale feature hierarchies, thereby enhancing the screen capture resistance of watermark systems. However, current deep watermark methods still face many challenges, such as excessive redundancy of watermark information, insufficient feature interaction between the watermark and the carrier image resulting in weakened semantic association, and the impact of non-differentiable noise and image distortion during screen capture on the extractability of watermarks.
[0005] Currently, the mainstream end-to-end watermarking frameworks have limitations in watermark redundancy control and feature fusion. Many methods adopt a mechanical watermark replication strategy to achieve the global distribution of watermarks, but lack adaptive redundancy control, resulting in poor integration of watermark features and image features, and affecting the robustness against screen capture. At the feature fusion level, existing methods mostly adopt a shallow fusion strategy of direct channel splicing, making it difficult to establish an effective association between watermark semantic features and image structure features, and failing to dynamically adjust watermark features according to the characteristics of image content, thus weakening the ability to resist screen capture attacks.
[0006] In addition, existing frameworks rely too much on spatial domain feature operations and ignore the advantages of frequency domain features. In contrast, frequency domain methods utilize the energy compression characteristics and the sensitivity of the human visual system, and show stronger robustness against attacks such as compression, blurring, and color distortion caused by screen capture. In terms of anti-noise interference, although adding a noise layer can improve the anti-interference ability of the encoded image, the DCT and quantization operations of JPEG compression in social media or the shooting device itself will block the reverse propagation chain of the neural network, restricting the encoder's optimization of the decoding loss. Although traditional analog training methods approximate the JPEG compression effect through differentiable DCT to maintain gradient conduction, there are still performance bottlenecks in high compression ratio scenarios. Summary of the Invention
[0007] Object of the Invention: The technical problem to be solved by the present invention is to provide a method and system for a screen capture-robust image watermark soft fusion network based on the UNet architecture in view of the deficiencies of the prior art.
[0008] The method includes: Step 1, preprocess the cover image to obtain the low-frequency component; Step 2, establish a multi-stage watermark extension sub-network to perform feature fusion on the watermark information tensor and the low-frequency component; Step 3, construct a watermark soft fusion network ProwterNet based on the Unet architecture: add discrete wavelet transform downsampling, cover self-attention CSA, and channel squeeze attention SE modules in the Unet network, and cascade message processors; first, replace all downsampling modules in the encoding stage of the Unet network with discrete wavelet transform downsampling, and add the channel squeeze attention SE after upsampling; secondly, add the cover self-attention CSA before the skip connection in the decoding stage; finally, receive the output of the cascade message processor, and the output includes the frequency domain information of the embedded watermark features and the watermark message, and fuse the frequency domain information of the embedded watermark features and the watermark message with the features of the skip connection and the features after continuous upsampling to achieve joint embedding in the spatial domain and the frequency domain; Step 4, design a non-differentiable noise simulation training process to train the watermark soft fusion network ProwterNet.
[0009] Step 1 includes: Step 1-1: Input the cover image into the discrete wavelet transform network, and decompose it into a first-level low-frequency component through the first-level wavelet transform , a first-level horizontal high-frequency component , a first-level vertical high-frequency component , and a first-level diagonal high-frequency component ; Step 1-2: Input the low-frequency component into the discrete wavelet transform network again for secondary decomposition to obtain a second-level low-frequency component , a second-level horizontal high-frequency component , a second-level vertical high-frequency component , a second-level vertical high-frequency component ; Step 1-3: Input the second-level low-frequency component into the discrete wavelet transform network to extract a third-level low-frequency component .
[0010] Step 2 includes: Step 2-1: Randomly generate watermark information of 64 bits , and expand the randomly generated watermark information into a watermark information tensor of through a fully connected layer ; Step 2-2: Duplicate the expanded watermark information tensor 4 times, denoted as , where ; Step 2-3: In the 0th layer message processor , receive the expanded watermark information tensor , without performing additional processing, and directly output the watermark information tensor ; Step 2-4: Concatenate the three low-frequency components , , obtained in Step 1 into the first layer message processor , the second layer message processor and the third layer message processor respectively; Step 2-5: The first layer message processor receives two inputs: the expanded watermark information tensor , and the first-level low-frequency component ; Step 2-6: The first layer message processor firstly takes the watermark information tensor Pass through the convolutional neural network to the watermark information tensor of the same size as the first-level low-frequency component ; then perform Hadamard product calculation on the watermark information tensor and the first-level low-frequency component , and generate the attention weight coefficient through the function, and use the attention weight to perform feature fusion on the watermark information tensor and the first-level low-frequency component and output the frequency-domain information feature with watermark characteristics after fusion ; ; Step 2-7, in the second-layer message processor and the third-layer message processor receive the extended watermark information tensor , the second-level low-frequency component and the extended watermark information tensor , the third-level low-frequency component , and repeat Step 2-6 to output the frequency-domain information feature with watermark characteristics after fusion and , and the formula for the specific fusion process is: , where represents 's convolutional neural network, is function, is the product of Hadamard product calculation; z takes values of 1, 2, 3; The message processor has a total of four layers. The 0th layer outputs the watermark information tensor , and the first layer processes the watermark information tensor , and finally outputs the frequency-domain information feature with watermark characteristics , and the second and third layers respectively output the frequency information feature with watermark characteristics .
[0011] Step 3 includes: Step 3-1, replace all the downsampling modules in the original encoding stage of the Unet network with discrete wavelet transform downsampling. The discrete wavelet downsampling includes: applying two-dimensional discrete wavelet transform to the natural image X with dimension , and decomposing X into four frequency-domain subgraphs, where H represents the length of the natural image X, W represents the width of the natural image X, and C represents the number of channels of the natural image X; The four frequency-domain subgraphs are the low-frequency component subgraph LL, the horizontal high-frequency component subgraph LH, the vertical high-frequency component subgraph HL, and the diagonal high-frequency component subgraph HH; the four frequency-domain subgraphs are concatenated along the channel dimension to construct a new feature map , with a size of ; apply a standard convolution operation for channel dimensionality reduction to re-extract and fuse the frequency-domain information, generating the downsampled image , with a size of ; Step 3-2, let the input feature image be , perform horizontal pooling on F in the horizontal direction and vertical pooling in the vertical direction to capture the axial global context information, where the horizontal pooling operation in the horizontal direction is performed along the column dimension, generating the global feature vector in the horizontal direction; the average pooling operation in the vertical direction is performed along the row dimension, generating the global feature vector in the vertical direction; apply broadcast addition to the feature vectors and to achieve global supplementation of the input feature , generating the global feature ; through the calibration function , perform a non-linear transformation on the generated global feature to obtain the calibrated global feature : , where and are learnable parameter matrices, represents the Relu activation function; Subsequently, perform secondary calibration through two different-direction bar convolutions: use the vertical bar convolution to calibrate the shape in the vertical direction and the horizontal bar convolution to calibrate the shape in the horizontal direction; Introduce the cover self-attention CSA in the encoder stage of the Unet network and send it to the decoder through skip connections, and add the channel squeeze-and-excitation attention SE module after each upsampling in the decoding stage; In the first layer of the Unet network, let be the feature extracted during the encoding process, be the watermark information feature output by the message processor , be the feature extracted during the decoding stage. First, obtain the attention weight coefficient after secondary feature calibration through the cover self-attention CSA, multiply and element-wise to obtain the feature , and then normalize the features through batch normalization (BN), and output the spatial domain features through the ReLU activation function , and splice the and the watermark information features and the features extracted by the encoder in the channel dimension to obtain the fused features . Then, use the channel attention SE module to calculate the attention weight coefficients of the features , multiply and element-wise to obtain the new features , and then perform upsampling and use it as the features extracted by the decoder in the second layer of the Unet network.
[0012] In step 4, the non-differentiable noise simulation training process includes the mini-batch training strategy Mini-Bath and the non-differentiable simulation noise layer DiffJpeg.
[0013] In step 4, the non-differentiable simulation noise layer DiffJpeg includes a forward propagation process and a backward propagation process. In the forward propagation, real rounding, flooring, and clipping operations are used for image attacks instead of simulated ones; in the backward propagation, an unconventional straight-through estimator (STE) is adopted, and polynomial approximation is used to replace the constant gradient. The formula is: where is the gradient of the pixel value during the backward propagation when the image encounters non-differentiable noise attacks.
[0014] To further improve the robustness of the model against non-differentiable noise attacks, the mini-batch training strategy Mini-Bath is adopted, and each time during training, it is randomly selected from the real non-differentiable noise attack layer JPEG, the no-noise attack layer Identity, and the non-differentiable simulation noise layer DiffJpeg.
[0015] In step 4, the following loss function L is adopted during the training process: where , , are weight coefficients and hyperparameters. In the present invention, they are set to 1, 10, and 0.0001 respectively; is the encoding loss, and the calculation formula is: where represents the cover image and the encoded image The mean square error represents the parameters of the encoder E, that is, the trainable weights and biases of the neural network. M represents the original watermark information. The purpose of the encoder is to make the encoded image and the cover image visually as similar as possible; is the encoder message loss, and the calculation formula is: , where represents the message predicted by the decoder and the mean square error of the original watermark information , represents the parameters of the decoder D, that is, the trainable weights and biases of the neural network, represents the image after noise attack; is the discriminator loss, and the calculation formula is: , where A is the discriminator, which approximates the original carrier image by generating random noise images. The first term is an adversarial loss, and the purpose is to make the discriminator A recognize that is tampered, so as to improve the detection performance. The second term , the purpose is to ensure that the original image can be correctly classified as untampered; are the parameters of the discriminator A.
[0016] The present invention also provides a robust image watermark soft fusion system based on the UNet architecture implemented according to the above method, including: Data preprocessing unit: used to receive the carrier image. After the carrier image is input into the data preprocessing unit, three low-frequency component features will be output, namely the first-level low-frequency component , the second-level low-frequency component and the third-level low-frequency component ; Message processing unit: used to receive the message and the low-frequency component features, and output the watermark information and watermark information with frequency domain features; Encoding unit: used to receive the watermark information block and the carrier image from the message processor, and output the encoded image through the watermark soft fusion network ProwterNet; Adversarial training unit: input the encoded image into the adversarial training unit, and perform adversarial training through perturbation to improve the encoding quality; Noise unit: Input the encoded image into a pre-established noise simulation layer, and optimize the robustness of the embedding and the visual quality of the image through backpropagation; for non-differentiable noise, design a non-differentiable noise simulation training process for training. Decoding and extraction unit: Input the noisy image into a pre-established decoding and extraction unit to extract one-dimensional watermark information.
[0017] The present invention also provides an electronic device, including a processor and a memory. The memory stores program code, and when the program code is executed by the processor, the processor is caused to execute the steps of the method.
[0018] The present invention also provides a storage medium storing a computer program or instruction, and when the computer program or instruction runs on a computer, the steps of the method are executed.
[0019] In order to improve the detectability of the watermark under screen capture conditions and overcome the defect of simple replication of traditional watermarks, the present invention first enables the watermark information to be adaptively extended into the image in stages through a diffusion sub-network, realizing a more stable watermark coding representation. Subsequently, based on the UNet structure, it replaces the traditional additive fusion method and introduces a frequency-domain to spatial-domain joint coding strategy, enabling the watermark information to be progressively embedded in different scales and frequency domains to significantly improve the robustness against screen capture. In addition, a cover self-attention mechanism is designed at the jump link. By adaptively allocating the attention weights of the cover image, it guides the watermark to make dynamic adjustments, enhancing the interaction between the watermark and the image features, so that the watermark still has strong extractability after screen capture. Aiming at the problem of gradient break caused by non-differentiable noise, the present invention proposes a DiffJpeg noise layer to optimize the robustness of the model under screen capture-related attacks such as JPEG compression.
[0020] The present invention has the following beneficial effects: The present invention uses ProwterNet to process the watermark information and the carrier image, realizing multi-scale cross-domain progressive embedding. Based on the common distortion situations in the screen capture scenario, the present invention evaluates the robustness of the model under the influence of different noises. During the screen capture process, the image will undergo different forms of degradation, such as compression, cropping, signal loss, and sensor noise, etc. Therefore, the present invention uses the common noise in traditional watermark methods for experiments to test the robustness of the model in the screen capture environment. The experiments include non-differentiable noise (such as JPEG compression) and differentiable noise (such as cropping, Dropout, Gaussian noise, etc.). Description of the Drawings
[0021] Figure 1 It is the overall framework diagram of the method of the present invention.
[0022] Figure 2 It is the encoder architecture diagram.
[0023] Figure 3 It is a wavelet downsampling graph.
[0024] Figure 4 It is a CSA module graph. Specific implementation manners
[0025] The following further specifically describes the present invention in conjunction with the accompanying drawings and specific implementation manners, and the above and / or other advantages of the present invention will become clearer.
[0026] An embodiment of the present invention provides a robust image watermark soft fusion network method based on the UNet architecture, including: Step 1, preprocess the cover image to obtain low-frequency components; Step 2, establish a multi-stage watermark extension sub-network, and perform feature fusion on the watermark information tensor and the low-frequency components; Step 3, construct a watermark soft fusion network ProwterNet based on the Unet architecture: add discrete wavelet transform downsampling, cover self-attention CSA, and channel squeeze attention SE modules in the Unet network, and cascade message processors; first, replace all downsampling modules in the encoding stage of the Unet network with discrete wavelet transform downsampling, and add channel squeeze attention SE after upsampling; secondly, add cover self-attention CSA before the skip connection in the decoding stage; finally, receive the output of the cascade message processor, and the output includes the frequency-domain information of the embedded watermark features and the watermark message, and fuse the frequency-domain information of the embedded watermark features and the watermark message with the features of the skip connection and the features after continuous upsampling to achieve joint embedding in the spatial domain and the frequency domain; Step 4, design a non-differentiable noise simulation training process to train the watermark soft fusion network ProwterNet.
[0027] Step 1 includes: Step 1-1, input the cover image into the discrete wavelet transform network, and decompose it into the first-level low-frequency component , the first-level horizontal high-frequency component , the first-level vertical high-frequency component , and the first-level diagonal high-frequency component ; Step 1-2, input the low-frequency component into the discrete wavelet transform network again for secondary decomposition to obtain the second-level low-frequency component , the second-level horizontal high-frequency component , the second-level vertical high-frequency component , the second-level vertical high-frequency component ; Step 1-3, input the second-level low-frequency component Extract the third-level low-frequency components in the input discrete wavelet transform network .
[0028] Step 2 includes: Step 2-1: Randomly generate watermark information of 64 bits, and expand the randomly generated watermark information into a watermark information tensor of through the fully connected layer ; ; Step 2-2: Duplicate the expanded watermark information tensor 4 times, denoted as , where ; ; Step 2-3: In the message processor of layer 0 , receive the expanded watermark information tensor , and directly output the watermark information tensor without additional processing; Step 2-4: Concatenate the three low-frequency components , , obtained in step 1 into the message processor of layer 1 , the message processor of layer 2 and the message processor of layer 3 respectively; Step 2-5: The message processor of layer 1 receives two inputs: the expanded watermark information tensor and the first-level low-frequency component ; Step 2-6: The message processor of layer 1 first expands the watermark information tensor through a convolutional neural network of to a watermark information tensor of the same size as the first-level low-frequency component , then performs a Hadamard product calculation on the watermark information tensor and the first-level low-frequency component , and generates an attention weight coefficient through the function, and uses the attention weight to perform feature fusion on the watermark information tensor and the first-level low-frequency component and outputs the frequency-domain information feature with watermark characteristics after fusion ; Step 2-7: In the message processor of layer 2 and the message processor of layer 3 respectively receive the expanded watermark information tensor and the secondary low-frequency component and the extended watermark information tensor and the tertiary low-frequency component , and repeat steps 2-6 to output the fused frequency-domain information features with watermark characteristics and , and the formula for the specific fusion process is: , where represents 's convolutional neural network, is function, is the Hadamard product calculation product; z takes values of 1, 2, 3; message processor has a total of four layers. The 0th layer outputs the watermark information tensor , and the first layer processes the watermark information tensor and finally outputs the frequency-domain information features with watermark characteristics . The second and third layers respectively output the frequency information features with watermark characteristics .
[0029] Step 3 includes: Step 3-1: Replace all the downsampling modules in the original encoding stage of the Unet network with discrete wavelet transform downsampling. The discrete wavelet downsampling includes: applying a two-dimensional discrete wavelet transform to the natural image X with dimensions to decompose X into four frequency-domain subgraphs, where H represents the length of the natural image X, W represents the width of the natural image X, and C represents the number of channels of the natural image X; The four frequency-domain subgraphs are respectively the low-frequency component subgraph LL, the horizontal high-frequency component subgraph LH, the vertical high-frequency component subgraph HL, and the diagonal high-frequency component subgraph HH; splicing the four frequency-domain subgraphs along the channel dimension to construct a new feature map , with a size of ; applying standard convolution operations for channel dimensionality reduction to re-extract and fuse the frequency-domain information to generate the downsampled image , with a size of ; Step 3-2: Let the input feature image be . Perform horizontal pooling on F in the horizontal direction and vertical pooling in the vertical direction to capture the axial global context information. The horizontal pooling operation in the horizontal direction is performed along the column dimension to generate the horizontal global feature vector ; the average pooling operation in the vertical direction is performed along the row dimension to generate the vertical global feature vector ; for the feature vector and Apply broadcast addition to achieve global supplementation of the input features and generate global features ; Through the calibration function , perform a non-linear transformation on the generated global features to obtain the calibrated global features : , where and are learnable parameter matrices, represents the Relu activation function; Subsequently, perform secondary calibration through two different-direction bar convolutions: use vertical bar convolution to calibrate the shape in the vertical direction, and use horizontal bar convolution to calibrate the shape in the horizontal direction; Introduce Cover Self-Attention (CSA) in the encoder stage of the Unet network and send it to the decoder through skip connections. Add a Channel Squeeze-and-Excitation (SE) module after each upsampling in the decoding stage; In the first layer of the Unet network, let be the features extracted during the encoding process, be the watermark information features output by the message processor , be the features extracted during the decoding stage. First, obtain the attention weight coefficients after secondary feature calibration through Cover Self-Attention (CSA). Multiply and element-wise to obtain the feature . Subsequently, normalize the feature through Batch Normalization (BN) and output the spatial domain feature through the ReLU activation function. Concatenate and the watermark information features , and the features extracted by the encoder in the channel dimension to obtain the fused feature . Calculate the attention weight coefficients of the feature using the Channel Squeeze-and-Excitation (SE) module. Multiply and element-wise to obtain the new feature . Subsequently, perform upsampling and use it as the features extracted by the decoder in the second layer of the Unet network.
[0030] In step 4, the non-differentiable noise simulation training process includes the mini-batch training strategy Mini-Bath and the non-differentiable simulation noise layer DiffJpeg.
[0031] In step 4, the non-differentiable simulation noise layer DiffJpeg includes a forward propagation process and a backward propagation process. In the forward propagation, real rounding, flooring, and clipping operations instead of simulated ones are used for image attacks; in the backward propagation, an unconventional straight-through estimator STE is adopted, and polynomial approximation is used to replace the constant gradient. The formula is: , where is the gradient of the pixel value during the backward propagation when the image encounters a non-differentiable noise attack.
[0032] To further improve the robustness of the model against non-differentiable noise attacks, a mini-batch training strategy Mini-Bath is adopted. Each time during training, it is randomly selected from the real non-differentiable noise attack layer JPEG, the no-noise attack layer Identity, and the non-differentiable simulation noise layer DiffJpeg.
[0033] In step 4, the following loss function L is adopted during the training process: , where , , are weight coefficients and hyperparameters. In the present invention, they are respectively set to 1, 10, and 0.0001; is the encoding loss, and the calculation formula is: , where represents the cover image and the encoded image 's mean square error, represents the parameters of the encoder E, that is, the trainable weights and biases of the neural network, and M represents the original watermark information; the purpose of the encoder is to make the encoded image and the cover image visually as similar as possible; is the encoder message loss, and the calculation formula is: , where represents the message predicted by the decoder and the original watermark information 's mean square error, represents the parameters of the decoder D, that is, the trainable weights and biases of the neural network, represents the image after being attacked by noise; is the discriminator loss, and its calculation formula is: , where A is the discriminator, which approximates the original carrier image by generating a random noise image; the first term is an adversarial loss, aiming to make the discriminator A recognize that is tampered, so as to improve the detection performance. The second term , aiming to ensure that the original image can be correctly classified as untampered; are the parameters of the discriminator A. Through this optimization process, more robust watermark embedding and extraction are achieved, the ability of the watermark to resist noise attacks is improved, and the visual quality of the encoded image is ensured at the same time.
[0034] The method of the present invention shows significant advantages in both robustness and invisibility. In the experimental evaluation, a variety of common noise attack means (including JPEG compression, Crop, Dropout, Gussian noise, etc.) were used to test the watermark system. The results show that the proposed method can maintain a low bit error rate and high peak signal-to-noise ratio PSNR and structural similarity SSIM under different distortion conditions, demonstrating excellent anti-screen capture ability. Especially under the high compression ratio JPEG attack, the proposed method can maintain the stable extraction of watermark information, and still maintain the integrity of the watermark in the more destructive scenarios such as cropping and Dropout, proving its wide applicability in practical applications.
[0035] Table 1
[0036]
[0037] Table 1 is the experimental data table of the Crop attack. The Crop attack is the influence of the perspective shift in screen capture. During the screen capture process, the change of the user's shooting angle or the adjustment of the viewfinder range may cause partial loss of the watermark information. Therefore, the present invention simulates different degrees of cropping distortion to investigate the adaptability of the model to the change of the screen capture perspective. In the experiment, the present invention selects four cropping ratios R = 0.3, 0.5, and 0.7, which respectively represent that the size of the cropped image is 30%, 50%, and 70% of the original image. The experimental results show that cropping causes significant damage to the integrity of the embedded information, but the model of the present invention can still maintain a 0% bit error rate at each cropping ratio. The measurement metrics Metric are selected as the peak signal-to-noise ratio PSNR and the structural similarity SSIM. This experiment shows that when maintaining PSNR and structural similarity SSIM at a high level, the model has extremely strong robustness in the screen capture environment, and can still effectively recover the embedded content even in the case of partial loss of watermark information. Among them, HiDDeN, MBRS, and Adaptor are the methods of different papers. The present invention selects these three methods as the control experimental groups to highlight the superiority of the present invention. The subsequent experiments will appear repeatedly.
[0038] Table 2
[0039]
[0040] Table 2 is for Gaussian noise interference. That is, in actual life, it is the sensor noise in screen capture. During the screen capture process, affected by factors such as sensor performance, low-light environment, or device jitter, the image may be interfered by random noise. In this invention, Gaussian noise is used for experiments to simulate the noise pollution situation in the screen capture environment and evaluate the robustness of the watermark information. In the experiment, Gaussian noise with different variances (0.001, 0.002, 0.005, and 0.010) is added to the image, and its influence is analyzed. The experimental results show that Gaussian noise will cause the embedded information to present a regular dot-like distribution. Especially when the noise intensity is high, the change in the information position is more obvious. This is due to the distribution characteristics of Gaussian noise, which makes the encoder more inclined to embed information in specific areas during the screen capture process. Although Gaussian noise will damage part of the image structure, the model of this invention still maintains a high watermark extractability, demonstrating good adaptability to the screen capture sensor noise.
[0041] The JPEG compression experiment is shown in Table 3. JPEG compression is the influence of lossy transmission during the screen capture process. For example, JPEG compression is a common image degradation factor during the screen capture process, especially obvious in social media dissemination or low-quality screen shooting. Its implementation process includes color mode conversion, data sampling, discrete cosine transform (DCT), frequency coefficient quantization, and encoding. Among them, due to the rounding operation in the quantization process, it is difficult for the gradient to effectively backpropagate, and it directly affects the extraction accuracy of the watermark information. In the experiment, different quality factors (Q) were used to test the model. Q = 30, Q = 50, and Q = 70 were respectively selected for JPEG compression attacks, and their influence on the watermark information was evaluated. It can be seen from the experimental results that as Q increases, the compression rate decreases, the image quality improves, PSNR and SSIM also increase, while the bit error rate (BER) gradually decreases. Comparing different methods under the condition of nearly the same bit error rate, the results show that the method of this invention performs excellently in terms of the compressed image quality, bit error rate control, and structure preservation. Especially in the screen capture environment, it can maintain good image readability while ensuring the robustness of the watermark.
[0042] Table 3
[0043]
[0044] Pixel loss attack Dropout, and the experimental results are shown in Table 4. Signal loss impact in screen capture During screen capture, due to issues such as screen refresh rate, pixel sampling, or signal interference, some pixels may be lost, resulting in damage to the integrity of the watermark information. To simulate this situation, the pixel loss attack Dropout was used in the experiment, that is, some pixels were randomly removed and replaced with pixels from the cover image to test the robustness of the model in the scenario of screen capture signal loss. In the experiment, Dropout ratios R = 0.3, 0.5, and 0.7 were selected, which represent 30%, 50%, and 70% of the remaining encoded image compared to the original image respectively. The results show that in the case of a high ratio of Dropout (such as R = 0.7), the method of the present invention still maintains a nearly perfect recovery effect, the bit error rate is significantly lower than other methods, and at the same time, SSIM and PSNR still remain at a relatively high level. This indicates that the method of the present invention can effectively resist signal loss during screen capture and performs excellently in terms of information integrity and visual quality.
[0045] Table 4
[0046]
[0047] The present invention provides a method and system for a robust anti-screen-capture image watermark soft fusion network based on the UNet architecture. There are many methods and ways to specifically implement this technical solution. The above description is only the preferred implementation manner of the present invention. It should be noted that for those of ordinary skill in the art in this technical field, without departing from the principle of the present invention, several improvements and refinements can be made, and these improvements and refinements should also be regarded as the protection scope of the present invention. Each component not clearly defined in this embodiment can be implemented using existing technologies.
Claims
1. A screen capture robust image watermarking soft fusion network method based on UNet architecture, characterized in that: The following steps are involved: Step 1, preprocessing the cover image to obtain low-frequency components; Step 2: Establish a multi-stage watermark extension sub-network to fuse the watermark information tensor with the low-frequency component; Step 3, construct a watermark soft fusion network ProwterNet based on the Unet architecture: add discrete wavelet transform downsampling, cover self-attention CSA and channel compression attention SE modules to the Unet network, and cascade message processors; first replace all downsampling modules in the encoding stage of the Unet network with discrete wavelet transform downsampling, and add channel compression attention SE after upsampling; secondly, add cover self-attention CSA before jump connection in the decoding stage; finally, receive the output of the cascade message processor, which includes the frequency domain information and watermark message of the embedded watermark feature, and fuse the frequency domain information and watermark message of the embedded watermark feature with the features of the jump link and the features after continuous upsampling to achieve joint embedding of the spatial domain and the frequency domain; Step 4: design a non-differentiable noise simulation training process to train the watermark soft fusion network ProwterNet.
2. The method according to claim 1, characterized in that Step 1 includes: Step 1-1: Input the cover image into the discrete wavelet transform network and decompose it into the first-level low-frequency component after the first layer of wavelet transform. , the first-level high-frequency component , the first-order vertical high-frequency component , the first-order diagonal high-frequency component ; Step 1-2: The low frequency components Input it into the discrete wavelet transform network again, perform secondary decomposition, and obtain the secondary low-frequency component , secondary horizontal high frequency component , secondary vertical high frequency component , secondary vertical high frequency component ; Step 1-3, the secondary low-frequency component Input discrete wavelet transform network to extract the third-level low-frequency components .
3. The method according to claim 2, characterized in that Step 2 includes: Step 2-1: Randomly generate a 64-bit watermark , through the fully connected layer, the randomly generated watermark information Expand to The watermark information tensor ; Step 2-2: Expand the watermark information tensor Make 4 copies, record as ,in ; Step 2-3, at the layer 0 message processor In the example above, we receive the expanded watermark information tensor , without additional processing, directly output the watermark information tensor ; Step 2-4: The three low-frequency components obtained in step 1 , , Cascaded to the first layer message processor , Layer 2 message processor and the third layer message processor middle; Step 2-5, first layer message processor Receive two inputs: the expanded watermark information tensor , the first-level low-frequency component ; Step 2-6, first layer message processor First, the watermark information tensor pass The convolutional neural network is extended to the first-level low-frequency component Watermark information tensor of the same size , and then the watermark information tensor and the first-order low-frequency component Perform Hadamard product calculations and pass Function generates attention weight coefficient , using attention weights to watermark information tensors and the first-order low-frequency component Perform feature fusion and output the frequency domain information features with watermark characteristics after fusion ; Step 2-7, in the second layer message processor and the third layer message processor Receive the expanded watermark information tensor respectively , secondary low frequency component And the expanded watermark information tensor , three-level low-frequency component , and repeat steps 2-6 to output the fused frequency domain information features with watermark characteristics and , the specific fusion process formula is: , in express Convolutional neural network, for function, Calculate the product for the Hadamard product; z takes values of 1, 2, and 3; Message Processor There are four layers in total, and the 0th layer outputs the watermark information tensor , the first layer is the watermark information tensor Processing is performed and finally the frequency domain information features with watermark characteristics are output The second and third layers output frequency information features with watermark characteristics. .
4. The method according to claim 3, characterized in that Step 3 includes: Step 3-1, all downsampling modules in the original encoding stage of the Unet network are replaced by discrete wavelet transform downsampling, wherein the discrete wavelet downsampling includes: Apply a two-dimensional discrete wavelet transform to the natural image X and decompose X into four frequency domain sub-images, where H represents the length of the natural image X, W represents the width of the natural image X, and C represents the number of channels of the natural image X; The four frequency domain sub-images are respectively a low frequency component sub-image LL, a horizontal high frequency component sub-image LH, a vertical high frequency component sub-image HL and a diagonal high frequency component sub-image HH; the four frequency domain sub-images are spliced along the channel dimension to construct a new feature map , size is ; Apply standard convolution operations to perform channel dimension reduction, re-extract and fuse frequency domain information, and generate downsampled images , the size is ; Step 3-2, let the input feature image be , horizontal pooling is performed on F in the horizontal direction and vertical pooling is performed in the vertical direction to capture the axial global context information, where the horizontal pooling operation is performed along the column dimension to generate a global feature vector in the horizontal direction ; The vertical average pooling operation is performed along the row dimension to generate a global feature vector in the vertical direction ; For the feature vector and Apply broadcast addition to implement input features The global complement of ; Through the calibration function , for the generated global features Perform nonlinear changes to obtain calibrated global features : , in and is the learnable parameter matrix, Represents the Relu activation function; Then, a second calibration is performed through two strip convolutions in different directions: the vertical shape is calibrated using the vertical strip convolution, and the horizontal shape is calibrated using the horizontal strip convolution; Cover self-attention CSA is introduced in the Unet network encoder stage and sent to the decoder through skip links. A channel compression attention SE module is added after each upsampling in the decoding stage. In the first layer of the Unet network, are the features extracted during the encoding process. For message processor Output watermark information features, For the features extracted in the decoding stage, the cover self-attention CSA is first used to obtain the attention weight coefficient after secondary feature calibration ,Will and Multiply elements by element to get features , and then batch normalization BN is used to normalize the features Normalize and output spatial features through ReLU activation function ,Will and watermark information characteristics , encoder extracts features Perform channel feature splicing to obtain fusion features , the channel attention SE module is used to calculate the feature The attention weight coefficient ,Will and Multiply elements by element to get new features , which is then upsampled and used as the feature extracted by the decoder in the second layer of the Unet network.
5. The method according to claim 4, characterized in that In step 4, the non-differentiable noise simulation training process includes a small batch training strategy Mini-Bath and a non-differentiable simulated noise layer DiffJpeg.
6. The method according to claim 5, characterized in that In step 4, the non-differentiable simulated noise layer DiffJpeg includes a forward propagation process and a backward propagation process. In the forward propagation, rounding, flooring and clipping operations are used for image attacks; in the backward propagation, a straight-through estimator STE is used, and a polynomial approximation is used to replace the constant gradient. The formula is: , in It is the gradient of the pixel value back-propagated when the image encounters non-differentiable noise attack; The mini-batch training strategy Mini-Bath is adopted, and each time during training, a random selection is made among the real non-differentiable noise attack layer JPEG, the noise-free attack layer Identity, and the non-differentiable simulated noise layer DiffJpeg.
7. The method according to claim 6, characterized in that In step 4, the following loss function L is used during training: , in , , is the weight coefficient; is the coding loss, which is calculated as: , in Indicates the cover image With the encoded image The mean square error, represents the parameters of the encoder E, and M represents the original watermark information; is the encoder message loss, calculated as: , in Indicates the message predicted by the decoder and original watermark information The mean square error, represents the parameters of the decoder D, Represents the image after noise attack; is the discriminator loss, calculated as: , Where A is the discriminator; are the parameters of the discriminator A.
8. A UNet-based robust image watermark soft fusion system for anti-screening according to the method described in any one of claims 1 to 7, characterized in that: include: Data preprocessing unit: used to receive the carrier image. After the carrier image is input into the data preprocessing unit, three low-frequency component features will be output, namely, the primary low-frequency component , secondary low frequency component and three-level low-frequency components ; Message processing unit: used to receive the message and low-frequency component features, and output watermark information and watermark information with frequency domain features; Encoding unit: used to receive the watermark information block and carrier image from the message processor, and obtain the encoded image through the watermark soft fusion network ProwterNet output; Adversarial training unit: Input the encoded image into the adversarial training unit, perform adversarial training through perturbation to improve the encoding quality; Noise unit: Input the encoded image into a pre-established noise simulation layer and optimize the embedding robustness and image visual quality through back-propagation; For non-differentiable noise, a non-differentiable noise simulation training process is designed for training; Decoding and extraction unit: The noisy image is input into the pre-established decoding and extraction unit to extract the one-dimensional watermark information.
9. An electronic device, characterized in that: The method comprises a processor and a memory, wherein the memory stores program codes, and when the program codes are executed by the processor, the processor executes the steps of the method according to any one of claims 1 to 7.
10. A storage medium, characterized in that: A computer program or instruction is stored, and when the computer program or instruction is run on a computer, the steps of the method according to any one of claims 1 to 7 are executed.
Citation Information
Patent Citations
JPEG (Joint Photographic Experts Group) compression robustness resistant image watermarking method based on multi-scale automatic encoder
CN116883222A
Robust screen watermark shooting method based on wavelet domain cascade network and reverse recovery
CN117455749A
Robust watermark embedding method and system based on information multi-dimensional embedding and texture guidance
CN117974412A
Cited By
Screen shooting resistant robust text image watermarking method fusing edge attention gating and multi-scale cavity convolution
CN120707366A
Generated content protection method, electronic equipment and storage medium
CN120744887A
Robust feature image watermark embedding method based on noise perception guidance
CN122089551A