Image watermark generation method, image watermark generation device and computer storage medium

Through the variational autoencoder and frequency domain attention mechanism combined with the differentiable minimum perceptible model, the watermark embedding intensity is adaptively adjusted, which solves the robustness problem of the digital image watermark system under composite attack, and achieves high inconsensuality and strong robustness.

CN120410831AActive Publication Date: 2025-08-01IFLYTEK CO LTD

Patent Information

Application Number
CN202510916718.3
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-07-03
Publication Date
2025-08-01
Estimated Expiration
2045-07-03

AI Technical Summary

Technical Problem

The prior art is difficult to design a digital image watermark system that takes into account both invisibility and strong robustness without affecting visual quality, especially when facing composite attacks, the robustness decreases sharply.

Method used

The variational autoencoder and frequency domain attention mechanism are used to combine differentiable minimum perceptible model to adaptively adjust the watermark embedding intensity, and the watermark energy distribution and image visual characteristics are achieved through an end-to-end deep convolution network, and a differentiable composite attack simulation module is built.

Benefits of technology

It significantly improves the robustness and invisibility of the watermark, can effectively resist JPEG compression, geometric deformation and composite attacks, reduces the bit error rate to below 5%, and increases the PSNR to above 40dB.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120410831A_ABST
    Figure CN120410831A_ABST
Patent Text Reader

Abstract

The invention provides an image watermark generation method, an image watermark generation device and a computer storage medium. The image watermark generation method comprises the following steps: acquiring original potential features of an original image; watermarking information with the same size as the original potential features is fused with the original potential features, and fused potential features are obtained; acquiring a pixel-level visual sensitivity mask parameter of the original image; adjusting the watermark intensity of the fused potential feature by using the pixel-level visual sensitivity mask parameter to obtain an adjusted potential feature; generating an image watermark according to the adjusted potential features; and generating a watermark image according to the image watermark and the original image. According to the image watermark generation method, the watermark embedding strength is adaptively adjusted by using the pixel-level visual sensitivity mask parameters, so that global collaborative adaptation of watermark energy distribution and image visual characteristics is realized, and the robustness of the watermark to composite attacks is remarkably improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present application relates to the field of digital watermark technology, and in particular to an image watermark generation method, an image watermark generation device, and a computer storage medium. Background Art

[0002] Digital image watermarking technology is a key research area in the field of information hiding. Its core goal is to embed identifying information into a carrier image without compromising visual quality, and to accurately extract that information during transmission and after potential attacks. This technology boasts strong visual imperceptibility and high traceability, making it widely applicable in media copyright management, judicial evidence collection, medical image security, and other fields.

[0003] In adversarial application scenarios (such as social media communication), since it is necessary to resist complex attacks such as value transformation (brightness / contrast adjustment, etc.), geometric transformation (rotation / scaling / cropping, etc.), and image masking (partial occlusion / mosaic), designing a watermark system that takes into account both imperceptibility (PSNR>40dB) and strong robustness (BER<1%) becomes a technical difficulty.

[0004] Traditional mainstream methods are based on frequency domain transforms (such as discrete cosine transform and wavelet transform) and use quantization coding techniques to embed watermarks by modulating coefficients in mid- and high-frequency bands. Although these methods have basic anti-attack capabilities, they can significantly increase the bit error rate of watermark extraction and sharply reduce robustness when faced with coordinated damage from strong JPEG compression (quality factor <50), geometric deformation (rotation >15° / scaling ±20%), non-uniform cropping (random cropping >30% of the area), and mixed attacks (such as compression + cropping + noise). Summary of the Invention

[0005] To solve the above technical problems, the present application proposes an image watermark generation method, an image watermark generation device and a computer storage medium.

[0006] To solve the above technical problems, the present application proposes a method for generating an image watermark, which includes: Obtain the original latent features of the original image; fusing the watermark information having the same size as the original latent feature with the original latent feature to obtain a fused latent feature; Obtaining pixel-level visual sensitivity mask parameters of the original image; adjusting the watermark intensity of the fused latent feature using the pixel-level visual sensitivity mask parameter to obtain the adjusted latent feature; generating an image watermark based on the adjusted latent features; A watermark image is generated according to the image watermark and the original image.

[0007] Among them, obtaining the original latent features of the original image includes: Extracting the latent space representation of the original image using a variational autoencoder; Performing a Fourier transform on the latent space representation to obtain the original latent features; Fusing the watermark information of the same size as the original latent features with the original latent features to obtain fused latent features, including: Fusing the watermark information of the same size as the original latent features with the original latent features, and reconstructing the fused latent features through an inverse Fourier transform.

[0008] Among them, obtaining the pixel-level visual sensitivity mask parameters of the original image includes: Inputting the original image into a differentiable minimum perceptible model to extract the pixel-level visual sensitivity mask parameters of the original image; Among them, the model parameters of the differentiable minimum perceptible model are iteratively optimized through the visual difference value between the watermark image and the original image.

[0009] Among them, the image watermarking method further includes: Extracting the high-dimensional feature map of the watermark image; Inputting the high-dimensional feature map into a two-stream decoding branch to obtain the watermark detection result of the watermark image; Among them, the model parameters of the two-stream decoding branch are iteratively optimized through the watermark detection result, and the model parameters of the differentiable minimum perceptible model and the model parameters of the two-stream decoding branch are jointly optimized in the same watermark generation and watermark detection stage.

[0010] Among them, the two-stream decoding branch includes a mask detection branch and a watermark decoding branch; Inputting the high-dimensional feature map into the two-stream decoding branch to obtain the watermark detection result of the watermark image, including; Inputting the high-dimensional feature map into the mask detection branch and the watermark decoding branch respectively; Extracting the multi-scale context features of the high-dimensional feature map through the mask detection branch; Fusing the shallow high-resolution features and the deep semantic features in the multi-scale context features to generate attention weights; Using the attention weights to output the classification result of whether the watermark image contains a watermark; Decomposing the high-dimensional feature map into low-frequency contours and high-frequency texture components through the watermark decoding branch; After the low-frequency contour and high-frequency texture components are hierarchically recombined and upsampled, a watermark prediction value is output.

[0011] Among them, the extraction of the high-dimensional feature map of the watermark image includes: The high-dimensional feature map of the watermark image is extracted by embedding deformable convolution.

[0012] Among them, the image watermarking method further includes: Extract the key semantic regions in the watermark image; Generate a binary mask on the image corresponding to the key semantic region to impose an attack on the watermark image and generate a perturbed image; Calculate the perturbation loss using the perturbed image and the watermark image; Iteratively optimize the differentiable minimum perceptible model according to the perturbation loss.

[0013] Among them, the image watermarking method further includes: Use a differentiable simulator to impose a discretization attack on the watermark image; And / or, use a spatial transformation network and bilinear interpolation to impose a differentiable geometric transformation attack on the watermark image to generate a perturbed image; Calculate the perturbation loss using the perturbed image and the watermark image; Iteratively optimize the differentiable minimum perceptible model according to the perturbation loss.

[0014] Among them, the adjustment of the watermark intensity of the fused latent feature using the pixel-level visual sensitivity mask parameter to obtain an adjusted latent feature includes: Use the pixel-level visual sensitivity mask parameter and a dynamic attenuation system to adjust the fused latent feature to obtain the adjusted latent feature.

[0015] To solve the above technical problems, the present application also proposes an image watermark generation device, which includes a memory and a processor coupled to the memory; among them, the memory is used to store program data, and the processor is used to execute the program data to implement the image watermark generation method as described above.

[0016] To solve the above technical problems, the present application also proposes a computer storage medium, which is used to store program data, and the program data, when executed by a computer, is used to implement the above image watermark generation method.

[0017] Compared with the prior art, the beneficial effects of the present application are as follows: Through the above image watermark generation method, the pixel-level visual sensitivity mask parameter is used to adaptively adjust the watermark embedding strength, realizing the global collaborative adaptation of the watermark energy distribution and the visual characteristics of the image, and significantly improving the robustness of the watermark against composite attacks; The present application calculates the pixel-level visual sensitivity mask parameter for the original image, realizes the generation of the adaptive mask parameter for each original image, and uses the adaptive mask parameter to adaptively adjust the watermark embedding strength, significantly improving the robustness of watermark generation on the premise of ensuring imperceptibility. BRIEF DESCRIPTION OF THE DRAWINGS

[0018] In order to more clearly illustrate the technical solutions in the embodiments of the present application, the following will briefly introduce the drawings required for the description of the embodiments. Obviously, the following drawings are only some embodiments of the present application. For those of ordinary skill in the art, without creative efforts, other drawings can be obtained according to these drawings. Among them: Figure 1 is a schematic flowchart of the first embodiment of the image watermark generation method provided by the present application; Figure 2 is a schematic diagram of a robust and imperceptible watermark embedding structure provided by the present application; Figure 3 is a schematic flowchart of the second embodiment of the image watermark generation method provided by the present application; Figure 4 is a schematic diagram of the scenario of attacking the watermarked image provided by the present application; Figure 5 is a schematic flowchart of the third embodiment of the image watermark generation method provided by the present application; Figure 6 is a schematic diagram of the scenario of watermark information detection and extraction provided by the present application; Figure 7 is Figure 5 a specific schematic flowchart of step S32 of the image watermark generation method shown; Figure 8 is a schematic structural diagram of an embodiment of the image watermark generation device provided by the present application; Figure 9 is a schematic structural diagram of an embodiment of the computer storage medium provided by the present application. DETAILED DESCRIPTION OF THE EMBODIMENTS

[0019] Next, the technical solutions in the embodiments of the present application will be clearly and completely described in conjunction with the accompanying drawings in the embodiments of the present application. Obviously, the described embodiments are only a part of the embodiments of the present application, rather than all the embodiments. Based on the embodiments in the present application, all other embodiments obtained by those of ordinary skill in the art without creative efforts shall fall within the protection scope of the present application.

[0020] The terms "first", "second", "third", "fourth", etc. (if any) in the specification, claims and above-mentioned drawings of the present application are used to distinguish similar objects, and do not necessarily need to be used to describe a specific order or sequence. It should be understood that such data can be interchanged under appropriate circumstances, so that the embodiments of the present application described herein, for example, can be implemented in an order other than those illustrated or described herein. In addition, the terms "comprising" and "having" and any variations thereof are intended to cover non-exclusive inclusion. For example, a process, method, system, product or device that includes a series of steps or units does not necessarily have to be limited to those steps or units clearly listed, but may include other steps or units not clearly listed or inherent to these processes, methods, products or devices.

[0021] In recent years, deep learning has brought innovative opportunities to watermarking technology. It automatically learns the image feature distribution through multi-level non-linear transformations, and realizes end-to-end watermark embedding and extraction. Some researchers use fully connected neural networks (Fully Connected Networks) to directly map the pixel domain and optimize the model by minimizing the watermark reconstruction error. Although such methods can capture local texture features, they lack the modeling of the joint spatial-frequency characteristics of images, resulting in it being difficult to balance the invisibility of watermarks (PSNR (Peak Signal-to-Noise Ratio) < 38dB) and the anti-attack ability.

[0022] Due to its local perception and weight sharing characteristics, the convolutional neural network (CNN) has become an efficient solution in the field of watermarking. It extracts spatio-frequency interleaved features through multiple convolutional layers, captures multi-scale context information using downsampling, and combines upsampling to achieve pixel-level watermark localization. Compared with traditional frequency domain methods, CNN can jointly optimize the watermark embedding strength and the adaptability of the image structure, and achieve a dynamic balance between the frequency domain energy distribution and the spatial domain visual quality.

[0023] The present application innovatively proposes a multi-scale adversarial coupling watermarking framework, which introduces an attention mechanism and an adversarial training strategy on the basis of a deep convolutional network. Through the joint gradient optimization of the embedder-detector, the invisibility of the watermarking system (PSNR > 38dB) and the anti-composite attack ability (bit error rate < 0.5%) are improved.

[0024] Therefore, this application proposes an end-to-end adversarial enhanced digital image watermarking system, which focuses on solving the following key technical problems: 1) How to achieve an adaptive match between the watermark embedding strength and the visual characteristics of the image, and improve the robustness while maintaining high imperceptibility (PSNR>38dB).

[0025] 2) How to construct a differentiable composite attack simulation module so that the system can learn to resist multi-dimensional attacks in real scenarios.

[0026] 3) How to design a global-local coupled decoding architecture to achieve accurate positioning and recovery of watermark information.

[0027] Through the end-to-end joint training of the VAE (Variational Autoencoder Encoder-Decoder) encoder-decoder, combined with the frequency domain attention mechanism and the JND (Just Noticeable Difference Perceptual Model) perceptual model, this application improves the watermark detection accuracy by more than 15% and reduces the bit error rate to less than 5% for composite attacks such as JPEG compression and geometric deformation while ensuring visual quality.

[0028] For details, please refer to Figure 1 and Figure 2 , Figure 1 is a schematic flow diagram of the first embodiment of the image watermark generation method provided by this application, Figure 2 is a schematic diagram of a robust and imperceptible watermark embedding structure provided by this application.

[0029] The image watermark generation method of this application is applied to an image watermark generation device. Among them, the image watermark generation device of this application can be a server, a terminal device, or a system composed of a server and a terminal device cooperating with each other. Correspondingly, each part included in the image watermark generation device, such as each unit, sub-unit, module, and sub-module, can be all set in the server, all set in the terminal device, or separately set in the server and the terminal device.

[0030] Furthermore, the above-mentioned server can be hardware or software. When the server is hardware, it can be implemented as a distributed server cluster composed of multiple servers or as a single server. When the server is software, it can be implemented as multiple software or software modules, such as software or software modules used to provide a distributed server, or as a single software or software module, which is not specifically limited here.

[0031] Such as Figure 1As shown in the figure, the specific steps are as follows: Step S11: Obtain the original latent features of the original image.

[0032] In the embodiment of the present application, the image watermark generation device extracts the original latent features of the original image.

[0033] In one embodiment, the image watermark generation device can scale any original image z to a fixed size, and then use the VAE Encoder to transform the image from the pixel space to the latent space to obtain the latent space representation of the image, denoted as . Among them, the AE Encoder (Variational Autoencoder) is one of the core components of the Variational Autoencoder (VAE), which is responsible for compressing the input data (such as images, texts, etc.) into a low-dimensional latent representation and outputting the probability distribution parameters (mean and variance) of the latent space.

[0034] Furthermore, the image watermark generation device can also use FFT (Fast Fourier Transform) to transform it into the Fourier space to obtain , that is, the original latent features.

[0035] In another embodiment, the image watermark generation device can also directly compress the original image into a low-dimensional latent vector through a standard autoencoder.

[0036] In yet another embodiment, the image watermark generation device can also add a sparsity constraint (such as L1 regularization) to the loss function through a Sparse Autoencoder to force the sparsity of the original image latent vector.

[0037] Step S12: Fuse the watermark information with the same size as the original latent features with the original latent features to obtain the fused latent features.

[0038] In the embodiment of the present application, the image watermark generation device encodes the watermark information into a feature representation with the same size as and then fuses it with , and reconstructs the latent features after embedding the watermark through IFFT (Inverse Fast Fourier Transform) , that is, the fused latent features.

[0039] Among them, the watermark information can be any one of the watermark information in the preset watermark library.

[0040] Specifically, when the image watermark generation device performs the fusion of watermark information and , it is necessary to balance invisibility (concealment), robustness (attack resistance), and capacity (embeddable information volume). The fusion schemes provided in this application include but are not limited to: 1. Embed the watermark in the DCT (Discrete Cosine Transform) coefficients of the original latent features.

[0041] 2. Embed the watermark in the high-frequency subbands (LH / HL) of the Haar / Daubechies wavelet decomposition of the original latent features.

[0042] 3. Deep learning-based latent space fusion, encoding the watermark as a small perturbation of the latent vector or injecting the watermark signal into the latent features during the denoising process, etc.

[0043] 4. Traditional image processing-based fusion, directly modifying the last few bits of the pixel values according to the watermark information, or performing block SVD (Singular Value Decomposition) on the image and embedding the watermark in the singular value matrix.

[0044] Among them, before the above watermark information is fused, it is necessary to extract the frequency-domain modulation map of the watermark information, that is, a visualization chart showing the embedding positions of the watermark in the frequency domain.

[0045] Step S13: Obtain the pixel-level visual sensitivity mask parameter of the original image.

[0046] In the embodiments of this application, as Figure 2 shown, the image watermark generation device extracts the pixel-level visual sensitivity mask parameter of the original image , by quantifying the perceptual differences of the human eye to noise or modification in different regions of the image, so as to guide operations such as watermark embedding, image compression, or enhancement to be performed in regions with low visual sensitivity, in order to maximize concealment or minimize visual distortion.

[0047] The pixel-level visual sensitivity mask parameter can assign a sensitivity weight value (usually in the range [0,1]) to each pixel or image block. The higher the value, the more sensitive the human eye is to the distortion in this region. During the adaptive process of guiding watermark embedding, it can guide the embedding of stronger watermark signals in regions with low sensitivity (low weight), improving the capacity and maintaining concealment.

[0048] The method for extracting the pixel-level visual sensitivity mask parameter is as follows: As Figure 2As shown, during the process of generating a watermark image, the image watermark generation device embeds a differentiable JND model (Just Noticeable Difference Model), that is, a differentiable minimum perceptible model, and uses this model to extract pixel-level visual sensitivity mask parameters of the original image. .

[0049] In another embodiment, the image watermark generation device can calculate pixel-level visual sensitivity mask parameters based on local image features. , for example, calculate the brightness sensitivity according to the fact that the human eye is more sensitive to dark area noise (Weber-Fechner's law), or calculate the texture masking degree by measuring the texture complexity through local variance or gradient.

[0050] Step S14: Adjust the watermark intensity of the fused latent features using the pixel-level visual sensitivity mask parameters to obtain adjusted latent features.

[0051] In the embodiment of the present application, the image watermark generation device dynamically adjusts the watermark intensity output by the decoder: , where is a dynamic attenuation coefficient, and the gradient magnitude of the watermark area is constrained to be consistent with the background distribution through adversarial loss.

[0052] Specifically, the expression of the adversarial loss is: . The main principle of training is to adjust the degree of difference in visual features between the watermark image generated after adjusting the watermark intensity and the original image, and the optimization direction is that the watermark image visually tends to be the same as the original image.

[0053] Step S15: Generate an image watermark according to the adjusted latent features.

[0054] Step S16: Generate a watermark image according to the image watermark and the original image.

[0055] In the embodiment of the present application, the image watermark generation device can obtain a visually lossless watermarked image by adding the original image and the robust imperceptible watermark output by the variational decoder pixel by pixel and scaling back to the original size. .

[0056] In this application, an image watermark generation device obtains the original latent features of an original image; fuses watermark information of the same size as the original latent features with the original latent features to obtain fused latent features; obtains the pixel-level visual sensitivity mask parameters of the original image; adjusts the watermark strength of the fused latent features by using the pixel-level visual sensitivity mask parameters to obtain adjusted latent features; generates an image watermark according to the adjusted latent features; and generates a watermarked image according to the image watermark and the original image. Through the above image watermark generation method, the pixel-level visual sensitivity mask parameters are used to adaptively adjust the watermark embedding strength, realizing the global collaborative adaptation of the watermark energy distribution and the visual characteristics of the image, and significantly improving the robustness of the watermark against composite attacks; in this application, the pixel-level visual sensitivity mask parameters are calculated for the original image, realizing the generation of adaptive mask parameters for each original image, and using the adaptive mask parameters to adaptively adjust the watermark embedding strength, significantly improving the robustness of watermark generation on the premise of ensuring imperceptibility.

[0057] This application proposes a watermark embedding architecture that fuses frequency domain attention modulation and human visual masking characteristics (JND). The image is mapped to the latent space through a VAE encoder, the watermark information is dynamically fused using frequency domain features, and a pixel-level visual sensitivity mask is generated in combination with a differentiable JND model to adaptively adjust the watermark embedding strength. This mechanism breaks through the limitations of traditional methods in optimizing in a single dimension of the spatial domain or the frequency domain, realizes the global collaborative adaptation of the watermark energy distribution and the visual characteristics of the image, and significantly improves the robustness of the watermark against composite attacks on the premise of ensuring imperceptibility with PSNR > 40dB.

[0058] Furthermore, this application also provides a solution for imposing composite attacks on the watermarked image. For details, please refer to Figure 3 and Figure 4 , Figure 3 is a schematic flowchart of the second embodiment of the image watermark generation method provided by this application, Figure 4 and

[0059] is a schematic diagram of the scenario of imposing attacks on the watermarked image provided by this application. Figure 3 As shown, the specific steps are as follows:

[0060] In the embodiment of this application, the image watermark generation device uses an object detection model (such as YOLO, Faster R-CNN) to extract key semantic regions (such as faces, texts, LOGOs) in the image.

[0061] Step S22: Generate a binary mask on the image corresponding to the key semantic region to impose an attack on the watermarked image.

[0062] In the embodiment of the present application, the image watermark generation device generates a binary mask in the key semantic region extracted in step S21, and this binary mask can adaptively erase the watermark in the detected region.

[0063] In addition, the image watermark generation device optimizes the mask position through gradient backpropagation, so that the attack is concentrated in the region that most affects the watermark decoding. The image damaged by the mask is expressed as: .

[0064] Furthermore, on the basis of the mask attack, the image watermark generation device further superimposes multi-dimensional composite perturbations to more realistically simulate the comprehensive attack effect under complex scenarios.

[0065] Specifically, since many traditional attack operations (such as geometric transformation, JPEG compression, etc.) are not differentiable themselves and will hinder end-to-end gradient optimization, a differentiable approximation alternative is adopted to achieve the differentiability of the attack process. For example, for geometric deformation operations (such as cropping, scaling, rotation, etc.), a spatial transformation network (STN, Spatial Transformer Network) combined with bilinear interpolation is used to achieve differentiable geometric transformation. STN is a differentiable module that can be embedded in a neural network and can automatically learn to perform spatial transformations (such as rotation, scaling, affine transformation, etc.) on the input image (or feature map) without manual annotation of transformation parameters, which can improve the robustness of the model to spatial changes in the input data (such as translation invariance).

[0066] Another example is that for discrete operations such as JPEG compression, the image watermark generation device uses a differentiable DiffJPEG simulator, that is, a differentiable simulator to approximate the compression effect of real JPEG. DiffJPEG is a differentiable simulator used to simulate the JPEG compression effect, mainly used in deep learning or image processing tasks to help the model better understand and process the distortion caused by JPEG compression during the training process.

[0067] Specifically, DiffJPEG reproduces the JPEG encoding / decoding process (such as discrete cosine transform DCT, quantization, Huffman encoding, etc.) through mathematical modeling, but in a differentiable way so that it can be embedded in a neural network.

[0068] Differentiability: Traditional JPEG is not differentiable (cannot be directly differentiated), while DiffJPEG enables the gradient to be backpropagated through operations such as approximate quantization, which is convenient for joint optimization with deep learning models.

[0069] Step S23: Calculate the perturbation loss using the perturbed image and the watermark image.

[0070] In the embodiment of the present application, the image watermark generation device combines the perturbed image, which is the image damaged by masking, with the watermark image in step S21 to generate a training sample, and uses this training sample to train the image watermark generation model, that is, the differentiable minimum perceptible model.

[0071] Specifically, the image watermark generation device calculates the perturbation loss of the training sample according to the feature difference between the perturbed image and the watermark image, and the loss consists of two parts: 1. Watermark extraction accuracy loss: Ensure that the watermark can still be correctly extracted after perturbation.

[0072] 2. Robustness constraint loss: Force the model to be insensitive to perturbations (such as feature consistency).

[0073] For the above watermark extraction loss, the image watermark generation device can calculate it through the binary cross-entropy loss function or the mean squared error loss function. For the above robustness constraint loss, the image watermark generation device can calculate it through the feature consistency loss function or the adversarial loss function.

[0074] Step S24: Iteratively optimize the differentiable minimum perceptible model according to the perturbation loss.

[0075] In the embodiment of the present application, the image watermark generation device calculates the gradient through the chain rule and updates the model parameters of the differentiable minimum perceptible model.

[0076] Furthermore, the present application also provides a scheme for imposing a composite attack on the watermark image. For details, please refer to Figure 5 and Figure 6 , Figure 5 is the flowchart of the third embodiment of the image watermark generation method provided by the present application, Figure 6 is the scenario diagram of watermark information detection and extraction provided by the present application.

[0077] As Figure 5 shown, the specific steps are as follows: Step S31: Extract the high-dimensional feature map of the watermark image.

[0078] In the embodiment of the present application, the image watermark generation device adopts an improved ResNet-34, replaces the standard convolutional layers in the 3rd and 4th residual stages by embedding deformable convolutions (Deformable Conv) to enhance the adaptability of the model to geometric deformations (such as distortion and stretching). The high-dimensional feature map output by layer4 is used as the shared input of the dual-stream decoding branch, taking into account both semantic abstraction and spatial detail retention.

[0079] In another embodiment, the image watermark generation device may also use the following common models to extract the high-dimensional feature map of the watermark image: VGG (VGG16 / VGG19), ResNet (ResNet50 / 101), EfficientNet, etc. Among them, VGG (VGG16 / VGG19) captures texture with shallow features and semantics with deep features; ResNet (ResNet50 / 101) has a residual structure to avoid gradient disappearance and is suitable for deep features; EfficientNet balances depth / width / resolution and efficiently extracts multi-scale features.

[0080] Step S32: Input the high-dimensional feature map into the two-stream decoding branch to obtain the watermark detection result of the watermark image.

[0081] In the embodiments of the present application, as Figure 6 shown, the image watermark generation device inputs the high-dimensional feature map into the mask detection branch and the watermark decoding branch respectively, and the two branches are responsible for different tasks of watermark detection. Among them, the model parameters of the two-stream decoding branch of the present application are iteratively optimized through the watermark detection result, and the model parameters of the differentiable minimum perceptible model and the model parameters of the two-stream decoding branch are jointly optimized in the same watermark generation and watermark detection stage, that is, the two-stream decoding branch and the differentiable minimum perceptible model can be synchronously trained in the sample training process of the same batch to achieve joint optimization.

[0082] The following continues to introduce the working process of the two-stream decoding branch for watermark detection: Specifically, please continue to refer to Figure 7 , Figure 7 which Figure 5 is the specific process schematic diagram of step S32 of the image watermark generation method shown.

[0083] As Figure 7 shown, the specific steps are as follows: Step S321: Input the high-dimensional feature map into the mask detection branch and the watermark decoding branch respectively.

[0084] In the embodiments of the present application, the image watermark generation device inputs the high-dimensional feature map into the mask detection branch and the watermark decoding branch as Figure 6 shown.

[0085] Step S322: Extract the multi-scale context features of the high-dimensional feature map through the mask detection branch.

[0086] In the embodiments of the present application, for the mask detection flow of the mask detection branch, the image watermark generation device adopts multi-scale context extraction, specifically including: Atrous Spatial Pyramid Pooling (ASPP), designing progressive dilation rates (1 / 3 / 6 / 9) to capture context information with different receptive fields.

[0087] Step S323: Fuse the shallow high-resolution features and deep semantic features in the multi-scale context features to generate attention weights.

[0088] In the embodiments of the present application, the image watermark generation device also uses a Feature Pyramid Network (FPN). Through the top-down path, it fuses the shallow high-resolution features and deep semantic features to improve the mask localization accuracy. Finally, it generates attention weights through spatial attention gating to suppress the background noise area.

[0089] Step S324: Use the attention weights to output the classification result of whether the watermark image contains a watermark.

[0090] In the embodiments of the present application, the mask detection branch finally outputs a binary mask map with the same resolution as the input (only detecting the part with the watermark), and uses BCELoss as the loss function to optimize the mask detection branch.

[0091] Specifically, pred_detection output by the mask detection branch corresponds to the prediction result of watermark detection, that is, the binary classification result (existence / non-existence) of determining whether there is a watermark in the image, or the confidence probability of the existence of the watermark (such as a value between 0 and 1).

[0092] For example, an output of 0.95 means that the image has a 95% probability of containing a watermark.

[0093] Function: Used to verify the detectability of the watermark, especially after adversarial attacks or distortion.

[0094] Usually compared with a threshold (such as >0.5 is determined as "there is a watermark").

[0095] Step S325: Decompose the high-dimensional feature map into low-frequency contours and high-frequency texture components through the watermark decoding branch.

[0096] In the embodiments of the present application, for the watermark decoding flow of the watermark decoding branch, the image watermark generation device first uses a multi-scale residual dense block nested skip connection to aggregate local features at different levels and retain high-frequency details. Then it uses Haar wavelet pyramid decomposition to decompose the feature map into low-frequency contours and high-frequency texture components.

[0097] Step S326: After the low-frequency contour and high-frequency texture components are hierarchically recombined, they are upsampled to output the watermark prediction value.

[0098] In the embodiment of the present application, the image watermark generation device realizes cross-band information complementarity through reversible wavelet upsampling after hierarchically recombining the low-frequency contour and high-frequency texture components. Finally, a prediction value with the number of output channels matching the watermark code length is output, and the loss function uses Binary Cross-Entropy Loss to optimize the watermark decoding branch.

[0099] Specifically, the pred code output by the watermark decoding branch corresponds to the prediction value of the watermark code, that is, in the watermark embedding or extraction stage, the prediction value of the watermark information (usually a binary sequence or matrix) hidden in the image by the model.

[0100] For example, if the original watermark is [1, 0, 1, 1], the pred_code may be the approximate value [0.9, 0.2, 0.8, 0.7] (before quantization) extracted by the model.

[0101] Function: Used to calculate the reconstruction error of the watermark (such as the mean square error with the original watermark).

[0102] In end-to-end training, the robustness of the watermark is improved by optimizing the accuracy of the pred_code.

[0103] In summary, in the overall architecture of the image watermark generation method provided by the present application, the total training loss of the differentiable minimum perceptible model, the mask detection branch, and the watermark decoding branch is specifically the weighted sum of the dual-path loss and the reconstruction loss: .

[0104] The present application proposes a fully end-to-end digital watermark system architecture, which integrates three key links of watermark embedding, attack simulation, and watermark extraction in a unified deep learning framework for joint optimization. This end-to-end design eliminates the performance bottlenecks caused by the independent optimization of each module in traditional methods and realizes the overall optimization of the entire process from the original image to watermark extraction.

[0105] To implement the above image watermark generation method, the present application also proposes an image watermark generation device. For details, please refer to Figure 8 , Figure 8 which is a schematic structural diagram of an embodiment of the image watermark generation device provided by the present application.

[0106] The image watermark generation device 400 in this embodiment includes a processor 41, a memory 42, an input / output device 43, and a bus 44.

[0107] The processor 41, the memory 42, and the input / output device 43 are respectively connected to the bus 44. Program data is stored in the memory 42, and the processor 41 is configured to execute the program data to implement the image watermark generation method described in the above embodiments.

[0108] In the embodiments of the present application, the processor 41 may also be referred to as a CPU (Central Processing Unit). The processor 41 may be an integrated circuit chip with signal processing capabilities. The processor 41 may also be a general-purpose processor, a digital signal processor (DSP, Digital Signal Process), an application specific integrated circuit (ASIC, Application Specific Integrated Circuit), a field programmable gate array (FPGA, Field Programmable Gate Array), or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components. The general-purpose processor may be a microprocessor or the processor 41 may also be any conventional processor, etc.

[0109] The present application also provides a computer storage medium. Please continue to refer to Figure 9 , Figure 9 FIG. is a schematic structural diagram of an embodiment of the computer storage medium provided by the present application. A computer program 61 is stored in the computer storage medium 600. When the computer program 61 is executed by a processor, it is used to implement the image watermark generation method described in the above embodiments.

[0110] When the embodiments of the present application are implemented in the form of software functional units and sold or used as independent products, they can be stored in a computer-readable storage medium. Based on such an understanding, the technical solution of the present application, in essence, or the part that contributes to the prior art, or all or part of the technical solution, can be embodied in the form of a software product. The computer software product is stored in a storage medium and includes several instructions for causing a computer device (which may be a personal computer, a server, or a network device, etc.) or a processor to execute all or part of the steps of the methods described in various embodiments of the present application. The foregoing storage medium includes: USB flash drives, mobile hard disks, read-only memories (ROM, Read-Only Memory), random access memories (RAM, Random Access Memory), magnetic disks, or optical disks, etc., which can store program codes.

[0111] The above are only the embodiments of the present application, and do not thus limit the patent scope of the present application. Any equivalent structure or equivalent process transformation made by using the content of the specification and drawings of the present application, or directly or indirectly applied in other related technical fields, shall be similarly included in the patent protection scope of the present application.

Claims

1. An image watermark generation method, characterized in that, The described image watermark generation method includes: Obtaining the original latent features of the original image; Fusing watermark information of the same size as the original latent features with the original latent features to obtain fused latent features; Obtaining the pixel-level visual sensitivity mask parameters of the original image; Adjusting the watermark intensity of the fused latent features using the pixel-level visual sensitivity mask parameters to obtain adjusted latent features; Generating an image watermark according to the adjusted latent features; Generating a watermarked image according to the image watermark and the original image.

2. The image watermark generation method according to claim 1, wherein: The obtaining of the original latent features of the original image includes: Extracting the latent space representation of the original image using a variational autoencoder; Obtaining the original latent features by performing a Fourier transform on the latent space representation; The fusing of watermark information of the same size as the original latent features with the original latent features to obtain fused latent features includes: Fusing watermark information of the same size as the original latent features with the original latent features and reconstructing the fused latent features through an inverse Fourier transform.

3. The image watermark generation method according to claim 1, wherein: The obtaining of the pixel-level visual sensitivity mask parameters of the original image includes: Inputting the original image into a differentiable minimum perceptible model to extract the pixel-level visual sensitivity mask parameters of the original image; Wherein, the model parameters of the differentiable minimum perceptible model are iteratively optimized through the visual difference value between the watermarked image and the original image.

4. The image watermark generation method according to claim 3, wherein: The image watermark generation method further includes: Extracting the high-dimensional feature map of the watermarked image; Inputting the high-dimensional feature map into a two-stream decoding branch to obtain the watermark detection result of the watermarked image; Wherein, the model parameters of the two-stream decoding branch are iteratively optimized through the watermark detection result, and the model parameters of the differentiable minimum perceptible model and the model parameters of the two-stream decoding branch are jointly optimized in the same watermark generation and watermark detection stage.

5. The image watermark generation method according to claim 4, Characterized in that, Wherein, the two-stream decoding branch includes a mask detection branch and a watermark decoding branch; The inputting of the high-dimensional feature map into the two-stream decoding branch to obtain the watermark detection result of the watermarked image includes; Inputting the high-dimensional feature map into the mask detection branch and the watermark decoding branch respectively; Extracting the multi-scale context features of the high-dimensional feature map through the mask detection branch; Fusing the shallow high-resolution features and the deep semantic features in the multi-scale context features to generate attention weights; Outputting the classification result of whether the watermarked image contains a watermark using the attention weights; Decomposing the high-dimensional feature map into low-frequency contours and high-frequency texture components through the watermark decoding branch; Outputting the watermark prediction value after hierarchically reorganizing and upsampling the low-frequency contours and high-frequency texture components.

6. The image watermark generation method according to claim 4, wherein: The extracting of the high-dimensional feature map of the watermarked image includes: Extract the high-dimensional feature map of the watermark image by embedding deformable convolutions.

7. The image watermark generation method according to claim 1, wherein the image watermark generation method further includes: extracting the key semantic regions in the watermark image; generating a binary mask on the image corresponding to the key semantic regions to impose an attack on the watermark image and generate a perturbed image; calculating a perturbation loss using the perturbed image and the watermark image; iteratively optimizing the differentiable minimum perceptible model according to the perturbation loss.

8. The image watermark generation method according to claim 1 or 7, wherein the image watermark generation method further includes: imposing a discretization attack on the watermark image using a differentiable simulator; and / or, imposing a differentiable geometric transformation attack on the watermark image using a spatial transformation network and bilinear interpolation to generate a perturbed image; calculating a perturbation loss using the perturbed image and the watermark image; iteratively optimizing the differentiable minimum perceptible model according to the perturbation loss.

9. The image watermark generation method according to claim 1, wherein the adjusting the watermark intensity of the fused latent feature using the pixel-level visual sensitivity mask parameter to obtain an adjusted latent feature includes: adjusting the fused latent feature using the pixel-level visual sensitivity mask parameter and a dynamic attenuation system to obtain the adjusted latent feature.

10. An image watermark generation device, characterized in that, The image watermark generation device includes a memory and a processor coupled to the memory; wherein, the memory is used to store program data, and the processor is used to execute the program data to implement the image watermark generation method according to any one of claims 1 to 9.

11. A computer storage medium, characterized in that, The computer storage medium is used to store program data, and the program data, when executed by a computer, is used to implement the image watermark generation method according to any one of claims 1 to 9.

Citation Information

Patent Citations

  • Self-coding color image robust watermark processing method based on visual perception

    CN114841846A

  • Robust image watermarking method and system based on hierarchical attention feature fusion

    CN115908095A

  • Image content protection method for Android platform

    CN116503231A

  • Real-time screen watermarking method and system for tracing leakage of screen shooting

    CN117574336A

  • Watermark information embedding method, device, storage medium and program product

    CN119784566A

Cited By

  • Anti-multi-attack medical image robust watermarking method based on Mamba architecture

    CN120931467A

  • Screen watermark embedding method and device based on deep learning

    CN121526866A

  • Spatial self-adaptive plug-and-play watermarking method and device for image editing traceability

    CN121544448A

  • Image editing traceable spatial adaptive plug-and-play watermarking method and device

    CN121544448B

  • Multi-region watermark embedding method and system

    CN121937275A