Image watermark generation method, image watermark generation device and computer storage medium

By combining the variational autoencoder and the differentiable minimally perceptible model, the watermark strength is adaptively adjusted, which solves the problems of robustness and imperceptibility of digital image watermarks under complex attacks and achieves efficient watermark extraction and detection.

CN120410831BActive Publication Date: 2025-09-12IFLYTEK CO LTD

Patent Information

Application Number
CN202510916718.3
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-07-03
Publication Date
2025-09-12
Estimated Expiration
2045-07-03

AI Technical Summary

Technical Problem

When facing complex attacks, the existing technology has difficulty in achieving both robustness and imperceptibility of digital image watermarks. Especially under attacks such as JPEG compression, geometric deformation and non-uniform cropping, the watermark extraction error rate increases significantly and the robustness decreases sharply.

Method used

A variational autoencoder is used to extract the potential features of the image. Combined with Fourier transform and differentiable least perceptible model, the watermark strength is adaptively adjusted through pixel-level visual sensitivity mask parameters. A multi-scale adversarially coupled watermark framework is constructed to achieve global coordinated adaptation of watermark energy distribution and image visual characteristics.

Benefits of technology

The robustness of the watermark against complex attacks has been significantly improved, the bit error rate has been reduced to below 5%, while maintaining high imperceptibility (PSNR>40dB), and the watermark detection accuracy has been increased by more than 15%.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120410831B_ABST
    Figure CN120410831B_ABST
Patent Text Reader

Abstract

The present application proposes an image watermark generation method, an image watermark generation device, and a computer storage medium. The image watermark generation method includes: obtaining the original latent features of the original image; fusing the watermark information of the same size as the original latent features with the original latent features to obtain fused latent features; obtaining the pixel-level visual sensitivity mask parameters of the original image; adjusting the watermark intensity of the fused latent features using the pixel-level visual sensitivity mask parameters to obtain the adjusted latent features; generating an image watermark based on the adjusted latent features; and generating a watermarked image based on the image watermark and the original image. Through the above-mentioned image watermark generation method, the pixel-level visual sensitivity mask parameters are used to adaptively adjust the watermark embedding strength, thereby achieving global coordinated adaptation of the watermark energy distribution and the image visual characteristics, and significantly improving the robustness of the watermark against complex attacks.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present application relates to the field of digital watermark technology, and in particular to an image watermark generation method, an image watermark generation device, and a computer storage medium. Background Art

[0002] Digital image watermarking technology is a key research area in the field of information hiding. Its core goal is to embed identifying information into a carrier image without compromising visual quality, and to accurately extract that information during transmission and after potential attacks. This technology boasts strong visual imperceptibility and high traceability, making it widely applicable in media copyright management, judicial evidence collection, medical image security, and other fields.

[0003] In adversarial application scenarios (such as social media communication), since it is necessary to resist complex attacks such as value transformation (brightness / contrast adjustment, etc.), geometric transformation (rotation / scaling / cropping, etc.), and image masking (partial occlusion / mosaic), designing a watermark system that takes into account both imperceptibility (PSNR>40dB) and strong robustness (BER<1%) becomes a technical difficulty.

[0004] Traditional mainstream methods are based on frequency domain transforms (such as discrete cosine transform and wavelet transform) and use quantization coding techniques to embed watermarks by modulating coefficients in mid- and high-frequency bands. Although these methods have basic anti-attack capabilities, they can significantly increase the bit error rate of watermark extraction and sharply reduce robustness when faced with coordinated damage from strong JPEG compression (quality factor <50), geometric deformation (rotation >15° / scaling ±20%), non-uniform cropping (random cropping >30% of the area), and mixed attacks (such as compression + cropping + noise). Summary of the Invention

[0005] To solve the above technical problems, the present application proposes an image watermark generation method, an image watermark generation device and a computer storage medium.

[0006] To solve the above technical problems, the present application proposes a method for generating an image watermark, which includes:

[0007] Obtain the original latent features of the original image;

[0008] fusing the watermark information having the same size as the original latent feature with the original latent feature to obtain a fused latent feature;

[0009] Obtaining pixel-level visual sensitivity mask parameters of the original image;

[0010] adjusting the watermark intensity of the fused latent feature using the pixel-level visual sensitivity mask parameter to obtain the adjusted latent feature;

[0011] generating an image watermark based on the adjusted latent features;

[0012] A watermark image is generated according to the image watermark and the original image.

[0013] The obtaining of the original potential features of the original image includes:

[0014] extracting a latent space representation of the original image using a variational autoencoder;

[0015] The latent space representation is subjected to Fourier transform to obtain the original latent features;

[0016] The step of fusing the watermark information having the same size as the original latent feature with the original latent feature to obtain the fused latent feature comprises:

[0017] The watermark information having the same size as the original latent feature is fused with the original latent feature, and the fused latent feature is reconstructed by inverse Fourier transform.

[0018] The step of obtaining pixel-level visual sensitivity mask parameters of the original image includes:

[0019] Inputting the original image into a differentiable minimally perceptible model to extract pixel-level visual sensitivity mask parameters of the original image;

[0020] The model parameters of the differentiable least perceptible model are iteratively optimized through the visual difference value between the watermark image and the original image.

[0021] Wherein, the image watermark method further includes:

[0022] Extracting a high-dimensional feature map of the watermark image;

[0023] Inputting the high-dimensional feature map into the dual-stream decoding branch to obtain a watermark detection result of the watermark image;

[0024] The model parameters of the dual-stream decoding branch are iteratively optimized based on the watermark detection result, and the model parameters of the differentiable minimally perceptible model and the model parameters of the dual-stream decoding branch are jointly optimized in the same watermark generation and watermark detection stages.

[0025] Wherein, the dual-stream decoding branch includes a mask detection branch and a watermark decoding branch;

[0026] The step of inputting the high-dimensional feature map into a dual-stream decoding branch to obtain a watermark detection result of the watermark image comprises:

[0027] Inputting the high-dimensional feature map into the mask detection branch and the watermark decoding branch respectively;

[0028] Extracting multi-scale context features of the high-dimensional feature map through the mask detection branch;

[0029] Fusing shallow high-resolution features and deep semantic features in the multi-scale contextual features to generate attention weights;

[0030] Outputting a classification result of whether the watermarked image contains a watermark using the attention weight;

[0031] Decomposing the high-dimensional feature map into low-frequency contour and high-frequency texture components through the watermark decoding branch;

[0032] The low-frequency contour and high-frequency texture components are hierarchically recombined and up-sampled to output a watermark prediction value.

[0033] The step of extracting a high-dimensional feature map of the watermark image includes:

[0034] A high-dimensional feature map of the watermark image is extracted by embedding deformable convolution.

[0035] Wherein, the image watermark method further includes:

[0036] Extracting key semantic areas in the watermark image;

[0037] Generating a binary mask on the image corresponding to the key semantic area to attack the watermark image and generate a perturbed image;

[0038] Calculating a disturbance loss using the disturbed image and the watermarked image;

[0039] A differentiable least perceptible model is iteratively optimized according to the perturbation loss.

[0040] Wherein, the image watermark method further includes:

[0041] Applying a discretization attack to the watermark image using a differentiable simulator;

[0042] and / or, applying a differentiable geometric transformation attack to the watermarked image using a spatial transformation network and bilinear interpolation to generate a perturbed image;

[0043] Calculating a disturbance loss using the disturbed image and the watermarked image;

[0044] A differentiable least perceptible model is iteratively optimized according to the perturbation loss.

[0045] The step of adjusting the watermark strength of the fused latent feature by using the pixel-level visual sensitivity mask parameter to obtain the adjusted latent feature includes:

[0046] The fused latent feature is adjusted using the pixel-level visual sensitivity mask parameter and a dynamic attenuation system to obtain the adjusted latent feature.

[0047] In order to solve the above technical problems, the present application also proposes an image watermark generation device, which includes a memory and a processor coupled to the memory; wherein the memory is used to store program data, and the processor is used to execute the program data to implement the image watermark generation method as described above.

[0048] In order to solve the above technical problems, the present application further proposes a computer storage medium, which is used to store program data. When the program data is executed by a computer, it is used to implement the above image watermark generation method.

[0049] Compared with the existing technology, the beneficial effects of the present application are: through the above-mentioned image watermark generation method, the pixel-level visual sensitivity mask parameters are used to adaptively adjust the watermark embedding strength, thereby achieving global coordinated adaptation of the watermark energy distribution and the image visual characteristics, and significantly improving the robustness of the watermark against complex attacks; the present application calculates the pixel-level visual sensitivity mask parameters for the original image, realizes the generation of adaptive mask parameters for each original image, and uses the adaptive mask parameters to adaptively adjust the watermark embedding strength, thereby significantly improving the robustness of watermark generation while ensuring imperceptibility. BRIEF DESCRIPTION OF THE DRAWINGS

[0050] In order to more clearly illustrate the technical solutions in the embodiments of the present application, the following briefly introduces the drawings required for describing the embodiments. Obviously, the drawings described below are only some embodiments of the present application. For those skilled in the art, other drawings can be obtained based on these drawings without inventive efforts. Among them:

[0051] Figure 1 This is a flowchart of the first embodiment of the image watermark generation method provided by this application;

[0052] Figure 2 is a schematic diagram of the robust and imperceptible watermark embedding structure provided by this application;

[0053] Figure 3 This is a flow chart of the second embodiment of the image watermark generation method provided by this application;

[0054] Figure 4 This is a schematic diagram of the scenario of attacking a watermarked image provided by this application;

[0055] Figure 5 This is a flowchart of the third embodiment of the image watermark generation method provided by this application;

[0056] Figure 6 This is a schematic diagram of the watermark information detection and extraction scenario provided by this application;

[0057] Figure 7 yes Figure 5 The specific flow diagram of step S32 of the image watermark generation method shown;

[0058] Figure 8 This is a structural diagram of an embodiment of an image watermark generation device provided by this application;

[0059] Figure 9 It is a structural diagram of an embodiment of a computer storage medium provided by this application. DETAILED DESCRIPTION

[0060] The following will be combined with the drawings in the embodiments of this application to clearly and completely describe the technical solutions in the embodiments of this application. Obviously, the embodiments described are only part of the embodiments of this application, not all of the embodiments. Based on the embodiments in this application, all other embodiments obtained by ordinary technicians in this field without making creative efforts are within the scope of protection of this application.

[0061] The terms "first," "second," "third," "fourth," and the like (if any) in the specification and claims of this application and in the accompanying drawings are used to distinguish similar objects and are not necessarily used to describe a particular order or precedence. It should be understood that the terms used in this manner are interchangeable where appropriate, so that the embodiments of the application described herein can, for example, be implemented in an order other than those illustrated or described herein. In addition, the terms "including" and "having," and any variations thereof, are intended to cover non-exclusive inclusions. For example, a process, method, system, product, or apparatus comprising a series of steps or elements is not necessarily limited to those steps or elements explicitly listed, but may include other steps or elements not explicitly listed or inherent to such process, method, product, or apparatus.

[0062] In recent years, deep learning has revolutionized watermarking technology. It automatically learns the distribution of image features through multi-level nonlinear transformations, enabling end-to-end watermark embedding and extraction. Some researchers have employed fully connected neural networks (FCNs) for direct mapping to the pixel domain, optimizing the model by minimizing watermark reconstruction error. While these approaches can capture local texture features, they lack the ability to model the joint spatial-frequency characteristics of the image, making it difficult to balance watermark concealment (PSNR < 38dB) with robustness against attacks.

[0063] Deep convolutional neural networks (CNNs), due to their local perception and weight sharing properties, have become a highly effective solution for watermarking. They extract interwoven spatial and frequency domain features through multi-level convolutional layers, capture multi-scale contextual information through downsampling, and achieve pixel-level watermark localization through upsampling. Compared to traditional frequency-domain methods, CNNs can jointly optimize watermark embedding strength and image structure adaptability, achieving a dynamic balance between frequency-domain energy distribution and spatial-domain visual quality.

[0064] This application innovatively proposes a multi-scale adversarial coupled watermarking framework, introduces an attention mechanism and adversarial training strategy based on a deep convolutional network, and improves the concealment (PSNR>38dB) and anti-composite attack capability (bit error rate <0.5%) of the watermarking system through joint gradient optimization of the embedder and detector.

[0065] Therefore, this application proposes an end-to-end adversarial enhanced digital image watermarking system, focusing on solving the following key technical problems:

[0066] 1) How to achieve adaptive matching of watermark embedding strength with image visual characteristics to improve robustness while maintaining high imperceptibility (PSNR>38dB).

[0067] 2) How to build a differentiable composite attack simulation module so that the system can learn to resist multi-dimensional attacks in real scenarios.

[0068] 3) How to design a global-local coupled decoding architecture to achieve accurate positioning and recovery of watermark information.

[0069] Through end-to-end joint training of the VAE (Variational Autoencoder Encoder-Decoder) encoder and decoder, combined with the frequency domain attention mechanism and the JND (Just Noticeable Difference Perceptual Model) perception model, this application improves the watermark detection accuracy of complex attacks such as JPEG compression and geometric deformation by more than 15% while ensuring visual quality, and reduces the bit error rate to below 5%.

[0070] Please refer to the following for details: Figure 1 and Figure 2 , Figure 1 This is a flow chart of the first embodiment of the image watermark generation method provided by this application. Figure 2 This is a schematic diagram of the robust and imperceptible watermark embedding structure provided by this application.

[0071] The image watermark generation method of the present application is applied to an image watermark generation device, wherein the image watermark generation device of the present application can be a server, a terminal device, or a system composed of a server and a terminal device. Accordingly, the various parts of the image watermark generation device, such as the various units, subunits, modules, and submodules, can be all provided in the server, all provided in the terminal device, or separately provided in the server and the terminal device.

[0072] Furthermore, the server described above may be either hardware or software. When the server is hardware, it may be implemented as a distributed server cluster consisting of multiple servers, or as a single server. When the server is software, it may be implemented as multiple software programs or software modules, such as software or software modules for providing a distributed server, or as a single software program or software module, without further limitation.

[0073] like Figure 1 As shown, the specific steps are as follows:

[0074] Step S11: Obtain original potential features of the original image.

[0075] In the embodiment of the present application, the image watermark generating device extracts the original latent features of the original image.

[0076] In one embodiment, the image watermark generation device can scale any original image z to a fixed size, and then use VAEEncoder to convert the image from pixel space to latent space to obtain the latent space representation of the image, which is denoted as Among them, AEEncoder (Variational Autoencoder) is one of the core components of the Variational Autoencoder (VAE), which is responsible for compressing input data (such as images, text, etc.) into a low-dimensional latent representation and outputting the probability distribution parameters (mean and variance) of the latent space.

[0077] Furthermore, the image watermark generating device can also Use FFT (Fast Fourier Transform) to transform into Fourier space to get , that is, the original latent features.

[0078] In another embodiment, the image watermark generating apparatus may also directly compress the original image into a low-dimensional latent vector through a standard autoencoder.

[0079] In another embodiment, the image watermark generation device may further add a sparsity constraint (such as L1 regularization) to the loss function through a sparse autoencoder to force the original image latent vector to be sparse.

[0080] Step S12: Fusing the watermark information having the same size as the original latent feature with the original latent feature to obtain a fused latent feature.

[0081] In the embodiment of the present application, the image watermark generating device encodes the watermark information into a The same feature representation is followed by Fusion is performed and the latent features after watermarking are reconstructed through IFFT (Inverse Fast Fourier Transform) , that is, fusing latent features.

[0082] The watermark information may be any watermark information in a preset watermark library.

[0083] Specifically, the image watermark generating device is performing watermark information and When integrating, it is necessary to balance concealment (invisibility), robustness (anti-attack capability) and capacity (amount of embedded information). The integration solutions provided in this application include but are not limited to:

[0084] 1. Embed a watermark in the DCT (Discrete Cosine Transform) coefficients of the original latent features.

[0085] 2. Embed the watermark in the high frequency subband (LH / HL) of the original latent feature Haar / Daubechies wavelet decomposition.

[0086] 3. Latent space fusion based on deep learning, encoding the watermark as a small perturbation of the latent vector or injecting the watermark signal into the latent features of the denoising process.

[0087] 4. Based on the fusion of traditional image processing, the last few digits of the pixel value are directly modified according to the watermark information, or the image is divided into blocks and SVD (Singular Value Decomposition) is performed to embed the watermark in the singular value matrix.

[0088] Among them, before the above watermark information is fused, it is necessary to extract the frequency domain modulation diagram of the watermark information, that is, a visual diagram showing the embedding position of the watermark in the frequency domain.

[0089] Step S13: Obtain pixel-level visual sensitivity mask parameters of the original image.

[0090] In the embodiments of this application, Figure 2 As shown, the image watermark generation device extracts the pixel-level visual sensitivity mask parameters of the original image By quantifying the differences in human eye perception of noise or modification in different areas of the image, it guides operations such as watermark embedding, image compression or enhancement to be performed in areas with low visual sensitivity to maximize concealment or minimize visual distortion.

[0091] Pixel-level visual sensitivity mask parameters Each pixel or image block can be assigned a sensitivity weight value (usually in the range [0,1]), where higher values ​​indicate that the human eye is more sensitive to distortion in that area. In the adaptive process of guiding watermark embedding, it is possible to guide the embedding of stronger watermark signals in less sensitive areas (low weights), thereby improving capacity while maintaining concealment.

[0092] Extract pixel-level visual sensitivity mask parameters The way is as follows:

[0093] like Figure 2 As shown in the figure, the image watermark generation device embeds a differentiable JND model (Just Noticeable Difference Model) in the process of generating the watermark image, which is a differentiable minimum perceptible model, and uses this model to extract the pixel-level visual sensitivity mask parameters of the original image. .

[0094] In another embodiment, the image watermark generation device can calculate pixel-level visual sensitivity mask parameters based on local image features. For example, the brightness sensitivity is calculated based on the fact that the human eye is more sensitive to noise in dark areas (Weber-Fechner law), or the texture masking is calculated by measuring the texture complexity through local variance or gradient.

[0095] Step S14: using the pixel-level visual sensitivity mask parameter to adjust the watermark intensity of the fused latent feature to obtain the adjusted latent feature.

[0096] In the embodiment of the present application, the image watermark generation device dynamically adjusts the watermark strength output by the decoder: ,in is the dynamic attenuation coefficient, and the adversarial loss is used to constrain the gradient amplitude of the watermark area to be consistent with the background distribution.

[0097] Specifically, the adversarial loss is expressed as: The main principle of training is to adjust the watermark intensity to determine the difference in visual features between the watermarked image and the original image. The optimization direction is to make the watermarked image visually similar to the original image.

[0098] Step S15: generating an image watermark according to the adjusted latent features.

[0099] Step S16: Generate a watermark image based on the image watermark and the original image.

[0100] In the embodiment of the present application, the image watermark generation device adds the original image and the robust imperceptible watermark output by the variational decoder pixel by pixel and scales it back to the original size to obtain a visually lossless watermarked image. .

[0101] In the present application, an image watermark generation device obtains the original latent features of the original image; fuses watermark information of the same size as the original latent features with the original latent features to obtain fused latent features; obtains pixel-level visual sensitivity mask parameters of the original image; uses the pixel-level visual sensitivity mask parameters to adjust the watermark strength of the fused latent features to obtain adjusted latent features; generates an image watermark based on the adjusted latent features; and generates a watermarked image based on the image watermark and the original image. Through the above-mentioned image watermark generation method, the pixel-level visual sensitivity mask parameters are used to adaptively adjust the watermark embedding strength, thereby achieving global coordinated adaptation of the watermark energy distribution and the image visual characteristics, significantly improving the robustness of the watermark against complex attacks; the present application calculates the pixel-level visual sensitivity mask parameters for the original image, realizes the generation of adaptive mask parameters for each original image, and uses the adaptive mask parameters to adaptively adjust the watermark embedding strength, significantly improving the robustness of watermark generation while ensuring imperceptibility.

[0102] This application proposes a watermark embedding architecture that combines frequency-domain attention modulation with the human visual masking property (JND). This architecture maps the image to a latent space using a VAE encoder, dynamically integrates the watermark information using frequency-domain features, and combines it with a differentiable JND model to generate a pixel-level visual sensitivity mask to adaptively adjust the watermark embedding strength. This mechanism overcomes the limitations of traditional methods that optimize in a single dimension in the spatial or frequency domain, achieving a global coordinated adaptation of the watermark energy distribution and the visual characteristics of the image. This significantly improves the watermark's robustness against complex attacks while maintaining an imperceptible PSNR of >40dB.

[0103] Furthermore, this application also provides a solution for applying a composite attack to a watermark image. Figure 3 and Figure 4 , Figure 3 This is a flow chart of the second embodiment of the image watermark generation method provided by this application. Figure 4 This is a schematic diagram of a scenario in which an attack is carried out on a watermarked image, as provided in this application.

[0104] like Figure 3 As shown, the specific steps are as follows:

[0105] Step S21: extracting key semantic areas in the watermark image.

[0106] In an embodiment of the present application, the image watermark generation device uses a target detection model (such as YOLO, Faster R-CNN) to extract key semantic areas (such as faces, text, and logos) in an image.

[0107] Step S22: Generate a binary mask on the image corresponding to the key semantic area to attack the watermark image.

[0108] In the embodiment of the present application, the image watermark generation device generates a binary mask in the key semantic area extracted in step S21, and the binary mask can adaptively erase the watermark of the detected area.

[0109] In addition, the image watermark generation device optimizes the mask position through gradient backpropagation, so that the attack is concentrated in the area that most affects the watermark decoding. The image damaged by the mask is represented as: .

[0110] Furthermore, the image watermark generation device further superimposes multi-dimensional composite disturbances on the basis of mask attacks to more realistically simulate the comprehensive attack effects in complex scenarios.

[0111] Specifically, because many traditional attack operations (such as geometric transformations and JPEG compression) are inherently non-differentiable and hinder end-to-end gradient optimization, differentiable approximate alternatives are employed to achieve differentiability in the attack process. For example, for geometric deformation operations (such as cropping, scaling, and rotation), a spatial transformer network (STN) is combined with bilinear interpolation to achieve differentiable geometric transformations. STN is a differentiable module that can be embedded in a neural network. It can automatically learn to perform spatial transformations (such as rotation, scaling, and affine transformations) on the input image (or feature map) without the need for manual annotation of transformation parameters. This improves the model's robustness to spatial variations in the input data (such as translation invariance).

[0112] For example, for discrete operations like JPEG compression, image watermark generation devices use a differentiable DiffJPEG simulator, a differentiable simulator that approximates the compression effects of real JPEGs. DiffJPEG is a differentiable simulator used to simulate the effects of JPEG compression. It is primarily used in deep learning or image processing tasks to help models better understand and handle the distortion introduced by JPEG compression during training.

[0113] Specifically, DiffJPEG reproduces the JPEG encoding / decoding process (such as discrete cosine transform DCT, quantization, Huffman coding and other steps) through mathematical modeling, but implements it in a differentiable way so that it can be embedded in a neural network.

[0114] Differentiability: Traditional JPEG is non-differentiable (cannot be directly differentiated), while DiffJPEG uses operations such as approximate quantization to allow gradients to be backpropagated, facilitating joint optimization with deep learning models.

[0115] Step S23: Calculate the disturbance loss using the disturbed image and the watermarked image.

[0116] In an embodiment of the present application, the image watermark generation device combines the masked image, i.e., the disturbed image, with the watermark image in step S21 to generate a training sample, and uses the training sample to train the image watermark generation model, i.e., the differentiable minimum perceptible model.

[0117] Specifically, the image watermark generation device calculates the perturbation loss of the training sample based on the feature difference between the perturbed image and the watermark image. The loss is composed of two parts:

[0118] 1. Loss of watermark extraction accuracy: Ensure that the watermark can still be correctly extracted after the disturbance.

[0119] 2. Robustness constraint loss: forces the model to be insensitive to perturbations (such as feature consistency).

[0120] For the above watermark extraction loss, the image watermark generation device can calculate it through the binary cross entropy loss function or the mean square error loss function. For the above robustness constraint loss, the image watermark generation device can calculate it through the feature consistency loss function or the adversarial loss function.

[0121] Step S24: iteratively optimize the differentiable minimum perceptible model according to the disturbance loss.

[0122] In an embodiment of the present application, the image watermark generation device calculates the gradient and updates the model parameters of the differentiable least perceptible model through the chain rule.

[0123] Furthermore, this application also provides a solution for applying a composite attack to a watermark image. Figure 5 and Figure 6 , Figure 5 This is a flow chart of the third embodiment of the image watermark generation method provided by this application. Figure 6 This is a schematic diagram of the scenario of watermark information detection and extraction provided by this application.

[0124] like Figure 5 As shown, the specific steps are as follows:

[0125] Step S31: extracting a high-dimensional feature map of the watermark image.

[0126] In this embodiment, the image watermark generation device uses an improved ResNet-34. By embedding deformable convolutions (DCEs) in place of the standard convolutional layers in the third and fourth residual stages, the model's adaptability to geometric deformations (such as distortion and stretching) is enhanced. The high-dimensional feature map output by layer 4 serves as the shared input for the dual-stream decoding branch, balancing semantic abstraction with the preservation of spatial detail.

[0127] In another embodiment, the image watermark generation device may also use the following commonly used models to extract high-dimensional feature maps of the watermark image: VGG (VGG16 / VGG19), ResNet (ResNet50 / 101), EfficientNet, etc. Among them, VGG (VGG16 / VGG19) shallow features capture texture and deep features capture semantics; ResNet (ResNet50 / 101) residual structure avoids gradient vanishing and is suitable for deep features; EfficientNet balances depth, width, and resolution to efficiently extract multi-scale features.

[0128] Step S32: input the high-dimensional feature map into the dual-stream decoding branch to obtain the watermark detection result of the watermark image.

[0129] In the embodiments of this application, Figure 6As shown, the image watermark generation device inputs the high-dimensional feature map into the mask detection branch and the watermark decoding branch, respectively. The two branches are responsible for different watermark detection tasks. Among them, the model parameters of the dual-stream decoding branch of the present application are iteratively optimized based on the watermark detection results, and the model parameters of the differentiable minimally perceptible model and the model parameters of the dual-stream decoding branch are jointly optimized during the same watermark generation and watermark detection stages. That is, the dual-stream decoding branch and the differentiable minimally perceptible model can be trained synchronously during the same batch of sample training to achieve joint optimization.

[0130] The following describes the working process of watermark detection in the dual-stream decoding branch:

[0131] Please continue to read for details Figure 7 , Figure 7 yes Figure 5 The specific flow chart of step S32 of the image watermark generation method is shown.

[0132] like Figure 7 As shown, the specific steps are as follows:

[0133] Step S321: input the high-dimensional feature map into the mask detection branch and the watermark decoding branch respectively.

[0134] In the embodiment of the present application, the image watermark generation device inputs the high-dimensional feature map into the following Figure 6 The mask detection branch and watermark decoding branch are shown.

[0135] Step S322: extracting multi-scale context features of the high-dimensional feature map through the mask detection branch.

[0136] In an embodiment of the present application, for the mask detection stream of the mask detection branch, the image watermark generation device adopts multi-scale context extraction, specifically including: Atrous Spatial Pyramid Pooling (ASPP), designing progressive dilation rates (1 / 3 / 6 / 9), and capturing context information of different receptive fields.

[0137] Step S323: Fuse the shallow high-resolution features and deep semantic features in the multi-scale contextual features to generate attention weights.

[0138] In this embodiment of the present application, the image watermark generation device uses a feature pyramid network (FPN) to fuse shallow high-resolution features with deep semantic features in a top-down manner to improve mask positioning accuracy. Finally, spatial attention gating is used to generate attention weights to suppress background noise areas.

[0139] Step S324: Use the attention weight to output the classification result of whether the watermark image contains a watermark.

[0140] In the embodiment of the present application, the mask detection branch finally outputs a binary mask image with the same resolution as the input (only the part with the watermark is detected), and uses BCELoss as the loss function to optimize the mask detection branch.

[0141] Specifically, the pred detection output by the mask detection branch corresponds to the prediction result of watermark detection, that is, the binary classification result of whether the image contains a watermark (presence / absence), or the confidence probability of the existence of the watermark (such as a value between 0 and 1).

[0142] For example, an output of 0.95 means there is a 95% probability that the image contains a watermark.

[0143] Purpose: Used to verify the detectability of watermarks, especially after attacks or distortion.

[0144] It is usually compared with a threshold (e.g. > 0.5 is considered "watermark present").

[0145] Step S325: decompose the high-dimensional feature map into low-frequency contour and high-frequency texture components through the watermark decoding branch.

[0146] In this embodiment of the present application, for the watermark decoding branch's watermark decoding stream, the image watermark generation device first uses multi-scale residual dense block nested skip connections to aggregate local features at different levels, preserving high-frequency details. It then uses Haar wavelet pyramid decomposition to decompose the feature map into low-frequency contours and high-frequency texture components.

[0147] Step S326: The low-frequency contour and high-frequency texture components are hierarchically recombined and up-sampled to output a watermark prediction value.

[0148] In this embodiment, the image watermark generation device hierarchically recombines low-frequency contour and high-frequency texture components, then uses reversible wavelet upsampling to achieve cross-band information complementarity. The final output is a prediction value whose number of channels matches the watermark code length. Binary Cross-Entropy Loss is used as the loss function to optimize the watermark decoding branch.

[0149] Specifically, the pred code output by the watermark decoding branch corresponds to the predicted value of the watermark encoding, that is, the model's predicted value of the watermark information (usually a binary sequence or matrix) hidden in the image during the watermark embedding or extraction stage.

[0150] For example, if the original watermark is [1, 0, 1, 1], pred_code may be the approximate value [0.9, 0.2, 0.8, 0.7] extracted by the model (before quantization).

[0151] Function: Used to calculate the reconstruction error of the watermark (such as the mean square error with the original watermark).

[0152] In end-to-end training, the watermark robustness is improved by optimizing the accuracy of pred_code.

[0153] In summary, in the overall architecture of the image watermark generation method provided by this application, the total training loss of the differentiable minimally perceptible model, the mask detection branch, and the watermark decoding branch is specifically the weighted sum of the two-path loss and the reconstruction loss: .

[0154] This application proposes a fully end-to-end digital watermarking system architecture that integrates the three key steps of watermark embedding, attack simulation, and watermark extraction into a unified deep learning framework for joint optimization. This end-to-end design eliminates the performance bottlenecks caused by the independent optimization of each module in traditional methods, achieving an optimal process from the original image to watermark extraction.

[0155] In order to implement the above-mentioned image watermark generation method, this application also proposes an image watermark generation device, please refer to Figure 8 , Figure 8 It is a structural diagram of an embodiment of an image watermark generation device provided by this application.

[0156] The image watermark generating apparatus 400 of this embodiment includes a processor 41 , a memory 42 , an input / output device 43 , and a bus 44 .

[0157] The processor 41 , memory 42 , and input / output device 43 are respectively connected to a bus 44 . The memory 42 stores program data, and the processor 41 is used to execute the program data to implement the image watermark generation method described in the above embodiment.

[0158] In the embodiments of the present application, the processor 41 may also be referred to as a CPU (Central Processing Unit). The processor 41 may be an integrated circuit chip with signal processing capabilities. The processor 41 may also be a general-purpose processor, a digital signal processor (DSP), an application-specific integrated circuit (ASIC), a field-programmable gate array (FPGA), or other programmable logic device, discrete gate or transistor logic device, or discrete hardware component. A general-purpose processor may be a microprocessor, or the processor 41 may be any conventional processor.

[0159] This application also provides a computer storage medium, please continue to refer to Figure 9 , Figure 9 6 is a schematic structural diagram of an embodiment of a computer storage medium provided in the present application. The computer storage medium 600 stores a computer program 61. When the computer program 61 is executed by a processor, it is used to implement the image watermark generation method of the above embodiment.

[0160] When the embodiments of the present application are implemented in the form of software functional units and sold or used as independent products, they can be stored in a computer-readable storage medium. Based on this understanding, the technical solution of the present application, or the part that contributes to the prior art, or all or part of the technical solution can be embodied in the form of a software product. The computer software product is stored in a storage medium and includes several instructions for enabling a computer device (which can be a personal computer, server, or network device, etc.) or a processor to execute all or part of the steps of the method described in each embodiment of the present application. The aforementioned storage medium includes various media that can store program code, such as a USB flash drive, a mobile hard disk, a read-only memory (ROM), a random access memory (RAM), a magnetic disk, or an optical disk.

[0161] The above description is only an implementation method of the present application and does not limit the patent scope of the present application. Any equivalent structure or equivalent process transformation made using the contents of the description and drawings of this application, or directly or indirectly applied in other related technical fields, are also included in the patent protection scope of the present application.

Claims

1. A method for generating an image watermark, characterized in that: The image watermark generation method comprises: Obtain the original latent features of the original image; fusing the watermark information having the same size as the original latent feature with the original latent feature to obtain a fused latent feature; Obtaining pixel-level visual sensitivity mask parameters of the original image; adjusting the watermark intensity of the fused latent feature using the pixel-level visual sensitivity mask parameter to obtain the adjusted latent feature; generating an image watermark based on the adjusted latent features; generating a watermark image according to the image watermark and the original image; The obtaining of pixel-level visual sensitivity mask parameters of the original image includes: Inputting the original image into a differentiable minimally perceptible model to extract pixel-level visual sensitivity mask parameters of the original image; wherein the model parameters of the differentiable minimum perceptible model are iteratively optimized by the visual difference value between the watermark image and the original image; The image watermark generation method further includes: Extracting a high-dimensional feature map of the watermark image; Inputting the high-dimensional feature map into the dual-stream decoding branch to obtain a watermark detection result of the watermark image; The model parameters of the dual-stream decoding branch are iteratively optimized based on the watermark detection result, and the model parameters of the differentiable minimally perceptible model and the model parameters of the dual-stream decoding branch are jointly optimized in the same watermark generation and watermark detection stages.

2. The image watermark generation method according to claim 1, characterized in that: The obtaining of original potential features of the original image includes: extracting a latent space representation of the original image using a variational autoencoder; The latent space representation is subjected to Fourier transform to obtain the original latent features; The step of fusing the watermark information having the same size as the original latent feature with the original latent feature to obtain the fused latent feature comprises: The watermark information having the same size as the original latent feature is fused with the original latent feature, and the fused latent feature is reconstructed by inverse Fourier transform.

3. The image watermark generation method according to claim 1, It is characterized in that Wherein, the dual-stream decoding branch includes a mask detection branch and a watermark decoding branch; The step of inputting the high-dimensional feature map into a dual-stream decoding branch to obtain a watermark detection result of the watermark image comprises: Inputting the high-dimensional feature map into the mask detection branch and the watermark decoding branch respectively; Extracting multi-scale context features of the high-dimensional feature map through the mask detection branch; Fusing shallow high-resolution features and deep semantic features in the multi-scale contextual features to generate attention weights; Outputting a classification result of whether the watermarked image contains a watermark using the attention weight; Decomposing the high-dimensional feature map into low-frequency contour and high-frequency texture components through the watermark decoding branch; The low-frequency contour and high-frequency texture components are hierarchically recombined and up-sampled to output a watermark prediction value.

4. The image watermark generation method according to claim 1, characterized in that: The extracting the high-dimensional feature map of the watermark image comprises: A high-dimensional feature map of the watermark image is extracted by embedding deformable convolution.

5. The image watermark generation method according to claim 1, characterized in that: The image watermark generation method further includes: Extracting key semantic areas in the watermark image; Generating a binary mask on the image corresponding to the key semantic area to attack the watermark image and generate a perturbed image; Calculating a disturbance loss using the disturbed image and the watermarked image; A differentiable least perceptible model is iteratively optimized according to the perturbation loss.

6. The image watermark generation method according to claim 1 or 5, characterized in that: The image watermark generation method further includes: Applying a discretization attack to the watermark image using a differentiable simulator; and / or, applying a differentiable geometric transformation attack to the watermarked image using a spatial transformation network and bilinear interpolation to generate a perturbed image; Calculating a disturbance loss using the disturbed image and the watermarked image; A differentiable least perceptible model is iteratively optimized according to the perturbation loss.

7. The image watermark generation method according to claim 1, characterized in that: The step of adjusting the watermark strength of the fused latent feature by using the pixel-level visual sensitivity mask parameter to obtain the adjusted latent feature comprises: The fused latent feature is adjusted using the pixel-level visual sensitivity mask parameter and the dynamic attenuation system to obtain the adjusted latent feature.

8. An image watermark generation device, characterized in that: The image watermark generating device includes a memory and a processor coupled to the memory; The memory is used to store program data, and the processor is used to execute the program data to implement the image watermark generation method according to any one of claims 1 to 7.

9. A computer storage medium, characterized in that The computer storage medium is used to store program data, and when the program data is executed by a computer, it is used to implement the image watermark generation method according to any one of claims 1 to 7.

Citation Information

Patent Citations

  • Real-time screen watermarking method and system for tracing leakage of screen shooting

    CN117574336A

  • Watermark information embedding method, device, storage medium and program product

    CN119784566A

Cited By

  • Method and system for generating image watermark based on potential variable optimization of diffusion model

    CN121458514A