Cross-medium physical image watermarking method and system based on micro-physical simulation
By constructing a virtual printing scene in rendering software and generating a dataset using a differentiable physics simulator, and fine-tuning the decoder to explicitly model the physical interactions of geometry, materials, and lighting, the problem of watermark signal loss in existing technologies is solved, achieving stable watermark extraction and efficient data generation in real-world scenarios.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2026-02-28
- Publication Date
- 2026-03-27
AI Technical Summary
Existing technologies lack explicit decoupling and simulation of geometry, material, and lighting in cross-media scenarios, resulting in easy loss of watermark signals, insufficient robustness, and difficulty in extracting robust features that remain unchanged in the environment.
A cross-media physical image watermarking method based on differentiable physics simulation is adopted. By constructing a virtual printing scene in rendering software, generating paired supervised datasets using a differentiable physics simulator, fine-tuning the decoder to decode the watermark data, and explicitly modeling the physical interaction of geometry, materials and lighting.
It significantly improves robustness in real-world scenarios, achieves end-to-end physical perception optimization, enhances the generalization ability of the decoder, provides an efficient data generation platform, and supports stable watermark extraction under different material and lighting conditions.
Smart Images

Figure CN121746153A_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the fields of information security and computer vision image processing technology, specifically to a cross-media physical image watermarking method and system based on differentiable physics simulation. Background Technology
[0002] With the rapid development of digital multimedia technology, copyright protection and traceability authentication of digital images have become increasingly important. Digital watermarking technology protects content by embedding imperceptible information into the carrier image. In practical applications, watermarked images often involve cross-media transmission, especially in the print-photograph process. Unlike compression or cropping in the pure digital domain, print-photographing is not just a simple signal conversion, but a complex physical imaging process. When a printed image is photographed, its final pixel representation is actually a comprehensive result of multiple factors, such as the shooting angle, paper surface material, and ambient light, after complex optical interactions and coupling. For example, the tilt of the shooting angle not only causes perspective distortion but also directly changes the reflective properties of the paper surface, such as producing highlights or shadows; different paper roughness also significantly affects the ink's color rendering and diffuse reflection behavior. However, most existing neural network watermarking technologies for cross-media scenarios fail to accurately reflect this real physical link, resulting in the watermark signal being easily lost. This is mainly due to the failure to effectively decouple and simulate some core physical factors. Existing simulation approaches mainly include: (1) Noise layer simulation based on two-dimensional image transformation (Encoder-Noise Layer-Decoder architecture). This type of method is currently the most mainstream approach (existing technologies using this method include HiDDeN, StegaStamp, and RIHOOP). This type of method inserts a differentiable "noise layer" between the encoder and decoder, forming an "encoder-noise layer-decoder" architecture, and simulates physical degradation by superimposing Gaussian noise, JPEG compression, random cropping, perspective transformation, or color dithering. This method essentially simplifies the physical world into a set of independently sampled two-dimensional random noise, ignoring the causal coupling between physical factors. For example, a change in shooting angle will inevitably lead to a change in surface illumination reflection, and simple two-dimensional transformation cannot simulate this linkage. Due to the lack of unified physical modeling of geometry, material, and lighting, the model trained by this method is difficult to learn environmentally invariant features, and the decoding robustness will significantly decrease when facing complex and varied lighting in the real world (such as specular reflection and shadow) and different paper materials (such as white cardboard and kraft paper). (2) Based on traditional rendering engines or hand-designed physical simulations, some methods in this approach attempt to simulate halftoning or specific geometric distortions using hand-designed pipelines. These traditional methods can usually only handle single degradation factors and cannot simulate complex appearance changes that depend on the viewpoint. More importantly, traditional graphics rendering pipelines or hand-designed simulation processes are usually non-differentiable, which blocks gradient backpropagation in neural networks, making it impossible for the encoder and decoder to perform end-to-end joint optimization through simulators, severely limiting the upper limit of system performance. (3) Simulation based on unpaired image transformation and generative adversarial networks (GANs).To address the data gap problem, some studies have used CycleGAN or adversarial perturbations to achieve distributional alignment between simulated and real images. While these methods are effective in terms of visual fidelity, they implicitly treat the physical world as random noise, lacking explicit decoupling and precise control over physical environment parameters (such as angle, distance, light intensity, and material reflectivity). Due to the large amount of spurious statistical correlations in the generated samples and the lack of physical constraints, the models have insufficient generalization ability when faced with unseen physical scenes. (4) Supervised learning simulation based on paired data. Some studies (such as CDTF) have attempted to collect real "digital original image - captured image" paired data and train distortion field models through supervised learning. These methods heavily rely on pixel-level aligned paired training data. However, cross-media physical processes are not repeatable (such as the slight shaking during handheld shooting), and obtaining large-scale, high-precision aligned data is extremely expensive and difficult. In addition, models trained on specific datasets are prone to overfitting, resulting in poor generalization ability when faced with shooting devices or environments outside the training set. In summary, existing technologies lack a method capable of decoupling the three core elements of geometry, material, and lighting, and accurately simulating the physical link of real printing-photography in a differentiable manner. This makes it difficult for existing watermarking systems to extract environmentally invariant robust features when faced with complex material properties, variable lighting environments, and dynamic shooting perspectives in the real physical world, thus limiting their practical application value. Therefore, how to construct a cross-media watermarking method with explicit physical property decoupling capability and end-to-end differentiable optimization support has become a key technical challenge that urgently needs to be solved in this field. Summary of the Invention
[0003] The technical problem to be solved by the present invention is to provide a cross-media physical image watermarking method and system based on differentiable physical simulation, which addresses the above-mentioned problems of the prior art. The present invention aims to solve the problems of insufficient robustness in real-world scenarios, insufficient adaptability to physical environments, and difficulty in accurately extracting watermark features that are not affected by the environment in the prior art.
[0004] To solve the above-mentioned technical problems, the technical solution adopted by the present invention is as follows: A cross-media physical image watermarking method based on differentiable physics simulation includes the following steps: S101: Import or generate the paper material texture map of the target paper in the rendering software, construct a virtual printing scene, and attach the physical image sample containing the watermark data to the surface of the paper material texture map to form a rendered image. Export the rendered image in NeRF standard format and a pairwise supervised dataset composed of different camera parameters. S102, using a differentiable physics simulator to generate batches of simulated printed image samples from paired supervised datasets; S103 uses a batch of simulated printed image samples to fine-tune the pre-trained decoder. S104, input the physical image into the finely tuned decoder to decode the watermark data in the physical image. If the decoded watermark data is inconsistent with the expected watermark data or the decoding fails, the input physical image is determined to be a forged physical image.
[0005] Optionally, in step S101, when importing or generating a paper texture map of the target paper in the rendering software, generating the paper texture map of the target paper refers to generating a paper texture map of the target paper based on the set albedo, roughness, metallicity and ink absorption rate of the paper material.
[0006] Optionally, in step S101, the generation of the physical image sample containing watermark data includes: normalizing and size standardizing the original host image sample, inputting the watermark data into a pre-trained encoder, and embedding the watermark data into the original host image sample through the encoder to obtain the physical image sample containing watermark data; the watermark data is binary data of a specified bit length.
[0007] Optionally, in step S101, when constructing the virtual printing scene and pasting the input physical image sample containing watermark data onto the paper material texture surface to form a rendered image, this includes configuring an outdoor HDR lighting environment and setting the camera distance, scaling the paper size so that the center of the paper material texture is located at the origin of the world coordinate system, and the paper size is located at the origin of the world coordinate system. Within the space, set the camera distance to determine the preset proportion of the paper texture image within the camera's field of view.
[0008] Optionally, in step S102, the differentiable physics simulator is a pre-trained Conditional Neurophysical Simulator (CNPS). The CNPS includes a geometry field module, a material field module, and a relighting module. The geometry field module consists of a radiation field module and a reflection field module. The radiation field module is used to predict the RGB color at the corresponding viewpoint based on the pose embedding features of the input rendered image and camera parameters. The reflection field module is used to predict the specular reflection RGB color and initial material parameters at the corresponding viewpoint based on the pose embedding features of the input rendered image and camera parameters. The material parameters include some or all of albedo, roughness, and metallicity, and infer the 3D mesh of the geometric surface of the rendered image. The material field module, based on the 3D mesh model output by the geometry field module, first obtains the pose embedding features of the input rendered image and camera parameters, as well as the input lighting descriptor, to infer and generate RGB images and more refined material parameters. The relighting module is used to synthesize simulated printed image samples at the corresponding viewpoint through differentiable rendering based on the RGB images and more refined material parameters.
[0009] Optionally, both the radiation field module and the reflection field module are multilayer perceptron (MLP); the material field module consists of multiple MLPs, each MLP corresponding to a material parameter, used to construct a conditional vector based on the pose embedding features of the rendered image and camera parameters, and the concatenation of lighting descriptors. The system generates more refined material parameters through inference. The lighting descriptor is the lighting parameter used by the rendering software to generate the rendered image, including ambient light intensity and proxy value of the main light direction. The relighting module includes an image encoder, a conditional encoder, and a multi-scale U-Net decoder. The image encoder includes a convolution module and a downsampling module connected in sequence. The input features of the image encoder are the H×W×8 RGB image output by the material field module and a more refined material parameter map, where H and W are the height and width of the RGB image, respectively. The conditional encoder includes a multilayer perceptron (MLP) and a multi-scale shaping module connected in sequence. The input features of the MLP are a multi-dimensional conditional vector formed by concatenating the pose embedding features of the sampling points and the features of the HDR lighting environment obtained through second-order spherical harmonic encoding. The multi-scale shaping module is used to deform the output features of the MLP to generate features of multiple scales as input to the multi-scale U-Net decoder. The multi-scale U-Net decoder is used to generate simulated printed image samples from corresponding viewpoints based on the input features of multiple scales.
[0010] Optionally, the calculation function expression for the pose embedding features of the rendered image and camera parameters is: ; ; in, Sampling points in a 3D geometric field for the input rendered image Latent geometric features. , and These are the axis vector factors of the X, Y, and Z axes of the 3D scene tensor, respectively. , and They are 3D scene tensors The axis matrix factors of the YZ axis, XZ axis, and XY axis. For splicing operations, Sampling points for camera parameters in a three-dimensional geometric field pose embedding features The dimension for encoding Fourier features.
[0011] Optionally, the loss function used by the geometric field module during training has the following expression: ; ; ; ; ; ; in, This is the loss function used by the geometry field module during training. For roughness-perceived reconstruction loss, For Eikonal's loss, Gaussian smoothing loss, For Hessian losses, For TV regularization loss, To cover up the damage, To stabilize the loss, , , and All are weighted parameters. For roughness, , and These are the output colors of the radiation field, the output colors of the reflection field, and the colors of the actual image, respectively. Represents the L1 norm; This indicates taking the mean of the set of points in the sampling space; Sampling points in a three-dimensional geometric field Signed distance field at location The gradient vector; Sampling points in a three-dimensional geometric field Signed distance field at location The gradient vector; It is the displacement vector; It is an L2 norm; Represents sampling points in a three-dimensional geometric field The Hessian matrix at that location; Denotes the Frobenius norm. The input rendered image, For a 3D scene tensor, The total variation norm is represented; the loss function used by the material field module during training has the following expression: ; ; in, This is the loss function used by the material field module during training. These are the weighting coefficients. For material regularization loss, , and Sampling points in the three-dimensional geometric field The material parameters include albedo, roughness, and metallicity; the loss function used by the relighting module during training has the following expression: ; ; in, This is the loss function used by the relighting module during training. The pixel-level L1 reconstruction loss is the predicted image generated by the relighting module. With the target real image The average of the absolute errors between pixels. In order to perceive loss, This represents a pre-trained feature extraction network. For the sample size, This represents the Euclidean distance in the feature space.
[0012] Optionally, in step S103, when fine-tuning the pre-trained decoder using a batch of simulated printed image samples, the fine-tuning includes employing an environment-consistent distillation strategy using both the student decoder and the teacher decoder: S201, initialize the pre-trained decoder as the student decoder and the teacher decoder; S202 uses batches of simulated printed image samples to train both the student and teacher decoders, and in each iteration of training, it employs a triple loss mechanism. The network parameters of the student decoder are updated to achieve environment-consistent distillation and learning environment-independent watermark features. The network parameters of the teacher decoder are updated using exponential moving average (EMA). This triple loss... The expression for the computation function is: ; ; ; ; in, , and For weight parameters, This represents a loss of visual consistency within the environment. For the loss of invariance between environments, Entropy loss weighted by confidence level For a set of perspectives, From the perspective The weight, Let KL divergence be the KL divergence. For the predicted bit probability distribution of the student decoder, Predict the probability distribution from the anchor point perspective of the teacher decoder. From the anchor point perspective, This serves as the reference distribution separator in KL divergence calculations. For the environment and perspective The confidence weights are as follows: The output entropy of the student decoder, For student decoders in the environment and perspective The probability distribution of the next predicted bit. For student decoders in the environment and perspective The probability distribution of the next predicted bit. For the environment and perspective The physical image decoded by the student decoder For the environment and perspective The physical image decoded by the student decoder; the function expression for updating the network parameters of the teacher decoder using the exponential moving average (EMA) is: ; in, For the network parameters of the teacher decoder, For the network parameters of the student decoder, This is the momentum coefficient.
[0013] The present invention also provides a cross-media physical image watermarking system based on differentiable physical simulation, comprising a microprocessor and a memory interconnected thereto, wherein the microprocessor is programmed or configured to execute the cross-media physical image watermarking method based on differentiable physical simulation.
[0014] Compared with existing technologies, the present invention can achieve the following beneficial effects: (1) Significantly improves robustness in real-world scenarios: The present invention explicitly models the physical interactions of geometry, materials, and lighting through a conditional neurophysical simulator, rather than simply superimposing noise. Experiments have shown that this method enables the finely tuned decoder to achieve better watermark extraction accuracy than the untune decoder model under different paper materials (such as coated paper and rough kraft paper), extreme shooting angles (up to 60 degrees of tilt), and complex lighting changes. (2) Achieves end-to-end physical perception optimization: Unlike black-box methods that rely on external rendering software (such as Blender), the present invention proposes a fully differentiable method. This means that the gradient of the physical environment can be directly backpropagated to the watermark network, guiding the watermark network to automatically adapt to the degradation laws of the physical world, thereby "learning" physics itself rather than just defending against noise. (3) Provides an efficient and standardized evaluation and data generation platform: The simulator of the present invention can generate massive amounts of training data with physical properties at low cost and high efficiency, overcoming the problem that real-world printing and shooting data collection is expensive and unreproducible. Meanwhile, it provides a standardized testing platform, which facilitates rigorous evaluation of different materials and lighting conditions in a controlled digital environment. (4) Enhanced generalization ability of the decoder: Through the Environment Consistent Distillation (ECD) strategy, this invention guides the model to learn core content representations that are independent of the environment, reducing dependence on specific scenes. This enables the decoder to not only stably adapt to various environments covered in training, but also maintain good watermark extraction performance under lighting and material conditions that are not explicitly trained, demonstrating strong scene adaptability. Attached Figure Description
[0015] Figure 1 This is a schematic diagram of the basic process of the method in an embodiment of the present invention.
[0016] Figure 2 This is a schematic diagram illustrating the working principle of the method in an embodiment of the present invention.
[0017] Figure 3 This is a schematic diagram illustrating the encoding and decoding working principles of the encoder and decoder in an embodiment of the present invention. Detailed Implementation
[0018] To enable those skilled in the art to better understand the technical solutions of the present invention, the technical solutions of the present invention will be further described in detail below with reference to the accompanying drawings of the embodiments of the present invention. Figure 1 and Figure 2 As shown, the cross-media physical image watermarking method based on differentiable physics simulation in this embodiment includes the following steps: S101: Import or generate the paper material texture map of the target paper in the rendering software, construct a virtual printing scene, and attach the physical image sample containing the watermark data to the surface of the paper material texture map to form a rendered image. Export the rendered image in NeRF standard format and a pairwise supervised dataset composed of different camera parameters. S102, using a differentiable physics simulator to generate batches of simulated printed image samples from paired supervised datasets; S103 uses a batch of simulated printed image samples to fine-tune the pre-trained decoder. S104, input the physical image into the finely tuned decoder to decode the watermark data in the physical image. If the decoded watermark data is inconsistent with the expected watermark data or the decoding fails, the input physical image is determined to be a forged physical image.
[0019] Using physical image samples containing watermark data as host images, and based on the NeRF standard synthetic data format (blender-generated images + camera parameters), through the entire process of pairwise supervised dataset generation, differentiable physics simulator reverse inference of geometric and material parameters, batch data simulation generation, and decoder fine-tuning, the watermark extraction effect of the decoder in real printing and shooting scenarios can be significantly improved. It can solve the following technical problems: (1) It can solve the problem that existing methods usually simplify physical degradation into independent two-dimensional image noise or transformation, ignoring the complex physical coupling relationship between shooting perspective, ambient lighting and paper material, resulting in insufficient robustness of the model in real scenarios. (2) It can solve the problem that the simulation method based on traditional rendering engines (such as Blender) has low computational efficiency and is not differentiable, and cannot achieve end-to-end joint optimization of watermark model and degradation simulator, which limits the model's adaptability to physical environment. (3) It can solve the problem that there is a lack of effective training supervision mechanism that can maintain feature consistency under different lighting, materials and perspectives, which makes it difficult for the decoder to extract watermark features that are not affected by the environment.
[0020] In step S101 of this embodiment, when importing or generating a paper texture map of the target paper in the rendering software, generating the paper texture map of the target paper refers to generating a paper texture map of the target paper based on the set albedo, roughness, metallicity, and ink absorption rate of the paper material. It should be noted that importing or generating a paper texture map of the target paper based on the set albedo, roughness, metallicity, and ink absorption rate of the paper material in the rendering software is a known existing method, so the details of its implementation or setting will not be described in detail here.
[0021] In step S101 of this embodiment, the generation of the physical image sample containing watermark data includes: normalizing the original host image sample (e.g., a commercial logo image) (mapping pixel values to the [0,1] range) and standardizing its size (e.g., adjusting it to 800×800 pixels), inputting the watermark data into a pre-trained encoder, and embedding the watermark data into the original host image sample through the encoder to obtain the physical image sample containing watermark data; the watermark data is binary data of a specified bit length, for example, in this embodiment, the watermark data is 100-bit binary watermark data, which can be represented as: It should be noted that the encoder and decoder for watermark data encoding and decoding can be selected as needed, such as a Stegastamp encoder and Stegastamp decoder. Figure 3 As shown, 100-bit binary watermark data and a physical image are input into the Stegastamp encoder E. The original physical image is embedded through the feature fusion and steganography mechanism of the Stegastamp encoder to obtain a physical image containing the watermark data. This ensures the watermark remains imperceptible while providing a base watermarked image for subsequent decoder fine-tuning. Finally, the physical image obtained from the printing-photographing process can be input into the Stegastamp decoder to decode the 100-bit binary watermark data from the physical image.
[0022] In step S101 of this embodiment, when constructing the virtual printing scene and pasting the physical image sample containing the watermark data onto the paper material texture surface to form a rendered image, this includes configuring the outdoor HDR lighting environment and setting the camera distance, scaling the paper size so that the center of the paper material texture is located at the origin of the world coordinate system, and the paper size is located at the origin of the world coordinate system. Within the space, the camera distance is set, and the paper material texture occupies a preset proportion in the camera's field of view. In this embodiment, the rendering software in step S101 can be various rendering software such as Blender. A physically realistic virtual scene is constructed using Blender, generating a pairwise supervised dataset in NeRF standard format (containing only Blender rendered images and camera parameters), providing a precise target for parameter inference of the differentiable physics simulator, ultimately serving the decoder's fine-tuning. The specific steps are as follows: A physical model of paper is constructed in Blender, matching the characteristics of real paper by using existing material textures or manually setting parameters such as albedo, roughness, metallicity, and ink absorption rate for each material; a virtual printing scene is built, displaying a logo image containing a watermark. Assemble the printed content onto the paper surface, configure a normal outdoor HDR lighting environment and set the camera distance, and scale the paper size so that the center of the paper model is located at the origin of the world coordinate system. The paper size needs to be within the range of the world coordinate system origin. Within a given space, with the camera distance set simultaneously, the paper occupies approximately 70% of the camera's field of view. Using the BlenderNeRF plugin, a NeRF standard format "image-camera parameter" paired supervised dataset is exported. For a logo and a material, only a few hundred images are needed. No additional explicit annotations of material, geometry, or other parameters are required; the subsequent model infers the geometric appearance and material parameters solely through the images and camera parameters.
[0023] In step S102 of this embodiment, the core objective of the differentiable physics simulator is to reverse-engineer the physical parameters of a paper material based on NeRF standard format data constructed from a material and a logo image, and to simulate the observed image after printing and shooting, providing a large number of physically realistic training samples for decoder fine-tuning. As an optional implementation, the differentiable physics simulator is a pre-trained Conditional Neurophysical Simulator (CNPS). The CNPS includes a geometry field module, a material field module, and a relighting module. The geometry field module consists of a radiation field module and a reflection field module. The geometry field module, using only the image and camera parameters, can infer the projection distortion of the paper's 3D geometry (surface deformation, 3D mesh) in relation to the shooting viewpoint, and estimate rough material parameters. The radiation field module is used to predict the RGB color at the corresponding viewpoint based on the pose embedding features of the input rendered image and camera parameters. The Reflection Field module predicts the RGB colors of specular reflections and initial material parameters (including albedo, roughness, and metallicity) from the input rendered image and camera parameters' pose embedding features. The Material Field module infers and generates more refined material parameters based on the input rendered image, camera parameters' pose embedding features, the RGB image output from the Geometry Field module, and the initial material parameters. The Material Field module, using only the image, camera parameters, the 3D mesh of the paper's geometric surface output from the Geometry Radiation Field module, and the input lighting descriptor, infers more refined material parameters such as albedo, roughness, and metallicity, characterizing the interaction between material and lighting. The Relighting module (a core component built into the Conditional Neurophysics Simulator CNPS) synthesizes simulated printed image samples from the corresponding viewpoint using differentiable rendering based on more refined material parameters. It receives geometric parameters, material parameters, and lighting descriptors, uses images rendered with the same parameters in Blender for supervised training, and rapidly synthesizes observation images under different environments through differentiable rendering, replacing traditional Blender rendering and ensuring rapid end-to-end generation by the Conditional Neurophysics Simulator CNPS.
[0024] In this embodiment, the calculation function expression for the pose embedding features of the rendered image and camera parameters is: ; ; in, Sampling points in a 3D geometric field for the input rendered image Potential geometric features , and These are the axis vector factors of the X, Y, and Z axes of the 3D scene tensor, respectively. , and They are 3D scene tensors The axis matrix factors of the YZ axis, XZ axis, and XY axis. For splicing operations, Sampling points for camera parameters in a three-dimensional geometric field The pose embedding features (representation enhanced by Fourier feature encoding). The Fourier feature encoding dimension is set, for example, B=6 in this embodiment, generating a 78-dimensional pose embedding feature. The radiation field module predicts the RGB color at the corresponding viewpoint based on the pose embedding features of the input rendered image and camera parameters, which can be represented as: ; in, Sampling points in a three-dimensional geometric field The signed distance field SDF at that location, Sampling points in a three-dimensional geometric field RGB color values at that location The material field module employs a multilayer perceptron (MLP). The reflection field module predicts the specular reflection RGB color and initial material parameters from the corresponding viewpoint based on the pose embedding features of the input rendered image and camera parameters. Its principle is similar to the radiation field module, both utilizing a multilayer perceptron (MLP). However, the output of the MLP includes not only the signed distance field (SDF) but also the specular reflection RGB color and initial material parameters (albedo a, roughness r, and metallicity m). Finally, the 3D mesh of the geometric surface is obtained through the SDF field. In this embodiment, the loss function used during training of the geometric field module includes constraints such as color reconstruction and physical priors to ensure the physical authenticity of the inference parameters. Specifically, the expression of the loss function is as follows: ; ; ; ; ; ; in, This is the loss function used by the geometry field module during training. For roughness-perceived reconstruction loss, For Eikonal's loss, Gaussian smoothing loss, For Hessian losses, For TV regularization loss, To cover up the damage, To stabilize the loss, , , and All are weighted parameters. For roughness, , and These are the output colors of the radiation field, the output colors of the reflection field, and the colors of the actual image, respectively. Represents the L1 norm; This indicates taking the mean of the set of points in the sampling space; Sampling points in a three-dimensional geometric field Signed distance field at location The gradient vector; Sampling points in a three-dimensional geometric field Signed distance field at location The gradient vector; It is the displacement vector; It is an L2 norm; Represents sampling points in a three-dimensional geometric field The Hessian matrix at that location; Denotes the Frobenius norm. The input rendered image, For a 3D scene tensor, The total variation norm is represented by . The Eikonal loss constrains the geometric properties of the signed range field (SDF). The Gaussian smoothing loss constrains the local consistency of the surface normal vectors. The TV regularization loss constrains the sparsity of the feature tensor. The occlusion loss regularizes the signed range field SDF in regions where light rays do not hit the surface, reducing artifacts in the void. The stabilization loss constrains the signed range field SDF to approximate a sphere or other prior shape early in training, preventing geometric collapse. The displacement vector is also mentioned. The aim is to minimize the difference in gradients between adjacent points; the Frobenius norm is used to constrain the curvature of the geometric surface to ensure smoothness. In this embodiment, the material field module is used to infer and generate more refined material parameters based on the pose embedding features of the input rendered image and camera parameters, as well as the 3D mesh and lighting descriptor of the geometric surface output by the geometry field module. Based on the 3D mesh model output by the geometry field module, high-precision surface points are obtained through ray intersection and signed distance field SDF sampling. The sampling points are... Potential geometric features Pose embedding features It is concatenated with the lighting descriptor (the lighting parameters used by the rendering software Blender to generate rendered images, including ambient light intensity and proxy values for the main light direction) to form a conditional vector. By inputting a material field module composed of multiple multilayer perceptrons (MLPs), RGB images and more refined material parameters (albedo, roughness, metallicity) can be generated through inference. Explicit material annotation is not required throughout the process. In this embodiment, the material field module includes multiple MLPs, each corresponding to a material parameter. This parameter is used to construct a conditional vector based on the pose embedding features of the rendered image and camera parameters, concatenated with lighting descriptors. To infer and generate more refined material parameters (including albedo, roughness, and metallicity), the lighting descriptor refers to the lighting parameters used by the rendering software to generate the rendered image, including ambient light intensity and principal ray direction proxy value. This multilayer perceptron (MLP) can be represented as:
[0025] in, Albedo, For roughness, For metallicity, The material field module is composed of multiple multilayer perceptron (MLP) modules. To sample points Potential geometric features Pose embedding features A conditional vector is formed by concatenating the lighting descriptor (the ray parameters used in Blender's data generation, including ambient light intensity and principal ray direction proxy values). Simultaneously, an RGB image from this viewpoint is synthesized using the predicted material parameters. The material field module employs color reconstruction loss during training. (Matching the rendered color with the RGB values of the Blender image based on the inferred material parameters) and material regularization loss (Ensuring spatial smoothness and cross-view consistency of material parameters). The loss function used by the material field module during training has the following expression: ; ; in, This is the loss function used by the material field module during training. These are the weighting coefficients. For material regularization loss, , and Sampling points in the three-dimensional geometric field The material parameters include albedo, roughness, and metallicity. By calculating the L1 norm of the spatial gradients of albedo, roughness, and metallicity, the aim is to constrain the smoothness of the spatial distribution of material parameters and avoid high-frequency noise interference; in this embodiment, wherein... =0.2, the material field module relies solely on image data in the NeRF format for supervision during training.
[0026] The relighting module is used to synthesize simulated printed image samples from a corresponding viewpoint using differentiable rendering based on RGB images and more refined material parameters. This module acts as a key hub connecting decoupled physical parameters and the final observed image, aiming to fit realistic lighting synthesis rules through deep neural networks. It receives more refined material parameters obtained from the geometry and material field modules and, based on explicit ambient lighting conditions, can quickly and faithfully synthesize observed images under specific lighting environments. In this embodiment, the relighting module includes an image encoder, a conditional encoder, and a multi-scale U-Net. The decoder includes an image encoder comprising a convolution module and a downsampling module connected in sequence. The input features of the image encoder are an H×W×8 RGB image output by the material field module and a more refined material parameter map, where H and W are the height and width of the RGB image, respectively. The conditional encoder comprises a multilayer perceptron (MLP) and a multiscale shaping module connected in sequence. The input features of the MLP are a multidimensional conditional vector formed by concatenating the pose embedding features of the sampling points and the features of the HDR lighting environment obtained through second-order spherical harmonic encoding. In this embodiment, the HDR lighting environment is encoded into 27 dimensions through second-order spherical harmonic (SH) encoding, and after concatenation, a 105-dimensional multidimensional conditional vector is formed. The multiscale shaping module is used to deform (reshape) the output features of the MLP to generate features of multiple scales as input to the multiscale U-Net decoder. The multiscale U-Net decoder (including skip connections, progressively upsampling to the original resolution) is used to generate simulated printed image samples (3-channel RGB observation images) from the corresponding viewpoint based on the input features of multiple scales. The supervised training of the relighting module uses the output of the material field in Blender, along with simulated images generated under different HDR lighting conditions as labels, to optimize the loss function. To ensure the physical realism of the synthesized image; after training, this module can replace the Blender generation process, achieving lightweight generation. In this embodiment, the loss function used by the relighting module during training has the following expression: ; ; in, This is the loss function used by the relighting module during training. The pixel-level L1 reconstruction loss is the predicted image generated by the relighting module. With the target real image The average of the absolute errors between pixels. In order to perceive loss, This represents a pre-trained feature extraction network. For the sample size, This represents the Euclidean distance in the feature space (used to improve the visual realism of the image). Finally, using the trained Conditional Neurophysical Simulator (CNPS), simulated printing-photographing images under different viewpoints and lighting conditions are generated in batches. This provides a large number of diverse samples for decoder fine-tuning. It iterates through preset physical environment parameters (material parameters output from the material field, synthesized RGB images, different HDR files, and camera parameters from different viewpoints) to form a complete environmental parameter combination space. The relighting module then synthesizes simulated printing-photographing image samples under the corresponding viewpoints. .
[0027] As an optional implementation, in step S103 of this embodiment, when fine-tuning the pre-trained StegaStamp decoder using a batch of simulated printed image samples, it includes using simulated images generated in batches by the Conditional Neurophysical Simulator CNPS, and employing a student decoder and teacher decoder to fine-tune the decoder using an Environment Consistent Distillation (ECD) strategy to make it learn environment-invariant features, significantly improving the watermark extraction effect in real-world scenarios. Specifically, using a student decoder and teacher decoder to fine-tune the decoder using an environment consistent distillation strategy includes: S201, initialize the pre-trained decoder as the student decoder and the teacher decoder; S202 uses batches of simulated printed image samples to train both the student and teacher decoders, and in each iteration of training, it employs a triple loss mechanism. The network parameters of the student decoder are updated to achieve environment-consistent distillation and learning environment-independent watermark features. The network parameters of the teacher decoder are updated using exponential moving average (EMA). This triple loss... The expression for the computation function is: ; ; ; ; in, , and For weight parameters, This represents a loss of visual consistency within the environment. For the loss of invariance between environments, Entropy loss weighted by confidence level For a set of perspectives, From the perspective The weight, Let KL divergence be the KL divergence. For the predicted bit probability distribution of the student decoder, Predict the probability distribution from the anchor point perspective of the teacher decoder. From the anchor point perspective, This serves as the reference distribution separator in KL divergence calculations. For the environment and perspective The confidence weights are as follows: The output entropy of the student decoder, For student decoders in the environment and perspective The probability distribution of the next predicted bit. For student decoders in the environment and perspective The probability distribution of the next predicted bit. For the environment and perspective The physical image decoded by the student decoder For the environment and perspective The physical image decoded by the student decoder For student decoders, For teacher decoders; loss of perspective consistency within the environment (weight parameters) =1.0) indicates a fixed environment Choose the anchor point view with the highest visibility. By constraining the decoding outputs of other perspectives through KL divergence, and comparing them with the teacher decoder... Consistent output from different viewpoints ensures stable decoding across different viewpoints within the same environment. (Loss of invariance across environments) (weight parameters) =0.8) Randomly sample two different environments , perspective , By minimizing the difference in the decoding bit distribution through KL divergence, consistent decoding results for the same watermark are ensured across different environments. Confidence-weighted entropy loss is used. ( =0.2) By combining the perspective confidence weight (By fusing prediction confidence, visibility, and specular mask), the output entropy of the student decoder is regularized to prevent the model from becoming overconfident in uncertain regions and to improve decoding stability.
[0028] In this embodiment, the function expression for updating the network parameters of the teacher decoder using the exponential moving average (EMA) is: ; in, For the network parameters of the teacher decoder, For the network parameters of the student decoder, This is the momentum coefficient. In this embodiment, the student decoder... The pre-trained StegaStamp decoder was used as the initialization target model for fine-tuning, and was ultimately used for watermark extraction in real-world scenarios; the teacher decoder... The structure is completely identical to that of the student decoder, and the initial parameters are copied from the student decoder. During training, an exponential moving average (EMA) is used for updates (decay coefficient = 0.999) to provide a stable soft target and prevent overfitting of the student decoder. The encoder parameters are fixed, and only the student decoder is updated. The weights are adjusted to ensure the embedded watermark information is not corrupted, and then simulated printed image samples generated by the conditional neurophysical simulator CNPS are input in batches. Calculate triple loss The weights of the student decoder are optimized through backpropagation. Finally, the original binary watermark data is extracted by inputting the real printed image into the trained student decoder.
[0029] To verify the cross-media physical image watermarking method based on differentiable physical simulation in this embodiment, this embodiment is evaluated on a virtual dataset in the NeRF standard format and a real print-shoot dataset to ensure the physical consistency and real scene adaptability of the method. Specifically, it includes: (1) Blender virtual dataset: using 5 common commercial logos, after superimposing watermarks, it is constructed based on 3 typical paper materials (uncoated A4 copy paper, coated cardboard, kraft paper). 110 pairs of "image-camera parameter" pairs of training samples are generated for each material to decompose each material in the geometric field and material field in CNPS, and 250 camera parameters and 12 different HDR files are used to synthesize images under new perspectives and new lighting based on the trained geometric field and material field. At the same time, based on the output of the geometric field and material field, Blender generates supervision data for the re-lighting network according to the same and decomposed material parameters. (2) Real Print-Shoot Dataset: Used to verify the performance improvement of the fine-tuned decoder. Real print-shoot data was used for testing. The same 5 pre-encoded commercial logos were used, covering 3 lighting types (sunlight, warm light, warm white LED), 4 lighting intensities (3500 lux, 2200 lux, 580 lux, 210 lux), 4 shooting angles (range 0°-60°, step size 10°), and 5 shooting distances (20cm-60cm, step size 10cm). iPhone XR was used. The evaluation indexes set in the experiment included: (1) Robustness index: Bit accuracy (ACC) was used to measure the watermark extraction accuracy. The degree of matching between the decoded bit string and the original secret information was calculated. The higher the ACC, the stronger the robustness. The calculation function expression of bit accuracy (ACC) is: ; in, The length of the watermark data (100 in this example). For indicator functions, and These are the actual value and predicted value of the i-th position of the watermark data, respectively. (2) Simulator consistency index: The simulation-to-real cross-entropy (S2R-CE) similarity is used to measure the physical consistency between the CNPS generated data and the real data. The higher the S2R-CE, the closer the simulation effect is to the real scene. The calculation function expression of simulation-to-real cross-entropy (S2R-CE) is: ; in, and These are the probability distributions of the real data and the probability distributions of the simulated generated data for the i-th position of the watermark data, respectively. All experiments were conducted on a computing platform equipped with an NVIDIA RTX A6000 (48GB video memory) and implemented based on the PyTorch framework. The specific settings are as follows: (1) Data partitioning: In the Blender virtual dataset, of the 110 pairs of paired training samples of the geometric field and the material field, 100 pairs are the training set, 10 pairs are the validation set, and 250 camera parameters can be used as the test set. The supervised data of the re-light network was divided into training and validation sets in an 8:2 ratio; (2) Training hyperparameters: the geometry field and material field training used the Adam optimizer with an initial learning rate of 1e-4, a weight decay of 1e-5, 180,000 training steps, and a batch size of 512; the re-light network used the Adam optimizer with an initial learning rate of 1e-4, a batch size of 8, and 150 training epochs; the decoder fine-tuning used the Adam optimizer with an initial learning rate of 1e-4, a weight decay of 1e-5, 100,000 training steps, and a batch size of 64. In this embodiment, a comparative experiment was conducted on three different decoders: StegaStamp decoder, Lim decoder, and Fang decoder, as well as the StegaStamp decoder fine-tuned using the method of this embodiment. The final results are shown in Tables 1 to 5.
[0030] Table 1: Simulator-Real Scene Consistency Index under Different Decoders for Data with Different Illumination and Materials
[0031] Table 2: Simulator-Real Scene Consistency Index under Different Illumination and Angle Data in Different Decoders
[0032] As shown in Tables 1 and 2, the high simulation-to-reality cross-entropy (S2R-CE) under three different decoders (StegaStamp, Lim, and Fang) also verifies the high consistency between the simulation data of the Conditional Neurophysical Simulator CNPS and the real physical scene.
[0033] Table 3: Performance comparison of the fine-tuned decoder under different materials and distances.
[0034] Table 4: Performance comparison of the fine-tuned decoder under different materials and angles.
[0035] Table 5: Performance comparison of the fine-tuned decoder under different materials and lighting conditions.
[0036] As shown in Tables 3, 4, and 5, compared to the benchmark StegaStamp encoder, the StegaStamp decoder finely tuned in this embodiment (“fine-tuned”) exhibits superior watermark extraction accuracy (bit accuracy ACC) under different shooting distances, angles, and lighting intensities. Especially in strong lighting (3500 lux) and complex degradation scenarios caused by materials, the finely tuned StegaStamp decoder effectively overcomes the lack of robustness in traditional methods.
[0037] In addition, in this embodiment, the ablation experiment only used the conditional neurophysical simulator CNPS to generate data for fine-tuning but without the environmental consistent distillation (ECD) strategy decoder (w / o ECD) to verify the effectiveness of the environmental consistent distillation (ECD) strategy. The results are shown in Tables 6, 7 and 8.
[0038] Table 6: Impact of ECD Strategy on the Performance of the Fine-tuned Decoder under Different Material and Distance Conditions
[0039] Table 7: Impact of ECD Strategy on the Performance of the Fine-tuned Decoder under Different Material and Angle Conditions
[0040] Table 8: Impact of ECD Strategy on the Performance of the Fine-tuned Decoder under Different Material and Light Intensity Conditions
[0041] The ablation experiments in Tables 6, 7, and 8 further validate the core role of the Environment Consistent Distillation (ECD) module. The model incorporating the ECD strategy outperformed the model without this module under all test conditions, demonstrating that this strategy effectively guides the decoder to remove environmental interference and learn robust features invariant to the environment. This ensures that the model can not only adapt to the simulated environment in the training data but also stably generalize to the varied physical conditions in the real world, thus guaranteeing high reliability for cross-media transmission. Therefore, the experimental results show that the cross-media physical image watermarking method based on differentiable physics simulation in this embodiment has significant performance advantages in real-world printing-photography scenarios.
[0042] In summary, the method in this embodiment selects a commercial logo image as the host carrier (physical image) for watermark embedding, uses the Neural Radiation Field (NeRF) standard synthetic data system as technical support, integrates Blender rendered images and camera parameter sets to construct a basic data framework, and significantly improves the stability of watermark extraction in real printing-shooting cross-media scenarios through the complete technical chain of Blender supervised sample construction, reverse solving of geometric and material parameters by the Conditional Neurophysical Simulator (CNPS) and batch generation of physically realistic simulation data, and iterative optimization of the decoder by the Environment Consistent Distillation (ECD) optimization strategy.
[0043] Those skilled in the art will understand that the technical solutions provided by this invention may take the form of a method, system, or computer program product. For example, this invention may provide a cross-media physical image watermarking system based on differentiable physical simulation, including an interconnected microprocessor and a memory, wherein the microprocessor is programmed or configured to execute the cross-media physical image watermarking method based on differentiable physical simulation. Furthermore, this invention may take the form of a completely hardware embodiment, a completely software embodiment, or an embodiment combining software and hardware aspects. Moreover, this invention may take the form of a computer program product implemented on one or more computer-readable storage media (including but not limited to disk storage, CD-ROM, optical storage, etc.) containing computer-usable program code. This invention is described with reference to flowchart illustrations and / or block diagrams of methods, apparatus (systems), and computer program products according to embodiments of the invention. It should be understood that each block of the flowchart illustrations and / or block diagrams, and combinations of blocks in the flowchart illustrations and / or block diagrams, can be implemented by computer program instructions. These computer program instructions can be provided to a processor of a general-purpose computer, special-purpose computer, embedded processor, or other programmable data processing device to produce a machine, such that the instructions, which are executable by the processor of the computer or other programmable data processing device, produce instructions for implementing the process. Figure 1 One or more processes and / or boxes Figure 1 The computer program instructions may also be stored in a computer-readable storage medium that can direct a computer or other programmable data processing device to operate in a particular manner, such that the instructions stored in the computer-readable storage medium produce an article of manufacture including instruction means, which are implemented in a process Figure 1 One or more processes and / or boxes Figure 1 The functions specified in one or more boxes. These computer program instructions may also be loaded onto a computer or other programmable data processing apparatus to cause a series of operational steps to be performed on the computer or other programmable apparatus to produce a computer-implemented process, thereby providing instructions that execute on the computer or other programmable apparatus for implementing the process. Figure 1One or more processes and / or boxes Figure 1 The steps of the function specified in one or more boxes.
[0044] The above description is merely a preferred embodiment of the present invention. The scope of protection of the present invention is not limited to the above embodiments. All technical solutions falling within the scope of the present invention's concept are within the scope of protection of the present invention. It should be noted that for those skilled in the art, any improvements and modifications made without departing from the principles of the present invention should also be considered within the scope of protection of the present invention.
Claims
1. A cross-media physical image watermarking method based on differentiable physics simulation, characterized in that, Includes the following steps: S101: Import or generate the paper material texture map of the target paper in the rendering software, construct a virtual printing scene, and attach the physical image sample containing the watermark data to the surface of the paper material texture map to form a rendered image. Export the rendered image in NeRF standard format and a pairwise supervised dataset composed of different camera parameters. S102 uses a differentiable physics simulator to generate batches of simulated printed image samples from paired supervised datasets. S103 uses a batch of simulated printed image samples to fine-tune the pre-trained decoder; S104, input the physical image into the finely tuned decoder to decode the watermark data in the physical image. If the decoded watermark data is inconsistent with the expected watermark data or the decoding fails, the input physical image is determined to be a forged physical image.
2. The cross-media physical image watermarking method based on differentiable physics simulation according to claim 1, characterized in that, In step S101, when importing or generating the paper texture map of the target paper in the rendering software, generating the paper texture map of the target paper refers to generating the paper texture map of the target paper based on the set paper material's albedo, roughness, metallicity, and ink absorption rate.
3. The cross-media physical image watermarking method based on differentiable physics simulation according to claim 1, characterized in that, In step S101, the generation of the physical image sample containing watermark data includes: normalizing and size standardizing the original host image sample, inputting the watermark data into a pre-trained encoder, and embedding the watermark data into the original host image sample through the encoder to obtain the physical image sample containing watermark data; the watermark data is binary data of a specified bit length.
4. The cross-media physical image watermarking method based on differentiable physics simulation according to claim 1, characterized in that, In step S101, when constructing the virtual printing scene and pasting the input physical image sample containing watermark data onto the paper material texture surface to form a rendered image, this includes configuring the outdoor HDR lighting environment and setting the camera distance, scaling the paper size so that the center of the paper material texture is located at the origin of the world coordinate system, and the paper size is located at the origin of the world coordinate system. Within the space, set the camera distance to determine the preset proportion of the paper texture image within the camera's field of view.
5. The cross-media physical image watermarking method based on differentiable physics simulation according to claim 1, characterized in that, In step S102, the differentiable physics simulator is a pre-trained Conditional Neurophysical Simulator (CNPS). The CNPS includes a geometry field module, a material field module, and a relighting module. The geometry field module consists of a radiation field module and a reflection field module. The radiation field module is used to predict the RGB color at the corresponding viewpoint based on the pose embedding features of the input rendered image and camera parameters. The reflection field module is used to predict the specular reflection RGB color and initial material parameters at the corresponding viewpoint based on the pose embedding features of the input rendered image and camera parameters. The material parameters include some or all of albedo, roughness, and metallicity, and infers the 3D mesh of the geometric surface of the rendered image. The material field module, based on the 3D mesh model output by the geometry field module, first obtains the pose embedding features of the input rendered image and camera parameters, as well as the input lighting descriptor, to infer and generate RGB images and more refined material parameters. The relighting module is used to synthesize simulated printed image samples at the corresponding viewpoint through differentiable rendering based on the RGB images and more refined material parameters.
6. The cross-media physical image watermarking method based on differentiable physics simulation according to claim 5, characterized in that, Both the radiation field module and the reflection field module are multilayer perceptron (MLP) modules; the material field module consists of multiple MLP modules, each corresponding to a material parameter, used to construct a conditional vector based on the pose embedding features of the rendered image and camera parameters, and the concatenation of lighting descriptors. The system generates more refined material parameters through inference. The lighting descriptor is the lighting parameter used by the rendering software to generate the rendered image, including ambient light intensity and proxy value of the main light direction. The relighting module includes an image encoder, a conditional encoder, and a multi-scale U-Net decoder. The image encoder includes a convolution module and a downsampling module connected in sequence. The input features of the image encoder are the H×W×8 RGB image output by the material field module and a more refined material parameter map, where H and W are the height and width of the RGB image, respectively. The conditional encoder includes a multilayer perceptron (MLP) and a multi-scale shaping module connected in sequence. The input features of the MLP are a multi-dimensional conditional vector formed by concatenating the pose embedding features of the sampling points and the features obtained by second-order spherical harmonic encoding of the HDR lighting environment. The multi-scale shaping module is used to deform the output features of the MLP to generate features of multiple scales as input to the multi-scale U-Net decoder. The multi-scale U-Net decoder is used to generate simulated printed image samples from corresponding viewpoints based on the input features of multiple scales.
7. The cross-media physical image watermarking method based on differentiable physics simulation according to claim 5, characterized in that, The calculation function expression for the pose embedding features of the rendered image and camera parameters is as follows: ; ; in, Sampling points in a 3D geometric field for the input rendered image Potential geometric features , and These are the axis vector factors of the X, Y, and Z axes of the 3D scene tensor, respectively. , and They are 3D scene tensors The axis matrix factors of the YZ axis, XZ axis, and XY axis. For splicing operations, Sampling points for camera parameters in a three-dimensional geometric field pose embedding features The dimension for encoding Fourier features.
8. The cross-media physical image watermarking method based on differentiable physics simulation according to claim 5, characterized in that, The loss function used by the geometry field module during training has the following expression: ; ; ; ; ; ; in, This is the loss function used by the geometry field module during training. For roughness-perceived reconstruction loss, For Eikonal's loss, Gaussian smoothing loss, For Hessian losses, For TV regularization loss, To cover up the damage, To stabilize the loss, , , and All are weighted parameters. For roughness, , and These are the output colors of the radiation field, the output colors of the reflection field, and the colors of the actual image, respectively. Represents the L1 norm; This indicates taking the mean of the set of points in the sampling space; Sampling points in a three-dimensional geometric field Signed distance field at location The gradient vector; Sampling points in a three-dimensional geometric field Signed distance field at location The gradient vector; It is the displacement vector; It is an L2 norm; Represents sampling points in a three-dimensional geometric field The Hessian matrix at that location; Denotes the Frobenius norm. The input rendered image, For a 3D scene tensor, The total variation norm is represented; the loss function used by the material field module during training has the following expression: ; ; in, This is the loss function used by the material field module during training. These are the weighting coefficients. For material regularization loss, , and Sampling points in the three-dimensional geometric field The material parameters include albedo, roughness, and metallicity; the loss function used by the relighting module during training has the following expression: ; ; in, This is the loss function used by the relighting module during training. The pixel-level L1 reconstruction loss is the predicted image generated by the relighting module. With the target real image The average of the absolute errors between pixels. In order to perceive loss, This represents a pre-trained feature extraction network. For the sample size, This represents the Euclidean distance in the feature space.
9. The cross-media physical image watermarking method based on differentiable physics simulation according to claim 1, characterized in that, In step S103, when fine-tuning the pre-trained decoder using a batch of simulated printed image samples, the fine-tuning includes employing an environment-consistent distillation strategy using both the student decoder and the teacher decoder: S201, initialize the pre-trained decoder as the student decoder and the teacher decoder; S202 uses batches of simulated printed image samples to train both the student and teacher decoders, and in each iteration of training, it employs a triple loss mechanism. The network parameters of the student decoder are updated to achieve environment-consistent distillation and learning environment-independent watermark features. The network parameters of the teacher decoder are updated using exponential moving average (EMA). This triple loss... The expression for the computation function is: ; ; ; ; in, , and For weight parameters, This represents a loss of visual consistency within the environment. For the loss of invariance between environments, Entropy loss weighted by confidence level For a set of perspectives, From the perspective The weight, Let KL divergence be the KL divergence. For the predicted bit probability distribution of the student decoder, Predict the probability distribution from the anchor point perspective of the teacher decoder. From the anchor point perspective, This serves as the reference distribution separator in KL divergence calculations. For the environment and perspective The confidence weights are as follows: The output entropy of the student decoder, For student decoders in the environment and perspective The probability distribution of the next predicted bit. For student decoders in the environment and perspective The probability distribution of the next predicted bit. For the environment and perspective The physical image decoded by the student decoder For the environment and perspective The physical image decoded by the student decoder; the function expression for updating the network parameters of the teacher decoder using the exponential moving average (EMA) is: ; in, For the network parameters of the teacher decoder, For the network parameters of the student decoder, This is the momentum coefficient.
10. A cross-media physical image watermarking system based on differentiable physics simulation, comprising an interconnected microprocessor and a memory, characterized in that, The microprocessor is programmed or configured to execute the cross-media physical image watermarking method based on differentiable physical simulation as described in any one of claims 1 to 9.
Citation Information
Patent Citations
Text image enhancement model, training method, enhancement method and electronic equipment
CN113177556A
Method and device for extracting watermark
CN120430923A
Fluorescent watermarking for partially reflecting / absorbing ink in ultraviolet spectrum
CN120547277A
Calligraphy and painting cultural relic technology photographing system
CN121069686A
Self-supervised image relighting method based on physical consistency
CN121458861A