Image protection method for balancing robustness and generation diversity

By employing two-stage encryption and tail-truncation sampling techniques, this method solves the problem of balancing watermark robustness and generation diversity in existing technologies. It achieves both robustness and generation diversity of watermarks without affecting image quality, providing a robust and covert solution for generated image detection and tracing.

CN121353048APending Publication Date: 2026-01-16UNIV OF SCI & TECH OF CHINA
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202511519291.X
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-10-23
Publication Date
2026-01-16

AI Technical Summary

Technical Problem

Existing NaW methods struggle to balance robustness of watermarking with diversity of generation, resulting in insufficient robustness or limited diversity of generated images, which negatively impacts user experience.

Method used

A two-stage encryption scheme using a preset master key and a random session key is adopted. Combined with tail-truncation sampling technology, the watermark noise vector is embedded into the image. By designing the tail-truncation sampling dimension and the watermark coding subspace, the diversity and robustness of image generation are maintained.

Benefits of technology

It achieves both robustness and diversity of watermark generation without affecting image quality, and provides a robust and covert solution for generated image detection and source tracing.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121353048A_ABST
    Figure CN121353048A_ABST
Patent Text Reader

Abstract

The invention discloses an image protection method for balancing robustness and generation diversity. The method comprises the following steps: encoding a randomly generated random session key by adopting a preset master key to obtain a random session key code; encoding a watermark to be embedded for image protection by using the random session key to obtain a watermark code; and generating a watermark noise vector based on the random session code and the watermark code, and embedding the watermark noise vector into a to-be-protected image. According to the method, a robust and hidden generation image detection and traceability scheme is realized, so that the watermark embedded in the image can give consideration to the robustness of the watermark and the diversity of image generation.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of image source tracing and processing technology, and in particular to an image protection method that balances robustness and generative diversity. Background Technology

[0002] In recent years, diffusion models have made groundbreaking progress in the field of generative artificial intelligence, especially in image synthesis. Thanks to their iterative denoising process, these models are able to generate highly realistic images. Furthermore, with the expansion of model architectures and training datasets, diffusion models have continuously improved in terms of generation quality and controllability, driving the rapid development of AI generation technology.

[0003] However, this rapid evolution has also brought new challenges: the authenticity of generated content is becoming increasingly difficult to verify, making diffusion models a powerful tool for spreading misinformation and manipulating public opinion. At the same time, service providers have invested massive amounts of data and computing power in training these models, creating an urgent need for effective mechanisms to protect their intellectual property from unauthorized use. Therefore, multiple factors have collectively driven the need to protect the copyright of AI-generated images and to trace their origins.

[0004] In existing technologies, watermarking can be used to protect and trace images generated by AI. Watermarking technology does indeed provide a fundamental solution for image protection and traceability and has shown considerable promise. Among various technical approaches, the Noise-as-Watermark (NaW) method stands out. Its core idea is to map the watermark codeword onto a standard Gaussian noise vector, using this vector as the initial noise for image generation, thus embedding watermark information while maintaining the noise distribution. The extraction stage only requires inversely reversing the diffusion process to recover the embedded noise and decode it to obtain the watermark. Furthermore, since the forward and reverse diffusion processes provide a reversible bijection between the high-entropy Gaussian vector and the low-entropy image manifold, the NaW method can achieve high-capacity, near-lossless embedding on a single image without additional training.

[0005] Currently, most NaW (Navigational Undefined) methods involve two core steps in their embedding process: discrete encoding and continuous sampling. The former converts the watermark message into codewords, while the latter maps these discrete binary codes to continuous noise. The challenge of discrete encoding lies in striking a balance between "randomness" (making the watermark unpredictable) and "robustness" (resisting various distortions)—these two aspects often conflict with each other. On the other hand, the key to continuous sampling is accurately preserving the initial Gaussian noise distribution to maintain imperceptibility.

[0006] However, existing NaW methods each have their own trade-offs: some use simple repeating codes to achieve high robustness, but lack randomness and mainly rely on the sampling process to introduce changes; others use pseudo-random error-correcting codes to improve undetectability, but have weak robustness and may fail under common inversion errors.

[0007] In summary, existing watermarking embedding methods severely limit the diversity of generated images in order to ensure the robustness of the watermark, leading to similar image layouts generated by the corresponding models and affecting user experience. Other watermarking embedding methods sacrifice robustness in order to achieve the cryptographically provable but undetectable effect of the watermark, making it difficult to resist complex adversarial distortions in actual deployment.

[0008] Therefore, the implementation schemes provided by the existing technologies mentioned above cannot simultaneously achieve robustness and generative diversity in embedding watermarks into images generated by diffusion models. Summary of the Invention

[0009] The purpose of this invention is to provide an image protection method that balances robustness and generative diversity, so that the embedded watermark can take into account both the robustness of the watermark and the diversity of image generation, thus solving the problems existing in the prior art.

[0010] The objective of this invention is achieved through the following technical solution:

[0011] An image preservation method balancing robustness and generative diversity includes:

[0012] A random session key is encoded using a preset master key to obtain a random session key encoding; and a watermark to be embedded for image protection is encoded using the random session key to obtain a watermark encoding.

[0013] A watermark noise vector is generated based on the random session coding and the watermark coding, and the watermark noise vector is embedded into the image to be protected.

[0014] The process of generating the watermark noise vector includes:

[0015] The initial noise vector is sampled and encoded using the master key to obtain a random session key vector of the corresponding dimension; the watermark to be embedded is sampled and encoded using the random session key to obtain a watermark vector for carrying watermark bits;

[0016] The watermark noise vector is constructed based on the random session key vector and the watermark vector.

[0017] The process of embedding the watermark noise vector into the image to be protected includes:

[0018] The initial noise vector is sampled using a tail-truncation sampling (TTS) method, specifically by sampling along the normal of each hyperplane and ensuring that the sampling point is at least a predetermined truncation threshold distance from the boundary. The predetermined truncation threshold is used to determine the watermark coding subspace for embedding the watermark noise vector, and the watermark noise vector is embedded into the watermark coding subspace.

[0019] The process of determining the watermark coding subspace for embedding the watermark noise vector includes:

[0020] The expected number of samples by truncating the sampling dimension at the tail. Determine the total dimension of the watermark encoding subspace, and the subspace dimension per bit. The watermark encoding subspace is determined by allocating subspace dimensions by bit; wherein, the desired number of tail-truncated sampling dimensions is determined. and the subspace dimension per bit The calculation formulas are as follows:

[0021] ;

[0022] in, The initial noise dimension, The predetermined cutoff threshold for TTS, This represents the length of the watermark.

[0023] The process of constructing the watermark noise vector includes:

[0024] The watermark noise vector is constructed by combining the symbol corresponding to the watermark bit contained in the watermark encoding and the magnitude of the initial noise vector in the tail-truncated sampling dimension, and retaining the original random sample in the center dimension representing the pure noise region in the initial noise vector.

[0025] In the process of constructing the watermark noise vector, the watermark noise vector is obtained. The formulas include:

[0026] ;

[0027] in, This represents the Hadamard product. It is by Bit vectors within the defined subspace Directional coding, ; Let m pseudo-random vectors be generated using the corresponding key of each layer as the seed of the pseudo-random number generator, and each pseudo-random vector... ; It is a binary mask used to... The space is divided into a watermark coding subspace and a random noise subspace.

[0028] The support sets of the pseudo-random vectors are pairwise disjoint and mutually orthogonal, and are used as the normals of the bit-coded hyperplane in the watermark embedding subspace, i.e.:

[0029] .

[0030] The binary mask The watermark encoding subspace is The random noise subspace is ;

[0031] And for the noise vector The application of tail-truncation sampling to each component includes:

[0032] ;

[0033] in, Indicates the interval superior The normalized truncated distribution.

[0034] Specifically, the method also includes a watermark decoding process, which includes:

[0035] Reconstructing the watermarked noise vector through diffusion inversion Then, the noise will be reconstructed. Divided into random session key encoding vectors With watermark encoding vector Then, the master key is used to encode the vector from the random session key. Recover the random session key Then use a random session key From the watermark encoding vector Decoding to obtain the watermark bit vector ;

[0036] The decoding process includes: regenerating the vector set using the corresponding key for each layer. ; Calculate the projection vector for With each The inner product, for The projection vector in the 3D real vector space The included first Projection components The calculation formula is:

[0037] ;

[0038] From the projection components The symbol directly recovers the watermark bit vector for:

[0039] ;

[0040] in, Acting on elements dimensional vector .

[0041] Furthermore, the method also includes a process for detecting and tracing the embedded watermark, wherein:

[0042] The detection process includes:

[0043] The L1 norm of the first-stage projection vector is used as the test statistic, i.e.:

[0044] ;

[0045] in, The number of bits in the random session key. It is Projected onto the first stage normal set The resulting projection vector;

[0046] Statistic With threshold If l exceeds d, the image is determined to have a watermark embedded; if l does not exceed d, the image is determined not to have a watermark embedded; the threshold d is pre-calibrated according to a preset target false alarm rate.

[0047] The source tracing process includes:

[0048] Restore the embedded watermark bit vector Then, the extracted watermark is compared with the registered identity watermark database to select the most matching account based on the similarity index, thereby determining the source of the image and realizing the source tracing operation.

[0049] Compared with existing technologies, the image protection method that balances robustness and generative diversity provided by this invention achieves a robust and covert generative image detection and tracing scheme, so that the watermark embedded in the image can take into account both the robustness of the watermark and the diversity of image generation; moreover, due to the use of distributed sampling techniques, the technical solution provided by this invention does not have an adverse impact on image quality during implementation. Attached Figure Description

[0050] To more clearly illustrate the technical solutions of the embodiments of the present invention, the drawings used in the following description of the embodiments will be briefly introduced. Obviously, the drawings described below are only some embodiments of the present invention. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.

[0051] Figure 1 This is a schematic diagram comparing existing truncated sampling with the tail-truncated sampling provided in the embodiments of the present invention;

[0052] Figure 2 This is a schematic diagram illustrating the implementation process of the image protection method provided in an embodiment of the present invention. Detailed Implementation

[0053] The technical solutions in the embodiments of the present invention will be clearly and completely described below with reference to the specific content of the present invention. Obviously, the described embodiments are only some embodiments of the present invention, and not all embodiments, which do not constitute a limitation of the present invention. Based on the embodiments of the present invention, all other embodiments obtained by those skilled in the art without creative effort are within the protection scope of the present invention.

[0054] First, the following explanations are provided for the terms that may be used in this article:

[0055] The term "and / or" means that either or both can be achieved simultaneously. For example, X and / or Y means that it includes both "X" or "Y" as well as the three cases of "X and Y".

[0056] The terms "comprising," "including," "containing," "having," or other similar semantic descriptions should be interpreted as non-exclusive inclusion. For example, including a technical feature element (such as raw material, component, ingredient, carrier, dosage form, material, size, part, component, mechanism, device, step, process, method, reaction conditions, processing conditions, parameter, algorithm, signal, data, product or article of manufacture, etc.) should be interpreted as including not only the expressly listed technical feature element, but also other technical feature elements that are not expressly listed and are well-known in the art.

[0057] The term "composed of" excludes any technical features not expressly listed. When used in a claim, it closes the claim to exclude all technical features other than those expressly listed, except for associated conventional impurities. If the term appears only in a clause of a claim, it limits the claim to the elements expressly listed in that clause; elements recited in other clauses are not excluded from the overall claim.

[0058] Unless otherwise explicitly specified or limited, the terms "installation," "connection," "linking," and "fixing," etc., should be interpreted broadly. For example, they can refer to fixed connections, detachable connections, or integral connections; they can refer to mechanical connections or electrical connections; they can refer to direct connections or indirect connections through an intermediate medium; and they can refer to the internal connection between two components. Those skilled in the art can understand the specific meaning of the above terms in this document according to the specific circumstances.

[0059] When concentration, temperature, pressure, size, or other parameters are expressed as numerical ranges, such ranges should be understood to specifically disclose all ranges formed by any pairing of upper limits, lower limits, or preferred values ​​within that range, regardless of whether the range is explicitly stated; for example, if the numerical range "2 to 8" is stated, then that range should be interpreted to include ranges such as "2 to 7", "2 to 6", "5 to 7", "3 to 4 and 6 to 7", "3 to 5 and 7", "2 and 5 to 7", etc. Unless otherwise stated, the numerical ranges described herein include both their endpoints and all integers and fractions within that range.

[0060] The terms “center,” “longitudinal,” “lateral,” “length,” “width,” “thickness,” “upper,” “lower,” “front,” “back,” “left,” “right,” “vertical,” “horizontal,” “top,” “bottom,” “inner,” “outer,” “clockwise,” and “counterclockwise” indicate the orientation or positional relationship based on the orientation or positional relationship shown in the accompanying drawings. They are used only for the convenience and simplification of description and do not imply that the device or component referred to must have a specific orientation, or be constructed and operated in a specific orientation. Therefore, they should not be construed as limitations on this document.

[0061] This invention addresses the problem of existing diffusion models struggling to balance the diversity and robustness of watermark generation in image watermarking. It proposes a method for image protection that balances robustness and generation diversity: a two-stage truncated watermarking method based on tail-truncation sampling, abbreviated as T2SMark. On one hand, it utilizes bilateral truncated sampling of a normally distributed system to eliminate easily disturbed intermediate regions, improving robustness. On the other hand, it injects controllable randomness through a two-stage hybrid encryption process to enhance the diversity of watermarked images, ultimately achieving a good balance between the two. This method is of great significance for solving the problem of tracing and evidence collection from generated images.

[0062] In the specific implementation process of this invention, the watermark embedding process in the image can be mainly achieved through the following technical processing solutions:

[0063] (1) The watermark is embedded in the symbol of the initial noise of the diffusion model by using distribution-preserving sampling. Since this processing method does not change its distribution, the watermark is lossless for a single image.

[0064] (2) Taking advantage of the normal distribution characteristics of the initial noise of the diffusion model, a tail-truncation sampling scheme was adopted to avoid embedding the watermark information into the distribution center region that is susceptible to noise and may be flipped, thereby improving the robustness of the watermark to various distortions.

[0065] (3) To address the issue of generation diversity, a two-stage hybrid encryption scheme is adopted. That is, a static master key is used to encrypt a dynamic session key that changes randomly each time it is generated, and the dynamic session key is also embedded in the image. Then, the dynamic session key is used to encrypt the real watermark content, injecting controllable randomness into the watermark codeword to enhance the generation diversity.

[0066] In summary, the image watermarking method proposed in this invention provides a robust and covert solution for generated image detection and tracing. It can balance the robustness of the watermark with the diversity of image generation, and because it uses a distribution-preserving sampling technique, it has no negative impact on image quality.

[0067] During the implementation of this invention, it was discovered that Gaussian samples closer to the origin are more prone to sign errors in the presence of noise. Therefore, this invention proposes T2SMark, a novel two-stage watermarking scheme based on Tail-Truncated Sampling (TTS). This differs from previous schemes that primarily focused on discrete coding strategies, such as... Figure 1 As shown, TTS no longer uses the method of simply mapping bits to positive / negative signs, but instead... Figure 1 The Gaussian distribution is divided into three regions: a region for bit 0, a region for bit 1, and an undecided region (i.e., ...). Figure 1 The middle part of the diagram on the right (As shown in the figure). In this way, information can be embedded only in the more reliable tail region; while the undetermined central region is randomly sampled to compensate for the generation diversity and maintain the overall distribution.

[0068] Furthermore, the two-stage framework of this invention introduces controllable randomness through layered key encryption. Specifically, in the first stage, a random session key is encrypted using a static master key to obtain a random session key encoding. In the second stage, the actual watermark bits are then encrypted using the random session key to generate the corresponding watermark encoding. This layered encryption randomizes both parts of the watermark codeword, thereby further ensuring the diversity of generation. Finally, this invention performs multi-dimensional projection on the reconstructed Gaussian noise to fully utilize its rich continuous information, thereby improving detection and decoding performance.

[0069] Regarding the above-mentioned Figure 1 The analysis and description provided below, in order to facilitate a further understanding of the present invention, will be combined with the appendix. Figure 2 The specific implementation process of the embodiments of the present invention will be described in detail.

[0070] like Figure 2 As shown, the specific implementation of T2SMark is intuitively demonstrated.

[0071] First, the initial noise vector is divided into two segments: the first segment encodes a random session key under the control of a fixed master key (i.e., a static master key), and the second segment uses the random session key to encode the watermark to be embedded for image protection (i.e., the actual watermark bits). In this way, since the processing of both segments depends on the random session key, the final watermark noise vector remains randomized.

[0072] Formally, in the single-stage case, it can be... The noise space is considered as a direct sum of a set of mutually orthogonal subspaces; the key selects a hyperplane passing through the origin in a pseudo-random manner to divide each subspace into two half-spaces for binary encoding; based on this, in this embodiment of the invention, tail-truncated sampling (TTS) is used for sampling, that is: sampling along the normal of each hyperplane, and ensuring that the sampling point maintains at least a predetermined truncation threshold distance from the boundary. This operation based on truncation thresholds can create a robust watermark coding subspace within a region consisting of vectors with large modulus lengths, enabling it to resist cross-boundary perturbations; while randomness is preserved in the remaining dimensions that do not carry bits.

[0073] In the extraction stage, the noise vector can be reconstructed by the inversion algorithm and projected onto the normal direction of the same set of hyperplanes; each bit is recovered by the sign of its projection, that is, the half space in which the sample is located is determined; at the same time, the L1 norm of the projection is used as the confidence measure for detection; it can be seen that this method can make full use of the continuous structure of the noise vector.

[0074] The specific implementation methods of the watermark encoding and decoding processes in the embodiments of the present invention will be further described in detail below.

[0075] (a) Watermark Encoding Process

[0076] This watermarking encoding refers to the complete process of mapping a watermark to a continuous watermark noise vector, and it is not just a discrete encoding step.

[0077] set up The initial noise dimension, This is the cutoff threshold for TTS. The length of the watermark. For bit vectors, The set of natural numbers;

[0078] During the watermark encoding stage, The initial noise (i.e., the noise in dimension 1) is divided into 1-dimensional noise. and The two segments (satisfying) This typically corresponds to splitting a vector into two segments based on channels; then, a sample is obtained using the master key. dimensional vector To encode random session keys To obtain the random session key encoding; then, using the random session key... Sample one dimensional vector To carry watermark encoded bits (i.e., the watermark to be embedded for image protection), to obtain the watermark encoding; finally, the random session key encoding and the watermark encoding are concatenated to obtain the watermarked noise vector. The watermark noise vector is then embedded into the image to be protected, wherein... This indicates vector concatenation, which allows watermarks to be embedded into the corresponding image for protection.

[0079] Based on the above description, the implementation process of the corresponding watermark encoding can specifically include:

[0080] (1) Based on truncation threshold Determine the expected number of tail sampling dimensions. and the subspace dimension per bit :

[0081] ;

[0082] (2) Using the corresponding key for each layer (i.e., different keys for different layers) as the seed for the pseudo-random number generator, generate... pseudo-random vectors Each of them The support sets of these vectors are pairwise disjoint and mutually orthogonal, and are used as the normals of the bit-coded hyperplane in the watermark embedding subspace, i.e.:

[0083] ;

[0084] (3) Construct a binary mask This mask can Each dimension is divided into a watermark encoding subspace ( ) and random noise subspace ( );

[0085] For the initial noise vector Each component is subjected to tail-truncation sampling (TTS) using the following formula:

[0086] ;

[0087] Among them, the truncated normal distribution Indicates the interval superior The normalized cutoff distribution, The mean, Standard deviation;

[0088] Reference Figure 1 The image on the right shows the " " represents the random noise subspace, i.e., the intermediate interval" The watermark encoding subspace corresponds to the left tail of "0" and the right tail of "1" in the diagram, i.e., the left and right tails. Section; at the same time, it can be seen that, Figure 1 The tail truncation sampling method provided by the embodiment of the present invention shown on the right side is similar to... Figure 1 The existing truncation sampling methods provided on the left side of the middle section are significantly different;

[0089] (4) Final watermarked noise vector By combining the symbol corresponding to the watermark bit contained in the watermark encoding with the initial noise vector in the tail-truncated sampling dimension. The amplitude is determined by preserving the original random samples in the center dimension representing the pure noise region in the initial noise vector to construct the following:

[0090] ;

[0091] in, This represents the Hadamard product. It is in the process of Bits within the defined subspace Directional coding.

[0092] The above processing steps can achieve the corresponding watermark encoding process and obtain the watermark noise vector used to embed the watermark in the image.

[0093] (II) Watermark Decoding Process

[0094] When extracting the watermark, layered decoding is required, which involves reconstructing the watermarked noise vector through diffusion inversion. Then, the noise will be reconstructed. Divided into and Then, first use the master key from recover reuse from Decode to obtain the corresponding watermark bit vector This allows for the extraction of the watermark; the corresponding random session key... It simultaneously plays two roles: it is both the embedded payload of the first segment and the key of the second segment, thus enabling... Globally inject randomness while maintaining generation diversity.

[0095] The corresponding decoding process to obtain the corresponding watermark bit vector may include: the specific decoding method of each layer depends on the sign characteristics of the vector inner product for direct projection decoding, that is, watermark decoding and extraction is performed by multi-dimensional projection of the reconstructed Gaussian noise.

[0096] Specifically, the vector set can be regenerated using the corresponding key for each layer. Calculate the projection vector. for With each Inner product:

[0097] ;

[0098] Then, the watermark bit vector is directly recovered from the sign of the projected component. :

[0099] ;

[0100] in, Acting on elements dimensional vector When each bit is encoded in dimension Less than When the repetition is reduced, TTS can achieve a higher signal-to-noise ratio (SNR), and theoretically, a lower bit error rate can be obtained under the assumption of additive white Gaussian noise (AWGN).

[0101] The above-described watermark encoding and decoding processes can effectively protect images. Furthermore, the image watermarking method provided in this embodiment offers a robust and covert solution for generated image detection and tracing, which balances the robustness of the watermark with the diversity of image generation.

[0102] In order to detect the embedding of watermarks and trace their origin, the following embodiments of the present invention also provide corresponding detection and tracing processes.

[0103] (1) The detection and processing procedure includes:

[0104] Considering the potential error propagation in the two-stage decoding, detection primarily relies on the first stage. Leveraging the larger projection norm provided by TTS, the L1 norm of the first-stage projection vector is used as the test statistic:

[0105] ;

[0106] in, The number of bits in the random session key. It is Projected onto the first stage normal set The resulting projection vector. The statistics. With threshold Compare and calibrate the threshold according to the predetermined target false alarm rate;

[0107] Specifically, based on the threshold calibration results, the statistics are... With threshold If l exceeds d, the image is determined to have a watermark embedded; if l does not exceed d, the image is determined not to have a watermark embedded. The corresponding threshold d can be pre-calibrated according to the target false alarm rate to control the probability of incorrect judgment.

[0108] (2) The traceability process includes:

[0109] For the source tracing process, a complete two-stage decoding is required to recover all embedded watermark bit vectors. Then, the deciphered watermark is compared with the registered identity watermark database. The most matching account is selected by similarity indicators such as match count or Hamming distance, thereby determining the source of the image (i.e., completing the source tracing process), and thus realizing the tracking and accountability for the misuse of the image.

[0110] To further verify the performance of the technical solutions provided in the embodiments of the present invention, and to facilitate understanding of the above-mentioned effects that the embodiments of the present invention can produce, experiments will be conducted below to illustrate the performance.

[0111] I. Experimental Setup

[0112] The image generation backbone was Stable Diffusion v2.1 (SD v2.1), configured with a guidance scale of 7.5, DDIM denoising steps of 50, and a fixed output resolution of 512×12. T2SMark used a 16-bit session key and a 256-bit watermark, with a truncation threshold of 0.674. A random session key was embedded into the first channel of the initial noise (this setting is derived from empirical conclusions drawn from parameter selection research based on embodiments of this invention). All experiments were conducted using PyTorch 2.4.1 on a single NVIDIA RTX A6000 GPU.

[0113] This study compares three existing methods: traditional post-processing (dwtDct, dwtDctSvd, RivaGAN), fine-tuning paradigm (StableSignature), and inversion paradigm (Tree-Ring, TRW; Gaussian Shading, GS; PRC-Watermark, PRCW, etc.). All inversion methods undergo a 10-step DDIM inversion; empty hints are used during inversion, and the guidance scale is fixed to 1 to simulate unknown hint conditions. To ensure capacity fairness, dwtDct, dwtDctSvd, GS, and PRCW are all embedded with 256 bits; RivaGAN and Stable Signature use 32 bits and 48 bits respectively (following their official implementations).

[0114] During the experiments, evaluations were conducted on the MS-COCO-2017 (COCO) and Stable-Diffusion-Prompt (SDP) datasets. Robustness was relatively stable under the given detection settings. time Bit-by-bit accuracy was compared under traceability settings. Specifically, for each method, 500 cues were sampled from the SDP training partition, generating 500 watermarked images. Nine distortions were applied, followed by detection and traceability evaluation. Generation diversity was assessed using LPIPS: For each non-post-processing method, 10 images were generated for each of the 1000 COCO test cues, with a fixed master key and watermark. The LPIPS for 45 unique image pairs for each cue were calculated, and the overall mean was reported (post-processing methods do not change the generation process and are therefore not included in this evaluation). For visual quality, CLIP scores and FID were reported. Ten independent trials were conducted: each trial fixed a key-watermark pair (simulating a single user), generating 1000 images using COCO test cues to calculate CLIP and FID; subsequently, a two-sample t-test was performed to evaluate the impact of the watermark on quality. The null hypothesis was that "the mean of watermarked and unwatermarked images is the same," and smaller t-values ​​indicated stronger support for the null hypothesis.

[0115] II. Main Results Obtained from the Experiment

[0116] Table 1 below presents the experimental results for each method, comparing the performance of T2SMark with other watermarking methods on SDv2.1. T2SMark performs best in traceability scenarios, while its detection performance is comparable to the best method, GS. In contrast, PRCW exhibits the worst traceability performance, with a detection TPR below 30% under adversarial conditions, making it difficult to meet the reliability requirements of practical deployments.

[0117] In terms of generation diversity (measured by LPIPS), PRCW scored the highest; T2SMark followed closely behind, with a difference of less than 1e-3, which is almost negligible. GS had the lowest diversity score, which is related to its use of fixed codewords for each user; StableSignature and TRW also showed a significant decrease in diversity.

[0118] Regarding image quality, only T2SMark and PRCW achieved no degradation in both benchmarks. While GS's CLIP score was competitive, its FID was significantly lower than the watermark-free baseline. Furthermore, GS's CLIP and FID standard deviations were significantly larger, indicating that its generation quality is more sensitive to specific watermarks and keys. This sensitivity can lead to user experience issues: inconsistent generation quality may occur when different accounts (corresponding to different watermarks / keys) participate.

[0119] Taking all indicators into account, the T2SMark technical solution provided in this embodiment of the invention achieves optimal overall balance.

[0120] Table 1

[0121]

[0122] III. The Undetectability of Embedded Watermarks

[0123] To evaluate the undetectability of the embedded watermark, a ResNet-18 classifier was trained to distinguish between watermarked and unwatermarked images. Four inversion-based methods were evaluated, and a fixed key and watermark were assigned to each method. The training set for each method consisted of 8000 watermarked images and 8000 clean images, with 500 images in each test set. Training lasted 10 epochs with a batch size of 128 and a learning rate of... .

[0124] Table 2 below shows the test accuracy of the corresponding undetectability, that is, it gives the comparison results of the undetectability of various generated image watermarks.

[0125] The results show that TRW and GS are relatively easier to detect. Although PRCW has the highest undetectability, it still cannot completely avoid being identified. T2SMark achieved the second-best performance, and is also difficult to detect, demonstrating a high degree of undetectability.

[0126] Table 2

[0127]

[0128] IV. Generalization

[0129] T2SMark was evaluated on the Stable Diffusion v3.5 Medium (SD v3.5M) model, along with other inversion-based watermarking methods. This model employs a DiT (Diffusion Transformer) as the denoising network and has a 16-channel latent space.

[0130] Table 3 below shows the performance comparison results of T2SMark and other watermarking methods on SD v2.1. The results indicate that SD v3.5M has stronger generation quality, thus all methods can maintain high visual fidelity. Compared with its performance on SDv2.1, TRW's robustness is significantly reduced, and GS shows a similar loss in diversity. PRCW's traceability is improved, but it still lags far behind the best methods. In contrast, T2SMark achieves the best balance between robustness and diversity, and its undetectability is almost indistinguishable from PRCW.

[0131] Table 3

[0132]

[0133] Through the above description of the embodiments, those skilled in the art can clearly understand that the above embodiments can be implemented by software, or by using software plus necessary general-purpose hardware platforms. Based on this understanding, the technical solutions of the above embodiments can be embodied in the form of a software product, which can be stored in a non-volatile storage medium (such as a CD-ROM, USB flash drive, mobile hard drive, etc.), including several instructions to cause a computer device (such as a personal computer, server, or network device, etc.) to execute the methods described in the various embodiments of the present invention.

[0134] Furthermore, the above description is merely a preferred embodiment of the present invention, but the scope of protection of the present invention is not limited thereto. Any variations or substitutions that can be easily conceived by those skilled in the art within the scope of the technology disclosed in the present invention should be included within the scope of protection of the present invention. Therefore, the scope of protection of the present invention should be determined by the scope of the claims. The information disclosed in the background section is intended only to enhance the understanding of the overall background technology of the present invention and should not be construed as an admission or implication in any way that such information constitutes prior art known to those skilled in the art.

Claims

1. An image protection method balancing robustness and generation diversity, characterized in that, The method comprises: encoding a randomly generated random session key with a preset master key to obtain a random session key encoding; and encoding a watermark to be embedded for image protection with the random session key to obtain a watermark encoding; generating a watermark noise vector based on the random session encoding and the watermark encoding, and embedding the watermark noise vector into an image to be protected.

2. The method of claim 1, wherein, The process of generating the watermark noise vector comprises: sample encoding the random session key with the master key on an initial noise vector to obtain a random session key vector of corresponding dimension; and sample encoding the watermark to be embedded with the random session key to obtain a watermark vector for carrying watermark bits; constructing the watermark noise vector based on the random session key vector and the watermark vector.

3. The method according to claim 1 or 2, characterized in that, The process of embedding the watermark noise vector into the image to be protected comprises: sample the initial noise vector in a tail truncation sampling (TTS) manner, which specifically means sampling along the normal of each hyperplane and keeping the sampling point at least a predetermined truncation threshold distance from the boundary; determine a watermark encoding subspace for embedding the watermark noise vector through the predetermined truncation threshold, and embed the watermark noise vector into the watermark encoding subspace.

4. The method of claim 3, wherein, The process of determining the watermark encoding subspace for embedding the watermark noise vector comprises: Truncating the tail to sample the expected number of dimensions determining the total dimension of the watermark encoding subspace, and determining the subspace dimension per bit allocating the subspace dimension per bit to determine the watermark encoding subspace; wherein the expected number of dimensions of the tail is determined and the subspace dimension per bit The calculation formula is respectively: ; wherein, is an initial noise dimension, is a predetermined truncation threshold for TTS, is a watermark length.

5. The method of claim 4, wherein, The process of constructing the watermark noise vector comprises: combine the signs corresponding to the watermark bits contained in the watermark encoding and the amplitude of the initial noise vector in the tail truncation sampling dimension, and keep the original random sample in the center dimension of the initial noise vector representing the pure noise region, to construct the watermark noise vector.

6. The method of claim 5, wherein, In the process of constructing the watermark noise vector, the watermark noise vector is obtained The formula includes: ; wherein, denotes the Hadamard product, is the directionality encoding of the bit vector within the subspace defined by ; ; are m pseudo-random vectors generated with the respective key of each layer as the seed of a pseudo-random number generator, and each pseudo-random vector ; is a binary mask used to divide the dimensions into a watermark encoding subspace and a random noise subspace.

7. The method of claim 6, wherein, The support sets of the pseudo-random vector are pairwise disjoint and orthogonal to each other, and are used as the normal of the bit encoding hyperplane in the watermark embedding subspace, i.e. 。 8. The method of claim 6, wherein, The binary mask The watermark encoding subspace is The random noise subspace is ; and applying tail-truncation sampling to each component of the noise vector includes: ; wherein, denotes the normalized truncated distribution over the interval .​ 9. The method of claim 6, wherein, The method further comprises a watermark decoding process, and the process comprises: Reconstructing the watermarked noise vector through diffusion inversion Then, the noise will be reconstructed. Divided into random session key encoding vectors With watermark encoding vector Then, the master key is used to encode the vector from the random session key. Recover the random session key Then use a random session key From the watermark encoding vector Decoding to obtain the watermark bit vector ; The decoding process includes: using a corresponding key of each layer to regenerate a vector set ; calculating a projection vector The inner product of each projection component A d-dimensional real vector space, the projection vector The calculation formula of the first projection component ​​​​ ; from the projection component of the sign directly recovering the watermark bit vector is: ; wherein act on vector .

10. The method of claim 9, wherein, The method further comprises a watermark embedding detection and tracing process, wherein: The detection process comprises: use the L1 norm of the first-stage projection vector as the test statistic, i.e. ; wherein, is the number of bits of the random session key, is the projection of onto the first-stage normal set resulting projection vector; The statistical quantity is compared with a threshold value If l exceeds d, it is determined that the image is embedded with a watermark; if l does not exceed d, it is determined that no watermark is embedded; the threshold value d is previously calibrated according to a preset target false alarm rate. The tracing process comprises: Restoring embedded watermark bit vector The extracted watermark is then compared with the registered identity watermark database to match the similarity index to select the most consistent account, thereby determining the image source and realizing the traceability operation.