A method for image copyright protection based on adversarial deep learning
By embedding intelligent perturbation strategies through adversarial deep learning, the problem of unauthorized image regeneration in generative artificial intelligence models is solved, efficient copyright protection is achieved, it is applicable to a variety of generative models, and reduces manual operation costs.
Patent Information
- Application Number
- CN202510998570.2
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-07-21
- Publication Date
- 2025-10-03
- Estimated Expiration
- 2045-07-21
AI Technical Summary
Existing technologies cannot effectively block unauthorized image regeneration in generative AI models, making copyright protection of original works difficult to achieve. In particular, existing watermark and copyright legislation has limitations in cross-border and cross-platform infringements.
By embedding intelligent perturbation strategies through adversarial deep learning, we use text descriptions and image feature vectors to generate imperceptible perturbations in the diffusion model, forming active adversarial capabilities, blocking infringement paths, and protecting image copyrights.
It effectively reduces the similarity between the generated image and the original image, severely damages local details, achieves wide adaptability to different generation models, improves copyright protection efficiency, and reduces manual intervention.
Smart Images

Figure SMS_11 
Figure SMS_97
Abstract
Description
Technical Field
[0001] The present invention belongs to the field of image processing technology, and specifically relates to an image copyright protection method based on adversarial deep learning. Background Art
[0002] With the breakthrough development of generative AI technology, deep generative models, represented by the diffusion model, have widely empowered the image creation field, enabling non-professional users to easily generate high-quality digital artworks. However, this widespread adoption of this technology has also posed a serious risk of theft of original content: malicious users can obtain copyrighted original images without authorization and, through the diffusion model's text-guided or image-guided generation capabilities, directly reproduce the original work's stylistic features or core visual elements. For example, attackers can exploit the diffusion model's fine-tuning mechanism to use original paintings as training material, extracting their brushstrokes, color schemes, and composition patterns, and then mass-produce derivative works with similar styles. In commercial design scenarios, plagiarists can use the image-to-image function to generate variant images based on original design drafts, circumventing copyright review. This behavior not only constitutes a substantial infringement of the original creator's intellectual property rights, but also leads to a vicious cycle in the art market where "bad money drives out good money," squeezing the living space of original artists and hindering the healthy development of the digital art ecosystem.
[0003] Currently, artists are primarily protected through legislation to protect the copyright of images, or by watermarking them. However, copyright legislation is subject to geographical restrictions and lengthy judicial processes, making it difficult for artists to quickly obtain evidence and pursue accountability for cross-border and cross-platform infringements. Watermarks only impose external constraints on the work, failing to penetrate deep into the model to block infringement paths. Consequently, conventional visible watermarks are easily ineffective, leaving original images at risk of misuse within generative AI systems.
[0004] Therefore, there is an urgent need to develop an active protection solution based on technological confrontation. By embedding intelligent perturbation strategies in images, when artists' works enter the generative artificial intelligence model, they can trigger the model failure mechanism or interfere with its core generation logic, thereby blocking unauthorized feature extraction and image regeneration from the root, and achieving a balance between copyright protection and technological development. Summary of the Invention
[0005] To address the shortcomings of the above-mentioned existing technologies, the present invention proposes a method for image copyright protection based on adversarial deep learning. Through adversarial deep learning, the protected images are embedded in intelligent perturbations to reduce the risk of original images being abused in generative artificial intelligence systems.
[0006] The technical solution provided by the present invention is as follows: a method for protecting image copyright based on adversarial deep learning, comprising the following steps:
[0007] S1. Provide a text description of the image to be protected;
[0008] S2, perform deep semantic analysis on the text description and then encode it into a text feature vector;
[0009] S3. Input the image to be protected into the variational autoencoder to parse the image content and output the image feature vector;
[0010] S4. Input the text feature vector and the image feature vector into the diffusion model and calculate the loss value;
[0011] S5. Add random noise to the image to be protected, combine the back propagation algorithm with the loss value of S4 to generate a disturbance, and superimpose the disturbance on the image to be protected; repeat this step multiple times to obtain an iterative guidance graph;
[0012] S6. Input the iterative guidance map into the variational autoencoder and output the iterative guidance map feature vector. Calculate the loss function of the image feature vector and the iterative guidance map feature vector, as well as the loss function of the two in the semantic space. Taking the distance between the two loss functions as the target, adjust the image through gradient backpropagation and perturbation superposition, and finally obtain a picture with copyright protection capabilities.
[0013] Compared with the prior art, the present invention has at least the following beneficial effects:
[0014] (1) By embedding imperceptible intelligent perturbations, the present invention blocks the technical path of infringers using diffusion models to extract styles and regenerate images at the model input level, thus forming an active countermeasure capability. Compared with existing image protection technologies, the method of the present invention is more effective. The generated image is less similar to the image to be protected, the damage to local details is more serious, and the protection of the image is more effective.
[0015] (2) The algorithm design does not rely on the architectural details of a specific generative artificial intelligence model. Whether it is the mainstream diffusion models such as StableDiffusion and Midjourney, or other image generation models based on GAN and Transformer, the optimization mechanism of semantic features and distance metrics can effectively resist feature theft and image regeneration, thereby reducing the technical cost of adapting different models.
[0016] (3) From text description generation and feature extraction to perturbation map output, the entire process does not require human intervention and can be integrated into image storage, transmission, or publishing platforms to achieve batch automated copyright protection for massive images, greatly improving the protection efficiency of artists and copyright holders and reducing the time and labor costs of manual operations. DETAILED DESCRIPTION
[0017] To make the objectives, technical solutions, and advantages of the embodiments of the present invention more clear, the technical solutions in the embodiments of the present invention are clearly and completely described below. Obviously, the described embodiments are only part of the embodiments of the present invention, not all of them. Based on the embodiments of the present invention, other embodiments obtained by ordinary technicians in this field without making any creative efforts are all within the scope of protection of the present invention.
[0018] A method for protecting image copyright based on adversarial deep learning, comprising the following steps:
[0019] S1. Provide a text description of the image to be protected.
[0020] In this step, the features of the image to be protected need to be described and text generated. This process can be done manually. For example, for an image, features can be directly entered: "a little girl wearing a red hat and a dress," "an ink painting depicting mountains and rivers." When performing manual description, the general features of the image to be protected must be accurately described.
[0021] Alternatively, an image description model may be used to provide a text description of the image to be protected. There are many models for describing images, such as the common BLIP and DeepDanbooru, which can be applied to this embodiment.
[0022] This embodiment takes DeepDanbooru as an example to provide a specific image description process.
[0023] S11. The picture to be protected Perform scaling to get the processed image , then the picture The pixel values are normalized to obtain the feature map , its height is H, width is W, and number of channels is C.
[0024] S1.2 Feature Map Input into the convolution layer of the DeepDanbooru model, and extract image features through multi-layer convolution. The processing process of one convolution is:
[0025] Where, For the The convolution kernel of the layer extracts features, For the Layer output channels The bias term, , For the The pixel space coordinates of the layer output feature map, , For the The pixel space coordinates of the layer convolution kernel, For the p Layer input channel index, For the Layer output channel index, For the The size of the convolution kernel used by the layer to extract features.
[0026] Since DeepDanbooru is based on the ResNet architecture, it also needs to be connected through residual blocks:
[0027] in, ReLU (·) is the activation function, Conv (·) is a convolution operation, BatchNorm (·) is the normalization function, It is the feature map of the residual structure calculated by the q layer. By calculating the multi-layer residual structure, the feature map after feature extraction can be obtained. ,in , , are the height, width, and number of channels of the feature map after feature extraction.
[0028] S1.3 The feature map generated by S1.2 After dimensionality increase, it is compressed into a one-dimensional feature vector:
[0029] in, k pix is the feature map Y DD The number of channels after dimensionality increase after 1×1 convolution processing, , It is the pixel space coordinate of the feature map after feature extraction. v DD (·) is the feature vector finally formed by the feature map after feature extraction.
[0030] S1.4 Use the fully connected layer to transform the feature vector v DD (·) Through the fully connected layer mapping, output each category The predicted value of each category is then applied to the Sigmoid function to obtain the confidence score of each category. :
[0031] in For category The corresponding weight vector in the fully connected layer, v DDis the eigenvector calculated in S1.3, is the bias term in the fully connected layer.
[0032] S1.5 Setting the threshold (Default is 0.6), filter out Category collection .
[0033] S1.6 The filtered categories are grouped together Convert the template into a natural language description. For example, if the final category set output by the model is {"ink painting", "mountains and rivers", "rivers"}, it can be directly connected by "," to generate the final description "ink painting, mountains and rivers, rivers" as a sentence describing the picture.
[0034] After the above steps, the text description will be used as the output of step S1 for semantic parsing in the subsequent CLIP model. The text description maintains a semantic connection with the original image content, providing a foundation for subsequent copyright protection mechanisms.
[0035] S2. Perform deep semantic analysis on the text description and then encode it into a text feature vector. In this step, the text description is converted into a feature vector to ensure the semantic consistency between the text and the image. The method includes the following steps:
[0036] S21. Divide the text description into multiple basic units, and encode them using a byte pair encoding method to obtain a token sequence, where the token sequence consists of multiple tokens;
[0037] In this step, after being divided into multiple basic units, they are encoded, the vocabulary is used to find the ID corresponding to each basic unit, and special tokens are added: [CLS] indicates the start of the sequence, and [SEP] indicates the end of the sequence. Finally, the determined token sequence is obtained , where t m The ID corresponding to the mth basic unit.
[0038] S22, use the embedding matrix to map each token in the token sequence and add the position encoding to obtain a dense vector;
[0039] In this step, each token in the token sequence is first encoded by One_hot, and then embedded by the matrix W e Map the token into the semantic space in the model:
[0040] Among them, One_hot is achieved by VThe operation of encoding the index corresponding to the word in the vector as 1 and the rest as 0; V is the vocabulary size, d is the embedding dimension. Thus, we get the word embedding vector sequence , then add the word embedding vector sequence to the position code to get the model input :
[0041] in p i Positional encoding aims to preserve sequence information. Each token is converted into a denser feature vector, placing tokens with similar meanings closer together in the semantic space (e.g., "sad" is closer to "sad" and farther from "happy"). Positional encoding is also superimposed to incorporate positional information within the sequence.
[0042] The essence of this step is to map discrete symbols into a continuous semantic space, which enables the machine to understand the semantics of the text through vector operations without manually defined rules, greatly improving the efficiency and accuracy of semantic representation.
[0043] S23. Use a multi-layer Transformer encoder to extract text content features of the dense vector. After repeating this step for a predetermined number of times, a text feature vector is obtained.
[0044] In this step, the multi-head attention mechanism and feedforward network in the multi-layer Transformer encoder are used to extract text content features and calculate the attention on each basic unit. The calculation method of this attention is as follows:
[0045] in, A content for attention to be paid to the focus on the text; Q content It is the query matrix of the text content, representing the "query intention" of the token at the current position to the tokens at other positions; K content is the key matrix of the text content, through Q content The multiplied score represents the attention weight of the current token to other tokens. V content is the value matrix of the text content, representing the content information of each token. After the attention weights are calculated, these weights will be combined with V content Matrix multiplication aggregates information from other tokens. d k is the vector dimension scaling factor, softmax (·) is the normalized exponential function.
[0046] Subsequently, the content features of deeper text are extracted through a multi-layer Transformer encoder, and the encoders are connected through residual connections:
[0047] in For l Attention of the points of interest on the layer text features, For l The text feature matrix output by the layer, MSA (·) is the multi-head attention operation, LayerNorm (·) is layer normalization. FFN (·) is the feedforward neural network operation.
[0048] Subsequently, it will be Input, repeat this operation multiple times until a predetermined number of times is reached, and finally a text feature matrix is obtained. The first column of this text feature matrix is taken as the overall feature vector of the text content, and it is normalized to a length unit to obtain the text feature vector. The number of repetitions can be set to 12 times. Of course, those skilled in the art can set a different number of repetitions according to actual circumstances.
[0049] In this step, a multimodal model is used to transform the output of S1 into a feature vector, ensuring the semantic consistency between the text and the image. The feature vector will be used for matching with image features and perturbation optimization in subsequent steps.
[0050] S3. Input the image to be protected into a variational autoencoder to parse the image content and output an image feature vector. This specifically includes the following steps:
[0051] S31, scaling the image to be protected and normalizing the pixel values to obtain an initial feature map; in this step, the image to be protected x 0 to perform scaling processing and obtain the processed image I resize , then the picture I resize Normalize the pixel values to obtain the processed initial feature map X;
[0052] S32. Input the initial feature map into the multi-layer convolution in the variational autoencoder for calculation to obtain the final feature map; the calculation expression of the final feature map is: , where ConvBlock L For the Lth convolution block operation, a convolution block operation includes a convolution operation, a normalization operation and the use of a ReLU activation function.
[0053] After the first few convolutional layers, local patterns such as edges, textures, and colors are obtained. After the next few convolutional layers, global information such as object components and abstract semantics is obtained. Through multiple layers of convolution, deep information from the input image can be extracted.
[0054] S33. Flatten the final feature map into a feature vector and input it into the dual-path fully connected layer of the variational autoencoder to obtain multiple latent space variables. Combined with the reparameterization technique, the image feature vector is generated.
[0055] In this step, the final feature map F is first flattened into a feature vector f Input the two-way fully connected layer, and generate the latent space variables from this inference μ and σ :
[0056] in FC μ (·), FC σ (·) are all fully connected layers of the model. The flattened vector is mapped to the mean vector of the latent space through the fully connected layer. μ , controls the center position of the encoding; generates the logarithmic variance vector through another fully connected layer , controls the variance of the encoding distribution.
[0057] Then, a reparameterization technique is used to generate a differentiable image feature vector:
[0058] in, z protect represents the image feature vector, is the random noise sampled , Multiply corresponding items element-wise.
[0059] After the above steps, VAE will input the image to be protected x 0 is converted to a feature vector in the latent space z protect ,This vector captures the semantic and structural information of the image and will be used to match it with text features in subsequent steps.
[0060] S4. Input the text feature vector and the image feature vector into the diffusion model and calculate the loss value; it includes the following sub-steps:
[0061] S41. Based on the Markov chain and the implicit diffusion model, Gaussian noise is added to the image feature vector according to the formula to obtain the image feature vector with the noise added;
[0062] The forward process of the diffusion model follows the Markov chain assumption and gradually adds Gaussian noise to the data, while the implicit diffusion model (LDM) adds noise in the latent space. x 0 in the latent space feature vector , No. The noise adding process can be expressed as:
[0063] in For pictures to be protected x 0 is the feature vector in the latent space, In the The vector after adding Gaussian noise, α t is the noise parameter of step t, which controls the noise intensity of each step to decrease with t. When t is larger, The closer it is to the vector distribution of Gaussian noise.
[0064] S42, inputting the text feature vector and the image feature vector with noise added into the U-Net network of the diffusion model to obtain predicted noise;
[0065] For the U-Net network, The text feature vector is 2.
[0066] S43. Substitute the predicted noise into the loss function of the diffusion model to obtain the loss value:
[0067] in, is the loss value; is the random noise sampled . The square of the L2 norm is used to calculate the distance between the sampled noise and the predicted noise.
[0068] The above process uses the features of text and images to penetrate into the specific model, so as to associate the subsequent disturbance generation with these features, which is a key link in copyright protection technology.
[0069] S5: Add random noise to the image to be protected, combine the back propagation algorithm with the loss value of S4 to generate a disturbance, and superimpose the disturbance on the image to be protected; repeat this step multiple times to obtain an iterative guidance graph. This includes the following sub-steps:
[0070] S51. Add random noise to the image to be protected. This process mainly introduces simple controllable uncertainty to the image to be protected, so that the image has a certain degree of randomness. , where The random noise added .
[0071] S52. Calculate the gradient direction of each pixel in the image after adding noise based on the back propagation algorithm, and then calculate the disturbance value based on the gradient direction: , where represents the disturbance value; Represents the image to be protected after adding random noise in the 1st to Tth time steps in the implicit diffusion model x protect The loss value after adding disturbance, that is, the image to be protected after adding random noise; sign (·) represents the sign function, ;
[0072] S53. Superimpose the disturbance value on the image after adding random noise to obtain the image after the initial disturbance: , where is a clipping function, which means that the value of the superimposed pixel is within [0, 255]. If it exceeds 255, it is discarded directly. x protect Represents the image to be protected after random noise is added;
[0073] S54. Using the image after the initial disturbance as input, repeat the operations of S52 to S53 for a predetermined number of times to obtain an iterative guidance graph. The predetermined number of times for this step can be set to 100 times, or can be set according to actual conditions.
[0074] S6 inputs the iterative guidance map into the variational autoencoder and outputs the iterative guidance map feature vector, calculates the loss function of the image feature vector and the iterative guidance map feature vector and the loss function of the two in the semantic space, takes the distance between the two loss functions as the target, adjusts the image through gradient backpropagation and perturbation superposition, and finally obtains a picture with copyright protection capability. The method includes the following steps:
[0075] S61. Based on the method of S3, obtain the characteristic vector of the iterative guidance graph; the characteristic vector of the iterative guidance graph obtained here can be used as one of the reference benchmarks for perturbation optimization.
[0076] S62. Calculate the loss function of the image feature vector and the iterative guidance graph feature vector as well as the loss function of the two in the semantic space: 、 , where and They represent the distance loss function between the feature vector of the image to be protected after the perturbation and the feature vector of the guide image, and the distance loss function between the feature vector of the image to be protected and the feature vector of the image to be protected after the perturbation; A feature vector representing the image to be protected; Zs The feature vector representing the iterative guidance graph; represents the first The initial value of the disturbance The image to be protected after adding random noise x protect Images of the same size and number of channels but with all pixel values set to zero, where size refers to width and height;
[0077] S63. Calculate the distance loss value of the two loss functions obtained in S62: , where Indicates the distance loss value;
[0078] S64. Using the distance loss value as the optimization target, adjust the image through gradient backpropagation and perturbation superposition: , where Indicates the i The perturbation after the second iteration optimization is added, and its initial value is the image to be protected after random noise is added.
[0079] S65 , repeating the operations of S62 to S64 until a predetermined number of times is reached, and outputting a picture with copyright protection capability.
[0080] To further illustrate the effectiveness of the method of the embodiment of the present invention, tests were performed using the method of the embodiment of the present invention and existing common image copyright protection methods based on images from an existing cartoon face dataset and a real face dataset. During the test, four indicators, namely FID, KID, SSIM, and LPIPS values, were used for comparison.
[0081] FID is a metric used to evaluate the quality and diversity of generative models. The lower the FID value, the more similar the distribution of generated images is to the distribution of real images. KID, similar to FID, is also a metric used to evaluate the quality of generative models, but KID takes into account details such as image edges (for example, if the image to be protected is a man's face, a dissimilar distribution means a dog's face was generated). SSIM is a metric used to measure the similarity between two images. It considers brightness, contrast, and structure. The closer the SSIM value is to 1, the more similar the two images are. LPIPS is an image similarity assessment metric based on deep learning. LPIPS is more consistent with human visual perception and can capture structural and content similarities in images. FID and KID consider the similarity between the generated image and the entire dataset, while SSIM and LPIPS consider the similarity between the original image and the generated image.
[0082] The FID and KID are calculated using the torch-fidelity library, the LPIPS is calculated using the lpips library, and the SSIM indicator is calculated using Skimage.
[0083] Since the method of the embodiment of the present invention is used for protecting images, the use of these four indicators needs to be reversed from the evaluation method in the generation model (for example, the lower the FID value, the better in the generation model, while the higher the FID value, the better when generating the protection image algorithm). The final results are shown in Table 1.
[0084] Table 1 Test results of the method of the embodiment of the present invention and the existing method
[0085]
[0086] As can be seen from the table, the method of this embodiment can make the distribution of the generated image less similar to the distribution of the real image, and at the same time reduce the similarity between the generated image and the image to be protected. Therefore, the method of this embodiment can make the destruction of the overall appearance and local details of the generated image when the protected image is placed in the generative model more effective than the traditional method.
[0087] The above description is merely a preferred embodiment of the present invention and does not constitute any form of limitation to the present invention. Although the present invention has been disclosed as a preferred embodiment as above, it is not intended to limit the present invention. Any technician familiar with this profession can make some changes or modifications to equivalent embodiments of the technical contents disclosed above without departing from the scope of the technical solution of the present invention. However, any simple modifications, equivalent changes and modifications made to the above embodiments based on the technical essence of the present invention without departing from the content of the technical solution of the present invention are still within the scope of the technical solution of the present invention.
Claims
1. A method for protecting image copyright based on adversarial deep learning, characterized in that: The following steps are involved: S1. Provide a text description of the image to be protected; S2, perform deep semantic analysis on the text description and then encode it into a text feature vector; S3. Input the image to be protected into the variational autoencoder to parse the image content and output the image feature vector; S4. Input the text feature vector and the image feature vector into the diffusion model and calculate the loss value; The following steps are included: S41. Based on the Markov chain and the implicit diffusion model, Gaussian noise is added to the image feature vector according to the formula to obtain the image feature vector with the noise added; S42, inputting the text feature vector and the image feature vector with noise added into the U-Net network of the diffusion model to obtain predicted noise; S43, substituting the predicted noise into the loss function of the diffusion model to obtain a loss value; S5. Add random noise to the image to be protected, combine the back propagation algorithm and the loss value to generate disturbance, and superimpose the disturbance on the image to be protected; Repeat this step multiple times to obtain an iterative guidance graph; S6. Input the iterative guidance map into the variational autoencoder and output the iterative guidance map feature vector. Calculate the loss function of the image feature vector and the iterative guidance map feature vector, as well as the loss function of the two in the semantic space. Taking the distance between the two loss functions as the target, adjust the image through gradient backpropagation and perturbation superposition, and finally obtain a picture with copyright protection capabilities.
2. The method according to claim 1, characterized in that In S1, a text description is provided by the user in a manner that reflects the content of the image to be protected.
3. The method according to claim 1, characterized in that In S1, a text description of the image to be protected is performed through the image description model.
4. The method according to claim 1, wherein In S2, the CLIP large model is used to perform deep semantic analysis of text descriptions.
5. The method according to claim 4, characterized in that The method of using the CLIP large model to perform deep semantic parsing of text descriptions includes the following steps: S21. Divide the text description into multiple basic units, and encode them using a byte pair encoding method to obtain a token sequence, where the token sequence consists of multiple tokens; S22, use the embedding matrix to map each token in the token sequence and add the position encoding to obtain a dense vector; S23. Use a multi-layer Transformer encoder to extract text content features of the dense vector. After repeating this step for a predetermined number of times, a text feature vector is obtained.
6. The method according to claim 1, wherein S3 includes the following steps: S31, scaling the image to be protected and normalizing the pixel values to obtain an initial feature map; S32, inputting the initial feature map into the multi-layer convolution in the variational autoencoder for calculation to obtain the final feature map; S33. Flatten the final feature map into a feature vector and input it into the dual-path fully connected layer of the variational autoencoder to obtain multiple latent space variables. Combined with the reparameterization technique, the image feature vector is generated.
7. The method according to claim 1, characterized in that S5 includes the following sub-steps: S51, adding random noise to the image to be protected; S52. Calculate the gradient direction of each pixel in the image after adding noise based on the back propagation algorithm, and then calculate the disturbance value based on the gradient direction: , where represents the disturbance value; Represents the image after adding noise at time steps 1 to T in the implicit diffusion model The loss value; sign (·) represents the sign function ; S53. Superimpose the disturbance value on the image after random noise to obtain the image after the initial disturbance: , where is the clipping function, which means the value of the superimposed pixel is within [0,255]; x protect Represents the image to be protected after random noise is added; S54: Using the image after the initial disturbance as input, repeat the operations of S52 to S53 for a predetermined number of times to obtain an iterative guidance graph.
8. The method according to claim 7, characterized in that S6 includes the following sub-steps: S61. Based on the method of S3, obtain the feature vector of the iterative guidance graph; S62. Calculate the loss function of the image feature vector and the iterative guidance graph feature vector as well as the loss function of the two in the semantic space: 、 , where and They represent the distance loss function between the feature vector of the image to be protected after the perturbation and the feature vector of the guide image, and the distance loss function between the feature vector of the image to be protected and the feature vector of the image to be protected after the perturbation; A feature vector representing the image to be protected; Z s represents the eigenvector of the iterative guided graph; VAE(·) represents the VAE eigenvector operator; represents the first The initial value of the disturbance is x protect An image with the same size and number of channels but with all pixel values set to zero; S63. Calculate the distance loss value of the two loss functions obtained in S62: , where Indicates the distance loss value; S64. Using the distance loss value as the optimization target, adjust the image through gradient backpropagation and perturbation superposition: , where Indicates the i The perturbation after the second iteration optimization is added, and its initial value is the image to be protected after adding random noise; S65 , repeating the operations of S62 to S64 until a predetermined number of times is reached, and outputting a picture with copyright protection capability.
Citation Information
Patent Citations
Image copyright protection method based on diffusion model and adversarial attack
CN116451184A
Zero-order optimization-based diffusion model artistic copyright protection method and device
CN119885113A