Generated image traceability detection method and system based on reconstruction

Through the reconstruction-based generation image traceability detection method, the image is extracted and processed by using a variational autoencoder and a deep learning network, the problem of insufficient space consumption and access performance of generated image traceability detection in the prior art is solved, and the traceability of generated image traceability and multi-class image traceability under black box conditions are realized.

CN120125977APending Publication Date: 2025-06-10INSTITUTE OF INFORMATION ENGINEERING CHINESE ACADEMY OF SCIENCES
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510062866.3
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-01-15
Publication Date
2025-06-10

AI Technical Summary

Technical Problem

The existing methods of generating image traceability detection have shortcomings in space consumption and access performance, especially when processing images generated by different generation models, it is difficult to achieve traceability under black box conditions.

Method used

The reconstruction-based image traceability detection method is used to reconstruct the image through a variational autoencoder, obtain the numerical difference feature matrix before and after reconstruction, and process these features using a convolutional neural network and Vision Transformer network to generate global features of the image, and finally realize traceability detection of the generated image.

Benefits of technology

This method can effectively reduce the overhead of data processing, support the traceability of images generated by GAN and diffusion models under black box conditions, and has a variety of traceability capabilities for generating images, especially in the case of data imbalance, which can still maintain good detection results.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120125977A_ABST
    Figure CN120125977A_ABST
Patent Text Reader

Abstract

The invention discloses a reconstruction-based generated image traceability detection method and system, and belongs to the technical field of machine learning. The method comprises the following steps: reconstructing a generated image, and obtaining a numerical value difference characteristic matrix before and after reconstruction; performing convolution and pooling operation on the numerical difference feature matrix by using a convolutional neural network to generate reconstruction features; segmenting the reconstruction feature into a plurality of reconstruction feature blocks, and inputting the reconstruction feature blocks added with the spatial position information into a Transform layer of a ViT network to obtain a global feature of the generated image; and classifying the global features of the generated image to obtain a traceability detection result of the generated image. According to the invention, traceability classification of the generated images can be realized.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention belongs to the technical field of machine learning, and particularly relates to a method and system for detecting the origin of generated images based on reconstruction. Background Art

[0002] Tracing the origin means finding and locating the source path of malicious behavior. In the field of detecting generated images, the goal of tracing the origin is to find and locate information such as the source path of the generated image and the identity of the initiator. Visual generation models, such as BigGAN and StableDiffusion, can generate highly realistic images based only on short text descriptions, which greatly simplifies the image creation process and improves the convenience of user use. However, the abuse of image generation models will accelerate the spread of harmful content in fake news. At the same time, the training and application of generation models may involve complex copyright issues. Therefore, detecting the origin of generated images has become a very important research field. Although there are various detectors for distinguishing real images from generated images, there is little research on tracing images back to their source models. Currently, there are the following three categories of methods for detecting the origin of generated images:

[0003] 1. Using watermark tracing: The origin of the generated content is traced by extracting the unique identifier (such as serial number, model ID, etc.) automatically embedded by the generation model during the process of generating images through specific algorithms or tools.

[0004] 2. Using fingerprint tracing: Since there are differences in the training data distribution, parameter settings, and training methods of different models, and there are also subtle differences in the image generation sampling process, the unique fingerprints of different models are extracted using algorithms or tools to achieve origin detection.

[0005] 3. Using reverse engineering method for tracing: The deployed local generation model is used to reprocess the image to be detected respectively, and the values before and after the processing are calculated and compared with the threshold of each model. If it is less than the threshold, it is determined that the image is generated by that model.

[0006] However, the above-mentioned solutions all have some deficiencies in terms of space consumption and access performance, which are specifically as follows:

[0007] 1. Using watermark tracing: When detecting images, the watermark of the model must be extracted, which increases the complexity of image reading. It is necessary to analyze the watermark embedding method of the generation model, and different methods are required to extract the watermark for different embedding methods. The operation calculation cost and time cost are relatively high when processing data in batches.

[0008] 2. Adopt fingerprint tracing: The effectiveness of fingerprint tracing depends on analyzing the subtle features in the generated content and extracting these distinguishing features, which usually include power spectral density, discrete cosine transform, autocorrelation, etc. However, if the generated content undergoes subsequent processing (such as compression, scaling, rotation, etc.), these features may be weakened or even lost. Therefore, how to design a fingerprint extraction method that can still maintain robustness under various image processing operations is the key challenge of fingerprint tracing technology.

[0009] 3. Reverse engineering method for tracing: This method usually relies on complex algorithms and models to analyze the internal structure of the generated image. This method requires multiple generation models to be deployed in advance, occupying a large amount of storage space. During the detection process, each model needs to process the image, resulting in a large amount of calculation. And it cannot trace the images generated by black-box models, only the locally deployed models.

[0010] In addition, most of the existing methods can only trace and detect the images generated by one type of base model, such as GAN or diffusion model, and cannot achieve the traceability of the images generated by GAN and diffusion models under black-box conditions. Summary of the Invention

[0011] The present invention provides a method and system for tracing and detecting generated images based on reconstruction. The method reconstructs the image and learns the feature distribution after reconstruction to achieve traceability classification of the generated image.

[0012] To achieve the above object, the technical solution of the present invention includes the following content.

[0013] A method for tracing and detecting generated images based on reconstruction, the method includes:

[0014] Reconstruct the generated image and obtain the numerical difference feature matrix before and after reconstruction;

[0015] Use a convolutional neural network to perform convolution and pooling operations on the numerical difference feature matrix to generate reconstruction features;

[0016] Slice the reconstruction features into multiple reconstruction feature blocks, and input the reconstruction feature blocks with added spatial position information into the Transformer layer of the ViT network to obtain the global features of the generated image;

[0017] Classify the global features of the generated image to obtain the tracing and detection result of the generated image.

[0018] Further, the reconstructing the generated image and obtaining the numerical difference feature matrix before and after reconstruction includes:

[0019] Reconstruct the generated image based on a variational autoencoder to obtain a reconstructed image;

[0020] Calculate the reconstruction distance value of each pixel in the generated image and the reconstructed image;

[0021] Generate a numerical difference feature matrix based on the reconstruction distance values of all pixels.

[0022] Further, the splitting the reconstruction features into multiple reconstruction feature blocks and inputting the reconstruction feature blocks with added spatial position information into the Transformer layer of the ViT network to obtain the global features of the generated image includes:

[0023] Split the reconstruction features into multiple reconstruction feature blocks of a fixed size;

[0024] Add the spatial position information of the reconstruction feature block to each feature block;

[0025] Combine the spatial position information with the reconstruction features through a position embedding operation to obtain the reconstruction feature blocks with added spatial position information;

[0026] Input the reconstruction feature blocks with added spatial position information into the Transformer layer of the ViT network to obtain the global features of the generated image.

[0027] Further, the method further includes:

[0028] Construct a training data set; wherein, each training data in the training data set includes: a generated image sample and the true label of the generated image sample;

[0029] Reconstruct the generated image sample and obtain the numerical difference feature matrix before and after the reconstruction of the generated image sample;

[0030] Perform convolution and pooling operations on the numerical difference feature matrix before and after the reconstruction of the generated image sample using a convolutional neural network to obtain the reconstruction features of the generated image sample;

[0031] Split the reconstruction features of the generated image sample into multiple reconstruction feature blocks and input the reconstruction feature blocks with added spatial position information into the Transformer layer of the ViT network to obtain the global features of the generated image sample;

[0032] Classify the global features of the generated image sample to obtain the prediction result of the generated image sample;

[0033] Use the cross-entropy loss function to calculate the error between the prediction result and the true label, and backpropagate the error to optimize the parameters of the convolutional neural network and the Transformer layer.

[0034] A reconstruction-based generated image traceability detection system, characterized in that the system includes:

[0035] A reconstruction module, configured to reconstruct the generated image and obtain a numerical difference feature matrix before and after reconstruction; perform convolution and pooling operations on the numerical difference feature matrix using a convolutional neural network to generate reconstruction features;

[0036] A discrimination module, configured to divide the reconstruction features into multiple reconstruction feature blocks, and input the reconstruction feature blocks with spatial position information into the Transformer layer of the ViT network to obtain the global features of the generated image; classify the global features of the generated image to obtain the traceability detection result of the generated image.

[0037] An electronic device, characterized in that the electronic device includes: a processor and a memory storing computer program instructions; when the processor executes the computer program instructions, the traceability detection method for generated images based on reconstruction described in any one of the above is implemented.

[0038] A computer-readable storage medium, characterized in that computer program instructions are stored on the computer-readable storage medium, and when the computer program instructions are executed by a processor, the traceability detection method for generated images based on reconstruction described in any one of the above is implemented.

[0039] A computer program product, characterized in that when the computer program product runs on a computer device, the computer device is caused to execute the traceability detection method for generated images based on reconstruction described in any one of the above.

[0040] Compared with the prior art, the present invention has at least the following beneficial effects.

[0041] 1) The present invention can effectively reduce the overhead of data processing and support the traceability of images generated by GAN and diffusion models under black box conditions.

[0042] 2) The present invention has the traceability ability for various generated images: the present invention can be used to trace images generated by different types of base generation models (including generative adversarial network GAN and diffusion model). By introducing a variational autoencoder (VAE) to reconstruct the input image and calculate the reconstruction error, the model can capture the feature differences left by different generation models during the image generation process.

[0043] 3) The present invention has the generalization ability for data imbalance: the model can maintain a good detection effect on the minority class generation models during the test phase. Even if some categories account for a small proportion in the test data, the model can still accurately identify the image source of the generation model. This ensures the practicality of the model system in real scenarios. BRIEF DESCRIPTION OF THE DRAWINGS

[0044] Figure 1 A flowchart of the traceability detection method for generated images based on reconstruction.

[0045] Figure 2 Schematic diagram of the reconstruction module.

[0046] Figure 3 Schematic diagram of the discrimination module. Specific implementation manners

[0047] The present invention will be further described in detail below with reference to the accompanying drawings. The examples given are only used to explain the present invention and are not intended to limit the scope of the present invention.

[0048] The main idea of the present invention is to design a reconstruction module and a discrimination module. Based on the reconstruction of the image by the reconstruction module, the numerical difference features and the reconstruction distribution features before and after reconstruction can be obtained. The discrimination module uses Vision Transformer to learn the fine-grained distribution features after reconstruction to achieve traceable classification of the generated images.

[0049] The method for traceability detection of generated images based on reconstruction of the present invention, as Figure 1 shown, includes the following steps 1 to 4.

[0050] Step 1: Reconstruct the generated image and obtain the numerical difference feature matrix before and after reconstruction.

[0051] The variational autoencoder VAE can encode the image into the latent space, and this latent space representation can capture the core texture and structure of the image. Therefore, VAE is used to reconstruct the image to amplify these differences to better distinguish images from different sources. To calculate the numerical difference before and after reconstruction, the present invention defines the reconstruction distance value of the pixel as: d i,j = x i,j - VAE(x i,j ), where x i,j is the pixel of the image, and VAE(x i,j ) is the reconstructed pixel value. Based on this definition of the reconstruction distance value of the pixel, the numerical difference feature matrix D rec (X) = X - VAE(X) can be obtained, where X = [x i,j .

[0052] Step 2: Use a convolutional neural network to perform convolution and pooling operations on the numerical difference feature matrix to generate reconstruction features.

[0053] To more intuitively display the distribution of the numerical differences, an initialized convolutional neural network (CNN) is used to process the above numerical difference feature matrix D rec(X) performs convolution and pooling operations to extract reconstructed features. Through layers of convolution and pooling operations, the CNN extracts a high-dimensional feature map representing the difference between the original image and the reconstructed features of the generative model, and these features can reflect the feature patterns of image generation.

[0054] Step 3: Split the reconstructed features into multiple reconstructed feature blocks, and input the reconstructed feature blocks with added spatial position information into the Transformer layer of the ViT network to obtain the global features of the generated image.

[0055] In one embodiment, the present invention splits the extracted reconstructed features into multiple blocks of a fixed size (such as a size of 8x8). Each feature block represents the features of a local area of the image. To enable the model to understand the position of each feature block in the image, position information is added to each feature block, and the spatial position information is combined with the reconstructed features through a Positional Embedding operation.

[0056] After that, the present invention inputs the reconstructed features with added spatial position information into the Transformer layer of the Vision Transformer (ViT). The Transformer learns the global feature relationships through the self-attention mechanism and can capture the subtle pattern differences generated by different generative models during image generation. After being processed by the multi-head self-attention mechanism and the feed-forward neural network, the output global features represent the overall characteristics of the generated image.

[0057] Specifically, the processing process of the Transformer layer can be summarized as: where Z l+1 is the output of the (l + 1)-th Transformer layer, Z l is the input of the (l + 1)-th Transformer layer, Q = Z l W Q K = Z l W K V = Z l W V where W Q W K W v are the weight matrices for generating queries, keys, and values respectively, is the dimension of the key vector.

[0058] Step 4: Classify the global features of the generated image to obtain the traceability detection result of the generated image.

[0059] The final output of the Transformer will pass through the MLP to calculate the probability distribution of the labels, and select the category with the highest probability as the attribution category of the generation model for the generated image, so as to realize image traceability.

[0060] In addition, when training the above convolutional neural network and Transformer layer, the present invention first constructs a training data set based on the image samples generated by each type of model and their true labels, then predicts the image samples based on the above method. Finally, the classification head maps the output features to the label probabilities of five types of generation models, and compares the output prediction probabilities with the true labels. The cross-entropy loss function is used to calculate the error between the prediction result and the true label, and the error is backpropagated to optimize the parameters of the CNN and Transformer layer, so that the model can better identify the images of different generation models.

[0061] To sum up, since both the diffusion model and GAN sample from the noise space to simulate the pixel distribution of real images. However, GAN relies on adversarial loss to train the generator, and directly maps the noise to the image through the adversarial loss. The diffusion model generates images through a series of denoising steps. Due to the differences in the training data distribution, parameter settings, and training methods of different models, there are also differences in the image generation sampling process. The generated images only simulate the real distribution, but are not exactly the same. Therefore, there are not only differences between the generated images and the real images, but also differences between the images generated by different models. And the present invention realizes the traceability of the generated images based on these unique features.

[0062] Based on the same concept, the present invention also provides a generation image traceability detection system based on reconstruction, and the system includes a Figure 2 reconstruction module as shown in Figure 3 and a discrimination module as shown in

[0063] The reconstruction module is used to reconstruct the generated image and obtain the numerical difference feature matrix before and after reconstruction; use the convolutional neural network to perform convolution and pooling operations on the numerical difference feature matrix to generate reconstruction features;

[0064] The discrimination module is used to divide the reconstruction features into multiple reconstruction feature blocks, and input the reconstruction feature blocks with spatial position information into the Transformer layer of the ViT network to obtain the global features of the generated image; classify the global features of the generated image to obtain the traceability detection result of the generated image.

[0065] Next, a specific experiment is used to illustrate the generation image traceability detection method based on reconstruction provided by the present invention. Among them, the table of this experiment is shown in Table 1.

[0066] Operating system Ubuntu 22.04.3 LTS Video memory 160G GPU NVIDIA GeForce RTX 4090

[0067] Table 1 Hardware Configuration

[0068] Experimental Design: To comprehensively evaluate the performance of the method, a dataset consisting of 11 subsets was used, including real images, images generated by diffusion models, and images generated by GANs. Seven of the generated image datasets were from the GenImage dataset, namely BigGAN, GLIDE, VQDM, SDV1.5, ADM, Midjourney, and Wukong. In addition, three advanced generated datasets were added: DALL-E 3, ProGAN, and StyleGAN. The real images were from the ImageNet dataset. All the data covered multiple categories such as animals, plants, vehicles, people, and daily items, aiming to be close to real-world application scenarios.

[0069] The VAE of the stuble-diffusion-v2 version was used as the reconstruction module. The feature extractor adopted a trainable CNN architecture. During the training and testing phases, the features were scaled to a size of 224×224. The Vision Transformer was used as the backbone network for classification, and the model was trained using the cross-entropy loss function. All 11 classes of data were input for training simultaneously, and only one model was trained for tracking.

[0070] The traceability experiment evaluated the performance using the overall accuracy, detection accuracy for each class, macro accuracy, recall rate, and F1 value. Considering that real images are much larger than generated images in real scenarios, to verify the effectiveness of the invention in real scenarios, a test data imbalance experiment was set up: the traceability detection model was normally trained, with 3000 images fixed for each category. To establish an imbalanced dataset, starting from a fixed number of 20,000 real images. Then, the number of generated images for each category was gradually increased from 100 to 1000, so that the ratio of a single type of generated image to real images increased from 1:200 to 1:20. During the entire experiment, the total number of generated images increased from 1,000 to 10,000, which means that the ratio of all generated images to real images increased from 1:20 to 1:2. The invention classified the images in these datasets, and the results of the traceability experiment are shown in Table 2. It can be seen that the invention achieved competitive results in tracing images generated by GANs and diffusion models, and only one classifier needed to be trained to achieve the model attribution of multi-class images.

[0071]

[0072]

[0073] Table 2 Statistical Results of the Traceability Experiment

[0074] Table 3 shows the statistics of the experimental results of data imbalance. It can be seen that even in the case of data imbalance, the invention still achieved excellent results. The macro-precision, macro-recall, and macro-F1 all exceeded 80%. The average accuracy and overall accuracy always remained at about 89%. These all verify the practicality of ReTD.

[0075]

[0076]

[0077] Statistics of the Experimental Results of Data Imbalance in Table 3

[0078] Although specific embodiments of the present invention are disclosed for illustrative purposes, which are intended to help understand the content of the present invention and implement it accordingly, those skilled in the art can understand that: without departing from the spirit and scope of the present invention and the appended claims, various substitutions, changes, and modifications are possible. Therefore, the present invention should not be limited to the content disclosed in the best embodiments, and the scope of protection required by the present invention shall be subject to the scope defined by the claims.

Claims

1. A reconstruction-based generated image source tracing detection method, characterized in that: The method comprises: Reconstruct the generated image and obtain the numerical difference feature matrix before and after the reconstruction; Use convolutional neural network to perform convolution and pooling operations on the numerical difference feature matrix to generate reconstruction features; The reconstructed features are divided into multiple reconstructed feature blocks, and the reconstructed feature blocks with spatial position information are input into the Transformer layer of the ViT network to obtain the global features of the generated image; The global features of the generated image are classified to obtain a source tracing detection result of the generated image.

2. The method according to claim 1, characterized in that The step of reconstructing the generated image and obtaining a numerical difference feature matrix before and after the reconstruction includes: Reconstruct the generated image based on the variational autoencoder to obtain a reconstructed image; Calculate the reconstruction distance value of each pixel in the generated image and the reconstructed image; Based on the reconstructed distance values ​​of all pixels, a numerical difference feature matrix is ​​generated.

3. The method according to claim 1, characterized in that The reconstructed features are divided into a plurality of reconstructed feature blocks, and the reconstructed feature blocks with spatial position information are input into the Transformer layer of the ViT network to obtain the global features of the generated image, including: Divide the reconstructed features into multiple reconstructed feature blocks of fixed size; Adding the spatial position information of the reconstructed feature block to each feature block; The spatial position information is combined with the reconstruction feature through the position embedding operation to obtain a reconstruction feature block with the spatial position information added; The reconstructed feature block with added spatial position information is input into the Transformer layer of the ViT network to obtain the global features of the generated image.

4. The method according to any one of claims 1 to 3, characterized in that: The method further comprises: Constructing a training data set; wherein each training data in the training data set includes: a generated image sample and a true label of the generated image sample; Reconstruct the generated image samples and obtain the numerical difference feature matrix of the generated image samples before and after reconstruction; The convolutional neural network is used to perform convolution and pooling operations on the numerical difference feature matrix of the generated image samples before and after reconstruction to obtain the reconstructed features of the generated image samples; The reconstructed features of the generated image samples are divided into multiple reconstructed feature blocks, and the reconstructed feature blocks with spatial position information are input into the Transformer layer of the ViT network to obtain the global features of the generated image samples; Classifying the global features of the generated image sample to obtain a prediction result of the generated image sample; The cross entropy loss function is used to calculate the error between the predicted result and the true label, and the error is back-propagated to optimize the parameters of the convolutional neural network and Transformer layer.

5. A reconstruction-based generated image tracing detection system, characterized in that: The system comprises: The reconstruction module is used to reconstruct the generated image and obtain the numerical difference feature matrix before and after the reconstruction; the convolutional neural network is used to perform convolution and pooling operations on the numerical difference feature matrix to generate reconstruction features; The discriminant module is used to divide the reconstructed features into multiple reconstructed feature blocks, and input the reconstructed feature blocks with spatial position information into the Transformer layer of the ViT network to obtain the global features of the generated image; classify the global features of the generated image to obtain the traceability detection result of the generated image.

6. An electronic device, characterized in that: The electronic device comprises: a processor and a memory storing computer program instructions; when the processor executes the computer program instructions, the generated image source tracing detection method based on reconstruction as described in any one of claims 1 to 4 is implemented.

7. A computer-readable storage medium, characterized in that: The computer-readable storage medium stores computer program instructions, and when the computer program instructions are executed by the processor, the reconstruction-based generated image tracing detection method according to any one of claims 1 to 4 is implemented.

8. A computer program product, characterized in that When the computer program product is run on a computer device, the computer device is enabled to execute the reconstruction-based generated image source tracing detection method as described in any one of claims 1 to 4.