Image authenticity detection method, system and equipment based on pre-training visual model

By using a pre-trained visual model-based method in image authenticity detection, combined with AI reconstruction and image block-level replacement combination, the problem of limited native features and network generalization capabilities of the prior art relying on pre-trained models is solved, and higher detection accuracy and generalization are achieved.

CN120070989AActive Publication Date: 2025-05-30ZHEJIANG UNIV
View PDF 5 Cites 0 Cited by

Patent Information

Application Number
CN202510148631.6
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-02-11
Publication Date
2025-05-30
Estimated Expiration
2045-02-11

AI Technical Summary

Technical Problem

The prior art has the limitations of relying on the native features of pre-trained models in image authenticity detection, and the problem that network generalization capabilities are limited by data sets.

Method used

The image authenticity detection method based on pre-trained visual model is adopted. By obtaining the training data set with labels, AI reconstruction and image block-level replacement combination, Transformer pre-trained model is loaded and linear layer is built, and iterative training is used to improve the generalization and robustness of the model.

Benefits of technology

It significantly improves the accuracy and generalization ability of image authenticity detection, and can increase the overall average accuracy rate to 95% while reaching 98% in the domain, which is 5% higher than the previous method.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120070989A_ABST
    Figure CN120070989A_ABST
Patent Text Reader

Abstract

The invention provides an image authenticity detection method, system and device based on a pre-training visual model. According to the method, feature extraction is performed on an input image by using a pre-trained visual model, then the image is input into a linear layer network for a binary classification task, and the authenticity probability of the image is output. In the training process, in addition to a real image and an AI generation image of a training set, AI reconstruction of the real image is added to serve as a stronger AI generation image; random mixing is performed on a real image and reconstruction at an image block level, and learning of the model on image authenticity features in locality is enhanced by performing comparative learning on image block level features during training; and meanwhile, a low-rank adaptation (Lora) fine tuning method is adopted instead of full-amount fine tuning, so that the training efficiency and effect of the model are improved. Through the above method, local image authenticity information is introduced to the pre-trained visual model, and the accuracy and generalization ability of image authenticity detection can be effectively improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the fields of computer vision and machine learning, and particularly to an image authenticity detection method, system and device based on a pre-trained vision model. Background Art

[0002] With the development of generative models such as Generative Adversarial Networks (GAN) and Diffusion Models, it has become increasingly easy to generate high-quality fake images. These generated images are almost indistinguishable from real images visually, posing a huge challenge to image authenticity detection. There are two existing methods for image authenticity detection. One is to classify through an MLP linear layer based on the features of a pre-trained model. The drawback of this type of method is that it relies on the native features of the pre-trained model, and the pre-trained model is generally an image model for instance recognition and segmentation, or a multimodal model for image-text pairs. The model itself does not have the concept of real / fake images. The other is to train a lightweight network from scratch for distinguishing real / generated images, such as a convolutional network. The drawback of this type of method is that the network is only trained on a real / generated detection dataset with a small amount of data, and its generalization ability is limited by the dataset.

[0003] A pre-trained vision model refers to a deep learning model pre-trained on a large-scale dataset. These models possess rich visual features and are usually used for downstream image segmentation tasks or image classification tasks. In the field of generated content detection, real images and AI-generated images have a certain degree of separability in the visual feature space of the pre-trained vision model because the pre-trained vision model is generally trained on a relatively large-scale real image dataset, and AI-generated images belong to out-of-domain data with a distribution inconsistent with that of real images. However, with the progress of image generation technology, such as the advent of new models like MidJourney and Stable DiffusionXL, the quality of image generation is also getting higher and higher, and there are certain limitations in directly using the visual features of the pre-trained vision model. Accordingly, numerous scholars have continuously proposed fine-tuning the pre-trained vision model to make it suitable for the generated detection task.

[0004] Pre-trained model fine-tuning refers to further training on the basis of a pre-trained model for a specific task and dataset to meet new task requirements. In the field of image authenticity detection, in order to make the pre-trained vision model more suitable for image authenticity features, some scholars have tried to add image authenticity information to the text and use contrastive learning to align image and text information. However, such methods attempt to directly learn the authenticity information of the whole image, and the feature granularity is not sufficient to identify the authenticity information of details, and there is still room for improvement. Summary of the Invention

[0005] The object of the present invention is to solve the above deficiencies in the prior art and provide an image authenticity detection method based on a pre-trained vision model.

[0006] The specific technical solutions adopted by the present invention are as follows:

[0007] In a first aspect, the present invention provides an image authenticity detection method based on a pre-trained vision model, which includes:

[0008] S1. Obtain a training data set with labels containing real images and AI-generated images, perform AI reconstruction on the real images in the training data set, add the obtained AI-reconstructed images to the training data set and label them as AI-generated images; at the same time, perform block-level replacement combination on the real images and the corresponding AI-reconstructed images, and add the generated combined images to the training data set and label them as AI-generated images;

[0009] S2. Load a pre-trained vision model based on Transformer as the feature extraction part and freeze the model weights, and construct a linear layer after the feature extraction part to perform a binary classification task of identifying whether the image is an AI-generated image; decompose each query matrix and value matrix in the pre-trained vision model into two low-rank matrices respectively, and thaw the parameters of the decomposed low-rank matrices as the learnable parameters in the pre-trained vision model;

[0010] S3. Use the training data set to iteratively train the image authenticity detection model composed of the pre-trained vision model and the linear layer. In each round of the iterative process, the input image first extracts image-level features and image-block-level features through the pre-trained vision model, then the linear layer outputs a class prediction based on the image-level features, and finally updates the learnable parameters in the pre-trained vision model and the linear layer by calculating the weighted sum of the image-level cross-entropy loss and the image-block-level contrast loss;

[0011] S4. For the image to be detected, input it into the trained image authenticity detection model, and the linear layer obtains the prediction result of whether the image to be detected belongs to an AI-generated image based on the image-level features extracted by the pre-trained vision model.

[0012] Preferably, the real images and AI-generated images originally included in the training data set are also processed by data augmentation to expand the sample size; the data augmentation processing includes one or more of random cropping, rotation, color jitter, and flipping.

[0013] Preferably, the AI reconstruction uses the Stable-Diffusion model.

[0014] Preferably, when generating the combined image, the real image and the corresponding AI reconstructed image are respectively divided into image patches of the same size, and then some image patches in one image are replaced with the image patches at the same position in the other image to generate the combined image.

[0015] Preferably, the pre-trained vision model adopts a CLIP model or a DINOv2 model.

[0016] Preferably, the rank of the low-rank matrix is 2 to 8, preferably 4.

[0017] Preferably, the form of the cross-entropy loss at the image level is:

[0018]

[0019] The form of the contrastive loss at the image patch level is:

[0020]

[0021] The form of the total loss function obtained by weighting the two losses is:

[0022] L = λ·L cross-entropy +(1 - λ)·L contrastive

[0023] where p i represents the probability that the i-th sample image output in the current training batch in the linear layer belongs to the AI-generated image; y i represents the label of the i-th sample image in the current training batch. The label value of 0 represents that the sample image is a real image, and the label value of 1 represents that the sample image is an AI-generated image; M is the size of the current training batch; D w (i) 2 represents the image patch-level feature distance value between the i-th pair of image patches, Y represents the source consistency flag value of the i-th pair of image patches. If the two image patches in the i-th pair of image patches both come from the real image or the AI-generated image, the value is 1, otherwise it is 0; λ represents the weight hyperparameter.

[0024] Preferably, the weight hyperparameter λ takes 0.7.

[0025] Preferably, the output dimension of the linear layer is 2, representing the probabilities that the input image belongs to the real image and the AI-generated image respectively. The two-dimensional probability distribution output by the linear layer needs to be normalized to the range of 0 to 1 through the Softmax layer.

[0026] Preferably, during the iterative training process, a batch training and dynamic learning rate adjustment strategy are adopted.

[0027] In a second aspect, the present invention provides an image authenticity detection system based on a pre-trained vision model, which includes:

[0028] A dataset generation module, configured to obtain a training dataset containing real images and AI-generated images with labels, perform AI reconstruction on the real images in the training dataset, add the obtained AI-reconstructed images to the training dataset and label them as AI-generated images; at the same time, perform block-level replacement combination on the real images and the corresponding AI-reconstructed images, and add the generated combined images to the training dataset and label them as AI-generated images;

[0029] A model setting module, configured to load a pre-trained vision model based on Transformer as a feature extraction part and freeze the model weights, and construct a linear layer after the feature extraction part to perform a binary classification task of identifying whether an image is an AI-generated image; decompose each query matrix and value matrix in the pre-trained vision model into two low-rank matrices respectively, and unfreeze the parameters of the decomposed low-rank matrices as learnable parameters in the pre-trained vision model;

[0030] A model training module, configured to iteratively train an image authenticity detection model composed of the pre-trained vision model and the linear layer using the training dataset. In each round of iteration, the input image first extracts image-level features and block-level features through the pre-trained vision model, then the linear layer outputs a class prediction based on the image-level features, and finally updates the learnable parameters in the pre-trained vision model and the linear layer by calculating the weighted sum of the image-level cross-entropy loss and the block-level contrast loss;

[0031] A detection module, configured to input the image to be detected into the trained image authenticity detection model, and obtain a prediction result of whether the image to be detected belongs to an AI-generated image based on the image-level features extracted by the linear layer from the pre-trained vision model.

[0032] In a third aspect, the present invention provides a computer electronic device, which includes a memory and a processor;

[0033] The memory is used to store a computer program;

[0034] The processor is configured to, when executing the computer program, implement the image authenticity detection method based on the pre-trained vision model as described in any one of the first aspects above.

[0035] The present invention has the following beneficial effects compared with the prior art:

[0036] The present invention comprehensively considers the accuracy and generalization ability of image authenticity detection, constructs a fine-tuning strategy for a pre-trained model based on contrastive learning and data augmentation, and plans the solution with the minimum cost for image authenticity detection. Compared with the prior art, considering the challenges such as the continuous improvement of the quality of AI-generated images and the inconsistent distribution of real images and AI-generated images, the present invention does not require additional parameters during inference. However, through data augmentation and contrastive learning of the present method, the generalization and robustness of the model are greatly improved. While the in-domain accuracy rate is 98%, the overall average accuracy rate is 95%, which is 5% higher than the prior method. BRIEF DESCRIPTION OF THE DRAWINGS

[0037] Figure 1 Schematic diagram of the steps of the method for detecting the authenticity of an image based on a pre-trained visual model;

[0038] Figure 2 Schematic diagram of the iterative training process;

[0039] Figure 3 Schematic diagram of the framework for detecting the authenticity of an image. DETAILED DESCRIPTION OF THE INVENTION

[0040] In order to make the above objects, features, and advantages of the present invention more obvious and understandable, the following detailed description of the specific embodiments of the present invention will be given with reference to the accompanying drawings. Many specific details are set forth in the following description in order to fully understand the present invention. However, the present invention can be implemented in many other ways different from those described herein, and those skilled in the art can make similar improvements without departing from the connotation of the present invention. Therefore, the present invention is not limited by the specific embodiments disclosed below. The technical features in various embodiments of the present invention can be combined correspondingly without conflict.

[0041] In the present invention, the so-called "image authenticity detection" refers to the process of using computer vision and machine learning technologies to analyze a given image to determine whether it is generated or modified by artificial intelligence technology. The image authenticity detection framework of the present invention uses a pre-trained vision model to extract features from the input image, obtaining a high-dimensional image feature vector; the extracted image feature vector is input into a linear layer network for a binary classification task, and the authenticity probability of the image is output. However, in addition to the real images and AI-generated images in the training set, the present invention adds AI reconstructions of real images as stronger AI-generated images; randomly mixes real images and reconstructions at the image patch level, and strengthens the model's learning of image authenticity features in terms of locality by performing contrastive learning on the patch-level features during training; at the same time, adopts the Low-Rank Adaptation (Lora) fine-tuning method instead of full-scale fine-tuning to improve the training efficiency and effect of the model. Through the above methods, the present invention introduces local image authenticity information into the pre-trained vision model, can effectively improve the accuracy and generalization ability of image authenticity detection, and is applicable to real-time detection of large-scale image data. The implementation method of the present invention will be specifically introduced below.

[0042] As Figure 1 shown, the present invention provides an image authenticity detection method based on a pre-trained vision model, which includes the following steps S1 to S4. The specific implementation of each step will be described in detail below.

[0043] S1. Obtain a training data set with labels containing real images and AI-generated images, perform AI reconstruction on the real images in the training data set, add the obtained AI-reconstructed images to the training data set and label them as AI-generated images; at the same time, perform patch-level replacement combination on the real images and the corresponding AI-reconstructed images, and add the generated combined images to the training data set and label them as AI-generated images.

[0044] In an embodiment of the present invention, the real images and AI-generated images originally included in the above training data set also need to be processed by data augmentation to expand the sample size to increase the generalization ability of the model. The data augmentation processing includes one or more of random cropping, rotation, color jitter, and flipping.

[0045] In an embodiment of the present invention, when performing AI reconstruction on the real images, the Stable-Diffusion model can be used to implement it.

[0046] In an embodiment of the present invention, the specific method for generating the combined image is: divide the real image and the corresponding AI-reconstructed image into image patches of the same size, and then replace some of the image patches in one image with the image patches at the same position in the other image to generate a combined image.

[0047] It should be noted that when performing image patch replacement, the specific replacement ratio can be adjusted according to the actual situation. In the embodiments of the present invention, 50% of the image patches in each real image need to be replaced with image patches in the AI-reconstructed image. It should be particularly noted that when performing image patch replacement, the replaced image patch and the replacing image patch need to be located at the same position in the image. The AI-reconstructed image used to replace the real image is output based on the real image by the Stable-Diffusion model. The process of this replacement combination can be achieved by randomly generating an image mask and then combining according to the mask.

[0048] S2. Load a pre-trained vision model based on Transformer as the feature extraction part and freeze the model weights, and construct a linear layer after the feature extraction part to perform a binary classification task of identifying whether the image is an AI-generated image; decompose each query matrix and value matrix in the pre-trained vision model into two low-rank matrices respectively, and unfreeze the decomposed low-rank matrices as learnable parameters in the pre-trained vision model.

[0049] In the embodiments of the present invention, the pre-trained vision model can be implemented using the CLIP model or the DINOv2 model. The CLIP model and the DINOv2 model belong to the prior art and will not be elaborated in detail here.

[0050] In this embodiment, all other learnable parameters in the pre-trained vision model except the query matrix and the value matrix are frozen and do not participate in the optimization during the subsequent iterative training process. Only the query matrix and the value matrix participate in the optimization of the training process. However, the query matrix and the value matrix are not directly optimized by learning, but need to be decomposed into two low-rank matrices and participate in the optimization in the form of these two low-rank matrices. Taking the query matrix Q as an example, it can be decomposed into matrix A and matrix B, that is, Q = A * B, and the ranks of matrix A and matrix B are lower than the rank of Q. In the embodiments of the present invention, the ranks of the low-rank matrices A and B can be set to 2 to 8, preferably 4. Assuming that the dimension of the query matrix Q is 1024 * 1024, it is preferably decomposed into a 1024 * 4 matrix A and a 4 * 1024 matrix B. Thus, unfreeze matrix A and matrix B and set them as learnable parameters, which can replace the original query matrix Q and participate in the learning optimization during the subsequent training process. Similarly, each value matrix V can also be decomposed into a 1024 * 4 learnable matrix A and a 4 * 1024 learnable matrix B and participate in the learning optimization during the subsequent training process.

[0051] In addition, in addition to the low-rank matrices obtained by decomposing the query matrix and the value matrix in the pre-trained vision model as learnable parameters, there are also learnable parameters in the linear layer constructed after the feature extraction part. Therefore, in the iterative training process of the present invention, it is necessary to optimize the low-rank matrices obtained by decomposing the query matrix and the value matrix in the pre-trained vision model, as well as the learnable parameters in the linear layer.

[0052] In an embodiment of the present invention, a linear layer is used to determine the authenticity of the input image, and this task is a binary classification task. Therefore, the output dimension of the linear layer is 2, which respectively represent the probabilities that the input image belongs to a real image and an AI-generated image. Specifically, the information at the first position in the output of the linear layer represents the probability that the image is real, and the information at the second position represents the probability that the image is AI-generated. In addition, the two-dimensional probability distribution output by the linear layer needs to be normalized to the range of 0 to 1 by the Softmax layer.

[0053] S3. Use the training data set obtained in the above S1 to perform iterative training on the image authenticity detection model composed of the pre-trained vision model and the linear layer. And in each round of iteration, the input image first extracts image-level features and image patch-level features through the pre-trained vision model, then the linear layer outputs a class prediction based on the image-level features, and finally updates the learnable parameters in the pre-trained vision model and the linear layer by calculating the weighted sum of the cross-entropy loss at the image level and the contrast loss at the image patch level.

[0054] In an embodiment of the present invention, the above cross-entropy loss at the image level has the following form:

[0055]

[0056] In an embodiment of the present invention, the above contrast loss at the image patch level needs to be calculated by pairwise sampling all the image patches in the current training batch. Each pair of image patches can come from the same image or different images, and the contrast loss has the following form:

[0057]

[0058] In an embodiment of the present invention, the total loss function obtained by weighting the above two losses has the following form:

[0059] L = λ·L cross-entropy +(1 - λ)·L contrastive

[0060] where p i represents the probability that the i-th sample image in the current training batch output by the linear layer belongs to an AI-generated image; y idenotes the label of the i-th sample image in the current training batch. A label value of 0 indicates that the sample image is a real image, and a label value of 1 indicates that the sample image is an AI-generated image; M is the size of the current training batch; D w (i) 2 denotes the image patch-level feature distance value between the i-th pair of image patches. Y denotes the source consistency flag value of the i-th pair of image patches. If both image patches in the i-th pair of image patches are from real images or AI-generated images, the value is 1; otherwise, it is 0. In the embodiments of the present invention, the image patch-level feature distance value can be calculated using the cosine distance. N is the number of image patch pairs sampled from the current training batch. In this embodiment, N is set to the maximum number of pairwise combinations of all image patches in the current training batch, that is, all image patch pairing combinations participate in the calculation of the contrastive loss. λ denotes the weight hyperparameter. This weight hyperparameter λ can be optimized and adjusted according to the actual situation. In the embodiments of the present invention, λ = 0.7 is taken.

[0061] The above total loss function is a linear combination of the cross-entropy loss and the contrastive loss. During the iterative optimization process, the optimizer can be used to update the learnable parameters in the pre-trained vision model and the linear layer based on the reverse gradient of the total loss function obtained in each round of iteration.

[0062] In the embodiments of the present invention, a batch training and dynamic learning rate adjustment strategy can be adopted during the iterative training process. At the same time, the maximum number of iterations G can be set. After each iteration, it is judged whether the maximum number of iterations G is reached. If the current iteration number is greater than the maximum number of iterations G, the iteration is stopped and the model parameters are saved to complete the model training; otherwise, the next iteration continues. The iterative process can be seen in Figure 2 shown.

[0063] S4. For the image to be detected, input it into the pre-trained vision model and the linear layer that have been trained. The pre-trained vision model extracts the image-level features and image patch-level features of the image to be detected, and inputs the image-level features into the binary classification linear layer to obtain the prediction result of whether the image to be detected belongs to an AI-generated image. This detection process is as Figure 3 shown.

[0064] The present invention is a general-purpose method for detecting image authenticity. In the framework of this method, the pre-trained vision model it is based on can theoretically be any vision model pre-trained on a large-scale image dataset. Through the confusion of real and fake images at the image patch level and the contrastive learning of image patch-level features, the visual features of the pre-trained vision model can be transferred to the image authenticity feature task, significantly improving the generalization and accuracy of the pre-trained vision model in the field of image authenticity detection. Moreover, the practical value of this method will increase with the enhancement of the pre-trained vision model. The present invention can be effectively applied to the image authenticity detection task and allows the appearance of generated images that did not appear during training during testing, having a wide range of application prospects.

[0065] It should be noted that the method steps of S1 to S4 above can essentially be implemented in the form of a computer program.

[0066] Therefore, based on the same inventive concept, the present invention also provides an image authenticity detection system based on a pre-trained vision model corresponding to the image authenticity detection method based on a pre-trained vision model provided in the above embodiment. This system includes:

[0067] A dataset generation module, which is used to obtain a training dataset containing real images and AI-generated images with labels, perform AI reconstruction on the real images in the training dataset, add the obtained AI-reconstructed images to the training dataset and label them as AI-generated images; at the same time, perform image patch-level replacement combination on the real images and the corresponding AI-reconstructed images, and add the generated combined images to the training dataset and label them as AI-generated images;

[0068] A model setting module, which is used to load a pre-trained vision model based on Transformer as the feature extraction part and freeze the model weights, and construct a linear layer after the feature extraction part to perform a binary classification task of identifying whether an image is an AI-generated image; decompose each query matrix and value matrix in the pre-trained vision model into two low-rank matrices respectively, and unfreeze the parameters of the decomposed low-rank matrices as the learnable parameters in the pre-trained vision model;

[0069] A model training module, which is used to iteratively train the image authenticity detection model composed of the pre-trained vision model and the linear layer using the training dataset. In each round of iteration, the input image first extracts image-level features and image patch-level features through the pre-trained vision model, then the linear layer outputs a class prediction based on the image-level features, and finally updates the learnable parameters in the pre-trained vision model and the linear layer by calculating the weighted sum of the image-level cross-entropy loss and the image patch-level contrastive loss;

[0070] The detection module is configured to input the image to be detected into a trained image authenticity detection model for the image to be detected, and the linear layer obtains a prediction result on whether the image to be detected belongs to an AI-generated image based on the image-level features extracted by the pre-trained vision model.

[0071] In the image authenticity detection system based on the pre-trained vision model in the above embodiment, the specific processes executed by each module can also refer to the specific steps of S1 to S4 described above, which will not be elaborated here.

[0072] Similarly, based on the same inventive concept, the present invention also provides a computer electronic device corresponding to the image authenticity detection method based on the pre-trained vision model provided in the above embodiment, which includes a memory and a processor;

[0073] The memory is used to store a computer program;

[0074] The processor is configured to implement the image authenticity detection method based on the pre-trained vision model as described above when executing the computer program;

[0075] In addition, when the logical instructions in the above memory are implemented in the form of software functional units and sold or used as independent products, they can be stored in a computer-readable storage medium. Based on such an understanding, the technical solution of the present invention, in essence, or the part that contributes to the prior art, or a part of this technical solution, can be embodied in the form of a software product. The computer software product is stored in a storage medium and includes several instructions for causing a computer device (which can be a personal computer, a server, or a network device, etc.) to execute all or part of the steps of the methods described in the various embodiments of the present invention.

[0076] It can be understood that the above storage medium may include a random access memory (RAM), and may also include a non-volatile memory (NVM), such as at least one disk memory. At the same time, the storage medium may also be various media such as a USB flash drive, a mobile hard disk, a magnetic disk, or an optical disc that can store program codes.

[0077] It can be understood that the above-mentioned processor can be a general-purpose processor, including a Central Processing Unit (CPU), a Network Processor (NP), etc.; it can also be a Digital Signal Processor (DSP), an Application Specific Integrated Circuit (ASIC), a Field-Programmable Gate Array (FPGA), or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components.

[0078] In addition, it should be noted that those skilled in the art can clearly understand that for the convenience and simplicity of description, the specific working process of the above-described system can refer to the corresponding process in the foregoing method embodiments, and will not be elaborated herein. In each of the embodiments provided in the present application, the division of steps or modules in the system and method is only a logical function division, and there may be other division methods in actual implementation. For example, multiple modules or steps can be combined or integrated together, and a module or step can also be split.

[0079] Next, the image authenticity detection method based on the pre-trained visual model described in S1 to S4 above will be tested on a specific data set to demonstrate its technical effects.

[0080] Embodiment

[0081] In this embodiment, two pre-trained vision models, the CLIP model or the DINOv2 model, are adopted. At the same time, to prove the generality of this embodiment, there are two cases: reconstructing using Stable-Diffusion v2 (sdv2) and reconstructing using Stable-Diffusion v1 (sdv1). Therefore, a total of four experimental groups, namely DINOv2+sdv2, CLIP+sdv2, DINOv2+sdv1, and CLIP+sdv1, are designed in this embodiment. They are respectively tested on the GenImage (Zhu M, Chen H, Yan Q, et al. Genimage: A million-scale benchmark for detecting ai-generated image [J]. Advances in Neural Information Processing Systems, 2024, 36.) dataset, trained on the Stable-Diffusion 1.4 dataset according to the testing method in the article, and the accuracy on other methods such as Midjourney and Big GAN is tested. Finally, the results of the four experimental groups are compared as shown in Table 1:

[0082] Table 1 Experimental data of algorithm comparison

[0083] Name of the generation method sd14 sd15 mj adm wukong glide vqdm gan avg DINOv2 + sdv2 97.94 97.94 95.97 92.06 98.02 95.07 95.47 86.70 94.89 CLIP + sdv2 98.36 98.03 97.97 94.05 98.19 94.12 97.93 92.51 96.39 DINOv2 + sdv1 98.97 98.71 89.27 91.97 98.80 94.17 97.82 92.73 95.30 CLIP + sdv1 98.77 98.51 91.19 90.79 98.72 89.03 98.75 95.30 95.13

[0084] The above results show that the present invention can strengthen the model's recognition of the information features added to the local part of the image by increasing the enhanced data that confuses real and generated images at the image patch granularity and performing contrastive learning at the image patch granularity. The present invention can transfer the visual features of the pre-trained vision model to the image authenticity feature task, significantly improving the generalization and robustness of the pre-trained vision model in various generation methods.

[0085] The above-described embodiments are only some preferred implementation solutions of the present invention, but are not intended to limit the present invention. Those of ordinary skill in the relevant technical field can make various changes and modifications without departing from the spirit and scope of the present invention. Therefore, all technical solutions obtained by adopting equivalent substitution or equivalent transformation fall within the protection scope of the present invention.

Claims

1. A method for detecting image authenticity based on a pre-trained visual model, characterized in that: include: S1. Obtain a labeled training data set containing real images and AI-generated images, perform AI reconstruction on the real images in the training data set, add the obtained AI-reconstructed images to the training data set and annotate them as AI-generated images; at the same time, replace and combine the real images and the corresponding AI-reconstructed images at the image block level, add the generated combined images to the training data set and annotate them as AI-generated images; S2. Load the Transformer-based pre-trained visual model as the feature extraction part and freeze the model weights, and build a linear layer after the feature extraction part to perform the binary classification task of identifying whether the image is an AI-generated image; decompose each query matrix and value matrix in the pre-trained visual model into two low-rank matrices, and unfreeze the parameters of the decomposed low-rank matrices as learnable parameters in the pre-trained visual model; S3, using the training data set to iteratively train the image authenticity detection model composed of the pre-trained visual model and the linear layer, in each round of iteration, the input image is firstly extracted with the pre-trained visual model to extract image-level features and image block-level features, and then the linear layer outputs the category prediction based on the image-level features, and finally the learnable parameters in the pre-trained visual model and the linear layer are updated by calculating the weighted sum of the image-level cross entropy loss and the image block-level contrast loss; S4. For the image to be detected, it is input into the trained image authenticity detection model, and the linear layer extracts the image-level features based on the pre-trained visual model to obtain the prediction result of whether the image to be detected is an AI-generated image.

2. The image authenticity detection method based on the pre-trained visual model according to claim 1, characterized in that: The real images and AI-generated images originally contained in the training data set are further processed by data enhancement to expand the sample size; the data enhancement processing includes one or more of random cropping, rotation, color jittering, and flipping.

3. The image authenticity detection method based on the pre-trained visual model according to claim 1, characterized in that: The AI ​​reconstruction adopts the Stable-Diffusion model.

4. The image authenticity detection method based on the pre-trained visual model according to claim 1, characterized in that: When generating the combined image, the real image and the corresponding AI reconstructed image are divided into image blocks of the same size, and then part of the image blocks in one image are replaced with image blocks at the same position in another image to generate a combined image.

5. The image authenticity detection method based on the pre-trained visual model according to claim 1, characterized in that: The pre-trained visual model adopts a CLIP model or a DINOv2 model.

6. The image authenticity detection method based on a pre-trained visual model according to claim 1, characterized in that: The rank of the low-rank matrix is ​​2-8, preferably 4.

7. The image authenticity detection method based on the pre-trained visual model according to claim 1, characterized in that: The image-level cross entropy loss is in the form of: The contrast loss at the image block level is in the form of: The total loss function obtained by weighting the two losses is: L=λ·L cross-entropy +(1-λ)·L contrastive Among them, p i represents the probability that the i-th sample image in the current training batch outputted by the linear layer belongs to the AI-generated image; y i Indicates the label of the i-th sample image in the current training batch. The label value of 0 represents the sample image is a real image, and the label value of 1 represents the sample image is an AI-generated image; M is the size of the current training batch; D w (i) 2 represents the image block level feature distance value between the i-th pair of image blocks, Y represents the source consistency flag value of the i-th pair of image blocks, if the two image blocks in the i-th pair of image blocks are derived from real images or AI generated images, the value is 1, otherwise it is 0; λ represents the weight hyperparameter.

8. The image authenticity detection method based on a pre-trained visual model according to claim 1, characterized in that: The iterative training process adopts batch training and dynamic learning rate adjustment strategy.

9. An image authenticity detection system based on a pre-trained visual model, characterized in that: include: The data set generation module is used to obtain a labeled training data set containing real images and AI-generated images, and perform AI reconstruction on the real images in the training data set, add the obtained AI-reconstructed images to the training data set and annotate them as AI-generated images; at the same time, replace and combine the real images with the corresponding AI-reconstructed images at the image block level, add the generated combined images to the training data set and annotate them as AI-generated images; The model setup module is used to load the Transformer-based pre-trained visual model as the feature extraction part and freeze the model weights, and build a linear layer after the feature extraction part to perform the binary classification task of identifying whether the image is an AI-generated image; each query matrix and value matrix in the pre-trained visual model is decomposed into two low-rank matrices, and the parameters of the decomposed low-rank matrices are unfrozen as learnable parameters in the pre-trained visual model; A model training module, used to iteratively train an image authenticity detection model composed of a pre-trained visual model and a linear layer using the training data set, wherein in each round of iteration, the input image is firstly subjected to image-level features and image block-level features extracted by the pre-trained visual model, and then the linear layer outputs a category prediction based on the image-level features, and finally the learnable parameters in the pre-trained visual model and the linear layer are updated by calculating the weighted sum of the image-level cross entropy loss and the image block-level contrast loss; The detection module is used to input the image to be detected into a trained image authenticity detection model, and the linear layer extracts the image-level features based on the pre-trained visual model to obtain a prediction result on whether the image to be detected is an AI-generated image.

10. A computer electronic device, characterized in that: including memory and processor; The memory is used to store computer programs; The processor is used to implement the image authenticity detection method based on the pre-trained visual model as described in any one of claims 1 to 8 when executing the computer program.

Citation Information

Patent Citations

  • Double-flow network image forgery detection method and system based on image block feature extraction

    CN113361474A

  • Video sensitive information detection method and system based on pre-training strategy

    CN117315537A

  • Defect detection meta-model construction method, defect detection method, equipment and medium

    CN117495786A

  • Method and data processing system for lossy encoding, transmission and decoding of images or videos

    CN117980914A

  • Freezing ViT feature fusion network-based AI generated image detection method

    CN118570599A