Generated image detection method, system and device and storage medium

By extracting significant conceptual terms and combining them with context optimization techniques, the performance of generated image detection is improved, addressing the shortcomings of existing methods in detecting unseen data distributions and complex scenes, and achieving efficient and robust generated image detection.

CN120852888AActive Publication Date: 2025-10-28UNIV OF SCI & TECH OF CHINA
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202511351370.4
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-09-22
Publication Date
2025-10-28
Estimated Expiration
2045-09-22

AI Technical Summary

Technical Problem

Existing methods for generating image detection have insufficient generalization ability when faced with unseen data distributions, and their reliance on visual features results in suboptimal detection performance, making them unable to effectively detect generated images.

Method used

By decoding the detection features of training images, salient concept words are extracted and learnable vectors are constructed. Image classification is performed by combining context optimization techniques and multimodal large models. Salient and non-salient concept words are used as cues to optimize the learnable vectors to improve detection performance.

Benefits of technology

It achieves efficient detection of unseen data distributions and complex scenes, is highly robust, can successfully detect images that have undergone real-world post-processing, has high generalization ability, and its average discrimination accuracy can reach over 95%.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120852888A_ABST
    Figure CN120852888A_ABST
Patent Text Reader

Abstract

The invention discloses a generated image detection method, system and device and a storage medium, which are corresponding schemes, in the scheme, cross-modal alignment of vision and text features is enhanced by extracting significant concept words and converting the significant concept words into learnable text prompts, multi-concept semantic information can be fused through subsequent fine adjustment, and the visual sense and text feature fusion efficiency is improved. The concept-level features of the generated image are captured, so that the detection performance of unseen data distribution and complex scenes is improved; in addition, training resources consumed during training are few, and plug-and-play is achieved; images subjected to common post-processing operation in the real world can be successfully detected, and the robustness is high; meanwhile, generated images with various different data distributions can be detected at the same time, generalization is high, practicability is high, and the average identification precision can reach 95% or above.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of generated image detection technology, and in particular to a generated image detection method, system, device and storage medium. Background Art

[0002] Generative AI images refer to images artificially synthesized using artificial intelligence technology. With the continuous development of generation technology, today's generated images are characterized by controllable content, realistic visual effects, and low generation barriers. However, the malicious misuse of generation technology could cause a series of serious social security problems. Therefore, efficient detection technologies are urgently needed.

[0003] Traditional image detection methods are mainly divided into those that train models using spatial or frequency domain features. Spatial domain-based methods detect generated artifacts, such as texture irregularities and edge distortions, at the pixel level. Frequency domain-based methods transform image data from the spatial domain to the spectral domain, examining the differences between generated and real images in the spectral domain to effectively distinguish between them. However, traditional methods suffer from poor generalization ability when faced with data distributions not seen during training. With the emergence of multimodal visual-language models, recent research has begun to leverage the rich prior knowledge gained from training on large-scale image-text datasets to detect generated images by aligning visual and textual features. Recent research uses the CLIP (Contrastive Language-Image Pre-trained) model as a feature extraction model and achieves excellent detection performance by training a linear classification head on it. However, this method does not utilize the CLIP text component and relies solely on visual features, which may lead to suboptimal performance, and the generated image detection performance needs further improvement.

[0004] In view of this, the present invention is hereby proposed. Summary of the Invention

[0005] The purpose of this invention is to provide a method, system, device, and storage medium for generating images, which can capture conceptual features of generated images, thereby improving the detection performance for unseen data distributions and complex scenes.

[0006] The objective of this invention is achieved through the following technical solution: A method for generating image detection, comprising: The detection features of each training image are decoded, and the significant concept word set is extracted by word frequency analysis and statistical extraction of the decoded text. Context optimization technology is introduced, which initializes context labels using salient concept words and related non-salient concept words, and constructs learnable vectors. Then, given context words are used as prompts. The prompts and training images are input into a pre-trained multimodal large model to perform image classification tasks, obtain prediction probabilities, and optimize the learnable vectors by calculating a loss function based on the prediction probabilities. The optimized learnable vectors are used to construct cueing, which is then input along with the image to be detected into a pre-trained multimodal large model to obtain the detection results.

[0007] A generative image detection system for implementing the aforementioned method includes: The salient concept extraction unit is used to decode the detection features of each training image and extract a salient concept word set by performing word frequency analysis and statistical extraction on the decoded text. The prompt fine-tuning unit is used to introduce context optimization technology. It initializes context labels using salient concept words and related non-salient concept words, and constructs learnable vectors. Then, it combines the given context words as prompts and inputs the prompts and training images into a pre-trained multimodal large model to perform image classification tasks, obtains prediction probabilities, and calculates a loss function based on the prediction probabilities to optimize the learnable vectors. The detection unit is used to construct cues using optimized learnable vectors and input them along with the image to be detected into a pre-trained multimodal large model to obtain the detection results.

[0008] A processing device includes: one or more processors; and a memory for storing one or more programs; When the one or more programs are executed by the one or more processors, the one or more processors implement the aforementioned method.

[0009] A readable storage medium storing a computer program that, when executed by a processor, implements the aforementioned method.

[0010] As can be seen from the technical solution provided by the present invention, by extracting significant concept words and converting them into learnable text prompts, the cross-modal alignment of visual features (image features) and text features is enhanced. Through subsequent fine-tuning, multi-concept semantic information can be fused to achieve the capture of concept-level features of generated images, thereby improving the detection performance for unseen data distributions and complex scenes. In addition, the present invention consumes less training resources during training and is plug-and-play. It can also successfully detect images that have undergone common post-processing operations in the real world, demonstrating strong robustness. At the same time, it can detect generated images with multiple different data distributions simultaneously, exhibiting high generalization and practicality, with an average discrimination accuracy of over 95%. Attached Figure Description

[0011] To more clearly illustrate the technical solutions of the embodiments of the present invention, the drawings used in the following description of the embodiments will be briefly introduced. Obviously, the drawings described below are only some embodiments of the present invention. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.

[0012] Figure 1 This is a flowchart of an image detection method provided in an embodiment of the present invention.

[0013] Figure 2 This is a schematic diagram of the overall architecture of an image detection method provided in an embodiment of the present invention.

[0014] Figure 3 This is a schematic diagram of a generated image detection system provided in an embodiment of the present invention.

[0015] Figure 4 This is a schematic diagram of a processing device provided in an embodiment of the present invention. Detailed Implementation

[0016] The technical solutions of the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of the present invention, and not all embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those skilled in the art without creative effort are within the protection scope of the present invention.

[0017] First, the following explanations are provided for the terms that may be used in this article: The terms "comprising," "including," "containing," "having," or other similar semantic descriptions should be interpreted as non-exclusive inclusion. For example, including a technical feature element (such as raw material, component, ingredient, carrier, dosage form, material, size, part, component, mechanism, device, step, process, method, reaction conditions, processing conditions, parameter, algorithm, signal, data, product or article of manufacture, etc.) should be interpreted as including not only the expressly listed technical feature element, but also other technical feature elements that are not expressly listed and are well-known in the art.

[0018] The term "composed of" excludes any technical features not expressly listed. When used in a claim, it closes the claim to exclude all technical features other than those expressly listed, except for associated conventional impurities. If the term appears only in a clause of a claim, it limits the claim to the elements expressly listed in that clause; elements recited in other clauses are not excluded from the overall claim.

[0019] Unless otherwise explicitly specified or limited, the terms "installation," "connection," "linking," and "fixing," etc., should be interpreted broadly. For example, they can refer to fixed connections, detachable connections, or integral connections; they can refer to mechanical connections or electrical connections; they can refer to direct connections or indirect connections through an intermediate medium; and they can refer to the internal connection between two components. Those skilled in the art can understand the specific meaning of the above terms in this document according to the specific circumstances.

[0020] The following provides a detailed description of the image generation detection method, system, device, and storage medium provided by this invention. Contents not described in detail in the embodiments of this invention are prior art known to those skilled in the art. Where specific conditions are not specified in the embodiments of this invention, they are performed according to conventional conditions in the art or conditions recommended by the manufacturer. Where the manufacturers of the reagents or instruments used in the embodiments of this invention are not specified, they are all conventional products that can be purchased commercially.

[0021] Example 1 This invention provides a method for generating image detection, such as... Figure 1 As shown, it mainly includes the following steps: Step 1: Extraction of salient concepts.

[0022] In this embodiment of the invention, the detection features of each training image are decoded, and significant concept words are extracted by word frequency analysis and statistical extraction of the decoded text. Specifically: for each training image in the real image training set and the generated image training set, image features are extracted and linearly weighted to obtain detection features. Based on the image description generation model, the corresponding text representation is decoded to form the text sets corresponding to the real image training set and the generated image training set, respectively. Then, word frequency analysis and statistical extraction are performed to obtain the significant concept word sets corresponding to the real image training set and the generated image training set, respectively.

[0023] In this embodiment of the invention, the step of performing word frequency analysis and statistical extraction to obtain the salient concept word sets corresponding to the real image training set and the generated image training set respectively includes: preprocessing the text sets corresponding to the real image training set and the text sets corresponding to the generated image training set respectively; using natural language processing technology to segment the two preprocessed text sets into words, and counting the frequency of each word, selecting the top few words with the highest frequency in the two preprocessed text sets as the salient concept word sets corresponding to the real image training set and the generated image training set respectively.

[0024] Step 2: Prompt for fine-tuning.

[0025] In this embodiment of the invention, a context optimization technique is introduced. Context labels are initialized using salient concept words and related non-salient concept words, and a learnable vector is constructed. Then, given context words are used as prompts. The prompts and training images are input into a pre-trained multimodal large model for image classification to obtain prediction probabilities. The loss function is calculated based on the prediction probabilities to optimize the learnable vector.

[0026] In this embodiment of the invention, context labels are initialized using salient concept words and related non-salient concept words, and learnable vectors are constructed, which are then combined with given context words as cues: ; ; Where t represents a prompt; A learnable vector consisting of M context labels. It is the m-th context tag, which is a vector with the same dimension as the word embedding. M is a hyperparameter specifying the number of context tags. The category label, used as contextual words, includes two categories: real and fake. Real means true, and fake means false. For text encoders in pre-trained multimodal large models; For a single salient concept word in the salient concept word set corresponding to the real image training set, To generate a single salient concept word from the salient concept word set corresponding to the image training set; or For a single non-salient concept word in the real image training set, To generate a single non-salient concept word from the image training set; a non-salient concept word is a word in the decoded text excluding salient concept words.

[0027] For example, the prompt could be: [Significant concept word] a photo of [real / fake]. In this example, "a photo of" is a non-concept word related to the significant concept word. Of course, in actual work, "a photo of" here is the corresponding vector form.

[0028] In this embodiment of the invention, the step of inputting the prompts and training images together into a pre-trained multimodal large model for image classification to obtain the predicted probabilities includes: The prompts are input into the text encoder of a pre-trained multimodal large model. In this process, training images are input into the image encoder of a pre-trained multimodal large model, combined with a text encoder. The predicted probability is calculated from the output of the image encoder and expressed as: ; in, This represents the training images calculated based on prompts and fine-tuning. Category The probability is given by 0 and 1, representing true and false respectively; K represents the number of categories (K=2), where k and i are both category labels. and The corresponding representation uses the categories corresponding to k and i as context words to construct hints; For temperature parameters, For image encoders from training images Image features extracted from the image; cos represents cosine similarity.

[0029] Preferably, to further enhance the pre-trained multimodal large model, the present invention also provides a prompt integration scheme, which involves converting multiple significant concept words into corresponding word vectors and injecting them into the learnable vector of each category to obtain multiple prompts corresponding to each category. The category label corresponding to each category is the corresponding context word. When the multimodal large model performs an image classification task, it integrates the text features corresponding to all prompts under each category and takes the average as the final text feature. Combined with the corresponding image features, the prediction probability is calculated.

[0030] Specifically, the process of integrating the text features corresponding to all prompts under each category and averaging them to obtain the final text features, and combining them with the corresponding image features to calculate the predicted probability, includes: for each prompt under each category, processing the text encoder in a pre-trained multimodal large model. Extract the corresponding text features; for each category, average and fuse all the text features corresponding to the prompts to obtain the final text features; using the final text features and the image features extracted by the image encoder of the pre-trained multimodal large model, calculate the prediction probability, expressed as: ; in, This represents the training images computed based on prompts. Category The probability is given by 0 and 1, representing true and false respectively; K represents the number of categories, where k and i are both category labels. and The corresponding representations are the final text features for categories k and i; For temperature parameters, For training images Image features; cos represents cosine similarity.

[0031] Step 3: Image detection.

[0032] The optimized learnable vectors are used to construct cue, which is then input into a pre-trained multimodal large model along with the image to be detected to obtain the detection result (i.e. whether it is a generated image).

[0033] In this embodiment of the invention, when constructing prompts using optimized learnable vectors, the prompts do not contain classification labels.

[0034] The above-mentioned solution provided by the embodiments of the present invention enhances the cross-modal alignment of visual and textual features in a multimodal large model by extracting significant concept words and converting them into learnable text prompts. With the help of prompts, it can fine-tune the fusion of multi-concept semantic information and achieve the capture of concept-level features of generated images, thereby improving the detection performance for unseen data distributions and complex scenes.

[0035] The solution provided in this invention can be used to detect generative artificial intelligence images. In implementation, it can be deployed on the backend servers of various video websites and online forums to effectively detect, label, and intercept generated images before publication, eliminating potential negative impacts. It can also be applied to security regulatory departments to analyze and collect evidence of suspicious images disseminated on online platforms, maintaining network security.

[0036] To more clearly demonstrate the technical solution and its effects provided by the present invention, the method provided by the embodiments of the present invention will be described in detail below with reference to specific examples.

[0037] like Figure 2 As shown, the overall architecture of the method is illustrated, which mainly includes three parts: salient concept extraction, hint fine-tuning, and hint integration. Each part will be described in detail below.

[0038] 1. Salient concept extraction.

[0039] like Figure 2 As shown above, the decoded text from the detection features of the training set is subjected to word frequency analysis and statistical extraction to obtain a significant concept word set.

[0040] Specifically: For each image in a given training set containing both real and generated images, image encoders are used to obtain their respective image features. These features are then linearly weighted using the parameter weights and biases of a linear classifier trained in recent research to obtain corresponding detection features. An image description generation model is then used to obtain the text representation corresponding to these features, resulting in a real-set detection feature text set (the text set corresponding to the real image training set) and a generated-set detection feature text set (the text set corresponding to the generated image training set). Preprocessing operations such as lowercase conversion, punctuation removal, and stop word removal are performed on each text set. Natural language processing techniques are then used to segment the preprocessed text, and the frequency of each word is counted. Based on the word frequency statistics, the top few most frequent words are extracted; these high-frequency words are considered salient concepts in their respective text sets.

[0041] To avoid overlap of salient concept words between the real and generated image training sets, and to further distinguish the salient concepts of the detection features in the real and generated sets, when selecting salient concept words for each set, words with a frequency difference of more than S1 (e.g., S1=2) times and still within the top S2 (e.g., S2=25) of the corresponding set's salient concepts are considered as the final salient concept words. The frequency difference is calculated as follows: the ratio of the generated image training set frequency to the real image training set frequency is ≥2 (for extracting salient concept words from the generated image training set), or the ratio of the real image training set frequency to the generated image training set frequency is ≥2 (for extracting salient concept words from the real image training set). For example, if "artificial" appears 100 times in the generated image training set and 40 times in the real image training set, the calculation is "100 / 40=2.5≥2", which meets the standard; therefore, it is extracted as a salient concept word corresponding to the generated image training set.

[0042] For example: the image encoder can use the pre-trained model ViT-L / 14, where ViT (VisionTransformer) is the visual transformer, L stands for Large, which means a larger model with more parameters and stronger expressive power among visual transformers of the same type, and 14 represents that the input image is segmented into 14x14 pixel image blocks; the image description generation model can be the open-source ClipCap model (a lightweight image description generation model based on CLIP).

[0043] also, Figure 2 The top section only provides examples of prominent concept words in English, such as person, woman, cat, table, and room. The specific prominent concept words and their language forms are determined by the content of the training set, and this invention does not impose any restrictions.

[0044] 2. Suggestion for fine-tuning.

[0045] In this embodiment of the invention, a pre-trained multimodal large model is fine-tuned based on prompts to make it suitable for image generation detection tasks.

[0046] Specifically, this invention introduces Context Optimization (CoOp) to improve the efficiency of pre-trained multimodal large models in image classification tasks. For example, the pre-trained multimodal large model can be a CLIP model. Unlike traditional cue templates, CoOp models context words using continuous vectors learned end-to-end from the data, while freezing the pre-trained parameters (…). Figure 2 CoOp uses snowflake symbols for labeling, thus avoiding manual fine-tuning of prompts. It appends learnable vectors to the context words of the prompts; these context words refer to the category labels of the given dataset, such as... Figure 2 As shown in the bottom right corner, the context words here are "real" and "fake." These learnable vectors can be initialized using pre-trained word embeddings. Specifically, for the text encoder... The prompt is designed as follows: ; ; in, It is the m-th context tag, which is a vector with the same dimension as the word embedding. M is a hyperparameter specifying the number of context tags. The category label, as a context word, can be placed at the end, the beginning, or the middle. The formula above provides an example of placing it at the end. The category label includes two categories: real and fake. Real means true and fake means false. This is a text encoder in a pre-trained multimodal large model.

[0047] In this embodiment of the invention, a prompt t is input into the text encoder. This yields a classification weight vector representing the visual concept. Combined with the image features extracted by the image encoder, the prediction probability is calculated in the following manner: ;

[0048] in, This represents the training images calculated based on prompts and fine-tuning. Category The probability is given by 0 and 1, representing true and false respectively; K represents the number of categories, where k and i are both category labels. and The corresponding representation uses the categories corresponding to k and i as context words to construct a hint, and each hint contains... Replace with the word embedding vectors corresponding to the class names (i.e., real, fake); For temperature parameters, For image encoders from training images Image features extracted from the image; cos represents cosine similarity.

[0049] Training is based on minimizing the cross-entropy classification loss, and the training objective is to maximize the score of the correct sample. The gradient can be obtained through the text encoder. The process involves backpropagation, leveraging the rich knowledge encoded in the parameters to optimize the learnable vector. Figure 2 (Use the flame symbol to mark it).

[0050] Since the specific working principle of the pre-trained multimodal large model and the subsequent training process can be implemented with reference to conventional techniques, they will not be elaborated further.

[0051] 3. Integration of prompts.

[0052] like Figure 2 As shown in the lower left, based on the prompt fine-tuning, a prompt ensemble method is proposed. This method converts multiple salient concept words into corresponding word vectors, which are then injected into the learnable context of each category according to preset rules for separate optimization and averaging. Specifically, as described in the aforementioned prompt fine-tuning section, by introducing CoOp to optimize the pre-trained multimodal large model for prompt fine-tuning, the model can adapt to specific downstream tasks, improving its efficiency in image classification tasks.

[0053] To further enhance the model's performance and optimize the text feature generation process, this invention introduces cue integration, such as... Figure 2 As shown in the lower left, multiple initial context phrases (i.e., the salient concept words extracted earlier) are first designed for each category. These context phrases are used to initialize a set of learnable context vectors, which are continuously optimized during training. For each context phrase, CoOp constructs a cue embedding (i.e., the cue t constructed earlier), and these cue embeddings are input into the text encoder. In this process, corresponding text features are generated. For each category, multiple cue embeddings are generated. ,in is the number of context phrases for each category, and these embeddings correspond to multiple context phrases. During forward propagation, the text features of these cue embeddings are averaged and fused to generate the next final text feature F for that category.

[0054] At this point, the formula for calculating the prediction probability becomes:

[0055] ;

[0056] in, This represents the training images computed based on prompts. Category The probability is given by 0 and 1, representing true and false respectively; K represents the number of categories, where k and i are both category labels. and The corresponding representations are the final text features for categories k and i; For temperature parameters, For training images Image features; cos represents cosine similarity.

[0057] The subsequent process is the same as the aforementioned fine-tuning section, and will not be repeated here.

[0058] By integrating prompts, the reliance on individual prompts in prompt fine-tuning is reduced. Different prompts embed different detection-related conceptual semantic information, and the resulting text features after averaging and fusion are more comprehensive and stable, better matching the features of the generated image in the feature space. This improves the accuracy of image-text similarity calculation, enhances classification and detection performance, and ultimately yields a general detector for generated images.

[0059] The above-mentioned solution provided in the embodiments of the present invention is based on a pre-trained multimodal large model text encoder, which adopts prompting fine-tuning, consumes less training resources, and is plug-and-play. The method can successfully detect images after common post-processing operations in the real world and has strong robustness. The method can simultaneously detect generated images with multiple different data distributions, has high generalization and strong practicality, and the average discrimination accuracy can reach more than 95%.

[0060] Through the above description of the embodiments, those skilled in the art can clearly understand that the above embodiments can be implemented by software, or by using software plus necessary general-purpose hardware platforms. Based on this understanding, the technical solutions of the above embodiments can be embodied in the form of a software product, which can be stored in a non-volatile storage medium (such as a CD-ROM, USB flash drive, mobile hard drive, etc.), including several instructions to cause a computer device (such as a personal computer, server, or network device, etc.) to execute the methods described in the various embodiments of the present invention.

[0061] Example 2 The present invention also provides a generative image detection system, which is mainly used to implement the method provided in the foregoing embodiments, such as... Figure 3 As shown, the system mainly includes: The salient concept extraction unit is used to decode the detection features of each training image and extract a salient concept word set by performing word frequency analysis and statistical extraction on the decoded text. The prompt fine-tuning unit is used to introduce context optimization technology. It initializes context labels using salient concept words and related non-salient concept words, and constructs learnable vectors. Then, it combines the given context words as prompts and inputs the prompts and training images into a pre-trained multimodal large model to perform image classification tasks, obtains prediction probabilities, and calculates a loss function based on the prediction probabilities to optimize the learnable vectors. The detection unit is used to construct cues using optimized learnable vectors and input them along with the image to be detected into a pre-trained multimodal large model to obtain the detection results.

[0062] Those skilled in the art will understand that, for the sake of convenience and brevity, the above-described division of functional modules is used as an example. In practical applications, the above functions can be assigned to different functional modules as needed, that is, the internal structure of the system can be divided into different functional modules to complete all or part of the functions described above.

[0063] Example 3 The present invention also provides a processing device, such as Figure 4 As shown, it mainly includes: one or more processors; a memory for storing one or more programs; wherein, when the one or more programs are executed by the one or more processors, the one or more processors implement the method provided in the foregoing embodiments.

[0064] Furthermore, the processing device also includes at least one input device and at least one output device; in the processing device, the processor, memory, input device, and output device are connected via a bus.

[0065] In this embodiment of the invention, the specific types of the memory, input device, and output device are not limited; for example: Input devices can be touchscreens, image acquisition devices, physical buttons, or mice, etc. The output device can be a display terminal; The memory can be random access memory (RAM) or non-volatile memory, such as disk storage.

[0066] Example 4 The present invention also provides a readable storage medium storing a computer program that, when executed by a processor, implements the method provided in the foregoing embodiments.

[0067] In this embodiment of the invention, the readable storage medium is a computer-readable storage medium and can be disposed in the aforementioned processing device, for example, as a memory in the processing device. Furthermore, the readable storage medium can also be any medium capable of storing program code, such as a USB flash drive, portable hard drive, read-only memory (ROM), magnetic disk, or optical disk.

[0068] The above description is merely a preferred embodiment of the present invention, but the scope of protection of the present invention is not limited thereto. Any variations or substitutions that can be easily conceived by those skilled in the art within the scope of the technology disclosed in the present invention should be included within the scope of protection of the present invention. Therefore, the scope of protection of the present invention should be determined by the scope of the claims. The information disclosed in the background section is intended only to enhance the understanding of the overall background technology of the present invention and should not be construed as an admission or implication in any way that such information constitutes prior art known to those skilled in the art.

Claims

1. A method for generating image detection, characterized in that, include: The detection features of each training image are decoded, and the significant concept word set is extracted by word frequency analysis and statistical extraction of the decoded text. Context optimization technology is introduced, which initializes context labels using salient concept words and related non-salient concept words, and constructs learnable vectors. Then, given context words are used as prompts. The prompts and training images are input into a pre-trained multimodal large model to perform image classification tasks, obtain prediction probabilities, and optimize the learnable vectors by calculating a loss function based on the prediction probabilities. The optimized learnable vectors are used to construct cueing, which is then input along with the image to be detected into a pre-trained multimodal large model to obtain the detection results.

2. The image generation detection method according to claim 1, characterized in that, The process of decoding the detection features of each training image and extracting a set of significant concept words through word frequency analysis and statistical extraction of the decoded text includes: For each training image in both the real image training set and the generated image training set, image features are extracted and linearly weighted to obtain detection features. Based on the image description generation model, the corresponding text representation is decoded and generated to form the text sets corresponding to the real image training set and the generated image training set, respectively. Then, word frequency analysis and statistical extraction are performed to obtain the salient concept word sets corresponding to the real image training set and the generated image training set, respectively.

3. The image generation detection method according to claim 2, characterized in that, The salient concept word sets corresponding to the real image training set and the generated image training set obtained by performing word frequency analysis and statistical extraction respectively include: The text sets corresponding to the real image training set and the text sets corresponding to the generated image training set are preprocessed separately. Natural language processing techniques were used to segment the two preprocessed text sets into words and count the frequency of each word. The top few words with the highest frequency in the two preprocessed text sets were selected as the salient concept word sets corresponding to the real image training set and the generated image training set, respectively.

4. The image generation detection method according to claim 1, characterized in that, Contextual tags are initialized using salient concept words and related non-salient concept words, and learnable vectors are constructed. These vectors are then combined with given context words as cues, as follows: ; ; Where t represents a prompt; A learnable vector consisting of M context labels. It is the m-th context tag. , The category label, used as contextual words, includes two categories: real and fake. Real means true, and fake means false. For text encoders in pre-trained multimodal large models; For a single salient concept word in the salient concept word set corresponding to the real image training set, To generate a single salient concept word from the salient concept word set corresponding to the image training set; or For a single non-salient concept word in the real image training set, To generate a single non-salient concept word from the image training set, a non-salient concept word is a word in the decoded text excluding salient concept words.

5. The image generation detection method according to claim 1, characterized in that, The step of inputting the prompts and training images together into a pre-trained multimodal large model for image classification to obtain the predicted probabilities includes: The prompts are input into the text encoder of a pre-trained multimodal large model. In this process, training images are input into the image encoder of a pre-trained multimodal large model, combined with a text encoder. The predicted probability is calculated from the output of the image encoder and expressed as: ; in, This represents the training images calculated based on prompts and fine-tuning. Category The probability is given by 0 and 1, representing true and false respectively; K represents the number of categories, where k and i are both category labels. and The corresponding representation uses the categories corresponding to k and i as context words to construct hints; For temperature parameters, For image encoders from training images Image features extracted from the image; cos represents cosine similarity.

6. The image generation detection method according to claim 1, characterized in that, The method also includes: introducing a cue integration method, converting multiple salient concept words into corresponding word vectors, injecting them into the learnable vector of each category, and obtaining multiple cuees corresponding to each category, wherein the category label corresponding to each category is the corresponding context word; When performing image classification tasks, the multimodal large model integrates the text features corresponding to all prompts under each category and takes the average as the final text feature. It then combines the corresponding image features to calculate the prediction probability.

7. The image generation detection method according to claim 6, characterized in that, The process of integrating the text features corresponding to all prompts under each category and averaging them to obtain the final text features, combined with the corresponding image features, to calculate the prediction probability includes: For each cue within each category, the text encoder in a pre-trained multimodal large model is used. Extract the corresponding text features; for each category, average and fuse the text features corresponding to all prompts to obtain the final text features; Using the final text features and the image features extracted by the image encoder of the pre-trained multimodal large model, the prediction probability is calculated and expressed as: ; in, This represents the training images computed based on prompts. Category The probability is given by 0 and 1, representing true and false respectively; K represents the number of categories, where k and i are both category labels. and The corresponding representations are the final text features for categories k and i; For temperature parameters, For training images Image features; cos represents cosine similarity.

8. A generated image detection system, characterized in that, To implement the method according to any one of claims 1 to 7, comprising: The salient concept extraction unit is used to decode the detection features of each training image and extract a salient concept word set by performing word frequency analysis and statistical extraction on the decoded text. The prompt fine-tuning unit is used to introduce context optimization technology. It initializes context labels using salient concept words and related non-salient concept words, and constructs learnable vectors. Then, it combines the given context words as prompts and inputs the prompts and training images into a pre-trained multimodal large model to perform image classification tasks, obtains prediction probabilities, and calculates a loss function based on the prediction probabilities to optimize the learnable vectors. The detection unit is used to construct cues using optimized learnable vectors and input them along with the image to be detected into a pre-trained multimodal large model to obtain the detection results.

9. A processing device, characterized in that, include: one or more processors; Memory, used to store one or more programs; Wherein, when the one or more programs are executed by the one or more processors, the one or more processors cause the one or more processors to implement the method as described in any one of claims 1 to 7.

10. A readable storage medium storing a computer program, characterized in that, When a computer program is executed by a processor, it implements the method as described in any one of claims 1 to 7.

Citation Information

Patent Citations

  • Image classification method based on cross-modal concept discovery and reasoning and intelligent terminal

    CN117115564A

  • Model training based on synthetic data

    US20230334834A1