An image generation method, system, device and storage medium

By decoding salient concept words and combining them with context optimization techniques, the alignment of visual and textual features is enhanced, which solves the problem of insufficient generalization ability of traditional detection methods under unseen data distributions and achieves efficient generated image detection.

CN120852888BActive Publication Date: 2025-11-25UNIV OF SCI & TECH OF CHINA
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202511351370.4
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-09-22
Publication Date
2025-11-25
Estimated Expiration
2045-09-22

AI Technical Summary

Technical Problem

Traditional generated image detection methods lack generalization ability when faced with unseen data distributions, and their detection performance based on visual features is suboptimal, making it difficult to effectively detect generated images.

Method used

By decoding the detection features of training images, extracting salient concept words and constructing learnable vectors, combining context optimization techniques and multimodal large models for image classification, and utilizing cross-modal alignment of visual and textual features, concept-level feature capture is achieved.

Benefits of technology

It improves detection performance for unseen data distributions and complex scenes, has strong robustness, can successfully detect images that have undergone real-world post-processing, has high generalization ability, and its average discrimination accuracy can reach over 95%.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120852888B_ABST
    Figure CN120852888B_ABST
Patent Text Reader

Abstract

The application discloses a kind of generation image detection method, system, equipment and storage medium, they are corresponding scheme, in scheme: by extracting significant concept word and converting into learnable text prompt, enhance the cross-modal alignment of visual and text features, by subsequent fine-tuning can fuse multi-concept semantic information, realize the capture of concept-level features of generated image, to improve the detection performance of unobserved data distribution and complex scene;In addition, the training resources consumed by the application during training are less, plug and play;And can successfully detect the image under the common post-processing operation of real world, robustness is strong;At the same time, a variety of different data distribution generated images can be detected simultaneously, generalization is high, practicality is strong, average discrimination accuracy can reach more than 95%.
Need to check novelty before this filing date? Find Prior Art

Description

TECHNICAL FIELD

[0001] The present application relates to the field of generated image detection technology, and in particular to a generated image detection method, system, device and storage medium. BACKGROUND

[0002] Generated artificial intelligence images refer to a kind of artificially synthesized images using artificial intelligence technology. With the continuous development of generation technology, today's generated images present the characteristics of controllable content, realistic visual effect and low generation threshold. Malicious misuse of generation technology may cause a series of serious social security problems. Therefore, efficient detection technology is urgently needed.

[0003] Traditional detection methods are mainly divided into training models through spatial domain or frequency domain features. The method based on spatial domain feature training detects artifacts such as texture irregularity and edge distortion through pixel-level features; the method based on frequency domain feature training transforms image data from spatial domain to frequency domain, checks the different characteristics shown by generated images and real images in the frequency domain, and then effectively distinguishes generated images and real images. However, the traditional detection method shows a lack of generalization ability when facing data distribution not seen in the training process. With the emergence of multi-modal visual language models, recent research has begun to use the rich prior knowledge obtained by training on large-scale image-text data sets, and through the way of aligning visual and text features to detect generated images. The latest research uses the CLIP (Contrastive Language-Image Pre-training) model as a feature extraction model, and trains a linear classification head on it to achieve excellent detection performance. However, this method does not use the CLIP text component, and only relies on visual features, which may lead to suboptimality, and the performance of generated image detection needs to be further strengthened.

[0004] In view of this, the present application is proposed. SUMMARY

[0005] The purpose of the present application is to provide a generated image detection method, system, device and storage medium, which can capture the concept-level features of generated images, thereby improving the detection performance of unseen data distribution and complex scenes.

[0006] The purpose of the present application is achieved by the following technical solutions:

[0007] A generated image detection method, comprising:

[0008] Decoding the detection features of each training image, and extracting a set of significant concept words by performing word frequency analysis and statistical extraction on the decoded text;

[0009] The context optimization technology is introduced, the significant concept word and the related non-significant concept word are used to initialize the context label and construct a learnable vector, the given context word is combined as a prompt, the prompt and the training image are input into the pre-trained multi-modal large model to perform an image classification task, a prediction probability is obtained, and a loss function is calculated based on the prediction probability to optimize the learnable vector;

[0010] The prompt is constructed by using the optimized learnable vector, and is input into the pre-trained multi-modal large model together with the image to be detected to obtain a detection result.

[0011] A generation image detection system is used to implement the foregoing method, and includes:

[0012] A significant concept extraction unit is configured to decode the detection features of each training image, and extract a significant concept word set by performing word frequency analysis and statistical extraction on the decoded text;

[0013] A prompt fine-tuning unit is configured to introduce the context optimization technology, use the significant concept word and the related non-significant concept word to initialize the context label, construct a learnable vector, combine the given context word as a prompt, input the prompt and the training image into the pre-trained multi-modal large model to perform an image classification task, obtain a prediction probability, and optimize the learnable vector by calculating a loss function based on the prediction probability.

[0014] A detection unit is configured to construct a prompt by using the optimized learnable vector, and input the prompt and the image to be detected into the pre-trained multi-modal large model to obtain a detection result.

[0015] A processing device includes one or more processors, and a memory for storing one or more programs;

[0016] When the one or more programs are executed by the one or more processors, the one or more processors implement the foregoing method.

[0017] A readable storage medium stores a computer program, and when the computer program is executed by a processor, the foregoing method is implemented.

[0018] It can be seen from the technical solutions provided by the application that the significant concept words are extracted and converted into learnable text prompts, the cross-modal alignment of visual features (image features) and text features is enhanced, the multi-concept semantic information can be fused through subsequent fine-tuning, the concept-level features of the generated images are captured, and the detection performance on unobserved data distribution and complex scenes is improved. In addition, the training resources consumed during the training of the application are less, and the application can be used immediately. The application can successfully detect images after common post-processing operations in the real world, has strong robustness, can detect generated images of multiple different data distributions at the same time, has high generalization and strong practicability, and the average discrimination accuracy can reach more than 95%. BRIEF DESCRIPTION OF DRAWINGS

[0019] In order to more clearly illustrate the technical solutions of the embodiments of the application, the drawings needed in the embodiment description will be briefly introduced. Obviously, the drawings in the following description are only some embodiments of the application, and other drawings can be obtained by those skilled in the art without creative labor.

[0020] Figure 1 A flowchart of a generated image detection method provided by the embodiment of the application.

[0021] Figure 2 A schematic diagram of the overall architecture of a generated image detection method provided by the embodiment of the application.

[0022] Figure 3 A schematic diagram of a generated image detection system provided by the embodiment of the application.

[0023] Figure 4 A schematic diagram of a processing device provided by the embodiment of the application. DETAILED DESCRIPTION

[0024] The technical solutions in the embodiments of the application will be described clearly and completely in combination with the drawings in the embodiments of the application. Obviously, the described embodiments are only some of the embodiments of the application, not all. Based on the embodiments of the application, all other embodiments obtained by those skilled in the art without creative labor are within the protection scope of the application.

[0025] First, the terms that can be used in this paper are described as follows:

[0026] The terms "comprising", "containing", "including", "having" or other similar semantic descriptions should be interpreted to be inclusive or open-ended, unless expressly specified otherwise. For example, the inclusion of an element (e.g., a raw material, a component, an ingredient, a carrier, a dosage form, a material, a dimension, a part, a component, a mechanism, a device, a step, a process, a method, a reaction condition, a processing condition, a parameter, an algorithm, a signal, data, a product, or an article, etc.) should be interpreted as including not only the explicitly recited element, but also other elements known to the art that are not explicitly recited.

[0027] The term "consisting of" means excluding any element not specifically recited. If the term is used in the claims, the term will make the claims closed, and will not include any element not specifically recited. If the term is used in a clause of the claims, it will limit only the clause in which it is used, and will not exclude elements recited in other clauses.

[0028] Unless expressly specified or limited otherwise, the terms "mounting", "connecting", "connecting", "fixing", and the like should be interpreted broadly, for example, can be fixedly connected, can be detachably connected, or integrally connected; can be mechanically connected, or electrically connected; can be directly connected, or indirectly connected through an intermediate medium, or can be connected internally between two elements. Those skilled in the art can understand the specific meaning of the above terms in this text according to the specific circumstances.

[0029] A method for generating image detection, a system, a device and a storage medium are described in detail below. The content not described in detail in the embodiments of the present application belongs to the prior art known to those skilled in the art. If no specific conditions are specified in the embodiments of the present application, the conventional conditions or the conditions recommended by the manufacturer are used. If no manufacturer of the reagent or instrument used in the embodiments of the present application is specified, it is a conventional product that can be purchased on the market.

[0030] Embodiment one

[0031] The embodiments of the present application provide a method for generating image detection, as shown in the following figure, which mainly includes the following steps: Figure 1

[0032] Step 1, significant concept extraction.

[0033] ​In the embodiment of the present application, the detection features of each training image are decoded, and the significant concept words are extracted by word frequency analysis and statistics on the decoded text. Specifically, for each training image in the real image training set and the generated image training set, image features are extracted respectively, and detection features are obtained by linear weighting, and the corresponding text representation is decoded based on the image description generation model, to form the text set corresponding to the real image training set and the generated image training set respectively, and then word frequency analysis and statistics are performed to obtain the significant concept word set corresponding to the real image training set and the generated image training set respectively.

[0034] In the embodiment of the present application, the significant concept word set corresponding to the real image training set and the generated image training set respectively obtained by the word frequency analysis and statistics includes: preprocessing the text set corresponding to the real image training set and the text set corresponding to the generated image training set respectively; using natural language processing technology to perform word segmentation on the two preprocessed text sets respectively, and counting the frequency of each word, and selecting the top several words with the highest frequency in the two preprocessed text sets respectively as the significant concept word set corresponding to the real image training set and the generated image training set respectively.

[0035] Step 2, prompt fine-tuning.

[0036] In the embodiment of the present application, the context optimization technology is introduced, the significant concept words and related non-significant concept words are used to initialize the context labels, and the learnable vector is constructed, and then the given context words are used as prompts, the prompts and the training images are input into the pre-trained multi-modal large model to perform image classification task, the prediction probability is obtained, and the loss function is calculated combined with the prediction probability to optimize the learnable vector.

[0037] In the embodiment of the present application, the significant concept words and related non-significant concept words are used to initialize the context labels, and the learnable vector is constructed, and then the given context words are used as prompts:

[0038] ;

[0039] ;

[0040] Wherein, t is a prompt; is a learnable vector composed of M context labels, is the mth context label, which is a vector with the same dimension as the word embedding, M is a hyperparameter that specifies the number of context labels, is a classification label, which is used as a context word and includes real and fake, real is true, and fake is false; is a text encoder in the pre-trained multi-modal large model; For a single salient concept word in the salient concept word set corresponding to the real image training set, To generate a single salient concept word from the salient concept word set corresponding to the image training set; or For a single non-salient concept word in the real image training set, To generate a single non-salient concept word from the image training set; a non-salient concept word is a word in the decoded text excluding salient concept words.

[0041] For example, the prompt could be: [Significant concept word] a photo of [real / fake]. In this example, "a photo of" is a non-concept word related to the significant concept word. Of course, in actual work, "a photo of" here is the corresponding vector form.

[0042] In this embodiment of the invention, the step of inputting the prompts and training images together into a pre-trained multimodal large model for image classification to obtain the predicted probabilities includes:

[0043] The prompts are input into the text encoder of a pre-trained multimodal large model. In this process, training images are input into the image encoder of a pre-trained multimodal large model, combined with a text encoder. The predicted probability is calculated from the output of the image encoder and expressed as:

[0044] ;

[0045] in, This represents the training image calculated based on prompts and fine-tuning. Category The probability is given by 0 and 1, representing true and false respectively; K represents the number of categories (K=2), where k and i are both category labels. and The corresponding representation uses the categories corresponding to k and i as context words to construct hints; For temperature parameters, For image encoders from training images Image features extracted from the image; cos represents cosine similarity.

[0046] Preferably, to further enhance the pre-trained multimodal large model, the present invention also provides a prompt integration scheme, which involves converting multiple significant concept words into corresponding word vectors and injecting them into the learnable vector of each category to obtain multiple prompts corresponding to each category. The category label corresponding to each category is the corresponding context word. When the multimodal large model performs an image classification task, it integrates the text features corresponding to all prompts under each category and takes the average as the final text feature. Combined with the corresponding image features, the prediction probability is calculated.

[0047] Specifically, the process of integrating the text features corresponding to all prompts under each category and averaging them to obtain the final text features, and combining them with the corresponding image features to calculate the predicted probability, includes: for each prompt under each category, processing the text encoder in a pre-trained multimodal large model. Extract the corresponding text features; for each category, average and fuse all the text features corresponding to the prompts to obtain the final text features; using the final text features and the image features extracted by the image encoder of the pre-trained multimodal large model, calculate the prediction probability, expressed as:

[0048] ;

[0049] in, This represents the training images computed based on prompts. Category The probability is given by 0 and 1, representing true and false respectively; K represents the number of categories, where k and i are both category labels. and The corresponding representations are the final text features for categories k and i; For temperature parameters, For training images Image features; cos represents cosine similarity.

[0050] Step 3: Image detection.

[0051] The optimized learnable vectors are used to construct cue, which is then input into a pre-trained multimodal large model along with the image to be detected to obtain the detection result (i.e. whether it is a generated image).

[0052] In this embodiment of the invention, when constructing prompts using optimized learnable vectors, the prompts do not contain classification labels.

[0053] The above-mentioned solution provided by the embodiments of the present invention enhances the cross-modal alignment of visual and textual features in a multimodal large model by extracting significant concept words and converting them into learnable text prompts. With the help of prompts, it can fine-tune the fusion of multi-concept semantic information and achieve the capture of concept-level features of generated images, thereby improving the detection performance for unseen data distributions and complex scenes.

[0054] The solution provided in this invention can be used to detect generative artificial intelligence images. In implementation, it can be deployed on the backend servers of various video websites and online forums to effectively detect, label, and intercept generated images before publication, eliminating potential negative impacts. It can also be applied to security regulatory departments to analyze and collect evidence of suspicious images disseminated on online platforms, maintaining network security.

[0055] In order to more clearly show the technical solutions provided by the present application and the technical effects produced, the method provided by the embodiments of the present application is described in detail below with specific examples.

[0056] As shown in the upper side, the text decoded from the detection features of the training set is subjected to word frequency analysis and statistical extraction to obtain a set of significant concept words. Figure 2 As shown in the upper side, the text decoded from the detection features of the training set is subjected to word frequency analysis and statistical extraction to obtain a set of significant concept words.

[0057] 1. Significant concept extraction.

[0058] As shown in the upper side, the text decoded from the detection features of the training set is subjected to word frequency analysis and statistical extraction to obtain a set of significant concept words. Figure 2 As shown in the upper side, the text decoded from the detection features of the training set is subjected to word frequency analysis and statistical extraction to obtain a set of significant concept words.

[0059] Specifically, for each image in the training set containing real images and generated images, the respective image features are obtained by using the image encoder, the respective detection features are obtained by linearly weighting the image features by the parameter weights and bias of the linear classifier trained in the latest research work, the text representation corresponding to the features is obtained by using the image description generation model, and then the real set detection feature text set (the text set corresponding to the real image training set) and the generated set detection feature text set (the text set corresponding to the generated image training set) are obtained. In the respective text sets, preprocessing operations such as lower case conversion, removal of punctuation marks and stop words are used, then the preprocessed text is segmented by using natural language processing technology, and the frequency of each word is counted. According to the word frequency counting result, the top several words with the highest frequency are extracted, and these high-frequency words are considered as the significant concepts in the respective text sets.

[0060] In order to avoid overlap of significant concept words in the real image training set and the generated image training set, and further distinguish the significant concepts of the detection features of the real set and the generated set, when selecting the respective significant concept words, the word elements that are more than S1 (for example, S1 = 2) times apart in different categories and still in the top S2 (for example, S2 = 25) words of the corresponding set are selected as the final significant concept words. The calculation method of the word frequency difference is: the frequency of the generated image training set / the frequency of the real image training set ≥ 2 (extract the significant concept words corresponding to the generated image training set), or the frequency of the real image training set / the frequency of the generated image training set ≥ 2 (extract the significant concept words corresponding to the real image training set). For example, if “artificial” appears 100 times in the generated image training set and 40 times in the real image training set, the calculation is “100 / 40=2.5≥2”, which meets the standard, and therefore, it is extracted as the significant concept word corresponding to the generated image training set.

[0061] Exemplary: the image encoder can use a pre-trained model ViT-L / 14, where ViT (Vision Transformer) is a visual transformer, L is a larger model with more parameters and stronger expression ability in the same type of visual transformer, and 14 represents that the input image is divided into 14x14 pixel image blocks; the image description generation model can select the open source ClipCap model (lightweight image description generation model based on CLIP).

[0062] In addition, Figure 2 The upper side only provides examples of significant concept words in part of English form such as person, woman, cat, table, room, etc., and specific significant concept words and their language forms are determined by the content of the training set, and the present application does not limit.

[0063] 2, prompt fine-tuning.

[0064] In the embodiments of the present application, the pre-trained multi-modal large model is fine-tuned based on the prompt to adapt it to the image detection task.

[0065] Specifically, the present application introduces context optimization (CoOp) to improve the efficiency of the pre-trained multi-modal large model in the image classification task. For example, the pre-trained multi-modal large model can select the CLIP model. Unlike traditional prompt templates, CoOp models the context words by using continuous vectors learned end-to-end from data while freezing the pre-trained parameters (the Figure 2 marked with snowflake symbols), thereby avoiding manual prompt fine-tuning. CoOp appends learnable vectors to the context words of the prompt, where the context words refer to the class labels of the given dataset, such as Figure 2 as shown on the right side, where the context words are real and fake. These learnable vectors can be initialized using pre-trained word embeddings. Specifically, for the prompt design of the text encoder is:

[0066] ;

[0067] ;

[0068] where, is the mth context token, which is a vector with the same dimension as the word embedding, M is a hyperparameter that specifies the number of context tokens, For the classification label, it can be placed at the end, at the beginning or in the middle as a context word, and the above formula provides an example of placing it at the end, the classification label includes two categories of real and fake, real is true, and fake is false; For the text encoder in the pre-trained multi-modal large model.

[0069] In the embodiment of the application, a prompt t is input into the text encoder to obtain a classification weight vector representing a visual concept, and the image features extracted by the image encoder are combined to calculate the prediction probability in the following manner:

[0070] ;

[0071] wherein, represents the probability of the class of the training image calculated based on prompt fine-tuning, 0 and 1 represent true and false; K represents the number of classes, k and i are both class labels, and correspond to prompts constructed by taking the class corresponding to k and i as the context word, the in each prompt is replaced by the word embedding vector corresponding to the class name (i.e., real, fake); is a temperature parameter, is the image feature extracted by the image encoder from the training image ; and cos represents the cosine similarity.

[0072] The training is based on minimizing the cross-entropy classification loss, and the training goal is to maximize the score of the correct sample, and the gradient can be back-propagated through the text encoder to optimize the learnable vector (marked with a flame symbol in Figure 2 ).

[0073] Considering the specific working principle of the pre-trained multi-modal large model and the subsequent training process, which can be implemented according to conventional techniques, further description is omitted.

[0074] 3. Prompt integration.

[0075] As Figure 2 shown on the lower left side, on the basis of prompt fine-tuning, a prompt integration method is proposed, in which a plurality of significant concept words are converted into corresponding word vectors, which are injected into the learnable context of each class according to a predetermined rule for separate optimization, combination and averaging. Specifically, as introduced in the foregoing prompt fine-tuning part, by introducing CoOp to optimize the prompt fine-tuned pre-trained multi-modal large model, the model can adapt to specific downstream tasks and improve its efficiency in image classification tasks. ​

[0076] To further enhance the model's performance and optimize the text feature generation process, this invention introduces cue integration, such as... Figure 2 As shown in the lower left, multiple initial context phrases (i.e., the salient concept words extracted earlier) are first designed for each category. These context phrases are used to initialize a set of learnable context vectors, which are continuously optimized during training. For each context phrase, CoOp constructs a cue embedding (i.e., the cue t constructed earlier), and these cue embeddings are input into the text encoder. In this process, corresponding text features are generated. For each category, multiple cue embeddings are generated. ,in is the number of context phrases for each category, and these embeddings correspond to multiple context phrases. During forward propagation, the text features of these cue embeddings are averaged and fused to generate the next final text feature F for that category.

[0077] At this point, the formula for calculating the prediction probability becomes:

[0078] ;

[0079] in, This represents the training images computed based on prompts. Category The probability is given by 0 and 1, representing true and false respectively; K represents the number of categories, where k and i are both category labels. and The corresponding representations are the final text features for categories k and i; For temperature parameters, For training images Image features; cos represents cosine similarity.

[0080] The subsequent process is the same as the aforementioned fine-tuning section, and will not be repeated here.

[0081] By integrating prompts, the reliance on individual prompts in prompt fine-tuning is reduced. Different prompts embed different detection-related conceptual semantic information, and the resulting text features after averaging and fusion are more comprehensive and stable, better matching the features of the generated image in the feature space. This improves the accuracy of image-text similarity calculation, enhances classification and detection performance, and ultimately yields a general detector for generated images.

[0082] The scheme provided by the embodiment of the present application is based on a text encoder of a pre-trained multi-modal large model, adopts prompt fine-tuning, consumes fewer training resources, and is plug-and-play; the method can successfully detect images that have undergone common post-processing operations in the real world, has strong robustness; the method can simultaneously detect generated images of multiple different data distributions, has high generalization and strong practicability, and the average discrimination accuracy can reach more than 95%.

[0083] Those skilled in the art can clearly understand from the above description of the embodiments that the above embodiments can be implemented by software, or can be implemented by means of software plus a necessary general hardware platform. Based on such understanding, the technical solutions of the above embodiments can be embodied in the form of a software product, which can be stored in a non-volatile storage medium (which can be a CD-ROM, a U disk, a mobile hard disk, etc.), and includes a plurality of instructions for causing a computer device (which can be a personal computer, a server, or a network device, etc.) to execute the methods described in the various embodiments of the present application.

[0084] Embodiment two

[0085] The present application also provides a generated image detection system, which is mainly used to implement the method provided by the above embodiments, as shown in the figure, the system mainly includes: Figure 3

[0086] A significant concept extraction unit is configured to decode the detection features of each training image, and extract a significant concept word set through word frequency analysis and statistical extraction on the decoded text.

[0087] A prompt fine-tuning unit is configured to introduce a context optimization technique, initialize context labels using significant concept words and related non-significant concept words, construct a learnable vector, and combine a given context word as a prompt. The prompt and the training image are input into the pre-trained multi-modal large model to perform an image classification task, obtain a prediction probability, and optimize the learnable vector by combining the prediction probability to calculate a loss function.

[0088] A detection unit is configured to construct a prompt using the optimized learnable vector, and input the prompt and a to-be-detected image into the pre-trained multi-modal large model to obtain a detection result.

[0089] Those skilled in the art can clearly understand that, for the convenience and brevity of description, only the division of the above functional modules is exemplified, and in actual applications, the above functions can be completed by different functional modules according to needs, that is, the internal structure of the system is divided into different functional modules to complete all or part of the functions described above.

[0090] Embodiment three

[0091] ​The application further provides a processing device, such as Figure 4 as shown in the figure, mainly comprising: one or more processors; a memory for storing one or more programs; wherein when the one or more programs are executed by the one or more processors, the one or more processors implement the method provided by the foregoing embodiments.

[0092] Further, the processing device further comprises at least one input device and at least one output device; in the processing device, the processor, the memory, the input device and the output device are connected through a bus.

[0093] In the embodiments of the application, the specific types of the memory, the input device and the output device are not limited; for example:

[0094] The input device can be a touch screen, an image acquisition device, a physical button or a mouse, etc.

[0095] The output device can be a display terminal.

[0096] The memory can be a random access memory (RAM) or a non-volatile memory such as a disk memory.

[0097] Embodiment four

[0098] The application further provides a readable storage medium storing a computer program, when the computer program is executed by a processor, the method provided by the foregoing embodiments is implemented.

[0099] In the embodiments of the application, the readable storage medium as the computer readable storage medium can be arranged in the foregoing processing device, for example, as the memory in the processing device. In addition, the readable storage medium can also be a U disk, a mobile hard disk, a read-only memory (ROM), a magnetic disk or an optical disk and various media capable of storing program codes.

[0100] The above is only the preferred specific implementation of the application, but the protection scope of the application is not limited to this, any person skilled in the art can easily think of changes or replacements within the technical range disclosed by the application, which should be covered in the protection scope of the application. Therefore, the protection scope of the application should be subject to the protection scope of the claims. The information disclosed in the background section of this document is only intended to deepen the understanding of the overall background of the application, and should not be regarded as acknowledging or implying in any form that the information constitutes the prior art known to those skilled in the art.

Claims

1. A method for generating image detection, characterized in that, include: The detection features of each training image are decoded, and the significant concept word set is extracted by word frequency analysis and statistical extraction of the decoded text. This includes: for each training image in the real image training set and the generated image training set, image features are extracted and linearly weighted to obtain detection features. Based on the image description generation model, the corresponding text representation is decoded to form the text set corresponding to the real image training set and the generated image training set, and then word frequency analysis and statistical extraction are performed to obtain the significant concept word set corresponding to the real image training set and the generated image training set, respectively. Context optimization technology is introduced, which initializes context labels using salient concept words and related non-salient concept words, and constructs learnable vectors. Then, given context words are used as prompts. The prompts and training images are input into a pre-trained multimodal large model to perform image classification tasks, obtain prediction probabilities, and optimize the learnable vectors by calculating a loss function based on the prediction probabilities. The optimized learnable vectors are used to construct cueing, which is then input along with the image to be detected into a pre-trained multimodal large model to obtain the detection result; Specifically, context labels are initialized using salient concept words and related non-salient concept words, and learnable vectors are constructed. These vectors are then combined with given context words as cues, as shown below: ; ; Where t represents a prompt; A learnable vector consisting of M context labels. It is the m-th context tag. , The category label, used as contextual words, includes two categories: real and fake. Real means true, and fake means false. For text encoders in pre-trained multimodal large models; For a single salient concept word in the salient concept word set corresponding to the real image training set, To generate a single salient concept word from the salient concept word set corresponding to the image training set; or For a single non-salient concept word in the real image training set, To generate a single non-salient concept word from the image training set, a non-salient concept word is a word in the decoded text excluding salient concept words.

2. The image generation detection method according to claim 1, characterized in that, The salient concept word sets corresponding to the real image training set and the generated image training set obtained by performing word frequency analysis and statistical extraction respectively include: The text sets corresponding to the real image training set and the text sets corresponding to the generated image training set are preprocessed separately. Natural language processing techniques were used to segment the two preprocessed text sets into words and count the frequency of each word. The top few words with the highest frequency in the two preprocessed text sets were selected as the salient concept word sets corresponding to the real image training set and the generated image training set, respectively.

3. The image generation detection method according to claim 1, characterized in that, The step of inputting the prompts and training images together into a pre-trained multimodal large model for image classification to obtain the predicted probabilities includes: The prompts are input into the text encoder of a pre-trained multimodal large model. In this process, training images are input into the image encoder of a pre-trained multimodal large model, combined with a text encoder. The predicted probability is calculated from the output of the image encoder and expressed as: ; in, This represents the training image calculated based on prompts and fine-tuning. Category The probability is given by 0 and 1, representing true and false respectively; K represents the number of categories, where k and i are both category labels. and The corresponding representation uses the categories corresponding to k and i as context words to construct hints; For temperature parameters, For image encoders from training images Image features extracted from the image; cos represents cosine similarity.

4. The image generation detection method according to claim 1, characterized in that, The method also includes: introducing a cue integration method, converting multiple salient concept words into corresponding word vectors, injecting them into the learnable vector of each category, and obtaining multiple cuees corresponding to each category, wherein the category label corresponding to each category is the corresponding context word; When performing image classification tasks, the multimodal large model integrates the text features corresponding to all prompts under each category and takes the average as the final text feature. It then combines the corresponding image features to calculate the prediction probability.

5. The image generation detection method according to claim 4, characterized in that, The process of integrating the text features corresponding to all prompts under each category and averaging them to obtain the final text features, combined with the corresponding image features, to calculate the prediction probability includes: For each cue within each category, the text encoder in a pre-trained multimodal large model is used. Extract the corresponding text features; for each category, average and fuse the text features corresponding to all prompts to obtain the final text features; Using the final text features and the image features extracted by the image encoder of the pre-trained multimodal large model, the prediction probability is calculated and expressed as: ; in, This represents the training images computed based on prompts. Category The probability is given by 0 and 1, representing true and false respectively; K represents the number of categories, where k and i are both category labels. and The corresponding representations are the final text features for categories k and i; For temperature parameters, For training images Image features; cos represents cosine similarity.

6. A generated image detection system, characterized in that, To implement the method according to any one of claims 1 to 5, comprising: The salient concept extraction unit is used to decode the detection features of each training image and extract a salient concept word set by performing word frequency analysis and statistical extraction on the decoded text. The prompt fine-tuning unit is used to introduce context optimization technology. It initializes context labels using salient concept words and related non-salient concept words, and constructs learnable vectors. Then, it combines the given context words as prompts and inputs the prompts and training images into a pre-trained multimodal large model to perform image classification tasks, obtains prediction probabilities, and calculates a loss function based on the prediction probabilities to optimize the learnable vectors. The detection unit is used to construct cues using optimized learnable vectors and input them along with the image to be detected into a pre-trained multimodal large model to obtain the detection results.

7. A processing device, characterized in that, include: One or more processors; Memory, used to store one or more programs; Wherein, when the one or more programs are executed by the one or more processors, the one or more processors cause the one or more processors to implement the method as described in any one of claims 1 to 5.

8. A readable storage medium storing a computer program, characterized in that, When a computer program is executed by a processor, it implements the method as described in any one of claims 1 to 5.