Multi-modal harmful information detection method based on thinking chain guidance

Through the multimodal harmful information detection method based on thinking chain guidance, the problem of insufficient type specificity in the prior art is solved, and more accurate and efficient harmful meme detection is achieved, which significantly improves the detection accuracy.

CN120220160APending Publication Date: 2025-06-27NAT UNIV OF DEFENSE TECH
View PDF 0 Cites 2 Cited by

Patent Information

Application Number
CN202510266201.4
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-03-07
Publication Date
2025-06-27

AI Technical Summary

Technical Problem

The prior art lacks type specificity in harmful meme detection, resulting in inaccurate information extraction and inefficient efficiency, and large models are prone to hallucinations and harmful avoidance when dealing with sensitive topics.

Method used

The multimodal harmful information detection method based on thinking chain guidance is adopted, memes are clustered through multimodal fusion features, representative memes are extracted for multimodal large model training, semantic extraction and detection models of different types of memes are designed, and multimodal large model scheduler based on specific key points is used for detection.

Benefits of technology

It significantly improves the detection accuracy, ensures full utilization of effective information, fully utilizes the reasoning ability of the thinking chain and the planning ability of the big model, and solves the impact of type specificity on meme information extraction.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120220160A_ABST
    Figure CN120220160A_ABST
Patent Text Reader

Abstract

The invention discloses a multi-modal harmful information detection method based on thinking chain guidance, and the method comprises the following steps: collecting model factor data which is multi-modal data composed of an image and an accompanying text; clustering the moduli by using the multi-modal fusion features, and extracting a representative moduli closest to a clustering center and inputting the representative moduli to a multi-modal large model; applying a thinking chain reasoning method to obtain high-distinction-degree specific key points of each type of moduli; designing semantic extraction and detection models of different types of moduli, and inputting the moduli into different detection models according to the types for model training; during testing, a multi-modal large model scheduler based on a specific key point is used for judging the attribution type of the model factor, so that a detection model of a corresponding class is scheduled, and the probability of harmfulness of the model factor is output. According to the method, the influence of type specificity on memetic information extraction is solved; and the reasoning ability of the thinking chain and the planning ability of the large model are fully exerted.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention belongs to the field of information technology, and particularly relates to a multi-modal harmful information detection method guided by a chain of thought. Background Art

[0002] With the rapid progress of Internet technology, social network ecosystems represented by various online network communities, online social tools, platforms and services have developed and improved rapidly, becoming an important channel for people to obtain a vast amount of information resources. The form of information presentation has gradually become diversified. Based on the previous single-modal information such as only relying on pictures, texts or audio, new multi-modal information represented by short videos and memes has emerged. Memes use images and short texts to convey specific information or viewpoints humorously or satirically. This kind of information can usually intuitively present richer content and attract the attention of social users more. Due to its simple and easy-to-understand characteristics, memes can spread rapidly on social media and become a popular form of communication.

[0003] However, while social media provides convenience for information sharing, it has also become a breeding ground for the spread of harmful information. The emergence of memes has provided a new carrier for harmful information. Although memes are usually considered humorous, when used intentionally in a social and cultural context, they become potential sources of harm. Harmful memes are usually defined as "multi-modal units composed of images and accompanying texts that have the potential to cause harm to individuals, organizations, communities or society as a whole".

[0004] Currently, the detection of harmful memes has gradually attracted more attention from researchers and has achieved remarkable results. Traditional research methods for harmful meme detection rely on image and text encoders to extract visual and text features, and attempt to analyze cross-modal features through complex visual and language models to interpret their meanings. These methods can utilize the complete information of memes to achieve the detection purpose. With the development of pre-trained vision-language models, recent research has utilized the multi-modal information processing capabilities of PVLMs to enhance the harmful meme detection effect. Visual question answering technologies represented by BLIP-2 and Flamingo can more comprehensively mine the semantic information of memes and help the detector understand the meme semantics. In recent years, large language models (LLMs) and multi-modal large language models (MLLMs) represented by ChatGPT have powerful knowledge bases and have shown strong reasoning and judgment abilities in multi-modal information processing tasks, and can effectively assist in the detection of harmful memes. Recent research also uses LLMs to improve the performance of PVLMs and adopts a dialogue mechanism to encourage PVLMs to generate high-quality meme semantic features.

[0005] However, the applicant has found that the existing technologies still have the following problems: There is a general lack of type specificity. The same information extraction and processing methods are applied to all memes. However, the subjects that different memes convey information are different, the emphases on extracting semantic information vary, and the associated background knowledge is different. For example, for memes with people as the main subject, when extracting semantic information, it is necessary to pay attention to people's expressions, actions, clothing, and language, etc.; for memes with animals as the main subject, it is necessary to pay attention to the types of animals, associated knowledge, and offensive words, etc. If the same processing method is adopted, it will lead to the lack of effective information and the redundancy of ineffective information during reasoning, making it difficult to grasp the key information, and the information processing efficiency is low, making it difficult to improve the detection effect. Hallucinations and harmful avoidance occur during the inference using large models. During the process of using large models for harmful meme detection, the output results of large models are restricted by moral and ethical norms and will avoid generating harmful remarks. Large models also have strong planning capabilities, which have not been applied in this field yet. Summary of the Invention

[0006] To address the above challenges, this application proposes a multi-modal harmful information detection method guided by the chain of thought. Starting from a multi-modal type sensor, memes are clustered using multi-modal fusion features, and the representative meme closest to the cluster center is extracted and input into a multi-modal large model. The chain of thought (CoT) reasoning method is applied to obtain highly discriminative specific key points for each type of meme. Based on this, semantic extraction and detection models for different types of memes are designed, and memes are input into different detection models by type for model training. During testing, a multi-modal large model scheduler based on specific key points is designed to determine the type to which the meme belongs, so as to schedule the corresponding type of detection model and output the harmful probability of the meme. Experiments show that this model can significantly improve the detection accuracy.

[0007] To achieve the above object, the multi-modal harmful information detection method disclosed in this application includes the following steps:

[0008] Collect meme data, where the meme is multi-modal data composed of images and accompanying text;

[0009] Cluster memes using multi-modal fusion features, and extract the representative meme closest to the cluster center and input it into a multi-modal large model;

[0010] Apply the chain of thought reasoning method to obtain highly discriminative specific key points for each type of meme;

[0011] Design semantic extraction and detection models for different types of memes, and input memes into different detection models by type for model training;

[0012] During testing, a multi-modal large model scheduler based on specific key points is used to determine the meme attribution type, so as to schedule the corresponding type of detection model and output the probability of meme harmfulness.

[0013] Preferably, in the harmful meme data, each meme M = {I, T} is a tuple representing an image I accompanied by text T, with a ground truth label:

[0014]

[0015] where and indicate harmless, while and indicate harmful; the detection model is used to determine whether the meme is harmful. The detection model generates a score s in the label space, where s0 is the probability score indicating that the meme is harmless, and s1 is the probability score indicating that the meme is harmful.

[0016] Preferably, the method of using multi-modal fusion features to cluster memes and extract the representative memes closest to the cluster center and input them into the multi-modal large model includes:

[0017] Optical character extraction is performed on the meme text. Using the PaddleOCR technology, and then BERT is used to encode the text to obtain the text feature T:

[0018] T = BERT(t OCR )

[0019] In the visual feature extraction stage, the meme is input into VGG19 to obtain the visual feature V;

[0020] Perform multi-modal feature alignment: transform the text feature T and the visual feature V into the same embedding space to obtain the aligned features and Concatenate to obtain the multi-modal fusion feature F:

[0021]

[0022] The multi-modal feature matrix is clustered using the K-Means method to output the category label to which the picture belongs;

[0023] When outputting the category label of the meme, at the same time output the n memes closest to each cluster center as the representative memes of this type.

[0024] Preferably, the method of thinking chain reasoning is applied to obtain the highly discriminative specific key points of each type of meme. Specifically, when implemented with the large model based on multi-modal thinking chain reasoning, the information received by the large model is the labeled representative memes of each type, and corresponding prompts are generated;

[0025] On the one hand, the output answer is applied to the scheduler as the basis for judging the meme attribution during testing. On the other hand, it is applied to the detector and organized into the key questions of visual question answering.

[0026] Preferably, for the semantic extraction and detection models of different types of memes, the memes are input into different detection models by category for model training, including:

[0027] First, use the information passed by the planner to design the prompts of BLIP-2; the answer output by the planner is further refined using the large model;

[0028] Thereby, a number of questions Q for each category are output, transmitted to the corresponding detector, and descriptive text answers A are output for visual question answering for each category of memes. Then, PaddleOCR is used to extract the text caption C on the meme and the prompt label S: It is [MASK]. Concatenate all text modality information as the representation M of the meme = [A, C, S], where [MASK] in the prompt template represents the hidden label words harmful and harmless; input M into RoBERTa and compare the probabilities of [MASK] being a positive example and a negative example;

[0029] Specifically, in the vocabulary space of the label words, the confidence score of each label word generated by RoBERTa is represented as:

[0030] p = Sigmoid(RoBERTa([M test , M harmless , M harmful )),

[0031] where p ∈ R 2 ; M test is the test label word, M harmless is the label word for the meme being harmless, M harmful is the label word for the meme being harmful, p0 and p1 are the scores for the meme being harmless and the meme being harmful respectively, and use the scores {p0, p1} as the predicted label words; input each category of meme dataset into each detector for training to obtain a combination of multi-type detectors.

[0032] Preferably, the implementation of the multi-modal large model scheduler is as follows:

[0033] In the test stage, the large model judges the attribution category of the meme to be tested based on the scheduling output by the planner and schedules the corresponding detector to detect the meme;

[0034] After analysis by the scheduler, the type label of the meme to be tested is output, and the meme is transmitted to the corresponding detector to complete the detection.

[0035] Compared with the prior art, the present application has the following beneficial effects:

[0036] For the first time, memes are divided into different types for separate detection, solving the influence of type specificity on meme information extraction, thus ensuring the full utilization of effective information.

[0037] Applying the multi-modal chain of thought and large models to harmful meme detection gives full play to the reasoning ability of the chain of thought and the planning ability of large models. BRIEF DESCRIPTION OF THE DRAWINGS

[0038] Figure 1 The overall framework diagram of this application;

[0039] Figure 2 The CoT-based multi-modal large model planner provided by an embodiment of this application;

[0040] Figure 3 The test process. DETAILED DESCRIPTION OF THE EMBODIMENTS

[0041] The present invention will be further described below with reference to the accompanying drawings, but the present invention is not limited in any way. Any transformation or replacement made based on the teachings of the present invention falls within the protection scope of the present invention.

[0042] Natural Language Processing (NLP) is an important direction in the fields of computer science and artificial intelligence. It studies various theories and methods that can achieve effective communication between humans and computers in natural language. Natural language processing is a science that combines linguistics, computer science, and mathematics. Therefore, the research in this field will involve natural language, that is, the language people use in daily life, so it has a close connection with the research of linguistics. Natural language processing technologies usually include text processing, semantic understanding, machine translation, robot question answering, knowledge graphs, and other technologies.

[0043] The technical solutions provided in the embodiments of this application involve technologies such as machine learning and natural language processing in artificial intelligence, and will be introduced and described specifically through the following embodiments.

[0044] Before introducing the embodiments of this application, some terms involved in this application will be explained.

[0045] 1. Harmful information detection: Before the prevalence of harmful memes, researchers discovered other forms of harmful information on the Internet and studied identification methods, including harmful remarks, images, videos, etc. In recent years, the detection methods have been continuously optimized. The early methods directly fused text and visual modal information, gradually noticed cross-modal consistency, combined social context, or used external factual evidence verification, and the detection accuracy has been continuously improved. Among them, the content-based false information detection method can be interoperable with the harmful meme detection method.

[0046] The emotional characteristics of text are closely related to the detection of harmful information. Sentiment analysis involves identifying emotions (positive, negative, neutral), sentiment types (sarcastic, funny, offensive, motivating), and their intensities. Negative emotions are often associated with harmful information.

[0047] 2. Harmful meme detection: Harmful meme detection is an emerging multi-modal classification task. Early studies mainly used visual and text encoders to extract multi-modal features and fuse them, further focusing on the interaction between multi-modalities, such as Hate-CLIPper, Memefier, and MHA-Meme. Some studies have tried to explore effective methods to enhance detection, such as data augmentation techniques.

[0048] With the development of PVLMs, researchers have used PVLMs to assist in enhancing the detection effect of harmful memes. By converting multi-modal information into single-modal information in the form of image-to-text, encoding single-modal features, and classifying based on them. The detection effect of this type of method is affected by the integrity and objectivity of the text description.

[0049] The emergence of LLMs has provided new ideas for harmful meme detection. LLMs have a rich context knowledge base and demonstrate excellent capabilities in complex reasoning. However, current methods of this type ignore the specificity of memes, and LLMs are prone to hallucinations and harmful avoidance when dealing with sensitive topics, which have not been avoided.

[0050] 3. Chain of thought reasoning: With the continuous update and growing scale of large language models, new capabilities have been discovered, such as in-context learning and chain of thought reasoning. The chain of thought is an effective step-by-step reasoning strategy. By adding a step-by-step reasoning process to the prompt, the reasoning ability of LLMs can be improved. Therefore, the chain of thought has gradually been applied to enhance the zero-shot and few-shot reasoning of large models. The chain of thought was initially only applied in the text modality. With the increasing popularity of multi-modal data, some works have started to extend the single-modal CoT to the multi-modal CoT. Currently, multi-modal chain of thought reasoning technology is developing rapidly, and harmful meme detection research has not yet applied this advanced technology.

[0051] Reference Figure 1, the framework of the multi-modal harmful information detection method guided by the chain of thought disclosed in this application includes four major modules, namely multi-modal feature fusion and clustering, multi-modal large model planner based on the chain of thought, multi-modal large model scheduler, and detector. Specifically, first, a fine-grained type division is performed on memes, and a multi-modal large model is used to extract the features of various memes. These features have a high degree of discrimination, which helps to better understand the essential differences between different types of memes. Based on this, the large model plans the training objectives for the detector to ensure that the detector can be trained specifically for different types of memes. In the test stage, the large language model determines the type of the meme to be tested and schedules the corresponding detector for detection, so as to achieve more accurate detection of harmful memes.

[0052] This application first defines the problem to be solved using mathematical symbols. In the harmful meme dataset, each meme M = {I, T} is a tuple representing an image I accompanied by text T, with a ground truth label

[0053]

[0054] where and represent harmless, while and represent harmful. The detection model is designed to determine whether a meme is harmful. The model generates a score s in the label space, where s0 is the probability score indicating that the meme is harmless, and s1 is the probability score indicating that the meme is harmful.

[0055] The core of this application is to utilize the planning and reasoning ability of the multi-modal large model, combined with the CoT method, to provide intelligent guidance for the design and training of the detector. The model is divided into four major modules. Starting from the extraction and fusion of multi-modal features of memes, multi-modal clustering of memes is performed, and then the labeled meme dataset is input into the multi-modal large model planner based on CoT to plan the training objectives of the detector. After reasoning, the prompts for BLIP-2 and the key points for the scheduler are output, and multiple detectors are trained according to different classes. During the test process, the large model scheduler determines the category to which the meme to be tested belongs based on the key points, schedules the corresponding detector to detect the meme, and outputs the predicted harmful probability. The specific process is as Figure 2 shown.

[0056] Currently, there is no relevant classified meme dataset available for reference or use. Therefore, in one embodiment, this application first extracts and clusters meme features. Memes are new types of multi-modal data, based on images and attached with text. When extracting their features, it is necessary to ensure that all elements are complete, and commonly used text and visual encoders are selected. To extract text features, first, optical character extraction is performed on the meme text using the PaddleOCR technology, and then the classic BERT is used to encode the text to obtain the text feature T.

[0057] T = BERT(t OCR )

[0058] In the visual feature extraction stage, the meme is input into VGG19 to obtain the visual feature V. Then, for multi-modal feature alignment, the text feature T and the visual feature V are transformed into the same embedding space to obtain the aligned features and which are concatenated to obtain the multi-modal fusion feature F.

[0059]

[0060] The multi-modal feature matrix is clustered using the K-Means method to output the category label of the picture. Since the subsequent tasks need to utilize the planning ability of the multi-modal large model, but the current technology is difficult to process thousands or tens of thousands of pictures simultaneously, to improve the planning efficiency, when outputting the category label of the meme, the n memes closest to each cluster center are simultaneously output as the representative memes of this type for subsequent work.

[0061] The chain of thought is a training method that introduces logical reasoning ability in large models. It can decompose a problem into a series of consecutive small problems, guiding the model to gradually obtain the final answer, imitating the thinking process of humans when solving problems. The multi-modal chain of thought combines text and visual modalities for application, realizing reasoning and answer generation. Experiments have confirmed that multi-modal CoT can utilize multi-modal information to generate better reasoning with a significant performance improvement. Therefore, in one embodiment, the present application applies the multi-modal chain of thought technology when planning the meme detector.

[0062] The present application aims to perform target planning for multiple types of detectors, clarify the key attention tasks of each type of detector, and design different training objectives based on the differences in meme themes and presentation forms. Specifically, when implemented with the large model based on multi-modal CoT, the information received by the large model is the representative meme of each type with labels, and the corresponding prompts are designed as follows:

[0063] [“Based on the above several different types of representative memes, please find their commonalities. And briefly summarize the key points to note when detecting harmfulness, with no more than three words in total. Explain step by step the basis for the extraction.”]

[0064] The last sentence is proven to be an important prompt for zero-shot CoT, which can guide the large model to generate results layer by layer. The reasoning process is as follows: The first step is to analyze the visual content; the second step is to analyze the text content; the third step is to combine the visual and text content to analyze and distinguish the special key points of each type of meme, no more than 3. After reasoning, the large model outputs the target answer, which is the specific key points extracted for each type.

[0065] [“Special points for each category are as follows:

[0066] Category 1: ...;

[0067] Category 2: ...;

[0068] …

[0069] Category k:…”]

[0070] The output answers are used in the scheduler as a basis for determining the attribution of memes during testing, and in the detector to be organized into key questions for visual question answering.

[0071] In one embodiment, the harmful meme detector is implemented as follows:

[0072] All detectors use the currently used basic architecture. Specifically, we first use the information passed by the planner to design the BLIP-2 prompt. However, the answer output by the planner cannot be used directly, so we use the large model to further refine it:

[0073] [“Based on the above points, what questions should be asked of BLIP-2 to fully extract this type of information and bring in external knowledge? Note that there are 3 questions in each direction, and the questions should not be specific to this meme, but to this type of meme. The topics of different classes may change, so the questions need to be further generalized.”]

[0074] Thus, several questions Q = [q1, q2, q3...] of each category are output and transmitted to the corresponding detector. Visual question answering is performed for each type of meme, and descriptive text answers A = [a1, a2, a3...] are output. Then PaddleOCR is used to extract the text subtitle C on the meme, and the prompt label S: It is [MASK]. All text modal information is concatenated as the representation of the meme M = [A, C, S]. The [MASK] of the prompt template represents the hidden label words harmful and harmless. M is input into RoBERTa to compare the probability of [MASK] being a positive example (harmless) and a negative example (harmful).

[0075] Specifically, in the vocabulary space of label words, the confidence score of each label word generated by RoBERTa is expressed as:

[0076] p = Sigmoid(RoBERTa([M test ,M harmless ,M harmful ])),

[0077] where p∈R 2Use the fractions {p0, p1} as the labeled words for our prediction. Input each type of meme dataset into each detector for training to obtain a combination of multi-type detectors.

[0078] In one embodiment, the implementation of the multi-modal large model scheduler is as follows:

[0079] In the test phase, the large model determines the category to which the meme to be tested belongs based on the scheduling output by the planner and schedules the corresponding detector to detect the meme. If the input meme belongs to the first category, the LLM will schedule the first-category detector to detect it. The specific transfer process is as Figure 3 shown.

[0080] The prompt for designing the scheduler is:

[0081] ["Based on the multi-modal content of the meme, determine which of the following categories it belongs to and gradually explain the basis. The following are the key points for distinguishing each category. Category 1:...; Category 2:...;... Category k:......"]

[0082] After analysis by the scheduler, the type label of the meme to be tested is output, and the meme is transmitted to the corresponding detector to complete the detection.

[0083] Compared with the prior art, the present application has the following beneficial effects:

[0084] For the first time, memes are divided into different types for separate detection, solving the influence of type specificity on meme information extraction, thereby ensuring the full utilization of effective information.

[0085] Applying the multi-modal chain of thought and large model to harmful meme detection gives full play to the reasoning ability of the chain of thought and the planning ability of the large model.

[0086] As used herein, the term "preferred" is intended to be used as an example, illustration, or exemplification. Any aspect or design described as "preferred" herein need not be construed as more advantageous than other aspects or designs. Instead, the use of the term "preferred" is intended to present concepts in a specific manner. As used in this application, the term "or" is intended to mean an inclusive "or" rather than an exclusive "or". That is, unless otherwise specified or clear from the context, "X uses A or B" is intended to naturally include any one of the permutations. That is, if X uses A; X uses B; or X uses both A and B, then "X uses A or B" is satisfied in any of the foregoing examples.

[0087] Moreover, although the present disclosure has been shown and described with respect to one or more implementations, equivalent variations and modifications will occur to those skilled in the art based on a reading and understanding of this specification and the drawings. The present disclosure includes all such modifications and variations and is limited only by the scope of the appended claims. In particular with respect to the various functions performed by the above-described components (e.g., elements, etc.), the terms used to describe such components are intended to correspond to any component that performs the specified function of the component (e.g., it is functionally equivalent), unless otherwise indicated, even if structurally different from the disclosed structure that performs the functions in the exemplary implementations of the present disclosure shown herein. In addition, although a particular feature of the present disclosure has been disclosed with respect to only one of several implementations, such a feature may be combined with one or other features of other implementations as may be desired and advantageous for a given or particular application. Moreover, insofar as the terms "comprises," "has," "contains," or any variation thereof are used in a particular embodiment or claim, such terms are intended to be inclusive in a manner similar to the term "includes."

[0088] Each functional unit in the embodiments of the present invention may be integrated into a processing module, or each unit may exist physically alone, or multiple or more than multiple units may be integrated into one module. The above-mentioned integrated module may be implemented in the form of hardware or in the form of a software functional module. When the integrated module is implemented in the form of a software functional module and sold or used as an independent product, it may also be stored in a computer-readable storage medium. The above-mentioned storage medium may be a read-only memory, a magnetic disk, an optical disk, or the like. Each of the above-mentioned devices or systems may execute the storage method in the corresponding method embodiment.

[0089] In summary, the above embodiments are one implementation manner of the present invention, but the implementation manner of the present invention is not limited by the above embodiments. Any other changes, modifications, substitutions, combinations, and simplifications made without departing from the spirit and principle of the present invention shall be equivalent replacement manners and are all included in the protection scope of the present invention.

Claims

1. A multimodal harmful information detection method based on thought chain guidance, characterized in that: The following steps are involved: Collecting meme data, wherein memes are multimodal data consisting of images and accompanying text; Use multimodal fusion features to cluster memes, and extract representative memes closest to the cluster center and input them into the multimodal large model; Apply the chain of thought reasoning method to obtain the key points of high discrimination for each type of meme; Design semantic extraction and detection models for different types of memes, and input memes into different detection models according to their categories for model training; During testing, a multimodal large model scheduler based on specific key points is used to determine the type of meme, thereby scheduling the detection model of the corresponding class and outputting the probability of the meme being harmful.

2. The multimodal harmful information detection method based on thought chain guidance according to claim 1 is characterized in that: In the harmful meme data, each meme M = {I, T} is a tuple representing an image I accompanied by text T with a ground truth label: in and It means harmless, but and Indicates harmful; the detection model is used to determine whether the meme is harmful. The detection model generates a score s in the label space, where s0 is the probability score indicating that the meme is harmless, and s1 is the probability score indicating that the meme is harmful.

3. The multimodal harmful information detection method based on thought chain guidance according to claim 2 is characterized in that: The method of clustering memes using multimodal fusion features and extracting representative memes closest to the cluster center and inputting them into the multimodal large model includes: Perform optical character extraction on the meme text using PaddleOCR technology, and then use BERT to encode the text to obtain the text feature T: T=BERT(t OCR ) In the visual feature extraction stage, the meme is input into VGG19 to obtain the visual feature V; Perform multimodal feature alignment: convert text features T and visual features V into the same embedding space to obtain aligned features and The multimodal fusion feature F is obtained by splicing: The multimodal feature matrix is ​​clustered using the K-Means method to output the category label to which the image belongs; When outputting the category label of the meme, the n memes closest to each cluster center are also output as representative memes of that type.

4. The multimodal harmful information detection method based on thought chain guidance according to claim 3 is characterized in that: Apply the chain of thought reasoning method to obtain the highly distinguishable specific key points of each type of meme. Specifically, when using a large model based on multimodal chain of thought reasoning, the information received by the large model is each type of representative meme with a label, and corresponding prompts are generated; The output answers are used in the scheduler as a basis for determining the attribution of memes during testing, and in the detector to be organized into key questions for visual question answering.

5. The multimodal harmful information detection method based on thought chain guidance according to claim 4 is characterized in that: The design of semantic extraction and detection models for different types of memes, and inputting memes into different detection models for model training according to their types, includes: First, we use the information delivered by the planner to design BLIP-2 prompts; the answers output by the planner are further refined using the large model; Several questions Q of each category are output and transmitted to the corresponding detector. For each category of meme, visual question answering is performed to output descriptive text answers A. Then PaddleOCR is used to extract the text subtitles C on the meme and the prompt label S: It is [MASK]. All text modal information is concatenated as the meme representation M = [A, C, S], where the prompt template [MASK] represents the hidden label words harmful and harmless; M is input into RoBERTa to compare the probability of [MASK] being a positive example and a negative example; Specifically, in the vocabulary space of label words, the confidence score of each label word generated by RoBERTa is expressed as: p=Sigmoid(RoBERTa([M test ,M harmless ,M harmful ])), where p∈R 2 ;M test is the label word of the test, M harmless is a harmless tag word for memes, M harmful is the label word of the harmful meme, p0 and p1 are the scores of the harmless meme and the harmful meme, and the scores {p0, p1} are used as the predicted label words; each type of meme data set is input into each detector for training to obtain a combination of multiple types of detectors.

6. The multimodal harmful information detection method based on thought chain guidance according to claim 5 is characterized in that: The implementation of the multimodal large model scheduler is as follows: In the testing phase, the large model determines the category of the meme to be tested based on the schedule output by the planner, and schedules the corresponding detector to detect the meme; After analysis by the scheduler, the type label of the meme to be tested is output, and the meme is transmitted to the corresponding detector to complete the detection.

Citation Information

Cited By

  • Video content security understanding method, system and device and storage medium

    CN120823549A

  • Method for identifying aircraft type in satellite image, electronic equipment, storage medium and program product

    CN120894704A