Harmful image detection and recognition system and method based on agent collaboration

Through a harmful image detection and recognition system based on agent collaboration, multiple agent modules are used to comprehensively analyze image content and identify harmful information, which solves the problem of difficult to detect complex and obscure bad image content in the prior art, and achieves a high accuracy and interpretability of harmful information recognition effect.

CN120070906AActive Publication Date: 2025-05-30BEIJING UNIV OF POSTS & TELECOMM

Patent Information

Application Number
CN202510560794.5
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-04-30
Publication Date
2025-05-30
Estimated Expiration
2045-04-30

AI Technical Summary

Technical Problem

The prior art is difficult to effectively detect and identify complex and obscure bad image content that may exist in generative artificial intelligence models.

Method used

Adopt a harmful image detection and recognition system based on agent collaboration, through the division of labor and cooperation of multiple agent modules, image metasemantic information analysis, harmful feature detection, deep semantic analysis, questioning verification and dynamic comprehensive analysis are carried out to build a comprehensive semantic knowledge graph to achieve comprehensive analysis of image content and accurate identification of harmful information.

Benefits of technology

It improves the accuracy and interpretability of identification of complex and obscure image content, enhances the system's anti-interference ability and adaptability to diversified application scenarios.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120070906A_ABST
    Figure CN120070906A_ABST
Patent Text Reader

Abstract

The invention relates to the technical field of image recognition, in particular to a harmful image detection and recognition system and method based on agent collaboration, and the system comprises a first agent module, a second agent module, a third agent module, a fourth agent module and a fifth agent module, the first agent module is used for analyzing image element semantic information, outputting structured data and constructing an image comprehensive semantic knowledge graph; and the second intelligent agent module is used for carrying out harmful feature detection and associated query on the image in combination with the structured data, and outputting a coarse-grained harmful semantic detection result and the confidence coefficient of the coarse-grained harmful semantic detection result. According to the harmful image detection and identification system based on intelligent agent cooperation, through hierarchical processing and cooperative work of the first to fifth intelligent agents, comprehensive analysis of the image is realized to obtain meta-semantic information, coarse-grained and fine-grained harmful semantic detection is performed on the meta-semantic information, and through multiple rounds of challenge and dynamic comprehensive analysis, the detection and identification efficiency of the harmful image is improved. Hidden harmful information can be accurately identified, so that the accuracy and interpretability of the harmful information are improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of image recognition technology, and in particular, to a harmful image detection and recognition system and method based on agent collaboration. Background Art

[0002] With the rapid development of generative artificial intelligence technology, text-to-image large models are widely used in various aspects such as ideological education, literary creation, and business design, which helps people intuitively display knowledge, express emotions, spread culture, or promote products. However, in practical applications, due to the possible malicious attacks on text-to-image large models by users. Since the application scenarios of text-to-image large models are diversified, the generation effects of bad images have the characteristics of wide sources, various styles, and deep hiding. Summary of the Invention

[0003] In view of this, the purpose of the present invention is to provide a harmful image detection and recognition system and method based on agent collaboration.

[0004] In a first aspect, an embodiment of the present invention provides a harmful image detection and recognition system based on agent collaboration, including: first to fifth agent modules with collaborative functions; The first agent module is used to analyze the meta-semantic information of the image, output structured data, and construct an image comprehensive semantic knowledge graph; The second agent module is used to detect harmful features of the image and perform associated queries in combination with the structured data, output a coarse-grained harmful semantic detection result and its confidence level, and dynamically update it to the image comprehensive semantic knowledge graph; The third agent module is used to deeply analyze the structured data and the coarse-grained harmful semantic detection result, output a fine-grained harmful semantic detection result, its confidence level, and dynamically update it to the image comprehensive semantic knowledge graph; The fourth agent module is used to question and verify the fine-grained harmful image semantic detection result, output various fine-grained harmful query records, network verification results and confidence levels, and dynamically update the query results to the image comprehensive semantic knowledge graph; The fifth agent module is used to perform dynamic comprehensive analysis on various harmful query results, and output harmful information recognition results, confidence levels, and judgment bases.

[0005] Combined with the first aspect, the first agent module includes: The person identity analysis module is used to perform preliminary recognition on the input person analysis prompt template and image based on a large model to output an image text description. When it is determined that there are people in the image, it is used to extract features of the people in the image based on a small model, analyze and output the person identity analysis result. It is also used to fuse the image text description output by the large model and the person identity analysis result output by the small model as initial meta-semantic information and construct an image comprehensive semantic knowledge graph; among them, the person identity analysis result includes at least one of group, gender, age, and occupation. The person emotion analysis module is used to, when it is determined that there are people in the image through preliminary recognition of the image, extract the expression features in the image based on a preset expression recognition small model and analyze based on the expression features to obtain the person emotion analysis result and the confidence level of the person emotion analysis result, and update the person emotion analysis result as the first meta-semantic information to the image comprehensive semantic knowledge graph; among them, the person identity analysis result includes at least one of group, gender, age, and occupation. The interaction relationship analysis module is used to, when it is determined that there are people in the image through preliminary recognition of the image, extract the action features in the image based on a preset action recognition small model and predict the interaction category in combination with the person emotion analysis result, output the interaction relationship analysis result and its confidence level, and update the interaction relationship analysis result as the second meta-semantic information to the image comprehensive semantic knowledge graph. The scene environment analysis module is used to perform environmental scene analysis on the image based on a preset scene recognition model and a light detection algorithm, output the environmental scene analysis result and its confidence level, the light analysis result and its confidence level, and update the environmental scene analysis result and the light analysis result as the third meta-semantic information to the image comprehensive semantic knowledge graph. The text information analysis module uses OCR technology to extract the text information in the image, uses a preset text recognition method and semantic analysis algorithm to analyze the text information in the image, obtains the text semantic information analysis result, and updates the image comprehensive semantic knowledge graph with the text semantic information analysis result as the fourth meta-semantic information.

[0006] Combined with the first aspect, the second intelligent agent module includes: The harmful element extraction module is used to screen harmful objects from the structured data output by the first intelligent agent module based on a small model, and output harmful object structured data; it is also used to perform harmful feature screening on the input harmful object detection prompt template and the image comprehensive semantic knowledge graph output by the first intelligent agent based on a large model to generate a harmful feature detection result and a confidence level, and update the harmful feature detection result as the fourth meta-semantic information to the image comprehensive semantic knowledge graph again.

[0007] A search and matching module, which is used to generate a search query associated with the harmful element based on the large model according to the extracted harmful element, call a preset external harmful element database for relevance matching, calculate the confidence level of the harmful element as a risk item, output the first parsing result as the fifth - element semantic information to update the image comprehensive semantic knowledge graph again, and generate a coarse - grained harmful semantic detection result by combining the fifth - element semantic information and the fourth - element semantic information.

[0008] Combined with the first aspect, the third intelligent agent module includes: A visual symbol parsing module, which detects harmful symbol information in the image based on a preset visual symbol recognition model, and performs in - depth parsing by combining the image comprehensive semantic knowledge graph information and the coarse - grained harmful semantic detection result output by the second intelligent agent module, and outputs a fine - grained harmful semantic detection result of the first specified category; A text metaphor mining module, which is used to perform in - depth parsing of the pun recognition features, historical mapping features, and regional cultural features in the image semantic information on the input metaphor analysis prompt template and the image based on the large model, outputs a fine - grained harmful semantic detection result of the second specified category, and updates the image comprehensive semantic knowledge graph with the fine - grained harmful semantic detection result of the second specified category as the sixth - element semantic information.

[0009] Combined with the first aspect, the fourth intelligent agent module includes: A multi - round debate base address module, when the fitness of the structured data, the coarse - grained harmful semantic detection result, the fine - grained harmful semantic detection result, and the harmful element is lower than the preset threshold, generates a questioning statement and distributes it to the first to third intelligent agent modules for re - parsing until the fitness is higher than the preset threshold, and determines the debate result, confidence level, and evidence chain information; A questioning analysis module, which is used to verify the consistency between the calculated image and the semantic information by using a text - from - image model, and inputs the fine - grained harmful semantic detection result of the second specified category and the questioning prompt template into the large model to perform a questioning analysis on the analyzed image comprehensive semantic knowledge graph, and obtains a fine - grained harmful semantic detection result of the third specified category; A questioning feedback mechanism module, when the fitness of the structured data, the coarse - grained harmful semantic information detection result, the fine - grained harmful semantic information detection result of the first specified category, and the fine - grained harmful semantic information detection result of the second specified category is lower than the preset threshold, generates a questioning question and dynamically adjusts the weight distribution based on the questioning result to dynamically update the image comprehensive semantic knowledge graph; An evidence chain construction module, which is used to output the coarse - grained harmful semantic detection result, the fine - grained harmful semantic detection result, and the harmful element association, and uses the large model to perform an online search on the output result to obtain authoritative information, and summarizes and forms a complete evidence chain composed of scattered evidence points; A contradiction analysis module, which is used to process structured data, coarse-grained harmful semantic detection results, and fine-grained harmful semantic detection results through multi-dimensional cross-validation using a large model to identify various harmful risk items.

[0010] Combined with the first aspect, the fifth intelligent agent module includes: A dynamic weight adjustment module, which is used to calculate the assignment weights of various harmful semantic detection results based on a preset dynamic weight algorithm; A risk quantification calculation module, which is used to perform a comprehensive scoring operation by combining the assignment weights of the fine-grained harmful semantic detection results, the type scores of the fine-grained harmful semantic detection results, and the scenario complexity function to obtain the confidence level of the fine-grained harmful semantic information detection results; A collaborative verification decision module, which is used to start a counter-questioning process for re-evidence collection when the confidence level is insufficient until a harmful information identification result with a confidence level meeting the preset requirements is obtained.

[0011] In the second aspect, an embodiment of the present invention further provides a harmful image detection and recognition method based on multi-agent collaboration, which is applied to the system as described above. The method includes: Analyze the meta-semantic information of the image through the first intelligent agent module, output structured data, and construct an image comprehensive semantic knowledge graph; Detect and perform associated queries on harmful elements of the image through the second intelligent agent module in combination with the structured data, output the coarse-grained harmful semantic detection results and their confidence levels, and dynamically update the image comprehensive semantic knowledge graph; Deeply analyze the structured data and the coarse-grained harmful semantic detection results through the third intelligent agent module, output the fine-grained harmful semantic detection results and their confidence levels, and dynamically update the image comprehensive semantic knowledge graph; Question and verify the coarse-grained harmful image semantic detection results and the fine-grained harmful image semantic detection results through the fourth intelligent agent module, output various harmful question records, verification evidence chains, network verification results, and confidence levels, and dynamically update the image comprehensive semantic knowledge graph; Perform dynamic comprehensive analysis on various harmful question results through the fifth intelligent agent module, and output harmful information identification results and judgment bases.

[0012] Combined with the second aspect, the steps for the fourth intelligent agent module to be used for questioning include: Calculate a certain type of risk score with the following formula: ; ; ; ; ; Among them, is the risk score of the detection result of the type of fine-grained harmful semantic information; and are both weight coefficients, and + ; is the evidence chain integrity score, ; is the multi-round interrogation consistency score, ; is the contradiction coefficient, ; is the network verification reliability score, ; is the total number of evidence points; represents the validity of the j th evidence point, takes a value of 0 or 1; is the result of the k-th round of interrogation, , is a natural number greater than 2; is the number of detected contradictions, Ntotal is the total number of checkpoints; is the weight coefficient, ; is the authoritative source ratio score, ; is the information consistency score, .

[0013] Combined with the first aspect, after the steps of the fourth agent module for interrogation, it includes: If , trigger multi-round interrogation, generate interrogation questions, and distribute them to the first to third agent modules, and then dynamically adjust the weight allocation based on the interrogation results; If , trigger the re-verification process, and update the evidence chain by searching for supplementary information through the network; If , determine that the current result is credible, and use the current result as the target result.

[0014] Combined with the second aspect, the steps of the fifth agent module for dynamically calculating the comprehensive risk score include: Calculate the comprehensive risk score with the following formula: ; ; Among them, is the comprehensive harmful risk score; is the Weight of harmful semantic detection results is the risk score of the detection result of the harmful semantic information of the = , is a natural number greater than 2; is the attenuation factor; is the implicit semantic weight; is the scene complexity function, 0. ; is the number of items and people in the image, .

[0015] The embodiments of the present invention bring the following beneficial effects: The harmful image detection and recognition system and method based on intelligent agent deep collaboration provided by this application, the system includes: the first to fifth intelligent agent modules with collaborative functions; the first intelligent agent module is used to analyze the meta-semantic information of the image, output structured data, and construct an image comprehensive semantic knowledge graph; the second intelligent agent module is used to detect harmful features and perform associated queries on the image in combination with the structured data, output the coarse-grained harmful semantic detection results and their confidence levels, and dynamically update the coarse-grained harmful semantic detection results to the image comprehensive semantic knowledge graph; the third intelligent agent module is used to deeply analyze the structured data and the coarse-grained harmful semantic detection results, output the fine-grained harmful semantic detection results and their confidence levels, and dynamically update the fine-grained harmful semantic detection results to the image comprehensive semantic knowledge graph; the fourth intelligent agent module is used to query and verify the fine-grained harmful image semantic detection results, output various fine-grained harmful query records, network verification results and confidence levels, and dynamically update the query results to the image comprehensive semantic knowledge graph; the fifth intelligent agent module is used to perform dynamic comprehensive analysis on various harmful query results, and output harmful information recognition results, confidence levels and evaluation bases.

[0016] The harmful image detection and recognition system based on intelligent agent collaboration provided by this application, through the hierarchical processing and collaborative work of the first to fifth intelligent agents, realizes the comprehensive analysis of the image to obtain meta-semantic information, and performs coarse-grained and fine-grained harmful semantic detections on the meta-semantic information. Through multiple rounds of queries and dynamic comprehensive analysis, it can accurately identify hidden harmful information, thereby improving the accuracy and interpretability of harmful information recognition.

[0017] Other features and advantages of the present invention will be described in the subsequent specification, and, in part, will be obvious from the specification, or will be understood by implementing the present invention. The objectives and other advantages of the present invention are achieved and obtained by the structures specifically pointed out in the specification, claims and drawings.

[0018] To make the above objects, features, and advantages of the present invention more obvious and understandable, the following provides preferred embodiments in conjunction with the accompanying drawings and describes them in detail as follows. Description of the Drawings

[0019] In order to more clearly illustrate the specific embodiments of the present invention or the technical solutions in the prior art, the following will briefly introduce the drawings required for use in the description of the specific embodiments or the prior art. Obviously, the drawings in the following description are some embodiments of the present invention. For those skilled in the art, without creative efforts, other drawings can also be obtained based on these drawings.

[0020] Figure 1 Schematic diagram of the composition of the harmful image detection and recognition system based on agent collaboration provided by the embodiment of the present invention; Figure 2 Schematic diagram of the structure of the first agent module provided by the embodiment of the present invention; Figure 3 Schematic diagram of the structure of the second agent module provided by the embodiment of the present invention; Figure 4 Schematic diagram of the structure of the third agent module provided by the embodiment of the present invention; Figure 5 Schematic diagram of the structure of the fourth agent module provided by the embodiment of the present invention; Figure 6 Schematic diagram of the structure of the fifth agent module provided by the embodiment of the present invention; Figure 7 Schematic diagram of the harmful image detection and recognition method based on agent collaboration provided by the embodiment of the present invention.

[0021] Reference Signs: 10 - First agent module, 11 - Person identity analysis module, 12 - Person emotion analysis module, 13 - Interaction relationship analysis module, 14 - Scene environment analysis module, 15 - Text information analysis module; 20 - Second agent module, 21 - Harmful element extraction module, 22 - Search and matching module; 30 - Third agent module, 31 - Visual symbol analysis module, 32 - Text metaphor mining module; 40 - Fourth agent module, 41 - Multi-round debate base module, 42 - Questioning analysis module, 43 - Questioning feedback mechanism module, 44 - Evidence chain construction module, 45 - Contradiction analysis module; 50 - Fifth agent module, 51 - Dynamic weight adjustment module, 52 - Risk quantification calculation module, 53 - Collaborative verification decision module. Detailed Embodiments

[0022] To make the objectives, technical solutions, and advantages of the embodiments of the present invention clearer, the technical solutions of the present invention will be clearly and completely described below with reference to the accompanying drawings. Apparently, the described embodiments are some, but not all, of the embodiments of the present invention. All other embodiments obtained by those skilled in the art based on the embodiments of the present invention without creative efforts fall within the scope of protection of the present invention.

[0023] To facilitate the understanding of this embodiment, the technical terms designed in this application will be briefly introduced below.

[0024] Meta-semantics is the description and definition of the specific semantics of an image itself, used to describe the essence, structure, and generation method of semantics. In image analysis, meta-semantic information not only focuses on specific objects (such as people and objects) in the image, but also on the relationships between these objects, the overall meaning, and their expressions in a specific scenario.

[0025] Structured data refers to data that has been processed or formatted and stored in a specific organizational manner, facilitating rapid and efficient reading and processing by computer systems, such as data stored in the form of tables, lists, or knowledge graphs.

[0026] After introducing the technical terms involved in this application, next, the application scenarios and design concepts of the embodiments of this application will be briefly introduced.

[0027] To prevent malicious behaviors such as malicious users sending, there are now various image review technologies, mainly focusing on the recognition of obvious items in the image or the semantic understanding of the overall view, but it is difficult to detect images with complex and implicit semantics.

[0028] Based on this, the embodiments of this application provide a harmful image detection and recognition system and method based on agent collaboration.

[0029] Embodiment 1 This application provides a harmful image detection and recognition system based on agent collaboration, combined with Figure 1 As shown, the system includes: a first agent module 10, a second agent module 20, a third agent module 30, a fourth agent module 40, and a fifth agent module 50 with collaborative functions.

[0030] The first agent module 10 is used to analyze the meta-semantic information of the image, output structured data, and construct an image comprehensive semantic knowledge graph.

[0031] The second agent module 20 is used to detect harmful features and perform associated queries on the image in combination with the structured data, output the coarse-grained harmful semantic detection results and their confidence levels, and dynamically update the coarse-grained harmful semantics to the image comprehensive semantic knowledge graph.

[0032] The third intelligent agent module 30 is used to deeply analyze the structured data and the coarse-grained harmful semantic detection results, output the fine-grained harmful semantic detection results and their confidence levels, and dynamically update the fine-grained harmful semantic detection results to the image comprehensive semantic knowledge graph.

[0033] The fourth intelligent agent module 40 is used to interrogate and verify the fine-grained harmful image semantic detection results, output various types of fine-grained harmful interrogation records, network verification results and confidence levels, and dynamically update the interrogation results to the image comprehensive semantic knowledge graph.

[0034] The fifth intelligent agent module 50 is used to perform dynamic comprehensive analysis by integrating various types of harmful interrogation results, and output harmful information identification results, confidence levels and judgment bases.

[0035] In this embodiment, through the division of labor and cooperation of the first to fifth intelligent agent modules, integrating various technical means such as coarse-grained detection and fine-grained analysis, dynamic knowledge graph update, and multi-round interrogation verification, the whole process from image meta-semantic parsing to comprehensive risk assessment is realized. By using mechanisms such as structured data transmission, cross-modal association, and dynamic weight adjustment, the accuracy and interpretability of harmful content identification are improved. Among them, the judgment basis is used to explain the judgment result.

[0036] Combined with the first aspect, the first intelligent agent module 10 includes: a person identity parsing module 11, a person emotion analysis module 12, an interaction relationship analysis module 13, a scene environment analysis module 14, and a text information parsing module 15, as Figure 2 shown.

[0037] In this embodiment, the first intelligent agent module 10 performs high-precision annotation on the objects (people, objects), scenes and relationships in the image, provides multi-level parsing from objects to scenes to relationships to comprehensively understand the image content, so as to convert complex image information into easy-to-process structured data. Among them, the structured data includes multi-dimensional parsing data. If no parsing data can be parsed for a certain dimension (such as person identity), the storage address of the parsing data corresponding to this dimension is left blank. Through the structured data, users can intuitively understand the specific content obtained by the first intelligent agent's preliminary parsing of the image.

[0038] The person identity parsing module 11 is used to perform preliminary recognition on the input person parsing prompt template and the image based on a large model to output an image text description, extract and parse the features of the people in the image based on a small model when it is determined that there are people in the image to output the person identity parsing result, and is also used to fuse the image text description output by the large model and the person identity parsing result output by the small model as the initial meta-semantic information and construct an image comprehensive semantic knowledge graph; among them, the person identity parsing result includes at least one of group, gender, age, and occupation.

[0039] Among them, the large model in the person identity parsing module 11 is used to perform a preliminary parsing of the image to output an image text description, which includes but is not limited to whether there are people in the image, the number and distribution of people, and the positional relationship between people; there can be multiple small models, such as FaceNet, DeepFace, Dlib, FairFace, InsightFace, etc. In this embodiment, the preset small models can be the InsightFace model and / or the FairFace model. The InsightFace model can be used to accurately identify the facial information of people and extract the identity characteristics of people, while the FairFace model can further refine the description of people, such as labeling the group to which the people belong (the group division in this embodiment is only based on technically measurable characteristics such as genotype, geographical origin, etc.) or a more specific age range. By using the InsightFace model and the FairFace model, accurate identification of the identity characteristics and specific attributes of people can be achieved in the image meta-semantic information parsing, and multi-dimensional person descriptions can be made, such as providing comprehensive person information including position, expression, age, gender, group, appearance characteristics, etc. In the actual application process, the InsightFace model can be placed in the front, first performing preliminary face detection and feature extraction, and further performing attribute analysis on the detected face features through the FairFace model to identify the group, gender, age, etc. of the corresponding person and the confidence level.

[0040] Preferably, before actual application, the FairFace model is iteratively trained on multi-ethnic data sets such as CelebA until fairness correction can be performed to improve the fairness of prediction.

[0041] The following gives an implementable example of the person parsing prompt template: Please carefully analyze the people, occupations, relationships between people, items, actions, environments, atmospheres, emotions, etc. in this picture, and store the analyzed contents item by item in a csv table. If these contents are not involved, fill in "not involved" in the corresponding item.

[0042] When the person emotion analysis module 12 is used to perform a preliminary identification of the image to determine that there are people in the image, based on a preset expression recognition small model, the expression features in the image are extracted and parsed based on the expression features to obtain the person emotion parsing result and the confidence level of the person emotion parsing result, and the person emotion parsing result is used as the first meta-semantic information to update the image comprehensive semantic knowledge graph.

[0043] Among them, when a person is detected in the image, the preset emotion recognition sub-model (in this embodiment, the ResNet-18 model trained based on the FER2013 dataset) is used to extract expression features, and emotion analysis, emotion classification, and confidence calculation are performed based on the expression features. The emotion classification result represents the judgment result of the model on the current emotion state of the person. Before actual application, the model is trained based on the FER2013 dataset. The FER2013 dataset is a very important public dataset in the field of emotion recognition and is widely used for training and testing emotion analysis models. There are approximately 35,887 face images in this dataset, and each image is labeled as one of the preset seven types of emotion labels. Initialize the ResNet-18 model, adjust the input layer to pixel grayscale images to match the dataset specifications, set the output layer dimension to 7 (corresponding to the above classification categories), and train the initial ResNet-18 model to obtain a human emotion recognition model that can accurately recognize human emotions.

[0044] When the interaction relationship analysis module 13 is used to initially identify the image and determine that there are people in the image, it extracts the action features in the image based on the preset action recognition sub-model and combines the human emotion analysis results to predict the interaction category, outputs the interaction relationship analysis result and its confidence, and updates the interaction relationship analysis result as the second meta-semantic information to the image comprehensive semantic knowledge graph.

[0045] Use the preset action recognition sub-model to extract the human action features in the image. These features may include main action description information such as body posture, gestures, and body orientation. Combine the human emotion analysis results output by the human emotion analysis module 12 with the action features, and use machine learning or deep learning models to monitor the interaction behavior parameters. Finally, output the interaction relationship analysis result, describing the interaction category and its possibility between the people in the image. Among them, the interaction relationship between people can be predefined social role pairs such as teacher-student, mother-daughter, or hierarchical relationship role pairs such as boss-employee, etc.

[0046] As an example, when the result of parsing the emotions of the person in the image is that the confidence of the first emotion label is greater than the preset confidence threshold (0.7), and at the same time, considering its action features: the radius of the hand trajectory is greater than the pixel threshold such as 30 pixels (indicating that he is waving), it can be inferred that the result of parsing the interaction relationship is a welcome behavior with a confidence of 80%. The interaction categories can also include: collaborative behavior, opposing behavior, etc. In this embodiment, the interaction behavior parameter feature of the collaborative behavior is that the detected person synchronization completion rate is greater than the threshold, such as 85%; the interaction behavior parameter feature of the opposing behavior is that the included angle of the motion vectors is greater than the threshold of 90° or the relevance of the person's motion is less than 0.5. It can be understood that the determination thresholds of the respective interaction behavior parameter features can be adjusted according to the detection requirements. Here, it is only an example and is not limited.

[0047] Furthermore, a preset action parsing model and a timing detection model are used to extract the action features and timing features in the image, parse the action trajectory, change trend and time series information, and generate a dynamic feature parsing result.

[0048] In this embodiment, an I3D model pre-trained on Kinetics-600 is used to detect multiple types of illegal actions. Among them, the I3D model defines 12 types of illegal action features with an action danger coefficient (value range from 0 to 1) greater than the set threshold to detect abnormal dynamics and potential illegal behaviors.

[0049] The scene environment analysis module 14 performs environmental scene parsing on the image based on a preset scene recognition model and a light detection algorithm, outputs the environmental scene parsing result and its confidence, the light parsing result and its confidence, and updates the environmental scene parsing result and the light parsing result as the third meta-semantic information to the image comprehensive semantic knowledge graph. Among them, the environmental scene parsing result includes but is not limited to indoor environment, natural environment, urban outdoor environment; the light parsing result includes but is not limited to natural light, artificial lighting, light intensity, weather conditions, light and shadow effects, shooting angle, composition method.

[0050] Among them, the overall scene features of the image are recognized through the scene recognition model. For example, it is judged whether the scene type is indoor or outdoor, environmental features such as urban or natural environment, light such as lighting and weather conditions, and environmental features, which helps to understand the background context and emotional atmosphere of the image.

[0051] In this embodiment, an EfficientNet-B3 model trained using the Places365 dataset is used. This model can identify 11 types of sensitive scenes, including indoor (bedroom, bathroom) and outdoor (street, wilderness), etc. The classification results may include but are not limited to: indoor scenes (such as bedroom, bathroom, office, etc.), natural environments (such as forest, grassland, lake, etc.), and urban outdoor environments (such as street, square, building complex, etc.). Subsequently, in combination with the scene classification results, the background context and emotional atmosphere of the image are deeply understood. For example, a bedroom scene may imply privacy, while a wilderness scene may convey emotions of loneliness or adventure.

[0052] Use Illumination-Net to detect the lighting conditions in the image, and judge the light source type through the consistency of shadow directions, such as natural light (sunlight) or artificial lighting (such as a ring fill light). The lighting detection results may include: natural lighting (such as sunlight, moonlight, etc.), artificial lighting (such as incandescent lamp, fluorescent lamp, ring fill light, etc.), lighting effects (such as soft shadow, strong shadow, etc.), and shooting angles (such as backlight, front light, etc.). One feasible example of the scene deconstruction prompt template given below: Please conduct a multi-level analysis of this image, including elements such as scene composition, character state, item distribution, behavior dynamics, environmental atmosphere, group interaction, etc. Organize the analysis results into a csv table by category, and mark items not involved as "not involved". Please pay special attention to the relevance between various elements.

[0053] The text information parsing module 15 uses OCR technology to extract the text information in the image, and uses a preset text recognition method and semantic parsing algorithm to parse the text information in the image to obtain the text semantic information parsing result, and updates the image comprehensive semantic knowledge graph with the text semantic information parsing result.

[0054] In this embodiment, the OCR technology is used to identify the text information in the image, including slogans, poster texts, advertising copy, etc. Based on the text content and font style, it is judged whether the text contains illegal content. In this embodiment, the PaddleOCR v3 technology is adopted to support multi-language text recognition, including vertical Arabic and Bengali deformed fonts. For blurred text, the convolutional recurrent neural network combined with the attention (CRNN + Attention) mechanism, SRGAN or diffusion model is used for restoration to improve the recognition accuracy; the RoBERTa is used to train the sensitive word classifier to identify the preset specific semantic information, and the regular expression is supported to dynamically update the word library to ensure sensitivity to newly emerging sensitive words; the overall color distribution, contrast, saturation and visual style (such as realistic, abstract, comic style, etc.) of the image are analyzed to infer the possible emotions or specific intentions conveyed by the image; 10,000 image samples containing human body regions are collected, and the HSV value range (H: 0-50°, S: 20-100%, V: 15-95%) of the skin color region in each sample is determined through manual annotation to form a skin color feature database, so as to pre-construct the SkinTone-10K dataset and quantify the skin color distribution through the HSV color space; the skin color pixels in the HSV color space of the input image are extracted, and the proportion in the lower half image region is calculated; the illegal judgment is carried out based on the preset conditions: when the total area proportion of the skin color region is greater than 30%, the lower half skin color proportion is less than 40%, and the overlap rate between the highlight region and the skin color region is greater than 60%, it is determined as illegal content; the VGG-19 model is used to extract deep features to detect the style consistency of AI-generated illegal images (such as the deviation of epidermal microstructure, the variation of cutin topology, the disorder characteristics of the dermal-epidermal interface, etc. existing in the images generated by DeepNude). After that, the above recognition results are integrated to generate a comprehensive text semantic information recognition result.

[0055] In this embodiment, the first intelligent agent module 10 performs multi-dimensional meta-semantic parsing and image visual parsing on the image to achieve a comprehensive breakdown of the image content, so as to provide basic data support for subsequent in-depth image recognition.

[0056] Combined with the first aspect, the second intelligent agent module 20 includes: a harmful element extraction module 21 and a search and matching module 22, as Figure 3 shown.

[0057] The harmful element extraction module 21 is used to screen harmful objects from the structured data output by the first intelligent agent module 10 based on a small model, and output structured data of harmful objects; it is also used to screen harmful features based on a large model for the input harmful object detection prompt template and the image comprehensive semantic knowledge graph output by the first intelligent agent module 10, generate harmful feature detection results and confidence levels, and use the harmful feature detection results as the fourth meta-semantic information to update the image comprehensive semantic knowledge graph again.

[0058] In this embodiment, harmful elements include harmful objects and harmful features. Harmful objects include, but are not limited to, safety risk items and special identifiers. Harmful features include, but are not limited to, cross-cultural cognitive difference features, safety risk visual features, abnormal content features, and non-neutral interaction features (including, but not limited to, inappropriate interactions).

[0059] Among them, by identifying the objects and symbols in the image, the main objects, the position distribution and usage status of the objects in the image are identified, so as to screen potential harmful objects.

[0060] As an implementable method, the YOLOv7 model is used to accurately detect harmful objects, and this model has high recognition accuracy; the CutMix strategy is adopted to synthesize the background, and the target object area and the random background image are synthesized into training samples at a mixing ratio of α = 0.4 to generate enhanced data with a size less than 10px × 10px, so as to improve the detection rate of harmful objects with a small volume below 10px, such as metal reflective areas with an aspect ratio greater than 5:1, specified trigger mechanism features, and irregular high-saturation red areas in the image (HSV color space: H ∈ [0,10] ∪ [350,360], S > 90%, V > 60%).

[0061] In addition, a special identifier feature library is pre-constructed. Specifically, a standardized database including 87 types of high-risk semantic special identifiers is constructed. Each type of special identifier contains the following metadata: visual feature vector (2048-dimensional feature extracted by Res-Net-50), multi-scenario semantic labels (such as the semantic labels of the same identifier are different in different scenarios), and risk level classification. The CLIP image-text matching model based on the ViT-B / 32 architecture is used to verify the image-text matching degree between the identifier to be detected and the special identifier feature library, so as to assist in judging its potential risks and adverse effects, and obtain an analysis result. This analysis result may include the following classified and encoded semantic categories: cultural identifier category, organizational identifier category, and commercial identifier category.

[0062] Harmful features include, but are not limited to, cross-cultural cognitive difference features, safety risk behavior features, abnormal content features, and non-neutral interaction features.

[0063] Among them, cross-cultural cognitive difference features include but are not limited to: high-frequency matching features in the preset face database, color gamut features specifying the RGB value range, edge features conforming to the CAD model of OpenStreetMap landmark buildings (error tolerance ±5 pixels), etc.; small models such as CNN convolutional neural network can be used and fused with the attention mechanism (Vision Transformer) for detection. For example, the detection is performed on the person identity parsing result in the structured data output by the first intelligent agent module 10 to detect whether there is such a harmful feature as the high-frequency matching feature in the preset face database.

[0064] For text semantics, the pre-trained language model BERT model can be used to identify relevant sensitive words (such as organizational structure identifiers, function identifiers, role descriptors) for specific domain entities and perform knowledge graph verification to calculate the confidence level. Among them, the knowledge graph is constructed in advance based on publicly verifiable standardized image data sources for the face database.

[0065] Among them, security risk visual features include but are not limited to: visually significant suppression texture (such as the Fourier spectrum energy distribution pattern with a low-frequency ratio < 30%), security risk object features (such as object features conforming to the external contour of security risk objects), security risk behavior features (such as the hyperextension structure of the metacarpophalangeal joint, with the joint angle between 160° and 180°), liquid substances with hue values in the low-wavelength visible spectral band region (corresponding to the HSV-H range of 0° to 20°) and liquid substances with hue values in the wavelength interval of 620 - 750 nm.

[0066] Among them, the illegal content includes but is not limited to: images with the proportion of pixel chromaticity values in the skin color clustering space exceeding the threshold, detected by shutdown motion trajectory recognition and expression recognition technologies; the OpenPose model can be used to detect the proportion of pixel chromaticity values in the skin color clustering space, the ResNet-50 model can be used for facial feature point analysis (such as the distance between eyebrows and eyes, the opening degree of the lips), and the Bi-LSTM model and the CRF model can be used to identify the metaphorical illegal semantic information in the text information recognition result.

[0067] Non-neutral interaction features include but are not limited to: including derogatory expressions, analysis of role right relationships, and the difference in the matching degree between images and text descriptions, etc.; the GCN model is combined with structured data to capture the local and global structural information between nodes in the graph to analyze the role right relationships and compare the matching degree differences between the image scene and the text description, enhancing the understanding and judgment of the image content. The RoBERTa-large model after transfer learning can also be used to detect derogatory expressions.

[0068] The search and matching module 22 is used to generate a search query associated with the harmful element based on the large model according to the extracted harmful element, call the preset external harmful element database for relevance matching, calculate the confidence level of the harmful element as a risk item, and use the first parsing result output as the fifth - element semantic information to update the image comprehensive semantic knowledge graph again. Then, combine the fifth - element semantic information and the fourth - element semantic information to generate a coarse - grained harmful semantic detection result. Based on the pre - configured search query generation rule, according to the detected sensitive element, generate a specific search query. For example, if the detected harmful element is the high - frequency matching feature in the preset face database, the query is "[Person's name] Recent group behavior clustering events". Call the search engine through the API, set search parameters (such as time range, geographical restriction) to ensure the relevance of the results, and obtain the search results. Then, based on natural language processing technology, analyze the search results to extract and supplement key information related to the harmful element, such as text summary carriers, event descriptions, or social media comments. Determine the confidence level based on the relevance between the key information and the harmful element. It can be understood that if the harmful element is a certain harmful object and the search result is a safety - hazard event, and there are a large number of such harmful objects in the safety - hazard event, at this time, increase the confidence level of the harmful element as a risk item. Then, use the classification model with preset rules to process the input fourth - element semantic information and fifth - element semantic information, and output the coarse - grained harmful semantic detection result and the confidence level.

[0069] In this way, after the first intelligent agent module 10 performs multi - dimensional semantic parsing on the image to extract structured data and construct the image comprehensive semantic knowledge graph, input the structured data into the second intelligent agent module 20. Based on the large model, combined with the harmful object detection prompt template, conduct a preliminary screening of violation features, call the retrieval - enhanced system to generate an associated query, and call the preset external harmful element database for relevance matching, output the first parsing result, and use the first parsing result as the fifth - element semantic information to update the image comprehensive semantic knowledge graph again. Combine the fourth - element semantic information and the fifth - element semantic information to generate a coarse - grained harmful semantic detection result.

[0070] Combined with the first aspect, the third intelligent agent module 30 includes: a visual symbol parsing module 31 and a text metaphor mining module 32, as Figure 4 shown.

[0071] The visual symbol parsing module 31 detects harmful symbol information in the image based on the preset visual symbol recognition model, and deeply analyzes the harmful symbol information in combination with the image comprehensive semantic knowledge graph information and the coarse - grained harmful semantic detection result output by the second intelligent agent module 20, and outputs a fine - grained harmful semantic detection result of the first specified category. Among them, visual symbols include, but are not limited to, predefined identifiers with specific cultural meanings, geometric features of clothing patterns, and specific joint movement trajectory features and iconic visual symbols in historical events that do not involve private data. For example, use a symbol recognition model (such as ResNet) to detect geometric configurations with cultural context dependence (such as a centrally symmetric starburst structure), primitive combinations with a matching degree > 90% in a predefined symbol library, and color coding that conforms to the highly salient color group in the Pantone standard; use the attention mechanism branch of ResNeXtt-101 for clothing semantics to detect cultural representative fabric semantics, head covering texture features, the relevance recognition of non-linear patterns in the facial area and non-neutral interaction scenarios. Among them, ResNeXt-101 is a deep convolutional neural network architecture that combines the residual connection of the ResNet model and the multi-branch structure of the Inception model to improve the model's expressive ability by increasing the width (cardinality) of the network; use the MediaPipe Holistic model for transfer learning, so as to be able to detect and track multiple key parts of the person in the image such as the face, hands, and body posture, calculate the dynamic angle change rate between joint points, output semantic labels when detecting specific joint configurations, and generate adaptive feedback signals when detecting predefined risk semantic labels to identify illegal action behaviors. Then, combine the above information detected with the cultural common sense graph to analyze its potential meaning, so as to determine the fine-grained harmful semantic detection result of the first specified category.

[0072] The text metaphor mining module 32 is used to deeply analyze the pun recognition features, iconic visual symbols in historical events, and regional cultural features in the image semantic information based on the large model for the input metaphor analysis prompt template and image, output the fine-grained harmful semantic detection result of the second specified category, and use the fine-grained harmful semantic detection result of the second specified category as the sixth meta-semantic information to update the image comprehensive semantic knowledge graph.

[0073] Among them, the fine-grained harmful semantic detection result of the second specified category includes abnormal puns, sensitive historical event keywords, and sensitive words of non-neutral categories. In this embodiment, the pun recognition model used is the ELECTRA model for feature extraction, matrix construction, and clustering analysis after early warning analysis using the constructed matrix to identify abnormal puns; use the large-scale pre-trained language model ERNIE3.0 for feature extraction, calculate the similarity between keywords and structured data with a defined keyword list (such as "specified agricultural raw materials"), and combine the context to determine whether it involves sensitive historical events; use the XLM-RoBERTa cross-language model for extraction, keyword matching, and context analysis in structured data to detect the degree of association with non-neutral categories, so as to determine sensitive words of non-neutral categories.

[0074] In this way, the third agent module 30 deeply analyzes the structured data output by the first agent module 10 and the coarse-grained harmful semantic detection results output by the second agent module 20, and uses advanced visual symbol parsing and text metaphor mining technologies to detect fine-grained harmful semantic information.

[0075] Combined with the first aspect, the fourth agent module 40 includes: a multi-round debate base module 41, a question analysis module 42, a question feedback mechanism module 43, an evidence chain construction module 44, and a contradiction analysis module 45, as Figure 5 shown.

[0076] When the fitness of the structured data, the coarse-grained harmful semantic detection results, the fine-grained semantic detection results, and the harmful elements is lower than the preset threshold, the multi-round debate base module 41 generates a question statement and distributes it to the first to third agent modules for re-analysis until the fitness is higher than the preset threshold, and determines the debate result, confidence level, and evidence chain information.

[0077] Combined with the output results of the first agent module 10, the second agent module 20, and the third agent module 30, locate the contradictory items (such as the existence of harmful objects, but the safety risk behavior characteristics are not detected), then generate questions based on predefined rules and generate questions through a preset natural language generation model, and then distribute the question statements to the associated agent modules (such as requiring the third agent module 30 to re-analyze cultural symbols). Among them, the contradictory items can be identification subject contradiction, identification behavior contradiction, identification emotion contradiction, identification scene contradiction, etc. For example, the contradictory item is an identification behavior contradiction. Specifically, safety risk behavior characteristics such as "trajectory curvature picture, sudden change in joint movement trajectory" are detected, but "non-linear distortion of acoustic features, strong activation of facial action units such as: based on the combined effect of the frontalis electromyogram signal strength EMG≥20μV and the perioral muscle tension value F≥5N, the emotion analysis result of generating the emotion feature vector V_emo is not detected. The question generated based on the predefined rules may be: The person on the left side of the picture performs a safety risk action. Wouldn't this cause harm to others?

[0078] The question analysis module 42 is used to verify the consistency between the calculated image and the semantic information by using a text generation model based on images, and input the fine-grained harmful semantic detection results of the second specified category and the question prompt template into a large model to question-analyze the comprehensive semantic knowledge graph of the analyzed image, and obtain the fine-grained harmful semantic detection results of the third specified category.

[0079] Among them, the cross-modal association network module uses the CLIP model to verify the consistency of the calculated image and text information, and based on the comprehensive semantic knowledge graph of the image output by the third intelligent agent, it verifies the suitability of the image to obtain the detection result of fine-grained harmful semantic information of the third specified category. In this embodiment, by using the CLIP model to encode the text in the image and structured data, the similarity between the image and the text is calculated based on a preset algorithm (such as the pre-similarity combined with the dynamic threshold adjustment algorithm), and in combination with the constructed association relationships containing a large number of cross-cultural context-specific elements for in-depth recognition, so as to extract the detection result of fine-grained harmful semantics of the third specified category. Among them, the cross-cultural context-specific element association relationship includes the one-to-one correspondence relationship among the subject, relationship, and object. In this way, through multi-modal semantic deconstruction, the structured data output by the first intelligent agent module 10, the detection result of coarse-grained harmful semantic information output by the second intelligent agent module 20, and the detection result of fine-grained harmful semantic information output by the third intelligent agent module 30 are deeply analyzed. By using the advanced visual symbol parsing and text metaphor mining technologies in the third intelligent agent module 30, combined with the CLIP model cross-modal alignment and cultural common sense graph, the fine recognition of identifiers with specific cultural meanings, the geometric features of clothing patterns, the specific joint movement trajectory features that do not involve privacy data, and the iconic visual symbols in historical events in the image is realized.

[0080] The question feedback mechanism module 43 is used to generate question problems and dynamically adjust the weight distribution based on the question results when the fitness of the structured data, the detection result of coarse-grained harmful semantic information, the detection result of fine-grained harmful semantic information of the first specified category, and the detection result of fine-grained harmful semantic information of the second specified category is lower than the preset threshold, so as to dynamically update the comprehensive semantic knowledge graph of the image.

[0081] When the first intelligent agent module 10 or the second intelligent agent module 20 detects a violation and harmful risk, but this violation and harmful risk does not match the detection results of the fine-grained harmful semantics of the first to third specified categories, it triggers the reverse verification of other intelligent agents (specifically referring to the first intelligent agent module 10 and the second intelligent agent module 20), forms a collaborative question mechanism, and dynamically coordinates the verification process through the question feedback mechanism, aiming to solve the contradictions or low-confidence problems in the detection results through reverse questioning and weight adjustment, and improve the accuracy and reliability of risk determination.

[0082] For example, the first intelligent agent module 10 detects a harmful object, but the third intelligent agent module 30 does not detect the opposing behavior characteristics, and the fitness between the two is low. At this time, question problems are generated such as: "Is the harmful object in a safe risk scenario and does the relationship between people need to be re-evaluated?

[0083] After that, the priority of the third agent module 30 is increased, forcing the first agent module 10 and the second agent module 20 to re-analyze, and comparing the results of the re-analysis with the CLIP model for cross-modal alignment (image-text similarity) to verify the consistency and update the global detection results. If there are still contradictions, multi-round questioning or manual review is triggered. Thus, a fine-grained semantic analysis result is generated by parsing the hidden background semantics to dynamically update the image comprehensive semantic knowledge graph.

[0084] The evidence chain construction module 44 is used to output and associate structured data, coarse-grained harmful semantic information detection results, fine-grained harmful semantic information detection results, and harmful element associations, and use the large model to perform online search on the output results to obtain authoritative information, and summarize and form a complete evidence chain composed of scattered evidence points connected in series.

[0085] During the multi-round questioning process, associate the analysis results and generate supplementary search queries when background information is needed, such as: query "[symbol] cultural interpretation". Subsequently, perform an online search, specifically by calling the historical database, cultural database, or news archives through the API interface to obtain authoritative information, and then perform data processing to extract relevant evidence and organize the evidence into a logical chain of evidence to obtain a complete evidence chain.

[0086] The contradiction analysis module 45 is used to process structured data, coarse-grained harmful semantic information detection results, and fine-grained harmful semantic information detection results through multi-dimensional cross-validation using the large model to identify various harmful risk items, and output various fine-grained harmful questioning records, network verification results, and confidence results.

[0087] After obtaining the complete evidence chain, combine the outputs of other agents, compare the relevance, consistency, and logical relationships between multiple pieces of evidence to verify whether there are contradictions. For example, if the background is a peace slogan, but a security risk event or security risk behavior appears, it can be judged that there is a contradiction, and various fine-grained harmful questioning records, network verification results, and confidence levels are output.

[0088] Combined with the first aspect, the fifth agent module 50 includes: a dynamic weight adjustment module 51, a risk quantification calculation module 52, and a collaborative verification decision module 53, as Figure 6 shown.

[0089] The dynamic weight adjustment module 51 is used to calculate the assignment weights of each fine-grained harmful semantic detection result based on a preset dynamic weight algorithm.

[0090] Based on the fine-grained harmful semantic detection results output by the fourth agent module 40, calculate the weights of various fine-grained harmful semantic detection results.

[0091] The risk quantification calculation module 52 is used to perform a comprehensive scoring operation by combining the assignment weights of the fine-grained harmful semantic detection results, the type scores of the fine-grained harmful semantic detection results, and the scene complexity function to obtain the confidence of the fine-grained harmful semantic detection results.

[0092] The collaborative verification decision module 53 is used to start a questioning process for re-evidence collection when the confidence is insufficient until the recognition result of the fine-grained harmful semantic detection result with a confidence meeting the preset requirements is obtained.

[0093] In a second aspect, the present application provides a harmful image detection and recognition method based on multi-agent collaboration, combined with Figure 7 As shown, the method includes: S110, analyzing the image meta-semantic information through the first agent module, outputting structured data, and constructing an image comprehensive semantic knowledge graph.

[0094] S120, detecting harmful elements in the image and performing associated queries on the image by combining the structured data through the second agent module, outputting the coarse-grained harmful semantic detection result and its confidence, and dynamically updating the image comprehensive semantic knowledge graph.

[0095] S130, deeply analyzing the structured data and the coarse-grained harmful semantic detection result through the third agent module, outputting the fine-grained harmful semantic detection result and its confidence, and dynamically updating the image comprehensive semantic knowledge graph.

[0096] S140, interrogating and verifying the coarse-grained harmful image semantic detection result and the fine-grained harmful image semantic detection result through the fourth agent module, outputting various harmful interrogation records, verification evidence chains, network verification results, and confidences, and dynamically updating the image comprehensive semantic knowledge graph.

[0097] S150, dynamically comprehensively analyzing various harmful interrogation results through the fifth agent module, and outputting the harmful information recognition result and the judgment basis.

[0098] The present invention discloses a harmful image detection and recognition system based on agent collaboration. The first agent module 10 performs multi-dimensional content analysis, the second agent module 20 performs harmful element detection, the third agent module 30 performs in-depth semantic analysis, the fourth agent module 40 performs multiple rounds of interrogation and verification, and finally the fifth agent module 50 makes a dynamic weighted determination. In this way, by adopting a multi-agent module division of labor and collaboration mechanism, combining the advantages of large model semantic understanding and small model precise detection associated with each agent module, the accuracy and interpretability of harmful information in images are improved, which is particularly suitable for processing harmful content detection scenarios with complex cultural backgrounds and implicit semantics.

[0099] In combination with the second aspect, the interrogation by the fourth intelligent agent module in step S140 specifically includes: Calculate a certain type of risk score using the following formula: ; ; ; ; ; where, is the risk score of the detection result of the -th type of fine-grained harmful semantic information, , is a natural number greater than 2; and are both weight coefficients, and + ; is the evidence chain integrity score, ; is the multi-round interrogation consistency score, ; is the contradiction coefficient, ; is the network verification reliability score, ; is the total number of expected evidence points; represents the validity of the j -th evidence point, takes a value of 0 or 1; is the result of the k-th round of interrogation, , is a natural number greater than 2; is the number of detected contradictions, Ntotal is the total number of checkpoints; is the weight coefficient, ; is the authoritative source ratio, ; is the information consistency, .

[0100] Among them, E is the evidence chain integrity score, and the evidence points are specifically composed of the following parts: The meta-semantic information parsing result of the first intelligent agent; The harmful element detection result of the second intelligent agent; The fine-grained semantic information analysis result of the third intelligent agent; The verification of authoritative information obtained by network search, where is the jThe validity of each evidence point, where J is the total number of expected evidence points.

[0101] is the contradiction coefficient, mainly used to measure the degree of inconsistency between the detection results of different intelligent agent modules.

[0102] Combined with the second aspect, step S130 includes: If , trigger multiple rounds of cross-examination, generate challenging questions, and distribute them to the first intelligent agent module, the second intelligent agent module, and the third intelligent agent module; If , trigger the re-verification process, and update the evidence chain by searching for supplementary information through the network; If , trigger the re-verification process, and update the evidence chain by searching for supplementary information through the network.

[0103] Among them, is the preset first threshold, is the preset second threshold, and .

[0104] After calculating the comprehensive score S based on the risk quantification calculation module 52, compare it with the preset first threshold and the preset second threshold . In this embodiment, = 0.6, = 0.8. Specifically, if S < 0.6, it is determined that the credibility of the current result is low. At this time, trigger the multi-round counter-question mechanism to generate challenging queries for the specific fine-grained harmful semantic detection results (for example: query "[Specific regulations on cross-cultural context sensitivity elements]"), execute network search, and after processing the search results, re-calculate the new comprehensive score S after passing through the dynamic weight adjustment module 51 and the risk quantification calculation module 52 until it is determined whether it is harmful content; if , it is determined that the credibility of the current result has not reached the preset standard but is close to the preset standard. At this time, trigger the re-verification process, and update the evidence chain by searching for supplementary information through the network to supplement or correct the current result to improve the accuracy of the result; if S≥0.8, it is determined that the credibility of the current result is high, and the current result can be output as the target result.

[0105] In this way, through the division of labor and cooperation among multiple intelligent agent modules, multi-dimensional, coarse-grained, and fine-grained analysis of image content is realized to improve the detection accuracy.

[0106] Combined with the second aspect, step S150 dynamically comprehensively analyzes various harmful cross-examination results through the fifth intelligent agent module 50, specifically including: Calculate the comprehensive risk score using the following formula: ; ; Among them, is the comprehensive score, is the weight of the detection result of the th type of harmful semantic information, is the category score corresponding to the detection result of the th type of harmful semantic information, , is a natural number greater than 2, is the attenuation factor; is the distance between the detection result of the th type of harmful semantic information and the context; is the implicit semantic weight; is the scenario complexity function, 0. ; is the number of items in the image, .

[0107] In the description of the present invention, the terms "first", "second", "third" are only used for descriptive purposes and cannot be construed as indicating or implying relative importance.

[0108] Finally, it should be noted that: the above embodiments are only specific embodiments of the present invention, which are used to illustrate the technical solutions of the present invention, rather than to limit it. The protection scope of the present invention is not limited thereto. Although the present invention has been described in detail with reference to the foregoing embodiments, those skilled in the art should understand that: any person skilled in the art within the technical scope disclosed by the present invention can still modify the technical solutions recorded in the foregoing embodiments, or can easily think of changes, or make equivalent replacements for some of the technical features; and these modifications, changes or replacements do not make the essence of the corresponding technical solutions deviate from the spirit and scope of the technical solutions of the embodiments of the present invention, and should all be covered within the protection scope of the present invention. Therefore, the protection scope of the present invention should be subject to the protection scope of the claims.

Claims

1. A harmful image detection and recognition system based on intelligent agent collaboration, characterized in that: include: Functionally coordinated first to fifth agent modules; The first agent module is used to analyze the image meta-semantic information, output structured data, and construct a comprehensive semantic knowledge graph of the image; The second agent module is used to perform harmful feature detection and association query on the image in combination with the structured data, output a coarse-grained harmful semantic detection result and its confidence, and dynamically update the coarse-grained harmful semantic detection result to the image comprehensive semantic knowledge graph; The third agent module is used to perform in-depth analysis on the structured data and the coarse-grained harmful semantic detection results, output fine-grained harmful semantic detection results and their confidence, and dynamically update the fine-grained harmful semantic detection results to the image comprehensive semantic knowledge graph; The fourth agent module is used to question and verify the fine-grained harmful image semantic detection results, output various fine-grained harmful query records, network verification results and confidence levels, and dynamically update the query results to the image comprehensive semantic knowledge graph; The fifth intelligent agent module is used to conduct dynamic comprehensive analysis on various harmful inquiry results, and output harmful information identification results, confidence levels and judgment basis.

2. The system according to claim 1, characterized in that The first agent module comprises: The character identity analysis module is used to perform preliminary recognition of the input character analysis prompt template and image based on the large model and output the image text description; when it is determined that there is a person in the image, the character in the image is extracted based on the small model, and the character identity analysis result is output; and it is also used to fuse the image text description output by the large model and the character identity analysis result output by the small model as initial meta-semantic information and construct an image comprehensive semantic knowledge graph; wherein the character identity analysis result includes at least one of group, gender, age, and occupation; A character emotion analysis module is used to perform preliminary recognition on the image to determine that there is a character in the image, extract the expression features in the image based on a preset expression recognition model, and perform analysis based on the expression features to obtain the character emotion analysis result and the confidence of the character emotion analysis result, and update the character emotion analysis result as the first element semantic information to the image comprehensive semantic knowledge graph; An interactive relationship analysis module is used to perform preliminary recognition on the image to determine that there is a person in the image, extract the action features in the image based on a preset action recognition model, and predict the interaction category in combination with the character emotion analysis result, output the interactive relationship analysis result and its confidence, and update the interactive relationship analysis result as the second-element semantic information to the image comprehensive semantic knowledge graph; A scene environment analysis module performs environment scene analysis on the image based on a preset scene recognition model and a lighting detection algorithm, outputs an environment scene analysis result and its confidence level, a lighting analysis result and its confidence level, and updates the environment scene analysis result and the lighting analysis result as third-dimensional semantic information to the image comprehensive semantic knowledge graph; The text information analysis module uses OCR technology to extract the text information in the image, uses a preset text recognition method and semantic analysis algorithm to analyze the text information in the image, obtains a text semantic information analysis result, and uses the text semantic information analysis result as the fourth element semantic information to update the image comprehensive semantic knowledge graph.

3. The system according to claim 2, characterized in that: The second agent module comprises: The harmful element extraction module is used to screen harmful objects based on the small model for the structured data input by the first intelligent agent module and output harmful object structured data; it is also used to screen harmful features based on the large model for the harmful object detection prompt template input and the image comprehensive semantic knowledge graph output by the first intelligent agent module, generate harmful feature detection results and confidence, and update the image comprehensive semantic knowledge graph again using the harmful feature detection results as the fourth element semantic information; A search and matching module is used to generate a search query associated with the harmful element based on the extracted harmful element based on the large model, and call a preset external harmful element database for correlation matching, calculate the confidence of the harmful element as a risk item, and output the first parsing result as the fifth-element semantic information to update the image comprehensive semantic knowledge graph again, and combine the fifth-element semantic information and the fourth-element semantic information to generate a coarse-grained harmful semantic detection result.

4. The system according to claim 1, characterized in that The third agent module comprises: A visual symbol parsing module detects harmful symbol information in the image based on a preset visual symbol recognition model, performs in-depth analysis based on the image comprehensive semantic knowledge graph information output by the second agent module and the coarse-grained harmful semantic detection result, and outputs a fine-grained harmful semantic detection result of a first specified category; The text metaphor mining module is used to perform in-depth analysis of the pun recognition features, historical allusion features and regional cultural features in the image semantic information of the input metaphor analysis prompt template and the image based on the large model, output the fine-grained harmful semantic detection results of the second specified category, and use the fine-grained harmful semantic detection results of the second specified category as the sixth meta-semantic information to update the image comprehensive semantic knowledge graph.

5. The system according to claim 4, characterized in that The fourth agent module comprises: The multi-round debate base module generates a question statement and distributes it to the first to third agent modules for re-analysis until the fitness is higher than the preset threshold, and determines the debate result, confidence and evidence chain information when the fitness of the structured data, the coarse-grained harmful semantic detection result, the fine-grained harmful semantic detection result and the harmful element is lower than the preset threshold; A query analysis module is used to verify the consistency between the calculated image and the semantic information by using a graph-to-text model, and input the fine-grained harmful semantic detection result of the second specified category and the query prompt template into the large model to perform query analysis on the analyzed image comprehensive semantic knowledge graph to obtain the fine-grained harmful semantic detection result of the third specified category; A questioning feedback mechanism module is used to generate a questioning question and dynamically improve and adjust the weight distribution based on the questioning result when the adaptability of the structured data, the coarse-grained harmful semantic information detection result, the fine-grained harmful semantic information detection result of the first specified category, and the fine-grained harmful semantic information detection result of the second specified category is lower than a preset threshold, so as to dynamically update the image comprehensive semantic knowledge graph; An evidence chain construction module is used to associate and output the coarse-grained harmful semantic detection results, the fine-grained harmful semantic detection results and harmful elements, and use a large model to search the output results online to obtain authoritative information, and summarize to form a complete evidence chain composed of scattered evidence points in series; The contradiction analysis module is used to process the structured data, the coarse-grained harmful semantic detection results and the fine-grained harmful semantic detection results through multi-dimensional cross-validation using a large model to identify multiple harmful risk items.

6. The system according to claim 5, characterized in that The fifth agent module includes: A dynamic weight adjustment module, used to calculate the assignment weights of various types of harmful semantic detection results based on a preset dynamic weight algorithm; A risk quantification calculation module, which is used to perform a comprehensive scoring operation based on the assigned weight of the fine-grained harmful semantic detection result, the type score of the fine-grained harmful semantic detection result, and the scenario complexity function to obtain the confidence of the fine-grained harmful semantic information detection result; The collaborative verification decision module is used to initiate a reverse questioning process to collect evidence again when the confidence level is insufficient, until a harmful information identification result with a confidence level that meets preset requirements is obtained.

7. A harmful image detection and recognition method based on multi-agent collaboration, characterized in that: Applied to the system according to any one of claims 1 to 6, the method comprises: Analyze the image meta-semantic information through the first agent module, output structured data and construct a comprehensive semantic knowledge graph of the image; Perform harmful element detection and association query on the image in combination with the structured data through the second agent module, output coarse-grained harmful semantic detection results and their confidence, and dynamically update the image comprehensive semantic knowledge graph; Performing in-depth analysis on the structured data and the coarse-grained harmful semantic detection results through a third agent module, outputting fine-grained harmful semantic detection results and their confidences, and dynamically updating the image comprehensive semantic knowledge graph; The fourth agent module questions and verifies the coarse-grained harmful image semantic detection results and the fine-grained harmful image semantic detection results, outputs various harmful query records, verification evidence chains, network verification results and confidence levels, and dynamically updates the image comprehensive semantic knowledge graph; The fifth intelligent agent module conducts dynamic comprehensive analysis on various harmful inquiry results and outputs harmful information identification results and judgment basis.

8. The method according to claim 7, characterized in that The steps for questioning by the fourth agent module include: The risk score of a certain category is calculated using the following formula: ; ; ; ; ; in, For the Risk scores for fine-grained harmful semantic information detection results; and are weight coefficients, and + ; Score the chain of evidence integrity. ; is the consistency score of multiple rounds of questioning, ; is the contradiction coefficient, ; is the network verification reliability score, ; is the total number of evidence points; To indicate the j The validity of each point of evidence The value of is 0 or 1; is the result of the k-th round of questioning, , is a natural number greater than 2; is the number of contradictions detected, Ntotal is the total number of checkpoints; is the weight coefficient, ; is the authoritative source ratio score, ; is the information consistency score, .

9. The method according to claim 8, characterized in that After the fourth agent module is used for the query step, it includes: like , triggering multiple rounds of questioning and generating questioning questions, which are then distributed to the first to third agent modules, and then the weight distribution is dynamically improved and adjusted based on the questioning results; like , triggering the re-verification process to update the chain of evidence by searching for supplementary information online; like , determine that the current result is credible and take the current result as the target result.

10. The method according to claim 7, characterized in that The steps for dynamically calculating the comprehensive risk score by the fifth agent module include: The comprehensive risk score is calculated using the following formula: ; ; in, To score the overall harmful risk; For the The weight of the harmful semantic detection results; For the Risk scores for fine-grained harmful semantic information detection results; = , is a natural number greater than 2; is the attenuation factor; is the implicit semantic weight; is the scene complexity function, 0. ; is the number of objects and people in the image, .

Citation Information

Patent Citations

  • Image recognition method and system based on agent map and readable storage medium

    CN114049493A

  • Semantic-based big data analysis system and method

    CN117312499A

  • Risk prediction method and apparatus, and device and storage medium

    WO2023065545A1

Cited By

  • Chinese webpage internationalization adaptation method and device based on large model and medium

    CN120315799A

  • Multi-agent automatic picture retouching system based on content analysis

    CN120612259A

  • Intelligent AI image real-time processing method and system for edge computing scene

    CN121660868A

  • Intelligent ai image real-time processing method and system for edge computing scenarios

    CN121660868B

  • Multimodal-based guardrail device and method for controlling inappropriate content in generative ai

    KR103000672B1