Safety behavior regulation matching method and system based on large model and knowledge graph

By constructing a multimodal large model and a knowledge graph-based method for matching security behavior regulations, this approach addresses the issues of missing compliance associations, insufficient scenario generalization capabilities, and model reliability in existing technologies. It enables automatic matching of security behaviors and compliance tracing, thereby improving recognition accuracy and management efficiency.

CN121682318APending Publication Date: 2026-03-17BEIJING UNIV OF TECH
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202511890821.1
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-12-15
Publication Date
2026-03-17

AI Technical Summary

Technical Problem

Existing security behavior recognition technologies suffer from several drawbacks in high-risk areas, including a lack of compliance correlation, insufficient scenario generalization ability, defects in dynamic risk assessment, and model reliability issues. These problems result in low efficiency and high false positive rates in compliance tracing.

Method used

We adopt a security behavior regulation matching method based on large models and knowledge graphs. By constructing multimodal large models and multimodal knowledge graphs, we can achieve automatic matching and compliance traceability of security behaviors. We combine security domain standards to construct multimodal knowledge graphs, use multimodal large models to process security specification manuals, and set up security scenario thinking chain templates to achieve full-process automation from behavior recognition to regulation matching.

Benefits of technology

It significantly improves the accuracy of identifying unsafe behaviors in complex scenarios, ensures the accuracy, traceability, and compliance of output results, shortens compliance traceability time, improves management efficiency, reduces the false judgment rate, and realizes full-process automation of security management.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121682318A_ABST
    Figure CN121682318A_ABST
Patent Text Reader

Abstract

The invention provides a safety behavior regulation matching method and system based on a large model and a knowledge graph, and relates to the technical field of safety intelligent management, and the method comprises the steps: enabling a multi-modal large model to receive a field image, and generating a behavior description and a candidate regulation direction; the multi-modal large model extracts visual features from a field image based on a thinking chain template, retrieves a multi-modal knowledge graph based on the visual features and behavior description, and screens a candidate term range based on the visual features and candidate term directions; the multi-modal large model screens candidate regulations from a relational database on the basis of the direction of the candidate regulations, converts the candidate regulations into multi-modal semantic vectors on the basis of visual features and behavior description, and screens the candidate regulations from a vector database; the multi-modal large model is based on a multi-modal knowledge graph, and the most matched security behavior regulation and image-text-regulation cross-modal interpretation are output. The method breaks through the limitation of single-mode recognition, improves the recognition precision of complex scenes, and achieves the cross-mode recognition of safety behaviors and the precise matching of regulations.
Need to check novelty before this filing date? Find Prior Art

Description

TECHNICAL FIELD

[0001] The present application relates to the technical field of safe intelligent management, and particularly to a safe behavior regulation matching method and system based on a large model and a knowledge graph. BACKGROUND

[0002] In high-risk fields such as construction and maritime operations, unsafe behavior recognition technology has evolved through three stages: (1) Single-modality sensor monitoring stage: In the early stage, local risks were monitored by single-modality devices such as infrared thermal imaging and vibration sensors, but there were problems such as high false positive rate and inability to identify complex behaviors. (2) Multi-source data fusion stage: (such as patent CN202510955079) proposed a multi-spectral sensor network based on 5G network, which fused environmental parameters (light, dust) and human joint motion trajectories, analyzed mechanical characteristics through a spatio-temporal graph convolution network (ST-GCN), and dynamically allocated risk weights combined with the FAHP algorithm. Although this system realizes a "collection-processing-analysis-decision" closed loop, it only outputs a risk level report and does not establish an association mechanism with safety regulations. (3) Vertical scene depth adaptation stage: (such as patent CN202510658825) for track construction scenes, an improved YOLOv8 algorithm is used to detect sitting and lying personnel, combined with a Segformer segmentation model and PCA principal component analysis to suppress false positives. Although this type of technology improves the recognition accuracy in specific scenarios, it relies on fixed hardware configurations (such as edge boxes), and when migrating to construction / maritime scenarios, the algorithm architecture needs to be reconfigured, and there is a lack of compliance analysis function.

[0003] Despite the progress made by the prior art, there are still the following key problems: (1) lack of compliance association: existing systems (such as patent CN202510955079) only divide risk levels (such as "high risk" and "medium risk"), and do not automatically match specific safety clauses (such as Article 5.1.3 of the "Safety Technology Code for High-altitude Operations in Building Construction" JGJ80-2016). Safety management personnel need to manually consult paper manuals, resulting in low efficiency of compliance tracing and an increase of 40% in average single event processing time. (4) Insufficient scene generalization capability: vertical scene detection systems (such as patent CN202510658825) are designed for track sitting behavior, and need to retrain the detection model (such as adjusting the YOLO backbone network parameters) when migrating to building construction, and cannot handle multi-modal interference (such as the contour of the human body being blurred under strong light, causing the false detection rate to rise to 18%). (3) Defects in dynamic risk assessment: risk assessment systems based on static rules (such as the FAHP algorithm) cannot adapt to scene differences. For example, the risk weight difference between 3m scaffolding operation and 5m high-altitude operation is not dynamically adjusted, resulting in a false positive rate of up to 22%. (4) Model reliability issues: general multi-modal large models (such as CLIP) are prone to hallucination in the safety field (such as misjudging "double-hook safety belts" as compliant), and the reasoning process lacks transparency, making it impossible to provide an auditable path for regulation matching, resulting in more than 90% of output results failing to pass safety audits.

[0004] In view of the above defects, it is urgent to provide a safety behavior regulation matching method to automatically match safety behavior regulations, realize the full-process automation of "behavior recognition → clause matching → rectification suggestion", shorten the compliance tracing time, and improve the management efficiency. SUMMARY

[0005] In view of the problems in the background art, the present application provides a safety behavior regulation matching method and system based on large models and knowledge graphs, which realizes the upgrade of safety management from "single recognition" to "precise recognition + compliance tracing".

[0006] To achieve the above-mentioned purpose, the present application provides a safety behavior regulation matching method based on large models and knowledge graphs, comprising: constructing a safety field multi-modal data set including images, texts and regulation labels based on safety scenes; Constructing a multi-modal large model, training the multi-modal large model based on the multi-modal data set, and optimizing the multi-modal large model based on an image-text alignment loss + regulation association loss objective function; Based on the multi-modal large model and the safety field specification, a multi-modal knowledge graph is constructed, the categories in the multi-modal knowledge graph include personnel, behavior, scene and regulation, each category is associated with multiple multi-modal attributes, and the defined relationships in the knowledge graph include execution, occurrence, association, inclusion and adaptation; constructing a relational database and a vector database based on the multimodal large model processing the safety regulation manual, the relational database storing structured information of multimodal regulation items, and the vector database storing multimodal semantic vectors of regulation original texts and applicable scene visual descriptions; setting a safety scene thinking chain template based on the multimodal knowledge graph; the multimodal large model receiving a scene image, generating a behavior description and a candidate regulation direction; the multimodal large model extracting visual features from the scene image based on the thinking chain template, retrieving the multimodal knowledge graph based on the visual features and the behavior description, and screening a candidate regulation range based on the visual features and the candidate regulation direction; the multimodal large model screening a candidate regulation from the relational database based on the candidate regulation direction, and converting the visual features and the behavior description into a multimodal semantic vector to screen a candidate regulation from the vector database; the multimodal large model outputting a most matched safety behavior regulation and a multimodal explanation of image-text-regulation based on the multimodal knowledge graph.

[0007] As a further improvement of the present application, the multimodal large model comprises: using an LLaVA-1.6 model as a basic multimodal large model, using a LoRA light-weight fine-tuning strategy, freezing the main parameters of the LLaVA-1.6 model, and training a low-rank matrix based on the multimodal data set; setting a fine-tuning target function as an image-text alignment loss + regulation association loss to obtain a fine-tuned multimodal large model.

[0008] As a further improvement of the present application, a multimodal knowledge graph is constructed based on the multimodal large model combined with safety field specifications, comprising: a multimodal ontology framework is constructed by the fine-tuned multimodal large model combined with safety field specifications, and the multimodal knowledge graph comprising categories of personnel, behavior, scene, regulation, and relationships of execution, occurrence, association, inclusion, and adaptation is formed based on the multimodal ontology framework and the multimodal data set.

[0009] As a further improvement of the present application, the safety regulation manual is processed based on the multimodal large model to construct a relational database and a vector database, comprising: text and tables in the safety regulation manual are extracted by the multimodal large model to generate typical image descriptions of regulation applicable scenes, and multimodal regulation items of text + visual descriptions are formed; structured information of the multimodal regulation items is stored to form a relational database, and fields include regulation unique identification, field, associated behavior type, regulation original text, applicable scene visual description, and applicable scene; All the regulations and their applicable scenarios are visually described and converted into multimodal semantic vectors to obtain a vector database. The fields include the unique identifier of the regulation and the multimodal semantic vector.

[0010] As a further improvement of the present invention, a security scenario thinking chain template is set based on the multimodal knowledge graph, including: Based on the logic of entities and relationships in multimodal knowledge graphs, a security scenario thinking chain template is set up. The logic of the thought chain template is as follows: first, visual features are extracted from the input image, and the core behavioral entities and scene entities are identified by combining the behavioral descriptions output by the multimodal big model; then, the association relationships of the behavioral entities are retrieved based on the multimodal knowledge graph, and candidate rules that fit the scene entities are filtered; finally, the matching range of the rules is obtained by combining the direction of the candidate rules output by the multimodal big model.

[0011] As a further improvement of the present invention, the scene entity includes scene category + details; The candidate rule directions output by the multimodal large model serve as the initial rule matching range; The matching range of the preliminary regulations is verified based on the details of the scenario. The consistency between the details of the scenario and the applicable conditions of the regulations is verified to obtain the matching range of the regulations.

[0012] As a further improvement of the present invention, the multimodal large model receives on-site images and generates behavioral descriptions and candidate rule directions; including: The multimodal big model retrieves similar behavior-scene data from the multimodal knowledge graph based on the scene images, and outputs structured behavior description text based on the similar behavior-scene data, including behavior category, subject, and scene features; The multimodal large model, based on built-in lightweight domain knowledge, outputs multiple candidate rule directions.

[0013] As a further improvement of the present invention, the multimodal large model extracts visual features from the scene image based on the thought chain template, retrieves the multimodal knowledge graph based on the visual features and the behavioral description, and filters the candidate rule range based on the visual features and the candidate rule direction; including: The thought chain template is input into the multimodal large model. The thought chain template guides the multimodal large model to process according to the distributed reasoning logic of the thought chain template and outputs the results.

[0014] As a further improvement of the present invention, the multimodal large model filters candidate rules from the relational database based on the candidate rule direction, and filters candidate rules from the vector database based on the visual features and behavioral descriptions converted into multimodal semantic vectors; including: The multimodal large model filters candidate clauses from the relational database based on behavior type matching and scene visual description matching; The multimodal big model converts image features and behavioral descriptions into multimodal semantic vectors, calculates cosine similarity with the multimodal semantic vectors of candidate rules in the vector database, and filters rules with a similarity greater than a preset threshold. As a further improvement of the present invention, the multimodal large model, based on the multimodal knowledge graph, outputs the most matching safety behavior regulations and cross-modal interpretations of image-text-regulation; including: The multimodal big model combines the rule-scene adaptation relationship in the multimodal knowledge graph to verify the applicability of candidate rules, calculates the final score of semantic similarity + scene adaptation, and outputs the rule with the highest score and the cross-modal interpretation of image-text-rule.

[0015] This invention also provides a safety behavior regulation matching system based on large models and knowledge graphs, including an input layer, a core layer, a support layer, and an output layer; The input layer is used for: Input on-site images, video streams, and safety regulations manuals into the multimodal large model; The core layer is used for: The multimodal large model receives on-site images and generates behavioral descriptions and candidate rule directions; The multimodal big model extracts visual features from the scene images based on the thought chain template, retrieves the multimodal knowledge graph based on the visual features and the behavior description, and filters the range of candidate rules based on the visual features and the candidate rule direction. The multimodal large model filters candidate rules from the relational database based on the candidate rule direction, and filters candidate rules from the vector database based on the visual features and behavioral descriptions converted into multimodal semantic vectors. The support layer is used for: A multimodal dataset for the security domain, including images, text, and rule labels, is constructed based on security scenarios. A multimodal large model is constructed, trained on a multimodal dataset, and optimized using an objective function that combines image-text alignment loss and rule association loss. Based on a multimodal big model combined with security domain standards, a multimodal knowledge graph is constructed. The categories in the multimodal knowledge graph include people, behaviors, scenarios and regulations. Each category is associated with multiple multimodal attributes. The relationships defined in the knowledge graph include execution, occurrence, association, inclusion and adaptation. Based on the aforementioned multimodal large model processing safety specification manual, a relational database and a vector database are constructed. The relational database stores the structured information of the multimodal regulation entries, and the vector database stores the original text of the regulations and multimodal semantic vectors of visual descriptions of applicable scenarios. A security scenario thinking chain template is set based on the aforementioned multimodal knowledge graph; The output layer is used for: The multimodal big model, based on the multimodal knowledge graph, outputs the most matching safety behavior regulations and cross-modal interpretations of image-text-regulation.

[0016] Compared with the prior art, the beneficial effects of the present invention are as follows: This invention overcomes the limitations of single-modal recognition and improves the recognition accuracy in complex scenes: Addressing the shortcomings of traditional systems that rely on single-modal visual models, have low recognition accuracy in complex scenes such as low light and personnel occlusion, and cannot resolve the issue of "synonymous but different words" (e.g., "safety rope" and "double-hook safety belt") through rule matching, this invention achieves cross-modal collaborative recognition of "image visual features - text behavioral description" through a domain-adapted multimodal large-scale model agent, significantly improving the recognition accuracy of unsafe behaviors in complex scenes.

[0017] This invention addresses the issues of general model adaptation and output reliability: For general multimodal large models in the security field, which suffer from poor domain adaptation and are prone to producing illusory outputs (such as fictitious regulations), this invention combines a security domain knowledge graph and a thought chain guidance mechanism. The "behavior-scenario-regulation" association logic in the graph constrains the model's reasoning process, ensuring that the behavioral descriptions output by the large model match the regulations accurately and traceably, thus meeting the compliance requirements of security management.

[0018] This invention constructs a collaborative closed-loop architecture that balances accuracy and interpretability: addressing the shortcomings of existing systems that "can only identify behavior without regulatory basis" or "have a black box reasoning process," it constructs a collaborative architecture of "multimodal large-scale model Agent + knowledge graph," forming a complete closed loop of "cross-modal behavior recognition → thinking chain reasoning to filter regulations → accurate matching to output basis," providing a smart solution for security management that combines recognition accuracy, compliance basis, and reasoning interpretability.

[0019] This invention proposes a collaborative architecture of "multimodal large-scale model Agent + knowledge graph," which achieves a compliance revolution: It automatically links 23 laws and regulations, including the "Production Safety Law" and the "Regulations on the Administration of Safety Production in Construction Projects," through a knowledge graph chain of thought, achieving full automation of the "behavior identification → clause matching → rectification suggestions" process, reducing compliance traceability time to within 5 minutes; it achieves breakthroughs in scenario universality: employing domain-adaptive multimodal large-scale models (such as ViLBERT fine-tuned based on construction / maritime data), the average recognition accuracy reaches 92.7% in 15 scenarios (high-altitude operations, hoisting operations, etc.), a 37% improvement over traditional methods; and it implements dynamic risk management: introducing a reinforcement learning mechanism to adjust the risk assessment model weights in real time based on environmental parameters (wind speed, humidity), reducing the misjudgment rate of high-risk operations to below 5%.

[0020] This invention marks a leap in safety management technology from "experience-driven" to "data-knowledge dual-driven." In terms of management, it constructs a closed loop of "identification-tracing-rectification," improving the efficiency of the PDCA cycle by 60%. In terms of economics, it is estimated that it can reduce losses due to work stoppages caused by violations by approximately RMB 1.2 million per project (taking a 100,000㎡ construction site as an example). In terms of society, it helps achieve the goal of reducing the rate of major accidents by 25%. Attached Figure Description

[0021] Figure 1 This is a schematic diagram of a multimodal large model fine-tuning process disclosed in an embodiment of the present invention; Figure 2 This is a diagram of a safety behavior rule matching system based on a large model and knowledge graph, disclosed in one embodiment of the present invention. Figure 3 This is a flowchart of a safety behavior regulation matching method based on a large model and knowledge graph, disclosed in one embodiment of the present invention. Figure 4 This is a flowchart of thought chain-guided cross-modal reasoning disclosed in one embodiment of the present invention. Detailed Implementation

[0022] To make the objectives, technical solutions, and advantages of the embodiments of the present invention clearer, the technical solutions of the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of the present invention, not all embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of the present invention.

[0023] The present invention will now be described in further detail with reference to the accompanying drawings: like Figure 3As shown, the safety behavior rule matching method based on large models and knowledge graphs provided by this invention includes: S1. Construct a multimodal dataset for the security domain, including images, text, and rule labels, based on security scenarios; For example, based on safety scenarios such as construction and maritime affairs, a domain-specific multimodal dataset of "image + text + regulation label" is constructed. The images come from on-site monitoring and public safety datasets (such as the crew unsafe behavior feature dataset), and the text contains "behavioral description + scenario features + regulation basis". A total of 1,000 high-quality samples were collected and labeled (800 of which were used for training and 200 for validation).

[0024] S2. Construct a multimodal large model. Train the multimodal large model based on the multimodal dataset, and optimize the multimodal large model based on the objective function of image-text alignment loss + rule association loss. in, like Figure 1 As shown, the LLaVA-1.6 (7B parameters) model is used as the basic multimodal large model because it supports cross-modal understanding of images and text and is open source and fine-tunable. The LoRA lightweight fine-tuning strategy is adopted to freeze the backbone parameters of the LLaVA-1.6 model and train a low-rank matrix (rank of 8) based on the multimodal dataset to reduce the computational cost. The fine-tuning objective function is set as image-text alignment loss + rule association loss. The image-text alignment loss ensures that the model can accurately describe the safety behavior in the image, and the rule association loss constrains the model output of the initial association clue between the behavior and the rule (such as "not wearing a safety rope → associated with JGJ59-2021"), thus obtaining the fine-tuned multimodal large model.

[0025] Furthermore, The finely tuned multimodal large model agent has three core functions. (1) Cross-modal recognition: Input on-site images and output structured “behavior description text” (including behavior category, subject, scene features, such as “construction workers on 3m scaffolding without safety ropes”). (2) Preliminary regulation association: Based on the built-in lightweight domain knowledge, output 1-3 candidate regulation directions (such as "JGJ59-2021, JGJ80-2016"). (3) Interactive feedback: Supports receiving “mind chain prompts” (such as “the impact of supplementary scene height on rule matching”) to optimize output results.

[0026] S3. Based on the multimodal large model and combined with safety domain specifications, construct a multimodal knowledge graph. The categories in the multimodal knowledge graph include personnel, behavior, scene, and regulations. Each category is associated with multiple multimodal attributes (for example, the behavior category is associated with "visual feature description" and "typical image URL", and the regulation category is associated with "original text of the clause" and "example image of the applicable scene"). The defined relationships in the knowledge graph include execution, occur in, association, inclusion, and adaptation; Among them, Construct a multimodal ontology framework by combining the fine-tuned multimodal large model with safety domain specifications, and form a multimodal knowledge graph including categories of personnel, behavior, scene, and regulations, and relationships of execution, occur in, association, inclusion, and adaptation based on the multimodal ontology framework and multimodal datasets.

[0027] Furthermore, Examples of various relationships: "execution" (personnel - behavior, such as <construction worker, executes, not wearing a safety rope>), "occur in" (behavior - scene, such as <not wearing a safety rope, occurs in, scaffolding scene>), "association" (behavior - regulation, such as <not wearing a safety rope, associates with, JGJ59 - 2021>), "inclusion" (scene - sub - scene, such as <scaffolding scene, includes, 3m scaffolding>), "adaptation" (regulation - scene, such as <JGJ59 - 2021 - 3.2.1, adapts to, ≥2m high - altitude scene>).

[0028] S4. Based on the multimodal large model to process the safety specification manual, construct a relational database and a vector database. The relational database stores the structured information of multimodal regulation entries, and the vector database stores the multimodal semantic vectors of the original text of the regulations and the visual descriptions of the applicable scenes; Among them, Extract the text and tables in the safety specification manual (PDF manual) through the OCR ability of the multimodal large model, generate typical image descriptions of the applicable scenes of the regulations (such as "Image description of the applicable scene of JGJ59 - 2021 - 3.2.1: Construction workers are working on a scaffolding above 2m, wearing double - hook safety belts"), and form multimodal regulation entries of text + visual description; Store the structured information of multimodal regulation entries to form a relational database. The fields include regulation_id (unique identifier of the regulation), domain (field), behavior_type (associated behavior type), content (original text of the regulation), visual_desc (visual description of the applicable scene), applicable_scene (applicable scene), and support precise filtering based on "behavior type + scene keywords"; Using a domain-fine-tuned Sentence-BERT model, the "content+visual_desc" of a regulation is converted into a 768-dimensional multimodal semantic vector. The stored fields include regulation_id (a unique identifier for the regulation) and multi_modal_vector (a multimodal semantic vector), resulting in a vector database that supports cross-modal semantic similarity calculation.

[0029] S5. Set up a security scenario thinking chain template based on multimodal knowledge graph; in, Based on the logic of entities and relationships in multimodal knowledge graphs, a security scenario thinking chain template is set up. The logic of the mind chain template is as follows: First, extract visual features from the input image, and combine the behavior description output by the multimodal big model to confirm the core behavior entities and scene entities; then, retrieve the association relationship of the behavior entities based on the multimodal knowledge graph, and filter the candidate rules that match the scene entities; finally, combine the direction of the candidate rules output by the multimodal big model to obtain the rule matching range.

[0030] Specifically, Based on the "entity-relationship" logic of multimodal knowledge graphs, a mind chain template specifically designed for security scenarios is presented, comprising three core steps, such as... Figure 4 As shown: Step 1 (Cross-modal feature extraction): "Extract visual features from the input image: [describe the people, behaviors, and scene details in the image, such as 'construction workers are on a 3m high scaffold with no safety rope restraint on their upper bodies']; combine the behavior description output by the multimodal agent: [Agent output text], and determine the core behavior entity as '[behavior category]' and the scene entity as '[scene category + details]'"; Step 2 (Knowledge Graph Retrieval and Reasoning): "Retrieve the associations of '[behavioral entities]' based on the knowledge graph: [List the entities that 'execution' occurs in relation to the 'association' relationship, such as 'not wearing a safety rope → associated with JGJ59-2021']; Filter candidate regulations that match '[scene entities]': [Filter based on the 'matching' relationship, such as '3m scaffolding scene → matched with regulations for heights ≥2m']" Step 3 (Preliminary Matching Judgment): "Based on the candidate rule directions output by the Agent, the rule matching scope is initially determined to be '[rule library name]'. Further verification of the consistency between scenario details and rule application conditions is required."

[0031] Collaboration between the thought chain and the multimodal large model agent: The thought chain template is used as a prompt input to the multimodal large model agent, guiding it to output structured results according to the "step-by-step reasoning" logic. For example, after inputting an image, the agent first outputs visual features and behavioral descriptions according to step 1, then retrieves knowledge graph related data according to step 2, and finally outputs the preliminary scope of rules according to step 3. Step 4: If there are ambiguities in the reasoning process (such as "scene height is not clear"), the multimodal large model agent can automatically provide feedback that "scaffolding height information in the image needs to be supplemented", thus realizing interactive reasoning.

[0032] S6. The multimodal large model receives on-site images and generates behavioral descriptions and candidate rule directions; in, The multimodal big model retrieves similar behavior-scene data from the multimodal knowledge graph based on on-site images, and outputs structured behavior description text based on the similar behavior-scene data, including behavior category, subject, and scene features; The multimodal large model is based on built-in lightweight domain knowledge and outputs 1-3 candidate regulation directions (such as "JGJ59-2021, JGJ80-2016").

[0033] S7. The multimodal large model extracts visual features from on-site images based on the thinking chain template, retrieves multimodal knowledge graphs based on visual features and behavioral descriptions, and filters the range of candidate regulations based on visual features and candidate regulation directions. in, The thought chain template takes a multimodal large model as input, guides the multimodal large model to process according to the distributed reasoning logic of the thought chain template, and outputs the results.

[0034] Specifically, After inputting an image, the multimodal big data agent first outputs visual features and behavioral descriptions in step 1, then retrieves related data from the knowledge graph in step 2, and finally outputs the preliminary scope of the rules in step 3. If there are ambiguities in the reasoning process (such as "the scene height is not clear"), the multimodal big data agent can automatically provide feedback that "the scaffolding height information in the image needs to be supplemented", thus realizing interactive reasoning.

[0035] S8. The multimodal large model selects candidate rules from a relational database based on candidate rule directions, converts visual features and behavioral descriptions into multimodal semantic vectors, and selects candidate rules from a vector database. in, The multimodal large model filters candidate clauses from the relational database based on behavior type matching and scene visual description matching (e.g., "not wearing a safety rope" + "visual description of 3m scaffolding", selecting 5-8 candidate clauses). The multimodal large model converts image features and behavioral descriptions into multimodal semantic vectors, calculates cosine similarity with the multimodal semantic vectors of candidate rules in the vector database, and filters rules with a similarity greater than a preset threshold (e.g., ≥0.8). S9. The multimodal large model is based on a multimodal knowledge graph and outputs the most matching safety behavior regulations and cross-modal interpretations of images, texts and regulations.

[0036] in, The multimodal big model combines the rule-scene adaptation relationship in the multimodal knowledge graph to verify the applicability of candidate rules, calculate the final score of semantic similarity + scene adaptation, and output the rule with the highest score (Top1) and the cross-modal interpretation of image-text-rule (e.g., "The 3m scaffolding scene in the image matches the ≥2m height applicable condition of JGJ59-2021-3.2.1, with semantic similarity of 0.85 and final matching score of 0.91").

[0037] like Figure 2 As shown, the present invention also provides a safety behavior rule matching system based on a large model and knowledge graph, including an input layer, a core layer, a support layer and an output layer; Input layer, used for: Input on-site images, video streams, and safety regulations manuals into the multimodal large model; Core layer, used for: The multimodal large model receives on-site images and generates behavioral descriptions and candidate rule directions; The multimodal large model extracts visual features from on-site images based on the thinking chain template, retrieves multimodal knowledge graphs based on visual features and behavioral descriptions, and filters the range of candidate regulations based on visual features and candidate regulation directions. The multimodal large model is based on candidate rule directions, filters candidate rules from relational databases, converts visual features and behavioral descriptions into multimodal semantic vectors, and filters candidate rules from vector databases; Support layer, used for: A multimodal dataset for the security domain, including images, text, and rule labels, is constructed based on security scenarios. Construct a multimodal large model, train the multimodal large model based on a multimodal dataset, and optimize the multimodal large model based on the objective function of image-text alignment loss + rule association loss; Based on a multimodal big model combined with security domain standards, a multimodal knowledge graph is constructed. The categories in the multimodal knowledge graph include people, behaviors, scenarios and regulations. Each category is associated with multiple multimodal attributes. The relationships defined in the knowledge graph include execution, occurrence, association, inclusion and adaptation. Based on the multimodal large model processing safety specification manual, a relational database and a vector database are constructed. The relational database stores the structured information of the multimodal regulation entries, and the vector database stores the multimodal semantic vectors of the original text of the regulations and the visual description of the applicable scenarios. Set up a security scenario thinking chain template based on multimodal knowledge graph; Output layer, used for: The multimodal big model is based on a multimodal knowledge graph and outputs the most matching safety behavior regulations and cross-modal interpretations of image-text-regulation. Example 1:

[0038] The method of this invention is applied to construction sites where safety ropes are not worn. The specific process includes: Step 1: Fine-tuning parameters of the multimodal agent: Based on LLaVA-1.6 (7B), LoRA fine-tuning was used (rank=8, learning rate=2e-4, training epochs=5). The training dataset contains 800 samples of "building safety image + text + regulations" and the validation set contains 200 samples. After fine-tuning, the agent's recognition accuracy for "not wearing a safety rope" reached 94.5%.

[0039] Step 2, Knowledge Graph Fusion Process: After the Agent inputs the image of "3m scaffolding without safety rope", it automatically retrieves the associated data of "without safety rope" in the knowledge graph (typical image URL, JGJ59-2021 text) as context to generate a behavior description: "Construction workers are working on 3m high scaffolding without safety rope, and the scene is unobstructed", initially associating the relevant regulations "JGJ59-2021, JGJ80-2016".

[0040] Step 3: Mind chain reasoning and rule matching. Mind chain Prompt input: Step 3.1: Extract visual features from the image: The construction worker is on a 3m scaffold with no safety rope on his upper body; Agent behavior description: [as described above], core behavior 'not wearing a safety rope', scene '3m scaffold'; Step 3.2: Search the map for 'No safety rope attached → associated with JGJ59-2021', and the scenario '3m scaffolding → applicable to ≥2m high-altitude regulations'; Step 3.3: Initial matching range 'JGJ59-2021'; Step 4, Matching Results: The Agent filters 5 candidate regulations of JGJ59-2021 from the relational database, calculates the multimodal semantic similarity (similarity between image + behavior description vector and regulation vector is 0.85), scene fit is 1.0 (3m≥2m), and finally matches the regulation JGJ59-2021-3.2.1 with a score of 0.91 and a response time of 1.2s. The output cross-modal interpretation is: "The 3m scaffolding scene in the image meets the applicable conditions of JGJ59-2021-3.2.1'≥2m high-altitude operations require safety ropes', with a semantic similarity of 0.85, and the match is accurate." Example 2:

[0041] The method of this invention is applied to construction work scenarios where safety helmets are not worn. The specific process includes: Step 1: Fine-tune the parameters of the multimodal agent: Same as in Example 1. After fine-tuning, the agent's accuracy in recognizing "not wearing a safety helmet" reaches 95.1%.

[0042] Step 2, Knowledge Graph Fusion Process: After the Agent inputs the image of "no safety helmet at construction site", it retrieves the related data of "no safety helmet" in the knowledge graph and generates a behavior description: "Construction workers are working on the construction site without wearing safety helmets. The scene is a ground construction area". The initial association is with the regulation direction "JGJ59-2021".

[0043] Mind chain reasoning and rule matching, Mind chain Prompt input: Step 3.1: Extract visual features: Construction worker is not wearing a safety helmet, scene is ground construction area; Agent description: [as described above], behavior 'not wearing a safety helmet', scene 'ground construction'; Step 3.2: Graph search 'Not wearing a safety helmet → associated with JGJ59-2021', scenario 'Ground construction → adapted to the full range of regulations for construction sites'; Step 3.3: Preliminary scope 'JGJ59-2021'; Step 4, Matching Results: The Agent filters 4 candidate clauses with a semantic similarity of 0.88 and a scene adaptability of 1.0. It matches clause JGJ59-2021-3.2.2 with a score of 0.94 and a response time of 1.0s. Explanation: "The image of ground construction scene adapts to JGJ59-2021-3.2.2 'All personnel on the construction site must wear safety helmets', with a high semantic matching degree." Example 3:

[0044] The specific process of applying the method of this invention in illegal hot work scenarios during maritime operations includes: Step 1: Fine-tuning parameters of the multimodal agent: Based on LLaVA-1.6 (7B), LoRA fine-tuning (rank=8, learning rate=2e-4, training epochs=5), the training dataset contains 800 samples of "maritime safety images + text + regulations". After fine-tuning, the agent's recognition accuracy for "illegal hot work" reaches 92.8%.

[0045] Step 2, Knowledge Graph Fusion Process: After the Agent inputs the image of "illegal hot work on deck", it retrieves the related data of "illegal hot work" in the knowledge graph and generates a description: "Crew members are using open flames on the ship's deck without any fire extinguishing equipment nearby, and the scenario is in a berthing state". The initial association with the regulation direction is "GB 50720-2011".

[0046] Step 3: Mind chain reasoning and rule matching. Mind chain Prompt input: Step 3.1: Extract visual features: open flame on deck, no fire extinguishing equipment, berthed status; Agent description: [as described above], behavior 'illegal open flame', scene 'berthed on deck'; Step 3.2: Graph search 'Illegal hot work → associated with GB 50720-2011', scenario 'Deck berthing → applicable to ship fire safety regulations'; Step 3.3: Preliminary scope 'GB 50720-2011'; Step 4, Matching Results: The Agent filters 6 candidate regulations, with a semantic similarity of 0.83 and a scene adaptability of 1.0. It matches GB 50720-2011-5.2.10 with a score of 0.90 and a response time of 1.3s. The explanation is: "The image of a hot work scene on a berthed deck meets the requirements of GB50720-2011-5.2.10 'Hot work on ship decks must be equipped with fire extinguishing equipment', and the semantic matching is accurate."

[0047] This invention revolves around "multimodal large-scale model agent and knowledge graph collaboration," and overcomes the limitations of existing security behavior regulation matching systems through four key technologies, achieving a comprehensive improvement in accuracy, reliability, and interpretability. The key technologies include: (1) Domain-Adaptive Multimodal Large Model Agent Construction: To address the shortcomings of general multimodal large models in the security domain, such as "low accuracy in identifying specific behaviors and easy output of illusory content," this invention adopts a combined strategy of "targeted dataset construction + LoRA lightweight fine-tuning." First, a security domain-specific multimodal dataset containing "on-site images + behavioral text descriptions + corresponding regulation labels" is constructed, covering 23 typical unsafe behaviors in scenarios such as construction and maritime operations. Then, based on the LLaVA-1.6 basic model, the backbone parameters are frozen, and only the low-rank matrix (rank set to 8) is trained. The model is optimized using a dual objective function of "image-text alignment loss + regulation association loss." This technique improves the agent's accuracy in identifying security scenario-specific behaviors such as "not wearing double-hook seat belts" and "illegal hot work on deck" by 35% compared to the general model. At the same time, the illusory output rate of fabricated regulations and incorrect associations is strictly controlled to below 2%, balancing model lightweighting and domain adaptability.

[0048] (2) Multimodal Knowledge Graph and Agent RAG Fusion Mechanism: To address the issue that large multimodal models relying on internal parameters are prone to outputting data detached from domain knowledge, this invention innovatively designs a "Knowledge Graph-RAG" collaborative architecture—constructing a multimodal knowledge graph containing four core categories: "personnel-behavior-scenario-regulation." Each entity is associated not only with textual attributes (such as the original text of the regulation and the definition of the behavior) but also with visual feature descriptions (such as typical image features of "not wearing a safety helmet"). This graph is used as an external retrieval library for the Agent. When the Agent receives on-site image input, it first retrieves multimodal data of "similar behavior-scenario" in the graph as context supplementation, and then generates behavior descriptions and regulation association results. This mechanism not only preserves the Agent's cross-modal understanding capabilities but also avoids illusory outputs through the domain knowledge constraints of the graph, improving the initial regulation association accuracy by 40% compared to models without RAG support, ensuring that the output results conform to safety regulations.

[0049] (3) Cross-modal reasoning guided by thought chain: In response to the shortcomings of traditional large models, such as "black box reasoning process and inability to trace decision logic", this invention designs a thought chain template specifically for safety scenarios to guide the Agent to output results in a "step-by-step reasoning + logical transparency" manner. The thought chain template includes three core steps: the first step, "visual feature extraction", requires the Agent to extract key visual information from images (such as "construction workers are on 3m scaffolding, and their upper bodies are not restrained by safety ropes"); the second step, "graph retrieval association", based on the extracted behavior and scene entities, retrieves the "behavior-regulation" association relationship in the knowledge graph (such as "not wearing a safety rope → associated with JGJ59-2021"); the third step, "preliminary matching judgment", combined with scene details to filter the scope of applicable regulations (such as "3m height → applicable to ≥2m height operation regulations"). Through this template, the Agent's reasoning process is presented in a traceable form of natural language, and the integrity and interpretability of the reasoning link reach 100%, which is convenient for safety management personnel to verify the decision logic and reduce the cost of misjudgment investigation.

[0050] (4) Multimodal collaborative regulation matching: To address the problems of traditional keyword matching, such as "inability to associate with visual scenes and large semantic deviations," this invention constructs a "dual-database collaboration + multimodal matching" mechanism. The relational database stores the structured information of regulations (such as regulation ID, applicable scenarios, and original text), supporting precise filtering based on "behavior type + scenario keywords." The vector database converts the "text content + visual description of applicable scenarios" of regulations into 768-dimensional multimodal semantic vectors, supporting cross-modal similarity calculation. The matching process is agent-driven: first, candidate regulations matching behavior types are filtered from the relational database; then, the cosine similarity between the multimodal vector of "on-site image + behavior description" and the candidate regulation vector is calculated; finally, the final result is obtained by weighting the scenario suitability (such as the rule judgment of "3m≥2m"). This technology achieves deep association between "visual scene and text regulation," improving the regulation matching accuracy by 32% compared to traditional keyword methods, and controlling the mismatch rate below 1.5%, meeting the precision requirements of security management.

[0051] Advantages of this invention: This invention, based on an innovative architecture of "multimodal large-scale model agent + knowledge graph + dual-database collaboration," conducted performance verification tests in three typical safety scenarios: building construction, maritime operations, and power maintenance (1500 test samples covering 23 common unsafe behaviors, with an RTX 3090 graphics card and 16GB of RAM). The test results show that this invention performs excellently in terms of recognition accuracy, rule matching, inference efficiency, and domain generalization, with the following effects: (1) High multimodal recognition accuracy and strong robustness in complex scenarios: This invention uses LoRA to lightly fine-tune the multimodal basic model (experimental conditions: 1440 "image-text-rule" exclusive samples in the training set, 10 rounds of fine-tuning, and low-rank matrix rank = 8). The overall recognition accuracy of the Agent for safe behavior after fine-tuning reaches 93.2%. For complex scenarios such as "low light (illuminance < 50 lux) and personnel occlusion (occlusion area < 30%)", the recognition robustness is still maintained at 89.3%, which can stably cope with the recognition needs of unsafe behavior in various on-site environments and avoid recognition failure caused by environmental interference. (2) Excellent matching performance and low mismatch rate: This invention adopts the "dual-database collaboration + multimodal matching" algorithm (experimental conditions: 360 verification samples in the test set, 1200 structured security regulations stored in the relational database, 768-dimensional multimodal semantic vectors stored in the vector database, and the similarity screening threshold set to 0.8). The matching accuracy of the regulations reaches 95.8%, and the scene adaptation rate is as high as 98%. At the same time, the mismatch rate is strictly controlled below 1.5%, which can accurately realize the automatic association of "unsafe behavior - security regulations", providing clear and reliable compliance basis for security management. (3) Good interpretability of reasoning and high decision-making efficiency: This invention guides reasoning through a three-step thinking chain template (experimental conditions: 10 senior safety management personnel participated in the logic understanding test, and each person independently processed 50 matching results). The reasoning process is 100% traceable, and the safety management personnel's understanding efficiency of the "behavior-regulation" matching logic is improved by 80%. Moreover, the system automatically outputs cross-modal interpretation of "key features of the image + original text of the regulation", eliminating the need for manual consultation of the regulation manual, reducing the rectification decision time from the traditional 30 minutes to 4 minutes, and greatly improving the safety management response speed. (4) Strong domain generalization ability and low adaptation cost: This invention achieves rapid cross-domain adaptation through "lightweight fine-tuning of multimodal agents + dynamic updating of knowledge graphs" (experimental conditions: when migrating to the power operation and maintenance scenario, only 100 power safety-specific samples need to be added to fine-tune the agent, and 50 graph entities and relationships need to be updated simultaneously). The model fine-tuning cost is reduced by 70% when migrating across domains, and the overall adaptation cycle is shortened to less than 3 days. It can flexibly cover multiple safety fields such as construction, maritime, power, and chemical industries, without the need for repeated development for a single field, which significantly reduces the cost of implementation.

[0052] The above are merely preferred embodiments of the present invention and are not intended to limit the present invention. Various modifications and variations can be made to the present invention by those skilled in the art. Any modifications, equivalent substitutions, improvements, etc., made within the spirit and principles of the present invention should be included within the scope of protection of the present invention.

Claims

1. A method for matching safety behavior regulations based on a large model and a knowledge graph, characterized in that, The method comprises the following steps: Constructing a safety field multi-modal data set including images, texts, and regulation labels based on a safety scenario; Constructing a multi-modal large model, training the multi-modal large model based on the multi-modal data set, and optimizing the multi-modal large model based on an image-text alignment loss + regulation correlation loss objective function; Based on the multi-modal large model and the safety field specification, a multi-modal knowledge graph is constructed, wherein the categories in the multi-modal knowledge graph include personnel, behavior, scene, and regulation, each category is associated with multiple multi-modal attributes, and the defined relationships in the knowledge graph include execution, occurrence, association, inclusion, and adaptation; Based on the multi-modal large model, a relational database and a vector database are constructed by processing a safety specification manual, the relational database stores the structured information of the multi-modal regulation items, and the vector database stores the multi-modal semantic vectors of the regulation original text and the applicable scene visual description; Based on the multi-modal knowledge graph, a safety scenario thinking chain template is set; The multi-modal large model receives a field image, generates a behavior description and a candidate regulation direction; The multi-modal large model extracts visual features from the field image based on the thinking chain template, retrieves the multi-modal knowledge graph based on the visual features and the behavior description, and filters a candidate regulation range based on the visual features and the candidate regulation direction; The multi-modal large model filters a candidate regulation from the relational database based on the candidate regulation direction, converts it into a multi-modal semantic vector based on the visual features and the behavior description, and filters a candidate regulation from the vector database; Based on the multi-modal knowledge graph, the multi-modal large model outputs the most matched safety behavior regulation and the image-text-regulation multimodal explanation.

2. The method of claim 1, wherein the method is based on a large model and a knowledge graph. The construction of the multi-modal large model comprises the following steps: adopting an LLaVA-1.6 model as a basic multi-modal large model, adopting a LoRA light-weight fine-tuning strategy, freezing the main parameters of the LLaVA-1.6 model, and training a low-rank matrix based on the multi-modal data set; The fine-tuning objective function is set as an image-text alignment loss + regulation correlation loss, and a fine-tuned multi-modal large model is obtained. 3.The method of claim 1, wherein, Based on the multi-modal large model and the safety field specification, a multi-modal knowledge graph is constructed, comprising: A multi-modal ontology framework is constructed by the fine-tuned multi-modal large model combined with the safety field specification, and the multi-modal knowledge graph including categories of personnel, behavior, scene, and regulation and relationships of execution, occurrence, association, inclusion, and adaptation is formed based on the multi-modal ontology framework and the multi-modal data set.

4. The method of claim 1, wherein the method is based on a large model and a knowledge graph. Based on the multi-modal large model, a relational database and a vector database are constructed by processing a safety specification manual, comprising: Texts and tables in the safety specification manual are extracted by the multi-modal large model to generate typical image descriptions of regulation applicable scenes, and multi-modal regulation items of texts + visual descriptions are formed; The structured information of the multi-modal regulation items is stored to form a relational database, and the fields include regulation unique identifier, field, associated behavior type, regulation original text, applicable scene visual description, and applicable scene. The regulation item and applicable scene visual description of all regulations are converted into a multi-modal semantic vector to obtain a vector database, and the fields include a regulation unique identifier and a multi-modal semantic vector. 5.The method of claim 1, wherein, The safety scene thinking chain template is set based on the multi-modal knowledge graph, including: The safety scene thinking chain template is set based on the logic of entities and relationships in the multi-modal knowledge graph; The logic of the thinking chain template is that visual features are extracted from the input image, the core behavior entity and the scene entity are confirmed in combination with the behavior description output by the multi-modal large model, the associated relationship of the behavior entity is retrieved based on the multi-modal knowledge graph, the candidate regulations suitable for the scene entity are screened, and the regulation matching range is obtained in combination with the candidate regulation direction output by the multi-modal large model.

6. The method of claim 1, wherein the method is based on a large model and a knowledge graph. The multi-modal large model receives the on-site image, generates the behavior description and the candidate regulation direction, including: The multi-modal large model retrieves similar behavior-scene data in the multi-modal knowledge graph according to the on-site image, and outputs a structured behavior description text including a behavior category, a subject and a scene feature based on the similar behavior-scene data; The multi-modal large model outputs multiple candidate regulation directions based on the built-in lightweight field knowledge.

7. The method of claim 5, wherein the method further comprises: determining a security behavior regulation based on the knowledge graph and the large model; and providing the determined security behavior regulation to the user. The multi-modal large model extracts visual features from the on-site image based on the thinking chain template, retrieves the multi-modal knowledge graph based on the visual features and the behavior description, and screens the candidate regulation range based on the visual features and the candidate regulation direction; including: The thinking chain template is input into the multi-modal large model, the thinking chain template guides the multi-modal large model to process according to the distribution reasoning logic of the thinking chain template, and outputs the result. 8.The method of claim 1, wherein, The multi-modal large model screens the candidate regulations from the relational database based on the candidate regulation direction, and screens the candidate regulations from the vector database based on the visual features and the behavior description; including: The multi-modal large model screens the candidate regulations that match the behavior type and the scene visual description from the relational database; The multi-modal large model converts the image features and the behavior description into a multi-modal semantic vector, calculates the cosine similarity of the multi-modal semantic vector of the candidate regulation in the vector database, and screens the regulations with a similarity greater than a preset similarity threshold.

9. The method of claim 8, wherein the method further comprises: The multi-modal large model outputs the most matched safety behavior regulation and the multi-modal explanation of image-text-regulation based on the multi-modal knowledge graph, including: The multi-modal large model verifies the applicable conditions of the candidate regulation in combination with the adaptation relationship of the regulation-scene in the multi-modal knowledge graph, calculates the final score of the semantic similarity and the scene adaptation degree, outputs the regulation with the highest score and the cross-modal explanation of image-text-regulation.

10. A system for matching safety behavior regulations based on large models and knowledge graphs, realizing the method for matching safety behavior regulations based on large models and knowledge graphs according to any one of claims 1-9, characterized in that: It includes an input layer, a core layer, a support layer and an output layer; The input layer is used to: input the on-site image, video stream and safety regulation manual into the multi-modal large model; The core layer is used to: the multi-modal large model receives the on-site image, generates the behavior description and the candidate regulation direction; The multi-modal large model extracts visual features from the live image based on the thought chain template, retrieves the multi-modal knowledge graph based on the visual features and the behavior description, and filters candidate regulation ranges based on the visual features and the candidate regulation direction; The multi-modal large model filters candidate regulations from the relational database based on the candidate regulation direction, converts multi-modal semantic vectors based on the visual features and the behavior description, and filters candidate regulations from the vector database; The support layer is used for: Based on the safety scene, a safety field multi-modal data set including images, texts, and regulation labels is constructed; A multi-modal large model is constructed, the multi-modal large model is trained based on the multi-modal data set, and the multi-modal large model is optimized based on an image-text alignment loss + regulation correlation loss objective function; Based on the multi-modal large model combined with safety field specifications, a multi-modal knowledge graph is constructed, the multi-modal knowledge graph includes classes such as personnel, behavior, scene, and regulation, each class is associated with multiple multi-modal attributes, and the knowledge graph defines relationships such as execution, occurrence, association, inclusion, and adaptation; Based on the multi-modal large model processing safety specification manuals, a relational database and a vector database are constructed, the relational database stores structured information of multi-modal regulation items, and the vector database stores multi-modal semantic vectors of regulation original texts and applicable scene visual descriptions; Based on the multi-modal knowledge graph, a safety scene thought chain template is set; The output layer is used for: The multi-modal large model outputs the most matched safety behavior regulation and the multi-modal explanation of image-text-regulation based on the multi-modal knowledge graph.

Citation Information

Patent Citations

  • Personnel track sitting and lying behavior detection method and device based on deep learning

    CN120580734A

  • Unsafe behavior identification method and system based on 5G and artificial intelligence

    CN120599194A