An emergency fire hazard detection method based on multi-modal AI large model recognition technology
By building a multimodal hidden danger knowledge base and multimodal large model technology, the problem of insufficient multimodal data fusion in fire hazard inspections is solved, and the accurate identification of fire hazards and the generation of rectification suggestions are achieved, which improves the intelligence and efficiency of fire hazard inspections.
Patent Information
- Application Number
- CN202411819760.5
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2024-12-11
- Publication Date
- 2025-07-18
- Estimated Expiration
- 2044-12-11
AI Technical Summary
In the fire hazard investigation, the existing technology has problems such as insufficient multimodal data fusion, lack of accuracy and comprehensiveness of hidden danger identification, and limited application of field knowledge. It is difficult to effectively combine images, texts and laws and regulations to conduct comprehensive and accurate hidden danger identification and rectification suggestions.
Build a multi-modal hidden danger knowledge base, combine deep learning and multi-modal large-modal model technology, and realize multi-dimensional analysis and accurate identification of fire hazards through adaptive image coding, LoRa fine-tuning and multi-hop reasoning, and generate rectification suggestions and legal basis.
It improves the intelligence and efficiency of fire hazard inspections, and can analyze hidden dangers in the image from multiple angles and levels to ensure the comprehensiveness and accuracy of hidden danger inspections, and generate legal and compliant rectification opinions.
Smart Images

Figure CN119646271B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of artificial intelligence, and specifically to an emergency fire hazard inspection method based on multi-modal AI large model recognition technology. Background Art
[0002] With the continuous advancement of the industrialization process, the factory production environment has become increasingly complex, and fire safety hazards have become important factors threatening production safety and the lives of employees. Traditional fire hazard inspection methods mainly rely on manual inspections or single-modal data analysis, and often have the following deficiencies:
[0003] 1. Insufficient multi-modal data fusion: Fire hazards usually exist in the form of images, texts, and other data, and existing technologies have weak processing capabilities for multi-modal data and are difficult to effectively fuse;
[0004] 2. Lack of accuracy and comprehensiveness in hazard identification: Traditional methods judge based on predefined hazard classifications and are difficult to handle new types or significantly changed hazards. The diversity and ambiguity of hazards make traditional methods difficult to adapt to complex scenarios in practical applications;
[0005] 3. Limited application of domain knowledge: The identification of fire hazards requires the combination of domain knowledge such as fire safety regulations and equipment operation specifications, and existing technologies have not fully and effectively applied these unstructured domain knowledge to the hazard inspection model;
[0006] Therefore, how to achieve a comprehensive and accurate identification of fire hazards and generate intelligent rectification suggestions in combination with laws and regulations has become an urgent problem to be solved by current technologies. Summary of the Invention
[0007] The purpose of the present invention is to provide an emergency fire hazard inspection method based on multi-modal AI large model recognition technology. By constructing a multi-modal hazard knowledge base and combining deep learning with multi-modal large model technology, the identification ability of hazard information is improved. Through multi-hop reasoning and knowledge retrieval enhanced generation technology, the present invention can accurately analyze fire hazards from multiple dimensions and automatically generate rectification suggestions and legal bases, improving the intelligence and efficiency of hazard inspection to solve the problems raised in the above background art.
[0008] To achieve the above purpose, the present invention provides the following technical solution: An emergency fire hazard inspection method based on multi-modal AI large model recognition technology, including the following steps:
[0009] Step 1: Data processing, collect historical hazard inspection data containing images as multi-modal data, uniformly organize the data, and label safety hazards with corresponding hazard levels;
[0010] Step 2: Image processing. An adaptive image coding mechanism is added to the image data to directly process images of various resolutions and aspect ratios;
[0011] Step 3: Fine-tuning the large model in the LoRa manner. Using the data collected and processed above, introduce LoRa layers in each layer of the base model, set training parameters. Except for the LoRa layers, freeze all other parameters in the base model and only update the parameters in the LoRa layers, so that the model can adapt to specific tasks more quickly and enable the large model to form a more comprehensive judgment system in the field of emergency fire protection;
[0012] Step 4: Vectorization of text data. Based on a deep learning embedding model, convert the text data in the historical hidden danger investigation data into numerical vector representations, and store this knowledge in the knowledge base in vector form to make the retrieval process more efficient;
[0013] Step 5: Hidden danger reasoning and retrieval. Design a multi-hop prompt template to perform multi-hop reasoning on the input pictures or videos for hidden dangers. Through a thinking chain that focuses on different detailed points in stages, conduct multi-dimensional analysis on the emerging hidden danger factors, comprehensively consider the information in each dimension, retrieve content similar to the results of hidden danger reasoning in the knowledge base, re-examine the results of hidden danger reasoning from multiple angles, and generate rectification opinions at the same time;
[0014] Step 6: Generate a reply. Use the results of hidden danger reasoning and the content retrieved from the knowledge base as context to give the large model a reply on safety hazards, the legal basis for the corresponding safety hazards, and the rectification opinions for the corresponding safety hazards.
[0015] Preferably, the historical hidden danger investigation data in Step 1 also includes text data such as hidden danger report documents, equipment operation specifications, industry standards, and relevant laws and regulations. Organize the text data into structured electronic data, where the text data of the laws and regulations is organized into a csv format, with one law in each cell. At the same time, label the hidden danger level and description of the safety hazards existing in each picture in the historical hidden danger investigation image data, and establish a mapping relationship between each law and regulation and the safety hazards.
[0016] Preferably, the specific method for the adaptive image coding mechanism to process pictures in Step 2 is as follows:
[0017] First, scale the picture to the specified size, encode it using VIT, and merge and flatten the encoded picture;
[0018] Secondly, divide the picture into several parts according to the set specifications, encode each part using VIT respectively, and merge and flatten the encoded pictures;
[0019] Finally, the flattened encoded images are vector - stitched twice to obtain the final image encoding.
[0020] Preferably, the specific method for converting the text data in the historical hidden - danger investigation data into numerical vector representation based on a deep - learning embedding model in Step 4 includes the following steps:
[0021] S101: Split the structured text data into meaningful text blocks, which can be words or sub - word units, and remove common stop words at the same time;
[0022] S102: Restore the text blocks to their basic forms, reduce the diversity of vocabulary, and add corresponding part - of - speech tags at the same time;
[0023] S103: Convert the processed text blocks into word vectors using a word - embedding model, store them in a vector knowledge base, and create an index for the knowledge base to enable retrieval and matching.
[0024] Preferably, the retrieval and matching of the knowledge base specifically includes the following method: Convert the user - input query into a vector in the same form as the knowledge base, use the cosine - similarity calculation method to find the text - data vector in the knowledge base that is most similar to the query vector, and sort the retrieved results from largest to smallest according to the similarity score.
[0025] Preferably, the hidden - danger reasoning and retrieval in Step 5 specifically includes the following steps:
[0026] S201: Identify the content in the image data, and multi - dimensionally and step - by - step describe the predicted hidden - danger content according to the identified image content and the constructed multi - hop prompt framework;
[0027] S202: Input the described predicted hidden - danger content into the knowledge base for similarity matching, output the content with the highest correlation score as the retrieved content, and determine the corresponding laws and regulations for each safety hidden - danger;
[0028] S203: Retrieve professional content from the knowledge base, and generate corresponding rectification opinions for the identified safety hidden - danger according to the corresponding laws and regulations.
[0029] Preferably, the specific method for inputting the predicted hidden - danger content into the knowledge base for similarity matching is:
[0030] Encode the step - by - step described predicted hidden - danger content and the corresponding safety - hidden - danger text blocks retrieved from the knowledge base together through a CrossEncoder model. The model generates a correlation score based on this pair of combined inputs, sorts according to the score of the model, and outputs the text block with the highest score as the final result of the retrieved content.
[0031] Preferably, the multi-hop prompt framework includes the following parts:
[0032] Problem statement: Describe the problem to be solved;
[0033] Initial information: Provide the necessary background information in the image content;
[0034] Inference steps: List the descriptions of the corresponding regions in the image content step by step according to the image content, and each step is equipped with the corresponding logical basis;
[0035] Auxiliary information: Introduce additional knowledge or data support;
[0036] Inference result: Output the inference result of each step and give a clear answer.
[0037] Preferably, a noise filtering mechanism is set in the retrieval matching process of the knowledge base. The specific method is as follows: Noise filtering is performed on the text data and image data before retrieval, and only high-quality data is retained for indexing and storage. After the user submits a query, the retrieval results are dynamically filtered to remove irrelevant information, and the noise filtering mechanism is continuously optimized according to the user's feedback.
[0038] In summary, the beneficial effects of the present invention are as follows:
[0039] 1. Through full investigation of the current status in the field of emergency fire protection, combined with the characteristics of fire hazards in factories and the characteristics of multi-modal data, the present invention collects multi-modal data such as images, texts, and relevant laws and regulations actually recorded during the hidden danger investigation process, uniformly organizes and annotates these data, establishes a multi-modal hidden danger knowledge base, and proposes to improve the recognition and analysis ability of hidden danger information through deep learning and multi-modal large model technologies. By using multi-modal thinking chain and retrieval enhanced generation technologies, the present invention realizes the function of accurately identifying hidden dangers driven by multi-modal data, helping users quickly and accurately discover potential safety hazards in the production environment.
[0040] 2. The present invention uses multi-hop reasoning and thinking chain technologies to decompose complex hidden danger positioning problems into a series of interpretable steps, making the decision-making process of the model more transparent, helping humans understand the basis for each hidden danger judgment, and enabling the analysis of hidden dangers in the picture from multiple angles and levels, rather than relying only on single information for judgment. Each reasoning step can gradually verify different assumptions or conditions, thereby enhancing the accuracy of the final judgment and comprehensively considering the information of each dimension for a comprehensive judgment.
[0041] Multi-hop reasoning guides the model to focus on different details through a phased thinking chain. In each step of reasoning, the model not only pays attention to the core elements of the current reasoning step but also re-analyzes and verifies the details in combination with the conclusions of the previous steps. This gradually refined process helps to improve the sensitivity of the model to subtle hidden dangers in the image. Brief Description of the Drawings
[0042] In order to more clearly illustrate the technical solutions in the embodiments of the invention or the prior art, the following will briefly introduce the drawings required for the description of the embodiments or the prior art. Obviously, the drawings in the following description are only some embodiments of the invention. For those of ordinary skill in the art, without creative efforts, other drawings can be obtained based on these drawings.
[0043] Figure 1 It is a schematic diagram of the process framework structure of an emergency fire hazard inspection method based on multi-modal AI large model recognition technology of the present invention;
[0044] Figure 2 It is a schematic diagram of the legal regulation segmentation of an emergency fire hazard inspection method based on multi-modal AI large model recognition technology of the present invention;
[0045] Figure 3 It is a schematic diagram of the description of safety hazards of an emergency fire hazard inspection method based on multi-modal AI large model recognition technology of the present invention;
[0046] Figure 4 It is a schematic diagram of the image coding process framework of an emergency fire hazard inspection method based on multi-modal AI large model recognition technology of the present invention;
[0047] Figure 5 It is a schematic diagram of the image coding display of an emergency fire hazard inspection method based on multi-modal AI large model recognition technology of the present invention. Detailed Description of the Embodiments
[0048] Now, the present invention will be further described in detail with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of the present invention, rather than all embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those of ordinary skill in the art without creative efforts belong to the scope of protection of the present invention. These drawings are all simplified schematic diagrams, which only illustrate the basic structure of the present invention in a schematic manner, and therefore only show the components related to the present invention.
[0049] To facilitate the understanding of the present invention, the present invention will be described more comprehensively with reference to the relevant drawings. Several embodiments of the present invention are given in the drawings. However, the present invention can be implemented in many different forms and is not limited to the embodiments described herein. On the contrary, the purpose of providing these embodiments is to make the disclosure of the present invention more thorough and comprehensive.
[0050] All features disclosed in this specification, or all steps in the disclosed methods or processes, except for mutually exclusive features and / or steps, can be combined in any way.
[0051] Any feature disclosed in this specification (including any additional claims, abstract, and drawings), unless specifically recited, may be replaced by other equivalent or similar-purpose alternative features. That is, unless specifically recited, each feature is merely an example of a series of equivalent or similar features.
[0052] In the present invention, unless otherwise clearly specified and defined, terms such as "installation", "connection", "attachment", "fixation", etc. shall be understood in a broad sense. For example, it may be a fixed connection, a detachable connection, or integrated; it may be a mechanical connection, a direct connection, or an indirect connection through an intermediate medium. It may be the communication inside at least two elements or the interaction relationship between at least two elements, unless otherwise clearly defined. For those of ordinary skill in the art, the specific meanings of the above terms in the present invention can be understood according to specific circumstances.
[0053] The intelligent identification of potential hazards based on multi-modal large models highly depends on the data volume of the potential hazard knowledge base and its coverage of different types of potential hazards. For newly added potential hazards or non-significant potential hazards not covered in the knowledge base, existing potential hazard identification systems often cannot give clear risk warnings, or only give general suggestions such as "it is recommended to manually recheck". This will not only reduce the trust of factory managers in the potential hazard identification system, but also may lead to potential hazards not being processed in a timely manner, increasing the risk of accidents.
[0054] Therefore, an embodiment provided by the present invention: an emergency fire hazard inspection method based on multi-modal AI large model recognition technology, which uses the existing potential hazard report documents, equipment operation specifications, fire safety rules and regulations, and relevant laws and regulations in the factory to generate diagnoses and suggestions corresponding to potential hazards from multi-modal data, further improving the accuracy and adaptability of potential hazard identification;
[0055] The following is combined with Figures 1-5 This embodiment will be described in detail, including the following steps:
[0056] The first step: data processing
[0057] Collect text data and image data including historical potential hazard inspection data, including potential hazard report documents, equipment operation specifications, industry standards, and relevant laws and regulations, as multi-modal data, uniformly organize the data and label the safety hazards corresponding to the hazard levels. Using the data actually recorded during the previous potential hazard inspection process enables the large model to better solve problems in actual inspection scenarios and better fit the real usage scenarios;
[0058] Specifically, the image data mainly comes from: obtaining images and videos from surveillance cameras, mobile devices (such as smartphones and tablets), drones and other devices, capturing scenes related to safety hazards, such as equipment failure, unsafe working conditions, blocked fire passages, etc., adding corresponding hazard descriptions and hazard level labels to the safety hazards in the images, and each picture or video clip should contain an accurate timestamp to track the time of the event. If possible, geotags can also be added to determine the location where the hazard occurred.
[0059] Text data can come from hidden danger report documents, checklists, maintenance logs, feedback from inspectors, accident investigation reports, etc. Text data should be converted into structured electronic data as much as possible to facilitate machine reading and processing;
[0060] It should be noted that the more important sources of text data for laws and regulations are national and local safety standards, industry guidelines, internal corporate safety policies, etc. It is necessary to keep the latest versions of laws, regulations and rules and regulations, and record the history of any updates or revisions, and organize the text data of laws and regulations into csv format, with one law as one cell.
[0061] Image data annotation includes:
[0062] Object detection: Label objects in images, such as equipment parts, workers, protective equipment, etc. For videos, labeling can be performed by frame or time period.
[0063] Scene understanding: Identify and annotate specific scenes shown in images or videos, such as workshop layout, hazardous areas, etc.
[0064] Behavior recognition: For video data, people’s behaviors can also be annotated, such as not wearing a helmet, illegal operation, etc.
[0065] Text data annotation includes:
[0066] Entity recognition: marking key entities in text, such as names of people, places, equipment names, time, quantity, etc.;
[0067] Relationship extraction: find the relationship between entities in the text, such as cause and effect, sequence, etc.
[0068] Sentiment analysis: assessing the sentiment expressed in text, especially when it comes to prosecutorial feedback;
[0069] Legal and regulatory data annotation includes
[0070] Clause analysis: Analyze laws, regulations and rules and regulations one by one to clarify their scope of application and requirements.
[0071] Compliance Checkpoint: Identify which specific clauses can be directly used for potential hazard investigation and rectification suggestions.
[0072] Associated Mapping: Establish the mapping relationship between regulatory clauses and actual potential hazards to help quickly locate relevant regulations.
[0073] Second step, adaptive processing of image processing
[0074] An adaptive image coding mechanism is added to the image data to directly process images of various resolutions and aspect ratios. Traditional image processing models usually require input images to have fixed resolutions and aspect ratios, or need to scale and crop the images before processing. This can lead to information loss or distortion, especially when the image content is complex or important details are located at the edges of the image. To support input images of any resolution and aspect ratio and ensure that the high-dimensional features of the image are not lost, the present invention designs an adaptive image coding mechanism. This mechanism can flexibly adapt to image data in various actual scenarios without manual intervention. This enables the system to process images or videos from various sources, greatly improving the generality and adaptability of the model. This mechanism can directly process images of various resolutions and aspect ratios, avoiding unnecessary image distortion or loss of important details, thereby improving the accurate recognition of image details.
[0075] Image and video data in the real world have a high degree of uncertainty and may come from different cameras, devices, or sensors with various formats. Supporting inputs of multiple resolutions and aspect ratios enables the model to flexibly handle different data sources, thereby enhancing the reliability and robustness of the model in various scenarios and ensuring accurate safety hazard analysis in various complex scenarios. Users do not need to care about the size, resolution, and aspect ratio of the image during actual operation, which greatly simplifies the usage process. Whether uploading high-resolution high-definition pictures or low-quality surveillance videos, the system can automatically adapt and perform effective analysis. Such a design improves the user experience, especially for non-professional users or staff without image processing skills, greatly reducing the usage threshold.
[0076] It is worth mentioning that in this embodiment, the specific processing method of the adaptive image coding mechanism for pictures is as follows:
[0077] First, the picture is scaled to the specified size, encoded using VIT, and the encoded pictures are merged and flattened; secondly, the picture is divided into several parts according to the set specification size, encoded using VIT respectively, and the encoded pictures are merged and flattened; finally, the two flattened encoded pictures are vector spliced to obtain the final picture encoding.
[0078] It can directly process images of various resolutions and aspect ratios, avoiding unnecessary image distortion or loss of important details, thereby improving the accurate recognition of image details. In reality, the resolutions and aspect ratios of images vary, especially in images captured by devices such as cameras, mobile phones, surveillance cameras, and drones, where there are significant differences in image quality and proportions. The present invention can support images from a variety of different sources, avoiding the forced unification of images in the preprocessing step, enabling images captured by different devices to be efficiently processed, enhancing the applicability of the system, and being able to process security monitoring images or videos of different sources and formats. Traditional methods usually require a large number of preprocessing steps, such as unifying the size, cropping, or scaling. These steps not only increase the workload but also may lead to the loss or distortion of image information. Supporting image input of any resolution and aspect ratio greatly simplifies this process, reduces the manual processing of data, and makes the system more efficient and automated during the processing.
[0079] Step 3: Fine-tune the large model in the LoRa manner
[0080] Using the data collected and processed above, fine-tune the large model in the LoRa manner to make the large model more suitable for the emergency fire protection field. LoRa is a parameter-efficient fine-tuning method that adapts to specific tasks by adjusting only a small part of the parameters in the model (i.e., low-rank matrices), thus greatly reducing the computational resources and time required for training.
[0081] First, introduce the LoRa layer: Introduce the LoRa layer in each layer of the base model. The LoRa layer consists of two low-rank matrices, which are used to adjust the row space and column space of the original weight matrix respectively. In this way, without changing the original weights, it can adapt to new tasks by learning a small number of parameters.
[0082] Set hyperparameters: Determine the rank of the LoRa layer (usually a relatively small rank, such as 4 or 8, can achieve good results), and other training parameters such as the learning rate and batch size can also be set.
[0083] Freeze most of the parameters: Except for the LoRa layer, freeze all other parameters in the base model and only update the parameters in the LoRa layer. This can significantly reduce the computational resources and time required for training.
[0084] Secondly, design specific tasks:
[0085] Classification task: For example, identify different types of safety hazards or emergency response levels according to the description.
[0086] Question-answering system: Develop a question-answering system that can answer questions about emergency procedures, laws and regulations, etc.
[0087] Text generation: Train the model to generate detailed emergency response plans, accident reports, or training content.
[0088] Multimodal tasks: Combine image and text data to enable the model to understand and analyze the relationship between visual information (such as photos of the fire scene) and text information (such as accident descriptions).
[0089] Finally, conduct training and optimization
[0090] Training process: Fine-tune the model using the prepared dataset. Since LoRA only needs to train a small number of parameters, the training speed is fast and it is not prone to overfitting.
[0091] Monitoring metrics: Closely monitor evaluation metrics such as loss function, accuracy, and F1 score during the training process to ensure continuous improvement of the model performance.
[0092] Validation and testing: Evaluate the performance of the model on independent validation and test sets to ensure its generalization ability.
[0093] By combining multimodal data such as images, text, and laws and regulations, the large model fine-tuned by LoRA can understand different types of information simultaneously. For example, images can provide visual clues of potential hazards, text data (such as safety inspection reports) can provide detailed descriptions and background information, while laws and regulations provide the legal and compliant standard basis for the model. This multimodal data fusion enables the model to conduct more comprehensive and accurate analysis. Traditional methods usually rely on a single modality (such as only using text or images for analysis), which may overlook important information in images or text. Multimodal learning allows the model to better understand different types of data and improves the model's performance in practical tasks. Through multimodal learning, the model can more comprehensively capture potential hazards in images and compare them with relevant texts and regulations, obtain more accurate rectification suggestions, efficiently integrate these heterogeneous knowledge sources, and form a more comprehensive judgment system;
[0094] One advantage of LoRA fine-tuning is that it can effectively reduce the need for labeled data. By leveraging pre-trained large models, the LoRA fine-tuning method can quickly improve the model performance based on a small amount of labeled data. Especially in the application of potential hazard investigation, due to the large volume and diversity of data, the workload of fully labeling all data is very large, and LoRA fine-tuning makes this process more efficient.
[0095] Step 4: Build a knowledge base
[0096] Based on the deep learning embedding model, convert the text data in the historical potential hazard investigation data into numerical vector representations, and store this knowledge in the knowledge base in vector form, making the retrieval process more efficient;
[0097] In this embodiment, the specific method for converting the text data in the historical potential hazard investigation data into a numerical vector representation based on a deep learning embedding model includes the following steps:
[0098] First is text preprocessing. The structured text data is segmented into meaningful text blocks, which can be words or sub-word units. Common stop words (such as "of", "is", etc.) are removed as they usually contribute less to semantics.
[0099] Secondly, the text blocks are restored to their basic forms (such as changing verbs back to their original forms) to reduce the lexical diversity, and at the same time, corresponding part-of-speech tags are added, which helps to understand the text structure more precisely.
[0100] Finally, the processed text blocks are converted into word vectors using a word embedding model: piccolo-large-zh, and stored in a vector number knowledge base. Common vector databases include Faiss, Annoy, HNSWLib, etc., which support fast approximate nearest neighbor search (ANN). Here we choose Faiss, create an index for the vector database to accelerate the retrieval process. The index can select different algorithms according to the application scenario, such as tree structures, hash tables, inverted indexes, etc.
[0101] The model can quickly find the regulations, standards, or inspection histories most relevant to the current potential hazard or problem from the knowledge base through similarity measurement in the vector space. The vectorized knowledge base uses a digital vector storage method and has high scalability. As laws, regulations, and industry standards increase, the knowledge base can be easily expanded without affecting the retrieval efficiency or model performance. The model can adapt to more regulatory data and adjust decision-making strategies according to the newly added information. Through the vectorized knowledge base, the model can generate rectification suggestions with legal basis for each potential hazard, ensuring that each potential hazard investigation is based on the same standard, avoiding the differences in human judgment, and ensuring the consistency and fairness of the application of regulations.
[0102] It is also worth mentioning that in this embodiment, the retrieval matching of the above-mentioned knowledge base specifically includes the following methods:
[0103] When a user enters a query, first convert it into the same vector representation as the knowledge base, search for the document vector most similar to the query vector in the vector database. The distance measurement methods used include cosine similarity, Euclidean distance, etc. Sort the retrieved results from largest to smallest according to the similarity score and return the most relevant documents or paragraphs.
[0104] Step 5: Potential hazard reasoning and retrieval
[0105] Design a multi-hop prompt template to perform multi-hop reasoning on potential hazards in the input image or video. Through a thinking chain that focuses on different details at different stages, guide the multi-dimensional analysis of the potential hazard factors that appear, comprehensively consider the information in each dimension, retrieve content similar to the results of the hazard reasoning in the knowledge base, re-examine the results of the hazard reasoning from multiple perspectives, verify the rationality of the reasoning results, and at the same time rely on the similar content in the knowledge base for self-correction, and re-examine the results of the hazard reasoning from multiple perspectives;
[0106] The specific method is as follows:
[0107] 1. Clarify the problem type. For the emergency fire protection field, it may include but is not limited to the following types of problems:
[0108] Hazard identification: Identify potential safety hazards based on the description;
[0109] Regulation query: Find laws, regulations, or industry standards related to specific situations.
[0110] 2. Construct a multi-hop prompt framework: The multi-hop prompt framework includes the following key parts:
[0111] Problem statement: Clearly describe the problem to be solved;
[0112] Initial information: Provide the necessary background information in the image content;
[0113] Reasoning steps: List the descriptions of the corresponding areas in the image content step by step according to the image content, and each step is equipped with the corresponding logical basis;
[0114] Auxiliary information: Introduce additional knowledge or data support when necessary;
[0115] Reasoning result: Output the reasoning result of each step and give a clear answer.
[0116] 3. Design a multi-hop prompt template
[0117] # Multi-hop prompt template
[0118] ## Problem statement
[0119] [Briefly describe the problem to be solved]
[0120] ## Initial information
[0121] [Provide the necessary background information, known conditions, or context]
[0122] Reasoning steps
[0123] # Step 1: [Description of the first step]
[0124] - **Operation**: [Specific operation or query]
[0125] - **Result**: [Record the result of the first step]
[0126] # Step 2: [Description of the second step]
[0127] - **Operation**: [Specific operation or query]
[0128] - **Result**: [Record the result of the second step] ...
[0130] # Step N: [Description of the Nth step]
[0131] - **Operation**: [Specific operation or query]
[0132] - **Result**: [Record the result of the Nth step]
[0133] ## Auxiliary Information
[0134] [Introduce additional knowledge, data, or reference materials]
[0135] ## Reasoning Results
[0136] [Input the content descriptions in each step into the knowledge base for retrieval to obtain corresponding potential safety hazards and rectification suggestions]
[0137] 4. Optimization and Expansion
[0138] Dynamic adjustment: Continuously optimize the prompt template according to the actual application scenario and user feedback, adding new reasoning steps or adjusting existing steps.
[0139] Automated generation: Develop tools or scripts to automatically generate multi-hop prompts, reducing the workload of manual writing.
[0140] Knowledge graph: Combine knowledge graph technology to structurally represent relevant regulations, standards, and cases, enhancing the accuracy and depth of reasoning.
[0141] User interaction: Allow users to input additional information or raise questions during the reasoning process, enabling the model to adjust and improve based on new information.
[0142] It is worth mentioning that in this embodiment, after completing the prompt framework template for multi-hop reasoning, the multi-hop reasoning and retrieval of potential hazards in image data specifically include the following steps:
[0143] Identify the content in the image data, and predict the potential hazard content in multiple dimensions and steps according to the identified image content and the constructed multi-hop prompt framework;
[0144] Input the described predicted potential hazard content into the knowledge base for retrieval to determine the corresponding laws and regulations for the corresponding safety hazards;
[0145] Input the generated safety hazards into the knowledge base for similarity matching. The CrossEncoder model combines the step-by-step described predicted hazard content with the corresponding safety hazard text blocks retrieved from the knowledge base for encoding. The model generates a relevance score based on this pair of combined inputs and sorts according to the model's score. The text block with the highest score is used as the final result to output the retrieved content, thereby obtaining the laws and regulations corresponding to each safety hazard in the image data;
[0146] Finally, obtain professional content from the knowledge base and generate corresponding rectification opinions for the identified safety hazards according to the corresponding laws and regulations.
[0147] It should be noted that in this embodiment, a noise filtering mechanism is set in the retrieval and matching process of the knowledge base to filter noise from the text data and image data before retrieval, and only high-quality data is retained for indexing and storage. After the user submits a query, the retrieval results are dynamically filtered to remove irrelevant information, and the noise filtering mechanism is continuously optimized according to the user's feedback;
[0148] The noise filtering mechanism can effectively identify and exclude content that is irrelevant or of low quality to the current inference task, thereby ensuring that the model only receives valuable knowledge. For example, for the task of safety hazard investigation, noise filtering can exclude those laws and regulations that are not relevant to specific hazard scenarios, making the retrieval results more in line with the task requirements. Especially in complex tasks (such as hazard detection), incorrect knowledge points may lead to misleading inference results of the model, thereby affecting the final decision. Through noise filtering, it can be ensured that the model obtains more accurate and high-quality inputs, improving the reliability of the inference results. Through the noise filtering mechanism, the amount of irrelevant information input to the model can be reduced, thereby reducing the computational burden in the subsequent inference stage. When the data input to the large model is more concise and relevant, the model can focus on key information during processing, improving the inference efficiency. The noise filtering mechanism can ensure that in cross-domain applications, the model is not interfered by noise information between different domains. This is especially important for the fusion of multi-modal data (such as images, texts, videos, etc.), because different types of data may contain different noise sources. Through noise filtering, the system can ensure that cross-domain information fusion is more efficient and accurate.
[0149] Step 6: Generate a reply. Finally, use the results of multi-hop reasoning and the content retrieved from the knowledge base as context for the large model to generate a reply regarding the safety hazards, the legal basis for the corresponding safety hazards, and the rectification opinions for the corresponding safety hazards:
[0150] The following is the content of the image data captured by all inspection personnel using mobile phones at 15:00 on December 4, 2024, in Area A of a certain production workshop: There is a large combustion boiler in the workshop, there is a small amount of coal cinder on the ground, and the surrounding environment is messy.
[0151] Therefore, a multi-hop prompt template is designed to perform multi-hop reasoning on potential hazards for the above image data.
[0152] # Hazard Identification
[0153] ## Problem Statement
[0154] Based on the provided on-site description, identify potential safety hazards, including but not limited to in-depth analysis in aspects such as equipment, materials, environment, and behavior.
[0155] ## Initial Information
[0156] - On-site description: There is a large combustion boiler in the workshop, there is a small amount of coal cinder on the ground, and the surrounding environment is messy.
[0157] - Time: 15:00 on December 4, 2024
[0158] - Location: Area A of the production workshop
[0159] - **Time alignment**: Ensure the timestamp synchronization of all sensor data.
[0160] # Step 1: Analyze the image or video using a fine-tuned multi-modal large model
[0161] - **Operation**: Identify environmental information, equipment information, behavior information, material information, and safety information in the picture or video.
[0162] - **Result**: The recognition result of the production picture or video, Environment: Dim workshop environment; Equipment: A combustion boiler; Behavior: No personnel behavior; Materials: Stacked messily; Safety: No safety warning information found.
[0163] # Step 2: Based on the information generated in the first step, identify potential safety hazards in the picture or video from multiple dimensions
[0164] - **Operation**: According to the environment, equipment, personnel behavior, materials, etc. generated in the first step, judge whether there are safety hazards in the picture.
[0165] - **Result**: Some potential safety hazards are found, Environment: Lack of proper lighting around; Equipment: Obvious corrosion and wear on the equipment surface; Behavior: No personnel behavior; Materials: Randomly stacked coal cinder on the ground; Safety: Lack of safety warning information.
[0166] # Step 3: Input the description of potential safety hazards into the knowledge base for retrieval, filter out irrelevant knowledge as noise, and output the corresponding valid legal and regulatory basis:
[0167] 1. Environment: Lack of proper lighting in the surroundings;
[0168] 2. Equipment: Obvious corrosion and wear on the surface of the equipment;
[0169] 3. Materials: Cinder is randomly piled on the ground;
[0170] 4. Safety: Lack of safety warning information;
[0171] Subsequently, input the content described in the above image or video into the knowledge base for similarity matching;
[0172]
[0173]
[0174] 1. "Building Construction Safety Management Specification 4.8.5", "Urban Road Maintenance Operation Regulations 12.0.8", "Technical Standard for Temporary Electricity Use Safety at Construction Sites of Buildings and Municipal Engineering 9.1.1";
[0175] 2. "Safety Production Standardization Specification for Machinery Manufacturing Enterprises 4.2.17.3", "Criteria for Judging Major Accident Hazards in Industrial and Trade Enterprises Article 6", "Safety Requirements for Cupolas and Cupola Charging Machines 5.6.2";
[0176] 3. "Measures of Hangzhou City for the Prevention and Control of Urban Dust Pollution Article 16", "Criteria for Judging Major Accident Hazards in Industrial and Trade Enterprises Article 6", "Safety Production Standardization Specification for Machinery Manufacturing Enterprises 4.2.17.3";
[0177] 4. "Criteria for Judging Major Accident Hazards in Industrial and Trade Enterprises Article 6", "Safety Production Standardization Specification for Machinery Manufacturing Enterprises 4.2.7";
[0178] # Step 4: Obtain professional knowledge from the knowledge base to determine the legal and regulatory basis for potential safety hazards.
[0179] - **Operation**: Check the content of relevant laws and regulations, confirm the legal basis for common potential safety hazards, and infer other possible potential safety hazards.
[0180] - **Result**: Finally, clarify the identified potential safety hazards:
[0181] 1. According to the requirements of "Building Construction Safety Management Specification" 4.8.5, the lack of proper lighting in the surroundings in the image is a potential safety hazard.
[0182] 2. According to the requirements of Article 4.2.17.3 of the "Safety Production Standardization Specification for Machinery Manufacturing Enterprises", obvious corrosion and wear on the surface of the boiler equipment in the image are potential safety hazards.
[0183] 3. According to Article 16 of the "Measures for the Prevention and Control of Urban Dust Pollution in Hangzhou", the random stacking of cinder on the ground in the image is a potential safety hazard.
[0184] 4. According to Article 6 of the "Judgment Criteria for Major Accident Hazards in Industrial and Trade Enterprises", the lack of safety signs in the image is a potential safety hazard.
[0185] # Step 5: Reasoning in Step 3, obtain professional knowledge from the knowledge base, and generate corresponding rectification opinions for the identified potential safety hazards.
[0186] - **Operation**: Check the content of relevant laws and regulations and potential safety hazards, and generate rectification opinions.
[0187] - **Result**: Generate rectification opinions for each potential hazard:
[0188] ## Auxiliary Information
[0189] - "Safety Management Specification for Construction Works"
[0190] - "Safety Production Standardization Specification for Machinery Manufacturing Enterprises"
[0191] - "Measures for the Prevention and Control of Urban Dust Pollution in Hangzhou"
[0192] - "Judgment Criteria for Major Accident Hazards in Industrial and Trade Enterprises"
[0193] ## Reasoning Result: Generate the final answer through the large model.
[0194] - **Operation**: Refer to the requirements of the corresponding data structure of the business system, and based on the language generation ability of the large model, generate the text of the final answer content.
[0195] # Step 6: Generate the final answer:
[0196] 1. Identification result: There is obvious corrosion and wear on the surface of the furnace body, especially around the furnace door.
[0197] Rectification opinion: Please check the daily maintenance records of the boiler furnace body structure to ensure that there are no potential safety hazards in the furnace body structure.
[0198] Legal basis: Article 4.2.17.3 of the "Safety Production Standardization Specification for Machinery Manufacturing Enterprises".
[0199] 2. Identification result: There are cinders and sundries on the ground, posing a risk of slipping or tripping.
[0200] Corrective suggestions: Please clean up the coal slag and debris on the ground in time to ensure that the ground is clean and tidy to prevent staff from slipping or tripping.
[0201] Legal basis: Article 16 of the Hangzhou Urban Dust Pollution Prevention and Control Management Measures.
[0202] 3. Identification results: There are no obvious safety signs or emergency stop devices.
[0203] Corrective suggestions: Please set up corresponding safety signs and equipment emergency stop devices as required.
[0204] Legal basis: Article 6 of the "Standards for Determining Major Accident Hazards in Industrial and Commercial Enterprises".
[0205] 4. Identification result: Lack of proper lighting around the furnace may affect operational safety.
[0206] Corrective suggestions: Please keep the surrounding light clean and not affect the safety of operation.
[0207] Legal basis: "Construction Safety Management Code" 4.8.5.
[0208] By retrieving content similar to the current reasoning results in the knowledge base, more relevant background information or case support can be provided to the large model. These similar knowledge points can be used as a reference for the model, allowing the model to make more professional and accurate judgments on the existing basis. The knowledge base contains information such as industry regulations, standards, and best practices related to hidden danger detection. The model can retrieve content similar to the reasoning results and combine it with domain-specific expertise for further analysis. This approach can enhance the sensitivity of the model in specific areas (such as safety hazard identification), enabling it to capture more details and potential hazards.
[0209] By introducing content similar to the current reasoning results, the model can re-examine the problem from multiple angles and verify whether the reasoning results are reasonable. This approach enhances the model's fault tolerance and enables it to correct potential misjudgments, especially when faced with ambiguous or uncertain information. The model can rely on similar content in the knowledge base to correct itself.
[0210] The above is only a specific implementation of the invention, but the protection scope of the invention is not limited to it. Any changes or substitutions that are not conceived through creative work should be included in the protection scope of the invention. Therefore, the protection scope of the invention should be based on the protection scope defined in the claims.
Claims
1. An emergency fire hazard inspection method based on multi-modal AI large model recognition technology, characterized in that: It includes the following steps: Step 1: Data processing. Collect historical hidden danger investigation data containing images as multi-modal data, uniformly organize the data, and label the safety hidden dangers with corresponding hidden danger levels; Step 2: Image processing. Add an adaptive image coding mechanism to the image data to directly process images with various resolutions and aspect ratios; Step 3: Fine-tune the large model in the LoRa way. Use the above-collected and processed data, introduce LoRa layers in each layer of the basic model, set training parameters, and only update the parameters in the LoRa layers to enable the model to quickly adapt to specific tasks; Step 4: Vectorize text data. Convert the text data in the historical hidden danger investigation data into a numerical vector representation based on a deep learning embedding model, store this knowledge in the knowledge base in vector form, and make the retrieval process more efficient; Step 5: Hidden danger reasoning and retrieval. Design a multi-hop prompt template to perform multi-hop reasoning on the input image data for hidden dangers. Guide the multi-dimensional analysis of the emerging hidden danger factors through a chain of thought with different attention details at different stages, comprehensively consider the information in each dimension, retrieve content similar to the result of hidden danger reasoning in the knowledge base, and re-examine the result of hidden danger reasoning from multiple angles. It specifically includes the following steps: Identify the content in the image data, and multi-dimensionally and step-by-step describe the predicted hidden danger content according to the identified image content and the constructed multi-hop prompt framework; Input the described predicted hidden danger content into the knowledge base for similarity matching. The specific method is as follows: Combine the step-by-step described predicted hidden danger content with the corresponding safety hidden danger text block retrieved in the knowledge base through the CrossEncoder model for encoding. The model generates a relevance score based on this pair of combined inputs, sorts according to the score of the model, and outputs the text block with the highest score as the final result of the retrieved content, and output the content with the highest relevance score as the retrieved content to determine the corresponding laws and regulations for each safety hidden danger; Retrieve professional content from the knowledge base, and generate corresponding rectification opinions for the identified safety hidden dangers according to the corresponding laws and regulations; Step 6: Generate a reply. Use the result of hidden danger reasoning and the content retrieved from the knowledge base as context to the large model to generate a reply.
2. The emergency fire hazard inspection method based on the multi-modal AI large model recognition technology according to claim 1, characterized in that: The historical hidden danger investigation data in Step 1 also includes text data of hidden danger report documents, equipment operation specifications, industry standards, and relevant laws and regulations. Organize the text data into structured electronic data, where the text data of the laws and regulations is organized into a csv format, with each law as a cell. At the same time, label the hidden danger levels and hidden danger descriptions for the safety hidden dangers existing in each picture in the historical hidden danger investigation image data, and establish a mapping relationship between each law and regulation and the safety hidden dangers.
3. The emergency fire hazard inspection method based on the multi-modal AI large model recognition technology according to claim 2, wherein: The specific processing method of the adaptive image coding mechanism for pictures in Step 2 includes the following steps: First, scale the picture to the specified size, encode it using VIT, and merge and flatten the encoded picture; Secondly, divide the picture into several parts according to the set specification size, encode each part using VIT respectively, and merge and flatten the encoded pictures; Finally, the flattened encoded images are vector - stitched twice to obtain the final image encoding.
4. The emergency fire hazard inspection method based on the multi-modal AI large model recognition technology according to claim 3, characterized in that: In step 4, the specific method of converting the text data in the historical hidden danger investigation data into a numerical vector representation based on a deep - learning embedding model includes the following steps: S101: Split the structured text data into meaningful text blocks and remove common stop words at the same time; S102: Restore the text blocks to their basic forms, reduce the diversity of vocabulary, and add corresponding part - of - speech tags at the same time; S103: Convert the processed text blocks into word vectors using a word - embedding model, store them in a vector database, create an index for the database, and enable the database to perform retrieval and matching.
5. The emergency fire hazard inspection method based on the multi-modal AI large model recognition technology according to claim 4, characterized in that: The retrieval and matching of the database specifically include the following method: Convert the user - input query into the same vector as the database, use the cosine - similarity calculation method to find the text - data vector most similar to the query vector in the database, and sort the retrieved results according to the similarity score.
6. The emergency fire hazard inspection method based on the multi-modal AI large model recognition technology according to claim 5, characterized in that: A noise - filtering mechanism is set in the retrieval and matching process of the database. The specific method is as follows: Pre - process the text data and image data before retrieval, remove irrelevant information, and only retain high - quality data for indexing and storage. After the user submits a query, dynamically filter the retrieval results to remove irrelevant information, and continuously optimize the noise - filtering mechanism according to the user's feedback.
7. The emergency fire hazard inspection method based on the multi-modal AI large model recognition technology according to claim 6, characterized in that: The content of the response generated in step 6 includes a description of the safety hazard, the corresponding laws and regulations for the safety hazard, and the corresponding rectification opinions for the safety hazard. Each rectification opinion corresponds to relevant laws and regulations.
Citation Information
Patent Citations
Knowledge reasoning method based on multi-modal knowledge graph
CN112288091A
Intelligent question answering method and system based on domain knowledge graph
CN117648984A