A safety production supervision system and method based on a multimodal large model

The multi-modal large model-based system automates safety hazard detection in enterprises, improving efficiency and reducing human workload while ensuring consistent and timely hazard identification and response.

CN119274142BActive Publication Date: 2025-07-15JIANGSU YUNSHEN INTELLIGENT SYST CO LTD

Patent Information

Application Number
CN202411785755.7
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2024-12-06
Publication Date
2025-07-15
Estimated Expiration
2044-12-06

AI Technical Summary

Technical Problem

Enterprise safety hazard inspections rely on manual inspections, resulting in untimely discovery and large workloads, fatigue of managers, and easy to omit safety issues.

Method used

The safety production supervision system based on multimodal large models is adopted, and through technologies such as multimedia information collection, visual large models, databases, text similarity comparison, image segmentation labeling and knowledge base search, it automatically detects and provides suggestions for damage rectification measures.

Benefits of technology

Achieve all-weather safety hazard monitoring, improve detection efficiency, reduce manual workload, promptly warning and deal with hidden dangers, and reduce labor costs.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119274142B_ABST
    Figure CN119274142B_ABST
Patent Text Reader

Abstract

The present invention discloses a safety production supervision system and method based on a multimodal large model, aiming to solve the problem of large workload of manual inspection for potential safety hazards in the existing enterprise production scenarios. By obtaining the image information and location information of the scene where potential hazards need to be identified, technologies such as vision large models and large language models are used to detect and mark potential safety hazards that may occur during the enterprise production process, and give rectification suggestions according to the corresponding safety regulations. The present invention can monitor potential safety hazards in the enterprise production process all-weather, significantly improve the efficiency of hazard monitoring, reduce the manual workload, lower the labor cost, and provide a basic guarantee for the safety production of enterprises.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention belongs to the technical field of work safety, and specifically relates to a work safety supervision system and method based on a multimodal large model. Background Art

[0002] During the daily production process of enterprises, there are potential safety hazards. In order to reduce and effectively avoid the occurrence of potential safety hazards, it is essential to detect and investigate potential hazards. Currently, the work of investigating potential safety hazards in enterprises mainly relies on safety management personnel to conduct on-site manual observations and investigations in accordance with the safety standards and specifications of each enterprise. However, the abilities of safety management personnel in different production enterprises to detect professional potential safety hazards vary, and at the same time, the huge workload of patrol inspections is also likely to cause fatigue among safety management personnel, resulting in inadequate detection of safety problems and easy omission of abnormal phenomena. Therefore, standardizing and automating the identification of production safety hazards and reducing the manual patrol inspection workload have become important requirements in enterprise work safety. Summary of the Invention

[0003] To solve the above problems, the present invention discloses a work safety supervision system and method based on a multimodal large model, which can monitor potential safety hazards in the enterprise production process all-weather, significantly improve the efficiency of potential hazard monitoring, reduce the manual workload, lower the labor cost, and provide a basic guarantee for the work safety of enterprises.

[0004] To achieve the above object, the technical solution of the present invention is as follows:

[0005] A work safety supervision system based on a multimodal large model, comprising:

[0006] A multimedia information collection terminal for collecting image information and location information in the enterprise production scenario, and the multimedia information collection terminal includes, but is not limited to, monitoring devices, smart phones, and inspection glasses with shooting and positioning functions;

[0007] A visual large model for detecting, identifying, and describing image content information and generating an image description;

[0008] A database for storing the corresponding relationship between location information and the potential hazard investigation scope, the corresponding relationship between potential hazard descriptions and potential hazard standard prompt words, and the finally detected potential hazard information;

[0009] A text similarity comparison module for receiving the image description and the potential hazard description and outputting the text similarity therebetween;

[0010] A threshold judgment module for judging whether a given text similarity is higher than a given threshold;

[0011] Standard Prompt Matching Module: Used to match the potential hazard standard prompts corresponding to the potential hazard description in the database. The potential hazard standard prompts are applicable to subsequent image segmentation annotation and knowledge base retrieval tasks;

[0012] Image Segmentation Annotation Module, including an image segmentation annotation model, used to segment and annotate potential hazard areas on potential hazard images based on potential hazard standard prompts;

[0013] Knowledge Base Retrieval Module, including a RAG knowledge base and a large language model, used to retrieve the source of potential hazards in safety standard documents in the RAG knowledge base, and use the large language model to summarize and output suggestions for potential hazard rectification measures;

[0014] Result Output Module, used to output and record the results of potential hazard investigation, including whether there are potential hazards, the time of potential hazard occurrence, the location of potential hazard occurrence, the name of potential hazard type, the potential hazard segmentation annotation image, suggestions for potential hazard rectification measures, and give early warning prompts.

[0015] Furthermore, for the Knowledge Base Retrieval Module, it includes a RAG knowledge base and a large language model:

[0016] For the RAG knowledge base, it contains different text segments formed after splitting the text in safety standard documents, which can be used to complete the semantic matching retrieval between each text segment and potential hazard standard prompts, and take the text content with the highest matching degree as the source information of the potential hazard in the safety standard document;

[0017] The large language model includes a locally deployed or cloud-deployed large language model, and analyzes and summarizes suggestions for potential hazard rectification measures based on the source information of the potential hazard in the safety standard document retrieved from the RAG knowledge base and preset guiding words.

[0018] Furthermore, the splitting of the safety standard document text in the RAG knowledge base: When facing text with multi-level heading calibration, first, based on the first-level headings, split the content under each first-level heading into corresponding numbers of text segments; then, based on the second-level headings, split the content under the second-level headings again, and so on, splitting the text content hierarchically according to different levels of headings, so that the split text segments include both long text segments containing the overall context and context of the text, and short text segments containing only information related to specific sub-topics, focusing on local key information, and can accurately match specific retrieval directions; when facing text without multi-level heading calibration, send the text to the large language model, and after the large language model completes the multi-level heading calibration, operate according to the above steps for splitting text with multi-level heading calibration.

[0019] The work safety supervision method implemented based on the above system can collect image information and location information in the enterprise production scenario, and use technologies such as visual analysis and natural language processing to detect whether there are potential safety hazards, segment and label the detected hazards, provide rectification measure suggestions, and give early warnings. The steps are as follows:

[0020] S1. The multimedia information collection terminal collects image information and location information in the enterprise production scenario;

[0021] S2. The image information is transmitted to the visual large model to generate an image description;

[0022] S3. Search in the database for the hazard descriptions within the hazard investigation scope that match the corresponding location information;

[0023] S4. The image description and the hazard description are transmitted to the text similarity comparison module to obtain the text similarity between the two;

[0024] S5. The text similarity is transmitted to the threshold judgment module, and it is judged whether there are hazards according to the relationship between the text similarity and the threshold;

[0025] S6. According to the hazard description of the existing hazard, match its corresponding hazard standard prompt words in the database;

[0026] S7. The image information and the hazard standard prompt words are transmitted to the image segmentation and labeling module for hazard area segmentation and labeling;

[0027] S8. The hazard standard prompt words are transmitted to the knowledge base retrieval module to retrieve the source of the hazard in the safety standard documents in the RAG knowledge base, and the large language model is used to summarize and output the hazard rectification measure suggestions;

[0028] S9. Output and record the hazard investigation results, including whether there are hazards, the hazard occurrence time, the hazard occurrence location, the hazard type name, the hazard segmentation and labeling image, the hazard rectification measure suggestions, and give an early warning to prompt the safety management personnel.

[0029] Further, in step S3, the hazard investigation scope to be checked corresponding to the location information is searched in the database. In the database, the input location information is retrieved, and the descriptions of multiple different hazards within the hazard investigation project scope matching the location information are output.

[0030] Further, in step S8, the safety standard documents are pre - segmented into text segments of different lengths in the RAG knowledge base according to a hierarchical segmentation rule, and a parameter N is set. The parameter N is the number of sources of potential hazards retrieved in the safety standard documents in the RAG knowledge base. When the potential hazard standard prompt word is input into the knowledge base retrieval module, the semantic similarity between the segmented text segments and the potential hazard standard prompt word is calculated, and the top N text segments with the highest semantic similarity to the potential hazard standard prompt word are input into the large language model.

[0031] The beneficial effects of the present invention are as follows:

[0032] The present invention detects, investigates, and disposes of potential safety hazards existing in the enterprise production process through the application of artificial - intelligence - based methods such as visual large models and large language models. It can monitor potential safety hazards in the enterprise production scenario all - weather, with high timeliness, a standard - unified output format, significantly improving the potential hazard detection frequency, reducing the manual workload, and providing a basic guarantee for the safe production of enterprises.

[0033] By obtaining the position information corresponding to the image information and querying the corresponding potential hazard investigation scope in the database, the scope of potential hazard investigation is narrowed, the time for potential hazard detection is shortened, the speed of potential hazard detection is increased, and potential hazards can be warned and disposed of more promptly. BRIEF DESCRIPTION OF THE DRAWINGS

[0034] Figure 1 It is a schematic diagram of the system structure of the present invention.

[0035] Figure 2 It is a schematic diagram of the process of the present invention. DETAILED DESCRIPTION OF THE INVENTION

[0036] The following further clarifies the present invention in conjunction with the drawings and specific embodiments. It should be understood that the following specific embodiments are only used to illustrate the present invention and not to limit the scope of the present invention.

[0037] A safety production supervision system and method based on a multi - modal large model according to the present invention aims at the problem that the investigation of potential safety hazards in enterprise production currently relies on manual labor and the inspection workload is large. It uses artificial - intelligence - based methods such as visual large models and large language models to detect and investigate potential safety hazards existing in the enterprise production process, which can reduce the manual workload and conduct all - weather monitoring and management of potential safety hazards.

[0038] As Figure 1 shown, the system structure of the present invention includes:

[0039] A multimedia information collection terminal, which is responsible for inputting image information into the visual large model and inputting position information into the database;

[0040] The visual large model is responsible for generating image description texts and transmitting the image description texts to the text similarity comparison module;

[0041] The database is responsible for transmitting the potential hazard description texts to be investigated to the text similarity comparison module;

[0042] The text similarity comparison module is responsible for calculating the text similarity between the image description text and the potential hazard description text to be investigated and transmitting it to the threshold judgment module;

[0043] The threshold judgment module is responsible for comparing whether the input text similarity is higher than the given threshold to judge whether there is a potential hazard;

[0044] The standard prompt word matching module is responsible for transmitting the standard prompt words corresponding to the potential hazard description text to the image segmentation annotation module and the knowledge base retrieval module;

[0045] The image segmentation annotation module is responsible for segmenting and annotating the potential hazard areas in the image;

[0046] The knowledge base retrieval module is responsible for retrieving the source of the potential hazard in the safety standard documents and outputting suggestions for potential hazard rectification measures;

[0047] The result output module is responsible for displaying the results of potential hazard investigation.

[0048] As Figure 2 shown, the specific steps for the implementation and operation of the present invention include:

[0049] S1. The multimedia information collection terminal collects image information and location information in the enterprise production scenario:

[0050] Use terminal devices that can obtain image information and location information, including but not limited to monitoring devices, smartphones, inspection glasses with shooting and positioning functions, etc. to collect image information and corresponding location information in the enterprise production scenario. For example, a certain image information is an image taken by the mobile phone in Warehouse No. 2 of the enterprise, and the location information is obtained by the positioning function of the mobile phone: "Warehouse No. 2".

[0051] S2. Transmit the image information to the visual large model to generate image description texts:

[0052] According to the image information collected by the multimedia information collection terminal, the visual large model such as CLIP generates image description texts corresponding to the image information. For example, "There is a fire extinguisher on the left side of the image, and there is rust on the fire extinguisher".

[0053] S3. Search in the database for potential hazard description texts within the scope of potential hazard investigation that match the corresponding location information:

[0054] According to the location information collected by the multimedia information acquisition terminal, search for the hidden danger description language within the scope of hidden danger investigation that matches the corresponding location information in the database. For example, if the location information is Warehouse No. 2, the hidden danger description languages within the corresponding scope of hidden danger investigation include "Fire-fighting equipment is damaged, rusty, or blocked", "The safety exit sign is tilted and fallen off".

[0055] S4. Input the image description language and the hidden danger description language into the text similarity comparison module to obtain the text similarity between them:

[0056] Perform a text similarity comparison between the image description language, such as "There is a fire extinguisher in the image and there is rust on the fire extinguisher", and the hidden danger description language, such as "The fire extinguisher is damaged, rusty, or blocked", "The safety exit sign is tilted and fallen off", calculate the semantic similarity degree between the texts. The text similarity comparison algorithm includes the Word2Vec algorithm, etc., to obtain the text similarity between them.

[0057] S5. Input the text similarity into the threshold judgment module, and judge whether there is a hidden danger according to the relationship between the text similarity and the threshold:

[0058] Compare the similarity between the image description language and the hidden danger description language with the preset text similarity threshold. If the similarity between the image description language and a certain hidden danger description language is greater than or equal to the threshold, it is determined that there is a hidden danger corresponding to the hidden danger description language in the image. If the similarity between the image description language and a certain hidden danger description language is less than the threshold, it is determined that there is no hidden danger corresponding to the hidden danger description language in the image. If there is no hidden danger, jump to step S9 to directly output that there is no hidden danger. If there is a hidden danger, perform the following steps.

[0059] S6. According to the hidden danger description language with hidden danger, match its corresponding hidden danger standard prompt word in the database:

[0060] Query and match the hidden danger standard prompt word corresponding to the hidden danger description language in the database according to the hidden danger description language. For example, if the hidden danger description language is "The fire extinguisher is rusty", the corresponding hidden danger standard prompt word is "The fire extinguisher box is rusty". The hidden danger standard prompt word can be better applied to subsequent image segmentation annotation and knowledge base retrieval tasks.

[0061] S7. Input the image information and the hidden danger standard prompt word into the image segmentation annotation module to perform hidden danger area segmentation and annotation:

[0062] Use an image segmentation annotation model, such as Grounded-SAM, to detect the pixel coordinate position of the hidden danger in the image according to the input image information and the hidden danger standard prompt word, add a dark semi-transparent mask to the hidden danger area, and annotate the hidden danger standard prompt word with text beside the hidden danger area.

[0063] S8. Input the potential hazard standard prompt words into the knowledge base retrieval module, retrieve the source of the potential hazard in the safety standard documents in the RAG knowledge base, and use the large language model to summarize and output suggestions for potential hazard rectification measures:

[0064] First, pre-segment the text of the safety standard documents according to a specific segmentation method. The segmentation method is multi-level heading loop nested text segmentation, including:

[0065] If the safety standard documents have been marked with multi-level headings (i.e., including first-level, second-level, third-level and above-level headings), start the hierarchical segmentation mechanism; first, based on the first-level headings, divide the content below each of them into independent text segments, and the number of segmented text segments is the same as the number of first-level headings; then, based on the first-level headings, divide the content below each second-level heading into independent text segments, and the number of segmented text segments is the same as the number of second-level headings; and so on, perform segmentation operations on subsequent levels of headings in turn until the last-level heading; if the safety standard documents have not been marked with multi-level headings (i.e., do not include first-level, second-level, third-level and above-level headings), then input the safety standard documents into the large language model, and with the help of the understanding and analysis capabilities of the large language model, mark multi-level headings for the text. After the heading marking is completed, perform operations according to the above-established segmentation process.

[0066] The RAG knowledge base will match the top N segments (N can be freely set) with the highest semantic similarity in the segmented text segments according to the input potential hazard standard prompt words, input these N segments of content into the large language model, and at the same time send guiding words, such as "The safety hazard detected by the user is: <potential hazard standard prompt words>, please give suggestions for potential hazard rectification measures according to the content retrieved from the knowledge base".

[0067] S9. Output and record the results of potential hazard investigation, including whether there is a potential hazard, the time of potential hazard occurrence, the location of potential hazard occurrence, the name of potential hazard type, the segmented and marked image of potential hazard, suggestions for potential hazard rectification measures, and give a warning prompt:

[0068] On the front end, including the web end and the client application interface, display whether there is a potential hazard. If there is no potential hazard, output that no potential hazard is detected; if there is a potential hazard, output that a potential hazard is detected, output the time of potential hazard occurrence, the location of potential hazard occurrence, the name of potential hazard type, the segmented and marked image of potential hazard, suggestions for potential hazard rectification measures, and store the above information in the database for subsequent viewing; at the same time, give a warning prompt to notify the person in charge of work safety to handle the safety hazard in time.

Claims

1. A safety production supervision system based on a multimodal large model, characterized in that, Including: A multimedia information collection terminal for collecting image information and location information in an enterprise production scenario. The multimedia information collection terminal includes monitoring devices, smartphones, and inspection glasses with shooting and positioning functions; A vision large model for receiving image information and generating image descriptions; A database storing the correspondence between location information and potential hazard inspection scopes, the correspondence between potential hazard descriptions and standard prompt words for potential hazards, and the finally detected potential hazard information; A text similarity comparison module for receiving the image description and the potential hazard description corresponding to the potential hazard inspection scope in the database and outputting the text similarity therebetween; A threshold judgment module for judging whether the text similarity output by the text similarity comparison module is higher than a given threshold; A standard prompt word matching module: for matching the standard prompt word for potential hazards corresponding to the potential hazard description in the database, and the standard prompt word for potential hazards is applicable to subsequent image segmentation annotation and knowledge base retrieval tasks; An image segmentation annotation module, including an image segmentation annotation model, for segmenting and annotating potential hazard areas on potential hazard images based on the standard prompt word for potential hazards; A knowledge base retrieval module, including a RAG knowledge base and a large language model, for retrieving the source of the potential hazard in the safety standard document in the RAG knowledge base and using the large language model to summarize and output suggestions for potential hazard rectification measures; A result output module for outputting and recording the potential hazard inspection results, including whether there are potential hazards, the occurrence time of potential hazards, the occurrence location of potential hazards, the name of potential hazard types, the potential hazard segmentation annotation images, the suggestions for potential hazard rectification measures, and giving early warning prompts.

2. The safety production supervision system based on the multimodal large model according to claim 1, characterized in that For the knowledge base retrieval module, it includes a RAG knowledge base, a large language model: For the RAG knowledge base, it contains different text segments formed by splitting the text in the safety standard document, for completing the semantic matching retrieval between each text segment and the standard prompt word for potential hazards, and taking the text content with the highest matching degree as the source information of the potential hazard in the safety standard document; The large language model, including a locally deployed or cloud-deployed large language model, analyzes and summarizes to obtain suggestions for potential hazard rectification measures based on the source information of the potential hazard in the safety standard document retrieved from the RAG knowledge base and the preset guiding words.

3. The safety production supervision system based on a multi-modal large model according to claim 2, characterized in that, Text segmentation of the safety standard document in the RAG knowledge base: When facing text with multi-level title calibration, first, based on the first-level titles, the content under each first-level title is split into corresponding numbers of text segments; then, based on the second-level titles, the content under the second-level titles is split again, and so on, splitting the text content hierarchically according to different levels of titles, so that among the split text segments, there are both long text segments containing the overall context and context connection of the text, and short text segments containing only information related to specific sub-topics, focusing on local key information, and capable of accurately adapting to specific retrieval directions; when facing text without multi-level title calibration, the text is sent into the large language model, and the large language model performs multi-level title calibration. After completion, the operation is carried out according to the above text segmentation steps for text with multi-level title calibration.

4. A supervision method for a safety production supervision system based on a multi-modal large model as described in any one of claims 1-3, characterized in that, Including the following steps: S1. The multimedia information collection terminal collects image information and location information in an enterprise production scenario; S2. Input the image information into the vision large model to generate an image description. S3. Search in the database for the hazard descriptions within the hazard investigation scope that match the corresponding location information. S4. Input the image description and the hazard description into the text similarity comparison module to obtain the text similarity between the two. S5. Input the text similarity into the threshold judgment module to determine whether there is a hazard based on the relationship between the text similarity and the threshold. S6. According to the hazard description of the existing hazard, match its corresponding hazard standard prompt words in the database. S7. Input the image information and the hazard standard prompt words into the image segmentation and annotation module for hazard area segmentation and annotation. S8. Input the hazard standard prompt words into the knowledge base retrieval module to retrieve the source of the hazard in the safety standard document in the RAG knowledge base, and use the large language model to summarize and output suggestions for hazard rectification measures. S9. Output and record the hazard investigation results, including whether there is a hazard, the hazard occurrence time, the hazard occurrence location, the hazard type name, the hazard segmentation and annotation image, the hazard rectification measure suggestions, and give a warning prompt.

5. The supervision method according to claim 4, characterized in that In step S3, the hazard scope to be investigated corresponding to the input location information is searched in the database. In the database, the input location information is retrieved, and the descriptions of multiple different hazards within the hazard investigation project scope matching the location information are output.

6. The supervision method according to claim 4, wherein In step S8, in the RAG knowledge base, the safety standard document is pre - segmented into text segments of different lengths according to the hierarchical segmentation rule, and a parameter N is set. The parameter N is the number of text segments in the RAG knowledge base for retrieving the hazard in the safety standard document. When the hazard standard prompt words are input into the knowledge base retrieval module, the semantic similarity between the segmented text segments and the hazard standard prompt words is calculated, and the top N text segments with the highest semantic similarity to the hazard standard prompt words in the text segments are input into the large language model.

Citation Information

Patent Citations

  • Potential safety hazard identifying and labeling method based on image-text conversion model

    CN117036778A

Cited By

  • Mine underground operation major hidden danger judgment and research method based on multi-modal large model fusion

    CN122798167A