General security risk monitoring method and system based on large model capability
By combining image scene recognition and knowledge retrieval technology, the security risk monitoring method with large model capabilities is adopted to solve the problems of narrow application scenarios, poor real-time and low accuracy in the existing technology, and efficient and intelligent security risk monitoring is achieved, which is suitable for diverse application scenarios.
Patent Information
- Application Number
- CN202510058786.0
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-01-15
- Publication Date
- 2025-05-09
AI Technical Summary
The existing security risk monitoring methods have problems such as narrow application scenarios, poor real-time, low accuracy, slow response speed and limited data processing capabilities. They are poor in versatility, making it difficult to adapt to diverse application scenarios.
A general security risk monitoring method based on large-model capabilities is adopted, combined with image scene recognition and knowledge retrieval technology, real-time or asynchronous polling monitoring and intelligent analysis of security risks in different fields is achieved through image recognition models, multimodal large models, RAG technology and large language models.
It significantly improves the ability to identify and manage potential risks, improves the accuracy, real-time and versatility of security risk monitoring, realizes high-efficiency intelligent monitoring, and supports diverse application scenarios.
Smart Images

Figure CN119964083A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of computer vision and intelligent information processing, and more specifically to a general security risk monitoring method and system based on large model capabilities. Background Art
[0002] At present, the existing security risk monitoring methods mainly rely on manual inspections, manual verification of dual-recording cameras, simple sensor data monitoring, or the use of single-point AI capabilities for image anomaly detection. The current risk monitoring methods have the following significant shortcomings: Narrow application scenarios: The existing dual-recording system mainly focuses on security needs. There is still no good algorithm and application solution for long-tail and low-frequency security risk monitoring in different business scenarios; Poor real-time performance: Manual inspections cannot achieve comprehensive monitoring and are prone to miss key events. Although sensor monitoring can monitor in real time, its coverage is limited and is easily affected by environmental interference; Low accuracy: Manual inspections rely on personnel experience and judgment, and due to the huge workload and high professional requirements in professional scenarios, misjudgments or missed judgments may occur. Sensor monitoring can only provide limited data and it is difficult to fully reflect the complex scenarios. security risks; Slow response speed: Traditional methods often take a long time to respond and process after discovering security risks, and are unable to take effective measures quickly; Limited data processing capabilities: Most of the existing monitoring systems lack strong data processing and analysis capabilities due to their early construction, making it difficult to extract valuable information from large amounts of data, and unable to achieve intelligent risk warning and management; High AI reasoning cost: The AI risk monitoring and early warning projects currently in use through special construction generally have the problems of single capabilities, high computing power requirements, many false alarms, and poor mobility. Data sharing across commissions and bureaus cannot share AI capabilities due to different business perspectives, and there is a problem of high AI reasoning cost; Poor versatility: Security risk monitoring needs in different fields vary, and existing monitoring methods and systems are often difficult to adapt to diverse application scenarios, lacking versatility and scalability.
[0003] Therefore, how to improve the accuracy, real-time and versatility of security risk monitoring and achieve efficient intelligent monitoring is an urgent problem that technical personnel in this field need to solve. Summary of the invention
[0004] In view of this, the present invention provides a general safety risk monitoring method and system based on large model capabilities. By combining image scene recognition and knowledge retrieval technology, it can flexibly respond to safety risk monitoring needs in different fields and effectively improve the safety management level and emergency response capabilities in various scenarios.
[0005] In order to achieve the above object, the present invention adopts the following technical solution:
[0006] A general safety risk monitoring method based on large model capabilities includes the following steps:
[0007] Step 1: Obtain video data and preprocess it to obtain the image to be recognized;
[0008] Step 2: Use the image recognition model and the multimodal large model to perform image scene recognition on the image to be recognized, and obtain the image scene label description;
[0009] Step 3: Use RAG technology to retrieve knowledge based on the image scene label description and the image to be identified, and recall knowledge;
[0010] Step 4: Use a large language model to generate security risk monitoring results based on the recalled knowledge and constructed risk warnings.
[0011] The technical effect of the above technical solution is that, by combining image scene recognition and knowledge retrieval technology, the ability to identify and manage potential risks can be significantly improved.
[0012] Preferably, the method for acquiring video data includes receiving video data from an external system and collecting video data using a surveillance camera.
[0013] Preferably, the process of preprocessing the video data includes:
[0014] Step 11: Decode the video data to obtain image frame sets and audio;
[0015] Step 12: Index the image frame set in time order, and extract the image frames of the set time as the images to be identified;
[0016] Step 13: Perform image processing on the image to be identified, including format conversion, scaling, and cropping, etc. To meet the analysis requirements and store in a suitable format, the image frame can be stored in image formats such as JPEG and PNG.
[0017] The technical effect of the above technical solution is that it can effectively decompose the video obtained from the surveillance camera or video file into independent video images and audio data, thereby providing data support for various applications such as computer vision analysis and security monitoring.
[0018] Preferably, the specific process of performing image scene recognition on the image to be recognized in step 2 is:
[0019] Step 21: Input the image to be identified and the image scene category label text pre-set based on security risks into the trained image recognition model, and output the highest image scene category prediction probability;
[0020] Step 22: If the output image scene category prediction probability is greater than the preset probability threshold, the image scene category label corresponding to the image to be identified is determined according to the image scene category prediction probability, and the image to be identified and the corresponding scene category label are respectively input into the trained multimodal large model to obtain the corresponding image scene label description; otherwise, the image to be identified is input into the trained multimodal large model, and the query is input at the same time to generate the image scene label description. The probability threshold can be set to a value in the range of (0,1), and the default value is 0.55. The greater the deviation of the probability threshold from the default value, the lower the sensitivity of recognition, and the smaller the absolute value deviation of the probability threshold from the default value, the higher the sensitivity of recognition.
[0021] Preferably, the image recognition model can be selected from the EVA-CLIP-8B model, and the specific implementation process of step 21 is:
[0022] Step 211: pre-set image scene category label text based on security risk, and perform text encoding to obtain text features;
[0023] Step 212: pre-processing the image to be identified and normalizing it;
[0024] Step 213: Input the normalized image and text features into the EVA-CLIP-8B model respectively, extract high-level visual features, and compare the high-level visual features with the text features corresponding to the preset image scene category label text to determine the image scene category prediction probability of each image scene category label, and output the highest image scene category prediction probability.
[0025] Preferably, the multimodal large model can use CogVLM2, BLIP3 or GLM4 model; the image to be identified is input into the trained multimodal large model, and the query is input at the same time, and the specific process of generating the image scene label description is as follows:
[0026] Step 221: input the image to be recognized into the multimodal large model;
[0027] Step 222: constructing a query input related to the image scene into a multimodal large model;
[0028] Step 223: The multimodal large model extracts image information from the image to be identified, combines the query with the multimodal processing mechanism, and generates a label description of the image scene. The query can be set as "what is the scene of the image?".
[0029] Preferably, the RAG technology in step 3 is implemented using the RAGFlow system, and the specific process includes:
[0030] Step 31: Input the image scene label description into the RAGFlow system, use matching algorithms and speech analysis algorithms to retrieve knowledge from preset business knowledge documents, and recall legal regulations or hidden danger identification rule knowledge;
[0031] Step 32: Normalize the image to be recognized and input it into the trained image encoder to generate an embedding vector. Input the embedding vector into the RAGFlow system and use the matching algorithm and speech analysis algorithm to retrieve knowledge from the preset scene-related knowledge base and recall the scene knowledge.
[0032] Preferably, step 31 also includes constructing a prompt word prompt and inputting it into the RAGFlow system, and the RAGFlow system performs knowledge retrieval in combination with the prompt word prompt.
[0033] Preferably, the specific process of step 4 is:
[0034] Step 41: Construct risk warning;
[0035] Step 42: Input the embedding vector, legal regulations or hidden danger identification rule knowledge, scenario knowledge and risk warning into the large language model to obtain the safety risk monitoring results.
[0036] Preferably, a risk warning text is set, and the risk warning text is rewritten into a risk warning based on a chain of thought, and the risk warning is converted into a structured framework prompt and then input into a large language model; the types of structured framework prompts include query structuring, workflow structuring and response structuring.
[0037] A general safety risk monitoring system based on large model capabilities, including an API interface, an image preprocessing module, an image scene recognition module, a knowledge retrieval module and a risk identification module;
[0038] The API interface receives video data and transmits it to the image preprocessing module;
[0039] The image preprocessing module preprocesses the video data, generates the image to be recognized and transmits it to the image scene recognition module and the knowledge retrieval module;
[0040] The image scene recognition module performs image scene recognition on the image to be recognized, obtains the image scene label description, and transmits it to the knowledge retrieval module;
[0041] The knowledge retrieval module performs knowledge retrieval based on the image scene label description and the image to be identified, recalls the knowledge and transmits it to the risk identification module;
[0042] The risk identification module constructs risk warnings and combines recall knowledge to identify risks and generate safety risk monitoring results.
[0043] Preferably, the image preprocessing module is loaded with video processing software, and the video data is decoded using the decoding library of the video processing software, and the video processing software is configured with an alignment algorithm, format conversion algorithm, image size adjustment algorithm, audio processing algorithm, etc. with a timestamp or other synchronization mechanism.
[0044] Preferably, the image scene recognition module loads the trained image recognition module and the multimodal large model, and stores different scene label description texts according to the business type of security risk monitoring.
[0045] Preferably, the knowledge retrieval module is loaded with the RAGFlow system for knowledge retrieval, and different knowledge documents are preset according to the business type of security risk monitoring.
[0046] Preferably, the risk identification module is provided with a prompt word engineering design unit, which presets risk prompts based on a thinking chain according to the business type of security risk monitoring.
[0047] It can be seen from the above technical solutions that, compared with the prior art, the present invention discloses a general safety risk monitoring method and system based on large model capabilities, which combines advanced image scene recognition technology and knowledge retrieval technology, aiming to achieve real-time or asynchronous polling monitoring and intelligent analysis of various safety risks, and can significantly improve the efficiency and effectiveness of safety risk monitoring. The present invention is applicable to safety risk monitoring needs in multiple fields including but not limited to fire safety, urban safety, industrial production safety, public facility safety, traffic management, etc. It has broad application prospects and significant practical value, and provides strong support for security assurance and risk management in various business scenarios. Specifically, the beneficial effects produced by the technical solution of the present invention include:
[0048] 1. Efficient and accurate recognition of different image scenes: The image scene recognition technology used is combined with advanced deep learning models, which can quickly and accurately identify different scenes in the video and give high-confidence prediction results to achieve real-time monitoring, ensuring that potential risks can be discovered in time at any time period and in any scenario. Compared with traditional methods, it has significantly improved accuracy and processing speed.
[0049] 2. Dynamic threshold adjustment and multimodal supplementation; when the prediction probability is lower than the preset threshold, the system can intelligently combine the image scene label description text generated by the multimodal large model to further enhance the robustness and accuracy of recognition. This dynamic adjustment and supplementation mechanism effectively overcomes the misjudgment or missed judgment problems that may occur in a single model.
[0050] 3. Enhanced knowledge retrieval based on RAG: By introducing RAG technology, efficient correlation retrieval with business knowledge documents is achieved, and key information such as laws and regulations related to image scenes, hidden danger identification rules, etc. can be quickly recalled. By using big data analysis and machine learning algorithms, the accuracy and reliability of safety risk identification can be improved, and the occurrence of misjudgments and missed judgments can be reduced. This link has greatly enriched the system's knowledge base and improved the comprehensiveness and professionalism of risk monitoring.
[0051] 4. Intelligent hidden danger identification; combining the recall knowledge and specific prompt words, the system can intelligently generate hidden danger identification results. This intelligent processing method can not only improve work efficiency, but also provide strong data support for decision makers.
[0052] 5. Structured rewriting and parameterized query: By applying prompt engineering technology and thinking chain method, the structured and parameterized rewriting of query statements is realized, the interaction method with the back-end knowledge base is optimized, and the query accuracy and response speed are improved.
[0053] 6. Friendly compatibility with non-real-time scenarios; It has friendly compatibility and support for both real-time and non-real-time scenarios.
[0054] 7. Support the construction of an intelligent early warning system that can trigger a response mechanism immediately after a security risk is discovered, quickly take effective countermeasures, and reduce the losses caused by the risk. BRIEF DESCRIPTION OF THE DRAWINGS
[0055] In order to more clearly illustrate the embodiments of the present invention or the technical solutions in the prior art, the drawings required for use in the embodiments or the description of the prior art will be briefly introduced below. Obviously, the drawings described below are only embodiments of the present invention. For ordinary technicians in this field, other drawings can be obtained based on the provided drawings without paying creative work.
[0056] Figure 1 A schematic diagram of a general safety risk monitoring method based on large model capabilities provided by the present invention;
[0057] Figure 2 A schematic diagram of an image to be identified in another embodiment provided by the present invention. DETAILED DESCRIPTION
[0058] The following will be combined with the drawings in the embodiments of the present invention to clearly and completely describe the technical solutions in the embodiments of the present invention. Obviously, the described embodiments are only part of the embodiments of the present invention, not all of the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by ordinary technicians in this field without creative work are within the scope of protection of the present invention.
[0059] The embodiment of the present invention discloses a general security risk monitoring method based on large model capabilities, such as Figure 1 As shown, the following steps are included:
[0060] S1: Obtain video data and perform preprocessing to obtain the image to be recognized;
[0061] S2: Use the image recognition model and the multimodal large model to perform image scene recognition on the image to be recognized and obtain the image scene label description;
[0062] S3: Use RAG technology to retrieve knowledge based on image scene label description and the image to be identified, and recall knowledge;
[0063] S4: Use a large language model to generate security risk monitoring results based on the recalled knowledge and constructed risk warnings.
[0064] Furthermore, the method for obtaining video data includes receiving video data from an external system and collecting video data using a surveillance camera; when receiving video data from an external system, the GB28181 standard protocol can be used to access the external system to obtain video data, or the video stream address of the external system can be collected to pull video data according to the video stream address; when collecting video data using a surveillance camera, the surveillance camera is connected to a switch or router, and port forwarding or VPN is configured to obtain video data collected by the surveillance camera through the video stream address.
[0065] Furthermore, the process of preprocessing the video data includes:
[0066] S11: Decode the video data to obtain an image frame set and audio;
[0067] S12: indexing the image frame set in time order, and extracting the image frames of a set time as the images to be identified;
[0068] S13: Process the image to be identified, including format conversion, scaling, and cropping, etc., to meet the analysis requirements and store it in a suitable format. The image frame can be stored in a picture format such as JPEG, PNG, etc.
[0069] Furthermore, an alignment algorithm based on a timestamp or other synchronization mechanism can be used to align the image frames and audio in the image frame set; extract the audio to be recognized corresponding to the image to be recognized in the audio; perform audio processing on the audio to be recognized, including denoising, enhancement, and format conversion, etc., to provide data support for various applications such as speech recognition and security monitoring. Audio data can be stored in audio formats such as WAV and MP3.
[0070] Furthermore, the specific process of performing image scene recognition on the image to be recognized in S2 is:
[0071] S21: Input the image to be recognized and the image scene category label text pre-set based on security risks into the trained image recognition model, and output the highest image scene category prediction probability;
[0072] S22: If the output image scene category prediction probability is greater than the preset probability threshold, the image scene category label corresponding to the image to be identified is determined according to the image scene category prediction probability, and the image to be identified and the corresponding scene category label are respectively input into the trained multimodal large model to obtain the corresponding image scene label description; otherwise, the image to be identified is input into the trained multimodal large model, and the query is input at the same time to generate the image scene label description. The probability threshold can be set to a value in the range of (0,1), and the default value is 0.55. The greater the deviation of the probability threshold from the default value, the lower the sensitivity of recognition, and the smaller the absolute value deviation of the probability threshold from the default value, the higher the sensitivity of recognition.
[0073] Furthermore, the image recognition model may use the EVA-CLIP-8B model, and the specific implementation process of S21 is as follows:
[0074] S211: pre-set image scene category label text based on security risks, and perform text encoding to obtain text features;
[0075] S212: pre-processing the image to be identified and normalizing it;
[0076] S213: Input the normalized image and text features into the EVA-CLIP-8B model respectively, extract high-level visual features, and compare the high-level visual features with the text features corresponding to the preset image scene category label text to determine the image scene category prediction probability of each image scene category label, and output the highest image scene category prediction probability.
[0077] Furthermore, the multimodal large model can use CogVLM2, BLIP3 or GLM4 model; the image to be identified is input into the trained multimodal large model, and the query is input at the same time, and the specific process of generating the image scene label description is as follows:
[0078] S221: Inputting the image to be recognized into the multimodal large model;
[0079] S222: constructing a query input related to the image scene into a multimodal large model;
[0080] S223: The multimodal large model extracts image information from the image to be identified, and generates a label description of the image scene by combining the query with the multimodal processing mechanism. The query can be set as "What is the scene of the image?".
[0081] Furthermore, the RAG technology in S3 is implemented using the RAGFlow system. The specific process includes:
[0082] S31: Input the image scene label description into the RAGFlow system, use matching algorithm and speech analysis algorithm to retrieve knowledge from the preset business knowledge documents, and recall the knowledge of laws and regulations or hidden danger identification rules;
[0083] S32: The image to be recognized is normalized and then input into the trained image encoder to generate an embedding vector, which is then input into the RAGFlow system. The matching algorithm and speech analysis algorithm are used to retrieve knowledge from a preset scene-related knowledge base to recall scene knowledge.
[0084] Furthermore, S31 also includes constructing a prompt word prompt and inputting it into the RAGFlow system, and the RAGFlow system performs knowledge retrieval in combination with the prompt word prompt.
[0085] Furthermore, the specific process of S4 is:
[0086] S41: Construct risk warning;
[0087] S42: Input embedding vectors, legal regulations or hidden danger identification rule knowledge, scenario knowledge and risk warnings into the large language model to obtain safety risk monitoring results.
[0088] Furthermore, the business knowledge document can contain a "Hazard Inspection and Correction Suggestion Record Form", which contains on-site hazards, hazard descriptions, hazard types, reference bases, correction suggestions, correction plans, etc. Improvement suggestions can be recalled through knowledge retrieval, and the large language model can output corresponding improvement suggestions while outputting safety risk monitoring results.
[0089] Furthermore, the risk warning text is set, and the risk warning is constructed based on the thinking chain. The risk warning is converted into a structured framework prompt and then input into the large language model; the types of structured framework prompts include query structuring, workflow structuring and response structuring.
[0090] Furthermore, the image recognition model may use edge detection algorithms, feature extraction algorithms, etc. to realize image scene recognition.
[0091] Furthermore, the image scene label description may adopt short text description, or long text description, existing database types, scene category libraries with existing coding specifications, and the like.
[0092] Furthermore, when performing knowledge retrieval, a knowledge retrieval method based on preset rules and keywords may be adopted.
[0093] Furthermore, the video data preprocessing, image scene recognition, knowledge retrieval, and hidden danger identification are distributed to multiple computing nodes respectively, and cloud computing resources are used to achieve efficient video processing.
[0094] Furthermore, the present invention is applicable to a variety of business scenarios, including but not limited to:
[0095] 1. Video surveillance of the conversation site: for the interrogation process in the conversation room, ensure the compliance and security of the conversation content; it is necessary to supervise and manage the conversation behavior in real time or after the fact to protect the legitimate rights and interests of both parties;
[0096] 2. Safety workshop monitoring: real-time monitoring of the operating status of production equipment, identification of potential mechanical failures and operating errors, analysis of workers' activities in dangerous areas, and prevention of work-related accidents;
[0097] 3. Large-scale building fire monitoring: identify fire sources and smoke, quickly locate the fire location, monitor the unobstructed conditions of evacuation passages and emergency exits, and guide personnel to evacuate;
[0098] 4. Bridge traffic safety monitoring: monitor bridge traffic conditions, identify illegal driving and dangerous behaviors, detect bridge vibration in real time, and warn of potential structural risks;
[0099] 5. Urban traffic monitoring: Analyze road traffic flow, identify traffic accidents and violations, and dispatch rescue forces in a timely manner;
[0100] 6. School and kindergarten security monitoring: real-time monitoring of the campus environment, ensuring the safety of teachers and students, and identifying and preventing unauthorized persons from entering;
[0101] 7. Internal monitoring: monitor the implementation of safety regulations in the office, identify and record violations, and provide evidence support for disciplinary review;
[0102] 8. Other application scenarios.
[0103] In a specific embodiment, assuming that a user uploads an image of a city street, the detailed process of generating a security risk monitoring result is as follows:
[0104] S1: Image upload and preprocessing: A user uploads an image of a city street; the image is preprocessed to ensure that it meets the input requirements of the image encoder;
[0105] S2: Use the pre-trained image encoder to process the pre-processed image and generate an embedding vector of fixed dimension;
[0106] S3: Input the embedding vector into the RAG system to retrieve relevant knowledge, such as traffic rules for urban streets, safety hazard identification and management regulations;
[0107] S4: Combine the embedding vector with the relevant knowledge of recall and input it into the large language model;
[0108] The large language model generates the following answer: "When driving on city streets, you must obey traffic rules and pay attention to the movements of pedestrians and other vehicles. Common safety hazards include areas without sidewalks and insufficient warning signs in construction areas. According to the city street management regulations, vehicles may not be parked randomly on both sides of the street."
[0109] In a specific embodiment, when setting a risk warning, the specific process is:
[0110] S1: Set a simple query, such as "Upload a picture, please determine whether there is any abnormal behavior?";
[0111] S2: Rewrite the simple query into a structured way to obtain risk warnings, for example, "The uploaded picture was taken during the interview in the interview room from an angle facing the interviewee. Please analyze whether the content shown in the picture contains one of the risk types listed below. The risk type consists of the risk type, risk type serial number, risk type name and risk type description. There are 5 types of risks, namely "1. Physical contact between people (physical contact between the interviewee and the investigator occurs, there is a compliance risk, and an alarm is issued)"; "2. Non-compliant items are detected (items that are not allowed to be brought in appear, there is a compliance risk, and an early warning is issued)"; "3. Suspected person falls abnormally (during the interview, the participant shows suspected falling behavior, there is a compliance risk, and an alarm is issued)"; "4. Suspected smoking behavior occurs (during the interview, the participant shows suspected smoking behavior (combined with physical contact), there is a compliance risk, and an alarm is issued)"; " 5. Large and dangerous movements (during the conversation, participants made large and dangerous body movements, there is a compliance risk, and an alarm was issued)"; "6. No obvious compliance risk (there is no obvious compliance risk in the picture of the conversation)". Please analyze and output: 1. Use text to understand and describe the content of the picture; 2. Is there a compliance risk in the conversation, which one, reply with the serial number and risk type name; if none of them, the assessment result is the risk type "6. No obvious compliance risk". 3. Provide the confidence level of which risk type it is, and the confidence level is expressed in percentage. 4. If it is analyzed to be one of the compliance risks, use the content of the risk type to supplement the description of this risk type, do not expand it, and only use the original text of the risk type. Arrange the above content into key-value format: {Risk type serial number: N; Name: N; Confidence level: N%; Risk type description: N.}";
[0112] S3: Rewrite the risk warning in a structured text based on the thought chain and convert it into a structured framework to obtain a structured framework prompt, for example:
[0113] The structured text is: "[Uploaded] [Picture] was [taken] from an angle [towards the interviewee] during the [interview in the interview room]. Please [analyze whether the content shown in the picture contains one of the risk types listed below]. [The risk type consists of the risk type, risk type serial number, risk type name and risk type description]. [There are 5 types of risks, namely "1. Physical contact between people (there is physical contact between the interviewee and the investigator, there is a compliance risk, and an alarm is issued)"; "2. Non-compliant items are detected (items that are not allowed to be brought in appear, there is a compliance risk, and an early warning is issued)"; "3. Suspected person falls abnormally (during the interview, the participant shows suspected falling behavior, there is a compliance risk, and an alarm is issued)"; "4. Suspected smoking behavior occurs (during the interview, the participant shows suspected smoking behavior (combined with physical contact), there is a compliance risk, and an alarm is issued)"; "5. Large and dangerous Action (During the conversation, the participants made large and dangerous body movements, there is a compliance risk, and an alarm was issued)"; "6. No obvious compliance risk (there is no obvious compliance risk in the picture of the conversation)". Please analyze and output: [1. Use text to understand and describe the content of the picture;] [2. Is there a compliance risk in the conversation? Which one is it? Reply with the serial number and risk type name; if none of them is true, the assessment result is the risk type "6. No obvious compliance risk".] [3. Provide the confidence level of which risk type it is. The confidence level is expressed in percentage.] [4. If it is analyzed to be one of the compliance risks, use the content of the risk type to supplement the description of this risk type. Do not expand it. Only use the original text of the risk type.] [Arrange the above content into key-value format: {Risk type serial number: N; Name: N; Confidence level: N%; Risk type description: N.}]";
[0114] The structured framework prompts for query structure are:
[0115]
[0116]
[0117]
[0118]
[0119] The structured framework of workflow structured prompts are:
[0120]
[0121]
[0122]
[0123] The structured framework prompts for the response structure are:
[0124]
[0125]
[0126] In a specific embodiment, the process of applying the method of the present invention to perform security identification is:
[0127] S1: Input content:
[0128] Upload a picture;
[0129] query0: Upload a picture, perform security risk monitoring and analysis, and return the analysis results.
[0130] Query 1: The query structure based on the thought chain is rewritten as follows:
[0131] query1:
[0132] The uploaded picture was taken during the interview in the interview room from an angle facing the interviewee. Please analyze whether the content shown in the picture is one of the risk types listed below. The risk type consists of the risk type, risk type serial number, risk type name and risk type description. There are 5 types of risk, namely "1. Physical contact between people (physical contact between the interviewee and the investigator occurs, there is a compliance risk, and an alarm is issued)"; "2. Non-compliant items are detected (items that are not allowed to be brought in appear, there is a compliance risk, and an early warning is issued)"; "3. Suspected person falls abnormally (during the interview, the participant appears to fall, there is a compliance risk, and an alarm is issued)"; "4. Suspected smoking behavior occurs (during the interview, the participant appears to smoke (combined with physical), there is a compliance risk, and an alarm is issued)"; "5. Large and dangerous movements (during the interview, the participant makes large and dangerous physical movements, there is a compliance risk, and an alarm is issued)"; "6. No obvious compliance risk (in the picture of the conversation, there is no obvious compliance risk)". Please analyze and output: 1. Use text to understand and describe the content of the picture; 2. Whether there is a compliance risk in the conversation, which type, reply with the serial number and risk type name; if none of them, the assessment result is risk type "6. No obvious compliance risk". 3. Provide the confidence level of which risk type it is, and the confidence level is expressed as a percentage. 4. If the analysis shows that it is one of the compliance risks, use the content of the risk type to supplement the description of this risk type, do not expand it, and only use the original text of the risk type. Arrange the above content into key-value format: {Risk type serial number: N; Name: N; Confidence level: N%; Risk type description: N.}
[0133] S2: Pictures:
[0134] S3: Output:
[0135] (1) Understanding and describing image content:
[0136] (2) Compliance risk analysis:
[0137] (3) Evaluation results
[0138] Risk Type No.: 6
[0139] Name: No obvious compliance risk
[0140] Confidence level: 100%
[0141] Risk type description: No obvious compliance risk
[0142] (4) Summary
[0143] According to the analysis of the image content, there are no compliance risks listed above. The assessment result is risk type "6, no obvious compliance risk" with a confidence level of 100%.
[0144] In a specific embodiment, the process of applying the method of the present invention to perform security identification is:
[0145] S1: Input:
[0146] Query input: Describe the image content, indicate the existing safety hazards, and provide the original text of the national technical specifications in the field of production safety and fire safety that should be followed for safety hazard identification as a reference.
[0147] S2: Image: Figure 2 shown.
[0148] S3: Output:
[0149] This image shows a messy workplace with multiple potential safety hazards. Below is a description of the image and an analysis of potential safety hazards:
[0150] (1) Description of image content
[0151] Clutter and tools: There are a lot of clutter and tools piled on the ground and the workbench, making it look very messy.
[0152] Gas cylinders: Several gas cylinders were randomly placed on the ground without any fixing measures.
[0153] Metal Parts: There is a large metal part on the ground, which may be a mechanical component.
[0154] Messy environment: The whole area appears very messy, and items are placed in disorder.
[0155] (2) Potential safety hazards
[0156] Improper storage of gas cylinders:
[0157] Gas cylinders should be stored upright and at least 10 meters away from open flames or high temperature sources.
[0158] Gas cylinders without anti-tipping measures may topple over during an earthquake or collision, causing gas leakage or explosion.
[0159] Items are piled up in a messy manner:
[0160] A cluttered environment can easily cause items to trip over or fall, causing personal injury.
[0161] Debris may block vision or passageways, increasing the risk of accidents.
[0162] Lack of safety signs:
[0163] There were no obvious safety signs or warning signs at the scene, and the staff might not be aware of the potential dangers in a timely manner.
[0164] Ground obstacles:
[0165] Metal parts and other obstacles on the floor can cause people to trip or slip.
[0166] (3) References
[0167] According to China's "Work Safety Law and Technical Specifications in the Field of Fire Safety", the following provisions are related to the above safety hazards:
[0168] "******* Production Safety Law":
[0169] Article 31: Production and operation units shall establish and improve a system for detecting and managing production safety accident hazards, and adopt technical and management measures to promptly discover and eliminate accident hazards.
[0170] Article 39: Workshops, shops and warehouses for the production, operation, storage and use of dangerous goods shall not be located in the same building as employee dormitories and shall maintain a safe distance from employee dormitories.
[0171] Regulations on the Safety Management of Hazardous Chemicals:
[0172] Article 20: Dangerous chemicals shall be stored in special warehouses, special sites or special storage rooms (hereinafter referred to as special warehouses) and shall be managed by dedicated personnel; highly toxic chemicals and other dangerous chemicals whose storage quantities constitute a major hazard source shall be stored separately in special warehouses and a two-person receiving and sending and two-person custody system shall be implemented.
[0173] Fire Protection Law:
[0174] Article 26: Fire-fighting equipment should be installed in crowded places as needed, and regular inspections and maintenance should be carried out to ensure that the fire-fighting equipment is in good condition and effective.
[0175] Article 27: The places where inflammable and explosive dangerous goods are stored shall comply with fire protection technical standards and management regulations, set up obvious fire prevention signs, and be equipped with necessary fire-fighting equipment.
[0176] (4) Recommended measures
[0177] Gas cylinder management:
[0178] Gas cylinders should be stored upright and equipped with anti-tip devices.
[0179] Gas cylinders should be kept away from open flames, high temperature sources and other flammable and explosive items.
[0180] Environmental arrangement:
[0181] Clean your workplace regularly to keep it tidy and tidy.
[0182] Items should be stored in categories and clearly labeled.
[0183] Safety signs:
[0184] Set up safety signs and warning signs at key locations.
[0185] Regularly inspect and maintain safety facilities.
[0186] Ground Management:
[0187] Clear ground obstacles and ensure that the passage is clear.
[0188] Check the floor regularly to make sure it is level to prevent tripping accidents.
[0189] Through the above measures, safety hazards can be effectively reduced and the safety of staff can be guaranteed.
[0190] On the other hand, a general safety risk monitoring system based on large model capabilities includes an API interface, an image preprocessing module, an image scene recognition module, a knowledge retrieval module, and a risk identification module;
[0191] The API interface receives video data and transmits it to the image preprocessing module;
[0192] The image preprocessing module preprocesses the video data, generates the image to be recognized and transmits it to the image scene recognition module and the knowledge retrieval module;
[0193] The image scene recognition module performs image scene recognition on the image to be recognized, obtains the image scene label description, and transmits it to the knowledge retrieval module;
[0194] The knowledge retrieval module performs knowledge retrieval based on the image scene label description and the image to be identified, recalls the knowledge and transmits it to the risk identification module;
[0195] The risk identification module constructs risk warnings and combines recall knowledge to identify risks and generate safety risk monitoring results.
[0196] Furthermore, the image preprocessing module is loaded with video processing software, and uses the decoding library of the video processing software to decode the video data, and uses the alignment algorithm, format conversion algorithm, image size adjustment algorithm, audio processing algorithm, etc. configured with timestamp or other synchronization mechanism by the video processing software.
[0197] Furthermore, the image scene recognition module loads the trained image recognition module and the multimodal large model, and stores different scene label description texts according to the business type of security risk monitoring.
[0198] Furthermore, the knowledge retrieval module is loaded with the RAGFlow system for knowledge retrieval, and different knowledge documents are preset according to the business type of security risk monitoring.
[0199] Furthermore, the risk identification module is provided with a prompt word engineering design unit, which presets risk prompts based on a thinking chain according to the business type of safety risk monitoring.
[0200] In this specification, each embodiment is described in a progressive manner, and each embodiment focuses on the differences from other embodiments. The same or similar parts between the embodiments can be referred to each other. For the device disclosed in the embodiment, since it corresponds to the method disclosed in the embodiment, the description is relatively simple, and the relevant parts can be referred to the method part.
[0201] The above description of the disclosed embodiments enables one skilled in the art to implement or use the present invention. Various modifications to these embodiments will be apparent to one skilled in the art, and the general principles defined herein may be implemented in other embodiments without departing from the spirit or scope of the present invention. Therefore, the present invention will not be limited to the embodiments shown herein, but rather to the widest scope consistent with the principles and novel features disclosed herein.
Claims
1. A general security risk monitoring method based on large model capabilities, characterized in that: The following steps are involved: Step 1: Obtain video data and preprocess it to obtain the image to be recognized; Step 2: Use the image recognition model and the multimodal large model to perform image scene recognition on the image to be recognized, and obtain the image scene label description; Step 3: Use RAG technology to retrieve knowledge based on the image scene label description and the image to be identified, and recall knowledge; Step 4: Use a large language model to generate security risk monitoring results based on the recalled knowledge and constructed risk warnings.
2. A general safety risk monitoring method based on large model capabilities according to claim 1, characterized in that: The process of preprocessing video data includes: Step 11: Decode the video data to obtain image frame sets and audio; Step 12: Index the image frame set in time order, and extract the image frames of the set time as the images to be identified; Step 13: Perform image processing on the image to be recognized, including format conversion, scaling and cropping.
3. A general safety risk monitoring method based on large model capability according to claim 1, characterized in that: The specific process of performing image scene recognition on the image to be recognized in step 2 is: Step 21: Input the image to be identified and the image scene category label text pre-set based on security risks into the trained image recognition model, and output the highest image scene category prediction probability; Step 22: If the output image scene category prediction probability is greater than a preset probability threshold, determine the image scene category label corresponding to the image to be identified according to the image scene category prediction probability, input the image to be identified and the corresponding scene category label into the trained multimodal large model respectively, and obtain the corresponding image scene label description; Otherwise, the image to be identified is input into the trained multimodal large model, and the query is input at the same time to generate an image scene label description.
4. A general safety risk monitoring method based on large model capability according to claim 3, characterized in that: The image recognition model uses the EVA-CLIP-8B model. The specific implementation process of step 21 is as follows: Step 211: pre-set image scene category label text based on security risk, and perform text encoding to obtain text features; Step 212: pre-processing the image to be identified and normalizing it; Step 213: Input the normalized image and text features into the EVA-CLIP-8B model respectively, extract high-level visual features, and compare the high-level visual features with the text features corresponding to the preset image scene category label text to determine the image scene category prediction probability of each image scene category label, and output the highest image scene category prediction probability.
5. A general safety risk monitoring method based on large model capability according to claim 3, characterized in that: The specific process of inputting the image to be identified into the trained multimodal large model and the query at the same time to generate the image scene label description is as follows: Step 221: input the image to be recognized into the multimodal large model; Step 222: constructing a query input related to the image scene into a multimodal large model; Step 223: The multimodal large model extracts image information from the image to be identified, combines the query with the multimodal processing mechanism, and generates an image scene label description.
6. A general safety risk monitoring method based on large model capability according to claim 1, characterized in that: The RAG technology in step 3 is implemented using the RAGFlow system. The specific process includes: Step 31: Input the image scene label description into the RAGFlow system, use matching algorithms and speech analysis algorithms to retrieve knowledge from preset business knowledge documents, and recall legal regulations or hidden danger identification rule knowledge; Step 32: Normalize the image to be recognized and input it into the trained image encoder to generate an embedding vector. Input the embedding vector into the RAGFlow system and use the matching algorithm and speech analysis algorithm to retrieve knowledge from the preset scene-related knowledge base and recall the scene knowledge.
7. A general safety risk monitoring method based on large model capability according to claim 6, characterized in that: Step 31 also includes constructing a prompt word prompt and inputting it into the RAGFlow system. The RAGFlow system performs knowledge retrieval in combination with the prompt word prompt.
8. A general safety risk monitoring method based on large model capability according to claim 6, characterized in that: The specific process of step 4 is: Step 41: Construct risk warning; Step 42: Input the embedding vector, legal regulations or hidden danger identification rule knowledge, scenario knowledge and risk warning into the large language model to obtain the safety risk monitoring results.
9. A general safety risk monitoring method based on large model capability according to claim 8, characterized in that: Set the risk warning text, and rewrite the risk warning text into risk warning based on the thinking chain, convert the risk warning into a structured framework prompt and input it into the large language model; the types of structured framework prompts include query structuring, workflow structuring and response structuring.
10. A general safety risk monitoring system based on large model capabilities, characterized in that: A general safety risk monitoring method based on large model capability applied to any one of claims 1 to 9, comprising an API interface, an image preprocessing module, an image scene recognition module, a knowledge retrieval module and a risk identification module; The API interface receives video data and transmits it to the image preprocessing module; The image preprocessing module preprocesses the video data, generates the image to be recognized and transmits it to the image scene recognition module and the knowledge retrieval module; The image scene recognition module performs image scene recognition on the image to be recognized, obtains the image scene label description, and transmits it to the knowledge retrieval module; The knowledge retrieval module performs knowledge retrieval based on the image scene label description and the image to be identified, recalls the knowledge and transmits it to the risk identification module; The risk identification module constructs risk warnings and combines recall knowledge to identify risks and generate safety risk monitoring results.
Citation Information
Cited By
Safety production intelligent monitoring method and device
CN120339027A
Intelligent data query method and system based on big data
CN120429479A
Intelligent data query method and system based on big data
CN120429479B
Unmanned aerial vehicle navigation reasoning method and system based on visual perception and large language model
CN120558243A
Safety early warning method, system and equipment for galvanization process production line and medium
CN120725453A