General security risk monitoring method and device based on image scene recognition knowledge retrieval enhancement, computer equipment and readable storage medium

Through the enhanced method of image scene recognition and knowledge retrieval, combined with RAG search engine and multimodal large model, the problems of low data processing efficiency and inaccurate scene recognition in traditional security monitoring methods are solved, and efficient and accurate security risk monitoring is achieved.

CN120356149AInactive Publication Date: 2025-07-22DARK MATTER ARTIFICIAL INTELLIGENT (BEIJING) TECHNOLOGY CO LTD
View PDF 0 Cites 3 Cited by

Patent Information

Application Number
CN202510382216.7
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-03-28
Publication Date
2025-07-22
Estimated Expiration
Not applicable · inactive patent

AI Technical Summary

Technical Problem

Traditional security monitoring methods have low data processing efficiency, inaccurate scene identification, and weak knowledge retrieval capabilities, which are difficult to meet the needs of accurate security risk monitoring in modern complex and changing environments.

Method used

Using an enhanced method based on image scene recognition knowledge retrieval, the scene recognition results are determined by obtaining video data, pre-trained image scene recognition model and multimodal large model, knowledge retrieval is carried out in combination with RAG search engine, knowledge retrieval is recalled, and knowledge of safety hazard rules is carried out, and hidden danger identification and safety suggestions are carried out. Finally, query structured rewriting based on thinking chain is carried out to obtain target security risk monitoring results.

Benefits of technology

It realizes efficient and accurate safety risk monitoring, can quickly identify potential hidden dangers and provide clear safety suggestions, improving the accuracy and efficiency of monitoring.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120356149A_ABST
    Figure CN120356149A_ABST
Patent Text Reader

Abstract

The invention discloses a general security risk monitoring method and device based on image scene recognition knowledge retrieval enhancement, computer equipment and a readable storage medium, and the method comprises the steps: firstly obtaining video data and extracting a video image, and then obtaining a scene recognition result through a pre-trained model and a multi-modal large model, the method comprises the following steps of: firstly, searching corresponding potential safety hazard rule knowledge by utilizing an RAG search engine, then carrying out potential hazard identification to obtain a result and a suggestion, and finally, carrying out query structured rewriting on related information based on a thinking chain to obtain a target safety risk monitoring result, thereby realizing efficient and accurate safety risk monitoring.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of security monitoring. Specifically, it relates to a general security risk monitoring method, device, computer device, and readable storage medium enhanced by image scene recognition knowledge retrieval. Background Art

[0002] With the development of society and the progress of technology, the importance of security risk monitoring has become increasingly prominent in various fields. However, traditional security monitoring methods have problems such as low data processing efficiency, inaccurate scene recognition, and weak knowledge retrieval ability, making it difficult to meet the precise monitoring requirements for general security risks in modern complex and changing environments. Therefore, an innovative and comprehensive security risk monitoring method is needed to address these challenges. Summary of the Invention

[0003] The purpose of the present invention is to provide a general security risk monitoring method, device, computer device, and readable storage medium enhanced by image scene recognition knowledge retrieval.

[0004] In a first aspect, an embodiment of the present invention provides a general security risk monitoring method enhanced by image scene recognition knowledge retrieval, including:

[0005] Obtain video data, and extract video images from the preprocessed video data;

[0006] Use a pre-trained image scene recognition model and a multimodal large model to determine the scene recognition result corresponding to the video image;

[0007] Use a RAG search engine to perform knowledge retrieval on the scene recognition result, and recall the safety hazard rule knowledge corresponding to the scene recognition result;

[0008] Perform hazard identification based on the video image and the safety hazard rule knowledge to obtain a hazard identification result and safety suggestions;

[0009] Perform query structured rewriting based on the chain of thought on the video image, scene recognition result, safety hazard rule knowledge, the hazard identification result, and the safety suggestions to obtain a target security risk monitoring result.

[0010] In a second aspect, an embodiment of the present invention provides a general security risk monitoring device enhanced by image scene recognition knowledge retrieval, including:

[0011] An acquisition module is used to acquire video data and extract video images from the preprocessed video data; determine the scene recognition results corresponding to the video images using a pre-trained image scene recognition model and a multimodal large model; perform knowledge retrieval on the scene recognition results using a RAG search engine to recall the safety hazard rule knowledge corresponding to the scene recognition results; perform hazard identification based on the video images and the safety hazard rule knowledge to obtain hazard identification results and safety recommendations;

[0012] The analysis module is used to perform query structured rewriting based on the thinking chain on the video image, scene recognition results, safety hazard rule knowledge, the hazard recognition results and the safety suggestions to obtain the target safety risk monitoring results.

[0013] In a third aspect, an embodiment of the present invention provides a computer device, comprising a processor and a non-volatile memory storing computer instructions, wherein when the computer instructions are executed by the processor, the computer device executes the method described in the first aspect.

[0014] In a fourth aspect, an embodiment of the present invention provides a readable storage medium, wherein the readable storage medium includes a computer program, and when the computer program is executed, the computer device where the readable storage medium is located is controlled to execute the method described in the first aspect.

[0015] Compared with the prior art, the beneficial effects provided by the present invention include: adopting the general safety risk monitoring method, device, computer equipment and readable storage medium based on image scene recognition knowledge retrieval enhancement provided by the embodiments of the present invention, including: first acquiring video data and extracting video images, and then using a pre-trained model and a multimodal large model to obtain scene recognition results, and then using the built RAG search engine to retrieve the corresponding safety hazard rule knowledge, and then based on this, perform hazard identification to obtain results and suggestions, and finally perform query structured rewriting of relevant information based on the thinking chain, so as to obtain the target safety risk monitoring results and achieve efficient and accurate safety risk monitoring. BRIEF DESCRIPTION OF THE DRAWINGS

[0016] In order to more clearly illustrate the technical solutions of the embodiments of the present invention, the accompanying drawings required for use in the embodiments are briefly introduced below. It should be understood that the following drawings only illustrate certain embodiments of the present invention and should not be regarded as limiting the scope. For those of ordinary skill in the art, other related drawings can also be obtained based on these drawings without creative work.

[0017] Figure 1 A schematic diagram of the steps of a general security risk monitoring method based on image scene recognition and knowledge retrieval enhancement provided by an embodiment of the present invention;

[0018] Figure 2 Schematic diagram of video image provided by an embodiment of the present invention;

[0019] Figure 3 Structural schematic block diagram of a general security risk monitoring device with enhanced knowledge retrieval based on image scene recognition provided by an embodiment of the present invention;

[0020] Figure 4 Structural schematic block diagram of a computer device provided by an embodiment of the present invention. Detailed implementation manners

[0021] To make the objectives, technical solutions, and advantages of the embodiments of the present invention clearer, the technical solutions in the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings in the embodiments of the present invention. Obviously, the described embodiments are some, but not all, of the embodiments of the present invention. Components of the embodiments of the present invention usually described and illustrated in the drawings here can be arranged and designed in various different configurations.

[0022] The following will describe in detail the specific implementation manners of the present invention with reference to the accompanying drawings.

[0023] To solve the technical problems in the foregoing background art, Figure 1 Flow schematic diagram of a general security risk monitoring method with enhanced knowledge retrieval based on image scene recognition provided by an embodiment of the present disclosure. The following will introduce in detail the general security risk monitoring method with enhanced knowledge retrieval based on image scene recognition.

[0024] Step S201: Obtain video data, and extract video images from the preprocessed video data;

[0025] Step S202: Use a pre-trained image scene recognition model and a multimodal large model to determine the scene recognition result corresponding to the video image;

[0026] Step S203: Use a RAG search engine to perform knowledge retrieval on the scene recognition result, and recall the safety hazard rule knowledge corresponding to the scene recognition result;

[0027] Step S204: Perform hazard identification based on the video image and the safety hazard rule knowledge to obtain a hazard identification result and safety suggestions;

[0028] Step S205: Perform query structured rewriting based on the chain of thought on the video image, scene recognition result, safety hazard rule knowledge, the hazard identification result, and the safety suggestions to obtain a target security risk monitoring result.

[0029] In an embodiment of the present invention, by way of example, assume that our server is deployed in a monitoring center of a large factory. A plurality of monitoring cameras are installed in the factory to monitor the conditions of key areas such as production workshops, warehouses, import and export areas, etc. in real time.

[0030] The server obtains the video data captured by these cameras through a pre-configured network connection. For example, one of the cameras is located beside the main production line in the production workshop, and its video stream is transmitted to the server through the local area network inside the factory in a specific protocol and format. After receiving the video data, the server immediately starts preprocessing. First, it imports and decodes the video data, converting it from a compressed format to the original image frames and audio samples. In this process, an efficient decoding library such as FFmpeg is used to ensure fast and accurate decoding. Next, video image extraction is performed. The server indexes the frames in the video stream in chronological order and extracts the image frames within a specific time period. These image frames may be several frames per second, depending on the monitoring requirements and system settings. For example, if we are interested in the situation of the workshop production line between 10:00 am and 10:10 am, the server will accurately extract the image frames during this period. At the same time, the server also performs some post-processing operations on the extracted video images, such as adjusting the size and resolution of the images to meet the input requirements of the subsequent image scene recognition model. The processed video images are stored in a specific storage area of the server for subsequent analysis. After completing the extraction and preprocessing of the video images, the server starts to perform scene recognition using a pre-trained image scene recognition model. Assume that the used image scene recognition model is based on deep learning and has been trained on a large amount of industrial scene image data. The server inputs the extracted video images into this model. The model first preprocesses the input images, such as resizing the images and normalizing them, in order to better extract image features. Then, the model extracts the high-level features of the images through structures such as convolutional neural networks and compares these features with the preset scene category labels. For example, for an image of a production workshop, the model may calculate the prediction probabilities of various scene categories, such as "normal production", "equipment failure", "personnel violation of operation", etc. Finally, the model outputs the category with the highest prediction probability, such as "normal production", and gives the corresponding probability value, assume it is 0.85. If this prediction probability is lower than a preset threshold (assume the threshold is 0.7), the server will input this image into a multimodal large model. The multimodal large model may integrate the visual information of the image, text description (if any), and other relevant multimodal information to generate a more accurate scene description text. For example, for this image with a low prediction probability, the description text generated by the multimodal large model may be "In the production workshop, the operating status of some equipment is slightly abnormal, but the overall production is still in progress". After obtaining the image scene recognition result, the server uses the RAG search engine for knowledge retrieval. Assume that the scene recognition result is "equipment failure", and the server inputs this result into the RAG search engine. The RAG search engine first performs semantic analysis and keyword extraction on the input scene description. For example, it extracts key information such as "equipment" and "failure".Then, the search engine searches in a pre-built business knowledge document library. This knowledge document library contains a large number of laws, regulations, equipment operation manuals, troubleshooting guides, etc. related to factory production safety. For example, the search engine may recall the following relevant safety hazard rule knowledge: 1. The procedures and requirements for emergency handling of equipment failures in the "Factory Equipment Maintenance and Management Measures". 2. The fault diagnosis manual for specific equipment, which details the possible types of faults and corresponding troubleshooting methods. 3. Case studies of safety accidents that may be caused by equipment failures and preventive measures. The search engine uses advanced matching algorithms and semantic analysis techniques to ensure that the recalled content is highly relevant and accurate to the input scenario description. After obtaining the safety hazard rule knowledge, the server combines it with the video images for hazard identification. First, the server further analyzes the video images, such as detecting the specific state of the equipment and the operation actions of personnel through image recognition technology. Then, these image analysis results are compared and comprehensively judged with the recalled safety hazard rule knowledge. Suppose there are obvious signs of damage to a certain component of the equipment in the video image, and the recalled knowledge mentions that such component damage may lead to production stoppages and even safety accidents. Based on this, the server obtains a hazard identification result of "Equipment component damage, there are production interruption and safety risks". At the same time, the server generates corresponding safety suggestions according to relevant safety rules and experience, such as "Immediately stop the equipment operation, arrange professional maintenance personnel for inspection, set warning signs during the maintenance period to prevent personnel from approaching". The server begins to perform a query structured rewrite based on the chain of thought for the various information obtained and generated previously.

[0031] First, it is clear that the action is to structure all relevant information, and the document content format is a comprehensive format including images, text, and data. The scenario type is determined as "equipment failure scenario in the factory production workshop". The tags include "equipment component damage", "risk of production interruption", etc. The task is to clearly and systematically organize and present the entire safety risk monitoring process and results. The task summary is "structuring and analyzing information related to equipment failures in the factory production workshop". The task rules clarify the structural and content requirements for each part. For example, the scenario description, potential hazard types, identification results, safety suggestions, etc. should all have clear classifications and presentations. The work process includes: 1. Conducting picture semantic understanding of video images to describe the status of the equipment and the surrounding environment in the images. 2. Performing multi-classification judgment to determine the specific type and severity of potential hazards. 3. Calculating the confidence level to indicate the reliability of the potential hazard judgment. 4. Explaining the potential hazard type by referring to relevant safety hazard rule knowledge. The response structure includes: 1. Response content header, briefly stating that it is the safety risk monitoring result regarding equipment failures in the factory production workshop. 2. Picture semantic understanding part, detailing the key information in the video images. 3. Multi-classification part, clearly indicating that the potential hazard type is the risk of production interruption caused by equipment component damage. 4. Confidence level part, giving a specific percentage, such as 80%, to represent the confidence level of this judgment. 5. Explanation part, referring to safety hazard rule knowledge to explain the possible consequences and countermeasures of such potential hazards. 6. Key-value part, organized as {"potential hazard type serial number": "1", "name": "equipment component damage", "confidence level": "80%", "potential hazard type description": "A key component of the equipment shows obvious damage, which may lead to production stoppages and safety accidents"}. Finally, the generated target safety risk monitoring result by the server is presented in a clear and structured format, facilitating relevant personnel to quickly understand and take corresponding measures. For example, the presented result may be as follows:

[0032] {

[0033] "Video image description": "In the production workshop, the [specific component] of the equipment is significantly damaged, and there are [relevant personnel] observing around.",

[0034] "Scenario recognition result": "Equipment failure",

[0035] "Safety hazard rule knowledge": "According to the 'Factory Equipment Maintenance and Management Measures', such component damage requires immediate shutdown for maintenance.",

[0036] "Hazard identification result": "Equipment component damage, there is a risk of production interruption and safety hazards",

[0037] "Safety advice": "Immediately stop the operation of the equipment, arrange for professional maintenance personnel to conduct inspections, and set warning signs during the maintenance period to prevent people from approaching.",

[0038] "Confidence level": "80%

[0039] }

[0040] In this way, relevant personnel can clearly understand the entire safety risk monitoring situation, thus quickly making decisions and taking actions to reduce potential safety risks and production losses.

[0041] In a possible implementation manner, the scene recognition result corresponding to the video image determined by using the pre-trained image scene recognition model and the multimodal large model can be implemented through the following examples.

[0042] Input the video image into the pre-trained CLIP image scene recognition model to obtain the scene type confidence levels of each preset scene corresponding to the video image;

[0043] Extract the target scene type confidence level with the highest execution degree among the multiple scene type execution degrees;

[0044] If the target scene type confidence level is higher than the preset confidence threshold, use the scene label of the target preset scene corresponding to the target scene type execution degree as the scene recognition result;

[0045] If the target scene type confidence level is lower than the preset confidence threshold, input the video image into the multimodal large model to obtain the descriptive text of the video image as the scene recognition result.

[0046] In an embodiment of the present invention, illustratively, in a monitoring system of a large shopping mall, the server is responsible for processing video images from various areas. The server receives a video image from the entrance of the shopping mall. First, this video image is input into a pre-trained CLIP image scene recognition model. This model has been trained on a large amount of shopping mall scene image data and can recognize a variety of preset scenes, such as "normal flow of people", "crowded people", "abnormal gathering", etc. The model analyzes the input video image and calculates the scene type confidence of each preset scene. For example, for this section of the video image at the entrance, the confidence of "normal flow of people" is calculated to be 0.6, the confidence of "crowded people" is 0.3, and the confidence of "abnormal gathering" is 0.1. The server extracts the highest target scene type confidence among multiple scene type confidences, that is, 0.6 for "normal flow of people". Next, the server compares this target scene type confidence with a preset confidence threshold. Assume that the preset confidence threshold is 0.5. Since 0.6 is higher than 0.5, the server determines that "normal flow of people" is the scene recognition result of the current video image. At another moment, the server received a video image from the promotion area of the mall. After inputting the CLIP image scene recognition model, the confidence of "normal shopping" was calculated to be 0.4, the confidence of "buying spree" was 0.3, and the confidence of "disorder" was 0.2. The highest confidence of the target scene type is 0.4 for "normal shopping". Since 0.4 is lower than the preset confidence threshold of 0.5, the server inputs this video image into the multimodal large model. The multimodal large model will conduct a deeper analysis and understanding of the video image and generate descriptive text, such as "There are many customers in the promotion area of the mall, but the shopping order is basically normal, and there are slightly more people in front of some shelves." This descriptive text is used as the scene recognition result of this video image. In the actual operation process, the server continuously receives and processes video images from all corners of the mall, and accurately determines the scene recognition result corresponding to each video image according to the above process. This helps mall managers to understand the conditions of different areas in the mall in a timely manner so that they can take corresponding management measures, such as adding staff to guide people in crowded or chaotic areas, or reasonably adjusting the display of goods in areas with normal traffic. Through this continuous and efficient scene recognition and processing, the server provides strong support for the safe and orderly operation of the mall. Such a scenario example fully demonstrates how the server uses the pre-trained image scene recognition model and multimodal large model to accurately determine the scene recognition results corresponding to the video image to meet the needs of practical applications.

[0047] In the embodiment of the present invention, the use of the RAG search engine to perform knowledge retrieval on the scene recognition result and recall the safety hazard rule knowledge corresponding to the scene recognition result can be implemented through the following examples.

[0048] Collect security risk-related data from a preset plurality of data sources based on the data layer; the preset plurality of data sources include logs, reports, news, and social media;

[0049] Based on the data layer, perform data cleaning and noise removal on the security risk-related data, and after performing standardization processing, store the obtained standardized security-related data;

[0050] Based on the retrieval layer, call a preset search engine to construct an inverted index and configure an optimized retrieval strategy; the optimized retrieval strategy includes dynamic index layer sealing, context caching mechanism, retrieval model driven by reinforcement learning, fine-tuning of retrieval components for datasets in specific domains, user feedback loop optimization, multi-step information snippet retrieval, and relevance ranking of documents and query information; the retrieval layer is configured with parallel computing mode and distributed retrieval mode;

[0051] Based on the production layer, determine an initial generation model, and perform pre-training including specific domain fine-tuning based on the standardized security-related data to obtain a trained generation model; the trained generation model is used to make decisions on the initial retrieval information output by the retrieval layer;

[0052] Based on the data layer, the retrieval layer, and the generation layer, construct the RAG search engine, input the scenario recognition result into the RAG search engine, and recall the security hazard rule knowledge corresponding to the scenario recognition result through a matching algorithm and semantic analysis in combination with prompt words.

[0053] In the embodiment of the present invention, the hazard identification is performed according to the video image and the security hazard rule knowledge to obtain a hazard identification result and a security suggestion, which can be implemented through the following examples.

[0054] Perform preprocessing on the video image to obtain a preprocessed video image;

[0055] Extract features from the preprocessed video image to obtain a video image vector;

[0056] Combine the video image vector and the security hazard rule knowledge and input them into a preset large language model to obtain the hazard identification result and the security suggestion.

[0057] In an embodiment of the present invention, exemplarily, in a monitoring system of a chemical plant, the server undertakes an important task of safety risk monitoring. The server first obtains a video image from the reactor area of the chemical plant. For better subsequent processing, the server preprocesses this video image. In the preprocessing stage, the server performs denoising operations on the video image to remove noise generated by factors such as electromagnetic interference in the chemical environment; it also adjusts the brightness and contrast of the image to make key equipment and operation details more clearly visible. After preprocessing, a clearer and more analyzable preprocessed video image is obtained. Next, the server extracts features from the preprocessed video image. By applying advanced image analysis algorithms, the server extracts a series of key features from the video image, such as the temperature display value of the reactor, the connection status of the pipelines, the action postures of the operators, etc., and converts these features into a video image vector. At the same time, the server obtains safety hazard rule knowledge related to the chemical plant reactor, which includes the normal operation parameter range of the reactor, the possible types of faults and their corresponding danger levels, as well as relevant safety operation specifications, etc. Then, the server combines the extracted video image vector and this safety hazard rule knowledge. For example, it compares the temperature value of the reactor in the video image with the normal temperature range specified in the safety hazard rule knowledge, and matches the actions of the operators with the requirements in the safety operation specifications. Subsequently, the server inputs the combined information into a preset large language model. This large language model has been trained with a large amount of chemical safety data and has powerful analysis and reasoning capabilities. After receiving the input information, the large language model conducts in-depth analysis and judgment. If the temperature of the reactor in the video image exceeds the upper limit specified in the safety hazard rule knowledge, and at the same time the actions of the operators do not conform to the safety specifications, the large language model will obtain a hazard identification result of "the reactor temperature is too high, the operators' operations are not standardized, and there is an explosion risk". At the same time, the large language model will also give corresponding safety suggestions, such as "immediately stop the operation of the reactor, evacuate the nearby personnel, check whether the cooling system is faulty, and conduct safety training for the operators". In another scenario, the server obtains a video image from the storage area of the chemical plant. After similar preprocessing and feature extraction steps, a video image vector regarding the pressure of the storage tank, the storage method of the materials, etc. is obtained. Combining relevant safety hazard rule knowledge, such as the pressure safety threshold of the storage tank, the requirements for classified storage of materials, etc., and inputting it into the large language model. The hazard identification result that the large language model may obtain is "the pressure of the storage tank is approaching the critical value, the materials are stored in a chaotic manner, and there is a leakage risk", and it gives safety suggestions of "adjust the pressure of the storage tank, reorganize the storage of the materials, and strengthen the operation of the ventilation equipment". In this way, the server can accurately identify the safety hazards in various areas of the chemical plant and provide targeted safety suggestions, playing an important role in ensuring the safe production of the chemical plant.

[0058] In an embodiment of the present invention, the query structure rewriting of the video image, scene recognition result, safety hazard rule knowledge, hazard recognition result, and safety suggestion based on the chain of thought to obtain the target safety risk monitoring result can be implemented through the following examples.

[0059] Perform query structure rewriting based on the chain of thought on the video image, scene recognition result, safety hazard rule knowledge, hazard recognition result, and safety suggestion to obtain structured data including a first-level structure name, a second-level structure name, a Chinese structure name, and content; the structured data includes query structured data, workflow structured data, and response structured data;

[0060] Use the structured data as the target safety risk monitoring result and display the target safety risk monitoring result on a preset interface.

[0061] In an embodiment of the present invention, exemplarily, in an urban traffic monitoring system, the server collects a series of video images and obtains relevant results after a series of processes. The server acquires a video image of an intersection and, through an image scene recognition model and a multimodal large model, determines that the scene recognition result is "traffic congestion scene during peak hours". At the same time, relevant safety hazard rule knowledge is retrieved using a RAG search engine, such as "specifications for traffic signal settings during peak hours" and "emergency evacuation measures when vehicles are congested". Based on the video image and this knowledge, hazard identification is carried out, and the hazard identification result is "some vehicles changing lanes illegally exacerbates congestion", and safety suggestions are given, such as "strengthen on-site command by traffic police and adjust the signal duration". Next, the server performs a query structured rewrite of this information based on the chain of thought. First is the description of the video image, "Vehicles are dense at the intersection, moving slowly, and there is a queuing phenomenon in some sections". The scene recognition result is "traffic congestion scene during peak hours". The safety hazard rule knowledge includes detailed articles on signal settings and evacuation measures. The hazard identification result clearly points out that "some vehicles changing lanes illegally exacerbates congestion". The safety suggestion is "strengthen on-site command by traffic police and adjust the signal duration". After the structured rewrite, structured data including the first-level structure name, second-level structure name, Chinese name of the structure, and content is obtained. In the query structured data, the first-level structure name is "Action", the second-level structure name is "Action", the Chinese name of the structure is "Obtain and process traffic monitoring information", and the content is "Collect video images of the intersection, conduct scene analysis and hazard identification". In the workflow structured data, the first-level structure name is "workflow", the second-level structure name is "Workflow", the Chinese name of the structure is "Traffic congestion monitoring and response process", and the content includes "Analyze video images, determine scene types, retrieve relevant rules, identify hazards, and put forward suggestions". In the response structured data, the first-level structure name is "response", the second-level structure name is "Response content", the Chinese name of the structure is "Traffic congestion monitoring results and response suggestions", and the content covers detailed information such as "video image description, scene recognition result, hazard identification result, safety suggestions, etc.". Finally, the server takes this structured data as the target safety risk monitoring result and displays it on a preset interface. On this interface, the staff of the traffic management department can clearly see various detailed information. For example, they can intuitively understand the specific congestion situation at the intersection, clearly know that this is a traffic congestion scene during peak hours, be aware that some vehicles changing lanes illegally is one of the reasons for the exacerbation of congestion, and obtain suggestions such as strengthening on-site command by traffic police and adjusting the signal duration.The server performs a chain-of-thought-based query structured rewrite on this information, generates corresponding structured data, and displays it on a preset interface, providing clear and organized safety risk monitoring results for the traffic management department to assist it in making timely and effective decisions and deployments.

[0062] In the embodiment of the present invention, for obtaining video data and extracting video images from the preprocessed video data, the following examples can be used for implementation.

[0063] Extract the video data from a preset video library;

[0064] Perform decoding, audio-visual separation and synchronization, and post-processing operations on the video data in sequence to obtain the video images.

[0065] In the embodiment of the present invention, by way of example, in a monitoring system of a large logistics park, the server is responsible for obtaining and processing video data to ensure the safety and operation efficiency of the park.

[0066] The server first extracts video data from a preset video library. This video library stores surveillance videos of various areas in the park, including warehouses, loading and unloading areas, entrances and exits, etc. For example, the server extracts a video data from inside the warehouse, which records the process of goods handling and storage.

[0067] After extracting the video data, the server starts a series of processing operations on it. First is decoding, which converts the video data from its original compressed format into processable image frames and audio samples. In this process, the server uses efficient decoding libraries and algorithms to ensure the accuracy and speed of decoding.

[0068] Next, perform audio-visual separation and synchronization. The server precisely separates the audio data and image data in the video and ensures their temporal synchronization through mechanisms such as timestamps. For the video inside the warehouse, the audio may include the sounds of handling equipment, conversations of staff, etc., and the images show the placement of goods, handling actions, etc.

[0069] After that is the post-processing operation. The server performs various processes on the separated images, such as format conversion, converting the images from the original format to a format more suitable for subsequent analysis, such as JPEG or PNG. It may also perform scaling and cropping to focus on key areas, such as the stacking height of goods, the distance between handling equipment and shelves, etc. At the same time, noise reduction and enhancement are performed on the audio data to hear important sound information more clearly.

[0070] After these processing operations, the server successfully extracts clear and accurate video images from the preprocessed video data. These video images can accurately reflect the real-time situation inside the warehouse, providing high-quality input data for subsequent scene recognition, hazard analysis, and other tasks.

[0071] In another scenario, the server extracts video data of the park entrance and exit from the video library. Similarly, through decoding, audio-visual separation and synchronization, and post-processing, the server obtains clear video images of vehicles and personnel entering and leaving at the entrance and exit. These images can be used to analyze traffic flow, identify abnormal behaviors, etc., which helps improve the safety management level and operation efficiency of the park.

[0072] In summary, by extracting video data from the preset video library and performing a series of professional processing operations, the server successfully extracts high-quality video images, laying a solid foundation for the entire security risk monitoring and operation management work.

[0073] In the embodiment of the present invention, the method further includes:

[0074] Constructing a standardized API interface;

[0075] Based on the standardized API interface, obtaining the video data from the target monitoring environment and storing it in the preset video library.

[0076] In the embodiment of the present invention, by way of example, in an intelligent factory, the server undertakes the important task of obtaining and managing video data.

[0077] The server first constructs a standardized API interface. This interface is designed very rigorously and normatively, and can communicate efficiently and stably with various different types of monitoring devices.

[0078] Taking the production workshop of the factory as an example, multiple high-definition cameras are installed in the workshop to monitor the operation of the production line, the operation specifications of workers, and the status of equipment in real time. These cameras are connected to the standardized API interface of the server through the network.

[0079] When the production line starts to operate, the cameras start to collect video data. Based on the constructed standardized API interface, the server can obtain video data from these cameras in real time.

[0080] After receiving the video data, the server will perform a series of processing and verification on it to ensure the integrity and accuracy of the data. If it is found that the data is missing or abnormal, a request will be sent to the camera through the API interface to re-obtain the relevant data.

[0081] The acquired video data will be stored in a preset video library. This video library has high-capacity and high-performance storage devices, capable of quickly writing and reading data.

[0082] For example, for the production line video data within a certain period, the server will classify and store it according to information such as time sequence and camera position for subsequent quick retrieval and call.

[0083] During the storage process, the server will also add relevant metadata to each video data, such as acquisition time, camera number, monitoring area, etc., further facilitating data management and query.

[0084] In addition to the production workshop, cameras in other target monitoring environments in the factory, such as warehouses and office areas, also transmit video data to the server in the same way and store it in the preset video library.

[0085] In this way, the server can obtain rich and comprehensive video data from various target monitoring environments in the factory, providing sufficient materials for subsequent analysis and processing, helping to timely discover potential safety hazards, quality problems and management loopholes in the production process, thereby improving the production efficiency and safety of the factory.

[0086] In another scenario, such as a surveillance system in a large shopping mall. There are numerous cameras distributed in the mall, covering key areas such as each floor, entrances and exits, and cashier desks.

[0087] The server continuously obtains video data from these cameras through a standardized API interface. Whether it is the flow of customers, the business status of stores or emergencies in public areas, they can all be recorded and transmitted in a timely manner.

[0088] These video data are also stored in an orderly manner in the preset video library, providing strong support for the operation management, security prevention and marketing strategy formulation of the mall.

[0089] In short, the server effectively obtains video data from the target monitoring environment through constructing a standardized API interface and stores it properly, laying a foundation for achieving comprehensive and accurate monitoring and analysis.

[0090] In order to more clearly describe the solution provided by the embodiments of the present invention, a relatively complete implementation method is provided below.

[0091] The technical solution of the present invention is described as follows:

[0092] Obtain video data;

[0093] Video preprocessing (decompose the acquired video into video images and audio data);

[0094] Image scene recognition process;

[0095] Compare the image with the scene labels; calculate the predicted probability for each category. The model will output the category with the highest predicted probability and the corresponding probability value. The predicted image scene label text is obtained in this step.

[0096] If the predicted probability is lower than the preset threshold, it means that the image does not match the set category labels sufficiently. For those that do not match, add the query, and the image scene label description text generated by the multimodal large model;

[0097] Knowledge retrieval enhancement; use the RAG technology; recall relevant laws, regulations, or hazard identification rules related to the image scene based on business knowledge documents;

[0098] Identify whether there are hazards; generate hazard identification and opinions based on the recalled knowledge and prompt words;

[0099] Structured rewriting of the query based on the chain of thought; apply prompt engineering technology and the chain of thought for the structuring and parameterization of the query;

[0100] Safety risk monitoring results and analysis;

[0101] Business system presentation: Integrate with the application system based on the above capabilities and present it on the system, so as to achieve strong general safety risk monitoring capabilities for multiple scenarios.

[0102] The entire technical solution consists of the following parts, as detailed below:

[0103] (1) Obtain video data

[0104] The video data acquisition method can either access the video data of the existing system or install cameras by oneself to obtain data.

[0105] There are two ways to access the video data of the existing system: One is to automatically access the system through the GB28181 standard; this method is exclusive. If the existing system has already used this method for docking, this method is not available. The other is to manually obtain the video stream address of the camera and configure it into the system; multiple systems can use this method to access simultaneously.

[0106] The main steps to install cameras by oneself and obtain video data:

[0107] Install surveillance cameras:

[0108] Select a suitable surveillance camera, such as a high-definition camera, a night vision camera, or a wide-angle camera, according to the specific environment and requirements of the conversation room.

[0109] Determine the installation location of the camera. It is usually recommended to install it at the entrance, exit of the conversation room, and areas with frequent problems.

[0110] Install the camera bracket and fix the camera at a suitable position such as the ceiling or wall, ensuring that the angle covers the area to be monitored.

[0111] Connect the power cable and network cable of the camera to ensure that the camera can be started and transmit data normally.

[0112] Configure the network connection:

[0113] Configure a static or dynamic IP address for the camera so that it can communicate with other devices in the network.

[0114] Connect the camera to the switch or router to ensure that the camera can access the Internet or local area network normally.

[0115] Configure port forwarding or VPN to access the surveillance camera from a remote location.

[0116] Conduct a network test on the camera to ensure smooth data transmission without obstacles.

[0117] Set up the storage device:

[0118] Select the method of local storage or cloud storage to save video data. Local storage includes digital video recorders (DVRs) or network-attached storage (NAS), while cloud storage relies on third-party services.

[0119] If local storage is selected, connect the storage device to the camera through the local area network and allocate appropriate storage space.

[0120] If cloud storage is selected, register for the relevant service and create a bucket, and at the same time configure data encryption and backup policies to ensure data security.

[0121] Obtain the video address:

[0122] Obtain the video stream address in the NVR, DVR or cloud storage service.

[0123] (2) Video preprocessing (decompose the obtained video into video images and audio data);

[0124] The process of decomposing the obtained video into video images and audio data is described in detail below:

[0125] Data import and decoding:

[0126] Import the recorded video data into the video processing software or development environment.

[0127] Use a decoding library (such as FFmpeg, GStreamer, etc.) to decode the video data. The decoding process will convert the compressed video encoding into the original image frames and audio samples.

[0128] Video image extraction:

[0129] After decoding, the video data is decomposed into a series of image frames. Each frame is a separate picture, usually represented in RGB, BGR, or other pixel formats.

[0130] Index the frames in the video stream in chronological order to facilitate the extraction of image frames at specific time points or within a specific time range.

[0131] Audio data extraction:

[0132] Meanwhile, the audio data in the video stream is extracted, which is usually saved in Pulse Code Modulation (PCM) or other audio encoding formats.

[0133] Extract the audio data and convert it into a processable format, such as waveform data or other digital audio formats, for further analysis and processing.

[0134] Synchronization separation:

[0135] Ensuring the temporal synchronization of image frames and audio data is very important, especially in scenarios where audio-visual analysis is required.

[0136] Use timestamps or other synchronization mechanisms to ensure that the image frames are temporally aligned with the corresponding audio samples.

[0137] Post-processing:

[0138] Perform post-processing on the extracted video images, such as format conversion, scaling, cropping, etc., to meet the requirements of specific applications.

[0139] Perform noise reduction, enhancement, format conversion, etc. on the extracted audio data to improve the audio quality or adapt to specific audio analysis tools.

[0140] Data storage:

[0141] Store the extracted video images and audio data in suitable formats respectively. For example, image frames can be stored in picture formats such as JPEG, PNG, etc., and audio data can be stored in audio formats such as WAV, MP3, etc.

[0142] Automated processing:

[0143] By writing scripts or using automated tools, automate the video data decomposition process to improve efficiency and reduce human errors.

[0144] Through the above steps, the video obtained from the surveillance camera or video file can be effectively decomposed into independent video images and audio data, thereby providing data support for various applications such as computer vision analysis, speech recognition, and security monitoring.

[0145] (3) Image scene recognition: Perform image scene recognition operations on the video images to identify the scene tags or descriptive texts of the video, and the threshold parameters for scene tag comparison can be set.

[0146] Through the operation in step (2), a single video file is shared into an audio data file and a video file. In the whole process, the extracted video images and audio data can be stored in suitable formats respectively. The video can be converted into pictures frame by frame; for example, image frames can be stored in picture formats such as JPEG and PNG, and audio data can be stored in audio formats such as WAV and MP3. In the current node of the present invention, the focus is on the scene recognition of video images, so the audio data is not further processed. If necessary, corresponding processing and recognition operations can also be carried out.

[0147] The scene recognition operation of video images mainly includes 2 steps:

[0148] Image comparison with scene tags; Receive an image to be recognized and the predicted tag text; The model compares the image with the preset scene category tags and calculates the predicted probability for each category. The model will output a category with the highest predicted probability and the corresponding probability value. The predicted image scene tag text obtained in this link is what is obtained.

[0149] If the predicted probability is lower than the preset threshold, it means that the image does not fully match the set category tags. At this time, the system inputs the image into a supported multimodal large model (such as CogVLM2, BLIP3 or GLM4) and asks the question "What is the scene of the image?" to generate the text information of the scene described by the image. The image scene tag description text generated through the multimodal large model obtained in this link is what is obtained.

[0150] Among them, the predicted image scene tag text will be described in detail in step (3.1); the image scene tag description text generated through the multimodal large model will be described in detail in step (3.2).

[0151] Explanation of setting the threshold parameter for image scene recognition: We set the threshold to a value in the range of (0,1); the default value is set to 0.55; the greater the deviation of the threshold from the default value, the lower the recognition sensitivity; the smaller the absolute deviation of the threshold from the default value, the higher the recognition sensitivity.

[0152] (3.1) Predicted image scene tag text

[0153] Here we use the CLIP model to implement it, and the process is as follows:

[0154] Input example:

[0155] Image input: An image file in JPEG format, with a size of 224x224 pixels, clearly showing natural scenery (such as mountains and lakes). Text input (for comparison and verification): Descriptive text: "Mountains and lakes in natural scenery". Processing process: Image preprocessing: The image is adjusted to the input size required by the model. The pixel values of the image are normalized. Feature extraction and matching: The EVA-CLIP-8B model extracts the high-level visual features of the image. The extracted features are compared and analyzed with the preset scene category labels. Output example: Prediction result: Category label: Natural scenery, Probability value: 0.95 (indicating that the probability of the image belonging to the "Natural scenery" category is 95%). Detailed output format: See the output example:

[0156] {

[0157] "input_image":"path / to / image.jpg",

[0158] "predicted_class":"Natural scenery",

[0159] "confidence_score":0.95

[0160] }

[0161] (3.2) Image scene label description text generated by the multimodal large model

[0162] If the prediction probability is lower than the preset threshold, it means that the image does not fully match the set category label. At this time, the system inputs the image into the supported multimodal large model (such as CogVLM2, BLIP3, or GLM4), and asks the question "What is the scene of the image?" to generate the text information of the scene described by the image.

[0163] The overall process is as follows:

[0164] Step 1: Upload a picture. The user uploads a picture to the system through the interface or API. The system receives and saves the picture file.

[0165] Step 2: Picture preprocessing. The system performs necessary preprocessing on the uploaded picture, such as resizing, cropping, denoising, etc., to ensure that the picture meets the input requirements of the MLLM model. The picture is converted in format, usually converted to a format acceptable to the model, such as JPEG or PNG.

[0166] Step 3: Invoke the MLLM model. (Such as CogVLM2, BLIP3, or GLM4) The system invokes the deployed MLLM (Multimodal Language Model) for inference. The preprocessed image is passed as input data to the MLLM model.

[0167] Step 4: Ask questions and generate scene descriptions. When invoking the MLLM model, append the question: "What is the scene of the image?"

[0168] The MLLM model combines the input image information and the question, and through its internal multimodal processing mechanism, generates the corresponding scene description text.

[0169] Step 5: Output the scene description. The MLLM model returns the generated scene description text to the system. The system receives and displays the scene description for output.

[0170] (4) Knowledge retrieval enhancement; Use the RAG technology; Recall relevant laws, regulations, or hazard identification rules related to the image scene based on business knowledge documents; In the field of safety risk monitoring, combining image scene recognition and knowledge retrieval technology can significantly improve the ability to identify and manage potential risks. Through the category labels predicted by the CLIP model or the scene description generated by the multimodal large model, combined with the RAGFlow system for knowledge retrieval, relevant laws, regulations, or hazard identification rules related to the scene can be recalled.

[0171] The overall process and steps are as follows:

[0172] Step 1: Input the scene description text generated by the model or the image scene label text predicted by comparing the image with the scene label; Input the category label "factory workshop" predicted by the CLIP model into the multimodal large model. The multimodal large model generates a detailed scene description text (for example: "A busy factory workshop with machinery and equipment in operation and workers bustling around the production line.").

[0173] Step 2: Knowledge retrieval enhancement

[0174] Input to the RAGFlow system: Input the class label "factory workshop" predicted by the CLIP model or the scene description text generated by the multimodal large model into the RAGFlow system. Efficiently retrieve relevant knowledge: The RAGFlow system utilizes its powerful knowledge retrieval ability to quickly recall relevant laws, regulations, or hazard identification rules for this scenario. For example, the RAGFlow system may retrieve the following relevant regulations and rules: "Factory Safety Management Measures"; "Operating Procedures for Mechanical Equipment"; "Factory Workshop Emergency Plan"; "Safety Protection Measures for Workers Working in Hazardous Areas"; Ensure relevance and accuracy: The RAGFlow system ensures that the recalled content is highly relevant and accurate to the input scene description through advanced matching algorithms and semantic analysis techniques.

[0175] Step 3: Output the result

[0176] Display relevant regulations and rules: The system presents the retrieved regulations and rules to the user in a clear and understandable manner. The user can view and download the relevant documents for further understanding and application. Through the prompt, the current image can be used more clearly, and within the scope of the business domain knowledge of the scene label, knowledge content such as the required laws, regulations, or hazard identification rules can be recalled.

[0177] (5) Identify whether there are hazards; Based on the recalled knowledge and prompt, generate hazard identification and opinions. This node is the output of (4), and is illustrated below through process description and examples:

[0178] Process description of using image embedding combined with RAG to recall knowledge:

[0179] Step 1: Image processing and embedding generation. Image upload: The user uploads an image containing a specific scene (e.g., a city street). Image preprocessing: The image is scaled to the input size required by the model and pixel value normalization is performed. Generate image embedding: Use a pre-trained image encoder (such as ResNet, VGG, etc.) to process the preprocessed image and generate an embedding vector of a fixed dimension.

[0180] Step 2: RAG recall relevant knowledge

[0181] Input RAG system: Input the generated image embedding vector into the RAG system. Efficiently retrieve relevant knowledge: The RAG system utilizes its powerful knowledge retrieval ability to quickly recall the knowledge base information related to this scenario. For example, the RAG system may retrieve the following relevant knowledge: traffic rules for urban streets; identification of safety hazards on urban streets; management regulations for urban streets; Ensure relevance and accuracy: The RAG system ensures that the recalled content is highly relevant and accurate to the input image embedding through advanced matching algorithms and semantic analysis techniques.

[0182] Step 3: Generate the response of the large model

[0183] Combine the image embedding with the recalled knowledge: Combine the image embedding vector with the recalled relevant knowledge and input it into a large language model (such as GPT-3, BERT, etc.). Generate an answer: The large language model utilizes its powerful generation ability to combine the image information and relevant knowledge to generate a detailed and accurate answer. The following is an example: Suppose the user uploads an image of an urban street, and the system generates the corresponding image embedding vector. The detailed process is as follows: Image upload and preprocessing: The user uploads an image of an urban street. The system preprocesses the image to ensure it meets the input requirements of the image encoder. Generate image embedding: After processing the image using a pre-trained image encoder, a fixed-dimensional embedding vector is generated. RAG recalls relevant knowledge: Input the image embedding vector into the RAG system. The RAG system retrieves relevant knowledge, such as traffic rules for urban streets, identification of safety hazards, and management regulations. Generate an answer: Combine the image embedding vector with the recalled relevant knowledge and input it into a large language model. The large language model generates the following answer: "When driving on urban streets, traffic rules need to be followed, and attention should be paid to the movements of pedestrians and other vehicles. Common safety hazards include areas without sidewalks and insufficient warning signs in construction areas. According to the urban street management regulations, vehicles are not allowed to park randomly on both sides of the street."

[0184] (6) Structured rewriting of the query based on the chain of thought; Apply prompting engineering techniques and the chain of thought for the structuring and parameterization of the query

[0185] (6.1) Structured rewriting of the query based on the chain of thought

[0186] The template for the structured rewriting of the query based on the chain of thought is as follows:

[0187] Simple query:

[0188] Upload an image and determine whether there is any abnormal behavior;

[0189] Request for structured rewriting:

[0190] The uploaded picture was taken from the angle facing the interviewee during a conversation in a certain interview room. Please analyze whether the content shown in the picture contains any of the following listed risk types. The risk types consist of a risk type number, a risk type name, and a risk type description. There are 5 risk types, namely: "1. Physical contact between people (physical contact between the interviewee and the investigator, there is a compliance risk, issue an alarm)"; "2. Detection of non-compliant items (items that are not allowed to be brought in are present, there is a compliance risk, issue a warning)"; "3. Suspected abnormal person falling (during the conversation, a participant shows a suspected falling behavior, there is a compliance risk, issue an alarm)"; "4. Appearance of a suspected smoking behavior (during the conversation, a participant shows a suspected smoking behavior (combined with body movements), there is a compliance risk, issue an alarm)"; "5. Large and dangerous movements (during the conversation, a participant shows large and dangerous body movements, there is a compliance risk, issue an alarm)"; "6. No obvious compliance risk (in the picture of the conversation, there is no obvious compliance risk)". Please analyze and output: 1. Use the text to understand and describe the content of the picture; 2. Whether there is a conversation compliance risk, which one it is, reply with the number and the risk type name; if none of them, the evaluation result is the risk type "6. No obvious compliance risk". 3. Provide the confidence level of which risk type it is, and the confidence level is expressed as a percentage. 4. If it is analyzed that it is one of the compliance risks, use the content of the risk type to supplement and explain this risk type, without expanding, only use the original text of the risk type. Organize the above content in key-value format: {Risk type number: N; Name: N; Confidence level: N%; Risk type description: N.}

[0191] After structured rewriting based on the chain of thought query:

[0192]

Upload

Picture

interrogating in a certain interview room

facing the interviewee

analyze whether the content shown in the picture contains any of the following risk types

The risk types consist of a risk type number, a risk type name, and a risk type description.

There are 5 risk types, namely: "1. Physical contact between people (physical contact between the interviewee and the investigator, there is a compliance risk, issue an alarm)"; "2. Detection of non-compliant items (items that are not allowed to be brought in are found, there is a compliance risk, issue a warning)"; "3. Suspicious person falling abnormally (during the interview, a participant shows a suspicious falling behavior, there is a compliance risk, issue an alarm)"; "4. Suspicious smoking behavior (during the interview, a participant shows a suspicious smoking behavior (combined with body movements), there is a compliance risk, issue an alarm)"; "5. Large and dangerous movements (during the interview, a participant shows large and dangerous body movements, there is a compliance risk, issue an alarm)"; "6. No obvious compliance risk (in the picture of the interview, there is no obvious compliance risk)".

Organize the above content in key-value format: {Risk type number: N; Name: N; Confidence level: N%; Risk type description: N.}

[0193] Table 1

[0194]

[0195]

[0196]

[0197]

[0198]

[0199] Table 2

[0200]

[0201]

[0202]

[0203] Table 3

[0204]

[0205]

[0206] Prompt engineering design:

[0207] <role>

[0208] You are a security risk assessment expert who assesses security risks based on images. You will arrange for the team to conduct the following rounds of security risk assessment work based on the uploaded images, determine whether there are security risks, identify risk factors, clarify the national standards and specification content clauses referred to in the risk assessment results, and cite the mandatory national standards and the titles of the standards. Furthermore, it is possible to predict accident and disaster chains, detect security risks in a timely manner, ensure the personal safety of citizens, and reduce the property losses of citizens and the public.

[0209] < / role>

[0210] <info>

[0211] - Author: liucaiyong

[0212] - Version: 20241115_v0.2

[0213] - Model: XXXX

[0214] - Usage: Conduct security risk assessment on uploaded images

[0215] < / info>

[0216] <scene_type>Safety risk assessment (here safety risk is a broad concept, including production safety risk, fire safety risk, bridge safety risk, urban safety risk, public safety risk, border inspection service compliance risk, a certain compliance safety risk, etc.)

[0217] < / scene_type>

[0218] <workflow>

[0219] You will conduct a security risk assessment based on the uploaded image according to the following requirements:

[0220] First, conduct image content understanding and description of the image;

[0221] Second, conduct intention understanding on the content understanding and description of the image, compare the keywords of the intention with the risk types in scene_typ to confirm scene_type, or bring scene_type into the query;

[0222] Third, introduce rag+scene_type data; extract slots if any;

[0223] Fourth, use the return of the second step and the return of the third step, use an external API, and use the function-call process; add the obtained response to query2;

[0224] Fifth, multiple agents; multiple rag processes;

[0225] response;

[0226] Image content understanding and description;

[0227] Areas of security risk assessment; {Production safety risk, Fire safety risk, Bridge safety risk, Urban safety risk, Public safety risk, Border control service compliance risk, A certain compliance safety risk,...}

[0228] Whether there is a security risk; Yes or No; Which type of security risk; Serial number + Risk type + Risk type description;

[0229] The confidence level of this judgment;

[0230] Technical specifications and clauses with risks; Articles of technical specifications;

[0231] The number and title of the technical specification;

[0232] Sixth, organize all the above content according to the template and conduct structured output.

[0233] < / workflow>

[0234] <printout>

[0235] Picture content understanding and description: {{{query+tag+scene_type+Picture semantic understanding}}}

[0236] Field of safety risk assessment: {{{scene_type}}}

[0237] Safety risk assessment result: KB recall {

[0238] {{{query+tag+scene_type+Picture semantic understanding}}}+Picture+{{{scene_type}}}

[0239] }

[0240] Confidence level of safety risk assessment result; [0,1]

[0241] Technical specification clauses or articles on which the safety risk assessment is based: {{{Part of KB recall}}}

[0242] Compulsory national standard number and title name of the technical specification: {{{Part of KB recall}}}

[0243] < / printout>

[0244] (7) Safety risk monitoring results and analysis;

[0245] This node is the execution result of (5) and (6), or the analysis of information. It will not be elaborated here.

[0246] 8) Business system presentation: Integrate with the application system based on the above capabilities and present on the system, so as to achieve a strong general safety risk monitoring ability for multiple scenarios.

[0247] This is the presentation on the business system side of (7). The technical solution of this patent further realizes the deep integration with the application system, ensuring the generality and practicality of the safety risk monitoring ability in various business scenarios.

[0248] The provided system integration method is as follows:

[0249] API Interface Docking: Provide standardized API interfaces, enabling this technical solution to be easily integrated into various business systems, whether it is an enterprise's internal security management platform or a third-party monitoring and early warning system. Through the API interface, real-time data transmission and instant feedback of processing results are achieved. Modular Design: Carefully divide each functional module of security risk monitoring. Each module is responsible for a specific task, such as video data processing, image scene recognition, hidden danger recognition, etc. This modular design not only improves the flexibility and maintainability of the system but also facilitates function expansion or upgrade according to actual needs. Scene Adaptation: Customized Configuration: Provide flexible configuration options for different business scenarios and security requirements. Users can adjust the thresholds for scene recognition, the scope of knowledge retrieval, and the rules for hidden danger recognition according to their own situations to ensure that the system can meet the monitoring needs in specific scenarios to the greatest extent. Cross-industry Application: This technical solution is not limited to a specific industry but has cross-industry generality. Whether it is urban security monitoring, industrial production environment monitoring, or public facility security management and other fields, this system can provide effective security risk monitoring services.

[0250] Specifically, the following implementation manners are also provided.

[0251] Input content:

[0252] Upload a picture;

[0253] query0: Upload a picture, please conduct security risk monitoring and analysis and return the analysis results.

[0254] query1: After the query is structured and rewritten based on the chain of thought, it is as follows:

[0255] query1:

[0256] The uploaded picture was taken from the angle facing the interviewee during an interrogation in a certain interview room. Please analyze whether the content shown in the picture contains any of the following listed risk types. The risk types consist of a risk type number, a risk type name, and a risk type description. There are 5 risk types, namely: "1. Physical contact between people (physical contact between the interviewee and the investigator, there is a compliance risk, issue an alarm)"; "2. Detection of non-compliant items (items not allowed to be brought in are present, there is a compliance risk, issue a warning)"; "3. Suspected abnormal person falling (during the interview, a participant shows a suspected falling behavior, there is a compliance risk, issue an alarm)"; "4. Appearance of suspected smoking behavior (during the interview, a participant shows a suspected smoking behavior (combined with body movements), there is a compliance risk, issue an alarm)"; "5. Large and dangerous movements (during the interview, a participant shows large and dangerous body movements, there is a compliance risk, issue an alarm)"; "6. No obvious compliance risk (in the picture of the interview, there is no obvious compliance risk)". Please analyze and output: 1. Use text to understand and describe the picture content; 2. Whether there is an interview compliance risk, and if so, which one, reply with the number and the risk type name; if none of them, the evaluation result is risk type "6. No obvious compliance risk". 3. Provide the confidence level of which risk type it is, and the confidence level is expressed as a percentage. 4. If it is analyzed that it is one of the compliance risks, use the content of the risk type to supplement and explain this risk type, without expanding, only using the original text of the risk type. Organize the above content in key-value format: {Risk type number: N; Name: N; Confidence level: N%; Risk type description: N.}

[0257] Picture: A certain picture.

[0258] Output:

[0259] Understanding and description of picture content

[0260] This picture shows a scene in an interview room, taken at 10:33:06 on Tuesday, November 14, 2023. There are three people sitting around a table, seemingly in the process of having a conversation or an interrogation. There are some items on the table, including a cup and some documents or equipment. The room is simply designed, with neutral colors for the walls and furniture.

[0261] Analysis of compliance risks

[0262] Based on the picture content, the following is the analysis of compliance risks:

[0263] Physical contact between people: There is no obvious physical contact between people in the picture, so there is no such risk.

[0264] Non - compliant item detected: No non - compliant items are seen in the picture, so there is no such risk.

[0265] Suspected abnormal person fall: The person's sitting posture is normal in the picture, and there are no obvious falls or abnormal behaviors, so there is no such risk.

[0266] Suspected smoking behavior occurred: No one is seen smoking in the picture, so there is no such risk.

[0267] Large - scale and dangerous actions: The person's actions in the picture are relatively stable, and there are no large - scale and dangerous actions, so there is no such risk.

[0268] Evaluation result

[0269] Risk type serial number: 6

[0270] Name: No obvious compliance risks

[0271] Confidence level: 100%

[0272] Risk type description: In the picture of the conversation, there are no obvious compliance risks

[0273] Summary:

[0274] Based on the analysis of the picture content, none of the above - listed compliance risks exist. The evaluation result is risk type "6. No obvious compliance risks", and the confidence level is 100%.

[0275] For another example, the following implementation manners are also provided:

[0276] query input:

[0277] Text: Describe the picture content, prompt the existing safety hazards, and provide the original text of the technical specifications in the fields of national work safety and fire safety followed for safety hazard discrimination as a reference basis.

[0278] Picture: A certain picture, please refer to it in combination Figure 2

[0279] Output:

[0280] This picture shows a messy workplace that may have various safety hazards. The following is a description of the picture content and an analysis of potential safety hazards:

[0281] Picture content description:

[0282] Miscellaneous items and tools: There are many miscellaneous items and tools piled up on the ground and workbench, looking very messy. Gas cylinders: Multiple gas cylinders are placed randomly on the ground without any fixing measures. Metal parts: There is a large metal part on the ground, which may be a mechanical component. Cluttered environment: The entire area looks very cluttered, and the items are placed disorderly. Potential safety hazards: Improper storage of gas cylinders: Gas cylinders should be stored upright and at least 10 meters away from open flames or heat sources. No anti-tipping measures are taken for the gas cylinders, which may tip over during an earthquake or collision, resulting in gas leakage or explosion. Messy stacking of items: A messy environment can easily cause items to trip or fall, causing personal injury.

[0283] Miscellaneous items may block the line of sight or passageways, increasing the risk of accidents. Lack of safety signs: No obvious safety signs or warning signs are seen at the site, and staff may not be able to timely understand potential hazards. Ground obstacles: The metal parts and other obstacles on the ground may cause people to trip or slip.

[0284] Reference basis:

[0285] According to the "Technical Specifications in the Field of Work Safety Law and Fire Safety" of China, the following articles are related to the above safety hazards: "Work Safety Law of the People's Republic of China": Article 31: Production and business operation entities shall establish and improve a system for investigating and managing potential work safety accidents, and adopt technical and management measures to timely discover and eliminate potential accidents. Article 39: Workshops, stores, warehouses for the production, operation, storage, and use of dangerous goods shall not be in the same building as employees' dormitories, and shall maintain a safe distance from employees' dormitories. "Regulations on the Safety Management of Hazardous Chemicals": Article 20: Hazardous chemicals shall be stored in special warehouses, special sites, or special storage rooms (collectively referred to as special warehouses) and be managed by special personnel; highly toxic chemicals and other hazardous chemicals whose storage quantity constitutes a major hazard source shall be stored separately in special warehouses and shall implement the system of double receipt and double custody by two persons. "Fire Protection Law": Article 26: Crowded places shall be equipped with fire-fighting equipment as needed, and regularly inspected and maintained to ensure that the fire-fighting equipment is in good condition and effective. Article 27: Places storing inflammable and explosive dangerous goods shall meet fire protection technical standards and management regulations, set obvious fire prevention signs, and be equipped with necessary fire-fighting equipment.

[0286] Suggested measures:

[0287] Gas cylinder management: Gas cylinders should be stored upright and anti-tipping devices should be used. Gas cylinders should be kept away from open flames, high-temperature sources, and other flammable and explosive items. Environmental tidying: Regularly clean the workplace to keep it clean, tidy, and orderly. Items should be stored in categories and clearly labeled. Safety signs: Set up safety signs and warning signs at key positions. Regularly inspect and maintain safety facilities. Floor management: Remove obstacles on the floor to ensure unobstructed passageways. Regularly check whether the floor is flat to prevent tripping accidents. Through the above measures, potential safety hazards can be effectively reduced and the safety of staff can be ensured.

[0288] In addition, in the video preprocessing stage, different technical solutions are used for audio-video separation operations. When converting video to images, different frame extraction strategies and encoding / decoding tools are used. In image encoding and text encoding, different encoding method solutions are used. In MLLM, multi-modal large models from different manufacturers with different parameter scales are used. Scene recognition based on traditional image processing techniques. Instead of using deep learning models, traditional image processing techniques (such as edge detection, feature extraction, etc.) are used for scene recognition. For the categories of scene recognition, a sentence description, a long text description, an existing database type, or a scene category library with an existing coding specification is used to replace the short text description label. Use rule-based knowledge retrieval. Instead of using the RAG technology, knowledge retrieval is performed based on preset rules and keywords. Hidden danger identification based on an expert system. Instead of relying on machine learning models, an expert system containing professional knowledge and rules is constructed for hidden danger identification. Cloud computing-based distributed processing. The processing and analysis tasks of video data are distributed to multiple computing nodes for processing, and cloud computing resources are utilized to achieve efficient video processing. Use prompt templates with different design ideas to implement the entire process. In terms of business data presentation, a more intuitive technology is used for presentation. For example, instead of directly presenting the results on the business system, the monitoring results are superimposed on the actual scene using AR technology for display to replace the traditional information-based business data presentation.

[0289] Please refer to Figure 3 , Figure 3 A general safety risk monitoring device 110 based on enhanced image scene recognition knowledge retrieval provided by an embodiment of the present invention, comprising:

[0290] An acquisition module 110, which acquires video data and extracts video images from the preprocessed video data; determines the scene recognition result corresponding to the video image by using a pre-trained image scene recognition model and a multi-modal large model; performs knowledge retrieval on the scene recognition result by using a RAG search engine to recall the safety hazard rule knowledge corresponding to the scene recognition result; performs hidden danger identification based on the video image and the safety hazard rule knowledge to obtain a hidden danger identification result and safety suggestions;

[0291] The analysis module 1102 performs a chain-of-thought-based query structured rewrite on the video image, scene recognition result, safety hazard rule knowledge, the hazard recognition result, and the safety suggestion to obtain the target safety risk monitoring result.

[0292] It should be noted that the implementation principle of the general safety risk monitoring device 110 enhanced by image scene recognition knowledge retrieval can refer to the implementation principle of the general safety risk monitoring method enhanced by image scene recognition knowledge retrieval, which will not be elaborated here. It should be understood that the division of each module of the above device is only a logical function division. In actual implementation, it can be fully or partially integrated into a physical entity, or physically separated. And these modules can all be implemented in the form of software called by processing elements; they can also all be implemented in the form of hardware; or some modules can be implemented in the form of software called by processing elements, and some modules can be implemented in the form of hardware. For example, the general safety risk monitoring device 110 enhanced by image scene recognition knowledge retrieval can be a separately established processing element, or can be integrated in a certain chip of the above device. In addition, it can also be stored in the memory of the above device in the form of program code, and called and executed by a certain processing element of the above device to perform the functions of the general safety risk monitoring device 110 enhanced by image scene recognition knowledge retrieval. The implementation of other modules is similar. In addition, these modules can be fully or partially integrated together, or can be independently implemented. The processing element described here can be an integrated circuit with signal processing capabilities. In the implementation process, each step of the above method or each of the above modules can be completed by the integrated logic circuit in the processor element or the instruction in the form of software.

[0293] For example, the above modules can be one or more integrated circuits configured to implement the above method, such as: one or more application specific integrated circuits (ASICs), or, one or more digital signal processors (DSPs), or, one or more field programmable gate arrays (FPGAs), etc. Again, when a certain module above is implemented in the form of a program code scheduled by a processing element, the processing element can be a general-purpose processor, such as a central processing unit (CPU) or other processors that can call program code. Again, these modules can be integrated together and implemented in the form of a system-on-a-chip (SOC).

[0294] The embodiment of the present invention provides a computer device 100, which includes a processor and a non-volatile memory storing computer instructions. When the computer instructions are executed by the processor, the computer device 100 executes the aforementioned general security risk monitoring device 110 based on image scene recognition knowledge retrieval enhancement. Figure 4 As shown, Figure 4 The computer device 100 provided in the embodiment of the present invention is a structural block diagram. The computer device 100 includes a general security risk monitoring device 110 based on image scene recognition knowledge retrieval enhancement, a memory 111, a processor 112 and a communication unit 113.

[0295] In order to realize data transmission or interaction, the memory 111, the processor 112 and the communication unit 113 are electrically connected to each other directly or indirectly. For example, the electrical connection between these elements can be realized through one or more communication buses or signal lines. The general security risk monitoring device 110 based on image scene recognition knowledge retrieval enhancement includes at least one software function module that can be stored in the memory 111 in the form of software or firmware or solidified in the operating system (OS) of the computer device 100. The processor 112 is used to execute the general security risk monitoring device 110 based on image scene recognition knowledge retrieval enhancement stored in the memory 111, such as the software function modules and computer programs included in the general security risk monitoring device 110 based on image scene recognition knowledge retrieval enhancement.

[0296] An embodiment of the present invention provides a readable storage medium, which includes a computer program. When the computer program is running, it controls the computer device where the readable storage medium is located to execute the aforementioned general security risk monitoring device 110 based on image scene recognition and knowledge retrieval enhancement.

[0297] For illustrative purposes, the foregoing description is made with reference to specific embodiments. However, the above illustrative discussion is not intended to be exhaustive or to limit the present disclosure to the precise form disclosed. Numerous modifications and variations are possible in accordance with the above teachings. These embodiments are selected and described in order to best illustrate the principles of the present disclosure and its practical application, so that those skilled in the art can best utilize the present disclosure and utilize various embodiments with different modifications to suit the intended specific application.

Claims

1. A general security risk monitoring method based on enhanced knowledge retrieval for image scene recognition, characterized in that, Including: Obtain video data and extract video images from the preprocessed video data; Use a pre-trained image scene recognition model and a multimodal large model to determine the scene recognition result corresponding to the video image; Use a RAG search engine to perform knowledge retrieval on the scene recognition result and recall the safety hazard rule knowledge corresponding to the scene recognition result; Perform hazard identification based on the video image and the safety hazard rule knowledge to obtain a hazard identification result and safety suggestions; Perform a query structured rewrite based on the chain of thought on the video image, scene recognition result, safety hazard rule knowledge, the hazard identification result, and the safety suggestions to obtain a target safety risk monitoring result.

2. The method according to claim 1, wherein The step of using a pre-trained image scene recognition model and a multimodal large model to determine the scene recognition result corresponding to the video image includes: Input the video image into a pre-trained CLIP image scene recognition model to obtain the scene type confidence levels of each preset scene corresponding to the video image; Extract the target scene type confidence level with the highest execution degree among multiple scene type execution degrees; If the target scene type confidence level is higher than a preset confidence threshold, use the scene label of the target preset scene corresponding to the target scene type execution degree as the scene recognition result; If the target scene type confidence level is lower than the preset confidence threshold, input the video image into the multimodal large model to obtain the descriptive text of the video image as the scene recognition result.

3. The method according to claim 1, characterized in that, The step of using a RAG search engine to perform knowledge retrieval on the scene recognition result and recall the safety hazard rule knowledge corresponding to the scene recognition result includes: Collect safety risk-related data from multiple preset data sources based on the data layer; the multiple preset data sources include logs, reports, news, and social media; Based on the data layer, perform data cleaning and noise removal on the safety risk-related data, and after standardization processing, store the obtained standardized safety-related data; Based on the retrieval layer, call a preset search engine to build an inverted index and configure an optimized retrieval strategy; the optimized retrieval strategy includes dynamic index layer sealing, context caching mechanism, reinforcement learning-driven retrieval model, fine-tuning of retrieval components for specific domain datasets, user feedback loop optimization, multi-step information segment retrieval, and relevance ranking of documents and query information; the retrieval layer is configured with parallel computing mode and distributed retrieval mode; Based on the production layer, determine an initial generation model and perform pre-training including specific domain fine-tuning based on the standardized safety-related data to obtain a trained generation model; the trained generation model is used to make decisions on the initial retrieval information output by the retrieval layer; Build the RAG search engine based on the data layer, the retrieval layer, and the generation layer, input the scene recognition result into the RAG search engine, and recall the safety hazard rule knowledge corresponding to the scene recognition result through a matching algorithm and semantic analysis in combination with prompt words.

4. The method according to claim 1, wherein Performing hazard identification based on the video image and the safety hazard rule knowledge to obtain a hazard identification result and safety suggestions, including: Preprocessing the video image to obtain a preprocessed video image; Performing feature extraction on the preprocessed video image to obtain a video image vector; Combining the video image vector and the safety hazard rule knowledge and inputting them into a preset large language model to obtain the hazard identification result and the safety suggestions.

5. The method according to claim 1, characterized in that, Performing a chain-of-thought based query structured rewrite on the video image, scene recognition result, safety hazard rule knowledge, the hazard identification result, and the safety suggestions to obtain a target safety risk monitoring result, including: Performing a chain-of-thought based query structured rewrite on the video image, scene recognition result, safety hazard rule knowledge, the hazard identification result, and the safety suggestions to obtain structured data including a first-level structure name, a second-level structure name, a Chinese structure name, and content; the structured data includes query structured data, workflow structured data, and response structured data; Using the structured data as the target safety risk monitoring result and displaying the target safety risk monitoring result on a preset interface.

6. The method according to claim 1, characterized in that Obtaining video data and extracting a video image from the preprocessed video data, including: Extracting the video data from a preset video library; Performing decoding, audio-visual separation and synchronization, and post-processing operations on the video data in sequence to obtain the video image.

7. The method according to claim 6, wherein The method further includes: Constructing a standardized API interface; Based on the standardized API interface, obtaining the video data from a target monitoring environment and storing it in the preset video library.

8. A general security risk monitoring device based on enhanced knowledge retrieval of image scene recognition, characterized in that, Including: An acquisition module, configured to obtain video data and extract a video image from the preprocessed video data; Determining a scene recognition result corresponding to the video image by using a pre-trained image scene recognition model and a multimodal large model; Performing knowledge retrieval on the scene recognition result by using a RAG search engine to recall safety hazard rule knowledge corresponding to the scene recognition result; Performing hazard identification based on the video image and the safety hazard rule knowledge to obtain a hazard identification result and safety suggestions; An analysis module, configured to perform a chain-of-thought based query structured rewrite on the video image, scene recognition result, safety hazard rule knowledge, the hazard identification result, and the safety suggestions to obtain a target safety risk monitoring result.

9. A computer device, characterized in that, The computer device includes a processor and a non-volatile memory storing computer instructions. When the computer instructions are executed by the processor, the computer device executes the method according to any one of claims 1-7.

10. A readable storage medium, characterized in that, The readable storage medium includes a computer program. When the computer program runs, it controls the computer device where the readable storage medium is located to execute the method according to any one of claims 1-7.

Citation Information

Cited By

  • Industrial hidden danger troubleshooting decision-making method based on multi-agent cooperation

    CN120725625A

  • An industrial hidden danger investigation and decision-making method based on multi-agent cooperation

    CN120725625B

  • Equipment inspection method and device

    CN121788959A