Unsafe behavior monitoring system
By designing an unsafe behavior monitoring system, using the synergy of multiple modules to achieve real-time automatic detection and early warning, and integrating the spatio-temporal information of video data, the existing safety monitoring methods are solved, and the monitoring effect and accuracy are significantly improved.
Patent Information
- Application Number
- CN202510133133.4
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-02-06
- Publication Date
- 2025-06-10
AI Technical Summary
The existing security monitoring methods are inefficient, with a large number of monitoring blind spots, making it difficult to effectively mine and utilize the spatio-temporal information in video data, resulting in poor monitoring results.
An unsafe behavior monitoring system was designed, including real-time monitoring module, early warning module, personnel safety file management module, site safety management module, dialogue query module and safety education and training module. Through the synergy of these modules, it is possible to automatically detect unsafe behaviors in real time and issue early warnings, integrate the time and space information of video data, and improve the accuracy of safety monitoring.
It significantly improves monitoring efficiency and monitoring effect, can effectively explore and utilize the value of video data, reduce the incidence of safety accidents, and improve the overall effect of multi-scene safety monitoring.
Smart Images

Figure CN120126070A_ABST
Abstract
Description
Technical Field
[0001] This application relates to the technical field of safety monitoring, and particularly to an unsafe behavior monitoring system. Background Art
[0002] With the rapid development of modernization, the demand for safety supervision in multiple scenarios such as the supervision of the production environment in construction sites and industrial parks and the safety supervision in public places is increasing day by day. Managers in these scenarios are faced with the huge challenge of ensuring the safety of construction workers and the environment while improving production and work efficiency. However, the current safety management methods have many deficiencies and are difficult to meet the complex and changing safety supervision requirements.
[0003] In the prior art, management methods based on surveillance videos, such as closed-circuit television monitoring, manual patrols, and regular inspections, although play a monitoring role to a certain extent, their inherent defects cannot be ignored. Such methods are often inefficient, have a large number of monitoring blind spots, and are prone to missing unsafe behaviors. The monitoring effect is greatly reduced, and the ability to analyze and summarize cross-surveillance video content is limited. It is difficult to effectively mine and utilize the spatio-temporal information in video data, thus limiting its application effect in multi-scenario safety monitoring. Summary of the Invention
[0004] The main purpose of this application is to overcome the shortcomings and deficiencies of the prior art, and provide an unsafe behavior monitoring system. Through the coordinated action of a real-time monitoring module, an early warning module, a personnel safety file management module, a site safety management module, a conversational query module, and an safety education and training module, it can automatically detect unsafe behaviors in real time and issue early warnings, improving the monitoring efficiency and monitoring effect; at the same time, through the conversational query module to integrate the spatio-temporal information in video data, it can effectively mine and utilize the value of video data, complete the analysis of video content based on different fine-grained levels, thereby improving the safety monitoring accuracy, and significantly enhancing the overall effect of multi-scenario safety monitoring.
[0005] To achieve the above object, this application adopts the following technical solutions:
[0006] In a first aspect, this application provides an unsafe behavior monitoring system, including: a real-time monitoring module, an early warning module, a personnel safety file management module, a site safety management module, a conversational query module, and an safety education and training module;
[0007] The real-time monitoring module is used to obtain the video images of the construction site scene in real time;
[0008] The early warning module is used to detect the unsafe behaviors in the video images of the construction site scene and give early warnings for the unsafe behaviors;
[0009] The personnel safety file management module is used to identify the identity information of construction workers and record their working status and unsafe behaviors;
[0010] The site safety management module is used to calculate the coverage area and blind area of cameras according to the floor plan information, camera positions and fields of view of the construction site, and mark the coverage area and the blind area on the floor plan;
[0011] The conversational query module is used to analyze the content of the user's query request, retrieve relevant video content according to the query request content, and output a retrieval result that meets the query requirements;
[0012] The safety education and training module is used to generate education and training content according to the detected unsafe behaviors.
[0013] As a preferred technical solution, the warning module uses the YOLOv5 model to detect unsafe behaviors in the video footage of the construction site scene, including:
[0014] Extract the deep features of each frame of the video footage through the backbone network in the YOLOv5 model;
[0015] Input the extracted deep features into the feature pyramid network in the YOLOv5 model, and the feature pyramid network fuses deep features of different scales to obtain fused features;
[0016] Input the fused features into the detection head in the YOLOv5 model, and the detection head outputs the prediction results of unsafe behaviors at different scales.
[0017] As a preferred technical solution, the output of the prediction results of unsafe behaviors at different scales includes:
[0018] Predict the offset of the unsafe behavior bounding box relative to the upper left corner of each frame of the image and the aspect ratio of the bounding box to the original image, and draw the bounding box in the original image according to the offset of the upper left corner and the aspect ratio to locate the position of the unsafe behavior;
[0019] Convert the original output into the probability distribution of all categories of unsafe behaviors through the softmax function, and select the category with the highest probability as the predicted unsafe behavior category;
[0020] Use the sigmoid function to normalize the original output of the network to between 0 and 1 to obtain the recognition confidence.
[0021] As a preferred technical solution, the personnel safety file management module is used to identify the identity information of construction workers and record their working status and unsafe behaviors, including:
[0022] After identifying an unsafe behavior, the video frame in which the unsafe behavior is identified is face-matched with the face picture database of construction workers through face recognition technology and person re-identification technology to obtain the ID of the construction worker who is matched, and the matching result, time, location, personnel information, and the record of the unsafe behavior are recorded in the staff database corresponding to the ID of the construction worker.
[0023] As a preferred technical solution, the site safety management module is used to calculate the coverage area and blind area of the camera according to the floor plan information, camera position, and field of view of the construction site, including:
[0024] Taking the position coordinates of each camera as the center point coordinates of the fan-shaped area;
[0025] Determining a fan-shaped area according to the center point coordinates, viewing angle, and maximum monitoring distance of each camera;
[0026] Using geometric methods to calculate the boundary points of the fan-shaped area;
[0027] Filling the fan-shaped area according to the boundary points to obtain the filled fan-shaped area, and the filled fan-shaped area is the field of view of the camera;
[0028] Creating a matrix with the same size as the floor plan, and initializing all elements to 0; wherein, each element in the matrix represents an area on the plane;
[0029] For the field of view of each camera, traversing all matrix elements covered by the field of view and setting the covered matrix elements to 1, indicating that the area has been covered;
[0030] Comparing the overlapping parts of different cameras to determine the area not covered by any camera's field of view as the blind area.
[0031] As a preferred technical solution, it also includes highlighting the currently detected unsafe behavior on the floor plan according to the warning information of the warning module.
[0032] As a preferred technical solution, the conversational query module is used to analyze the content of the user's query request, retrieve relevant video content according to the query request content, and output a retrieval result that meets the query requirements, including:
[0033] Extracting the visual features in each frame of the video picture;
[0034] Detecting the people and relevant features of the people appearing in each frame of the video picture;
[0035] Segment the video frame using a segmentation strategy based on the differences in the visual features and the dynamic changes in the number of people, to obtain multiple video segments;
[0036] Select the same number of video frames from the multiple video segments at equal intervals as key frames;
[0037] Perform visual understanding on the image features in the key frames, generate a retrieval result of the first natural language description of the segment and return it to the user. The retrieval result of the first natural language description includes the people appearing in the video segment, their behaviors, the scene background, and the interaction information between people.
[0038] As a preferred technical solution, after completing the extraction of the visual features and the segmentation of the video frame, it further includes organizing the multiple video segments obtained by the segmentation into a multi-level video tree structure to achieve a hierarchical video representation from global to local, including:
[0039] Use the complete video frame as the video segment of the root node, and divide the video segment into multiple sub-segments through a segmentation strategy; among them, the information of each sub-segment is stored as a sub-node in the tree, and the root node serves as the parent node of the sub-node;
[0040] Apply the segmentation strategy to each sub-segment again to further divide the sub-segment into small sub-segments; among them, the small sub-segments are added as sub-nodes of the current node to the tree structure in a recursive manner until a certain segment cannot be further divided;
[0041] Among them, the start and end times of the segment in the video, content annotations, the construction worker IDs included in the current segment and their work, and all sub-nodes linked to the current node are stored in each node.
[0042] As a preferred technical solution, it further includes:
[0043] The conversational query module is used to analyze the content of the user's query request, perform retrieval across multiple video contents according to the query request content, and output a retrieval result that meets the query requirements. Specifically:
[0044] Construct a multi-level video tree structure for each video according to multiple videos and their related time and location information;
[0045] Screen relevant video tree structures from the multi-level video tree structure according to the time and location information mentioned in the user's query;
[0046] Use the root nodes of the screened relevant video tree structures as the initial node list;
[0047] Input the video description content of all nodes in the initial node list into a pre-established language model for preliminary query analysis;
[0048] If the description content of the current node meets the user's query request, generate a retrieval result with a second natural language description that meets the user's query request and return it to the user; if the description content of the current node does not meet the user's query request, recommend child nodes to be added to a new node list based on the description content and context content for iterative query;
[0049] In each round of list query, the pre-established language model performs a relevance analysis of the current node information with the query request content;
[0050] After several rounds of iterative query, according to the relevance analysis, integrate the information from multiple nodes and multiple video tree structures, generate a retrieval result with a second natural language description that meets the user's query request and return it to the user.
[0051] As a preferred technical solution, it further includes a security report generation module;
[0052] The security report generation module is used to generate a security report according to the detection result of the early warning module or according to the retrieval result of the conversational query module.
[0053] In summary, compared with the prior art, the effective effects brought by the technical solution provided by this application at least include:
[0054] This application proposes an unsafe behavior monitoring system, including a real-time monitoring module for obtaining video images of the construction site scene in real time; an early warning module for detecting unsafe behaviors in the video images of the construction site scene and giving early warnings about the unsafe behaviors; a personnel safety file management module for identifying the identity information of construction workers and recording the working status and unsafe behaviors of the construction workers; a site safety management module for calculating the coverage area and blind area of cameras based on the floor plan information, camera positions, and fields of view of the construction site, and marking the coverage area and the blind area on the floor plan; a conversational query module for analyzing the content of the user's query request, retrieving relevant video content according to the content of the query request, and outputting a retrieval result that meets the query requirements; and a safety education and training module for generating education and training content based on the detected unsafe behaviors. Through the collaborative action of the real-time monitoring module, early warning module, personnel safety file management module, site safety management module, conversational query module, and safety education and training module, this application can automatically detect unsafe behaviors in real time and give early warnings, improving the monitoring efficiency and monitoring effect; at the same time, by integrating the spatio-temporal information in multiple video data through the conversational query module, effectively mining and utilizing the value of video data, completing the analysis of video content based on different fine-grained levels, thereby improving the efficiency and accuracy of safety monitoring, reducing the incidence of safety accidents, and significantly enhancing the overall effect of multi-scene safety monitoring. Description of the Drawings
[0055] To more clearly illustrate the technical solutions in the embodiments of this application, the following will briefly introduce the drawings required for the description of the embodiments. Obviously, the following-described drawings are only some embodiments of this application. For those of ordinary skill in the art, without creative efforts, other drawings can also be obtained based on these drawings.
[0056] Figure 1 It is a block diagram of a module of an unsafe behavior monitoring system provided by an embodiment of this application;
[0057] Figure 2 It is a flowchart of the early warning module detecting unsafe behaviors through the YOLOv5 model provided by an embodiment of this application. Detailed Embodiments
[0058] To enable those skilled in the art to better understand the solution of this application, the following will clearly and completely describe the technical solutions in the embodiments of this application in conjunction with the drawings in the embodiments of this application. Obviously, the described embodiments are only some embodiments of this application, rather than all embodiments. Based on the embodiments in this application, all other embodiments obtained by those skilled in the art without creative efforts belong to the scope of protection of this application.
[0059] References to "embodiments" in this application mean that the specific features, structures, or characteristics described in connection with the embodiments can be included in at least one embodiment of this application. The phrase appears in various places in the specification and does not necessarily refer to the same embodiment each time, nor are they independent or alternative embodiments mutually exclusive of other embodiments. Those skilled in the art will explicitly and implicitly understand that the embodiments described in this application can be combined with other embodiments.
[0060] Embodiment:
[0061] Please refer to Figure 1 , in an embodiment of this application, an unsafe behavior monitoring system is provided, including a real-time monitoring module, an early warning module, a personnel safety file management module, a site safety management module, a conversational query module, and a safety education and training module;
[0062] The real-time monitoring module is used to obtain the video images of the construction site scene in real time;
[0063] The early warning module is used to detect unsafe behaviors in the video images of the construction site scene and give early warnings for the unsafe behaviors;
[0064] The personnel safety file management module is used to identify the identity information of construction workers and record the working status and unsafe behaviors of construction workers;
[0065] The site safety management module is used to calculate the coverage area and blind area of the camera according to the floor plan information, camera position, and field of view of the construction site, and mark the coverage area and the blind area on the floor plan;
[0066] The conversational query module is used to analyze the content of the user's query request, retrieve relevant video content according to the content of the query request, and output a retrieval result that meets the query requirements;
[0067] The safety education and training module is used to generate education and training content according to the detected unsafe behaviors.
[0068] Among them, the real-time monitoring module, the early warning module, the personnel safety file management module, the site safety management module, and the conversational query module are the manager function modules in the monitoring system.
[0069] The safety education and training module is the construction worker function module in the monitoring system.
[0070] The monitoring system in the embodiments of this application supports adding a single camera or batch importing cameras by inputting the connection parameters and information of the cameras, so as to implement an integrated monitoring system for multiple cameras. Among them, the monitoring system allows users to customize the monitoring display layout and has interactive functions such as click-to-zoom to provide a more intuitive and convenient monitoring experience. In addition, the monitoring system uses advanced computer vision-related technologies and relies on the powerful computing power provided by the deployed edge computing devices to achieve round-the-clock real-time monitoring of the scene environment and safety conditions, automatically identify and highlight warning of risk behaviors, and magnify the functions of relevant cameras.
[0071] Furthermore, in this application, the success rate of obtaining the video image by inputting the camera connection parameters is not less than 98%; the average processing time for batch importing 100 camera information is within 3 minutes, improving the operation efficiency.
[0072] Specifically, the input for camera connection and import is the Ezviz Cloud key parameters filled in by the user when adding a camera, or the Excel table (multiple Ezviz Cloud key parameters) uploaded during batch addition; among them, the Ezviz Cloud key parameters include: Appkey (application key), AppSecret (the secret part of the application key), DeviceSerial (device serial number), and ChannelNO (channel number); in addition, the input for automatic warning and page interaction is multiple HLS (HTTP Live Streaming) video streams transmitted after connecting to the remote camera. After the camera is successfully connected, the HLS (HTTP Live Streaming) video stream corresponding to the camera will be output for subsequent monitoring and analysis. The output of automatic warning and page interaction is the HLS video stream processed by using the detection model for detection and opencv-python, including the bounding boxes of identified unsafe behaviors, the categories of unsafe behaviors, and the recognition confidence.
[0073] More specifically, the HLS video stream obtained after the successful connection of the Ezviz Cloud key parameters filled in by the user is subjected to recognition and detection by the YOLOv5 model fine-tuned with the dataset, and the bounding boxes of identified unsafe behaviors, the categories of unsafe behaviors, and the recognition confidence are obtained. The HLS video stream and the recognition results are spliced using opencv-python to obtain the processed HLS video stream.
[0074] Among them, HLS refers to the protocol used for network streaming media transmission.
[0075] As one of the implementation manners, the warning module uses the YOLOv5 model to detect the unsafe behaviors in the video images of the construction site scene.
[0076] Among them, the YOLOv5 model is optimized using model pruning technology after being fine-tuned with a large number of unsafe behavior datasets, with the YOLOv5 basic object detection model as the pre-training model. Compared with the original YOLOv5 model, the fine-tuned and optimized YOLOv5 model has made the following improvements:
[0077] (1) Higher detection accuracy for unsafe behaviors: The fine-tuned model learns knowledge about safety detection from a large number of unsafe behavior data samples, and can more accurately detect and identify specific unsafe behavior targets existing in the dataset, reducing the occurrence of missed detections and false detections. (2) Improved robustness of the model and tolerance to inputs: Through fine-tuning, the diversity of model inputs has been greatly increased, further strengthening the model's feature extraction ability, enabling the fine-tuned model to better adapt to the application scenarios of unsafe behavior detection and enhancing the robustness of the model. At the same time, the fine-tuned model is more tolerant of inputs, and can still detect unsafe behaviors relatively accurately even when the input is occluded and has low visibility.
[0078] (3) Faster training convergence speed: Since fine-tuning is carried out on the basis of a pre-training model, the weight parameters of the model are initialized to relatively reasonable values, instead of starting training from zero initial values like the original YOLOv5 model. Therefore, the time required for the fine-tuned model to converge during training is shorter than that of training a basic YOLOv5 model from scratch.
[0079] (4) Smaller model size and faster inference speed: In addition to fine-tuning, the improved YOLOv5 model deployed also uses the illustrated model pruning optimization technology. After optimization, the parameter size of the model has been reduced, and at the same time, the inference speed is faster than that of the original model.
[0080] As one of the implementation manners, please refer to Figure 2 , the warning module uses the YOLOv5 model to detect unsafe behaviors in the video images of the construction site scene, including:
[0081] Extracting deep features of each frame image in the video image through the backbone network in the YOLOv5 model;
[0082] Inputting the extracted deep features into the feature pyramid network in the YOLOv5 model, and the feature pyramid network fuses deep features of different scales to obtain fused features;
[0083] Inputting the fused features into the detection head in the YOLOv5 model, and the detection head outputs prediction results of unsafe behaviors at different scales.
[0084] As one of the implementation manners, the output of prediction results of unsafe behaviors at different scales includes:
[0085] Predict the offset of the unsafe behavior bounding box relative to the upper left corner of each frame of the image, as well as the aspect ratio of the bounding box and the original image. Draw the bounding box in the original image according to the offset of the upper left corner and the aspect ratio to locate the position of the unsafe behavior;
[0086] Through the softmax function, convert the original output into the probability distribution of all categories of unsafe behaviors, and select the category with the highest probability as the predicted unsafe behavior category;
[0087] Use the sigmoid function to normalize the original output of the network between 0 and 1 to obtain the recognition confidence.
[0088] Specifically, the YOLOv5 model in the embodiments of the present application consists of three parts: a backbone network, a feature pyramid network, and a detection head. Among them, the backbone network is used to extract the features of the image. YOLOv5 uses CSPDarknet53 as the backbone network to extract the deep features of the image through a series of operations such as convolutional layers and residual blocks; the feature pyramid network is used to fuse features of different scales to better detect targets of different sizes. YOLOv5 combines PANet to further enhance the feature fusion ability; the detection head outputs predictions at different scales, including bounding boxes, categories, and confidences.
[0089] The present application fine-tunes the model with the labeled common unsafe behavior data, supports the user to select the types of unsafe behaviors to be detected, and thus completes the detection to achieve the generalization of the usage scenario.
[0090] After detecting the unsafe behavior, according to the severity of the unsafe behavior, the monitoring system will classify the unsafe behavior, and its classification includes:
[0091] 1), Severe levels 1-2 (may cause mass deaths and injuries), and its unsafe behaviors include: stacking loads on the edge of the construction site foundation pit, overtopping of the external scaffolding at the construction site, smoking or using open flames at the gas station, chemical leakage at the chemical plant, smoking or using open flames in the mine.
[0092] 2), Relatively severe levels 3-4 (prone to cause casualty accidents), and its unsafe behaviors include: cross-operation within the lifting radius at the construction site, water accumulation in the construction site foundation pit, water accumulation in the tower crane foundation at the construction site, water accumulation in the scaffolding foundation at the construction site, failure to wear a safety rope during high-altitude / edge operations at the construction site, uneven lengths of lifted objects at the construction site, lack of protection around the construction site foundation pit, uncovered holes at the construction site, lack of upper protection for the operation platform at the construction site, failure to hang the safety net at the construction site, deformation of the scaffolding members at the construction site, failure to handle vehicle failures at the gas station, improper storage of dangerous goods at the gas station, excessive dust concentration in the mine, failure to wear protective clothing during the maintenance of power facilities, open flames or smoke indoors.
[0093] 3) Severe levels 5 - 6 (may cause casualty accidents), and their unsafe behaviors include: not wearing safety helmets at the construction site, trailing construction site cables on the ground, missing protective covers for the cable connection posts of construction site welding machines, smoking at the construction site, not setting up warning signs at gas stations, swimming in reservoirs, and fatigue driving by drivers.
[0094] 4) Civilization hazards below level 7: exposed soil at the construction site not covered, large crowds gathering, and not wearing protective gear during sports.
[0095] By obtaining video images of the construction site scene through the real - time monitoring module and combining with the early warning module to conduct real - time detection and early warning of unsafe behaviors, it can quickly respond to potential safety hazards, avoid the occurrence of accidents. The automated monitoring and early warning mechanism greatly improves the monitoring efficiency and accuracy compared with traditional manual patrols and regular inspections, and reduces safety accidents caused by human negligence.
[0096] The monitoring system supports users to input personal information (such as name, height, and face images, etc.) of construction workers in a single or batch manner into the personnel safety file database, and supports adding, deleting, modifying, and querying construction worker information. According to the database information, through technologies such as face recognition and pedestrian re - identification, the identities of construction workers are automatically matched and identified, and their working status and detected unsafe behaviors (such as not wearing seat belts, staying in dangerous areas, illegal operations, etc.) are completely recorded in the safety file. The platform automatically counts the safety performance based on the records and feeds back relevant statistical data (such as lists of the best, the worst, and those with the greatest progress, etc.), presenting them to users in the form of interactive queries and data visualization, so that managers can carry out training in a timely manner and establish management mechanisms.
[0097] As one of the implementation methods, the personnel safety file management module is used to identify the identity information of construction workers and record the working status and unsafe behaviors of construction workers, including:
[0098] After identifying an unsafe behavior, the video frame of the identified unsafe behavior is face - matched with the face image database of construction workers through face recognition technology and pedestrian re - identification technology to obtain the ID of the matched construction worker, and the matching result, time, location, personnel information, and unsafe behavior record are stored in the staff database corresponding to the construction worker ID.
[0099] Specifically, the input for personnel matching and behavior recording is the face images of construction workers added by users or uploaded during batch addition, as well as the video frames after fine-tuning the YOLOv5 model for recognition and annotation. The input for database management and data visualization is the basic information and face images of construction workers added by users or uploaded during batch addition, as well as the count statistics of unsafe behavior records. The output of personnel matching and behavior recording is the construction worker ID matched by face recognition and person re-identification; the output of database management and data visualization is the interactive Html web page screen display. The implementation steps include: collecting image and video data containing construction workers; annotating the positions of faces and pedestrians to determine the identity information of each person; installing Python and related dependency libraries such as OpenCV, dlib, and pytorch in the monitoring system and downloading pre-trained models for face recognition and person re-identification; secondly, using the YOLOv5 model to detect pedestrians in the video data; using deep learning-based models such as OSNet and ResNet for pedestrian feature extraction; storing the features of known construction workers in the database; finally, comparing the detected pedestrian features with the features of personnel in the database to find the matching results; recording the matching results in the database, including time, location, personnel information, and unsafe behavior; and automatically triggering an early warning mechanism according to preset rules to notify relevant personnel to take necessary actions.
[0100] As one of the implementation manners, the site safety management module is used to calculate the coverage area and blind area of the camera according to the floor plan information, camera position, and field of view of the construction site, including:
[0101] Taking the position coordinates of each camera as the center point coordinates of the fan-shaped area;
[0102] Determining a fan-shaped area according to the center point coordinates, viewing angle (θ), and maximum monitoring distance (which can be converted into radius R) of each camera;
[0103] Using geometric methods to calculate the boundary points of the fan-shaped area;
[0104] Filling the fan-shaped area according to the boundary points to obtain the filled fan-shaped area, and the filled fan-shaped area is the field of view range of the camera;
[0105] Creating a matrix or grid map with the same size as the floor plan, and initializing all elements to 0; where each element in the matrix represents a region on the plane;
[0106] For the field of view range of each camera, traversing all matrix elements covered by the field of view range and setting the covered matrix elements to 1, indicating that the area has been covered; the area with a value of 0 in the matrix or grid map is the uncovered area;
[0107] Compare the overlapping parts of different cameras and determine the areas not covered by any camera's field of view as blind spots.
[0108] Furthermore, the blind spots are areas that cannot be covered between the fields of view of cameras, which may be caused by perspective occlusion or perspective gaps. Therefore, by calculating the overlapping parts of different cameras' fields of view, duplicate calculation of covered areas can be avoided. Finally, the areas not covered by any camera's field of view are identified and defined as blind spots.
[0109] After obtaining the covered areas and uncovered areas of the cameras, use different colors to distinguish and display the covered areas and uncovered areas of each camera on the plane map; among them, highlight the blind spots with special marks or colors for quick identification.
[0110] As one of the implementation manners, it further includes highlighting the currently detected unsafe behaviors on the plane map according to the warning information of the warning module.
[0111] In the embodiments of the present application, through the import of the plane map and the automatic calculation of the monitored coverage area, blind spots can be automatically detected and calculated, and a warning can be sent to the user in a timely manner, prompting the manager to strengthen the management of specific areas, thereby improving the comprehensiveness and timeliness of safety management.
[0112] In the embodiments of the present application, the conversational query module aims to retrieve video content through natural language query, with particular attention to the extraction and analysis of person information, so as to meet the user's other person-related query needs in addition to the analysis of specific unsafe behavior records, and combine with the real-time monitoring module and the warning module to broaden the usage scenarios. At the same time, through the query scope of multiple videos, the insight into the spatio-temporal correlation information of multiple videos is realized. The user can initiate a query request in the form of text input, and use a large language model (LLM) to parse the query content to accurately understand the user's needs for video information. By screening the videos that meet the requirements to form a list, and searching for the content that meets the conditions layer by layer in the tree structure, and generating a summary output with the help of the LLM, a concise answer that meets the query requirements is finally provided.
[0113] As one of the implementation manners, the conversational query module is used to analyze the content of the user's query request, retrieve relevant video content according to the query request content, and output a retrieval result that meets the query requirements, including:
[0114] To construct a hierarchical video representation, first perform visual feature extraction and segmentation processing on the video. This process realizes the preliminary segmentation of long videos by extracting the visual and semantic information of the video, combining dynamic change detection and the change in the number of people. The specific steps are as follows:
[0115] S1. Use the ViCLIP model to extract the visual features in each frame of the video, generating a frame-level semantic embedding representation;
[0116] Among them, the semantic embedding representation captures the visual information and scene semantics within the frame, providing a solid foundation for subsequent segmentation;
[0117] S2. Adopt a Tracking model to detect the people and relevant features of the people that appear in each frame of the video, including time, spatial position, and person ID;
[0118] S3. According to the differences in the visual features and the dynamic changes in the number of people, use a segmentation strategy to segment the video frames to obtain multiple video segments; specifically, when any of the following conditions is met, the video segment is segmented into a new segment:
[0119] The visual feature difference between two adjacent frames reaches a set threshold, indicating a significant change in the video content;
[0120] The visual difference between the current frame and the first frame of the segment reaches the set threshold;
[0121] The number of people appearing in the video changes, such as the addition or reduction of key people;
[0122] Through the above segmentation strategy, a long video can be divided into video segments with more consistent content, and the visual features and semantic information in each segment are more similar. This segmentation method can not only capture significant scene changes but also effectively handle semantic switches caused by the dynamic activities of people, thus providing high-quality basic data for the subsequent construction of a multi-level video tree structure.
[0123] After completing the video feature extraction and segmentation, one of the key tasks is to understand the content of each video segment. This process aims to generate an accurate text description to summarize the people, behaviors, scenes, and their contexts in the video segment, helping the system understand the semantic content of the video.
[0124] S4. Select the same number of video frames from the multiple video segments at equal intervals as key frames;
[0125] S5. Use sharegpt4video to perform visual understanding on the image features in the key frames, generating a retrieval result of the first natural language description of the segment and returning it to the user. The retrieval result of the first natural language description includes information such as the people who appear in the video segment, their behaviors, the scene background, and the interactions between people. At the same time, this processing method has a large key frame selection step size for the segments at the top of the video tree and a small step size for the bottom segments, thus realizing the transformation of information from coarse-grained to fine-grained, which is convenient for the model to be efficient and accurate.
[0126] After the extraction of the visual features and the segmentation of the video frames, it further includes organizing the multiple video segments obtained from the segmentation into a multi-level video tree structure to achieve a hierarchical video representation from global to local. Specifically, the construction of the multi-level video tree follows the steps below:
[0127] Regarding the video frame of the complete video as the video segment of the root node, divide the video segment into multiple sub-segments through a segmentation strategy; among them, the information of each sub-segment is stored as a sub-node in the tree, and the root node serves as the parent node of the sub-nodes; apply the segmentation strategy to each sub-segment again to further divide the sub-segment into smaller sub-segments; among them, the smaller sub-segments are added to the tree structure as the sub-nodes of the current node in a recursive manner until a certain segment cannot be further divided; finally, each video is organized into a tree structure with the complete video as the root node. The top-level nodes represent the global information of the video, while the bottom-level nodes provide a more refined content representation based on the key frames of the refined segments.
[0128] Among them, in each node, store the start and end times of the segment in the video, content annotations, the ID of the construction workers included in the current segment and their work, and all the sub-nodes linked to the current node.
[0129] Through the recursive video tree structure construction method, a long video can be refined layer by layer into a multi-granularity representation with a clear structure, thus taking into account both the global overview and local details, and supporting efficient multi-video semantic query and reasoning.
[0130] It further includes: the conversational query module, which is used to analyze the content of the user's query request, retrieve across multiple video contents according to the query request content, and output the retrieval results that meet the query requirements. Specifically:
[0131] 1. Construct a multi-level video tree structure for each video according to multiple videos and their related time and location information;
[0132] The time and location information of each video and its corresponding video tree are used as candidate items to input into the query model as the basic data for subsequent queries. This step ensures that the system can select appropriate video data for queries according to the spatio-temporal requirements of the question.
[0133] 2. Screen the relevant video tree structures from the multi-level video tree structure according to the time and location information mentioned in the user's query; this step ensures that only the videos closely related to the query content are selected by comparing with the time stamps and location information of the videos, thus effectively reducing the interference of irrelevant information on the query results.
[0134] 3. Use the root nodes of the screened relevant video tree structures as the initial node list;
[0135] The root node contains the global information of the video, such as the time range of the video and the segment description generated based on key frames. These root nodes provide the basic information for subsequent queries and are the starting points for multi-level reasoning.
[0136] 4. Input the video description content of all nodes in the initial node list into a pre-established language model for preliminary query analysis;
[0137] 5. If the description content of the current node meets the user's query request, generate a retrieval result in the second natural language description that meets the user's query request and return it to the user; if the description content of the current node does not meet the user's query request, recommend child nodes to be added to the new node list based on the description content and context content (i.e., suggest which child nodes of the current node may contain potential key information and indicate the child nodes that may need attention), and perform iterative queries until sufficient information is extracted and a retrieval result in the second natural language description that meets the user's query request can be generated;
[0138] 6. In each round of list query, the pre-established language model performs a relevance analysis of the current node information with the content of the query request;
[0139] 7. After several rounds of iterative queries, according to the relevance analysis, integrate the information from multiple nodes and multiple video trees, generate a retrieval result in the second natural language description that meets the user's query request and return it to the user. This process ensures the efficient integration of cross-video information and can answer more complex cross-video query questions.
[0140] Through the above steps, the platform can efficiently extract and integrate relevant information from multiple video trees to ensure that accurate answers are provided to users. During the entire query process, the query strategy can be dynamically adjusted, and the spatio-temporal information of the video is used to optimize the query process and reduce the interference of irrelevant content. In addition, through iterative queries, the query scope can be gradually refined when the initial information is insufficient, and finally high-quality answers can be generated; this method is applicable to multi-video scenarios with complex spatio-temporal associations and can effectively support cross-perspective and cross-time video content analysis and question-and-answer tasks.
[0141] As one of the implementation manners, the manager function module in the monitoring system of the present application further includes a security report generation module;
[0142] Among them, the security report generation module is used to generate a security report according to the detection result of the early warning module or according to the retrieval result of the conversational query module;
[0143] Furthermore, the security report generation module can combine the detection results of the early warning function module, utilize the empowerment of the language model for past surveillance video data, automatically generate a security graphic report, and provide rectification suggestions.
[0144] Moreover, the security report generation module can also extract key information from the video and, in combination with the content of the conversational query, organize this information into a JSON format for output. After obtaining the output JSON format data, according to the user's layout customization requirements, a structured layout description is generated by designing DSL. The user can input some instructions, such as generating a report containing a title, table, and icon, and then generate the corresponding DSL code according to these instructions. The DSL structure is similar to the JSON format and contains information such as node type, position, and size. The specific node types are as follows:
[0145] CoordinateNode: Used to define the coordinate system of the entire layout and its sub-nodes; ContainerNode: Used to display text content, such as the title of the report or data description; IconNode: Used to display icons; ImageNode: Used to display pictures; TableNode: Used to display data tables. Through these node types, the DSL describes each part of the layout. For example, the title of the report, statistical tables, involved images, and detailed information on unsafe behaviors will all be defined in the DSL.
[0146] The generated DSL layout code will be processed by a parser, which converts the layout represented by the DSL into a corresponding tree structure, and each node in the tree structure will be parsed according to its type and position. To ensure a reasonable layout, the system will also apply some constraints, such as ensuring that tables or pictures do not exceed the visible range of the page.
[0147] After parsing and processing the DSL layout, the layout tree structure will be rendered into a PDF format report. The renderer will call the corresponding rendering method according to the type of each node; for example: ContainerNode: Used to draw the title and text description of the report; ImageNode: Used to display pictures of key figures; TableNode: Used to display statistical data, such as the number of special individuals detected by the camera and the occurrence of dangerous events. Finally, all nodes will be drawn on the PDF page to generate a graphic and text combined security report. The security report not only contains key information such as personnel, vehicles, and dangerous events in the video, but also can display relevant images and detailed descriptions, enabling users to intuitively see the analysis results.
[0148] The automatic generation of the graphic and text security report improves the intuitiveness of the feedback, enabling managers to more intuitively understand the security management situation.
[0149] The accuracy of the safety report output by this application reaches over 95%, including the practicality of the rectification suggestions; the report generation time is completed within 1 hour after the end of each day.
[0150] As one of the implementation manners, the safety education and training module of this application is used to generate education and training content according to the detected unsafe behaviors. The safety management interaction form is enriched through the safety education and training module, and the management effect is improved.
[0151] In summary, this application provides an unsafe behavior monitoring system, which can achieve the following effects:
[0152] (1) By implementing functions such as floor plan import, automatic calculation of the monitored coverage area, and automatic detection and calculation of blind spots in area division, etc., blind spots are automatically detected and calculated, and at the same time, a blind spot warning is issued to the user to urge the user to strengthen management and improve equipment in specific areas; (2) In the process of safety management, the monitoring system of this application only requires the input of monitoring cameras for monitoring equipment, and does not require other sensors as assistance, with relatively low requirements for equipment quality and low usage threshold costs. (3) This application realizes a conversational query of the safety management results with the help of a large language model, which is beneficial for personalized acquisition of relevant data. At the same time, it realizes the automatic generation of graphic reports, improves the intuitiveness of feedback, and improves the feedback efficiency. (4) Considering the bilateral monitoring of the monitor and the construction personnel, this application provides feedback on the management results to the monitor, and provides personal unsafe behavior records and personalized safety training content generation for the construction personnel, enriching the safety management interaction form. (5) By constructing a structured video tree structure for long monitoring videos, content analysis according to different fine-grained contents is realized, and with the help of a large language model, search of video information from rough to detailed is realized to gain insights into video information; at the same time, through cross-video video content query, content analysis and capture are completed under different spatio-temporal backgrounds, and then spatio-temporal laws are summarized through insights. Compared with the single-video query method, it has a wider query range and is more systematic in terms of person recognition and behavior analysis and summary.
[0153] It should be noted that for the foregoing embodiments, for the sake of simple description, they are all expressed as a series of action combinations. However, those skilled in the art should know that this application is not limited by the described action sequence, because according to this application, certain steps can be performed in other sequences or simultaneously.
[0154] Those of ordinary skill in the art can understand that implementing all or part of the processes in the above embodiments can be achieved by instructing relevant hardware through a computer program. The program can be stored in a non-volatile computer-readable storage medium. When the program is executed, it can include the processes of the embodiments of the above methods. Among them, any reference to a memory, storage, database, or other medium used in the embodiments provided in this application can include non-volatile and / or volatile memories. Non-volatile memory can include read-only memory (ROM), programmable ROM (PROM), electrically programmable ROM (EPROM), electrically erasable programmable ROM (EEPROM), or flash memory. Volatile memory can include random access memory (RAM) or external cache memory. By way of illustration and not limitation, RAM is available in various forms, such as static RAM (SRAM), dynamic RAM (DRAM), synchronous DRAM (SDRAM), double data rate SDRAM (DDR SDRAM), enhanced SDRAM (ESDRAM), synchronous link DRAM (SLDRAM), Rambus direct RAM (RDRAM), direct memory bus dynamic RAM (DRDRAM), and Rambus dynamic RAM (RDRAM), etc.
[0155] The technical features of the above embodiments can be combined arbitrarily. For the sake of concise description, not all possible combinations of the technical features in the above embodiments are described. However, as long as there is no contradiction in the combination of these technical features, it should be considered as the scope described in this specification.
[0156] The above embodiments are preferred embodiments of the present application, but the embodiments of the present application are not limited to the above embodiments. Any other changes, modifications, substitutions, combinations, and simplifications made without departing from the spirit and principle of the present application shall be equivalent replacement methods and are all included in the protection scope of the present application.
Claims
1. An unsafe behavior monitoring system, characterized in that: include: Real-time monitoring module, early warning module, personnel safety file management module, site safety management module, conversational query module and safety education and training module; The real-time monitoring module is used to obtain video images of the construction site scene in real time; The early warning module is used to detect unsafe behaviors in the video images of the construction site scene and issue an early warning for the unsafe behaviors; The personnel safety file management module is used to identify the identity information of construction personnel and record the working status and unsafe behaviors of construction personnel; The site safety management module is used to calculate the coverage area and blind spot of the camera according to the plan information of the construction site, the position and field of view of the camera, and mark the coverage area and the blind spot on the plan; The conversational query module is used to analyze the query request content of the user, retrieve relevant video content according to the query request content, and output the search results that meet the query requirements; The safety education and training module is used to generate education and training content based on the detected unsafe behaviors.
2. The unsafe behavior monitoring system according to claim 1, characterized in that: The early warning module uses the YOLOv5 model to detect unsafe behaviors in the video footage of the construction site scene, including: Extracting deep features of each frame of the video through the backbone network in the YOLOv5 model; The extracted deep features are input into the feature pyramid network in the YOLOv5 model, and the feature pyramid network fuses the deep features of different scales to obtain fused features; The fused features are input into the detection head in the YOLOv5 model, and the detection head outputs prediction results of unsafe behaviors at different scales.
3. The unsafe behavior monitoring system according to claim 1, characterized in that: The output is the prediction results of unsafe behaviors at different scales, including: Predicting the offset of the unsafe behavior bounding box relative to the upper left corner of each frame image and the aspect ratio of the bounding box and the original image, and drawing the bounding box in the original image according to the offset of the upper left corner and the aspect ratio to locate the position of the unsafe behavior; The softmax function is used to convert the original output into the probability distribution of all categories of unsafe behaviors, and the category with the highest probability is selected as the predicted unsafe behavior category; Using the sigmoid function, the original output of the network is normalized to between 0 and 1 to obtain the recognition confidence.
4. The unsafe behavior monitoring system according to claim 1, characterized in that: The personnel safety file management module is used to identify the identity information of construction personnel and record the working status and unsafe behaviors of construction personnel, including: When an unsafe behavior is identified, the video frame of the unsafe behavior is matched with the facial image database of the construction workers through face recognition technology and pedestrian re-identification technology to obtain the matched construction worker ID, and the matching result, time, location, personnel information and unsafe behavior record are recorded in the staff database corresponding to the construction worker ID.
5. The unsafe behavior monitoring system according to claim 1, characterized in that: The site safety management module is used to calculate the coverage area and blind spot of the camera according to the floor plan information of the construction site, the camera position and the field of view, including: The position coordinates of each camera are used as the center point coordinates of the fan-shaped area; Determine a sector area according to the center point coordinates, viewing angle, and maximum monitoring distance of each camera; Calculate the boundary points of the sector area using a geometric method; Filling the sector-shaped area according to the boundary points to obtain a filled sector-shaped area, wherein the filled sector-shaped area is the field of view of the camera; Create a matrix with the same size as the plane image, and initialize all elements to 0; wherein each element in the matrix represents an area on the plane; For the field of view of each camera, traverse all matrix elements covered by the field of view, and set the covered matrix elements to 1, indicating that the area has been covered; Compare the overlapping parts of different cameras to determine the areas not covered by any camera's field of view as blind spots.
6. The unsafe behavior monitoring system according to claim 5, characterized in that: The method also includes highlighting the currently detected unsafe behavior on the plan view according to the warning information of the warning module.
7. The unsafe behavior monitoring system according to claim 1, characterized in that: The conversational query module is used to analyze the query request content of the user, retrieve relevant video content according to the query request content, and output the search results that meet the query requirements, including: Extract visual features from each video frame; Detect the characters and their related features in each frame of video; According to the difference of the visual features and the dynamic change of the number of characters, the video screen is segmented using a segmentation strategy to obtain multiple video clips; Selecting the same number of video frames in the multiple video clips as key frames with equal step length; Visual understanding is performed on the image features in the key frame, and a retrieval result of a first natural language description of the clip is generated and returned to the user. The retrieval result of the first natural language description includes characters appearing in the video clip and their behaviors, scene background, and interaction information between characters.
8. The unsafe behavior monitoring system according to claim 7, characterized in that: After completing the extraction of the visual features and the segmentation of the video screen, the method further includes organizing the multiple video clips obtained by the segmentation into a multi-level video tree structure to achieve hierarchical video representation from global to local, including: The complete video screen is used as a video segment of the root node, and the video segment is divided into multiple sub-segments through a segmentation strategy; wherein the information of each sub-segment is stored as a child node in the tree, and the root node is used as the parent node of the child node; Apply the segmentation strategy to each of the sub-segments again to further divide the sub-segments into smaller sub-segments; wherein the smaller sub-segments are added to the tree structure as child nodes of the current node, and the process is performed recursively until a segment cannot be further segmented; Each node stores the start and end time of the segment in the video, content annotations, the ID of the construction worker included in the current segment and his / her work, and all child nodes linked to the current node.
9. The unsafe behavior monitoring system according to claim 1, characterized in that: Also includes: The conversational query module is used to analyze the query request content of the user, perform a search across multiple video contents according to the query request content, and output a search result that meets the query requirements, specifically: According to multiple videos and their related time and location information, a multi-level video tree structure is constructed for each video; Filtering relevant video tree structures from the multi-level video tree structure according to the time and location information mentioned in the user query; The root node of the selected related video tree structure is used as the initial node list; Inputting the video description contents of all nodes in the initial node list into a pre-established language model for preliminary query analysis; If the description content of the current node meets the user's query request, a search result of the second natural language description that meets the user's query request is generated and returned to the user; if the description content of the current node does not meet the user's query request, based on the description content and context content, a recommended child node is added to a new node list for iterative query; In each round of list query, the pre-established language model performs a correlation analysis between the current node information and the query request content; After several rounds of iterative queries, information from multiple nodes and multiple video tree structures is integrated according to the correlation analysis to generate a search result of the second natural language description that meets the user's query request and return it to the user.
10. The unsafe behavior monitoring system according to claim 1, characterized in that: It also includes a safety report generation module; The safety report generation module is used to generate a safety report according to the detection result of the early warning module or according to the retrieval result of the conversational query module.
Citation Information
Cited By
Multi-modal large model-based worker safety protection equipment missing detection method
CN121190756A
Worker safety protective equipment absence detection method based on multi-modal large model
CN121190756B