A method and system for construction site safety monitoring based on dynamic and static detection
By collecting video streams in real time on the construction site and using neural network groups to perform dynamic and static detection of construction personnel, the problem of omissions in the monitoring of construction personnel in the existing technology is solved, and real-time monitoring and risk warning of the safety status of construction personnel on the construction site is achieved.
Patent Information
- Application Number
- CN202210080857.3
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2022-01-24
- Publication Date
- 2025-06-17
- Estimated Expiration
- 2042-01-24
AI Technical Summary
The existing construction site safety monitoring methods mainly focus on construction equipment and scenarios, ignore the actual construction status of construction personnel, resulting in monitoring omissions and increasing construction risks.
The construction site safety monitoring method based on dynamic and static detection is adopted. The cameras set in different areas of the construction site collect video streams in real time, and the construction personnel are dynamically detected and statically detected by preset neural network groups (including YOLO static neural network and LSTM dynamic neural network) are used to monitor the safety status of construction personnel in real time.
By monitoring the dynamic and static state of construction personnel in real time, abnormal situations can be discovered in a timely manner, avoid construction accidents, ensure the safety of construction personnel, and improve the reliability of the project.
Smart Images

Figure CN114596518B_ABST
Abstract
Description
Background Art
[0002] In recent years, the process of urbanization in China has accelerated, and urban construction is in full swing. Construction sites can be seen in every corner of the city. The construction site is intricate, with a large number of construction equipment and personnel, and there are many security loopholes. Therefore, the safety supervision of the construction site is one of the key concerns in the field of engineering construction.
[0003] The currently commonly used method for monitoring the safety of construction sites is to obtain the scene video of the construction site, and then call the neural network that has completed deep learning to monitor and analyze the construction scene and construction equipment in the construction site video, so as to determine whether the construction site meets the safety specifications.
[0004] However, the currently commonly used monitoring method has the following technical problems: The existing monitoring method only takes the construction equipment and construction scene of the construction site as the monitoring objects, ignoring the actual construction conditions of each construction worker, resulting in omissions in monitoring; moreover, in actual construction, the safety of construction workers is the primary criterion for construction site safety. If an accident occurs to a construction worker during construction (for example, an operation accident or a personnel conflict accident), it is easy to trigger a construction crisis, resulting in the project being unable to be carried out normally or even suspended, increasing the construction risk. Summary of the Invention
[0005] The present invention provides a construction site safety monitoring method, system and safety monitoring system based on dynamic and static detection. The method monitors the dynamic and static safety of construction workers, thereby avoiding delays in the project due to construction accidents of construction workers, ensuring the safety of construction workers and improving the reliability of the project.
[0006] The first aspect of the embodiment of the present invention provides a construction site safety monitoring method based on dynamic and static detection, and the method includes:
[0007] Call the cameras set in different areas of the construction site to collect the monitoring video stream in real time, and receive the monitoring instructions input by the user;
[0008] Based on the monitoring instructions, control the preset neural network group to perform dynamic detection or static detection on the construction workers in the monitoring video stream. The preset neural network group includes: the YOLO static neural network for static detection and the LSTM dynamic neural network for dynamic detection;
[0009] When the detection result of the dynamic detection is abnormal or the detection result of the static detection is abnormal, visually display the construction workers in the monitoring video stream and trigger a monitoring alarm.
[0010] In a possible implementation manner of the first aspect, the dynamic detection includes:
[0011] Identify the human bone structure of each frame of video image in the monitored video stream, and extract bone marker points from the human bone structure to obtain a set of bone dot matrices;
[0012] Input the set of bone dot matrices into the LSTM dynamic neural network to obtain continuous action sequences between different humans in the monitored video stream;
[0013] Determine the dynamic actions of the construction workers based on the continuous action sequences;
[0014] When the dynamic actions are different from the preset set of actions, the detection result of the dynamic detection is abnormal.
[0015] In a possible implementation manner of the first aspect, the static detection includes:
[0016] Divide the monitored video stream into several video data blocks;
[0017] Based on a preset recognition frame, extract the person information of each frame of video image in the video data block, and calculate the person change rate of the construction worker or object within the duration corresponding to the video data block by using the person information;
[0018] When the person change rate is greater than the preset change rate, the detection result of the static detection is abnormal.
[0019] In a possible implementation manner of the first aspect, the dynamic training operation of the LSTM dynamic neural network is specifically:
[0020] Obtain a first video stream to be trained that has been preprocessed;
[0021] Select several first feature images containing abnormal construction workers from the first video stream to be trained, and add feature identifiers to each of the first feature images. The feature identifiers include the category, position, and status information of the construction worker;
[0022] Through the human bone key point detection technology, convert the construction workers in each of the first feature images into point-line models to obtain several human model images, and arrange the several human model images into a human model sequence;
[0023] Input the human model sequence into the LSTM network for iterative training to obtain the LSTM dynamic neural network.
[0024] In a possible implementation manner of the first aspect, the static training operation of the YOLO static neural network is specifically:
[0025] Obtain a second video stream to be trained that has been preprocessed;
[0026] Screen a number of second feature images containing abnormal construction workers from the second video stream to be trained, and add feature identifiers to each of the second feature images. The feature identifiers include the category, location, and status information of the construction workers;
[0027] Input the several second feature images into the YOLO neural network one by one for iterative training to obtain a YOLO static neural network.
[0028] In a possible implementation manner of the first aspect, the method further includes:
[0029] Obtain several evaluation results of the YOLO static neural network or the LSTM dynamic neural network;
[0030] Calculate the several evaluation results through the YOLOv5 algorithm to obtain result evaluation parameters, where the result evaluation parameters include: accuracy rate, no-miss detection rate, and comprehensive accuracy rate;
[0031] When the accuracy rate is greater than the preset accuracy rate, the no-miss detection rate is greater than the preset no-miss detection rate, and the comprehensive accuracy rate is greater than the preset comprehensive rate, determine that the YOLO static neural network or the LSTM dynamic neural network meets the detection requirements;
[0032] When the accuracy rate is less than the preset accuracy rate, the no-miss detection rate is less than the preset no-miss detection rate, or the comprehensive accuracy rate is less than the preset comprehensive rate, perform network optimization on the YOLO static neural network or the LSTM dynamic neural network.
[0033] In a possible implementation manner of the first aspect, the training method of the network optimization includes: adjusting the network structure, adjusting the network depth, adjusting the feature map dimension, and setting multi-scale training.
[0034] A second aspect of the embodiments of the present invention provides a construction site safety monitoring system based on dynamic and static detection. The system includes:
[0035] A on-site perception module, which is used to call cameras set in different areas of the construction site to collect monitoring video streams in real time and receive monitoring instructions input by users;
[0036] A detection and recognition module, which is used to control a preset neural network group based on the monitoring instructions to perform dynamic detection or static detection on construction workers in the monitoring video stream. The preset neural network group includes: a YOLO static neural network for static detection and an LSTM dynamic neural network for dynamic detection;
[0037] A visualization and alarm module, which is used to visually display the construction workers in the monitored video stream and trigger a monitoring alarm when the detection result of the dynamic detection is abnormal or the detection result of the static detection is abnormal.
[0038] Compared with the prior art, a construction site safety monitoring method and system based on dynamic and static detection provided by an embodiment of the present invention have the following beneficial effects: The present invention can perform real-time monitoring on a construction site, and use the real-time monitored video to perform dynamic safety detection and static safety monitoring on construction workers to determine whether there are abnormalities among the construction workers at the construction site, thereby avoiding the situation of delaying the project due to construction accidents of construction workers, ensuring the safety of construction workers, and improving the reliability of the project. Description of the Drawings
[0039] Figure 1 is a flowchart of a construction site safety monitoring method based on dynamic and static detection provided by an embodiment of the present invention;
[0040] Figure 2 is a schematic diagram for identifying a continuous action sequence provided by an embodiment of the present invention;
[0041] Figure 3 is a flowchart of neural network training provided by an embodiment of the present invention;
[0042] Figure 4 is an operation flowchart of a construction site safety monitoring method based on dynamic and static detection provided by an embodiment of the present invention;
[0043] Figure 5 is a schematic diagram of the structure of a construction site safety monitoring system based on dynamic and static detection provided by an embodiment of the present invention. Detailed Embodiments
[0044] Next, the technical solutions in the embodiments of the present invention will be clearly and completely described in conjunction with the drawings in the embodiments of the present invention. Obviously, the described embodiments are only a part of the embodiments of the present invention, rather than all the embodiments. All other embodiments obtained by those of ordinary skill in the art based on the embodiments of the present invention without making creative efforts shall fall within the protection scope of the present invention.
[0045] The currently commonly used monitoring methods have the following technical problems: The existing monitoring methods only take the construction equipment and construction scenes of the construction site as the monitoring objects, ignoring the actual construction conditions of each construction worker, resulting in omissions in monitoring; moreover, during actual construction, the safety of construction workers is the primary criterion for construction site safety. If an accident occurs to a construction worker during construction (for example, an operation accident or a personnel conflict accident), it is easy to trigger a construction crisis, resulting in the project being unable to be executed normally or even stopped, increasing the construction risk.
[0046] To solve the above problems, the following will introduce and illustrate in detail a construction site safety monitoring method based on dynamic and static detection provided by an embodiment of the present application through the following specific embodiments.
[0047] In one embodiment, the method can be applicable to a background server, which can be communicatively connected to multiple cameras and can also be connected to a user terminal. The user terminal can be a control machine or a terminal computer. The user terminal can be placed in the monitoring room of the construction site for users to operate.
[0048] Referring to Figure 1 , a flowchart of a construction site safety monitoring method based on dynamic and static detection provided by an embodiment of the present invention is shown.
[0049] Among them, by way of example, the construction site safety monitoring method based on dynamic and static detection may include:
[0050] S11. Call cameras set in different areas of the construction site to collect monitoring video streams in real time and receive monitoring instructions input by the user.
[0051] In one embodiment, multiple cameras can be pre-set in different places or different areas of the construction site, or multiple cameras can be set in the same place or the same area. The cameras can collect the real-time conditions of the area in real time, so that the monitoring video stream of the area can be obtained in real time. Among them, the monitoring video stream can include the environment and construction workers in the area.
[0052] The monitoring instruction can be an instruction input by the user when operating the user terminal. The monitoring instruction can be used to select the monitoring type. For example, the monitoring scenario, construction equipment or objects can be selected, or the personal status, working status or construction actions of construction workers can also be monitored, etc. The user can perform corresponding operations according to their actual needs.
[0053] S12. Based on the monitoring instruction, control a preset neural network group to perform dynamic detection or static detection on the construction workers in the monitoring video stream. The preset neural network group includes: a YOLO static neural network for static detection and an LSTM dynamic neural network for dynamic detection.
[0054] After receiving the monitoring instruction, the specific monitoring type can be determined based on the monitoring instruction, so that the preset neural network group can be called according to the monitoring type to perform static detection or dynamic detection on the construction workers. Among them, static detection can detect the status of construction workers or the status of objects in the construction site, and dynamic detection can be used to detect the construction conditions or construction operations of construction workers.
[0055] Through two different detections, the safety of construction workers can be monitored in real time, which can not only safeguard the personal interests of construction workers, but also reduce the probability of construction crises caused by accidents of construction workers during the construction process, and ensure the construction progress.
[0056] In an alternative embodiment, the dynamic detection includes:
[0057] Sub-step S121: Identify the human skeletal structure of each frame of video image in the monitored video stream, and extract skeletal landmark points from the human skeletal structure to obtain a set of skeletal dot matrices.
[0058] Sub-step S122: Input the set of skeletal dot matrices into the LSTM dynamic neural network to obtain the continuous action sequences between different humans in the monitored video stream.
[0059] Sub-step S123: Determine the dynamic actions of construction workers based on the continuous action sequences.
[0060] Sub-step S124: When the dynamic actions are different from the preset set of actions, the detection result of the dynamic detection is abnormal.
[0061] In an embodiment, the set of actions may include multiple construction actions preset by users.
[0062] Refer to Figure 2 , which shows a schematic diagram of the recognition of continuous action sequences provided by an embodiment of the present invention.
[0063] In actual operation, the human skeletal structure of each frame of video image in the monitored video stream can be first identified through human skeletal key point detection technology, and then the human skeletal structure of each frame of video image can be extracted, and the skeletal landmark points of the human skeletal structure of each frame of video image can be collected. The skeletal landmark points of each frame of video image are used to construct a set of dot matrices to obtain a set of skeletal dot matrices. Then, the set of skeletal dot matrices can be input into the LSTM dynamic neural network in the form of a continuous data sequence according to the order of video images, so that the LSTM dynamic neural network can use the skeletal dot matrix as the detection object to detect the action form of the individual dot matrix of the construction worker and the relative relationship between different individual dot matrices, forming a continuous action sequence, as specifically shown in Figure 2 . Then, based on the continuous action sequences, determine the dynamic actions (such as falling, fighting, or carrying bricks, etc.) performed by the construction workers within the corresponding duration of the monitored video stream. Finally, the recognized dynamic actions can be matched with multiple construction actions in the preset set of actions. If the dynamic actions do not match any of the construction actions, it can be considered that the construction workers may not be working, and it can be determined that the detection result of the dynamic detection is abnormal.
[0064] Optionally, the action set may include multiple non-construction actions preset by the user.
[0065] It is also possible to match the identified dynamic action with multiple non-construction actions within the preset action set. If the dynamic action matches at least one non-construction action, it can be considered that the construction worker is not working and may perform a dangerous action, and it can be determined that the detection result of the dynamic detection is abnormal.
[0066] It should be noted that when performing dynamic recognition by inputting into the LSTM network, an image reading window can be defined. The calculation method for the window width W is as follows: If the required recognition time period for the index is t (s) and the frame rate of the camera is p frames / s, then the window width W = t × p.
[0067] In one embodiment, the static detection includes:
[0068] Sub-step S125: Divide the monitoring video stream into several video data blocks.
[0069] Among them, the number of divisions can be determined according to the capacity size of the monitoring video stream. If the capacity of the monitoring video stream is large, multiple video data blocks can be divided. If the capacity of the monitoring video stream is small, smaller video data blocks can be divided.
[0070] Sub-step S126: Based on the preset recognition frame, extract the person information of each frame of video image in the video data block, and calculate the person change rate of the construction worker or object within the time period corresponding to the video data block by using the person information.
[0071] In one embodiment, the person information includes object information and human body information. Among them, the object information includes the information of various equipment at the construction site, such as the placement position of the fire extinguisher, the inclination angle of the crane, the number of cement bags, etc. The human body information includes various status information of the construction workers, such as whether the construction workers wear safety helmets, whether the construction workers wear work clothes, the number of construction workers, the construction positions of the construction workers, etc.
[0072] Optionally, the size of the recognition frame can be adjusted according to actual needs.
[0073] In actual operation, the recognition frame can be used to perform information features on each frame of video image, so that the person information in the video image can be extracted. Then, within the time period corresponding to the video data block, the change rate of the person information can be obtained to get the person change rate.
[0074] For example, the recognition frame detects a fire extinguisher inside the construction site, leaning against the windowsill on the second floor, and detects the length and width of the fire extinguisher within the recognition frame. Then, calculate the difference between the length and width of the fire extinguisher within the recognition frame in the first video image of the video data block and the length and width of the fire extinguisher within the recognition frame in the last video image, and divide the difference in length and width by the duration corresponding to the video data block to obtain the human change rate regarding the placement state of the fire extinguisher.
[0075] For another example, the recognition frame detects a construction worker inside the construction site and determines that this construction worker is wearing a safety helmet. The occupied area of the safety helmet can be collected. Then, calculate the difference between the occupied area of the safety helmet in the first video image of the video data block and the occupied area of the safety helmet in the last video image, and divide the difference in occupied area by the duration corresponding to the video data block to obtain the human change rate regarding wearing a safety helmet.
[0076] Sub-step S127: When the human change rate is greater than the preset change rate, the detection result of the static detection is abnormal.
[0077] Combined with the above examples, under normal circumstances, it is consistent with the length-width ratio of the oxygen cylinder in the upright state. Correspondingly, the human change rate of the placement state of the fire extinguisher is 0, the human change rate is less than the preset change rate, and the fire extinguisher has not changed; if it is tilted or toppled, the human change rate of the placement state of the fire extinguisher is greater than 0 and exceeds the preset change rate, then it can be determined that the fire extinguisher has changed within the duration corresponding to the video data block, and an abnormality may occur at the construction site, so the detection result of the static detection is abnormal.
[0078] Correspondingly, following the above examples, if the construction worker always wears a safety helmet, the difference in the occupied area of the safety helmet is 0, and the human change rate regarding wearing a safety helmet is also 0, and the human change rate regarding wearing a safety helmet is less than the preset change rate; if the construction worker takes off the safety helmet midway, then the human change rate regarding wearing a safety helmet is greater than 0 and may exceed the preset change rate, then it can be determined that the construction worker has taken off the safety helmet and the construction worker may be in danger, so the detection result of the static detection can also be abnormal.
[0079] Through static detection, each device and construction worker inside the construction site can be detected in real time to avoid the situation of construction interruption caused by equipment or personnel abnormalities.
[0080] In order to improve the detection ability of the LSTM dynamic neural network, in one embodiment, the dynamic training operation of the LSTM dynamic neural network is specifically as follows:
[0081] S21: Obtain the first video stream to be trained that has been preprocessed.
[0082] S22. Screen several first feature images containing abnormal construction workers from the first video stream to be trained, and add feature identifiers to each of the first feature images. The feature identifiers include the category, location, and status information of the construction workers.
[0083] Among them, the feature identifiers can be manually added by the user.
[0084] S23. Through the human skeleton key point detection technology, convert the construction workers in each of the first feature images into point-line models to obtain several human model images, and arrange the several human model images into a human model sequence.
[0085] S24. Input the human model sequence into the LSTM network for iterative training to obtain an LSTM dynamic neural network.
[0086] Refer to Figure 3 , which shows a schematic flowchart of neural network training provided by an embodiment of the present invention.
[0087] During training, after obtaining the first video stream to be trained, the first video stream to be trained can be divided into two video streams for training and detection. The division method can be adjusted according to actual needs. Preferably, it can be divided according to a ratio of 7:3.
[0088] The neural network identifies through the experience learned during the training process. As long as the training set contains non-compliant target objects and corresponding labels, the network can learn their features accordingly, so as to complete the self-identification of similar situations.
[0089] In one embodiment, the preprocessing includes: video stream cleaning, image-signature association and correspondence, resolution screening, and image classification.
[0090] Through preprocessing, pictures that do not meet the training requirements in the data can be eliminated, including pictures without target objects, pictures with too low resolution, and pictures with unconventional target object positions; during the preprocessing process, it is mainly necessary to correspond the file names of the pictures and labels one by one to meet the input format of the network.
[0091] The specific static training operation of the YOLO static neural network is as follows:
[0092] S31. Obtain a second video stream to be trained that has undergone preprocessing.
[0093] S32. Screen several second feature images containing abnormal construction workers from the second video stream to be trained, and add feature identifiers to each of the second feature images. The feature identifiers include the category, location, and status information of the construction workers.
[0094] S33. Input several of the second feature images into the YOLO neural network one by one for iterative training to obtain a YOLO static neural network.
[0095] In one embodiment, the preprocessing includes: video stream cleaning, image-signature association, resolution screening, and image classification.
[0096] Optionally, the training of this neural network can also divide the second video stream to be trained into two video streams for training and detection.
[0097] It should be noted that after the training is completed, the two trained neural network models can be written into the background server to form a neural network group, enabling them to run at high speed on the server to provide the computing power support required for real-time monitoring.
[0098] S13. When the detection result of the dynamic detection is abnormal or the detection result of the static detection is abnormal, visually display the construction workers in the monitoring video stream and trigger a monitoring alarm.
[0099] In order to promptly notify the management personnel at the construction site to perform corresponding maintenance and management operations, when the detection result of the dynamic detection is abnormal or the detection result of the static detection is abnormal, the monitoring video stream with the abnormality can be highlighted. Optionally, a highlighted box can be added to the monitoring video stream to prompt the management personnel to perform corresponding management and response operations, and at the same time trigger and control the user terminal to issue an alarm to prompt the management personnel.
[0100] After training the two neural networks and putting them into application, it is necessary to measure their detection effects. Optionally, in one embodiment, the method may further include:
[0101] S14. Obtain several evaluation results of the YOLO static neural network or the LSTM dynamic neural network.
[0102] Optionally, the evaluation results for a period of time can be obtained, and the evaluation results are correct or incorrect.
[0103] In actual operation, after an abnormality occurs, the evaluation results of the management personnel can be received and the evaluation results for a period of time can be counted to obtain several evaluation results of the YOLO static neural network or the LSTM dynamic neural network.
[0104] S15. Calculate the result evaluation parameters from the several evaluation results through the YOLOv5 algorithm. The result evaluation parameters include: accuracy rate, no-miss detection rate, and comprehensive accuracy rate.
[0105] In one embodiment, the accuracy rate is P, the no-miss detection rate is R, and the comprehensive accuracy rate is mAP.
[0106] P = The number of correctly recognized targets / The number of all recognized targets;
[0107] R = The number of correctly recognized targets / The number of expected recognized targets;
[0108] There are different calculation methods for mAP. Generally, it is as follows: After monitoring for a period of time, for each of the n R values obtained, calculate the highest P that can be achieved with this as the lower limit, and then obtain n highest Ps; average the n Ps to obtain the AP value; after calculating the AP values of all recognition categories, average them again to obtain the final mAP value.
[0109] In one embodiment, P, R, and mAP are all calculated by the YOLOv5 algorithm.
[0110] Suppose a total of 20 times are monitored. Among them, the neural network recognizes 15, 10 of which are correct and 5 are incorrect. Then:
[0111] Accuracy P = 10 / 15 = 0.66, which means the network has a 66% certainty that each recognition result is correct and there is no false detection;
[0112] Recall rate R without missed detection = 10 / 20 = 0.5, which means the network has a 50% certainty that it can correctly recognize the expected targets without missing any.
[0113] S16. If the accuracy is greater than the preset accuracy, the recall rate without missed detection is greater than the preset recall rate without missed detection, and the comprehensive accuracy is greater than the preset comprehensive rate, determine that the YOLO static neural network or the LSTM dynamic neural network meets the detection requirements.
[0114] S17. If the accuracy is less than the preset accuracy, the recall rate without missed detection is less than the preset recall rate without missed detection, or the comprehensive accuracy is less than the preset comprehensive rate, perform network optimization on the YOLO static neural network or the LSTM dynamic neural network.
[0115] In one embodiment, if the accuracy is greater than the preset accuracy, the recall rate without missed detection is greater than the preset recall rate without missed detection, and the comprehensive accuracy is greater than the preset comprehensive rate, determine that the monitoring results of the YOLO static neural network or the LSTM dynamic neural network meet the actual detection requirements and the monitoring is accurate; otherwise, if the accuracy is less than the preset accuracy, the recall rate without missed detection is less than the preset recall rate without missed detection, or the comprehensive accuracy is less than the preset comprehensive rate, determine that the monitoring results of the YOLO static neural network or the LSTM dynamic neural network do not meet the actual detection requirements and network optimization is required for the two neural networks.
[0116] In one embodiment, the training methods for the network optimization include: adjusting the network structure, adjusting the network depth, adjusting the feature map dimension, and setting multi-scale training.
[0117] Through network optimization, the monitoring accuracy of two neural networks can be effectively improved.
[0118] Refer to Figure 4 , which shows the operation flowchart of a construction site safety monitoring method based on dynamic and static detection provided by an embodiment of the present invention.
[0119] Specifically, neural network training can be performed first, and the trained neural network can be written into the background server. Then, the construction site monitoring video stream can be collected in real time through a camera, and the trained neural network can be called by the background server to monitor and identify the monitoring video stream to determine whether there is an abnormality at the construction site. If there is an abnormality at the construction site, the abnormal monitoring video stream can be visually displayed and an alarm can be sent to notify the management personnel to perform corresponding operations.
[0120] In this embodiment, the embodiment of the present invention provides a construction site safety monitoring method based on dynamic and static detection, and its beneficial effect is that: the present invention can perform real-time monitoring on the construction site, and use the real-time monitoring video to perform dynamic safety detection and static safety monitoring on the construction personnel to determine whether there is an abnormality of the construction personnel at the construction site, so as to avoid the situation of delaying the project due to construction accidents of the construction personnel, ensure the safety of the construction personnel, and improve the reliability of the project.
[0121] The embodiment of the present invention also provides a construction site safety monitoring system based on dynamic and static detection. Refer to Figure 5 , which shows the structural schematic diagram of a construction site safety monitoring system based on dynamic and static detection provided by an embodiment of the present invention.
[0122] Among them, by way of example, the construction site safety monitoring system based on dynamic and static detection may include:
[0123] The on-site perception module 501 is used to call the cameras set in different areas of the construction site to collect the monitoring video stream in real time and receive the monitoring instructions input by the user;
[0124] The detection and recognition module 502 is used to control a preset neural network group based on the monitoring instructions to perform dynamic detection or static detection on the construction personnel in the monitoring video stream. The preset neural network group includes: the YOLO static neural network for static detection and the LSTM dynamic neural network for dynamic detection;
[0125] The visualization and alarm module 503 is used to visually display the construction personnel in the monitoring video stream and trigger a monitoring alarm when the detection result of the dynamic detection is abnormal or the detection result of the static detection is abnormal.
[0126] Optionally, the dynamic detection includes:
[0127] Identify the human bone structure of each frame of video image in the monitored video stream, and extract bone marker points from the human bone structure to obtain a set of bone dot matrices;
[0128] Input the set of bone dot matrices into the LSTM dynamic neural network to obtain the continuous action sequences between different humans in the monitored video stream;
[0129] Determine the dynamic actions of the construction workers based on the continuous action sequences;
[0130] When the dynamic actions are different from the preset set of actions, the detection result of the dynamic detection is abnormal.
[0131] Optionally, the static detection includes:
[0132] Divide the monitored video stream into several video data blocks;
[0133] Based on a preset recognition frame, extract the person information of each frame of video image in the video data block, and calculate the person change rate of the construction worker or object within the time duration corresponding to the video data block using the person information;
[0134] When the person change rate is greater than the preset change rate, the detection result of the static detection is abnormal.
[0135] Optionally, the dynamic training operation of the LSTM dynamic neural network is specifically as follows:
[0136] Obtain a first video stream to be trained that has been preprocessed;
[0137] Select several first feature images containing abnormal construction workers from the first video stream to be trained, and add feature identifiers to each of the first feature images. The feature identifiers include the category, location, and status information of the construction worker;
[0138] Through the human bone key point detection technology, convert the construction workers in each of the first feature images into point-line models to obtain several human model images, and arrange the several human model images into a human model sequence;
[0139] Input the human model sequence into the LSTM network for iterative training to obtain the LSTM dynamic neural network.
[0140] Optionally, the static training operation of the YOLO static neural network is specifically as follows:
[0141] Obtain a second video stream to be trained that has been preprocessed;
[0142] Screen a number of second feature images containing abnormal construction workers from the second video stream to be trained, and add feature identifiers to each of the second feature images. The feature identifiers include the category, location, and status information of the construction workers;
[0143] Input the number of second feature images into the YOLO neural network one by one for iterative training to obtain a YOLO static neural network.
[0144] Optionally, the system further includes:
[0145] An acquisition module for acquiring a number of evaluation results of the YOLO static neural network or the LSTM dynamic neural network;
[0146] A calculation module for calculating the number of evaluation results through the YOLOv5 algorithm to obtain result evaluation parameters, where the result evaluation parameters include: accuracy, no-miss detection rate, and comprehensive accuracy;
[0147] A compliance module for determining that the YOLO static neural network or the LSTM dynamic neural network meets the detection requirements when the accuracy is greater than the preset accuracy, the no-miss detection rate is greater than the preset no-miss detection rate, and the comprehensive accuracy is greater than the preset comprehensive rate;
[0148] An optimization module for performing network optimization on the YOLO static neural network or the LSTM dynamic neural network when the accuracy is less than the preset accuracy, the no-miss detection rate is less than the preset no-miss detection rate, or the comprehensive accuracy is less than the preset comprehensive rate.
[0149] Optionally, the training methods for the network optimization include: adjusting the network structure, adjusting the network depth, adjusting the feature map dimension, and setting multi-scale training optionally.
[0150] Those skilled in the art can clearly understand that for the convenience of description and simplicity, the specific working process of the system described above can refer to the corresponding process in the foregoing method embodiment, and will not be elaborated here.
[0151] Furthermore, an embodiment of the present application further provides an electronic device, including: a memory, a processor, and a computer program stored on the memory and executable on the processor. When the processor executes the program, it implements the method for monitoring construction site safety based on static and dynamic detection as described in the above embodiment.
[0152] Furthermore, an embodiment of the present application further provides a computer-readable storage medium, where the computer-readable storage medium stores computer-executable instructions, and the computer-executable instructions are used to make a computer execute the method for monitoring construction site safety based on static and dynamic detection as described in the above embodiment.
[0153] The above are the preferred embodiments of the present invention. It should be noted that for those of ordinary skill in the art, without departing from the principle of the present invention, several improvements and refinements can be made, and these improvements and refinements are also regarded as the protection scope of the present invention.
Claims
1. A construction site safety monitoring method based on dynamic and static detection, characterized in that, The method includes: Invoking cameras set in different areas of the construction site to collect monitoring video streams in real time and receiving monitoring instructions input by the user; Based on the monitoring instructions, controlling a preset neural network group to perform dynamic detection or static detection on the construction workers in the monitoring video stream. The preset neural network group includes: a YOLO static neural network for static detection and an LSTM dynamic neural network for dynamic detection; When the detection result of the dynamic detection is abnormal or the detection result of the static detection is abnormal, visually display the construction workers in the monitoring video stream and trigger a monitoring alarm; The dynamic detection includes: Identifying the human skeletal structure of each frame of video image in the monitoring video stream and extracting skeletal marker points from the human skeletal structure to obtain a skeletal dot matrix set; Inputting the skeletal dot matrix set into the LSTM dynamic neural network to obtain a continuous action sequence between different humans in the monitoring video stream; Determining the dynamic actions of the construction workers based on the continuous action sequence; When the dynamic actions are different from a preset action set, the detection result of the dynamic detection is abnormal; The static detection includes: Dividing the monitoring video stream into several video data blocks; Based on a preset recognition frame, extracting the person information of each frame of video image in the video data block and calculating the person change rate of the construction workers or objects within the duration corresponding to the video data block using the person information; When the person change rate is greater than a preset change rate, the detection result of the static detection is abnormal.
2. The construction site safety monitoring method based on dynamic and static detection according to claim 1, characterized in that, The specific dynamic training operation of the LSTM dynamic neural network is: Obtaining a first video stream to be trained that has been preprocessed; Selecting several first feature images containing abnormal construction workers from the first video stream to be trained and adding feature identifiers to each of the first feature images. The feature identifiers include the category, position, and status information of the construction workers; Through human skeletal key point detection technology, converting the construction workers in each of the first feature images into a point-line model to obtain several human model images, and arranging the several human model images into a human model sequence; Inputting the human model sequence into the LSTM network for iterative training to obtain the LSTM dynamic neural network.
3. The construction site safety monitoring method based on dynamic and static detection according to claim 1, characterized in that, The specific static training operation of the YOLO static neural network is: Obtaining a second video stream to be trained that has been preprocessed; Selecting several second feature images containing abnormal construction workers from the second video stream to be trained and adding feature identifiers to each of the second feature images. The feature identifiers include the category, position, and status information of the construction workers; Inputting the several second feature images into the YOLO neural network one by one for iterative training to obtain the YOLO static neural network.
4. The construction site safety monitoring method based on dynamic and static detection according to claim 1, characterized in that, The method further includes: Obtaining several evaluation results of the YOLO static neural network or the LSTM dynamic neural network; Calculating the several evaluation results through the YOLOv5 algorithm to obtain result evaluation parameters, where the result evaluation parameters include: accuracy rate, no-miss detection rate, and comprehensive accuracy rate; When the accuracy rate is greater than the preset accuracy rate, the non-missing detection rate is greater than the preset non-missing detection rate, and the comprehensive accuracy rate is greater than the preset comprehensive rate, it is determined that the YOLO static neural network or the LSTM dynamic neural network meets the detection requirements; When the accuracy rate is less than the preset accuracy rate, the non-missing detection rate is less than the preset non-missing detection rate, or the comprehensive accuracy rate is less than the preset comprehensive rate, network optimization is performed on the YOLO static neural network or the LSTM dynamic neural network.
5. The construction site safety monitoring method based on dynamic and static detection according to claim 4, characterized in that, The training methods for the network optimization include: adjusting the network structure, adjusting the network depth, adjusting the feature map dimension, and setting multi-scale training.
6. A construction site safety monitoring system based on dynamic and static detection, characterized in that, The system includes: A field perception module, which is used to call cameras set in different areas of the construction site to collect monitoring video streams in real time and receive monitoring instructions input by the user; A detection and recognition module, which is used to control a preset neural network group based on the monitoring instructions to perform dynamic detection or static detection on construction workers in the monitoring video stream. The preset neural network group includes: a YOLO static neural network for static detection and an LSTM dynamic neural network for dynamic detection; A visualization and alarm module, which is used to visually display construction workers in the monitoring video stream and trigger a monitoring alarm when the detection result of the dynamic detection is abnormal or the detection result of the static detection is abnormal; The dynamic detection includes: Identifying the human bone structure of each frame of video image in the monitoring video stream, and extracting bone marker points from the human bone structure to obtain a set of bone dot matrices; Inputting the set of bone dot matrices into the LSTM dynamic neural network to obtain a continuous action sequence between different humans in the monitoring video stream; Determining the dynamic actions of construction workers based on the continuous action sequence; When the dynamic actions are different from the preset set of actions, the detection result of the dynamic detection is abnormal; The static detection includes: Dividing the monitoring video stream into several video data blocks; Extracting the person information of each frame of video image in the video data block based on a preset recognition frame, and calculating the person change rate of construction workers or objects within the time period corresponding to the video data block using the person information; When the person change rate is greater than the preset change rate, the detection result of the static detection is abnormal.
7. An electronic device, comprising: A memory, a processor, and a computer program stored on the memory and executable on the processor, wherein the processor implements the method for monitoring construction site safety based on dynamic and static detection according to any one of claims 1-5 when executing the program.
8. A computer-readable storage medium, characterized in that, The computer-readable storage medium stores computer-executable instructions, and the computer-executable instructions are used to cause a computer to execute the method for monitoring construction site safety based on dynamic and static detection according to any one of claims 1-5.
Citation Information
Patent Citations
Two-dimensional-inclination-criterion-based safety helmet for detecting abnormality of head pose
CN106530614A
Human body behavior detection method, monitoring equipment, electronic equipment and medium
CN113869127A
AI-based construction site abnormal actor identification and alarm system
CN211293956U