Construction site fire behavior detection and judgment method and system based on deep learning
By using a cross-layer fusion backbone network and a path aggregation network, combined with the optimal transmission allocation strategy and data enhancement technology, the problem of small and low target detection rate in single-stage object detection methods is solved, and high accuracy detection and real-time monitoring of fire behavior at the construction site are achieved.
Patent Information
- Application Number
- CN202211674265.0
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2022-12-26
- Publication Date
- 2025-07-25
- Estimated Expiration
- 2042-12-26
AI Technical Summary
The existing single-stage target detection method has a low detection rate for small targets, especially when the computing power of edge equipment is low, it is difficult to meet the actual needs of fire behavior detection at construction sites.
Two object detection networks are built using the cross-layer fusion backbone network CSPDarknet and the path aggregation network PANet, which are used to detect conventional sizes and small targets respectively, and are trained through the optimal transmission allocation strategy OTA, combining stochastic gradient descent method and data augmentation technology to improve the detection accuracy of small targets.
The detection rate of small targets is improved, high-accuracy identification and real-time monitoring of targets of different sizes is achieved, and 24-hour uninterrupted abnormal event alarm prompts are met. It is suitable for end-to-end construction site fire behavior detection systems.
Smart Images

Figure CN116206253B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to a method and system for detecting and judging the behavior of using fire at a construction site based on deep learning, and belongs to the technical field of target detection. Background Art
[0002] Target detection refers to identifying and locating a specified object from an image, and is an important application in the field of computer vision. The application of target detection methods is extremely extensive, especially playing an indispensable role in fields such as industrial monitoring security, autonomous driving, robot vision, and military detection. In recent years, with the development of deep learning algorithms, target detection methods have also made rapid progress, and have provided new solutions for the detection of the behavior of using fire at construction sites in the industrial field.
[0003] With the rapid development of deep learning technology, most current target detection algorithms are based on deep learning. The target detection methods based on deep learning can be divided into two categories: one is the two-stage target detection method, and the specific approach is to first extract candidate regions where targets may exist, and then use a convolutional neural network (CNN) to extract features; the other is the single-stage target detection method, which is an end-to-end detection method that omits the region selection stage and only performs CNN calculations to regress to the image target category and location information.
[0004] However, the existing single-stage target detection methods still have the problem of low accuracy in detecting small targets in engineering applications. This is because the computing power of edge devices is low, the speed of transmitting the original image to the training network is slow, and it is necessary to resample and reduce the resolution of the size of the original image. For small detection targets, the reduction of resolution will bring more difficulties to the detection process. For example, for the detection of smoking by construction workers at a construction site, if the existing algorithm is used for model training, the detection rate of the obtained model will be far lower than the actual application requirements. Summary of the Invention
[0005] The technical problem to be solved by the present invention is that the detection rate of small targets by the existing single-stage target detection method is low.
[0006] To solve the above technical problem, a technical solution of the present invention is to provide a method for detecting and judging the behavior of using fire at a construction site based on deep learning, which is characterized by including the following steps:
[0007] Step 1: Collect image data of various real fire scenarios from the real operation area of the construction site operation, and perform rectangular box annotation on each image to mark targets of regular size and small targets. Among them, the targets of regular size include welding sparks, cutting sparks, open flames, safety helmets, safety clothing, people, oxygen cylinders, acetylene cylinders, and fire extinguishers, and the small target is a cigarette;
[0008] Step 2: After preprocessing all the image data marked in Step 1, a training dataset and a test set are obtained;
[0009] Step 3: Construct and train an object detection network:
[0010] Step 301: Use the cross-layer fusion backbone network CSPDarknet and the path aggregation network PANet to build two object detection networks respectively: One object detection network is used to detect objects of regular size, defined as Object Detection Network One; the other object detection network is used to detect small objects, that is, to detect whether there are people smoking at the construction site, defined as Object Detection Network Two;
[0011] Step 302: Adopt the optimal transport assignment strategy OTA as the positive and negative sample assignment strategy, input the training dataset obtained in Step 2 into Object Detection Network One and Object Detection Network Two, and use the stochastic gradient descent method to train Object Detection Network One and Object Detection Network Two;
[0012] During training, the loss functions of Object Detection Network One and Object Detection Network Two are as follows:
[0013]
[0014] In the formula, L cls is the classification loss, L reg is the localization loss, L obj is the object confidence, λ is the balance coefficient of the localization loss, N pos is the anchor point classified as a positive sample;
[0015] Step 303: Use the test set to test the trained Object Detection Network One and Object Detection Network Two. If the requirements are met, the training of Object Detection Network One and Object Detection Network Two is completed. If the requirements are not met, return to Step 302 to retrain Object Detection Network One and Object Detection Network Two;
[0016] Step 4: After obtaining the real-time image data, use the trained Object Detection Network One and Object Detection Network Two to detect the objects of regular size and small objects in the real-time image data respectively, where:
[0017] The detection of objects of regular size includes the following steps:
[0018] Reduce the resolution of the real-time image to the set size and then input it into Object Detection Network One for detection to obtain the detection results of objects of regular size;
[0019] The detection of small objects includes the following steps:
[0020] Step 401: Reduce the resolution of the real-time image to a set size and then input it into the first target detection network for detection, and extract the prediction boxes of the personnel category after non-maximum suppression in the first target detection network;
[0021] Step 402: Map the prediction boxes to the resolution size of the original real-time image;
[0022] Step 403: Input the image processed in Step 402 into the second target detection network, and the second target detection network detects whether there are small targets.
[0023] Preferably, in Step 2, the preprocessing of the data includes the following steps:
[0024] Step 201: Reduce the resolution of the image data to a preset size;
[0025] Step 202: Perform data augmentation on the real construction site hot work scenario dataset composed of all image data;
[0026] Step 203: Divide the real construction site hot work scenario data after data augmentation into a training dataset and a test set according to a certain ratio.
[0027] Preferably, in Step 202, the methods used for data augmentation include: HSV color transformation, mix-up fusion augmentation, and multi-scale scaling.
[0028] Preferably, in Step 302, the classification loss L cls and the object confidence L obj both use binary cross-entropy loss, and its expression is as follows:
[0029]
[0030] In the formula, x i represents the i-th training sample in the training dataset, n represents the total number of training samples in the training dataset, y i ∈{0,1} represents the sample label of the training sample x i , f(x i ,) represents the prediction output of the first target detection network or the second target detection network based on the training sample x i , and θ is the model parameter.
[0031] Preferably, in Step 302, the localization loss L reg uses the intersection over union loss IOULoss, and its expression is as follows:
[0032] IOULoss = 1 - IOU
[0033] In the formula, It is the ratio of the intersection to the union of the ground truth box A and the predicted box B.
[0034] Another technical solution of the present invention is to provide a construction site hot work behavior detection and judgment system based on deep learning, which adopts the aforementioned construction site hot work behavior detection and judgment method, and is characterized in that it includes a hot work personnel face verification module and a hot work detection module, wherein:
[0035] The hot work personnel face verification module is used to verify the identity of the personnel who hold the construction site hot work permit and apply for hot work based on face recognition; if the face check is unsuccessful, hot work is not allowed. Only after the face check is successful can the relevant personnel receive hot work equipment for construction and activate the hot work detection module;
[0036] During the construction process, the hot work detection module uses the target detection network one and the target detection network two to detect regular-sized targets and small targets, and judges whether there are the following violations based on the detection results: detecting accidental flames, hot work personnel smoking, not wearing safety helmets and safety clothing, not equipped with fire extinguishers, personnel breaking in, and the distance between oxygen cylinders and acetylene cylinders being too close; if there are the aforementioned violations, the hot work detection module continuously conducts voice broadcasts until no violations are detected, and uploads video and image evidence for retention.
[0037] Preferably, during the construction process, the hot work personnel face verification module is triggered at regular intervals for a face check.
[0038] The present invention is applicable to detecting targets of different sizes, focuses on improving the detection rate of small targets, accurately identifies and alarms abnormal events detected, and meets 24-hour uninterrupted real-time police situation monitoring. It is an end-to-end detection system. BRIEF DESCRIPTION OF THE DRAWINGS
[0039] Figure 1 It is a flowchart of the construction site hot work behavior detection method provided by the present invention;
[0040] Figure 2 It is a training flowchart of the construction site hot work behavior detection method provided by the present invention;
[0041] Figure 3 It is an inference flowchart of the construction site hot work behavior detection method provided by the present invention. DETAILED DESCRIPTION OF THE INVENTION
[0042] The present invention will be further described below in conjunction with specific embodiments. It should be understood that these embodiments are only used to illustrate the present invention and are not intended to limit the scope of the present invention. In addition, it should be understood that after reading the content taught by the present invention, those skilled in the art can make various changes or modifications to the present invention, and these equivalent forms fall within the scope limited by the appended claims of the application equally.
[0043] like Figure 1 As shown, a method for detecting and judging fire behavior on a construction site based on deep learning disclosed in this embodiment includes the following steps:
[0044] Step 1. Use a camera to obtain image data. In this embodiment, a camera is used to collect image data of various real fire scenes from the real working area of the construction site. After clarifying the target to be detected, use the target detection data set annotation tool LabelImg to annotate each image with a rectangular frame to mark the target. In this embodiment, the targets to be detected are divided into regular-sized targets and small targets, among which regular-sized targets include welding sparks, cutting sparks, open flames, safety helmets, safety clothing, people, oxygen cylinders, acetylene cylinders, and fire extinguishers. The small target is a cigarette, that is, to detect whether there are people smoking on the construction site.
[0045] Step 2: Preprocess all image data marked in step 1:
[0046] Due to the low computing power of edge devices and the slow speed of transmitting original images, the image data size is adjusted during data preprocessing, and the resolution of the image data is adjusted from 1024×1920 to 320×640. Then, the real construction site fire scene data set composed of all image data is enhanced, and the methods used include: hsv color transformation, mix-up fusion enhancement and multi-scale scaling. Finally, the real construction site fire scene data after data enhancement is divided into a training data set (80%) and a test set (20%).
[0047] Combination Figure 2 , Step 3, build and train the target detection network:
[0048] Step 301, use the cross-layer fusion backbone network CSPDarknet + path aggregation network PANet to build two target detection networks respectively: one target detection network is used to detect targets of regular size, defined as target detection network one; the other target detection network is used to detect small targets, that is, to detect whether there are people smoking on the construction site, defined as target detection network two.
[0049] There is no essential difference in structure between the target detection network 1 and the target detection network 2. In this embodiment, the channel width and network depth of the target detection network 2 are slightly smaller than those of the target detection network 1.
[0050] Step 302: Use the optimal transport assignment strategy OTA as the positive and negative sample assignment strategy, input the training dataset obtained in Step 2 into Object Detection Network 1 and Object Detection Network 2, and use the stochastic gradient descent method to train Object Detection Network 1 and Object Detection Network 2.
[0051] During training, the loss functions of Object Detection Network 1 and Object Detection Network 2 are as follows:
[0052]
[0053] In the formula, L cls is the classification loss, L reg is the localization loss, L obj is the object confidence, λ is the balance coefficient of the localization loss, N pos is the anchor point classified as a positive sample.
[0054] The classification loss L cls and the object confidence L obj both use the binary cross-entropy loss (BCELoss), and its expression is as follows:
[0055]
[0056] In the formula, x i represents the i-th training sample in the training dataset, n represents the total number of training samples in the training dataset, y i ∈{0,1} represents the sample label of the training sample x i , f(x i ,) represents the prediction output of Object Detection Network 1 or Object Detection Network 2 based on the training sample x i , and θ is the model parameter.
[0057] The localization loss L reg uses the intersection over union loss IOULoss, and its expression is as follows:
[0058] IOULoss = 1 - IOU
[0059] In the formula, is the ratio of the intersection and union of the ground truth box A and the predicted box B. When the ground truth box A and the predicted box B completely overlap, IOU is 1.
[0060] The target detection network one and the target detection network two are trained for 100 rounds based on the training dataset. The training optimizer used is the SGD optimizer, with momentum = 0.9 and weight decay = 5e-4. The warm-up cosine annealing learning rate decay strategy is adopted, with an initial learning rate of 5e-3 and a warm-up round of 5 epochs.
[0061] Step 303: Use the test set to test the trained target detection network one and the target detection network two. If the requirements are met, the training of the target detection network one and the target detection network two is completed. If the requirements are not met, return to step 302 to retrain the target detection network one and the target detection network two.
[0062] Step 4: After obtaining the real-time image data, use the trained target detection network one and the target detection network two to detect regular-sized targets and small targets in the real-time image data respectively.
[0063] Combined with Figure 3 , the detection of regular-sized targets includes the following steps:
[0064] Adjust the resolution of the real-time image to 320×640 and then input it into the target detection network one for detection.
[0065] The detection of small targets includes the following steps:
[0066] Step 401: Adjust the resolution of the real-time image to 320×640 and then input it into the target detection network one for detection, and extract the prediction boxes of the personnel category after non-maximum suppression in the target detection network one;
[0067] Step 402: Map the position of the prediction box at the 320×640 resolution to the 1024×1920 resolution of the original real-time image;
[0068] Step 403: Input the image processed in step 402 into the target detection network two, and the target detection network two detects whether there are small targets.
[0069] The present invention also discloses a construction site hot work behavior detection and judgment system based on deep learning, including a hot work personnel face verification module and a hot work detection module, wherein:
[0070] The hot work personnel face verification module is used to verify the identity of the personnel who hold the construction site hot work permit and apply for hot work based on face recognition. If the face verification is unsuccessful, hot work is not allowed. Only after the face verification is successful can the relevant personnel receive the hot work equipment for construction and activate the hot work detection module. During the construction process, the hot work personnel face verification module is triggered every once in a while for a face check to ensure the matching of the construction personnel and the number of operators.
[0071] During the construction process, the hot work detection module uses the above-mentioned target detection network 1 and target detection network 2 to detect regular-sized targets and small targets, and judges whether there are the following violations based on the detection results: accidental flames are detected, the person working the hot work is smoking, not wearing a safety helmet and safety clothing, not equipped with a fire extinguisher, people breaking in, and the oxygen cylinder and acetylene cylinder are too close. If the above situations exist, the hot work detection module will continue to make voice broadcasts until no violations are detected, and upload video and image evidence for preservation.
Claims
1. A method for detecting and judging the behavior of using fire at a construction site based on deep learning, characterized in that, The following steps are involved: Step 1: collect image data of various real fire scenes from the real working area of the construction site, mark each image with a rectangular frame, and mark regular-sized targets and small targets. Regular-sized targets include welding sparks, cutting sparks, open flames, safety helmets, safety clothing, people, oxygen cylinders, acetylene cylinders, and fire extinguishers, and small targets are cigarettes; Step 2: After preprocessing all the image data annotated in step 1, a training data set and a test set are obtained; Step 3: Build and train the target detection network: Step 301: Use the cross-layer fusion backbone network CSPDarknet and the path aggregation network PANet to build two target detection networks respectively: one target detection network is used to detect targets of regular size, defined as target detection network 1; the other target detection network is used to detect small targets, that is, to detect whether there are people smoking on the construction site, defined as target detection network 2; Step 302: Use the optimal transmission allocation strategy OTA as the positive and negative sample allocation strategy, input the training data set obtained in step 2 into the target detection network 1 and the target detection network 2, and use the stochastic gradient descent method to train the target detection network 1 and the target detection network 2; During training, the loss functions of target detection network 1 and target detection network 2 are shown as follows: Wherein, L cls is the classification loss, L reg is the localization loss, L obj is the object confidence, λ is the balance coefficient of the localization loss, and N pos is the anchor point classified as a positive sample; Step 303: Use the test set to test the trained target detection network 1 and target detection network 2. If the requirements are met, the training of target detection network 1 and target detection network 2 is completed. If the requirements are not met, return to step 302 to retrain target detection network 1 and target detection network 2. Step 4: After obtaining the real-time image data, the trained target detection network 1 and the trained target detection network 2 are used to detect the regular-sized targets and small targets in the real-time image data respectively, wherein: Detection of regular-sized objects consists of the following steps: The resolution of the real-time image is reduced to a set size and then input into the target detection network 1 for detection to obtain the detection result of the target of regular size; The detection of small targets includes the following steps: Step 401: Reduce the resolution of the real-time image to a set size and input it into the target detection network 1 for detection, and extract the prediction box of the person category after non-maximum suppression in the target detection network 1; Step 402: Map the prediction frame to the resolution size of the original real-time image; Step 403: Input the image processed in step 402 into the target detection network 2, and the target detection network 2 detects whether there is a small target.
2. The method for detecting and judging the behavior of welding and cutting on construction sites based on deep learning according to claim 1, characterized in that, In step 2, the data preprocessing includes the following steps: Step 201, reducing the resolution of the image data to a preset size; Step 202, performing data enhancement on a real construction site fire scene data set consisting of all image data; Step 203: Divide the real construction site fire scene data after data enhancement into a training data set and a test set according to a certain ratio.
3. The method for detecting and judging the behavior of using fire at a construction site based on deep learning according to claim 2, characterized in that, In step 202, the methods used for data enhancement include: HSV color transformation, mix-up fusion enhancement and multi-scale scaling.
4. The method for detecting and judging the behavior of using fire at the construction site based on deep learning according to claim 1, characterized in that, In step 302, the classification loss L cls and the target confidence L obj both use binary cross-entropy loss, and its expression is as follows: where \(x\) i represents the \(i\)-th training sample in the training dataset, \(n\) represents the total number of training samples in the training dataset, \(y\) i \(\in \{0, 1\}\) represents the sample label of the training sample \(x\) i , \(f(x\) i , ) represents the predicted output of Object Detection Network One or Object Detection Network Two based on the training sample \(x\) i , and \(\theta\) is the model parameter.
5. The method for detecting and judging the behavior of welding and cutting operations on a construction site based on deep learning according to claim 1, characterized in that, In step 302, the localization loss L reg uses the Intersection over Union Loss (IOULoss), and its expression is shown as follows: IOULoss = 1-IOU Wherein, is the ratio of the intersection and union of the ground truth box A and the predicted box B.
6. A construction site hot work behavior detection and judgment system based on deep learning, which adopts the construction site hot work behavior detection and judgment method as described in claim 1, and is characterized in that, It includes a face verification module for hot work personnel and a hot work detection module, where: The face verification module for hot work personnel is used to verify the identity of personnel who hold a hot work permit for the construction site and apply for hot work based on face recognition; if the face check is unsuccessful, hot work is not allowed. Only after the face check is successful can the relevant personnel receive hot work equipment for construction and activate the hot work detection module; During the construction process, the hot work detection module uses the first target detection network and the second target detection network to detect regular-sized targets and small targets, and judges whether there are the following violations based on the detection results: detecting accidental flames, hot work personnel smoking, not wearing safety helmets and safety clothing, not equipped with fire extinguishers, personnel breaking in, and the distance between oxygen cylinders and acetylene cylinders being too close; if there are the aforementioned violations, the hot work detection module will continue to conduct voice broadcasts until no violations are detected, and upload video and image evidence for retention.
7. The on-site fire behavior detection and judgment system based on deep learning according to claim 6, characterized in that, During the construction process, the face verification module for hot work personnel is triggered every once in a while for a face check.
Citation Information
Patent Citations
Construction site image safety helmet detection method based on deep learning
CN110263686A
Safety helmet wearing identification method based on deep learning
CN110728223A