Construction site smoke and fire detection method based on images
By using the YOLOv11 network and the upsampling convolution module EUCB, a firework detection model is built, which solves the problem of smoke and flame recognition in complex environments of the construction site, and efficient and rapid fire detection is achieved, which improves detection accuracy.
Patent Information
- Application Number
- CN202510370576.5
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-03-27
- Publication Date
- 2025-06-27
- Estimated Expiration
- Not applicable · inactive patent
AI Technical Summary
Fire hazards at construction sites are high, and existing fireworks detection technology is difficult to achieve efficient and rapid smoke and flame recognition in complex construction environments.
The fire image data set is trained using the YOLOv11 network to build a firework detection model, improve detection accuracy through upsampling convolution module EUCB, and accelerate model convergence using the loss function L.
It realizes efficient and rapid smoke and flame recognition in construction site environments, and improves the average detection accuracy.
Smart Images

Figure CN120219850A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of smoke and fire detection in construction sites. Background Art
[0002] With the acceleration of the urbanization process, the number and scale of construction sites have been increasing year by year, and the potential fire hazards have also increased accordingly. Combustible materials, welding operations, and irregular electricity use are often present at construction sites, significantly raising the fire risk. In recent years, deep learning, especially convolutional neural networks (CNNs), has made remarkable progress in image processing and computer vision tasks. Object detection algorithms such as YOLO (You Only Look Once) can achieve real-time and accurate detection of multiple targets. Applying deep learning technology to smoke and fire detection can greatly improve the detection speed and accuracy, adapt to the complex construction site environment, and achieve automated and intelligent fire warnings. Summary of the Invention
[0003] Based on this, in view of the deficiencies of the prior art, the present invention proposes an image-based method for smoke and fire detection in construction sites.
[0004] To achieve the above object, the solution of the present invention is: (1) Use the YOLOv11 network to train a fire image dataset to obtain a smoke and fire detection model; (2) Collect real-time images and input them into the smoke and fire detection model to output the detection results.
[0005] Further, the step (1) includes the following steps:
[0006] Step 11: Construct a YOLOv11 network model;
[0007] The YOLOv11 model includes a backbone network, a neck network, and a detection head network; the backbone network includes a first convolution module, a second convolution module, a first C3k2 module, a third convolution module, a second C3k2 module, a fourth convolution module, a third C3k2 module, a fifth convolution module, a fourth C3k2 module, an SPPF module, and a C2PSA module connected in sequence; the neck network includes a first EUCB module, a first splicing module, a fifth C3k2 module, a second EUCB module, a second splicing module, a sixth C3k2 module, a sixth convolution module, a third splicing module, a seventh C3k2 module, a seventh convolution module, a fourth splicing module, and an eighth C3k2 module connected in sequence; the first splicing module splices the features output by the third C3k2 module and the features output by the first EUCB module and inputs them into the fifth C3k2 module, the second splicing module splices the features output by the second C3k2 module and the features output by the second EUCB module and inputs them into the sixth C3k2 module, the third splicing module splices the features output by the fifth C3k2 module and the features output by the second EUCB module and inputs them into the seventh C3k2 module, the fourth splicing module splices the features output by the C2PSA module and the features output by the seventh convolution module and inputs them into the eighth C3k2 module; the features output by the sixth C3k2 module, the seventh C3k2 module, and the eighth C3k2 module are input into the detection head network; the detection head network contains three detection heads; each detection head contains a location regression branch and a classification branch, the input of the first detection head is the sixth C3k2 module, and the output location regression branch includes an eighth convolution module, a ninth convolution module, and a tenth convolution module connected in sequence for location prediction, and the output classification branch of the first detection head includes a first depthwise separable convolution module, an eleventh convolution module, a second depthwise separable convolution module, a twelfth convolution module, and a thirteenth convolution module connected in sequence for classification prediction; the structures of the second detection head and the third detection head are the same as that of the first detection head;
[0008] The EUCB module includes a 2x upsampling layer, a DWC layer, a batch normalization layer, an activation function, and a convolution layer connected in sequence;
[0009] Step 12: Import the orbit dataset into the model for training, use the SGD optimization method to iterate the model parameters until the loss function L converges or reaches the predetermined number of iterations to obtain the orbit recognition model.
[0010] Furthermore, the loss function L includes a bounding box regression loss function L Powerful-IoU , a classification loss function L cls , and the formula is as follows:
[0011] L = α·L Powerful-IoU + β·L cls
[0012] Among them, α and β are weight coefficients, and the formula of the Powerful-IoU loss function is as follows:
[0013]
[0014] Among them, λ is a hyperparameter, and L pIoU represents the pIoU loss,
[0015]
[0016] Among them, IoU inner is the intersection over union of the predicted bounding box and the ground truth bounding box, and p is the quality parameter of the predicted bounding box.
[0017]
[0018] Among them, represents the abscissa, ordinate, width, and height of the center point of the ground truth bounding box; (x, y, w, h) represents the abscissa, ordinate, width, and height of the center point of the predicted bounding box, and r is the scale factor;
[0019] The formula of the L cls loss function is as follows:
[0020]
[0021] Among them, 1 obj represents whether the target is included, and p(c) represents the probability that the target belongs to c. represents whether the ground truth bounding box label belongs to class c;
[0022] A method for detecting smoke and fire in a construction site based on images proposed by the present invention aims to achieve efficient and rapid smoke and flame recognition. By using the upsampling convolutional module EUCB, the detection accuracy is improved. Through the loss function L, the model can converge faster and better, and the average precision is improved. Brief Description of the Drawings
[0023] Figure 1 is the flow chart of the detection process of the present invention.
[0024] Figure 2 is the structural diagram of the network of the present invention.
[0025] Figure 3 is the structural diagram of the EUCB module of the present invention.
[0026] Figure 4 is the schematic diagram of the smoke and fire target detection of the present invention. Detailed Embodiments
[0027] The following clearly and completely describes the technical solutions in the embodiments of the present invention with reference to the accompanying drawings in the embodiments of the present invention. Obviously, the described embodiments are only a part of the embodiments of the present invention, rather than all embodiments. All other embodiments obtained by those of ordinary skill in the art based on the embodiments of the present invention without creative efforts belong to the scope of protection of the present invention.
[0028] For the convenience of understanding by those skilled in the art, an embodiment is now used to illustrate a method for detecting smoke and fire in a construction site based on images, which includes the following steps: (1) Use the YOLOv11 network to train the fire image data set to obtain a smoke and fire detection model; (2) Collect real-time images and input them into the smoke and fire detection model to output the detection results.
[0029] Specifically, the fire data set in step (1) consists of pictures and labels. The categories include flame, smoke, and others, and the corresponding labels are fire, smoke, and other. This embodiment includes 20,000 pictures and corresponding labels, and the division ratio of the training set and the test set is 8:2.
[0030] Specifically, the smoke and fire detection model based on the yolov11 network constructed in step (11) includes a backbone network, a neck network, and a detection head network, and the structure is as Figure 2 shown, and the structure of the EUCB module is as Figure 3 shown.
[0031] Specifically, in the network training method of step (12), the weight coefficients α and β of the loss function are respectively set to 0.6 and 0.4; the scale factor r is set to 0.5; the number of iterations is 500, and the learning rate is 0.001.
[0032] Specifically, the detection result of the embodiment is as Figure 4 shown.
Claims
1. A method for detecting smoke and fire at a construction site based on an image, characterized in that: include: Step 1: Use the YOLOv11 network to train the fire image dataset and obtain a smoke and fire detection model; Step 2: Collect real-time images, input them into the fireworks detection model, and output the fireworks detection results.
2. The method according to claim 1, characterized in that: Step 1: Establishing the fireworks detection model includes: Step 11, build the YOLOv11 network model; The YOLOv11 model includes a backbone network, a neck network and a detection head network; the backbone network includes a first convolution module, a second convolution module, a first C3k2 module, a third convolution module, a second C3k2 module, a fourth convolution module, a third C3k2 module, a fifth convolution module, a fourth C3k2 module, an SPPF module and a C2PSA module connected in sequence; the neck network includes a first EUCB module, a first splicing module, a fifth C3k2 module, a second EUCB module, a second splicing module, a sixth C3k2 module, a sixth convolution module, a third splicing module, a seventh C3k2 module, a seventh convolution module, a fourth splicing module and an eighth C3k2 module connected in sequence; the first splicing module splices the features output by the third C3k2 module with the features output by the first EUCB module and inputs them into the fifth C3k2 module, and the second splicing module splices the features output by the second C3k2 module with the features output by the second EUCB module and inputs them into the sixth C3k2 module block, the third splicing module splices the features output by the fifth C3k2 module with the features output by the second EUCB module and inputs them into the seventh C3k2 module, the fourth splicing module splices the features output by the C2PSA module with the features output by the seventh convolution module and inputs them into the eighth C3k2 module; the features output by the sixth C3k2 module, the seventh C3k2 module and the eighth C3k2 module are input into the detection head network; the detection head network comprises three detection heads; the detection head comprises a position regression branch and a classification branch, the input of the first detection head is the sixth C3k2 module, the position regression branch of the first detection head comprises the eighth convolution module, the ninth convolution module and the tenth convolution module connected in sequence for position prediction, the classification branch of the first detection head comprises the first depthwise separable convolution module, the eleventh convolution module, the second depthwise separable convolution module, the twelfth convolution module and the thirteenth convolution module connected in sequence for classification prediction; the structures of the second detection head and the third detection head are the same as those of the first detection head; The EUCB module includes a 2x upsampling layer, a DWC layer, a batch normalization layer, an activation function and a convolutional layer connected in sequence; Step 12: Import the track data set into the model for training, and use the SGD optimization method to iterate the model parameters until the loss function L converges or reaches a predetermined number of iterations to obtain a track recognition model.
3. The method according to claim 2, characterized in that: The loss function L includes the bounding box regression loss function L Powerful-IoU , classification loss function L cls , the formula is as follows: L=α·L Powerful-IoU +β·L cls Among them, α, β are weight coefficients, and the Powerful-IoU loss function formula is as follows: Among them, λ is a hyperparameter, L pIoU represents the pIoU loss, Among them, IoU inner is the intersection-over-union ratio of the predicted box and the true box, p is the quality parameter of the predicted box, in, represents the horizontal coordinate, vertical coordinate, width and height of the center point of the real box; (x, y, w, h) represents the horizontal coordinate, vertical coordinate, width and height of the center point of the predicted box, r is the scale factor, The L cls The loss function formula is as follows: Among them, 1 obj Indicates whether the target is included, p(c) indicates the probability that the target belongs to c, Indicates whether the true box label belongs to category c.
Citation Information
Cited By
Complex underwater environment fish size estimation system and method based on key point detection
CN120452029A