Fire smoke detection method based on improved YOLO v5 model

By improving the YOLO v5 model, using the lcnet_100 lightweight network and DPSA and EMA attention modules, combined with the SCMPDIoU loss function, the problems of large amount of parameters and low detection accuracy of the YOLO v5 model are solved, and efficient, accurate and real-time detection of fire smoke detection is achieved.

CN120259960APending Publication Date: 2025-07-04HANGZHOU DIANZI UNIV +1
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510310637.9
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-03-17
Publication Date
2025-07-04

AI Technical Summary

Technical Problem

The existing YOLO v5 network model has large parameters, complex calculation volume and low detection accuracy, making it difficult to meet the early accurate detection and rapid response requirements of fire smoke.

Method used

The lcnet_100 lightweight convolutional network is used as the backbone network, and the dynamic channel hybrid space attention mechanism DPSA is added, and the EMA attention module is added to the Neck part, replacing the loss function as the SCMPDIoU loss function, and optimizing the model training process.

Benefits of technology

While reducing network parameters and calculation amount, the average accuracy and robustness of fire smoke detection are improved, and are suitable for real-time detection of embedded devices.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120259960A_ABST
    Figure CN120259960A_ABST
Patent Text Reader

Abstract

The invention discloses a fire smoke detection method based on an improved YOLO v5 model, and the method comprises the following steps: firstly obtaining fire smoke pictures in different scenes, and constructing a data set; then, a YOLO v5s model is used as a basic network for improvement, so that a fire smoke detection model is constructed; dividing the data set into a training set, a test set and a verification set; training and optimizing the improved YOLO v5s model by using a training set, and verifying by using a verification set to obtain an optimal fire smoke detection model; and performing fire smoke detection by using the test set through the pre-trained fire smoke detection model. According to the method, by improving the YOLO v5s model, the calculation amount and the model size are reduced, and the accuracy, the regression rate and the average detection precision are improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of computer vision, and specifically refers to a fire and smoke detection method based on an improved YOLO v5 model. Background Art

[0002] Existing fire detection technologies mainly rely on smoke sensors and temperature sensors to judge fires by monitoring the smoke concentration and temperature changes in the environment. However, these methods are easily affected by the external environment and have problems such as high false alarm rates and frequent missed alarms.

[0003] At the same time, unreasonable sensor layout may lead to detection blind spots, and the high cost of network construction and maintenance limits its large-scale application. Traditional methods are difficult to meet the requirements of early and accurate detection and rapid response to fire smoke.

[0004] In contrast, computer vision-based fire detection technology has become an important development direction for fire detection due to its non-contact, real-time, and applicability in complex scenarios. Especially the one-stage object detection network YOLO, with its high detection accuracy, fast inference ability, and end-to-end characteristics, makes it more suitable for early detection of fire smoke.

[0005] The existing YOLO v5 network model has a large number of network parameters and low detection accuracy. Therefore, this improved YOLO v5 model is proposed to detect fire smoke. Summary of the Invention

[0006] The object of the present invention is to address the problems of the YOLO v5 network model, such as a large number of network parameters, complex calculations, and low detection accuracy. A fire and smoke detection method based on an improved YOLO v5 model is proposed. This model improves the average detection accuracy while reducing network parameters and calculations.

[0007] To solve the above technical problems, the technical solution of the present invention is as follows:

[0008] A fire and smoke detection method based on an improved YOLO v5 model, comprising the following steps:

[0009] S1. Obtain fire and smoke pictures in different scenarios and construct a dataset;

[0010] S2. Improve the backbone network of the YOLO v5 network model, use the lcnet_100 lightweight convolutional network as the backbone network, and add a dynamic channel mixing spatial attention mechanism DPSA;

[0011] S3. Improve the neck part of the YOLO v5s network, and add EMA attention modules to the small and medium target detection layers in the Neck part of the YOLO v5s network respectively;

[0012] S4. Replace the loss function of the YOLO v5s network with the SCMPDIoU loss function optimized based on CIoU and MPDIoU;

[0013] S5. For the improved model, use the above dataset for training adjustment to obtain the optimal improved YOLO v5s model, and obtain the fire and smoke detection effect.

[0014] Preferably, in the process of step S1, by referring to the public dataset and collecting fire and smoke pictures in various scenarios, and using the LabelImg annotation tool, mark the positions of the flames and smoke in the pictures to generate a YOLO format dataset.

[0015] Preferably, in the process of step S1, divide the dataset into a training set, a validation set, and a test set according to 7:2:1 to ensure the rationality of the model training, tuning, and evaluation processes.

[0016] Preferably, in the process of step S2, because the CSPDarknet53 backbone network of the backbone network has a high computational complexity and a large number of parameters, the lcnet_100 lightweight convolutional network is used as the backbone network to improve the inference speed of the network while reducing the number of network parameters and the amount of computation.

[0017] Preferably, in the process of step S2, by adding a DPSA attention module to the backbone network, more attention can be concentrated on the high-information regions, thereby improving the feature extraction ability of the network.

[0018] Preferably, in the process of step S2, for the PSA attention mechanism, a channel attention mechanism is introduced, and 1*1 convolution is used to achieve dynamic interaction between channels, and the BatchNorm batch normalization layer is used to stabilize the training process of the model, thereby enhancing the gradient fluidity after the channel attention mechanism. For the spatial attention mechanism, global spatial features are extracted by cascading 7×7 convolution, and spatial attention weights are generated in combination with the Sigmoid activation function.

[0019] Preferably, in the process of step S3, adding an EMA attention module to the small and medium target detection layers in the Neck part can dynamically adjust the weights of different channels, enhance the attention to important features, and improve the feature representation ability of small and medium targets, enhance the robustness of the model and improve the positioning accuracy and detection accuracy of the network for small and medium targets.

[0020] Preferably, during step S3, an EMA attention module is separately introduced into the small and medium target detection layers in the Neck part. The EMA attention module splices two feature maps in the height direction and processes them through a shared 1×1 convolution, so as to maintain the original dimension output of the feature map. The output of the processed 1×1 convolution branch is decomposed into two vectors, and then two independent non-linear Sigmoid functions are used to fit the linear convolution result to obtain a 2D binary distribution. Then, through element-wise multiplication operations, the two channel attention maps within each group are fused, thus realizing the interaction and fusion of cross-channel features. By using a 3×3 convolution, the model can capture multi-scale features, thereby expanding the feature space. In this way, the EMA module can not only measure the importance of different channels by encoding the information between channels, but also improve the accuracy and diversity of feature expression.

[0021] Preferably, during step S4, by replacing the loss function of YOLO v5s with the improved SCMPDIoU loss function, the object detection model performs more accurately and stably when facing objects of different scales, positions, and shapes, and can effectively improve the detection performance.

[0022] Preferably, during step S4, traditional IoU calculation only focuses on the overlapping area of two bounding boxes, while the improved SCMPDIoU also introduces the distance between the centers of the two bounding boxes as an additional factor. By introducing the distance between the centers, the influence of position error on the model performance can be reduced, thereby improving the detection accuracy, especially when the spatial position of the object changes greatly.

[0023] Preferably, during step S4, the improved SCMPDIoU can effectively solve the matching problem caused by the changes in the shape and size of the object by introducing the diagonal length difference, and improve the detection performance of multi-scale objects, especially in the case of large scale differences.

[0024] Preferably, during step S5, the improved yolo v5s model is trained multiple times on the labeled YOLO format dataset to obtain the best weight file, and the test set is predicted through the weight file to analyze the prediction results.

[0025] Preferably, during step S5, comparative analysis is carried out on indicators such as the computational amount, number of parameters, accuracy, and average detection precision mAP of the improved yolo v5s model, so as to adjust and optimize the improved model.

[0026] The present invention has the following characteristics and beneficial effects:

[0027] With the above technical solutions, first, LCNet_100 is used as the backbone network to replace the traditional YOLOv5 backbone network, reducing the number of model parameters. Then, combined with the DPSA module, more accurate feature representations can be obtained to improve the average detection accuracy of the model. The EMA attention module enhances the adaptability of the fire and smoke detection model to multi-scale targets by combining feature grouping, parallel subnet structure design, and cross-space learning strategies. The improved loss function can further optimize the shape and position of the prediction boxes to make them more suitable for the detection of complex and irregular targets, especially the irregular target shapes in fire and smoke detection, thus improving the detection performance and accuracy. As can be seen from the above table, either a single improvement method or a combination of multiple improvement methods can improve the average detection accuracy of the model. Brief Description of the Drawings

[0028] In order to more clearly illustrate the technical solutions in the embodiments of the present invention or the prior art, the following will briefly introduce the drawings required for the description of the embodiments or the prior art. Obviously, the drawings in the following description are only some embodiments of the present invention. For those of ordinary skill in the art, without creative efforts, other drawings can be obtained based on these drawings.

[0029] Figure 1 It is the overall implementation flowchart of the improved YOLO v5 model of the present invention.

[0030] Figure 2 It is the network structure diagram of DPSA.

[0031] Figure 3 It is the network structure diagram of the EMA attention module.

[0032] Figure 4 It is the network structure diagram of the improved SCMPDIoU.

[0033] Figure 5 It is the overall structure diagram of the improved YOLO v5 network.

[0034] Figure 6 It is the effect diagram of the improved network model for detecting the fire and smoke dataset. Detailed Embodiments

[0035] It should be noted that, without conflict, the embodiments in the present invention and the features in the embodiments can be combined with each other.

[0036] In the description of the present invention, it should be understood that the orientation or positional relationships indicated by the terms "center", "longitudinal", "lateral", "upper", "lower", "front", "rear", "left", "right", "vertical", "horizontal", "top", "bottom", "inner", "outer", etc. are based on the orientation or positional relationships shown in the drawings. These are only for the convenience of describing the present invention and simplifying the description, rather than indicating or implying that the device or element referred to must have a specific orientation, be constructed and operated in a specific orientation. Therefore, it should not be construed as a limitation to the present invention. In addition, the terms "first", "second", etc. are only used for descriptive purposes and cannot be understood as indicating or implying relative importance or implicitly specifying the quantity of the indicated technical features. Thus, features defined with "first", "second", etc. may explicitly or implicitly include one or more of such features. In the description of the present invention, unless otherwise specified, the meaning of "a plurality" is two or more.

[0037] In the description of the present invention, it should be noted that unless otherwise clearly specified and defined, the terms "installed", "connected", "coupled" should be understood in a broad sense. For example, it can be a fixed connection, a detachable connection, or an integral connection; it can be a mechanical connection or an electrical connection; it can be directly connected or indirectly connected through an intermediate medium, and it can be the communication inside two elements. For those of ordinary skill in the art, the specific meanings of the above terms in the present invention can be understood through specific circumstances.

[0038] The present invention provides a fire and smoke detection method based on an improved YOLO v5 model, as Figure 1 shown, including the following steps:

[0039] S1. Obtain fire and smoke pictures in different scenarios and construct a dataset;

[0040] In this embodiment, there are a total of 6750 pictures in the constructed dataset. The targets of fire and smoke in the pictures are marked through a labeling file by LabelImg to generate a labeling file of txt type.

[0041] S2. Construct a fire and smoke detection model. In this embodiment, YOLO v5s is selected as the basic network for improvement because, compared with other versions such as YOLO v5m and YOLO v5l, the YOLO v5s network model has fewer parameters, is suitable for running on embedded devices with less resources, and the inference speed of YOLO v5s is relatively fast, which is suitable for tasks with high real-time requirements such as fire and smoke detection.

[0042] Among them, the YOLO v5s model includes a backbone network, a neck network, and a detection head network.

[0043] Specifically, as Figure 5 shown, replace CSPDarknet53 in the backbone network with the lcnet_100 lightweight convolutional network as the basic module to construct a lightweight network structure. Use the H-Swish activation function to improve performance with almost no increase in inference time. Add an SE module and large convolutional kernels at the end of the network to balance accuracy and speed. Add a large-dimensional 1x1 convolution after global average pooling to enhance the feature combination ability. Therefore, using the lcnet_100 network as the backbone feature extraction network of the improved YOLOv5 has a lighter network model, higher accuracy, and faster inference speed.

[0044] Furthermore, add a Dynamic Channel Mixing Spatial Attention mechanism (DPSA) module to the last layer of the backbone network, which can enhance the feature extraction ability.

[0045] Specifically, as Figure 2 shown, the DPSA module includes two branches, namely the dynamic channel branch and the spatial attention branch. The two use a weighted fusion mechanism to form the final DPSA attention module. The expression of the DPSA module is as follows:

[0046] DPSA(x) = Z ch +(Z sp +1)*PSA

[0047] where Z ch is the dynamic channel branch, Z sp is the spatial attention branch, and PSA is the attention output of the PSA module.

[0048] The calculation formula of the channel branch Z ch is the output result of the dynamic channel branch, Z sp is the output result of the spatial attention branch. The output of the dynamic channel branch is obtained by 1×1 convolution calculation plus BatchNorm to get Z ch , and the spatial attention branch is obtained by 7×7 convolution calculation plus BatchNorm and Sigmoid activation function to get Z sp .

[0049] It should be noted that DPSA introduces a combination of dynamic channels and spatial attention on the basis of PSA. The dynamic channel mechanism is realized by using grouped 1x1 convolution for dynamic interaction between channels, enhancing the model's adaptability to feature interaction between different channels and flexibly selecting the channels that need to be concerned. The spatial attention weights are generated by using a convolutional module and Sigmoid activation to control the activation degree of different spatial positions to highlight the key areas.

[0050] Furthermore, an EMA attention module is added to the small and medium target detection layers of the neck network of the YOLO v5s model respectively.

[0051] The EMA attention module divides the input feature map \(X\in\mathbb{R}\) C×H×W into several sub-features to learn different semantic features. Through these attention weight descriptors, the model can strengthen the key features related to the task.

[0052] For \(X\in\mathbb{R}\) C×H×W , where \(X\) is the input feature map, \(C\) is the number of channels, \(H\) is the height of the feature map, and \(W\) is the width of the feature map.

[0053] As shown in the appendix Figure 3 , the EMA mechanism uses three parallel paths to extract the attention weight descriptors of the feature map. Two of the parallel paths are in the 1×1 branch, and the third path is the 3×3 branch. To capture the dependencies between all channels and reduce the computational budget, each 1×1 branch uses two 1D global average pooling operations to perform 1D global average pooling in the two-dimensional space (height and width), while the 3×3 branch uses a single 3×3 convolutional kernel to capture the multi-scale feature representation and adds a 2D global average pooling operation after the 3×3 convolution.

[0054] By introducing a cross-space information aggregation method, the outputs of the 1×1 branch and the 3×3 branch are combined. First, two 1D global poolings are used to encode the global spatial information of the output of the 1×1 branch, and the output of the 3×3 branch is converted to the corresponding dimension. The outputs of the two are fused through a matrix dot product operation to generate the first spatial attention map.

[0055] Similarly, 2D global pooling is used to encode the global spatial information of the 3×3 branch, while the 1×1 branch is converted to the required shape before the joint activation of the channel features to obtain the second spatial attention map. Finally, the two spatial attention maps are weighted and summed and passed through the Sigmoid function to obtain the final feature map.

[0056] Among them, the formula for the 2D global pooling operation is:

[0057]

[0058] In addition, the loss function of the improved YOLO v5 network is replaced with the improved S-CMPDIoU loss function.

[0059] As shown in the appendix Figure 4As shown, S-CMPDIoU is based on calculating the IoU between the predicted bounding box and the ground truth bounding box. By calculating the diagonal lengths of the two bounding frames and the distance between the center points of the two bounding boxes, it subtracts the squared term of the normalized center point distance and the squared term of the normalized diagonal length difference from the IoU. The improved S-CMPDIoU loss function provides a more comprehensive measure of the similarity of bounding boxes by combining the degree of overlap, the center point distance, and the scale difference.

[0060] Specifically, the coordinates of the predicted bounding box b1 are (b1x1, b1y1), (b1x2, b1y2), and the coordinates of the ground truth bounding box are (b2x1, b1y1) and (b2x2, b2y2).

[0061] Then the widths and heights of the predicted bounding box b1 and the ground truth bounding box b2 are w1, h1, w2, h2 respectively. The diagonal distances of the predicted bounding box b1 and the ground truth bounding box b2 are d1, d2.

[0062] w1 = b1x2 - b1x1, h1 = b1y2 - b1y1

[0063] w2 = b2x2 - b2x1, h2 = b2y2 - b2y1

[0064]

[0065] Then the distance between the center points of the predicted bounding box b1 and the ground truth bounding box b2 is calculated as:

[0066]

[0067] C represents the square of the diagonal length of the minimum bounding rectangle of the two boxes, which is used for normalization, where C x1 , C x2 , C y1 , C y2 is used to represent the boundary coordinates of the minimum bounding rectangle:

[0068] C = (C x2 - C x1 ) 2 + (C y2 - C y1 )

[0069] Finally, SCMPDIoU is expressed as:

[0070]

[0071] It can be understood that this SCMPDIoU combines the intersection over union, the center point distance, and the scale difference of the boxes, providing a more comprehensive measure of the similarity of bounding boxes. This makes SCMPDIoU particularly suitable for high-precision object detection tasks and is suitable for scenarios with slender objects or a large span of aspect ratios.

[0072] The overall network block diagram of the improved YOLO v5 network model after the above steps is shown in the appendix Figure 5 as follows.

[0073] S3. Divide the dataset into a training set, a test set, and a validation set.

[0074] All image files and label files are divided in the manner of a training set, a validation set, and a test set. There are two categories in the overall dataset, namely Fire and Smoke. In this embodiment, the dataset is divided into a training set, a validation set, and a test set in a ratio of 7:2:1.

[0075] S4. Use the training set to train and optimize the improved YOLO v5s model, and use the validation set to verify it to obtain the optimal fire and smoke detection model;

[0076] Specifically, the improved YOLO v5 model is verified without using pre-trained weight parameters. The experimental environment is configured with the operating system win11, the graphics card is 3060ti, the python version is 3.8, and the pytorch version is 1.8, and 100 epochs are trained. The evaluation metrics used are map_0.5, accuracy P, recall R, and weight size.

[0077] S5. Use the test set to perform fire and smoke detection through the pre-trained fire and smoke detection model.

[0078] In this embodiment, the test results of the fire and smoke detection model provided and the simulation results of the YOLO v5s model

[0079]

[0080]

[0081] Table 1 shows the ablation experiment comparison of the fire and smoke detection model provided by the present invention

[0082]

[0083] Table 2 shows the simulation comparison between the fire and smoke detection model provided by the present invention and the YOLO v5s model

[0084] As can be seen from Table 1 and Table 2, the present invention discloses a fire smoke detection model improved based on YOLOv5. This model constitutes the backbone network by introducing the lightweight network lcnet_100 and the DPSA attention module, adding the EMA attention module to the small and medium target detection layer of Neck, and increasing the improved SCMPDIoU loss function. On the basis of reducing the model weight by 13.14% and the total number of parameters by 18.37%, the mAP is increased by 3.3%, the Precision is increased by 4.6%, and the Recall is increased by 2.4%. This optimization not only improves the detection accuracy and generalization ability of the model, but also maintains the lightweight characteristics, providing a good balance between performance and efficiency for real-time detection and embedded deployment. Therefore, the method of this embodiment is applicable to embedded devices and can achieve real-time detection of fire smoke while reducing the occupancy of computing resources.

[0085] The above has described in detail the embodiments of the present invention in conjunction with the accompanying drawings, but the present invention is not limited to the described embodiments. For those skilled in the art, without departing from the principle and spirit of the present invention, various changes, modifications, substitutions and variations to these embodiments including components still fall within the protection scope of the present invention.

Claims

1. A fire smoke detection method based on an improved YOLO v5 model, characterized in that, It includes the following steps: S1. Obtain fire smoke pictures under different scenarios and construct a dataset; S2. Construct a fire smoke detection model. The fire smoke detection model uses the YOLO v5s model as the basic network. The YOLO v5s model includes a backbone network, a neck network, and a detection head network. Replace the CSPDarknet53 in the backbone network with the lcnet_100 lightweight convolutional network, and add a DPSA module with a dynamic channel attention and spatial attention mechanism at the last layer of the backbone network; add an EMA attention module to the small and medium target detection layers in the neck network; S3. Divide the dataset into a training set, a test set, and a validation set; S4. Use the training set to train and optimize the improved YOLO v5s model, and use the validation set to verify it to obtain the optimal fire smoke detection model; S5. Use the test set to perform fire smoke detection through the pre-trained fire smoke detection model.

2. The fire smoke detection method according to claim 1, characterized in that, The lcnet_100 lightweight convolutional network uses depthwise separable convolution as the basic module to construct a lightweight network structure, uses the H-Swish activation function, adds an SE module and a 5×5 convolutional kernel at the tail, and adds a large-dimension 1x1 convolution after the global average pooling of the basic module.

3. The fire smoke detection method according to claim 1, wherein The DPSA module is realized by adding a weighted fusion of dynamic channel attention and spatial attention branches on the basis of PSA. Among them, the dynamic channel branch uses a 1×1 convolution and a batch normalization layer to realize dynamic interaction between channels, and the spatial attention branch uses a 7×7 convolution and a batch normalization layer and uses the Sigmoid activation function to generate spatial weights. The expression is as follows: DPSA(x) = Z ch +(Z sp +1)*PSA Among them, Z ch is the dynamic channel branch, Z sp is the spatial attention branch, and PSA is the attention output of the PSA module.

4. The fire smoke detection method according to claim 1, characterized in that, The EMA attention module divides the input feature map into several sub-feature groups and extracts attention weight descriptors using three parallel paths. The three parallel paths include two 1×1 convolution branches and one 3×3 convolution branch, and combines the outputs of the three parallel paths by combining a cross-space information aggregation method.

5. The fire smoke detection method according to claim 4, wherein Each 1×1 convolution branch uses two 1D global average pooling operations. The 3×3 convolution branch uses a single 3×3 convolutional kernel to capture multi-scale feature representations, and adds a 2D global average pooling operation after the 3×3 convolution.

6. The fire smoke detection method according to claim 5, characterized in that, The combination method of the outputs of the three parallel paths is as follows: Use two 1D global average poolings to encode the global spatial information of the output of the 1×1 branch, and convert the output of the 3×3 branch into the corresponding dimension, and fuse the outputs of the two through a matrix dot product operation to generate the first spatial attention map; Use 2D global average pooling to encode the global spatial information of the 3×3 branch, while the 1×1 branch is converted into the required shape before the joint activation of the channel features to obtain the second spatial attention map; Finally, obtain the final feature map by weighted summing the two spatial attention maps and passing through the Sigmoid function.

7. The fire smoke detection method according to claim 6, characterized in that, The formula for the 2D global pooling operation is: Among them, X c is the input feature map, H and W are the height and width of the feature map respectively, and X c (i, j) is the value at the i-th row and j-th column in the feature map, and Z c is the output after 2D global pooling.

8. The fire smoke detection method according to claim 1, characterized in that, Replace the loss function of the fire smoke detection model YOLOv5s model with the scale-aware S-CMPDIoU loss function.

9. The fire smoke detection method according to claim 8, wherein The S-CMPDIoU loss function is calculated as follows by combining the Intersection over Union (IoU) of bounding boxes, the normalized term of the distance between the center points, and the normalized term of the difference in diagonal lengths: where d is the distance between the center points of the predicted bounding box and the ground truth bounding box, d1 and d2 are the diagonal lengths of the two bounding boxes respectively, and C is the square of the diagonal length of the minimum bounding rectangle for normalization.

10. The fire smoke detection method according to claim 1, characterized in that, The dataset is divided into a training set, a validation set, and a test set in a ratio of 7:2:1, and the bounding boxes of fire and smoke targets are annotated using the LabelImg tool to generate annotation files in YOLO format.

11. The fire smoke detection method according to claim 1, wherein The method is applicable to embedded devices and can achieve real-time detection of fire and smoke while reducing the consumption of computing resources.