A lightweight flame detection network
By improving the YOLOv5 network and adopting lightweight technologies such as ShuffleNet V2 and SE modules, the problems of high computing costs and strict performance requirements of existing flame detection technologies on embedded devices are solved, and efficient flame detection on embedded devices is achieved.
Patent Information
- Application Number
- CN202310351788.X
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2023-04-03
- Publication Date
- 2025-06-13
- Estimated Expiration
- 2043-04-03
AI Technical Summary
The existing flame target detection technology has high calculation costs and strict requirements on equipment performance, making it difficult to achieve efficient flame detection on embedded devices with limited resources.
The lightweight flame detection network based on YOLOv5 is adopted, and the lightweight network structure of ShuffleNet V2 is used to reconstruct the YOLOv5 backbone network, and the attention module SE module and downsampling structure StemBlock are added to the feature extraction to improve the feature extraction capability of the network.
It realizes efficient flame detection on embedded devices, reduces the number of parameters and calculations of the network, improves detection speed and accuracy, and is suitable for deployment on various embedded devices.
Smart Images

Figure CN116258948B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to flame detection technology, and specifically relates to a lightweight flame detection network for embedded devices, which is improved based on YOLOv5. Background Art
[0002] At present, with the development of technology and the advancement of the modernization process, the frequency of fires is increasing continuously. The occurrence of fires not only destroys human property resources but also threatens human lives. If a fire can be detected in time at the early stage of its occurrence and the origin of the fire can be found and extinguished in time, most of the losses caused by fires can be reduced. Therefore, the research on fire detection has important research significance.
[0003] In traditional flame detection, currently widely used fire alarms are all based on sensor recognition, such as temperature sensors, smoke sensors, etc. When the temperature or smoke reaches a threshold value, the alarm will be triggered. This method has a certain delay and usually misses the best time for extinguishing the fire. Moreover, these sensor-based fire monitoring systems are also easily affected by other factors. For example, a smoke sensor may be affected by dust and fog, resulting in the possibility of false alarms.
[0004] With the rapid development of deep learning and computer vision, video surveillance has gradually become popular, and fire video surveillance technology has emerged accordingly. Detecting and warning fires by using video surveillance will not be affected by factors such as temperature and air flow. However, most of the existing early flame detection technologies based on video surveillance need to rely on high-performance, high-power-consuming, and expensive servers, and only by using a Graphics Processing Unit (GPU) can the requirement of real-time monitoring be achieved. For occasions that require fire monitoring, the requirements of high-performance devices are not always met, and the high cost will also hinder the popularization of video-based flame detection.
[0005] With the continuous development of embedded processors, it has become possible to apply an early flame detection system to embedded devices. Embedded devices have the advantages of low cost and low power consumption, and can be better popularized and applied to the public. Moreover, the operation and arrangement of embedded devices are more flexible, and can be arranged in some areas that are difficult to reach by traditional detection methods. For example, for the monitoring of forest fires, it is unrealistic to arrange sensors on a large scale, but if embedded devices such as cameras are arranged on drones for patrol monitoring, a larger-scale detection task can be completed. If the flame detection technology is run on embedded devices, the video surveillance-based flame detection technology can be more widely popularized, which is very meaningful for the development of fire warning technology.
[0006] Early image-based fire detection methods usually rely on handcrafted pixel features and heuristics, involving different thresholds. Although it may work well in controlled environments, these rule-based methods require continuous threshold adjustment and are relatively difficult to work in real-world environments.
[0007] In the past few years, the emergence of deep learning has allowed for automatic feature extraction. With the development of deep learning, Convolutional Neural Networks (CNNs) have brought significant improvements in flame detection. Such progress is now state-of-the-art in many video / image-related tasks. The shift from handcrafted features makes fire detection methods more robust and adaptable to real-world settings. Some CNN-based examples include [1]-[3].
[0008] In a recent research branch, the YOLO framework has been widely adopted for fire monitoring. YOLO v3 achieved the most satisfactory performance. In addition, a multi-scale output mechanism and channel attention have been combined with YOLO v3 to increase network generalization, and the regression box loss in YOLO v4 has been improved to enhance small-scale fire detection results.
[0009] The computational cost of the YOLO series is usually higher than that of other advanced lightweight networks (such as the MobileNet series and the ShuffleNet series), and it requires a high-performance platform to complete. However, due to space and cost issues, the demand for high-performance devices cannot be met in many places where flame detection technology needs to be deployed. There is a higher demand for optimizing the performance of existing flame detection technology on more cost-effective and flexible embedded devices. Summary of the Invention
[0010] The main objective of the present invention is to address the problems of the relatively high computational cost of existing flame target detection technologies and the relatively strict performance requirements for devices. A lightweight flame target detection network based on the improvement of YOLOv5 is proposed; a lightweight network is used to reconstruct the backbone network of YOLOv5; a lightweight network structure similar to ShuffleNet V2 is used to reconstruct the backbone network of the YOLOv5 network; the feature extraction ability of the lightweight backbone network is enhanced by adding an attention module SE module and a downsampling structure StemBlock to make up for the accuracy loss caused by network lightweighting; the network model generated through training can be deployed on various embedded platforms.
[0011] The technical solution adopted by the present invention is: a lightweight flame detection network, including:
[0012] S1, overall replacement of the backbone network;
[0013] The ReLU activation function is used in each basic unit, as shown in Equation (1):
[0014] ReLU activation function:
[0015]
[0016] S2, use StemBlock for downsampling; the StemBlock structure is as Figure 3 shown: This structure first performs a convolution operation on the input feature map with a convolution kernel size of 3x3. Its main purpose is to change the number of channels of the feature map; then the network structure is divided into two branches, and the feature map is also divided into two parts. One part of the feature map undergoes max pooling, and the other part of the feature map first undergoes a 1x1 convolution to reduce the number of channels by half, and then a 3x3 convolution with a stride of 2 is performed to achieve the second downsampling; the output results of the two branches are concatenated along the channel dimension, and finally a 1x1 convolution is performed again to restore the number of channels;
[0017] S3, add the SE module into the unit structure of ShuffleNet V2 to form the basic unit structure of this network;
[0018] Sigmoid activation function:
[0019]
[0020] S4, the Neck network structure after the backbone network mainly performs feature fusion. This structure still adopts the same structure as the YOLOv5 network, using the FPN+PAN structure with the CSP structure added;
[0021] The activation function used in this structure is the Reaky-ReLU function, where
[0022]
[0023] S5, the loss calculation of this network consists of three parts:
[0024] Classes loss, classification loss, using BCE loss, note that only the classification loss of positive samples is calculated;
[0025] Objectness loss, obj loss, still using BCE loss. Note that here obj refers to the CIoU between the target bounding box predicted by the network and the GT Box; the obj loss of all samples is calculated;
[0026] Location loss, localization loss, using CIoU loss, note that only the localization loss of positive samples is calculated;
[0027] The calculation formula is:
[0028] Loss=λ 1 L cls +λ 2 L obj +λ 3 L loc (4)
[0029] where λ 1 , λ 2 , λ 3 is the balance coefficient;
[0030] In addition, losses at different scales are balanced; this refers to using different weights for the obj losses on three prediction feature layers; in the source code, the weight used for the prediction feature layer for predicting small targets is 4.0, the weight used for the prediction feature layer for predicting medium targets is 1.0, and the weight used for the prediction feature layer for predicting large targets is 0.4;
[0031] The calculation formula is:
[0032]
[0033] S6. The construction and training of this network model are based on the pytorch framework, and the following strategies are adopted in model training:
[0034] The network model obtained through training is in the pt model format; through a model format conversion tool, the pt model is converted into a model of other formats and deployed to an embedded platform or deployed to an Android phone.
[0035] Advantages of the present invention:
[0036] The network of the present invention has:
[0037] Lightweight optimization of the network. Compared with the original YOLOv5 network, this network algorithm reduces a large number of parameters and computational amounts, has a faster detection speed, and is more suitable for deployment on various embedded devices;
[0038] The StemBlock structure and SE module added to this network model can both improve the feature extraction ability of the network and can make up for the decrease in detection accuracy caused by using a lightweight backbone network;
[0039] This network model can be deployed to other embedded platforms through format conversion methods.
[0040] In addition to the purposes, features, and advantages described above, the present invention has other purposes, features, and advantages. The following will refer to the figures for a further detailed description of the present invention. Brief Description of the Drawings
[0041] The drawings forming a part of this application are used to provide a further understanding of the present invention. The schematic embodiments of the present invention and their descriptions are used to explain the present invention and do not constitute an improper limitation to the present invention.
[0042] Figure 1 is the original YOLOv5 model of the present invention;
[0043] Figure 2 is the improved algorithm model of the present invention;
[0044] Figure 3 is the StemBlock module of the present invention;
[0045] Figure 4 are the downsampling unit and the basic unit of the present invention;
[0046] Figure 5 is the SE module of the present invention;
[0047] Figure 6 are the partial model test results of the present invention;
[0048] Figure 7 is the operation result of the model of the present invention on the mobile phone. Detailed Description of the Preferred Embodiments
[0049] In order to make the objectives, technical solutions and advantages of the present invention clearer, the present invention will be further described in detail below with reference to the accompanying drawings and embodiments. It should be understood that the specific embodiments described herein are only used to explain the present invention and are not used to limit the present invention.
[0050] The algorithm of the present invention is a new network model obtained by reconstructing the backbone network based on the YOLOv5 ( Figure 1 ), and its basic structure is as Figure 2 shown. Among them, the input at the input end is a picture of 640*640*3, the structure used in the downsampling layer is StemBlock ( Figure 3 ), and the subsequent feature extraction network is composed of the basic unit of ShuffleNet V2 with the SE module added ( Figure 4 upper part). The Neck part of the network remains consistent with the structure of YOLOv5, using the FPN+PAN structure with the CSP structure added to perform the work of feature fusion, and output three different-scale feature images at the output end to detect targets of different sizes.
[0051] The overall process design of the algorithm of the present invention is as Figure 2 shown.
[0052] (1) Overall replacement of the backbone network. Looking at the overall network structure of YOLOv5, there is room for lightweight modification. Compared with its original backbone network, the DarkNet-53 network structure, the lightweight network structure of ShuffleNet V2 can significantly reduce the number of network parameters and the amount of computation. The part from the "Focus" part after the input end to the "SPP + CSP" part in the original network ( Figure 1 ) is overall replaced with Figure 2 "StemBlock" + "SHE1" in
[0053] + "SHE2" + "SHE3" structure. Among them, "SHE1" is composed of a downsampling unit ( Figure 4 ) and 3 basic units ( Figure 6 ), "SHE2" is composed of a downsampling unit ( Figure 4 ) and 7 basic units ( Figure 6 ), and "SHE3" is composed of a downsampling unit ( Figure 4 ) and 3 basic units ( Figure 6 ).
[0054] The structural units in the "SHEx" structure layer are constructed by referring to the basic unit structure of ShuffleNetV2, which mainly consists of 3x3 depth convolution, 1x1 convolution, batch normalization BN layer, and ReLU activation function as shown in Equation (1).
[0055] The improved backbone network structure is as shown in Figure 2 , including several structural modules such as "StemBlock" + "SHE1" + "SHE2" + "SHE3". Among them, "StemBlock" is a module for downsampling the input. The three "SHE" structural modules are three convolutional layers using lightweight convolutional structures, which are composed of basic lightweight convolutional units. The specific structure is:
[0056] This convolutional unit is constructed by referring to the basic unit of ShuffleNetV2, and an attention module SE module is inserted into this structure. The structure of the basic unit is as shown in Figure 4As shown, it can be divided into a downsampling unit with a stride of 2 and a general unit with a stride of 1. In the unit with a stride of 1, first, the input feature map is split into two branches with equal number of channels in the channel dimension (channel split). The left branch performs an identity mapping, and the right branch contains three consecutive convolutions (1x1 convolution, 3x3 DW convolution, 1x1 convolution). The two branches are concatenated together through a concat operation, and then a channel shuffle is performed to ensure information exchange between the two branches. For the downsampling module with a stride of 2, the channel split operation is no longer performed. Instead, a copy of the input is directly made in each branch. Each branch has a downsampling DW convolution with a stride of 2. Finally, after concatenating them together through concat, a channel transformation is performed. The spatial size of the obtained output feature map is halved, but the number of channels is doubled.
[0057] The three "SHE" structural layers are composed of different numbers of basic units. Among them, "SHE1" is composed of 1 downsampling unit and 3 general units; "SHE2" is composed of 1 downsampling unit and 7 general units; "SHE3" is composed of 1 downsampling unit and 3 general units.
[0058] The ReLU activation function is adopted in the basic units as shown in Equation (1).
[0059] ReLU activation function:
[0060]
[0061] (2) Compared with the Focus downsampling structure in the original network and the convolution plus max-pooling downsampling structure of ShuffleNet V2, using StemBlock( Figure 3 ) for downsampling has better performance in terms of feature extraction ability and computational complexity.
[0062] (3) SENet is short for Squeeze-and-Excitation Networks, which can be called Squeeze and Excitation Networks in Chinese( Figure 5 ). The idea of the SE module it proposed is simple and easy to implement, and it can be easily loaded into the existing network model framework. SENet mainly learns the correlation between channels, filters out the attention for channels, and can achieve good results with only a slight increase in computational complexity. Adding the SE module into the unit structure of ShuffleNetV2 forms the basic unit structure of this network( Figure 4 upper part).
[0063] Among them, the structure of "global pooling layer + fully connected layer + ReLU activation function (Equation 1) + fully connected layer + Sigmoid activation function (Equation (2)) + Scale (channel weight multiplication)" added after the 3x3 depth convolution is the SE module.
[0064] Among them, the SE module is added after the 3x3 depth convolution; the structure of the SE module is as Figure 5 shown: A branch is led out from the input. In the branch, the input is passed through global pooling, a fully connected layer FC, a ReLU activation function, a fully connected layer FC, and a Sigmoid activation function to obtain a channel weight; the channel weight values calculated by the SE module are respectively multiplied (scale) with the two-dimensional matrices of the corresponding channels of the original feature map, and the obtained results are output.
[0065] Sigmoid activation function:
[0066]
[0067] (4) The Neck network structure after the backbone network mainly performs feature fusion. This structure still adopts the same structure as the YOLOv5 network, and adopts the FPN+PAN structure with the CSP structure added.
[0068] The FPN+PAN structure with the CSP structure added is:
[0069] The CSP module solves the problem of large computational complexity in inference from the perspective of network structure design. The method it adopts is to divide the feature map of the basic layer into two parts, and then merge them through a cross-stage hierarchical structure, which can ensure the accuracy while reducing the computational complexity. The specific structure of the CSP module is as Figure 2 shown: The input is divided into two branches. Branch one performs three-layer convolution-batch normalization (BN)-leakyReLU activation function, and then performs another convolution; branch two performs one convolution, and then the two branches are concatenated through a concat operation, and then passed through batch normalization-leakyReLU activation function-convolution-batch normalization-leakyReLU activation function to obtain the output.
[0070] The activation function adopted in this structure is the Reaky-ReLU function, where
[0071]
[0072] The FPN structure is a top-down feature pyramid network that uses feature maps of different levels for prediction. The specific structure is: Feature maps are extracted from different levels of the backbone network, such as Figure 2As shown, feature maps are extracted from several structural layers of "SHE1", "SHE2", and "SHE3" respectively. After upsampling the top-level feature map (i.e., the smallest feature map output by SHE3), it is concatenated with the extracted feature map of the corresponding size to obtain triple feature map outputs of different sizes.
[0073] However, only semantic information is enhanced in FPN, and no location information is transmitted. In response to this, the PAN structure is adopted. A bottom-up pyramid structure is added after FPN to supplement FPN and transmit the bottom-level location features upwards. The specific structure is as Figure 2 shown: First, the bottom layer in the FPA structure is used as the bottom output of the new pyramid PAN structure. Then, the bottom output of the new pyramid structure is downsampled through convolution and concatenated with the feature map of the same size in FPA to obtain the middle output of the new pyramid. Finally, similar steps are repeated. The middle output of the new pyramid is downsampled and concatenated with the feature map of the same size in FPN to obtain the upper output of the new pyramid. Finally, three-layer outputs of the network are obtained.
[0074] (5) The loss calculation function at the output end still uses the same method as YOLOv5. It mainly consists of three parts:
[0075] Classes loss, classification loss, which uses BCE loss. Note that only the classification loss of positive samples is calculated.
[0076] Objectness loss, obj loss, which still uses BCE loss. Note that here obj refers to the CIoU between the target bounding box predicted by the network and the GT Box. The obj loss of all samples is calculated here.
[0077] Location loss, location loss, which uses CIoU loss. Note that only the location loss of positive samples is calculated.
[0078] The calculation formula is:
[0079] Loss = λ 1 L cls + λ 2 L obj + λ 3 L loc (4)
[0081] Among them, λ 1 , λ 2 , λ 3 are balance coefficients.
[0082] In addition, losses at different scales are balanced. Here, it refers to using different weights for the obj losses on the three prediction feature layers. In the source code, the weight used for the prediction feature layer for predicting small objects is 4.0, the weight used for the prediction feature layer for predicting medium objects is 1.0, and the weight used for the prediction feature layer for predicting large objects is 0.4.
[0083] The calculation formula is:
[0084]
[0085] (6) The construction and training of this network model are based on the pytorch library functions, and the format of the network model obtained through training is the pt model. It is possible to convert the pt model into models of other formats through a model format conversion tool and deploy it to an embedded platform. For example, after converting the pt model into an onnx model, and then converting the onnx model into an ncnn model (ncnn is a high-performance neural network forward calculation framework optimized for mobile devices), the model can be deployed to an Android phone.
[0086] Multi-scale training, Multi-scale training(0.5~1.5x); Auto-adjusting anchor boxes, AutoAnchor(For training custom data), when training your own dataset, you can re-cluster to generate the Anchors template using the K-means method according to the objects in your own dataset; Warmup and Cosine LR scheduler, perform warm-up training first for 3 epochs before training, and then adopt the cosine annealing learning rate decay strategy. Some data augmentation strategies are also used: mosaic data augmentation strategy, with a ratio set to 1.0; mixup data augmentation strategy, with a ratio set to 0.1; horizontal flipping strategy, with a ratio set to 0.5;
[0087] Some hyperparameters for training are set: the initial learning rate lr is 0.01; the cosine annealing learning rate parameter is set to 0.2, that is, the minimum learning rate is (0.01 * 0.2), the batch size is set to 8; the image input size is set to [640, 640].
[0088] The format of the network model obtained through training is the pt model; through a model format conversion tool, convert the pt model into models of other formats and deploy it to an embedded platform or deploy the model to an Android phone.
[0089] The overall structure of the improved network is as Figure 2 shown. The main improvement operations are in the backbone network part, which includes the "Stem Block" downsampling layer ( Figure 3);composed of lightweight convolutional basic units ( Figure 5 ) incorporated with the attention mechanism SE module ( Figure 4 ) to form the feature extraction layer "SHE" structure layer. The part after the backbone network basically retains the structure of the original network, and this part of the structure is called the "Neck" layer, which is composed of the FPN+PAN structure to obtain multi-scale feature layers for prediction. Through the "Neck" layer, the network finally obtains three outputs of different sizes.
[0090] The present invention
[0091] (1) Reconstruct the backbone part of the YOLOv5 network using a lightweight network.
[0092] (2) Add the SE module into the units of the lightweight network ShuffleNet V2 network to form the basic unit structure of this network.
[0093] The above are only the preferred embodiments of the present invention and are not intended to limit the present invention. Any modifications, equivalent replacements, improvements, etc. made within the spirit and principle of the present invention shall be included within the protection scope of the present invention.
Claims
1. A lightweight flame detection network, characterized in that, it includes: The input of the lightweight flame detection network is a 640*640*3 picture; S1, overall replacement of the backbone network; The ReLU activation function is adopted in each basic unit, as shown in Equation (1): ReLU activation function: ReLU(x)= (1); S2, use StemBlock for downsampling; StemBlock structure: This structure first performs a convolution operation with a 3 x 3 convolution kernel on the input feature map, and its main purpose is to change the number of channels of the feature map; then the network structure is divided into two branches, and the feature map is also divided into two parts. One part of the feature map undergoes max pooling, and the other part of the feature map first undergoes a 1 x 1 convolution to reduce the number of channels by half, and then undergoes a 3 x 3 convolution with a stride of 2 to achieve the second downsampling; the output results of the two branches are concatenated along the channel dimension, and finally a 1 x 1 convolution is performed again to restore the number of channels; S3, add the SE module into the unit structure of ShuffleNet V2 to form the basic unit structure of this network; Sigmoid activation function: S(x) = (2) S4, the Neck network structure after the backbone network mainly performs feature fusion. This structure still adopts the same structure as the YOLOv5 network, and adopts the FPN+PAN structure with the CSP structure added; The activation function adopted in this structure is the Reaky-ReLU function, where LeakyReLU(x) (3) S5, the loss calculation of this network consists of three parts: Classes loss, classification loss, which adopts BCE loss and only calculates the classification loss of positive samples; Objectness loss, obj loss, which still adopts BCE loss. Here, obj refers to the CIoU between the target bounding box predicted by the network and the GT Box; the obj loss of all samples is calculated; Location loss, localization loss, which adopts CIoU loss and only calculates the localization loss of positive samples; The calculation formula is: (4) Among them λ 1 , λ 2 , λ 3 is the balance coefficient; In addition, the loss of different scales is balanced, which means that different weights are adopted for the obj losses on the three prediction feature layers; in the source code, the weight adopted for the prediction feature layer for predicting small targets is 4.0, the weight adopted for the prediction feature layer for predicting medium targets is 1.0, and the weight adopted for the prediction feature layer for predicting large targets is 0.4; The calculation formula is: (5) S6, the construction and training of this network model are based on the pytorch framework. The following strategies are adopted in model training: The network model format obtained through training is a pt model; through the model format conversion tool, the pt model is converted into models of other formats, and the models of other formats are deployed on the embedded platform or Android mobile phone.
Citation Information
Patent Citations
Lightweight unmanned ship target detection method
CN115661657A
General target detection method for adaptive attention guidance mechanism
WO2021139069A1