Infrared flame detection method and system based on deep learning
By improving the YOLOv8 model, building a BiFPN module and fusing Swin Transformer and CondConv2D, adding the SPPELAN module and a dynamic non-monotonic focus mechanism, and adding a fourth output layer, solving the problems of slow flame detection speed and low accuracy in the existing technology, and achieving fast and accurate flame detection and positioning.
Patent Information
- Application Number
- CN202510303100.X
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-03-14
- Publication Date
- 2025-06-27
AI Technical Summary
The prior art has problems in flame detection with slow detection speed, low accuracy and large model parameters, which is difficult to meet the needs of real-time application scenarios.
Improve the YOLOv8 model, build a general vision converter BiFPN module, fuse Swin Transformer and CondConv2D into the neck network, add SPPELAN module and a dynamic non-monotonic focus mechanism, and add a fourth output layer to the target detection layer to reduce model complexity and improve detection accuracy.
It realizes fast and accurate detection and positioning of infrared flame image targets, improves the feature extraction ability and real-time performance of the model, and reduces the computational complexity of the model.
Smart Images

Figure CN120219966A_ABST
Abstract
Description
Technical Field
[0001] The present invention belongs to the technical fields of deep learning and object detection, and particularly relates to an infrared flame detection method and system based on deep learning. Background Art
[0002] Human use of fire can be traced back to the era of human civilization. With the development of human civilization, in today's society, from small-scale family life to large-scale industrial technology, and even the development of national energy, all are inseparable from the use of fire. However, the development of things has two sides. While humans use fire to promote the development of industrial technology and accelerate urban construction, it also means that once a fire occurs, the resulting hazards and losses will be greater. Different from forest fires that occur in sparsely populated areas and pose little threat to human life safety, due to the dense urban buildings, urban fires occur in a relatively enclosed environment. Once a fire occurs, it will quickly heat up and ignite all combustibles, thus quickly passing from the initial stage of the fire to the development stage of the fire and directly entering the stage of intense combustion of the fire.
[0003] Flame detection and positioning technology has broad application prospects in the fields of fire monitoring, intelligent fire protection systems, and emergency rescue. By quickly and accurately detecting and positioning the flame, accurate fire source information can be provided, providing strong conditions for subsequent operations.
[0004] Two-stage object detection algorithms are represented by Faster R-CNN, and one-stage object detection algorithms are represented by the YOLO series. Although two-stage object detectors have relatively accurate position information and perform well in detecting small objects, due to the large number of components and the long overall process, the detection speed is slow and cannot meet real-time application scenarios. One-stage object detectors have the characteristics of simple structure and high computational efficiency, and at the same time have good detection accuracy. However, while improving the accuracy, the number of model parameters is increased, and usually due to the computing power limitation of embedded devices, the real-time performance is not good. If a lightweight network is added, the number of model parameters and the complexity of the calculation can be reduced, but this will result in a decrease in accuracy. Summary of the Invention
[0005] The present invention aims to solve the deficiencies of the prior art and provides the following solutions:
[0006] An infrared flame detection method based on deep learning, comprising the following steps:
[0007] Obtain infrared flame image sets in different scenarios and label the infrared flame image sets;
[0008] Perform Mosaic data augmentation on the labeled infrared flame image sets to obtain an augmented image set;
[0009] Improve the YOLOv8 model and train the improved model based on the enhanced image set to obtain a flame detection model;
[0010] Use the flame detection model to detect and locate the flame to obtain the ignition position.
[0011] Preferably, the method for improving the YOLOv8 model includes:
[0012] Construct a general vision transformer BiFPN module;
[0013] Fuse Swin Transformer and CondConv2D in the neck network of the YOLOv8 model to construct a neck feature extraction network Swin Transformer-CondConv2D module;
[0014] Add the SPPELAN module and a dynamic non-monotonic focusing mechanism;
[0015] Add a fourth output layer in the object detection layer of the YOLOv8 model.
[0016] Preferably, the constructed general vision transformer BiFPN module includes: a feature pyramid generation unit, a feature integration unit, and a feature fusion unit;
[0017] The feature pyramid generation unit is used to extract features from multiple layers of the backbone network to generate a feature pyramid;
[0018] The feature integration unit introduces bidirectional connections between adjacent levels of the feature pyramid and integrates the feature information of different levels of the feature pyramid through the bidirectional connections;
[0019] The feature fusion unit uses a weighted feature fusion mechanism to fuse the feature information.
[0020] Preferably, the constructed neck feature extraction network Swin Transformer-CondConv2D module includes: a SwinTransformer unit and a CondConv2D dynamic convolution unit;
[0021] The Swin Transformer unit includes: a Patch Partition sub-unit, a W-MSA sub-unit, and a SW-MSA sub-unit. Both the W-MSA sub-unit and the SW-MSA sub-unit are composed of 1 normalization layer, 1 attention module, 1 normalization layer, and 1 MLP layer.
[0022] Preferably, the SPPELAN module consists of an SPP layer and a local attention mechanism, and the dynamic non-monotonic focusing mechanism consists of a Wise-IoU loss function.
[0023] Preferably, the object detection layer with a fourth output layer includes: a first output layer P3, a second output layer P4, a third output layer P5, and a fourth output layer P2;
[0024] The detection feature map corresponding to the first output layer P3 has a size of 80×80 and is used to detect objects larger than 8×8;
[0025] The detection feature map corresponding to the second output layer P4 has a size of 40×40 and is used to detect objects larger than 16×16;
[0026] The detection feature map corresponding to the third output layer P5 has a size of 20×20 and is used to detect objects larger than 32×32;
[0027] The detection feature map corresponding to the fourth output layer P2 has a size of 160×160 and is used to detect objects larger than 4×4.
[0028] Preferably, the method for obtaining the fire location includes:
[0029] Collect the flame RGB image of the area to be detected, align the depth information with the flame RGB image, and obtain the three-dimensional coordinates of the image center point of the flame RGB image;
[0030] Use the flame detection model to identify the flame RGB image, draw the prediction box of the target through the Plot function, and calculate the coordinates of the center point of the prediction box;
[0031] Calculate the three-dimensional coordinates of the identified flame based on the three-dimensional coordinates of the image center point and the coordinates of the center point of the prediction box to obtain the fire location.
[0032] The present invention also provides an infrared flame detection system based on deep learning. The detection system applies the detection method described in any one of the above, and includes: an image annotation module, an image enhancement module, a model improvement module, and a detection module;
[0033] The image annotation module is used to obtain an infrared flame image set under different scenarios and annotate the infrared flame image set;
[0034] The image enhancement module is used to perform Mosaic data enhancement on the annotated infrared flame image set to obtain an enhanced image set;
[0035] The model improvement module is used to improve the YOLOv8 model and train the improved model based on the enhanced image set to obtain a flame detection model;
[0036] The detection module uses the flame detection model to detect and locate the flame to obtain the fire location.
[0037] Compared with the prior art, the beneficial effects of the present invention are as follows:
[0038] The present invention improves the traditional YOLOv8 model. First, a relevant data set is constructed by combining a web crawler and a public infrared flame image set, avoiding the indirect fusion of information at different levels through continuous iteration. A novel general vision transformer BiFPN module is constructed, and the dynamic convolution CondConv2D and Swin Transformer are fused in the neck network to construct a new backbone feature extraction network SwinTransformer-CondConv2D module to reduce the model complexity; a SPPELAN module is proposed, which adds random factors to improve the generalization performance of the model; a dynamic non-monotonic focusing mechanism is proposed, and a Wise-IoU (WIoU) loss function is designed. The dynamic non-monotonic focusing mechanism uses the "degree of deviation" to replace IoU for quality evaluation of anchor boxes and provides a wise gradient gain allocation strategy. This strategy reduces the competitiveness of high-quality anchor boxes while also reducing the harmful gradients generated by low-quality examples, and improves the feature extraction ability of the model while slightly increasing the model complexity; a fourth output layer is added based on the YOLO V8 object detection layer to further process the image features, and finally the object detection performance is evaluated by the computational amount, average precision, and inference speed, and then an improved flame detection model is constructed. On the basis of improving the recognition accuracy and feature extraction performance, this model can quickly and accurately detect and locate the infrared flame image target. Finally, after detecting the target flame, the fire extinguishing is carried out on the target area by using its own fire water cannon device. Description of the Drawings
[0039] In order to more clearly illustrate the technical solutions of the present invention, the drawings required for use in the embodiments will be briefly introduced below. Obviously, the drawings in the following description are only some embodiments of the present invention. For those of ordinary skill in the art, other drawings can be obtained based on these drawings without creative efforts.
[0040] Figure 1 It is a schematic flowchart of the method according to the embodiment of the present invention;
[0041] Figure 2Schematic diagram of the neck feature extraction network Swin Transformer-CondConv2D module according to an embodiment of the present invention;
[0042] Figure 3 Schematic diagram of the structure of the reconnaissance and strike integrated robot according to an embodiment of the present invention;
[0043] Explanation of reference numerals:
[0044] 1. Top camera; 2. Front dual-lens pan-tilt; 3. Signal receiver; 4. LiDAR; 5. Obstacle avoidance radar; 6. Chassis system. Detailed implementation manners
[0045] Next, the technical solutions in the embodiments of the present invention will be clearly and completely described in conjunction with the accompanying drawings in the embodiments of the present invention. Obviously, the described embodiments are only a part of the embodiments of the present invention, rather than all the embodiments. All other embodiments obtained by those of ordinary skill in the art based on the embodiments of the present invention without creative efforts shall fall within the protection scope of the present invention.
[0046] To make the above objects, features, and advantages of the present invention more obvious and understandable, the present invention will be further described in detail below in conjunction with the accompanying drawings and specific implementation manners.
[0047] Embodiment 1
[0048] In this embodiment, as Figure 1 shown, an infrared flame detection method based on deep learning includes the following steps:
[0049] S1. Obtain infrared flame image sets in different scenarios and annotate the infrared flame image sets.
[0050] In this embodiment, an infrared flame image set is obtained, and flame scene images in different scenarios are saved by using a web crawler and an open-source data set. The saving format is the VOC format, and the images in the VOC format are annotated to obtain an annotated infrared flame image set.
[0051] S2. Perform Mosaic data augmentation on the annotated infrared flame image set to obtain an augmented image set.
[0052] In this embodiment, the obtained VOC format labels and images are converted into txt files in the YOLO format label through a script and used as the initial training set and validation set; Mosaic data augmentation and adaptive anchor box calculation are performed on the converted image set. Among them, Mosaic splices the images in a way of random scaling, cropping, and arranging, which improves the detection effect on small target flames.
[0053] S3. Improve the YOLOv8 model and train the improved model based on the enhanced image set to obtain a flame detection model.
[0054] The methods for improving the YOLOv8 model include: constructing a general vision transformer BiFPN module; fusing SwinTransformer and CondConv2D in the neck network of the YOLOv8 model to construct a neck feature extraction network SwinTransformer-CondConv2D module; adding the SPPELAN module and a dynamic non-monotonic focusing mechanism; adding a fourth output layer in the object detection layer of the YOLOv8 model.
[0055] The constructed general vision transformer BiFPN module includes: a feature pyramid generation unit, a feature integration unit, and a feature fusion unit. The feature pyramid generation unit is used to extract features from multiple layers of the backbone network to generate a feature pyramid; the feature integration unit introduces bidirectional connections between adjacent levels of the feature pyramid and integrates the feature information of different levels of the feature pyramid through the bidirectional connections; the feature fusion unit adopts a weighted feature fusion mechanism to fuse the feature information.
[0056] In this embodiment, the feature pyramid generation unit network generates a feature pyramid by extracting features from multiple layers of the backbone network (usually a convolutional neural network such as ResNet). Different from the traditional FPN, the BiFPN in the feature integration unit introduces bidirectional connections between adjacent levels of the feature pyramid. This means that information can flow from higher-level features to lower-level features (top-down path), and can also flow from lower-level features to higher-level features (bottom-up path); the bidirectional connections allow the integration of information from different levels of the feature pyramid in both directions. This integration helps to effectively capture multi-scale features; the feature fusion unit adopts a weighted feature fusion mechanism to combine features at different levels, and the fusion weights are learned during training to ensure optimal feature integration. The bidirectional connections in BiFPN help to better capture feature representations at different scales and improve the network's ability to process objects of different sizes and complexities. This is particularly important in object detection tasks because the sizes of objects in images can vary significantly.
[0057] The constructed neck feature extraction network's Swin Transformer-CondConv2D module includes: a Swin Transformer unit and a CondConv2D dynamic convolution unit; the Swin Transformer unit includes: a PatchPartition sub-unit, a W-MSA sub-unit, and a SW-MSA sub-unit. Both the W-MSA sub-unit and the SW-MSA sub-unit are composed of 1 normalization layer, 1 attention module, 1 normalization layer, and 1 MLP layer.
[0058] In this embodiment, the Patch Partition sub-unit is similar to ViT. The input image H×W×3 is divided into non-overlapping patch sets through Patch Partition, where the size of each patch is 4×4. Then the feature dimension of each patch is 4×4×3 = 48, and the number of patch blocks is H / 4×W / 4. By default, given an image of 224×224×3, after patch partition, the size of the image is 56×56×48 (56 = 224 / 4, 48 = 16*3, 3 is the number of RGB channels); in the Swin Transformer unit, the Window MSA (W-MSA sub-unit) and the Shift Window MSA (SW-MSA sub-unit) are used to replace the standard multi-head self-attention (MSA) module used in ViT. Each sub-unit consists of a normalization layer, an attention module, another normalization layer, and an MLP layer.
[0059] The SPPELAN module is composed of an SPP layer and a local attention mechanism. The dynamic non-monotonic focusing mechanism is composed of a Wise-IoU loss function.
[0060] In this embodiment, the SPPELAN model is also introduced, that is, Spatial Pyramid Pooling with Enhanced Local Attention Network. The SPPELAN module combines SPP with a local attention mechanism, aiming to further improve the accuracy of object detection. The local attention mechanism is a method that can focus on the local region information in an image. By introducing the local attention mechanism, SPPELAN can capture more refined object features at different scales, thereby improving the accuracy of object detection. At the same time, due to the small scope of action of the local attention mechanism and relatively small computational amount, it will not increase the computational burden too much. By combining SPP with the local attention mechanism, SPPELAN not only inherits the advantages of SPP but also can further improve the performance of object detection.
[0061] The dynamic non - monotonic focusing mechanism consists of the Wise - IoU loss function. In this embodiment, object detection, as a core issue in computer vision, its detection performance depends on the design of the loss function. The bounding box loss function, as an important part of the object detection loss function, its good definition will bring significant performance improvement to the object detection model. Most recent research assumes that the examples in the training data have high quality and is committed to strengthening the fitting ability of the bounding box loss. However, we notice that the object detection training set contains low - quality examples. If we blindly strengthen the regression of the bounding box to low - quality examples, it will obviously harm the improvement of the model detection performance. Focal - EIoU v1 was proposed to solve this problem, but due to its static focusing mechanism, it does not fully exploit the potential of the non - monotonic focusing mechanism. Based on this view, this embodiment proposes a dynamic non - monotonic focusing mechanism and designs the Wise - IoU (WIoU). The dynamic non - monotonic focusing mechanism uses "outlier degree" to replace IoU for quality assessment of anchor boxes and provides a wise gradient gain allocation strategy. This strategy reduces the competitiveness of high - quality anchor boxes while also reducing the harmful gradients generated by low - quality examples. This enables WIoU to focus on anchor boxes of ordinary quality and improve the overall performance of the detector.
[0062] The object detection layer with the fourth output layer includes: the first output layer P3, the second output layer P4, the third output layer P5, and the fourth output layer P2; the detection feature map corresponding to the first output layer P3 has a size of 80×80 and is used to detect objects with a size of more than 8×8; the detection feature map corresponding to the second output layer P4 has a size of 40×40 and is used to detect objects with a size of more than 16×16; the detection feature map corresponding to the third output layer P5 has a size of 20×20 and is used to detect objects with a size of more than 32×32; the detection feature map corresponding to the fourth output layer P2 has a size of 160×160 and is used to detect objects with a size of more than 4×4.
[0063] After that, the improved YOLOv8 model is iteratively trained using the training set and the validation set to obtain the flame detection model.
[0064] S4. Use the flame detection model to detect and locate the flame to obtain the fire location.
[0065] The method for obtaining the fire location includes: collecting the flame RGB image of the area to be detected, aligning the depth information with the flame RGB image to obtain the three - dimensional coordinates of the image center point of the flame RGB image; using the flame detection model to identify the flame RGB image, drawing the prediction box of the target through the Plot function, and calculating the coordinates of the center point of the prediction box; calculating the three - dimensional coordinates of the identified flame based on the three - dimensional coordinates of the image center point and the coordinates of the center point of the prediction box to obtain the fire location.
[0066] Embodiment 2
[0067] In this embodiment, an infrared flame detection system based on deep learning includes: an image annotation module, an image enhancement module, a model improvement module, and a detection module.
[0068] The image annotation module is used to obtain infrared flame image sets in different scenarios and annotate the infrared flame image sets.
[0069] In this embodiment, the image annotation module obtains an infrared flame image set, saves the flame scene images in different scenarios using a web crawler and an open-source data set in the VOC format, annotates the VOC-format images, and obtains an annotated infrared flame image set.
[0070] The image enhancement module is used to perform Mosaic data enhancement on the annotated infrared flame image set to obtain an enhanced image set.
[0071] In this embodiment, the image enhancement module converts the obtained VOC-format labels and images into txt files in the YOLO format through a script and uses them as the initial training set and validation set; performs Mosaic data enhancement and adaptive anchor box calculation on the converted image set, where Mosaic is to splice images in a way of random scaling, cropping, and arranging, which improves the detection effect of small target flames.
[0072] The model improvement module is used to improve the YOLOv8 model and train the improved model based on the enhanced image set to obtain a flame detection model.
[0073] The method for improving the YOLOv8 model includes: constructing a general vision transformer BiFPN module; fusing SwinTransformer and CondConv2D in the neck network of the YOLOv8 model to construct a neck feature extraction network SwinTransformer-CondConv2D module; adding an SPPELAN module and a dynamic non-monotonic focusing mechanism; adding a fourth output layer in the object detection layer of the YOLOv8 model.
[0074] The constructed general vision transformer BiFPN module includes: a feature pyramid generation unit, a feature integration unit, and a feature fusion unit. The feature pyramid generation unit is used to extract features from multiple layers of the backbone network to generate a feature pyramid; the feature integration unit introduces bidirectional connections between adjacent levels of the feature pyramid and integrates the feature information of different levels of the feature pyramid through the bidirectional connections; the feature fusion unit uses a weighted feature fusion mechanism to fuse the feature information.
[0075] In this embodiment, the feature pyramid generation unit network generates a feature pyramid by extracting features from multiple layers of a backbone network (usually a convolutional neural network such as ResNet). Different from the traditional FPN, the BiFPN introduces bidirectional connections between adjacent levels of the feature pyramid. This means that information can flow from higher-level features to lower-level features (top-down path), and can also flow from lower-level features to higher-level features (bottom-up path); the bidirectional connection allows the integration of information from different levels of the feature pyramid in both directions. This integration helps to effectively capture multi-scale features; the feature fusion unit adopts a weighted feature fusion mechanism to combine features at different levels, and the weights for fusion are learned during training, ensuring optimal feature integration. The bidirectional connection in BiFPN helps to better capture feature representations at different scales, improving the network's ability to process objects of different sizes and complexities. This is particularly important in object detection tasks because the sizes of objects in an image can vary significantly.
[0076] The constructed neck feature extraction network Swin Transformer-CondConv2D module is as Figure 2 shown, including: a Swin Transformer unit and a CondConv2D dynamic convolution unit; the Swin Transformer unit includes: a PatchPartition sub-unit, a W-MSA sub-unit, and a SW-MSA sub-unit. Both the W-MSA sub-unit and the SW-MSA sub-unit are composed of 1 normalization layer, 1 attention module, 1 normalization layer, and 1 MLP layer.
[0077] In this embodiment, the Patch Partition sub-unit is similar to ViT. The input image H×W×3 is divided into a non-overlapping patch set through Patch Partition, where the size of each patch is 4×4. Then the feature dimension of each patch is 4×4×3 = 48, and the number of patch blocks is H / 4×W / 4. By default, for a given 224×224×3 image, after patch partition, the size of the image is 56×56×48 (56 = 224 / 4, 48 = 16*3, 3 is the number of RGB channels); in the Swin Transformer unit, the Window MSA (W-MSA sub-unit) and the Shift Window MSA (SW-MSA sub-unit) are used to replace the standard multi-head self-attention (MSA) module used in ViT. Each sub-unit consists of a normalization layer, an attention module, another normalization layer, and an MLP layer.
[0078] The SPPELAN module consists of the SPP layer and the local attention mechanism, and the dynamic non-monotonic focusing mechanism consists of the Wise-IoU loss function.
[0079] In this embodiment, the SPPELAN model is also introduced, that is, Spatial Pyramid Pooling with Enhanced Local Attention Network. The SPPELAN module combines SPP with the local attention mechanism, aiming to further improve the accuracy of object detection. The local attention mechanism is a method that can focus on the local region information in the image. By introducing the local attention mechanism, SPPELAN can capture more refined object features at different scales, thereby improving the accuracy of object detection. At the same time, due to the small scope of action of the local attention mechanism and relatively small computational amount, it will not increase the computational burden too much. By combining SPP with the local attention mechanism, SPPELAN not only inherits the advantages of SPP but also can further improve the performance of object detection.
[0080] The dynamic non-monotonic focusing mechanism consists of the Wise-IoU loss function. In this embodiment, object detection, as the core problem of computer vision, its detection performance depends on the design of the loss function. The bounding box loss function, as an important part of the object detection loss function, its good definition will bring a significant performance improvement to the object detection model. Most recent studies assume that the examples in the training data have high quality and are committed to strengthening the fitting ability of the bounding box loss. However, we note that the object detection training set contains low-quality examples. If we blindly strengthen the regression of the bounding box to low-quality examples, it will obviously harm the improvement of the model detection performance. Focal-EIoU v1 was proposed to solve this problem, but due to its static focusing mechanism, it does not fully exploit the potential of the non-monotonic focusing mechanism. Based on this view, this embodiment proposes a dynamic non-monotonic focusing mechanism and designs Wise-IoU (WIoU). The dynamic non-monotonic focusing mechanism uses "outlier degree" to replace IoU to evaluate the quality of anchor boxes and provides a wise gradient gain allocation strategy. This strategy reduces the competitiveness of high-quality anchor boxes while also reducing the harmful gradients generated by low-quality examples. This enables WIoU to focus on anchor boxes of ordinary quality and improve the overall performance of the detector.
[0081] The object detection layer with the fourth output layer includes: the first output layer P3, the second output layer P4, the third output layer P5, and the fourth output layer P2; the size of the detection feature map corresponding to the first output layer P3 is 80×80, which is used to detect objects larger than 8×8; the size of the detection feature map corresponding to the second output layer P4 is 40×40, which is used to detect objects larger than 16×16; the size of the detection feature map corresponding to the third output layer P5 is 20×20, which is used to detect objects larger than 32×32; the size of the detection feature map corresponding to the fourth output layer P2 is 160×160, which is used to detect objects larger than 4×4.
[0082] After that, the improved YOLOv8 model is iteratively trained using the training set and the validation set to obtain the flame detection model.
[0083] The detection module uses the flame detection model to detect and locate the flame to obtain the ignition position.
[0084] The working process of the detection module includes: collecting the flame RGB image of the area to be detected, aligning the depth information with the flame RGB image to obtain the three-dimensional coordinates of the image center point of the flame RGB image; using the flame detection model to identify the flame RGB image, drawing the prediction box of the target through the Plot function, and calculating the coordinates of the center point of the prediction box; calculating the three-dimensional coordinates of the identified flame based on the three-dimensional coordinates of the image center point and the coordinates of the center point of the prediction box to obtain the ignition position.
[0085] Embodiment III
[0086] In this embodiment, an integrated reconnaissance and strike robot is also provided. This robot performs flame detection and fire extinguishing based on the above infrared flame detection method, as Figure 3 shown, including: a top camera 1, a front dual-lens pan-tilt 2, a signal receiver 3, a lidar 4, an obstacle avoidance radar 5, and a chassis system 6.
[0087] The top camera 1 is connected to the main control unit through a high-bandwidth interface to transmit video data. The image data of the top camera 1 can be used in coordination with the front-end dual-lens pan-tilt 2 to enhance the field of view and recognition accuracy. The front-end dual-lens pan-tilt 2 is connected to the main control unit through a mechanical pan-tilt control RS-485. The image data transmission is sent to the main control system through a video transmission line. The signal receiver 3 is connected to the main control system through wireless communication to receive external control instructions, positioning data, navigation information, etc. The lidar 4 is connected to the main control system through a high-speed interface to transmit environmental scan data. The point cloud data of the lidar 4 will be used by the processing system for obstacle recognition and environmental modeling. The obstacle avoidance radar 5 is connected to the main control system through an electrical interface to transmit data on real-time obstacle detection. The main control system makes path planning and obstacle avoidance decisions based on this data. The chassis system 6 is connected to the main control unit through a control interface, receives control commands, and feeds back motion state information.
[0088] The embodiments described above are only descriptions of the preferred embodiments of the present invention and do not limit the scope of the present invention. Without departing from the design spirit of the present invention, various deformations and improvements made by those of ordinary skill in the art to the technical solutions of the present invention shall fall within the protection scope determined by the claims of the present invention.
Claims
1. A deep learning-based infrared flame detection method, characterized in that: The following steps are involved: Acquire infrared flame image sets in different scenes, and annotate the infrared flame image sets; Performing Mosaic data enhancement on the labeled infrared flame image set to obtain an enhanced image set; The YOLOv8 model is improved, and the improved model is trained based on the enhanced image set to obtain a flame detection model; The flame detection model is used to detect and locate the flame to obtain the ignition position.
2. According to the deep learning-based infrared flame detection method of claim 1, it is characterized in that: Methods for improving the YOLOv8 model include: Build the BiFPN module, a general vision converter; The Swin Transformer and CondConv2D are integrated into the neck network of the YOLOv8 model to construct the neck feature extraction network Swin Transformer-CondConv2D module. Add SPPELAN module and dynamic non-monotonic focusing mechanism; Add a fourth output layer to the object detection layer of the YOLOv8 model.
3. According to the deep learning-based infrared flame detection method of claim 2, it is characterized in that: The constructed universal visual converter BiFPN module includes: a feature pyramid generation unit, a feature integration unit and a feature fusion unit; The feature pyramid generation unit is used to extract features from multiple layers of the backbone network to generate a feature pyramid; The feature integration unit introduces bidirectional connections between adjacent levels of the feature pyramid, and integrates feature information at different levels of the feature pyramid through the bidirectional connections; The feature fusion unit adopts a weighted feature fusion mechanism to fuse the feature information.
4. According to the deep learning-based infrared flame detection method of claim 2, it is characterized in that: The constructed neck feature extraction network Swin Transformer-CondConv2D module includes: Swin Transformer unit and CondConv2D dynamic convolution unit; The Swin Transformer unit includes: a Patch Partition subunit, a W-MSA subunit and a SW-MSA subunit, and the W-MSA subunit and the SW-MSA subunit are both composed of 1 normalization layer, 1 attention module, 1 normalization layer and 1 MLP layer.
5. The infrared flame detection method based on deep learning according to claim 2 is characterized in that: The SPPELAN module consists of an SPP layer and a local attention mechanism, and the dynamic non-monotonic focusing mechanism consists of a Wise-IoU loss function.
6. The infrared flame detection method based on deep learning according to claim 2 is characterized in that: The target detection layer with the fourth output layer added includes: a first output layer P3, a second output layer P4, a third output layer P5 and a fourth output layer P2; The detection feature map size corresponding to the first output layer P3 is 80×80, which is used to detect targets with a size larger than 8×8; The detection feature map size corresponding to the second output layer P4 is 40×40, which is used to detect targets with a size larger than 16×16; The detection feature map size corresponding to the third output layer P5 is 20×20, which is used to detect targets with a size of 32×32 or larger; The detection feature map size corresponding to the fourth output layer P2 is 160×160, which is used to detect targets with a size larger than 4×4.
7. The infrared flame detection method based on deep learning according to claim 1 is characterized in that: The method for obtaining the ignition position includes: Collecting a flame RGB image of the area to be detected, aligning the depth information with the flame RGB image, and obtaining the three-dimensional coordinates of the image center point of the flame RGB image; The flame RGB image is recognized by using the flame detection model, a prediction box of the target is drawn by using the Plot function, and the coordinates of the prediction box center point of the prediction box are calculated; The three-dimensional coordinates of the identified flame are calculated based on the three-dimensional coordinates of the center point of the image and the coordinates of the center point of the prediction frame to obtain the ignition position.
8. An infrared flame detection system based on deep learning, the detection system applying the detection method according to any one of claims 1 to 7, characterized in that: include: Image annotation module, image enhancement module, model improvement module and detection module; The image annotation module is used to obtain infrared flame image sets in different scenes and annotate the infrared flame image sets; The image enhancement module is used to perform Mosaic data enhancement on the annotated infrared flame image set to obtain an enhanced image set; The model improvement module is used to improve the YOLOv8 model, and train the improved model based on the enhanced image set to obtain a flame detection model; The detection module detects and locates the flame using the flame detection model to obtain the ignition position.
Citation Information
Cited By
Visual perception method and system for tunnel drilling blast holes
CN120852517A