A Visible Light Forest Fire Detection Method Based on a Lightweight Anchor-Free Detection Model

By adopting a lightweight anchor-free detection model in forest fire detection, combined with YOLOX and MobileNetv3g, the problem of limited equipment configuration is solved and efficient and accurate fire detection is achieved.

CN115423998BActive Publication Date: 2025-06-24XIDIAN UNIV
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202211027153.6
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2022-08-25
Publication Date
2025-06-24
Estimated Expiration
2042-08-25

AI Technical Summary

Technical Problem

Existing deep learning forest fire detection methods are difficult to achieve efficient detection when device configuration and computing capabilities are limited, especially on mobile devices.

Method used

Using a lightweight anchor-free detection model method, the network architecture combined with YOLOX and MobileNetv3g is used to reduce the complexity of the model, and adaptive spatial feature fusion is carried out through the ASFF module to improve detection accuracy.

Benefits of technology

It realizes efficient detection of forest fires on mobile devices, improves detection accuracy and computing efficiency, and reduces model complexity.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN115423998B_ABST
    Figure CN115423998B_ABST
Patent Text Reader

Abstract

A visible light forest fire detection method based on a lightweight anchor-free detection model, annotate the position and category information of the flames in each image of the infrared forest fire image dataset to obtain a flame dataset, and divide it into a training set, a validation set and a test set; construct an improved lightweight anchor-free detection neural network, use the backbone feature extraction network MobileNetv3g and multi-scale max pooling operation to extract fire features, and use D-PANet to perform enhanced feature fusion on the fire features from deep to shallow for the feature layers; then input it into the ASFF module for adaptive spatial feature fusion, and use the decoupled head to decode and predict the flame target category information and position regression information for multi-scale prediction to obtain the prediction results; perform score sorting and non-maximum suppression screening on the prediction results, and screen out the prediction box with the highest score and meeting the confidence level within a certain area, thereby obtaining the final fire prediction result. The present invention has good detection effects.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention belongs to the technical field of deep learning image processing, and particularly relates to a visible light forest fire detection method based on a lightweight anchor-free detection model. Background Technique

[0002] Forest fires are natural disasters with strong suddenness, great destructive power, wide occurrence range and difficult extinguishment. They not only burn down forest trees and damage the forest environment, but also seriously affect the stability and balance of the forest ecosystem and threaten the life and property safety of the people. With the continuous development of imaging technology and computer vision technology, image-based fire detection technology has become a major research hotspot. It can not only meet the requirements of long-distance fire detection, but also provide detailed fire information, which is conducive to the timely discovery and extinguishment of forest fires. The fire target features in visible light images are very complex, and it is difficult to obtain ideal detection results using traditional algorithms. In recent years, deep learning methods have shown excellent performance in the field of object detection, providing new ideas for forest fire detection. Among the commonly used deep learning algorithms, the YOLO series algorithms are widely used in the field of object detection due to their incomparable speed and accuracy.

[0003] Although using deep learning-based methods for forest fire detection can obtain relatively high detection accuracy, since the deep learning model needs to be trained, the configuration requirements will be relatively high. However, forest fire detection systems are often carried on mobile devices such as inspection platforms or drones, and the configuration requirements for the devices cannot be too high. To address the above problems, the present invention takes the latest anchor-free structure YOLOX in the YOLO series as the basic framework, introduces the improved lightweight network MobileNetv3g, and establishes a lightweight anchor-free visible light forest fire detection model, so that the model can be better carried on mobile devices and can obtain relatively accurate detection results. Summary of the Invention

[0004] In view of this, the main purpose of the present invention is to provide a visible light forest fire detection method based on a lightweight anchor-free detection model.

[0005] To achieve the above object, the technical solution of the present invention is realized as follows:

[0006] A forest fire detection method based on a lightweight anchor-free detection model, comprising the following steps:

[0007] Step 1, obtain an infrared forest fire image dataset, mark the position and category information of the flames in each image to obtain a flame dataset, preprocess each marked image, and divide it into a training set, a validation set and a test set;

[0008] Step 2: Construct an improved lightweight anchor-free detection model. Use the backbone feature extraction network MobileNetv3g to extract fire features from the input image, obtaining three effective feature layers: Feature1, Feature2, and Feature3;

[0009] Step 3: Perform multi-scale max pooling on the effective feature layer Feature3 through the Spatial Pyramid Pooling module SPP;

[0010] Step 4: Input the effective feature layers Feature1, Feature2, and Feature3 into the enhanced feature extraction network D-PANet. Perform enhanced feature fusion on the feature layers from deep to shallow, and then use downsampling to transfer the position information of the shallow layer into the deep feature, which is fused with the semantic information of the deep layer, outputting the feature layers Level1, Level2, and Level3;

[0011] Step 5: Input the feature layers Level1, Level2, and Level3 into the Adaptive Spatial Feature Fusion module ASFF for adaptive spatial feature fusion, and then use 1×1 convolution to compress to the original number of channels, outputting the feature layers ASFF-1, ASFF-2, and ASFF-3;

[0012] Step 6: Input the feature layers ASFF-1, ASFF-2, and ASFF-3 into the prediction network for multi-scale prediction. Use the decoupled head to decode the predicted object category information and location regression information to obtain the prediction results;

[0013] Step 7: Perform score sorting and non-maximum suppression screening on the prediction results. Screen out the prediction boxes with the highest scores and meeting the confidence level within a certain area to obtain the final fire prediction results.

[0014] Compared with the prior art, the present invention can effectively reduce the model complexity and improve the model detection accuracy for problems such as complex image backgrounds and limited computing capabilities of mobile devices carried by forest fire detection systems, and has good detection effects. BRIEF DESCRIPTION OF THE DRAWINGS

[0015] Figure 1 It is a diagram of the MobileNetv3g-DAF-YOLOX model in the embodiment of the present invention.

[0016] Figure 2 It is a schematic flow diagram of the present invention.

[0017] Figure 3 It is an example diagram of the basic structure of MobileNetv3g.

[0018] Figure 4 It is a structural diagram of the DAF-PANet framework.

[0019] Figure 5 It is the structural diagram of the ASFF module.

[0020] Figure 6 It is the schematic diagram of flame target marking.

[0021] Figure 7 It is the P-R curve of the detection results of the algorithm of the present invention and the comparative algorithm, where (a) is the comparison with the anchor-based algorithm, and (b) is the comparison with the anchor-free algorithm.

[0022] Figure 8 It is the comparison diagram of the detection results of the Anchor-based algorithm, where (a) is the ground truth map of object detection in forest fire images, (b) is the fire detection result map of the YOLOv3 algorithm, (c) is the fire detection result map of the SSD300 algorithm, (d) is the fire detection result map of the Faster R-CNN algorithm, and (e) is the schematic diagram of the fire detection result of the algorithm of the present invention.

[0023] Figure 9 It is the comparison diagram of the detection results of the Anchor-free algorithm, where (a) is the ground truth map of object detection in forest fire images, (b) is the fire detection result map of the FCOS algorithm, (c) is the fire detection result map of the CenterNet algorithm, and (d) is the schematic diagram of the fire detection result of the algorithm of the present invention. Detailed implementation manners

[0024] In order to make the objectives, technical solutions and advantages of the present invention clearer and more understandable, the present invention will be further described in detail below with reference to the accompanying drawings and embodiments. It should be understood that the specific embodiments described herein are only used to explain the present invention and are not used to limit the present invention.

[0025] The embodiment of the present invention provides a visible light forest fire detection method based on a lightweight anchor-free detection model. The corresponding fire detection model is MobileNetv3g-DAF-YOLOX, and the structure is as Figure 1 shown. First, the latest anchor-free object detection framework YOLOX is used as the basic architecture. Secondly, considering that the forest fire monitoring system is usually carried on mobile devices such as inspection platforms and drones, and the configuration requirements cannot be too high, the backbone network of the YOLOX model is replaced with the improved lightweight network MobileNetv3g; then, depthwise separable convolutions are used to simplify PANet and reduce the complexity of the model; finally, the ASFF module is introduced to adaptively learn the weight parameters of feature fusion at each level to achieve the purpose of consistent feature extraction.

[0026] Specifically, refer to Figure 2, the visible light forest fire detection method based on the lightweight anchor-free detection model of the present invention includes the following steps:

[0027] Step 1: Obtain an infrared forest fire image dataset, annotate the positions and category information of the flames in each image to obtain a flame dataset, preprocess each annotated image, and divide it into a training set, a validation set, and a test set.

[0028] Exemplarily, in this step, the dataset used contains 2,514 forest fire images, of which 2,003 are from the FLAME public dataset proposed by scholars from Northern Arizona University, and 511 are from the fire datasets provided by Bilkent University and Durham University, which are jointly used as the dataset of the present invention. At the same time, the forest fire dataset is randomly divided into three parts: a training set, a validation set, and a test set, with a ratio of 7:2:1. Then, data augmentation strategies such as mosaic, random affine transformation, random flipping, and contrast change are used to process the images in the forest fire dataset, and the image size is adjusted to a unified pixel size, such as 640×640. Finally, the annotation software LabelImg is used to annotate the flames in the images.

[0029] Step 2: Construct an improved lightweight anchor-free detection model, use the backbone feature extraction network MobileNetv3g to extract fire features from the input image, and obtain three effective feature layers Feature1, Feature2, and Feature3.

[0030] Exemplarily, in this step, MobileNetv3g is an improved version of MobileNetv3, and its basic structure is as Figure 3 shown. First, use a one-dimensional convolution to replace the fully connected layer of the squeeze-and-excitation network (SENet) attention mechanism module, then add a common convolution module before the depthwise separable convolution, and combine the output results of the two to fully obtain semantic information, making the detection results more accurate.

[0031] The specific improvement method: Using the anchor-free detection model YOLOX as the basic framework, first replace the backbone network of YOLOX with the lightweight network MobileNetv3, so that the detection model can be carried on mobile devices such as drones; then, improve the MobileNetv3 network by combining the efficient channel attention network (ECANet) and the GhostNet idea to construct the MobileNetv3g backbone network, so as to retain multi-scale feature information while making the model lightweight; finally, use depthwise separable convolution to simplify the path aggregation network, further reduce the computational amount of the model, and introduce the adaptive spatial feature fusion module ASFF to reduce the inconsistency between different scale feature maps.

[0032] The MobileNetv3g of the present invention includes the following components:

[0033] (a) ECANet attention mechanism: After global average pooling, ECANet obtains local cross-channel interaction information of each channel and its adjacent k channels through a one-dimensional convolution of size k. ECANet adaptively determines the convolution kernel size k by given number of channels C, specifically:

[0034]

[0035] In the formula, represents the nearest odd number to There is a non-linear mapping relationship φ between k and C: C = φ(k) = 2 γ*k-b , where γ and b are constant coefficients. In the present invention, γ and b are set to 2 and 1 respectively.

[0036] (b) Improved depthwise separable convolution structure: The MobileNetv3 is mainly improved by referring to the structure of GhostNet. First, a common 3×3 convolution is used to map the number of channels of the input feature map to half of the original number of channels, then the depthwise separable convolution is used to further extract features, and finally the feature maps output by the two parts are added together.

[0037] (c) h-swish activation function: The h-swish activation function is an improvement of the swish activation function in the MobileNetv3 network. Compared with the swish function, the h-swish function has a faster calculation speed and better numerical stability. The expression of the h-swish activation function is:

[0038]

[0039] In the formula, ReLU6(x + 3) represents the ReLU6 activation function, that is, the maximum output value of ReLU is limited to 6, which can be expressed as ReLU6(x) = min(6, max(0, x)).

[0040] The backbone feature extraction network MobileNetv3g of the present invention is divided into six stages, namely Stage1, Stage2, Stage3, Stage4, Stage5 and Stage6. The overall architecture parameters of MobileNetv3g are designed as shown in Table 1 below:

[0041] Table 1 MobileNetv3g network architecture parameter table

[0042]

[0043]

[0044] In Table 1, the "Input" column represents the input feature layer size and number of channels of each structural block of MobileNetv3, the "Output" column represents the output feature layer size and number of channels of each structural block of MobileNetv3, and the "Operator" column represents the type of structural block that the feature layer passes through. "bneck" represents an inverted residual structure with a linear bottleneck, which includes a convolutional layer, a BN layer, and an activation layer. The "expsize" column represents the number of feature channels after the increase in the inverted residual structure, the "c" column represents the number of channels actually output by each structural block, the "ECA" column represents whether the ECA attention module is introduced in this structural block, the "NL" column represents the type of activation function, where "HS" represents the h-swish activation function, "RE" represents the ReLU activation function, and the "s" column represents the stride of each structural block; "NBN" means that the BN layer is not applicable.

[0045] In the present invention, the output feature layers of Stage3, Stage5, and Stage6 are taken as the effective feature layers Feature1, Feature2, and Feature3, and are input into the enhanced feature extraction network of YOLOX.

[0046] Step 3: Perform multi-scale max pooling processing on the effective feature layer Feature3 through the SPP module.

[0047] Step 4: Input the feature layers Feature1, Feature2, and Feature3 into the D-PANet module, perform enhanced feature fusion on the feature layers from deep to shallow, and then use downsampling to transfer the position information of the shallow layer into the deep feature, and fuse it with the semantic information of the deep layer to output the feature layers Level1, Level2, and Level3.

[0048] Exemplarily, in this step, the three effective feature layers Feature1, Feature2, and Feature3 are input into the D-PANet to achieve path aggregation. The basic structure of the D-PANet is as Figure 4 shown. It uses depthwise separable convolutions to replace the 3×3 ordinary convolutions in the path aggregation network PANet, thereby reducing the complexity of the model. The specific process is as follows: The D-PANet is interspersed among the three effective feature layers Feature1, Feature2, and Feature3, repeatedly extracts features, upsamples the underlying features, stacks them with the same-dimensional feature layer of the previous layer, downsamples the high-level features, and stacks them with the same-dimensional feature layer of the next layer to achieve feature fusion, and finally outputs three feature layers Level1, Level2, and Level3.

[0049] Step 5: Input the feature layers Level1, Level2, and Level3 into the ASFF module for adaptive spatial feature fusion, and then use 1×1 convolution to compress them to the original number of channels, outputting the feature layers ASFF-1, ASFF-2, and ASFF-3. By adaptively learning the spatial weights of feature fusion, more hierarchical features can be extracted, thereby improving the detection accuracy.

[0050] Since the output number of channels expands to twice the input number of channels after the ASFF module adds the feature layers, it is necessary to first use 1×1 convolution to compress the feature layers back to the original number of channels before inputting them into the prediction network.

[0051] Different from traditional multi-level feature fusion methods based on element summation or concatenation, the ASFF module alleviates the inconsistency problem by adaptively learning the spatial fusion weights of feature maps at each scale. The ASFF module can be divided into two parts: unified scaling and adaptive spatial feature fusion. Its basic structure is as Figure 5 shown as follows:

[0052] (a) In the unified scaling part, upsampling and downsampling operations are mainly used to adjust the feature layers. For upsampling, first use 1×1 convolution to compress the number of channels to the same as Level1, and then use interpolation upsampling to increase the resolution; for 2-fold downsampling, directly use a 3×3 convolution with a stride of 2 to modify the size and number of channels of the feature layer; for 4-fold downsampling, add a max pooling layer with a stride of 2 before the convolutional layer.

[0053] (b) In the adaptive spatial feature fusion part, let X n (n ∈ {1, 2, 3}) be the feature layer Leveln output by PANet, and Y l (l ∈ {1, 2, 3}) be the new feature ASFF-l obtained by X 1 , X 2 and X 3 passing through the ASFF module. Multiply X 1 , X 2 , X 3 by the weight parameters α l , β l , γ l respectively and sum them to obtain the fused feature Y l . The fusion process can be expressed as:

[0054]

[0055] In the formula, represents the feature vector of the ASFF output feature map Y l at (i, j), Denote the feature vector at (i, j) of the feature map obtained by scaling Leveln to Levell. respectively represent the spatial weight parameters from three feature layers to Levell, and through the Softmax function, and

[0056] Adding the ASFF module to the enhanced feature extraction network can effectively filter out the features belonging to other layers by multiplying the features of each layer by the corresponding weight parameters and then summing them, and only retain the feature information useful for the current layer, so that the extracted features are hierarchical and the detection effect of the model is better.

[0057] Step 6: Input the feature layers ASFF-1, ASFF-2, and ASFF-3 into the prediction network for multi-scale prediction, and use the decoupled head to decode the predicted object category information and location regression information to obtain the prediction results.

[0058] Exemplarily, in this step, the prediction network uses the prediction network of YOLOX, which mainly includes a new decoupled head, the anchor-free idea, and the SimOTA dynamic positive sample matching method. It specifically includes the following parts.

[0059] (9a) New decoupled head: In object detection, the conflict between the classification and regression tasks is a well-known problem. The YOLOX network divides the YOLO detection head into two parts for classification and regression. First, a 1×1 convolutional layer is used to reduce the channel size, then two parallel branches are respectively used for the classification and regression tasks, and finally they are integrated together during prediction.

[0060] (9b) Anchor-Free idea: The anchor-free mechanism significantly reduces the number of design parameters that need to be heuristically adjusted and abandons techniques such as anchor clustering and grid sensitivity, making the detector simpler. The specific process is as follows: The number of predictions at each position is reduced from 3 to 1, and it can directly predict four values, namely the two offsets to the upper left corner of the grid, the height and width of the prediction box. At the same time, the center position of each object is designated as a positive sample, and a scale range is preset to assign the level of the feature pyramid to the object. This modification reduces the number of parameters of the detector, makes the detection speed faster, and can obtain better detection performance.

[0061] (9c) SimOTA Dynamic Positive Sample Matching Method: OTA (Optimal Transport Assignment) mainly analyzes label assignment from a global perspective, formulates the assignment process as an OT (Optimal Transport) problem, and has good performance in the current label assignment strategy. Simplify OTA to the dynamic top-k strategy SimOTA, which reduces the training time by obtaining an approximate solution and reduces the additional hyperparameters in the Sinkhorn-Knopp algorithm. The specific process is as follows: SimOTA first calculates the pairwise matching degree, which can be expressed as the cost or quality of each prediction-ground truth pair. In simOTA, the cost between the ground truth g i and the prediction p j is calculated as:

[0062]

[0063] where λ represents the balance coefficient, and represent the classification loss and regression loss between the ground truth g i and the prediction p j respectively.

[0064] Then, for the ground truth g i , select the top k predictions with the minimum cost within the fixed central region as positive samples. Finally, assign the corresponding grids of these positive predictions as positive, and the remaining grids as negative. Among them, the value of k is dynamic and changes with the ground-truth.

[0065] Step 7: Sort the prediction results by score and perform non-maximum suppression filtering to select the prediction box with the highest score and meeting the confidence within a certain region, thereby obtaining the final fire prediction result.

[0066] The following further describes the effect of the present invention in combination with simulation experiments.

[0067] 1. Simulation Conditions:

[0068] The hardware environment of the simulation experiment of the present invention is AMD Ryzen 5 5600H with Radeon Graphics CPU and NVIDIA GeForce GTX 2080Ti, and the Ubuntu 18.04 operating system. The simulation software is: the deep learning framework is Pytorch1.4.0, and the algorithm program is written in Python3.7.

[0069] 2. Experimental Content: To demonstrate the effectiveness of the visible light forest fire detection method based on the lightweight anchor-free detection model, the dataset used in this invention contains 2,514 forest fire images, of which 2,003 are from the FLAME public dataset proposed by scholars from Northern Arizona University. The resolution of the images in the FLAME dataset is 3840×2160, which was collected by drones during the burning of dead wood designated in the Arizona pine forest. Considering that the background of the fire images in the FLAME dataset is relatively single, 511 images containing forest fires were selected from the fire datasets provided by Bilkent University and Durham University and used together with the FLAME dataset as the dataset for the experiments of this invention.

[0070] Before training the fire detection model, the adopted forest fire dataset is first randomly divided into three parts: training set, validation set, and test set, with a ratio of 7:2:1. Then, data augmentation strategies such as mosaic, random affine transformation, random flipping, and contrast change are used to process the images in the forest fire dataset, and the size of the images input into the model is adjusted to 640×640 pixels. Finally, the flame in the images is labeled using the annotation software LabelImg, as Figure 6 shown.

[0071] During the training process, the number of images in a batch is set to 8, the SGD (Stochastic Gradient Descent) optimizer is used, the momentum is set to 0.9, the decay is set to 0.0005, the initial learning rate is set to 0.01, the number of iterations is 300, and the cosine learning rate decay strategy is applied.

[0072] To verify the performance of the algorithm of this invention, the five currently most advanced fire detection algorithms are used as comparison algorithms, namely YOLOv3, SSD300, Faster R-CNN, FCOS, and CenterNet algorithms, and the detection effects of each algorithm on the forest fire dataset are compared. As shown in Table 1, the comparative experiments are mainly divided into two parts: experiments using the anchor-based method and experiments using the anchor-free method. Among them, the anchor-based method includes the YOLOv3, SSD300, and Faster R-CNN methods, and the anchor-free method includes the FCOS and CenterNet methods. Figure 7 is the P-R curve of the detection results of the algorithm of this invention and the comparison algorithms, Figure 8 is the comparison chart of the detection results of the anchor-based method, Figure 9 is the comparison chart of the detection results of the anchor-free method.

[0073] Comparison of Evaluation Metrics for Detection Results of Different Algorithms

[0074]

[0075] In the comparative experiment using the anchor - based method, the detection results of the algorithm of the present invention were compared with those of YOLOv3, SSD300, and Faster R - CNN algorithms. As shown in Table 2, the two - stage detection method Faster R - CNN has the highest detection accuracy, with an AP of 0.9470, only 0.0129 lower than the method of the present invention. However, the FLOPs of Faster R - CNN is 61.09G, much larger than 9.67G of the present invention method. Figure 7 (a)'s P - R curve further proves the superiority of the detection method of the present invention.

[0076] From Figure 8 (b), it can be seen that the YOLOv3 algorithm has a problem of missed detection; due to the large model, Figure 8 (c) shows that the SSD300 algorithm has model overfitting during training, so some backgrounds will be detected as targets in the test set; Figure 8 (d) shows that the FasterR - CNN algorithm also has a missed detection phenomenon when detecting small targets. According to Table 2 and Figure 8 (e), it can be seen that the method studied in the present invention performs relatively better in terms of accuracy and computational complexity.

[0077] In the comparative experiment using the anchor - free method, the detection results of the algorithm of the present invention were compared with those of FCOS and CenterNet algorithms. As can be seen from Table 2, the CenterNet algorithm has a relatively high recall rate of 0.9039, but it is 0.0124 lower than the algorithm of the present invention; FCOS has a relatively high accuracy rate of 0.8942, but it is 0.0409 lower than the algorithm of the present invention. The accuracy rate and AP of the algorithm of the present invention reach 0.9351 and 0.9589 respectively, both higher than the comparative algorithms. At the same time, the FLOPs of the algorithm studied in the present invention is less than half of that of CenterNet, indicating that the detection efficiency of this algorithm is relatively higher.

[0078] From Figure 9 (d), it can be seen that the algorithm of the present invention can better detect small targets in complex backgrounds. And from Figure 7 (b)'s P - R curve, it can be seen that the area enclosed by the P - R curve of the algorithm of the present invention is the largest, that is to say, the detection accuracy of the algorithm of the present invention is the highest and the performance is the best.

Claims

1. A forest fire detection method based on a lightweight anchor-free detection model, characterized in that, The steps include the following: Step 1: Obtain an infrared forest fire image dataset, annotate the positions and category information of the flames in each image to obtain a flame dataset, preprocess each annotated image, and divide it into a training set, a validation set, and a test set; Step 2: Construct an improved lightweight anchor-free detection model. Use the backbone feature extraction network MobileNetv3g to extract fire features from the input image to obtain three effective feature layers Feature1, Feature2, and Feature3. The improved lightweight anchor-free detection model uses the anchor-free detection model YOLOX as the basic framework. First, replace the backbone network of YOLOX with the lightweight network MobileNetv3; Then, improve the MobileNetv3 network by combining the efficient channel attention network and the GhostNet idea to construct the backbone feature extraction network MobileNetv3g, so as to retain multi-scale feature information while making the model lightweight. Finally, use depthwise separable convolutions to simplify the path aggregation network to further reduce the computational amount of the model, and introduce an adaptive spatial feature fusion module to alleviate the inconsistency between feature maps of different scales. MobileNetv3g is an improvement of MobileNetv3. First, use a one-dimensional convolution to replace the fully connected layer of the squeeze-and-excitation network attention mechanism module, then add a common convolution module before the depthwise separable convolution, and combine the output results of the two to fully obtain semantic information, making the detection results more accurate; Step 4: Perform multi-scale max pooling processing on the effective feature layer Feature3 through the spatial pyramid pooling module SPP; Step 5: Input the effective feature layers Feature1, Feature2, and Feature3 into the enhanced feature extraction network D-PANet to perform enhanced feature fusion on the feature layers from deep to shallow, and then use downsampling to transfer the position information of the shallow layer to the deep feature and fuse it with the semantic information of the deep layer to output the feature layers Level1, Level2, and Level3; Step 6: Input the feature layers Level1, Level2, and Level3 into the adaptive spatial feature fusion module ASFF for adaptive spatial feature fusion, and then use 1×1 convolutions to compress to the original number of channels to output the feature layers ASFF-1, ASFF-2, and ASFF-3; Step 7: Input the feature layers ASFF-1, ASFF-2, and ASFF-3 into the prediction network for multi-scale prediction, and use a decoupled head to decode the predicted target category information and position regression information to obtain the prediction results; Step 8: Perform score sorting and non-maximum suppression screening on the prediction results, screen out the prediction box with the highest score and meeting the confidence level within a certain area, and thus obtain the final fire prediction result.

2. The forest fire detection method based on a lightweight anchor-free detection model according to claim 1, wherein In step 1, LabelImg, a labeling software, is used for labeling. For the flame dataset, data augmentation strategies such as mosaic, random affine transformation, random flipping, and contrast change are adopted, and the image size is adjusted to 640×640 pixels.

3. The forest fire detection method based on a lightweight anchor-free detection model according to claim 1, wherein The MobileNetv3g includes the following components: (4a) ECANet attention mechanism: After global average pooling, ECANet obtains local cross-channel interaction information of each channel and its adjacent k channels through a one-dimensional convolution of size k. ECANet adaptively determines the convolution kernel size k given the number of channels C, specifically: In the formula, in the formula, represents the nearest odd number, and there is a non-linear mapping relationship φ between k and C: , where γ and b are constant coefficients; (4b) Improved depthwise separable convolution structure: The structure of MobileNetv3 is improved by referring to the structure of GhostNet. First, a common 3×3 convolution is used to map the number of channels of the input feature map to half of the original number of channels, then depthwise separable convolution is used to further extract features, and finally, the feature maps output by the two parts are added together; (4c) h-swish activation function: The expression is: In the formula, represents the activation function, that is, the maximum output value of the restricted ReLU is 6.

4. The forest fire detection method based on a lightweight anchor-free detection model according to claim 1, wherein In step 2, MobileNetv3g is divided into six stages, namely Stage1, Stage2, Stage3, Stage4, Stage5, and Stage6. The output feature layers of Stage3, Stage5, and Stage6 are taken as the effective feature layers Feature1, Feature2, and Feature3, and are fed into the enhanced feature extraction network of YOLOX.

5. The forest fire detection method based on a lightweight anchor-free detection model according to claim 1, wherein In step 4, the three effective feature layers Feature1, Feature2, and Feature3 are input into D-PANet to achieve path aggregation. D-PANet uses depthwise separable convolution to replace the 3×3 common convolution in the path aggregation network PANet, thereby reducing the complexity of the model. The specific process is as follows: D-PANet is doped in the three effective feature layers Feature1, Feature2, and Feature3, repeatedly extracts features, upsamples the bottom-layer features, stacks them with the same-dimensional feature layer of the previous layer, downsamples the high-layer features, and stacks them with the same-dimensional feature layer of the next layer to achieve feature fusion, and finally outputs three feature layers Level1, Level2, and Level3.

6. The forest fire detection method based on a lightweight anchor-free detection model according to claim 1, wherein, In step 5, the adaptive spatial feature fusion method of ASFF is as follows: The feature layer Leveln output by PANet is , and The new feature ASFF-l obtained through the ASFF module. Multiply , , by the weight parameters , , respectively and sum them to obtain the fused feature . The fusion process is expressed as: In the formula, represents the output feature map of ASFF at the feature vector, represents the feature map obtained by scaling Leveln to Levell at the feature vector, , , respectively represent the spatial weight parameters from three feature layers to Levell, and through the Softmax function, , and .

7. The forest fire detection method based on a lightweight anchor-free detection model according to claim 1, characterized in that In step 5, ASFF uses upsampling and downsampling operations to compress and adjust the feature layers. For upsampling, first, a 1×1 convolution is used to compress the number of channels to the same as that of Levell, and then interpolation upsampling is used to increase the resolution; for 2-fold downsampling, a convolution of size 3×3 and stride 2 is directly used to modify the size and number of channels of the feature layer; for 4-fold downsampling, a max pooling layer with a stride of 2 is added before the convolutional layer.

8. The forest fire detection method based on a lightweight anchor-free detection model according to claim 1, wherein, In step 6, the prediction network uses the prediction network of YOLOX, which includes a new decoupled head, the anchor-free idea, and the SimOTA dynamic positive sample matching method.

9. The forest fire detection method based on a lightweight anchor-free detection model according to claim 8, characterized in that, The new decoupled head means that the YOLOX network divides the YOLO detection head into a classification part and a regression part. First, a 1×1 convolutional layer is used to reduce the channel size. Then, two parallel branches are used for the classification and regression tasks respectively. Finally, they are integrated together during prediction. The anchor-free idea means that the number of predictions at each position is reduced from 3 to 1, and it can directly predict four values, namely the two offsets to the upper left corner of the grid, the height and width of the predicted bounding box. At the same time, the center position of each object is designated as a positive sample, and a scale range is preset to assign the level of the feature pyramid for the object. The SimOTA dynamic positive sample matching method refers to: First, calculate the pairwise matching degree, which is expressed as the cost or quality of each prediction-ground truth pair, and the ground truth value g i and the predicted value p j The cost between them is calculated as: where λ represents the balance coefficient, and respectively represent the classification loss and regression loss between the true value g i and the predicted value p j ; Then, for the true value g i , select the top k predictions with the smallest cost within the fixed central region as positive samples; Finally, the corresponding grid of the positive prediction is designated as positive, and the remaining grids are negative.

Citation Information

Patent Citations

  • Multi-class forest scene image segmentation method, storage medium and equipment

    CN111062950A

  • Rapid pest detection method based on improved YOLO V4

    CN114220035A