Deep learning lightweight efficient forest grass live wire segmentation method based on unmanned aerial vehicle image
By building a lightweight and efficient forest and grass fire line segmentation model LiFE-BONet, using MobileViT2-Net and depth feature enhancement boundary enhancement module, the problem of low model computing efficiency and insufficient edge detection accuracy during the spread of the drone forest fire is solved, and efficient and accurate live line segmentation is achieved, which is suitable for equipment with limited resources.
Patent Information
- Application Number
- CN202510575183.8
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-05-06
- Publication Date
- 2025-08-12
AI Technical Summary
During the spread of existing drone forest fires, the model calculation efficiency is low, the edge detection accuracy is insufficient, and the semantic segmentation network increases with the increase in depth in complex scenarios leads to information loss and insufficient segmentation accuracy.
Build an improved lightweight and efficient forest and grass fire line segmentation model LiFE-BONet, adopting the lightweight MobileViT2-Net as the feature extraction backbone network, and introducing deep feature extraction enhancement edge and boundary enhancement modules, including deep separable convolution and residual-Sobel Laplace modules, optimizing the network structure to improve segmentation effect.
While reducing the amount of parameters, the accuracy and speed of live line segmentation are improved, the feature extraction capability of the model in complex scenarios is enhanced, and it is suitable for embedded devices with limited resources, real-time live line segmentation during forest fire spread.
Smart Images

Figure CN120472167A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of deep neural network technology, and in particular to a deep learning, lightweight and efficient forest and grass fire line segmentation method based on drone images. Background Art
[0002] Deep learning technology is widely used in forest fire monitoring and can be used to identify and segment forest fires. Compared to traditional machine learning and image processing methods, deep learning technology is more effective in forest fire monitoring. By inputting drone images of fire characteristics in forest fires into a deep learning model, the model is trained to learn these characteristics and ultimately achieve fire extraction. Common deep learning-based recognition and segmentation models include Convolutional Neural Networks (CNNs) and Fully Convolutional Neural Networks (FCNs), which can effectively process large amounts of image data and accurately extract features in complex environments.
[0003] Currently, numerous computer vision-based models and methods are being used to identify forest fires and facilitate timely response. However, these studies focus primarily on the initial stages of a forest fire, with few algorithms specifically examining the spread of fire fronts after a fire has occurred. After a forest fire breaks out, it gradually spreads outward along its front. Efficiently identifying the location and size of a fire front can provide scientific guidance for the allocation of firefighting resources and personnel, effectively curbing the spread of a forest fire and minimizing the loss of forest resources.
[0004] The purpose of this invention is to construct a UAV forest fire spread dataset to train an improved lightweight and efficient semantic segmentation model for real-time segmentation and extraction of fire lines during forest fire spread. The aim is to ensure the lightweight of the model while taking into account the model processing speed and the effect of fire line segmentation, and provide scientific guidance for forest fire rescue. Summary of the Invention
[0005] To address the problems of low model computation efficiency, insufficient edge detection accuracy, and insufficient segmentation accuracy caused by loss of semantic information as the depth of the semantic segmentation network increases in UAV forest and grass fireline segmentation, the present invention provides a lightweight and efficient forest and grass fireline segmentation method based on deep learning of UAV images, including the following steps:
[0006] Step 1: Construct a UAV forest and grass fire line dataset;
[0007] The UAV forest and grass fire line dataset includes: BurnedAreaUAV data, field burning data and public forest fire aerial photography data;
[0008] The UAV forest and grass fire line dataset is feature-annotated using Labelme to construct labels;
[0009] The labels include a background label Ground and a fireline label FireLine;
[0010] Step 2: Build an improved lightweight and efficient forest-grass fireline segmentation model LiFE-BONet;
[0011] The improved lightweight and efficient forest and grass fire line segmentation model is constructed using the classic segmentation network U-Net as the framework;
[0012] The framework includes an encoder for feature extraction and a decoder for feature reconstruction;
[0013] The specific improvements of the improved lightweight and efficient forest and grass fire line segmentation model include: deep feature extraction to enhance edges, boundary enhancement module and model lightweighting;
[0014] Input the UAV forest and grass fire line dataset, perform feature extraction processing through the encoder and the deep feature extraction enhancement edge, and output the feature pixels corresponding to the fire line label FireLine;
[0015] The feature pixels are input into the decoder, and then the boundary enhancement module outputs a feature map of the same size as the input UAV forest and grass fire line dataset;
[0016] Step 3, evaluating the comprehensive performance of the improved lightweight and efficient forest-grass fire line segmentation model LiFE-BONet;
[0017] U-Net and LiFE-BONet were trained and validated on the training and validation sets of the UAV forest and grassland fire dataset using the same training environment and hyperparameters. The predicted intersection-over-union ratio, parameter count, running time and visualization results of U-Net and LiFE-BONet were compared on the test set to evaluate the comprehensive performance of LiFE-BONet.
[0018] Furthermore, in step 1, the label is a binary classification label;
[0019] The background label Ground is used to classify forest and grass fire background pixels as the classification basis;
[0020] The fire line label FireLine is used to distinguish open fire pixels in forest and grass fires based on fire line characteristics.
[0021] Furthermore, in step 2, the specific process of lightweighting the model is as follows:
[0022] Step 2.1: Replace the original U-Net encoder used for feature extraction with the MobileViT2-Net backbone network;
[0023] Step 2.2: Keep the original first convolutional layer structure of U-Net;
[0024] Step 2.3: Replace the ordinary convolution of the downsampling part of the original structure with grouped convolution;
[0025] Furthermore, in step 2, the specific process of extracting and enhancing edges with deep features is as follows:
[0026] In the high-dimensional feature extraction layer of the improved encoder, an auxiliary feature enhancement edge is introduced. This edge starts from the last layer of shallow feature extraction and performs channel sorting through depth-wise separable convolution. During the depth downsampling process, the auxiliary edge is continuously fused with the original features and continuously accumulates semantic features of each depth layer, so that the model retains richer semantic information in the final downsampling stage, allowing the model to better distinguish hot pixels in the encoding stage.
[0027] Furthermore, in step 2, the boundary enhancement module includes two functions: enhancing the feature map boundary individually and constraining the model as a whole.
[0028] Furthermore, the specific process of enhancing the feature map boundary separately is as follows:
[0029] The specific process of enhancing the feature map boundary separately is:
[0030] At the shallow feature output end of the improved decoder, a residual-Sobel Laplace module is integrated. The residual-Sobel Laplace module first performs grayscale processing on the shallow feature map obtained by the improved encoder, and then inputs it into the Sobel operator for boundary contour extraction. The result processed by the Sobel operator is then input into the Laplace operator for further processing. Finally, the output of the Laplace operator is residually connected with the original input to obtain the final boundary enhancement feature map.
[0031] Furthermore, the specific process of constraining the model as a whole is as follows:
[0032] The enhanced boundary prediction features are generated by the residual-Sobel Laplace module, and the binary boundary classification cross entropy loss is calculated with the true boundary contour. The obtained boundary auxiliary loss is weightedly superimposed with the main loss function, and the constraints on the overall prediction results of the model are achieved through joint gradient backpropagation.
[0033] Furthermore, the encoder input image size is 512×512, the number of channels is 3, and the image is downsampled by 32 times.
[0034] Furthermore, the decoder outputs a feature map with a size of 512×512 and a channel number of 2, and then uses a Sigmoid activation function to map each type of pixel value to a probability interval of 0 to 1 to finally obtain a probability distribution feature map of each category.
[0035] Furthermore, in step 3, the training environment and hyperparameters are set as follows: Windows 10 operating system, NVIDIA GeForce GTX 4090 20G graphics card; under the PyTorch framework with Python version 3.11, the torch and cuda versions are torch-2.3.0+cu121; Adm is used as the optimizer, batch = 16, number of iterations epoch = 150, learning rate lr = 0.001, and weight decay weight_decay = 0.001.
[0036] The beneficial effects of the present invention are as follows: (1) The improved lightweight and efficient forest and grass fire line segmentation model LiFE-BONet can segment fire lines based on the relationship between real labels and real images, demonstrating the effectiveness of the model in the process of forest fire spread. By adopting a lightweight feature extraction backbone (Mobilevit2-Net), global convolution replacement, and structural optimization, the method significantly reduces the number of parameters compared to the baseline model (U-Net), highlighting the advantage of lightweight model.
[0037] (2) The improved lightweight and efficient forest-grass fireline segmentation model LiFE-BONet introduces high-dimensional information feature enhancement edge (En-Feature) in the high-dimensional feature extraction layer to supplement the model's lost information due to network depth, enabling the model to learn more feature details to improve the segmentation effect.
[0038] (3) The improved lightweight and efficient forest and grass fire line segmentation model LiFE-BONet integrates the residual-Sobel Laplace (SL) boundary enhancement module in the shallow feature extraction layer to improve the model's ability to segment feature boundaries, making the model's predicted feature boundary contours closer to the real shape.
[0039] (4) The improved lightweight and efficient forest and grass fire line segmentation model LiFE-BONet showed good performance on the UAV forest and grass fire line dataset test set, demonstrating the generalization ability of the method. BRIEF DESCRIPTION OF THE DRAWINGS
[0040] Figure 1 This is a flowchart of a lightweight and efficient forest and grass fire line segmentation method based on deep learning of UAV images;
[0041] Figure 2 This is a structural diagram of the baseline model U-Net;
[0042] Figure 3 This is a schematic diagram of the structure of the improved lightweight and efficient forest-grass fireline segmentation model LiFE-BONet;
[0043] Figure 4 Diagram showing how edges are enhanced for deep features.
[0044] Figure 5 This is the working principle diagram of the residual-Sobel Laplace (SL) boundary enhancement module;
[0045] Figure 6 is the Loss and Miou curve;
[0046] Figure 7 This is a comparison chart of the comprehensive effects of the ablation test;
[0047] Figure 8 A visual comparison of the baseline model and the lightweight and efficient forest-grass fireline segmentation model LiFE-BONet; DETAILED DESCRIPTION
[0048] To make the technical solutions and advantages of the embodiments of the present invention more clearly understood, exemplary embodiments of the present invention are further described in detail below with reference to the accompanying drawings. It should be noted that the embodiments described are only a portion of the embodiments of the present invention, and are not an exhaustive list of all embodiments. It should be noted that the embodiments of the present invention and the features thereof may be combined with each other unless they conflict.
[0049] Example 1, combined Figure 1 This embodiment describes a lightweight and efficient forest and grass fire line segmentation method based on deep learning of drone images. The method is implemented by the following steps:
[0050] Step 1: Construct a UAV forest and grass fire line dataset;
[0051] The UAV forest and grass fire line dataset includes: BurnedAreaUAV data, field burning data and public forest fire aerial photography data;
[0052] The UAV forest and grass fire line dataset is feature-annotated using Labelme to construct labels;
[0053] The labels include a background label Ground and a fireline label FireLine;
[0054] Specifically, the UAV forest and grassland fire front dataset was manually screened to obtain a total of 1,898 real-life UAV images of forest fire fronts from various scenarios and angles. Data annotation was performed in Labeme. After the fire front annotation was completed, the resulting JSON file was converted into a PNG mask format using a Python script. Since this dataset contains images of various sizes, including 3840×2160, 1920×1080, and 1280×720, these sizes were uniformly resized to 512×512 to adapt the data to the model input channels during subsequent training, validation, and testing. Furthermore, to reduce the risk of model overfitting due to the large data size, a series of data augmentation techniques were employed. These included horizontal and vertical flipping (using ImageOps.flip to vertically mirror the image and label); horizontal mirroring (using ImageOps.mirror to horizontally mirror the image and label); 45-degree rotation (rotating the image and the corresponding label by 45 degrees); and random noise addition (adding a random noise value to each pixel, retaining the original label as its characteristics remain unchanged). Contrast adjustment, sharpening, mosaicing, and a 135-degree rotation were also performed, similar to the above. The final data set after data augmentation has 15184 (1898×8) images. Finally, the data set is divided into training set, validation set and test set in a ratio of 6:2:2, resulting in 8107, 2697 and 2697 images respectively;
[0055] Step 2: Build an improved lightweight and efficient forest-grass fireline segmentation model LiFE-BONet;
[0056] The improved lightweight and efficient forest and grass fire line segmentation model is constructed using the classic segmentation network U-Net as the framework;
[0057] The framework includes an encoder for feature extraction and a decoder for feature reconstruction;
[0058] The specific improvements of the improved lightweight and efficient forest and grass fire line segmentation model include: deep feature extraction to enhance edges, boundary enhancement module and model lightweighting;
[0059] Input the UAV forest and grass fire line dataset, perform feature extraction processing through the encoder and the deep feature extraction enhancement edge, and output the feature pixels corresponding to the fire line label FireLine;
[0060] The feature pixels are input into the decoder, and then the boundary enhancement module outputs a feature map of the same size as the input UAV forest and grass fire line dataset;
[0061] Specifically, the improved lightweight and efficient forest and grass fire line segmentation model LiFE-BONet first selects the segmentation model U-Net as the baseline framework, and its network structure is as follows Figure 2 As shown, it is then lightweight optimized and the designed deep feature enhancement edge and boundary enhancement modules are introduced to obtain the final lightweight and efficient forest and grass fire line segmentation model. The final model structure is as follows Figure 3 As shown, the comprehensive performance of the lightweight and efficient forest and grass fire line segmentation model is finally evaluated, and it is compared with the baseline model (U-Net) and module ablation for indicator evaluation (Miou, Precision, Recall, F1, Speed, ChamferDistance, Parameters).
[0062] The baseline model (U-Net) encoder is replaced with the lightweight Mobilevit2-Net. The feature extraction of the backbone feature extraction network Mobilevit2-Net has 6 stages, where each stage represents a different feature extraction layer and each stage contains several MV2Blocks and MobileViTBlocks. In order to adapt to the input channel and size of the backbone extraction network, the original U-Net first layer is retained and named stem_conv. Its function is to retain lower-dimensional boundary contour information while organizing the channel, so that the model can learn more low-dimensional features and reduce the number of model parameters to achieve model lightweight. Since the introduced feature extraction network has one more stage than the original convolutional feature extraction network and retains the stage1 of the original network (that is, the stem_conv of this network), the encoding part of the final model has 7 stages (1 stem_conv and 6 stages).
[0063] Step 3, evaluating the comprehensive performance of the improved lightweight and efficient forest-grass fire line segmentation model LiFE-BONet;
[0064] The same training environment and hyperparameters were used to train U-Net and LiFE-BONet on the training and validation sets of the UAV forest and grassland fire line dataset. The predicted intersection-over-union ratio, parameter count, running time and visualization results of U-Net and LiFE-BONet were compared on the test set to evaluate the comprehensive performance of the lightweight and efficient forest and grassland fire line segmentation model.
[0065] In step 1, the label is a binary classification label;
[0066] The background label Ground is used to classify forest and grass fire background pixels as the classification basis;
[0067] The fire line label FireLine is used to distinguish open fire pixels in forest and grass fires based on fire line characteristics.
[0068] In step 2, the specific process of lightweighting the model is as follows:
[0069] Step 2.1: Replace the original U-Net encoder used for feature extraction with the MobileViT2-Net backbone network;
[0070] Step 2.2: Keep the original first convolutional layer structure of U-Net;
[0071] Step 2.3: Replace the ordinary convolution of the downsampling part of the original structure with grouped convolution;
[0072] In step 2, the specific process of extracting and strengthening edges with deep features is as follows:
[0073] In the high-dimensional feature extraction layer of the improved encoder, an auxiliary feature enhancement edge is introduced. This edge starts from the last layer of shallow feature extraction and performs channel sorting through depth-wise separable convolution. During the depth downsampling process, the auxiliary edge is continuously fused with the original features and continuously accumulates semantic features of each depth layer, so that the model retains richer semantic information in the final downsampling stage, allowing the model to better distinguish hot pixels in the encoding stage.
[0074] Specifically, the working principle of the depth feature enhanced edge (En-Feature) is as follows Figure 4 As shown in the figure, the depth feature enhancement edge (En-Feature) is mainly used to make up for the information loss problem caused by the deepening of the network, especially in the complex background of forest fires. It aims to use auxiliary edges to retain the features of each depth stage (stage3-stage7), and then enhance the main features to improve the segmentation effect. The feature enhancement edge starts from stage2, retains the low-dimensional boundary contour information of the fire line, and then continuously accumulates the feature information of the next stage, and then fuses the features with each depth stage to obtain the depth enhancement feature. During the fusion process, due to the influence of the image size, the depth separable convolution (DW_Conv) is used to process the size of the feature map in order to adapt to the feature size of each depth stage.
[0075] In step 2, the boundary enhancement module includes two functions: enhancing the feature map boundary individually and constraining the model as a whole.
[0076] The specific process of enhancing the feature map boundary separately is:
[0077] At the shallow feature output end of the improved decoder, a residual-Sobel Laplace module is integrated. The residual-Sobel Laplace module first performs grayscale processing on the shallow feature map obtained by the improved encoder, and then inputs it into the Sobel operator for boundary contour extraction. The result processed by the Sobel operator is then input into the Laplace operator for further processing. Finally, the output of the Laplace operator is residually connected with the original input to obtain the final boundary enhancement feature map.
[0078] The specific process of constraining the model as a whole is as follows:
[0079] The enhanced boundary prediction features are generated by the residual-Sobel Laplace module, and the binary boundary classification cross entropy loss is calculated with the true boundary contour. The obtained boundary auxiliary loss is weightedly superimposed with the main loss function, and the constraints on the overall prediction results of the model are achieved through joint gradient backpropagation.
[0080] The working principle of the boundary enhancement module is as follows Figure 5 As shown, the boundary enhancement module (SL) is a cascade of MaxPooling, Sobel and Laplace operators. The Sobel operator is a first-order partial derivative, which can effectively detect the edges of the image and provide structural information in the image. This helps the model identify the contours of objects, thereby improving the accuracy of segmentation. The Laplace operator is a second-order partial derivative, which can capture higher-order features and emphasize regional changes, which is particularly important for distinguishing objects of different categories. It can better handle noise and details and help the model learn more complex boundaries; the boundary contour information extracted from the shallow features is highlighted by MaxPooling and sent to the subsequent two operators for operation, which can provide richer edge and structural information in the boundary feature extraction stage, thereby improving the performance of the semantic segmentation model. The operation process is illustrated in (a) in the figure. The module first performs grayscale processing on the input feature map, and then convolves it with the Gx and Gy masks in (b) to obtain the horizontal boundary feature S_X and the vertical boundary feature S_Y. Finally, Gs is obtained through the formula in (c) in the figure. This is the operation process of the Sobel operator. The subsequent Laplace operator operation is similar. It only needs to replace the mask. Finally, the original boundary contour is added to its output through the residual to obtain the final enhanced boundary contour feature map.
[0081] The encoder input image size is 512×512, the number of channels is 3, and it is downsampled by 32 times.
[0082] The decoder outputs a feature map with a size of 512×512 and a channel number of 2, and then uses a Sigmoid activation function to map each type of pixel value to a probability interval of 0 to 1 to finally obtain a probability distribution feature map of each category.
[0083] In step 3, the training environment and hyperparameters are set as follows: Windows 10 operating system, NVIDIA GeForce GTX 4090 20G graphics card; Python version 3.11 under the PyTorch framework, where the torch and cuda versions are torch-2.3.0+cu121; Adm is used as the optimizer, batch size is 16, number of iterations is 150, learning rate is lr is 0.001, and weight decay is weight_decay is 0.001.
[0084] Specifically, after building a lightweight and efficient forest and grass fire line segmentation model, we use the training set for training and the validation set for verification, and finally get the loss curve and Miou curve, as shown below: Figure 6 As shown in the figure, the results show that the model can learn the fireline features. With the increase of epochs, the loss gradually converges without overfitting. Miou gradually increases to about 0.82 and tends to be stable. The model can also be generalized to the validation set, proving the effectiveness of the model.
[0085] pass Figure 7It can be seen that according to the comprehensive index comparison results with the baseline model and module ablation on the test set, by replacing the lightweight and efficient backbone feature extraction network Mobilevit2-Net and optimizing the upsampling part and convolution, the model parameter volume is reduced from 31.04M to 1.39M and other indicators have no obvious changes except Speed and Chamfer Distance, indicating that the lightweight model has a good effect on fire line segmentation, but is lacking in processing speed and boundary contour processing; with the introduction of deep feature enhanced edge (En-Feature), its Miou, Precision, F1 have improved, and Chamfer Distance has increased. The distance decreases, indicating a significant improvement in the model's performance on the segmentation task. However, the processing speed and number of parameters increase due to the presence of a depthwise separable convolution (DW_Conv) on this edge, which increases the number of parameters by 0.29. This increased computational effort also slows down the processing speed, but overall, good segmentation results are achieved at the expense of small parameter and processing speed. Finally, the addition of the boundary enhancement module (SL) improves all metrics without increasing the number of parameters, but still slows down the processing speed, which has no significant impact on real-time performance. Overall, the model achieves significant improvements in multiple performance metrics, such as M-IoU, Precision, Recall, and F1-Score. This demonstrates that the proposed lightweight optimization, deep feature enhancement edge (En-Feature), and boundary enhancement module (SL) play a significant role in improving model performance. Despite the increase in inference speed, the optimized model maintains good real-time processing capabilities, meeting the requirements of most real-time applications. The significant reduction in the number of parameters improves the model's deployment efficiency, making it suitable for a wider range of application scenarios, especially in embedded devices or resource-constrained environments.
[0086] The comparison of visualization results before and after improvement is shown in the figure. Figure 8 It can be seen that the improved segmentation results are closer to the true label in terms of both external morphological contours and internal details. Although the baseline model can also accurately segment the fire line, it still lacks internal details and external contour shapes.
[0087] A comprehensive evaluation of the lightweight and efficient forest and grass fire line segmentation model LiFE-BONet described in this embodiment shows that the model is effective in extracting fire lines in real time during the spread of forest fires. Compared with traditional segmentation models, it can more accurately segment and identify the position and size of fire lines while maintaining a low number of parameters, effectively reducing the demand for computing resources and making its application more flexible.
[0088] The technical features of the above-mentioned embodiments can be combined arbitrarily. In order to make the description concise, not all possible combinations of the technical features in the above-mentioned embodiments are described. However, as long as there is no contradiction in the combination of these technical features, they should be considered to be within the scope of this specification.
[0089] The above-described embodiments merely illustrate several implementations of the present invention, and while their descriptions are relatively specific and detailed, they should not be construed as limiting the scope of the patent. It should be noted that a person skilled in the art would be able to make numerous variations and improvements without departing from the spirit of the present invention, all of which fall within the scope of protection of the present invention. Therefore, the scope of protection of the patent for this invention shall be determined by the appended claims.
Claims
1. A lightweight and efficient forest and grass fire line segmentation method based on deep learning of UAV images, characterized by: The following steps are involved: Step 1: Construct a UAV forest and grass fire line dataset; The UAV forest and grass fire line dataset includes: BurnedAreaUAV data, field burning data and public forest fire aerial photography data; The UAV forest and grass fire line dataset is feature-annotated using Labelme to construct labels; The labels include a background label Ground and a fireline label FireLine; Step 2: Build an improved lightweight and efficient forest-grass fireline segmentation model LiFE-BONet; The improved lightweight and efficient forest and grass fire line segmentation model is constructed using the classic segmentation network U-Net as the framework; The framework includes an encoder for feature extraction and a decoder for feature reconstruction; The specific improvements of the improved lightweight and efficient forest and grass fire line segmentation model include: deep feature extraction to enhance edges, boundary enhancement module and model lightweighting; Input the UAV forest and grass fire line dataset, perform feature extraction processing through the encoder and the deep feature extraction enhancement edge, and output the feature pixels corresponding to the fire line label FireLine; The feature pixels are input into the decoder, and then the boundary enhancement module outputs a feature map of the same size as the input UAV forest and grass fire line dataset; Step 3, evaluating the comprehensive performance of the improved lightweight and efficient forest-grass fire line segmentation model LiFE-BONet; U-Net and LiFE-BONet were trained and validated on the training and validation sets of the UAV forest and grassland fire dataset using the same training environment and hyperparameters. The predicted intersection-over-union ratio, parameter count, running time and visualization results of U-Net and LiFE-BONet were compared on the test set to evaluate the comprehensive performance of LiFE-BONet.
2. The method for light-weight and efficient forest and grass fire line segmentation based on deep learning of drone images according to claim 1 is characterized in that: In step 1, the label is a binary classification label; The background label Ground is used to classify forest and grass fire background pixels as a classification basis; The fire line label FireLine is used to distinguish open fire pixels in forest and grass fires based on fire line characteristics.
3. The method for light-weight and efficient forest and grass fire line segmentation based on deep learning of drone images according to claim 1 is characterized in that: In step 2, the specific process of lightweighting the model is as follows: Step 2.1: Replace the original U-Net encoder used for feature extraction with the MobileViT2-Net backbone network; Step 2.2: Keep the original first convolutional layer structure of U-Net; Step 2.3: Replace the ordinary convolution of the downsampling part of the original structure with grouped convolution; Step 2.4: Replace the deconvolution of the upsampling part of the original structure with bilinear interpolation.
4. The method for light-weight and efficient forest and grass fire line segmentation based on deep learning of drone images according to claim 1 is characterized in that: In step 2, the specific process of extracting and enhancing edges with deep features is as follows: In the high-dimensional feature extraction layer of the improved encoder, an auxiliary feature enhancement edge is introduced. This edge starts from the last layer of shallow feature extraction and performs channel sorting through depth-wise separable convolution. During the depth downsampling process, the auxiliary edge is continuously fused with the original features and continuously accumulates semantic features of each depth layer, so that the model retains richer semantic information in the final downsampling stage, allowing the model to better distinguish hot pixels in the encoding stage.
5. The method for light-weight and efficient forest and grass fire line segmentation based on deep learning of drone images according to claim 1 is characterized in that: In step 2, the boundary enhancement module includes two functions: enhancing the feature map boundary individually and constraining the model as a whole.
6. The method for light-weight and efficient forest and grass fire line segmentation based on deep learning of drone images according to claim 5 is characterized in that: The specific process of enhancing the feature map boundary separately is: At the shallow feature output end of the improved decoder, a residual-Sobel Laplace module is integrated. The residual-Sobel Laplace module first performs grayscale processing on the shallow feature map obtained by the improved encoder, and then inputs it into the Sobel operator for boundary contour extraction. The result processed by the Sobel operator is then input into the Laplace operator for further processing. Finally, the output of the Laplace operator is residually connected with the original input to obtain the final boundary enhancement feature map.
7. The method for light-weight and efficient forest and grass fire line segmentation based on deep learning of drone images according to claim 5 is characterized in that: The specific process of constraining the model as a whole is as follows: The enhanced boundary prediction features are generated by the residual-Sobel Laplace module, and the binary boundary classification cross entropy loss is calculated with the true boundary contour. The obtained boundary auxiliary loss is weightedly superimposed with the main loss function, and the constraints on the overall prediction results of the model are achieved through joint gradient backpropagation.
8. The method for light-weight and efficient forest and grass fire line segmentation based on deep learning of drone images according to claim 1 is characterized in that: The encoder input image size is 512×512, the number of channels is 3, and it is downsampled by 32 times.
9. The method for light-weight and efficient forest and grass fire line segmentation based on deep learning of drone images according to claim 1 is characterized in that: The decoder outputs a feature map with a size of 512×512 and a channel number of 2, and then uses a Sigmoid activation function to map each type of pixel value to a probability interval of 0 to 1 to finally obtain a probability distribution feature map of each category.
10. The method for light-weight and efficient forest and grass fire line segmentation based on deep learning of drone images according to claim 1 is characterized in that: In step 3, the training environment and hyperparameters are set as follows: Windows 10 operating system, NVIDIA GeForce GTX 4090 20G graphics card; Python version 3.11 under the PyTorch framework, where the torch and cuda versions are torch-2.3.0+cu121; Adm is used as the optimizer, batch size is 16, number of iterations is 150, learning rate is lr is 0.001, and weight decay is weight_decay is 0.001.