A remote sensing vegetation extraction method and device based on deep learning semantic segmentation

By using deep learning semantic segmentation methods, combined with encoder-decoder structures and multi-scale modules, a semantic segmentation network was constructed, which solved the problems of accuracy and efficiency in vegetation extraction from remote sensing images, and achieved efficient vegetation information extraction and evaluation.

CN116385875BActive Publication Date: 2026-04-14宁夏回族自治区遥感调查院(高分辨率对地观测系统宁夏数据与应用中心)
View PDF 1 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2023-03-24
Publication Date
2026-04-14

AI Technical Summary

Technical Problem

Existing technologies lack accurate and efficient methods for extracting vegetation from remote sensing images based on deep learning semantic segmentation, resulting in manual visual interpretation being labor-intensive and resource-intensive, and failing to meet time and accuracy requirements.

Method used

A semantic segmentation method based on deep learning is adopted. By using an encoder-decoder structure and an attention fusion edge detection model, combined with multi-scale modules and various loss functions, a semantic segmentation network is constructed to accurately extract farmland, grassland, forest land and orchard. The random forest algorithm is used for multi-source data interpretation.

Benefits of technology

It achieves efficient and accurate vegetation extraction from remote sensing images, reduces the need for manual interpretation, improves extraction efficiency and accuracy, and can generate machine interpretation results and perform vectorized evaluation.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN116385875B_ABST
    Figure CN116385875B_ABST
Patent Text Reader

Abstract

The present application relates to the technical field of remote sensing image recognition, in particular to a remote sensing vegetation extraction method and device based on deep learning semantic segmentation. The remote sensing vegetation extraction method based on deep learning semantic segmentation comprises the following steps: performing data processing on original remote sensing images collected by Gaofen-2 and Landsat8 satellites to obtain training data and semantic labels; improving an encoder-decoder structure to establish a vegetation extraction model to be trained; training the vegetation extraction model to be trained by using the training data and the semantic labels to obtain a vegetation extraction model; inputting remote sensing images to be extracted into the vegetation extraction model to obtain vegetation extraction information; and performing multi-source data interpretation on the vegetation extraction information to obtain vegetation information. The present application is a precise and efficient remote sensing image vegetation extraction method based on deep learning semantic segmentation.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of remote sensing image recognition technology, and in particular to a method and apparatus for extracting remote sensing vegetation based on deep learning semantic segmentation. Background Technology

[0002] Vegetation extraction from remote sensing images involves labeling different vegetation types based on their similarities and differences in characteristics. Effective and timely classification and extraction of vegetation is of great significance in natural resource management and ecological environmental protection.

[0003] Vegetation extraction enables the analysis of vegetation conditions in a region, providing crucial data support for land surveys, forest censuses, and urban planning. Due to the high-resolution, multispectral nature of remote sensing imagery, vegetation extraction is an extremely challenging task. Currently, vegetation extraction from remote sensing images primarily relies on manual visual interpretation. This method is resource-intensive, requiring interpreters with specialized knowledge and experience for annotation, and sometimes necessitates field surveys to extract and determine targets. Consequently, it struggles to meet the time and accuracy requirements of vegetation classification tasks.

[0004] In recent years, there has been an increasing number of cases using semantic segmentation to solve remote sensing tasks. Convolutional neural networks can effectively extract objects from images, greatly improving image recognition speed and efficiency. Due to the rapid development of deep learning research, significant progress has been made in many areas. Pixel-based image segmentation algorithms employ deep learning techniques, specifically semantic segmentation. In the field of remote sensing, the use of semantic segmentation for interpretation has become a research hotspot.

[0005] In the existing technology, there is a lack of a precise and efficient method for extracting vegetation from remote sensing images based on deep learning semantic segmentation. Summary of the Invention

[0006] This invention provides a method and apparatus for remote sensing vegetation extraction based on deep learning semantic segmentation. The technical solution is as follows:

[0007] On the one hand, a remote sensing vegetation extraction method based on deep learning semantic segmentation is provided. This method is implemented by an electronic device and includes:

[0008] Data processing was performed on the raw remote sensing images acquired by Gaofen-2 and Landsat-8 satellites to obtain training data and semantic labels;

[0009] An improved encoder-decoder structure was used to establish a vegetation extraction model to be trained.

[0010] The vegetation extraction model is trained using the training data and the semantic labels to obtain the vegetation extraction model.

[0011] Input the remote sensing image to be extracted into the vegetation extraction model to obtain vegetation extraction information;

[0012] The extracted vegetation information is analyzed using multi-source data to obtain vegetation information.

[0013] Optionally, the data processing of the raw remote sensing images acquired by Gaofen-2 and Landsat-8 satellites to obtain training data and semantic labels includes:

[0014] The original remote sensing images acquired by Gaofen-2 and Landsat-8 satellites were manually annotated to obtain semantic tags; the semantic tags include cultivated land, grassland, forest land and orchard.

[0015] The normalized vegetation index (NDI) is calculated based on the original remote sensing images acquired by Gaofen-2 and Landsat-8 satellites and the semantic tags; the NDI is used as a spectral feature.

[0016] Based on the original remote sensing images acquired by Gaofen-2 and Landsat-8 satellites, image homogeneity, image mean, and image correlation are calculated; these image homogeneity, image mean, and image correlation are then used as texture features.

[0017] The spectral features and the texture features are added to the original remote sensing image to obtain fused data;

[0018] Data augmentation is performed on the fused data to obtain training data.

[0019] The vegetation extraction model to be trained includes a vegetation extraction sub-model for cultivated land and grassland and a vegetation extraction sub-model for forest and orchard.

[0020] The sub-model for extracting vegetation from cultivated land and grassland is used to extract vegetation from training data with semantic labels of cultivated land and grassland. The sub-model for extracting vegetation from cultivated land and grassland is constructed based on an encoder-decoder structure and combined with an attention fusion edge detection model.

[0021] The sub-model for extracting vegetation from forest and orchard land is used to extract vegetation from training data with semantic labels of forest and orchard. The sub-model for extracting vegetation from forest and orchard land is constructed based on an encoder-decoder structure, combined with multi-scale modules and dilated convolutional pyramid pooling modules.

[0022] Optionally, training the vegetation extraction model using the training data and the semantic labels to obtain the vegetation extraction model includes:

[0023] Based on the semantic labels, the training data is divided into cultivated land and grassland training data and forest and orchard training data;

[0024] The cultivated land and grassland training data are input into the vegetation extraction model for training to obtain the cultivated land and grassland loss function; when the cultivated land and grassland loss function converges, the cultivated land and grassland vegetation extraction sub-model is obtained.

[0025] The woodland and orchard training data are input into the vegetation extraction model for training to obtain the woodland and orchard loss function; when the woodland and orchard loss function converges, the woodland and orchard vegetation extraction sub-model is obtained.

[0026] A vegetation extraction model is obtained based on the cultivated land and grassland vegetation extraction sub-model and the forest and orchard output vegetation extraction sub-model.

[0027] The farmland and grassland loss function is obtained by weighting the binary cross-entropy loss function, the IOU loss function, and the structural loss function.

[0028] The woodland and orchard loss function is obtained by linear combination of the Lovasz loss function and the binary cross-entropy loss function.

[0029] Optionally, the step of performing multi-source data analysis on the extracted vegetation information to obtain vegetation information includes:

[0030] A multi-source data vector map is obtained by overlaying the preset multi-source heterogeneous layer data and the extracted vegetation information.

[0031] The multi-source data vector map is rasterized to obtain a multi-source raster vector map.

[0032] Based on the multi-source raster vector map, vegetation information is obtained through interpretation using the random forest algorithm.

[0033] On the other hand, a remote sensing vegetation extraction device based on deep learning semantic segmentation is provided. This device is applied to a remote sensing vegetation extraction method based on deep learning semantic segmentation. The device includes:

[0034] The data acquisition module is used to process raw remote sensing images acquired by Gaofen-2 and Landsat-8 satellites to obtain training data and semantic labels.

[0035] The model building module is used to build a vegetation extraction model to be trained based on the improved encoder-decoder structure.

[0036] The model training module is used to train the vegetation extraction model to be trained using the training data and the semantic labels to obtain the vegetation extraction model.

[0037] The information extraction module is used to input the remote sensing image to be extracted into the vegetation extraction model to obtain vegetation extraction information;

[0038] The data interpretation module is used to perform multi-source data interpretation on the vegetation extraction information to obtain vegetation information.

[0039] Optionally, the data acquisition module is further configured to:

[0040] The original remote sensing images acquired by Gaofen-2 and Landsat-8 satellites were manually annotated to obtain semantic tags; the semantic tags include cultivated land, grassland, forest land and orchard.

[0041] The normalized vegetation index (NDI) is calculated based on the original remote sensing images acquired by Gaofen-2 and Landsat-8 satellites and the semantic tags; the NDI is used as a spectral feature.

[0042] Based on the original remote sensing images acquired by Gaofen-2 and Landsat-8 satellites, image homogeneity, image mean, and image correlation are calculated; these image homogeneity, image mean, and image correlation are then used as texture features.

[0043] The spectral features and the texture features are added to the original remote sensing image to obtain fused data;

[0044] Data augmentation is performed on the fused data to obtain training data.

[0045] The vegetation extraction model to be trained includes a vegetation extraction sub-model for cultivated land and grassland and a vegetation extraction sub-model for forest and orchard.

[0046] The sub-model for extracting vegetation from cultivated land and grassland is used to extract vegetation from training data with semantic labels of cultivated land and grassland. The sub-model for extracting vegetation from cultivated land and grassland is constructed based on an encoder-decoder structure and combined with an attention fusion edge detection model.

[0047] The sub-model for extracting vegetation from forest and orchard land is used to extract vegetation from training data with semantic labels of forest and orchard. The sub-model for extracting vegetation from forest and orchard land is constructed based on an encoder-decoder structure, combined with multi-scale modules and dilated convolutional pyramid pooling modules.

[0048] Optionally, the model training module is further configured to:

[0049] Based on the semantic labels, the training data is divided into cultivated land and grassland training data and forest and orchard training data;

[0050] The cultivated land and grassland training data are input into the vegetation extraction model for training to obtain the cultivated land and grassland loss function; when the cultivated land and grassland loss function converges, the cultivated land and grassland vegetation extraction sub-model is obtained.

[0051] The woodland and orchard training data are input into the vegetation extraction model for training to obtain the woodland and orchard loss function; when the woodland and orchard loss function converges, the woodland and orchard vegetation extraction sub-model is obtained.

[0052] A vegetation extraction model is obtained based on the cultivated land and grassland vegetation extraction sub-model and the forest and orchard output vegetation extraction sub-model.

[0053] The farmland and grassland loss function is obtained by weighting the binary cross-entropy loss function, the IOU loss function, and the structural loss function.

[0054] The woodland and orchard loss function is obtained by linear combination of the Lovasz loss function and the binary cross-entropy loss function.

[0055] Optionally, the data interpretation module is further configured to:

[0056] A multi-source data vector map is obtained by overlaying the preset multi-source heterogeneous layer data and the extracted vegetation information.

[0057] The multi-source data vector map is rasterized to obtain a multi-source raster vector map.

[0058] Based on the multi-source raster vector map, vegetation information is obtained through interpretation using the random forest algorithm.

[0059] On the other hand, an electronic device is provided, comprising a processor and a memory, wherein the memory stores at least one instruction, which is loaded and executed by the processor to implement the aforementioned remote sensing vegetation extraction method based on deep learning semantic segmentation.

[0060] On the other hand, a computer-readable storage medium is provided, wherein at least one instruction is stored in the storage medium, the at least one instruction being loaded and executed by a processor to implement the above-described remote sensing vegetation extraction method based on deep learning semantic segmentation.

[0061] The beneficial effects of the technical solutions provided in the embodiments of the present invention include at least the following:

[0062] This invention proposes a remote sensing vegetation extraction method based on deep learning semantic segmentation. It employs a deep learning algorithm, an encoder-decoder architecture, and integrates an attention mechanism module with edge detection technology to construct a semantic segmentation network based on attention fusion and edge detection for farmland and grassland extraction. A multi-scale module is added to the encoder-decoder for woodland and orchard extraction. Multiple loss functions are used to assist model training, and the trained model is used to test vegetation extraction on a test area image. The machine interpretation results are obtained and vectorized. A random forest algorithm combined with multiple vector feature layers is used for further interpretation of vegetation distribution, and the results are evaluated in conjunction with the actual vegetation distribution. This invention provides a precise and efficient remote sensing image vegetation extraction method based on deep learning semantic segmentation. Attached Figure Description

[0063] To more clearly illustrate the technical solutions in the embodiments of the present invention, the accompanying drawings used in the description of the embodiments will be briefly introduced below. Obviously, the accompanying drawings described below are only some embodiments of the present invention. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.

[0064] Figure 1 This is a flowchart of a remote sensing vegetation extraction method based on deep learning semantic segmentation provided in an embodiment of the present invention;

[0065] Figure 2 This is a block diagram of a remote sensing vegetation extraction device based on deep learning semantic segmentation provided in an embodiment of the present invention;

[0066] Figure 3 This is a schematic diagram of the structure of an electronic device provided in an embodiment of the present invention. Detailed Implementation

[0067] To make the technical problems, technical solutions and advantages of the present invention clearer, a detailed description will be given below in conjunction with the accompanying drawings and specific embodiments.

[0068] This invention provides a remote sensing vegetation extraction method based on deep learning semantic segmentation. This method can be implemented by an electronic device, which can be a terminal or a server. Figure 1 The flowchart shown is a remote sensing vegetation extraction method based on deep learning semantic segmentation. The processing flow of this method may include the following steps:

[0069] S1. Process the raw remote sensing images acquired by Gaofen-2 and Landsat-8 satellites to obtain training data and semantic labels.

[0070] Optionally, the raw remote sensing images acquired by Gaofen-2 and Landsat-8 satellites are processed to obtain training data and semantic labels, including:

[0071] The original remote sensing images acquired by Gaofen-2 and Landsat-8 satellites were manually annotated to obtain semantic tags; the semantic tags include cultivated land, grassland, forest land and orchard.

[0072] The normalized vegetation index (NDI) was calculated based on the original remote sensing images and semantic tags collected by Gaofen-2 and Landsat-8 satellites; the NDI was used as a spectral feature.

[0073] Based on the original remote sensing images acquired by Gaofen-2 and Landsat-8 satellites, image homogeneity, image mean, and image correlation were calculated; these three factors were then used as texture features.

[0074] Spectral and texture features are added to the original remote sensing image to obtain fused data;

[0075] Data augmentation is performed on the fused data to obtain training data.

[0076] In one feasible implementation, data feature processing is performed on remote sensing images acquired from Gaofen-2 and Landsat-8. The directly acquired data features include a 4-channel image consisting of three visible light channels and a single near-infrared channel superimposed. Based on spectral analysis of the processed images, the Normalized Digital Vegetation Index (NDVI) is calculated as a vegetation spectral feature. The calculation formula is shown in the following formula (1):

[0077]

[0078] Wherein, NIR and R are the near-infrared and visible red band numbers, respectively.

[0079] The texture feature extraction process selected correlation, mean, and homogeneity as texture features for different texture statistics of cultivated land, grassland, woodland, and orchard to distinguish different vegetation types.

[0080] The spectral features and texture features are used as independent input channels, and combined with the original 4-channel image, a total of 6 channels are fused to obtain fused data.

[0081] The final fused data has a total of 6 channels of source data; the labels are marked according to the number of vegetation categories, that is, 4 labels are marked according to cultivated land, grassland, forest land and orchard, with label values ​​of 0 and 1, 0 for background and 1 for vegetation objects to be extracted.

[0082] An overlapping sliding window was used for image cropping to reduce loss from cropping edges and increase sample size; data augmentation was performed using random inversion and Gaussian noise. Training data was constructed by randomly sampling data from the augmented data. The sliding window size was 1024×1024, the overlap rate was set to 0.2, and the step size was calculated based on the overlap rate and cropping size.

[0083] The training data processed in the above steps is divided into a training set, a validation set, and a test set; the training set and validation set are used for model training, and the test set is used for model performance evaluation.

[0084] S2. Based on the encoder-decoder structure, an improved vegetation extraction model is established to be trained.

[0085] Among them, the vegetation extraction model to be trained includes a vegetation extraction sub-model for cultivated land and grassland and a vegetation extraction sub-model for forest land and orchard.

[0086] Among them, the sub-model for extracting vegetation from cultivated land and grassland is used to extract vegetation from training data with semantic labels of cultivated land and grassland; the sub-model for extracting vegetation from cultivated land and grassland is constructed based on an encoder-decoder structure and combined with an attention fusion edge detection model.

[0087] In one feasible implementation, the attention fusion edge detection model proposed for cultivated land and grassland vegetation objects is implemented as follows:

[0088] Residual Network 34 (ResNet34) is used as the backbone network in the encoding stage. The encoder stage is divided into 6 parts. The first part is a convolutional layer with 6 input channels and 64 output channels, followed by batch normalization and ReLU activation function. The second to fifth parts are the same as the first to fourth layers of ResNet34, with downsampling performed once in each layer. The sixth part is a non-supplementary max pooling stage with 512 input and 512 output channels, which goes through 5 downsampling processes. The final output feature map resolution is 1 / 32 of the original sample data. A coordinated attention mechanism is embedded in each part of the encoding block. The channel attention mechanism is decomposed into two directions, image height and width, and then merged to coordinate the spatial localization ability of the prediction result and the capture of long-range dependencies.

[0089] In the BasicBlock module of the ResNet34 encoder during the encoding phase, a coordinated attention mechanism module is incorporated; its input is x. c The height and width are H and W respectively, x c Let represent the feature map of the c-th channel in the model. The coordinated attention mechanism decomposes the channel attention into two directions: image height and image width, and then merges the feature maps generated in the two directions. It takes into account both accurate localization in one direction and long-range dependency in the other direction. The mathematical expressions of the output in the two directions are shown in equations (2) and (3) below:

[0090]

[0091]

[0092] in, This represents the output at height h located at the c-th channel, i.e., the attention along the height direction; This represents the output with width w located at the c-th channel, i.e., the attention along the width direction.

[0093] In the decoding stage, the feature map output from the downsampling stage is acquired. In each decoding block of the upsampling stage, two parallel operations are performed: the input X is upsampled at different scales, Y1 is upsampled at twice the resolution using bilinear interpolation, and Y2 is directly output via two dilated convolutions and one regular convolution, then processed by the sigmoid function. These outputs serve as edge outputs in the entire network and are used in the loss function calculation along with the ground truth labels. A total of six edge outputs are generated during the entire network training process. These six edge outputs are incorporated into the loss function calculation to enable multi-scale and multi-level training of the network structure.

[0094] Among them, the training sub-model for extracting vegetation from forest land and orchard land is used to extract vegetation from training data with semantic labels of forest land and orchard land. The training sub-model for extracting vegetation from forest land and orchard land is constructed based on an encoder-decoder structure, combined with multi-scale modules and dilated convolutional pyramid pooling modules.

[0095] In one feasible implementation, the multi-scale module proposed for woodland and orchard vegetation objects is implemented as follows:

[0096] A feature pyramid structure is constructed and a channel attention mechanism is integrated to assist in feature extraction. The connection layer between the encoder and decoder is replaced with a dilated convolutional pyramid pooling module to improve the overall receptive field of the network while taking into account the learning and extraction of detailed information.

[0097] The decoder module also uses the same edge outputs as the decoder stage of the cultivated land and grassland model, and adds the edge outputs of a total of 6 parts to the loss function calculation; the loss function adopts a linear combination of Lovasz loss and binary cross-entropy loss to solve the problem of imbalance between foreground and background distribution.

[0098] S3. Using training data and semantic labels, train the vegetation extraction model to obtain the vegetation extraction model.

[0099] Optionally, the vegetation extraction model to be trained is trained using training data and semantic labels to obtain the vegetation extraction model, including:

[0100] Based on semantic labels, the training data is divided into cultivated land and grassland training data and forest and orchard training data;

[0101] The training data of cultivated land and grassland is input into the vegetation extraction model for training to obtain the cultivated land and grassland loss function; when the cultivated land and grassland loss function converges, the cultivated land and grassland vegetation extraction sub-model is obtained.

[0102] The training data of woodland and orchard is input into the vegetation extraction model for training to obtain the woodland and orchard loss function; when the woodland and orchard loss function converges, the woodland and orchard vegetation extraction sub-model is obtained.

[0103] Based on the vegetation extraction sub-models for cultivated land and grassland and forest and orchard, a vegetation extraction model is obtained.

[0104] In one feasible implementation, the two vegetation extraction models established in the above steps are independent of each other and are trained using cultivated land / grassland / woodland / orchard labels respectively; the training set composed of a portion of the dataset is used for training, and the weights of the loss function are adjusted.

[0105] The training process uses the Adam optimizer with an initial learning rate of 0.001 and smoothing constants of 0.9 and 0.999. The learning rate is dynamically adjusted in multiple rounds of learning using a multi-interval adjustment method.

[0106] The farmland and grassland loss function is obtained by weighting the binary cross-entropy loss function, IOU loss function, and structural loss function together.

[0107] In one feasible implementation, the vegetation extraction model for cultivated land and grassland is trained, and the specific implementation method is as follows:

[0108] An edge output is added after the convolutional layer. Multiple edge outputs of different depths and sizes are fused together using a 1x1 convolutional kernel and a bilinear interpolation process, and used as the basis for calculating the loss function.

[0109] The loss function combines the binary cross-entropy loss (BCELoss), the soft intersection over union loss (Soft-IOULoss), and the structural similarity loss (SSIMLoss). By training on the edge outputs of all levels, a multi-scale and multi-level learning process is achieved. The mathematical calculation formulas are shown in equations (4) and (5) below:

[0110]

[0111] Where y is a binary label 0 or 1, and p(y) is the probability that the output belongs to the label y.

[0112]

[0113] IOULoss measures the quality of a predicted bounding box by calculating the ratio of the intersection and union of the predicted bounding box and the ground truth bounding box, i.e., IOU, where r and c represent pixel coordinates, G is the ground truth value, and S is the predicted value.

[0114] The structural similarity index includes three aspects: illumination l, contrast c, and structure s. The relevant formulas for calculating the loss function are shown in equations (6), (7), (8), (9), and (10) below:

[0115]

[0116]

[0117]

[0118]

[0119] Loss ssim =1-SSIM(x,y) (10)

[0120] Where μ and σ represent the mean and variance of the images x and y, respectively.

[0121] The loss function for cultivated land and grassland is a linear weighted sum of cross-entropy, structural similarity, and IOU loss, where α and β are the weighting coefficients, and its calculation formula is shown in equation (11) below:

[0122] Loss = Loss BCE +αLoss SSIM +βLoss IOU (11)

[0123] The woodland and orchard loss function is obtained by linear combination of the Lovasz loss function and the binary cross-entropy loss function.

[0124] In one feasible implementation, the dilated convolutional pyramid pooling module consists of five parts: a 1×1 convolutional kernel, three 3×3 dilated convolutional kernels with different dilation ratios, and dilated convolutional pooling. The dilated convolutional pooling is an adaptive pooling method that compresses the feature maps of each channel to 1×1 and connects them to the 1×1 convolutional kernel. The outputs of each part are connected and then convolved by a 1×1 convolution to output as a multi-scale feature module.

[0125] Lovasz loss is designed for IOU optimization and is a Lovasz extension of Jaccard loss. The relevant calculation formulas are shown in equations (12), (13), (14), and (15) below:

[0126]

[0127]

[0128]

[0129]

[0130] in, It is an extension of Lovasz for IoU loss, where F represents the value of the scoring function output for the i-th pixel.

[0131] S4. Input the remote sensing image to be extracted into the vegetation extraction model to obtain vegetation extraction information.

[0132] In one feasible implementation, the present invention mainly establishes a relatively complete vegetation extraction dataset based on Gaofen-2 and Landsat-8 multi-band remote sensing satellite image data of Ningxia region, a semantic segmentation network based on attention mechanism and edge detection designed for cultivated land and grassland, a semantic segmentation network based on multi-scale modules designed for orchards and woodlands, a comprehensive interpretation process using random forest method combined with multiple vector layers, and a final verification and evaluation based on the actual distribution.

[0133] Input the remote sensing image to be extracted into the vegetation extraction model established in the above steps to extract vegetation information.

[0134] S5. Perform multi-source data interpretation on the extracted vegetation information to obtain vegetation information.

[0135] Optionally, the vegetation extraction information is subjected to multi-source data interpretation to obtain vegetation information, including:

[0136] Vector overlay is performed based on preset multi-source heterogeneous layer data and vegetation extraction information to obtain a multi-source data vector map;

[0137] Rasterize the multi-source data vector image to obtain a multi-source raster vector image;

[0138] Vegetation information is obtained by interpreting multi-source raster vector maps using a random forest algorithm.

[0139] In one feasible implementation, the image to be predicted is input into the model and the prediction result is obtained, then vectorized. The resulting vector layer is then fused with six types of multi-source data vector layers related to vegetation distribution, namely soil type, slope, landform, vegetation classification, land use, and strata, to form a multi-source heterogeneous data body for each year.

[0140] A total of 7 layers of vector data were constructed into fishing net units with a size of 5 meters × 5 meters and rasterized. All layers were clipped, and point feature vectors were generated at the geometric center of each fishing net unit. The attributes of the geometric center were used to represent various feature attributes within the raster range.

[0141] Attribute values ​​are encoded using one-hot codes, and the data is trained using a random forest algorithm. The multi-source data obtained in the above steps are used to help determine whether the vegetation extraction results are true. If the probability is lower than the threshold, it is considered a false detection. The threshold is set to 0.9. Predictions that do not meet the judgment verification are discarded and reassembled.

[0142] The vectorized layer based on the re-stitched prediction results is combined with the actual vegetation distribution for evaluation. The accuracy index of the extracted vegetation is calculated and evaluated to obtain vegetation information.

[0143] This invention proposes a remote sensing vegetation extraction method based on deep learning semantic segmentation. It employs a deep learning algorithm, an encoder-decoder architecture, and integrates an attention mechanism module with edge detection technology to construct a semantic segmentation network based on attention fusion and edge detection for farmland and grassland extraction. A multi-scale module is added to the encoder-decoder for woodland and orchard extraction. Multiple loss functions are used to assist model training, and the trained model is used to test vegetation extraction on a test area image. The machine interpretation results are obtained and vectorized. A random forest algorithm combined with multiple vector feature layers is used for further interpretation of vegetation distribution, and the results are evaluated in conjunction with the actual vegetation distribution. This invention provides a precise and efficient remote sensing image vegetation extraction method based on deep learning semantic segmentation.

[0144] Figure 2 This is a block diagram of a remote sensing vegetation extraction device based on deep learning semantic segmentation, according to an exemplary embodiment. (Refer to...) Figure 2 The device includes:

[0145] The data acquisition module 210 is used to process the raw remote sensing images acquired by Gaofen-2 and Landsat-8 satellites to obtain training data and semantic labels.

[0146] The model building module 220 is used to build a vegetation extraction model to be trained based on the improved encoder-decoder structure.

[0147] The model training module 230 is used to train the vegetation extraction model to be trained using training data and semantic labels, so as to obtain the vegetation extraction model.

[0148] Information extraction module 240 is used to input the remote sensing image to be extracted into the vegetation extraction model to obtain vegetation extraction information;

[0149] The data interpretation module 250 is used to interpret multi-source data of vegetation extraction information to obtain vegetation information.

[0150] Optionally, the data acquisition module 210 is further used for:

[0151] The original remote sensing images acquired by Gaofen-2 and Landsat-8 satellites were manually annotated to obtain semantic tags; the semantic tags include cultivated land, grassland, forest land and orchard.

[0152] The normalized vegetation index (NDI) was calculated based on the original remote sensing images and semantic tags collected by Gaofen-2 and Landsat-8 satellites; the NDI was used as a spectral feature.

[0153] Based on the original remote sensing images acquired by Gaofen-2 and Landsat-8 satellites, image homogeneity, image mean, and image correlation were calculated; these three factors were then used as texture features.

[0154] Spectral and texture features are added to the original remote sensing image to obtain fused data;

[0155] Data augmentation is performed on the fused data to obtain training data.

[0156] Among them, the vegetation extraction model to be trained includes a vegetation extraction sub-model for cultivated land and grassland and a vegetation extraction sub-model for forest land and orchard.

[0157] Among them, the sub-model for extracting vegetation from cultivated land and grassland is used to extract vegetation from training data with semantic labels of cultivated land and grassland; the sub-model for extracting vegetation from cultivated land and grassland is constructed based on an encoder-decoder structure and combined with an attention fusion edge detection model.

[0158] Among them, the training sub-model for extracting vegetation from forest land and orchard land is used to extract vegetation from training data with semantic labels of forest land and orchard land. The training sub-model for extracting vegetation from forest land and orchard land is constructed based on an encoder-decoder structure, combined with multi-scale modules and dilated convolutional pyramid pooling modules.

[0159] Optionally, the model training module 230 is further used for:

[0160] Based on semantic labels, the training data is divided into cultivated land and grassland training data and forest and orchard training data;

[0161] The training data of cultivated land and grassland is input into the vegetation extraction model for training to obtain the cultivated land and grassland loss function; when the cultivated land and grassland loss function converges, the cultivated land and grassland vegetation extraction sub-model is obtained.

[0162] The training data of woodland and orchard is input into the vegetation extraction model for training to obtain the woodland and orchard loss function; when the woodland and orchard loss function converges, the woodland and orchard vegetation extraction sub-model is obtained.

[0163] Based on the vegetation extraction sub-models for cultivated land and grassland and forest and orchard, a vegetation extraction model is obtained.

[0164] The farmland and grassland loss function is obtained by weighting the binary cross-entropy loss function, IOU loss function, and structural loss function together.

[0165] The woodland and orchard loss function is obtained by linear combination of the Lovasz loss function and the binary cross-entropy loss function.

[0166] Optionally, the data interpretation module 250 is further used for:

[0167] Vector overlay is performed based on preset multi-source heterogeneous layer data and vegetation extraction information to obtain a multi-source data vector map;

[0168] Rasterize the multi-source data vector image to obtain a multi-source raster vector image;

[0169] Vegetation information is obtained by interpreting multi-source raster vector maps using a random forest algorithm.

[0170] This invention proposes a remote sensing vegetation extraction method based on deep learning semantic segmentation. It employs a deep learning algorithm, an encoder-decoder architecture, and integrates an attention mechanism module with edge detection technology to construct a semantic segmentation network based on attention fusion and edge detection for farmland and grassland extraction. A multi-scale module is added to the encoder-decoder for woodland and orchard extraction. Multiple loss functions are used to assist model training, and the trained model is used to test vegetation extraction on a test area image. The machine interpretation results are obtained and vectorized. A random forest algorithm combined with multiple vector feature layers is used for further interpretation of vegetation distribution, and the results are evaluated in conjunction with the actual vegetation distribution. This invention provides a precise and efficient remote sensing image vegetation extraction method based on deep learning semantic segmentation.

[0171] Figure 3 This is a schematic diagram of the structure of an electronic device 300 provided in an embodiment of the present invention. The electronic device 300 may vary considerably due to different configurations or performance. It may include one or more central processing units (CPUs) 301 and one or more memories 302. The memory 302 stores at least one instruction, which is loaded and executed by the processor 301 to implement the steps of the above-mentioned remote sensing vegetation extraction method based on deep learning semantic segmentation.

[0172] In an exemplary embodiment, a computer-readable storage medium is also provided, such as a memory including instructions that can be executed by a processor in a terminal to complete the aforementioned remote sensing vegetation extraction method based on deep learning semantic segmentation. For example, the computer-readable storage medium may be a ROM, random access memory (RAM), CD-ROM, magnetic tape, floppy disk, or optical data storage device.

[0173] Those skilled in the art will understand that all or part of the steps of the above embodiments can be implemented by hardware or by a program instructing related hardware. The program can be stored in a computer-readable storage medium, such as a read-only memory, a disk, or an optical disk.

[0174] The above description is only a preferred embodiment of the present invention and is not intended to limit the present invention. Any modifications, equivalent substitutions, improvements, etc., made within the spirit and principles of the present invention should be included within the protection scope of the present invention.

Claims

1. A remote sensing vegetation extraction method based on deep learning semantic segmentation, characterized in that, The method includes: Data processing was performed on the raw remote sensing images acquired by Gaofen-2 and Landsat-8 satellites to obtain training data and semantic labels; An improved encoder-decoder structure was used to establish a vegetation extraction model to be trained. The vegetation extraction model is trained using the training data and the semantic labels to obtain the vegetation extraction model. The vegetation extraction model to be trained includes a vegetation extraction sub-model for cultivated land and grassland and a vegetation extraction sub-model for forest and orchard. The sub-model for extracting vegetation from cultivated land and grassland is used to extract vegetation from training data with semantic labels of cultivated land and grassland. The sub-model for extracting vegetation from cultivated land and grassland is constructed based on an encoder-decoder structure and combined with an attention fusion edge detection model. The sub-model for extracting vegetation from forest land and orchard land is used to extract vegetation from training data with semantic labels of forest land and orchard land. The sub-model for extracting vegetation from forest land and orchard land is constructed based on an encoder-decoder structure, combined with multi-scale modules and dilated convolutional pyramid pooling modules. The step of training the vegetation extraction model using the training data and the semantic labels to obtain the vegetation extraction model includes: Based on the semantic labels, the training data is divided into cultivated land and grassland training data and forest and orchard training data; The cultivated land and grassland training data are input into the vegetation extraction model for training to obtain the cultivated land and grassland loss function; when the cultivated land and grassland loss function converges, the cultivated land and grassland vegetation extraction sub-model is obtained. The woodland and orchard training data are input into the vegetation extraction model for training to obtain the woodland and orchard loss function; when the woodland and orchard loss function converges, the woodland and orchard vegetation extraction sub-model is obtained. Based on the cultivated land and grassland vegetation extraction sub-model and the forest land and orchard output vegetation extraction sub-model, a vegetation extraction model is obtained; The farmland and grassland loss function is obtained by connecting the binary cross-entropy loss function, the IOU loss function, and the structural loss function in a weighted manner. The woodland / orchard loss function is obtained by a linear combination of the Lovasz loss function and the binary cross-entropy loss function. Input the remote sensing image to be extracted into the vegetation extraction model to obtain vegetation extraction information; The extracted vegetation information is analyzed using multi-source data to obtain vegetation information.

2. The remote sensing vegetation extraction method based on deep learning semantic segmentation according to claim 1, characterized in that, The process of processing raw remote sensing images acquired by Gaofen-2 and Landsat-8 satellites to obtain training data and semantic labels includes: The original remote sensing images acquired by Gaofen-2 and Landsat-8 satellites were manually annotated to obtain semantic tags; the semantic tags include cultivated land, grassland, forest land and orchard. The normalized vegetation index (NDI) is calculated based on the original remote sensing images acquired by Gaofen-2 and Landsat-8 satellites and the semantic tags; the NDI is used as a spectral feature. Based on the original remote sensing images acquired by Gaofen-2 and Landsat-8 satellites, image homogeneity, image mean, and image correlation are calculated; these image homogeneity, image mean, and image correlation are then used as texture features. The spectral features and the texture features are added to the original remote sensing image to obtain fused data; Data augmentation is performed on the fused data to obtain training data.

3. The remote sensing vegetation extraction method based on deep learning semantic segmentation according to claim 1, characterized in that, The step of performing multi-source data analysis on the extracted vegetation information to obtain vegetation information includes: A multi-source data vector map is obtained by overlaying the preset multi-source heterogeneous layer data and the extracted vegetation information. The multi-source data vector map is rasterized to obtain a multi-source raster vector map. Based on the multi-source raster vector map, vegetation information is obtained through interpretation using the random forest algorithm.

4. A remote sensing vegetation extraction device based on deep learning semantic segmentation, characterized in that, The device includes: The data acquisition module is used to process raw remote sensing images acquired by Gaofen-2 and Landsat-8 satellites to obtain training data and semantic labels. The model building module is used to build a vegetation extraction model to be trained based on the improved encoder-decoder structure. The model training module is used to train the vegetation extraction model to be trained using the training data and the semantic labels to obtain the vegetation extraction model. The vegetation extraction model to be trained includes a vegetation extraction sub-model for cultivated land and grassland and a vegetation extraction sub-model for forest and orchard. The sub-model for extracting vegetation from cultivated land and grassland is used to extract vegetation from training data with semantic labels of cultivated land and grassland. The sub-model for extracting vegetation from cultivated land and grassland is constructed based on an encoder-decoder structure and combined with an attention fusion edge detection model. The sub-model for extracting vegetation from forest land and orchard land is used to extract vegetation from training data with semantic labels of forest land and orchard land. The sub-model for extracting vegetation from forest land and orchard land is constructed based on an encoder-decoder structure, combined with multi-scale modules and dilated convolutional pyramid pooling modules. The step of training the vegetation extraction model using the training data and the semantic labels to obtain the vegetation extraction model includes: Based on the semantic labels, the training data is divided into cultivated land and grassland training data and forest and orchard training data; The cultivated land and grassland training data are input into the vegetation extraction model for training to obtain the cultivated land and grassland loss function; when the cultivated land and grassland loss function converges, the cultivated land and grassland vegetation extraction sub-model is obtained. The woodland and orchard training data are input into the vegetation extraction model for training to obtain the woodland and orchard loss function; when the woodland and orchard loss function converges, the woodland and orchard vegetation extraction sub-model is obtained. Based on the cultivated land and grassland vegetation extraction sub-model and the forest land and orchard output vegetation extraction sub-model, a vegetation extraction model is obtained; The farmland and grassland loss function is obtained by connecting the binary cross-entropy loss function, the IOU loss function, and the structural loss function in a weighted manner. The woodland / orchard loss function is obtained by a linear combination of the Lovasz loss function and the binary cross-entropy loss function. The information extraction module is used to input the remote sensing image to be extracted into the vegetation extraction model to obtain vegetation extraction information; The data interpretation module is used to perform multi-source data interpretation on the vegetation extraction information to obtain vegetation information.

Citation Information

Patent Citations

  • Model training method, woodland change detection method, system, and apparatus, and medium

    WO2022252799A1