Remote sensing winter wheat identification method based on multi-task learning model

By introducing the multi-task learning model MCFormer with NDVI and LST features, the problems of unclear and misclassification of winter wheat edges in remote sensing images are solved, and the synchronous improvement of winter wheat extraction accuracy and boundary recognition are achieved.

CN120339829APending Publication Date: 2025-07-18ANHUI UNIV
View PDF 0 Cites 1 Cited by

Patent Information

Application Number
CN202510393762.0
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-03-31
Publication Date
2025-07-18

AI Technical Summary

Technical Problem

When existing remote sensing technology extracts winter wheat in high-resolution remote sensing images, omissions caused by unclear edges and dense planting and misclassification problems caused by intra-class diversity cannot be effectively solved.

Method used

Using MCFormer based on a multi-task learning model, combined with NDVI and LST features, the semantic segmentation and boundary detection of winter wheat is realized through the spatial detail feature module, the global information feature module and the feature fusion module. CBAM is used to capture the global information in the space, and a dual-path structure is built for multi-task learning.

Benefits of technology

It significantly improves the extraction accuracy and boundary recognition ability of winter wheat, reduces misclassification, enhances the distinction between winter wheat and non-crop land, and provides more accurate segmentation results.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120339829A_ABST
    Figure CN120339829A_ABST
Patent Text Reader

Abstract

The invention relates to a remote sensing winter wheat identification method based on a multi-task learning model, and aims to solve the problems of inaccurate winter wheat edge extraction in a high-resolution remote sensing image and omission, misclassification and the like caused by over-dense distribution and intra-class diversity in an existing classical semantic segmentation model. According to the invention, a Vision Transform (MCFormer) based on multi-task learning is designed, and a semantic segmentation task and a semantic boundary task are associated together, so that the accuracy of boundary detection is remarkably improved. NDVI and LST are introduced as sample wavebands, so that the spectral characteristics of the winter wheat are enhanced.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to a method for remotely sensing winter wheat identification based on a multi-task learning model, belonging to the technical field of remote sensing agricultural monitoring. Background Art

[0002] Timely, efficient, and accurate acquisition of winter wheat planting area information and spatial distribution is crucial for government management departments to adjust agricultural planting structures, formulate agricultural food policies, and ensure national food security.

[0003] Remote sensing, as a non-destructive way to monitor crops in time and space, remote sensing images with different spatial resolutions provide an effective means for winter wheat mapping. Current research on remote sensing of winter wheat mainly relies on Sentinel, MODIS, and Landsat data. However, winter wheat has the characteristics of small patches (meter level) and high spatial heterogeneity, especially in China. Sentinel (10m), MODIS (250m), and Landsat (30m) may have difficulty in effectively detecting these small fragmented plots. In recent years, high-resolution images represented by GF-2 data (China's high-resolution series, resource series, Planet, GeoEye-1, WordView series, etc.) have been widely applied to remote sensing monitoring, which can provide clearer spectral information of winter wheat.

[0004] Deep learning is currently paving a new path for remote sensing analysis. Deep learning methods for remote sensing crop classification can be divided into two categories according to different research objectives: pixel-based classification methods and semantic segmentation-based methods. Pixel-based crop classification methods use individual pixels in remote sensing images as training samples for the CNN model to achieve crop classification for each pixel point. However, pixel-based classification methods only utilize the spectral information of individual pixels and limited neighborhood information, and cannot consider the spectral information and spatial context information of surrounding pixels. The extraction results often have serious problems such as salt-and-pepper noise and internal structure fragmentation. Semantic segmentation assigns clearer and more accurate semantic concepts to each pixel on the basis of image segmentation. Compared with pixel-based classification methods, semantic segmentation can effectively capture object boundaries and details and provide relatively accurate segmentation results. However, semantic segmentation methods are weak in capturing global context information and dealing with complex scenes.

[0005] Transformer was first applied to the field of natural language processing. Inspired by the significant success of the Transformer architecture in the NLP field, researchers have recently applied Transformer to computer vision (CV) tasks. VisionTransformer (ViT) is the first visual Transformer designed with positional encoding, which can complement the loss of positional information caused by sequential input. ViT can capture long-range dependencies and global context information, which makes it more advantageous in dealing with complex scenarios. Theoretically, it is highly feasible to use ViT for remote sensing winter wheat extraction. However, to improve the performance of the model in remote sensing images, not only the selection and optimization of the model architecture are required, but also the enhancement of sample features is crucial.

[0006] The Normalized Difference Vegetation Index (NDVI) is one of the most common indices reflecting vegetation characteristics. NDVI can directly enhance the spectral characteristics of winter wheat, making it easier to distinguish winter wheat from other surface features (such as buildings, bare land, water bodies, etc.) in remote sensing images. The Land Surface Temperature Index (LST) can inversely enhance the thermal radiation characteristics of other surface features (such as buildings, bare land, etc.), making the temperature changes of these non-vegetation areas more significant, thereby indirectly enhancing the contrast effect between winter wheat and its surrounding environment. Through this contrast, the characteristics of winter wheat also become more obvious and are easier to identify and distinguish in remote sensing images. However, most current studies on NDVI and LST focus on the traditional machine learning field, and rarely combine them as sample features directly with deep learning methods.

[0007] The above-mentioned research has made fruitful progress in remote sensing winter wheat mapping. However, existing methods still have problems such as unclear edges, omissions caused by dense planting, and misclassifications caused by intra-class diversity when using remote sensing images to extract winter wheat. Therefore, how to significantly improve the ability of boundary recognition while maintaining accuracy requires the study of new deep learning models that can simultaneously improve accuracy and boundary recognition ability. Summary of the Invention

[0008] The purpose of the present invention is to solve the problems existing in the above-mentioned prior art and propose a remote sensing winter wheat recognition method based on a multi-task learning model, which can simultaneously improve the extraction accuracy of winter wheat and the boundary recognition ability.

[0009] A remote sensing winter wheat recognition method based on a multi-task learning model proposed by the present invention is characterized by including the following steps:

[0010] Step 1: Obtain the multi-spectral remote sensing image and the land surface temperature data LST in spring in the target area, calculate the Normalized Difference Vegetation Index (NDVI) of the pixels in the multi-spectral remote sensing image, and fuse the NDVI, green light band, and land surface temperature data LST to generate a PNG image;

[0011] Step 2: Select winter wheat samples and background samples in the PNG image by visual identification with the help of the multi-spectral remote sensing image, set type labels, perform mask inversion on the winter wheat image samples and background samples, and input them into the MCFormer model together for multi-task learning, while training the internal features and boundary features of the target;

[0012] Step 3: Input the PNG image in Step 1 into the trained MCFormer model for winter wheat identification and winter wheat boundary extraction.

[0013] Furthermore, the MCFormer model includes:

[0014] - Spatial-detailed Feature Module (SFM), which is used to perform several convolution operations on the input data to generate a spatial-detailed feature map;

[0015] - Global-information Feature Module (GFM), which is used to process the input data and generate 4 global-information feature maps of different sizes; and

[0016] - Feature Fusion Module (FFM), which is used to fuse the spatial-detailed feature map and the 4 global-information feature maps to generate the final semantic features.

[0017] Furthermore, the spatial-detailed feature module has six standard 3×3 convolution layers, and each convolution layer is equipped with batch normalization operation and ReLU6 activation function. Among them, the stride of the convolution kernel of each convolution layer gradually increases, gradually expanding the channel dimension of the spatial feature map, and the size of the output feature map is 1 / 4 of the original image.

[0018] Furthermore, the Global-information Feature Module (GFM) includes a Patch Embedding (PE) module, three Patch Merging (PM) modules, and four WFB modules. The PE module, PM modules, and WFB modules are arranged at intervals and transfer data in sequence. The WFB module includes a convolutional attention module and a convolutional multi-layer perceptron. The convolutional attention module is used to generate a channel attention map and obtain a spatial attention map based on it, and the convolutional multi-layer perceptron is used to generate a global information feature map.

[0019] Furthermore, the attention module includes a channel attention module and a spatial attention module. The channel attention module uses global average pooling and global max pooling methods to obtain two vectors when compressing the spatial dimension of the feature map, and then inputs these two vectors into a shared multi-layer perceptron composed of a hidden layer and a multi-layer perceptron for processing to obtain two features Maxout and Avgout respectively, and then calculates the channel attention map through the sigmoid function. Summing each element in combination with the channel attention map and outputting a feature vector; the spatial attention module performs global max pooling and global average pooling on the feature vector, connects them to generate an effective feature descriptor, performs a convolution operation, and then performs a sigmoid operation to generate a spatial attention map.

[0020] Furthermore, the convolutional multi-layer perceptron includes a 3×3 convolution, batch normalization, and ReLU6 activation function that process data in sequence; a depth 3×3 convolution, batch normalization, and ReLU6 activation function; a 1×1 convolution and batch normalization, and finally generates the global information feature map.

[0021] Furthermore, the feature fusion module performs feature fusion through a feature pyramid network. First, perform 1×1 convolution on the 4 global information feature maps generated by the global information feature module; then, through a 3×3 convolution, batch normalization, and ReLU6 activation function, and then perform an upsampling operation and an addition operation to achieve global information feature fusion; finally, combine the fused global information feature with the spatial detail feature map of the spatial detail feature module to generate the final winter wheat semantic feature, that is, the winter wheat extraction result and the winter wheat boundary line extraction result.

[0022] Furthermore, the NDVI value of the pixel in the remote sensing image is calculated by the following formula:

[0023]

[0024] In the formula, Red represents the red light band of the remote sensing image, and NIR represents the near-infrared light of the remote sensing image.

[0025] Furthermore, in the first step, the processed LST image data corresponding to the selected remote sensing image time is obtained from USGS, the LST image data is resampled to the same spatial resolution as the selected remote sensing image, and the red light band of the original remote sensing image is replaced with the NDVI value and the blue light band is replaced with the LST value by using band replacement in ArcGIS, thereby obtaining an enhanced feature winter wheat sample library.

[0026] The present invention introduces NDVI and LST directly into winter wheat samples. LST can reflect the characteristics of the surface thermal environment, and there are significant differences in the thermal radiation characteristics of different land types, making the distinction between winter wheat and non-crop land types more obvious, thus effectively reducing misclassification and improving the accuracy of remote sensing segmentation. The MCFormer model based on multi-task learning is designed, and a multi-task learning training framework of semantic task plus boundary task is proposed to improve the accuracy of field edge recognition through a boundary constraint mechanism. On this basis, a joint method of spatial information-global information is constructed inside the model, using deep convolution to retain texture details and combining Vision Transformer to capture long-range dependencies. Specifically, this study uses a dual-path structure to construct MCFormer. One path uses a stacked convolutional layer to preserve rich spatial details. The other path uses Transformer blocks to construct a ViT backbone that captures global context information. Here, the convolutional block attention module (CBAM) is introduced because of their high portability and versatility, and the spatial global information presented in the data is effectively captured through channel and spatial attention mechanisms.

[0027] The main contributions of the present invention can be summarized as follows:

[0028] (1) NDVI and LST are introduced as bands into the sample image, enriching the spectral characteristics of the winter wheat sample image.

[0029] (2) A multi-task learning Vision Transformer for winter wheat mapping, called MCFormer, is proposed. The semantic segmentation task and the semantic boundary task are associated together, enabling the model to learn the boundary texture characteristics of the sample while learning the internal characteristics of the original sample.

[0030] (3) The MCFormer is constructed using a spatial path and a global path, and the convolutional block attention module (CBAM) is used to replace the window-based multi-head self-attention mechanism (W-MHSA), making it more effective in capturing the spatial global information of winter wheat. Description of the Drawings

[0031] The present invention will be further described below with reference to the accompanying drawings.

[0032] Figure 1Obtain the structural schematic diagram of the remote sensing winter wheat recognition method based on the multi-task learning model of the present invention.

[0033] Figure 2 It is the overall architecture diagram of the MCFormer model.

[0034] Figure 3 It is the patch embedding, patch fusion, and WFB architecture diagram of the global information feature module.

[0035] Figure 4 Schematic diagram of the winter wheat extraction result and the winter wheat boundary line extraction result. Specific implementation mode

[0036] The present invention will be further described below in conjunction with the accompanying drawings and specific embodiments.

[0037] In this embodiment, a certain province in China is selected as the target area for remote sensing winter wheat extraction. The remote sensing winter wheat recognition method based on the multi-task learning model in the embodiment of the present invention includes the following steps:

[0038] First step, obtain the multi-spectral remote sensing images (GF-2 remote sensing images) of the target area in April and May, and perform processing such as radiometric calibration, atmospheric correction, and orthorectification on the remote sensing images. Radiometric calibration converts the original image into physical radiance values, and atmospheric correction is based on the radiative transfer model to eliminate the influence of atmospheric scattering and absorption to ensure the accuracy of spectral features. Orthorectification is used to eliminate geometric distortion and ensure the spatial accuracy of the image. Then, the normalized difference vegetation index NDVI of the selected remote sensing image is calculated using the following formula:

[0039]

[0040] In the formula, Red represents the red band of the remote sensing image, and NIR represents the near-infrared light of the remote sensing image.

[0041] Obtain the processed land surface temperature data LST corresponding to the time of the selected remote sensing image from USGS, and resample the land surface temperature data LST to the same spatial resolution as the selected remote sensing image.

[0042] In the data fusion stage, the processed land surface temperature data LST and the normalized difference vegetation index NDVI are integrated into the selected remote sensing image in the form of bands, thus forming a new image data containing a total of five bands: R, G, B, NDVI, and LST. In order to improve the effect of the model, in this embodiment, when converting to a PNG image, the NDVI band replaces the original red band, and the LST band replaces the blue band. Experimental results show that the effect of using the LST to replace the blue band is better than the scheme of replacing the green band. The normalized difference vegetation index NDVI, the green band, and the land surface temperature data LST are fused to generate a PNG image.

[0043] After the image synthesis was completed, the image was cropped to generate samples with a size of 1024×1024 pixels. A total of 1243 samples were cropped, among which 747 samples were used as the training data set, 248 as the validation data set, and another 248 as the test data set. By fusing multi-spectral data and land surface temperature information, more abundant feature information can be provided for the winter wheat mapping task. The samples are shown as Figure 1 shown below.

[0044] Step 2: Select winter wheat samples and background samples in the PNG image through visual recognition with the help of multi-spectral remote sensing images and set type labels. Semantic boundary detection aims to identify the pixel categories of the target boundary. As a dual task of semantic segmentation, the goal of semantic boundary detection is to identify the image pixels on the object boundary. The present invention couples the semantic segmentation and semantic boundary detection tasks and realizes the interaction between the two tasks by sharing latent semantics. Specifically, the model extracts the boundary pixels of the input mask to generate a boundary mask, and simultaneously trains the internal and boundary features of the target. Finally, the predicted boundary mask is used as the outer contour line constraint mask to achieve the effect of boundary constraint. To achieve this goal, the present invention performs mask inversion on the winter wheat image samples after the band synthesis operation, and the original mask and the inverted mask are simultaneously input into the MCFormer model for multi-task learning, and the internal and boundary features of the target are trained simultaneously.

[0045] In this embodiment, the overall architecture of the MCFormer model consists of Figure 2 as shown below, including: a spatial detail feature module, a global information feature module, and a feature fusion module. Among them, the spatial detail feature module is used to perform several convolution operations on the input data to generate a spatial detail feature map, the global information feature module is used to process the input data and generate 4 global information feature maps with different sizes, and the feature fusion module is used to perform feature fusion on the spatial detail feature map and the 4 global information feature maps to generate the final semantic feature. The spatial detail feature module uses convolution operations to generate a spatial detail feature map, while the global information feature module generates 4 global information feature maps with different sizes. Finally, the spatial detail feature map and the 4 global information feature maps enter the feature fusion module for feature fusion to generate the final semantic feature.

[0046] Specifically, as shown in Figure 2As shown, the spatial detail feature module has six standard 3×3 convolutional layers, and each convolutional layer is equipped with a batch normalization operation and a ReLU6 activation function. Among them, the stride of the convolutional kernel of each convolutional layer gradually increases, gradually expanding the channel dimension of the spatial feature map. The purpose of doing this is to obtain sufficient spatial information in the winter wheat image to generate a high-resolution spatial detail feature map, preserve more spatial detail features of the winter wheat, and finally the size of the output feature map is 1 / 4 of the original image.

[0047] As Figure 2 shown, the global information feature module includes a patch embedding module, three patch merging modules, and four WFB modules. The patch embedding module, patch merging modules, and WFB modules are arranged at intervals and perform data transfer in sequence.

[0048] As Figure 3 shown, patch embedding divides the input image into fixed-size tiles through two 3×3 convolutional layers, flattens each tile into a vector, and maps it to a low-dimensional space, reducing the image resolution to 1 / 4 of the original. After the convolutional layer, there are a batch normalization and a ReLU6 activation function. The relative position relationship between patches is enhanced through 3×3 depth convolution and residual connection. Patch merging first uses batch normalization and then downsamples the image to 1 / 2 of the original through a 2×2 convolutional layer. Three patch merging modules are used in the global information feature module to gradually reduce the image resolution to 1 / 8, 1 / 16, and 1 / 32. After each module, the position relationship between patches is also enhanced through 3×3 depth convolution and residual connection.

[0049] The WFB (Wheat Former Block) module consists of batch normalization, a convolutional attention module, a convolutional multi-layer perceptron, and a residual connection. The convolutional attention module is used to generate a channel attention map and obtain a spatial attention map accordingly. The convolutional multi-layer perceptron is used to generate a global information feature map. The attention module includes a channel attention module and a spatial attention module. The channel attention module uses global average pooling and global max pooling methods to obtain two vectors when compressing the spatial dimension of the feature map, and then inputs these two vectors into a shared multi-layer perceptron composed of a hidden layer and a multi-layer perceptron for processing to obtain two features Maxout and Avgout respectively, and then calculates the channel attention map through a sigmoid function. The sum of each element is combined with the channel attention map and the feature vector is output; the spatial attention module performs global max pooling and global average pooling on the feature vector, connects them to generate an effective feature descriptor, and performs a convolution operation with a 7×7 convolutional kernel, and then performs a sigmoid operation to generate spatial attention features to obtain a spatial attention map. The convolutional multi-layer perceptron combines convolutional operations to enhance local feature extraction and non-linear expression capabilities, which helps to better model complex image features.

[0050] The convolutional multi-layer perceptron includes a 3×3 convolution, batch normalization, and ReLU6 activation function that process data sequentially; a depth 3×3 convolution, batch normalization, and ReLU6 activation function; and a 1×1 convolution and batch normalization, finally generating the global information feature map. The convolutional multi-layer perceptron is a deep learning model architecture that combines convolutional operations with a multi-layer perceptron (MLP), aiming to enhance the ability of image feature extraction. Convolutional operations (CNN) extract features from images through local receptive fields and can automatically learn low-level local features such as edges and textures. However, it has certain limitations in modeling global features and dealing with complex non-linear relationships. On the other hand, MLP can effectively perform complex non-linear feature mapping through fully connected layers, but it ignores the local spatial structure in images. Therefore, it is often less flexible than convolutional operations when capturing fine-grained features. C-MLP combines the advantages of convolution and MLP, being able to efficiently extract local features through convolutional operations and enhance the ability to model complex features through the non-linear mapping ability of MLP.

[0051] To better fuse the winter wheat spatial detail feature map generated by the spatial detail feature module and the 4 global semantic feature maps generated by the global information feature module, the embodiment of the present invention uses a feature pyramid network for feature fusion. First, perform 1×1 convolution on the 4 global information feature maps generated by the global information feature module to unify the number of channels to 384; then, through 3×3 convolution, batch normalization, and ReLU6 activation function, and then perform upsampling operation and addition operation to achieve global information feature fusion; finally, combine the fused global information feature with the spatial detail feature map of the spatial detail feature module to generate the final winter wheat semantic feature, that is, the winter wheat extraction result and the winter wheat boundary line extraction result. The extraction results and recognition results are shown in Figure 4 。

[0052] To improve the accuracy of winter wheat plot boundary extraction, this embodiment introduces a multi-task learning boundary constraint technique and uses a joint loss to train the model. The formula for the joint loss function L is as follows:

[0053]

[0054] represents the cross-entropy loss between the predicted label and the true label.

[0055] represents the dice loss between the predicted label and the true label.

[0056] L(Y) represents the Laplacian convolution of the predicted label.

[0057] represents the Laplacian convolution of the true label.

[0058] Indicates the binary cross-entropy loss under the Laplacian convolution of the predicted label and the true label.

[0059] Where Y and represent the predicted label and the true label respectively, and L bce is the cross-entropy loss. The cross-entropy loss function is a loss function used to measure the difference between the predicted probability distribution of the model and the true probability distribution. L dice represents the dice loss, which is used to handle the imbalance problem between the background and winter wheat pixels. The dice loss is a loss function for image segmentation tasks. It optimizes the model by calculating the overlapping part between the prediction and the true result, especially suitable for the case of class imbalance, used to evaluate the similarity of two samples, with a value range between 0 and 1, and the larger the value, the more similar. L represents the Laplacian convolution, which is used to extract the constructed boundary of the predicted label and the true label. The binary cross-entropy loss L bce is used to calculate the loss of the boundary pixels. The binary cross-entropy loss is a commonly used loss function in binary classification problems, used to measure the difference between the predicted probability of the model and the actual label. This loss function is particularly suitable for the case where the output is a probability value and the label is 0 or 1.

[0060] The total loss L of multi-task learning S is expressed as:

[0061] L S = L1 + L2

[0062] Where L1 and L2 are the losses of different tasks respectively. Through multi-task learning, the model can more accurately extract the boundary of winter wheat plots during the training phase.

[0063] Step 3: Input the PNG image in the first step into the trained MCFormer model for winter wheat recognition and winter wheat boundary extraction. The extraction result is similar to the test result ( Figure 4 ).

[0064] In addition to the above embodiments, the present invention may have other embodiments. All technical solutions formed by equivalent replacement or equivalent transformation fall within the protection scope required by the present invention.

Claims

1. A method for remotely sensing winter wheat identification based on a multi-task learning model, characterized in that It includes the following steps: The first step: Obtain the multi-spectral remote sensing image and the land surface temperature data LST in spring in the target area, calculate the normalized difference vegetation index NDVI of the pixels in the multi-spectral remote sensing image, and fuse the normalized difference vegetation index NDVI, the green light band and the land surface temperature data LST to generate a PNG image; The second step: Select winter wheat samples and background samples in the PNG image by visual recognition with the help of the multi-spectral remote sensing image, set type labels, perform mask inversion on the winter wheat image samples and background samples, and input them into the MCFormer model together for multi-task learning, and train the internal features and boundary features of the target at the same time; The third step: Input the PNG image in the first step into the trained MCFormer model for winter wheat recognition and winter wheat boundary extraction.

2. The remote sensing winter wheat identification method based on the multi-task learning model according to claim 1, wherein The MCFormer model includes: - A spatial detail feature module for performing several convolution operations on the input data to generate a spatial detail feature map; - A global information feature module for processing the input data and generating 4 global information feature maps of different sizes; and - A feature fusion module for fusing the spatial detail feature map and the 4 global information feature maps to generate the final semantic feature.

3. The remote sensing winter wheat recognition method based on the multi-task learning model according to claim 2, characterized in that: The spatial detail feature module has six standard 3×3 convolutional layers, and each convolutional layer is equipped with a batch normalization operation and a ReLU6 activation function. Among them, the stride of the convolutional kernel of the convolutional layer increases gradually each time, gradually expanding the channel dimension of the spatial feature map, and the size of the output feature map is 1 / 4 of the original image.

4. The remote sensing winter wheat identification method based on the multi-task learning model according to claim 2, characterized in that: The global information feature module includes a patch embedding module, three patch merging modules and four WFB modules. The patch embedding module, the patch merging module and the WFB module are arranged at intervals and perform data transmission in sequence. The WFB module includes a convolutional attention module and a convolutional multi-layer perceptron. The convolutional attention module is used to generate a channel attention map and obtain a spatial attention map accordingly. The convolutional multi-layer perceptron is used to generate a global information feature map.

5. The remote sensing winter wheat identification method based on the multi-task learning model according to claim 4, characterized in that: The attention module includes a channel attention module and a spatial attention module. The channel attention module uses the global average pooling and global max pooling methods to obtain two vectors when compressing the spatial dimension of the feature map, and then inputs these two vectors into a shared multi-layer perceptron composed of a hidden layer and a multi-layer perceptron for processing to obtain two features Maxout and Avgout respectively, and then calculates the channel attention map through the sigmoid function, sums each element in combination with the channel attention map and outputs a feature vector; The spatial attention module performs global max pooling and global average pooling on the feature vector, connects them to generate an effective feature descriptor, performs a convolution operation, and then performs a sigmoid operation to generate a spatial attention map.

6. The remote sensing winter wheat identification method based on the multi-task learning model according to claim 4, wherein: The convolutional multi-layer perceptron includes a 3×3 convolution, a batch normalization and a ReLU6 activation function for processing the data in sequence; a depth 3×3 convolution, a batch normalization and a ReLU6 activation function; a 1×1 convolution and a batch normalization, and finally generates the global information feature map.

7. The remote sensing winter wheat recognition method based on the multi-task learning model according to claim 2, characterized in that: The feature fusion module performs feature fusion through a feature pyramid network. First, 1×1 convolution is performed on the 4 global information feature maps generated by the global information feature module. Then, through 3×3 convolution, batch normalization, and ReLU6 activation function, and then upsampling operation and addition operation are performed to achieve global information feature fusion. Finally, the fused global information feature is combined with the spatial detail feature map of the spatial detail feature module to generate the final winter wheat semantic feature, that is, the winter wheat extraction result and the winter wheat boundary line extraction result.

8. The remote sensing winter wheat identification method based on the multi-task learning model according to claim 1, characterized in that: The NDVI value of the pixel in the remote sensing image is calculated by the following formula: In the formula, Red represents the red light band of the remote sensing image, and NIR represents the near-infrared light of the remote sensing image.

9. The remote sensing winter wheat recognition method based on a multi-task learning model according to claim 1, characterized in that: In the first step, the processed LST image data corresponding to the selected remote sensing image time is obtained from USGS, and the LST image data is resampled to the same spatial resolution as the selected remote sensing image. The red light band of the original remote sensing image is replaced with the NDVI value, and the blue light band is replaced with the LST value by using band replacement in ArcGIS, thus obtaining an enhanced winter wheat sample library.

10. The remote sensing winter wheat identification method based on the multi-task learning model according to claim 1, wherein: During model training, a joint loss function is used for multi-task learning boundary constraint. The formula of the joint loss function L is as follows: Represents the cross-entropy loss between the predicted label and the true label; Dice loss representing the predicted label and the true label; L(Y) represents the Laplacian convolution of the predicted label; Laplacian convolution representing the true label; Indicates the binary cross-entropy loss under the Laplace convolution of the predicted label and the true label; where Y and represent the predicted label and the true label respectively, and L bce is the cross-entropy loss function, and L dice represents the dice loss, L represents the Laplacian convolution, and the binary cross-entropy loss L bce is used to calculate the loss of the boundary pixels. The total loss L S of multi-task learning is expressed as: L S = L1 + L2 where L1 and L2 are the losses of different tasks respectively.

Citation Information

Cited By

  • Titanium alloy tissue multi-scale feature extraction method

    CN121544912A