A Sparse Sample Crop Mapping Method Based on Weakly Supervised Semantic Segmentation
By adopting a weakly supervised semantic segmentation method in crop mapping, combining the time-sequence semantic segmentation model and multi-task weakly supervised learning network, using sparse samples and image edge labels to generate spatiotemporal and spatial spectral features and dense pseudo-labels, the problem of low crop mapping accuracy caused by sparse samples is solved, and more efficient and accurate crop recognition is achieved.
Patent Information
- Application Number
- CN202411271106.5
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2024-09-11
- Publication Date
- 2025-05-27
- Estimated Expiration
- 2044-09-11
AI Technical Summary
The prior art is difficult to efficiently obtain high-quality samples with sufficient quantity and dense distribution, resulting in low crop mapping accuracy, especially in large-area crop remote sensing mapping.
A sparse sample crop mapping method based on weakly supervised semantic segmentation is adopted. Through a time-sequential semantic segmentation model and a multi-task weak supervised learning network, combined with a convolutional neural network and a long and short-term memory network, a sparse point crop label and image edge label are used to generate space-time spectral features and dense pseudo-labels, and supervised training is carried out to improve crop recognition accuracy.
It effectively improves the accuracy of crop category identification based on sparse point crop labels, enhances the perception of crop spatial structure information, improves the generalization ability of the model under sparse and few samples, and provides more reliable technical support for large-scale remote sensing agricultural monitoring.
Smart Images

Figure CN119205976B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of agricultural remote sensing ground object information extraction, and in particular to a sparse sample crop mapping method based on weakly supervised semantic segmentation, which is used for sparse small sample crop mapping. Background Art
[0002] The growing population has greatly increased the demand for food. However, blind expansion of cultivated land and increased farming intensity may have adverse effects on the ecological environment. Therefore, it is crucial to effectively balance food production and ecological environmental protection to achieve sustainable agricultural development. Timely grasping of crop spatial distribution information is the key to achieving sustainable agriculture for assessing food production, formulating agricultural production policies, and optimizing agricultural resource allocation. Traditionally, crop spatial distribution information is mainly obtained through field surveys by census personnel. However, the high cost of capital and manpower investment has led to a long update cycle and poor effectiveness of census data, which is difficult to meet actual needs.
[0003] The rapid development of remote sensing satellite technology has provided an effective means for large-scale automated and efficient identification of crops. Existing remote sensing mapping studies mostly use threshold methods or machine learning to mine the unique phenological characteristics of crop growth stages to distinguish different crop categories. Among them, the threshold method is based on phenological prior knowledge and uses single or multi-period remote sensing image features to construct classification decision rules and thresholds for crop identification. However, the threshold method is often limited to the spectral characteristics of a single period and is difficult to effectively deal with the phenological shift problem caused by differences in climate, topography, soil conditions and agricultural practices. In contrast, machine learning methods, such as decision trees, random forests and support vector machines, can automatically establish mathematical relationships between remote sensing features and target crop categories, improving the robustness of the model to deal with phenological migration problems. However, traditional machine learning algorithms usually treat multi-temporal information as multiple independent data, ignoring their temporal dependencies, and are difficult to deal with the "same spectrum, different objects" problem. In recent years, deep learning methods, such as recurrent neural networks (RNNs), convolutional neural networks (CNNs) and self-attention networks, have been widely used in various crop mapping applications with their powerful high-dimensional feature mining capabilities. However, deep learning models are highly dependent on a large number of high-quality dense surface labels. Existing field sample collection strategies are limited by terrain and sampling costs, and can only obtain sparse point labels, which is difficult to meet the training requirements of deep learning models, resulting in reduced crop recognition accuracy. Therefore, how to design a suitable weakly supervised learning framework to efficiently utilize limited sparse weak samples is the key to achieving efficient automatic crop recognition. Summary of the invention
[0004] The object of the present invention is to provide a sparse sample crop mapping method based on weakly supervised semantic segmentation for the above problems existing in the prior art, and to solve the problems that it is difficult to efficiently obtain a sufficient number of high-quality samples with dense distribution in large-area crop remote sensing mapping, the actual sample distribution is sparse in space and the quantity is insufficient, resulting in low crop mapping accuracy and other problems.
[0005] In order to achieve the above object, the following is the technical solution of the present invention:
[0006] A sparse sample crop mapping method based on weakly supervised semantic segmentation, comprising the following steps:
[0007] Step S1: Obtain the monthly composite temporal reflectance data in the study area;
[0008] Step S2: Obtain the sparse point-like crop labels in the study area, and process the monthly composite temporal reflectance data to obtain the image edge labels;
[0009] Step S3: Construct a temporal semantic segmentation model, and the temporal semantic segmentation model generates spatio-temporal spectral features based on the monthly composite temporal reflectance data;
[0010] Step S4: Construct a multi-task weakly supervised learning network, and the multi-task weakly supervised learning network includes a Segmentation task head module, an Edge task head module, and an Expansion task head module.
[0011] In the Segmentation task head module: Based on the sparse point-like crop labels and spatio-temporal spectral features, generate the predicted probability values of different crop categories for all sampling points in the study area;
[0012] In the Edge task head module: Based on the image edge labels and spatio-temporal spectral features, generate the predicted edge intensities for all sampling points in the study area;
[0013] In the Expansion task head module: Based on the dense pseudo-labels constructed from the image edge labels and the predicted probability values whose predicted probability values in the Segmentation task head module are greater than a set adjustable threshold, generate the final predicted probability values of different crop categories for all sample points in the study area;
[0014] Step S5: Construct a total loss function, and train the temporal semantic segmentation model and the multi-task weakly supervised learning network;
[0015] Step S6: Obtain the monthly composite temporal reflectance data, image edge labels, and sparse point-like crop labels of the area to be predicted, and use the temporal semantic segmentation model and the multi-task weakly supervised learning network to obtain the final crop categories of all sample points in the area to be predicted.
[0016] As described above, the monthly synthetic temporal reflectivity data in step S1 is obtained based on the following steps:
[0017] Collect Sentinel-2 remote sensing images of the study area, use the Sentinel-2 cloud probability product provided by the GEE platform, remove the pixels with a corresponding cloud probability higher than 65%, obtain the cloud-removed Sentinel-2 reflectivity data, and perform median synthesis on the cloud-removed Sentinel-2 reflectivity data to obtain the monthly synthetic temporal reflectivity data.
[0018] As described above, the image edge labels in step S2 are obtained based on the following steps:
[0019] Process the monthly synthetic temporal reflectivity data based on Sobel filtering to generate the edge gradient map for the corresponding month, normalize the edge gradient map to generate the edge intensity map for each month, and perform mean synthesis on the edge intensity map to generate the image edge labels.
[0020] As described above, the temporal semantic segmentation model in step S3 includes a multi-scale spatial feature encoder, a temporal feature encoder, and a multi-stage feature decoder.
[0021] In the multi-scale spatial feature encoder: Different numbers of cascaded DownConv modules are used to extract spatial feature maps of different scales corresponding to the synthetic reflectivity data of each month. The spatial feature maps obtained at the same scale and different months are concatenated along the time dimension to obtain the temporal-spatial feature map at this scale. Generating the temporal-spatial feature maps at each scale is the multi-scale temporal-spatial feature map.
[0022] In the temporal feature encoder: The temporal-spatial feature map is converted into a three-dimensional vector of dimension (HW)×T×C, where T is the number of months of the temporal-spatial feature map, C is the number of feature channels, H and W are the height and width of the temporal-spatial feature map. The three-dimensional vector passes through two layers of bidirectional BiLSTM modules to generate the spatio-temporal feature sequence corresponding to the scale. The spatio-temporal feature sequence corresponding to the scale is input into a 1×1 convolutional layer with a softmax activation function to obtain the attention weights for different months. The features of different months in the spatio-temporal feature sequence corresponding to the scale are multiplied by the corresponding attention weights and then weighted and summed to generate the spatio-temporal feature map corresponding to the scale.
[0023] The multi-stage feature decoder contains an atrous spatial pyramid pooling module and multiple joint attention mechanism modules. n is the serial number of the scale, and the value range of n is 1 to N, where N is the total number of scales, and the scales from 1 to N become deeper in turn. The spatio-temporal feature map corresponding to scale N is input into the atrous spatial pyramid pooling module. After the feature map output by the atrous spatial pyramid pooling module passes through the upsampling module corresponding to scale N, the decoded feature map corresponding to scale N is obtained.
[0024] For scales 1 to (N - 1): After concatenating the spatio-temporal feature map corresponding to scale n and the decoded feature map corresponding to scale n + 1, feature fusion is performed through the corresponding joint attention mechanism module, and then through the upsampling module corresponding to scale n to obtain the decoded feature map corresponding to scale n;
[0025] The decoded feature map corresponding to scale 1 serves as the spatio-temporal spectral feature output by the temporal semantic segmentation model.
[0026] As described above, the atrous spatial pyramid pooling module includes three atrous convolution modules and one pooling module, and the joint attention mechanism module includes a channel attention module and a spatial attention module.
[0027] As described above, the total loss function L is based on the following formula:
[0028] L = L seg + L exp + λ e L edge + λ c L con
[0029] Where, L seg is the loss function of the Segmentation task head module, L edge is the edge loss function of the Edge task head module, L exp is the loss function of the Expansion task head module, L con is the consistency loss function between the prediction probability values of different crop categories by the Expansion task head module and the Segmentation task head module, λ e and λ c are the weight coefficients of the edge detection loss and the consistency loss, respectively.
[0030] As described above, the loss function L seg of the Segmentation task head module is based on the following formula:
[0031]
[0032] Where, n g and k respectively represent the total number of sample points of the sparse dot-like crop label and the number of crop categories, is the prediction probability value of the c-th crop category of the i-th sample point of the sparse dot-like label by the Segmentation task head module, y i,c is the actual probability value of the c-th crop category of the i-th sample point of the sparse dot-like crop label, and γ is an adjustable coefficient with a value range greater than zero.
[0033] As described above, the edge loss function L of the Edge task head module edge is based on the following formula:
[0034]
[0035] where j represents the serial number of the sample point in the study area, is the predicted edge intensity of the Edge task head module for the j-th sample point, and b j is the actual image edge intensity of the j-th sample point of the image edge label, is the first type of edge loss function, is the second type of edge loss function.
[0036] As described above, the loss function L of the Expansion task head module exp and the consistency loss L con are based on the following formula:
[0037]
[0038] where |||| 2 represents the square value calculation, n e represents the number of non-zero sample points of the dense pseudo-label, m is the serial number of the non-zero sample point of the dense pseudo-label, is the predicted probability value of the c-th crop category for the non-zero m-th sample point of the dense pseudo-label by the Expansion task head module, E m,c is the probability value of the c-th crop category for the non-zero m-th sample point of the dense pseudo-label, is the predicted probability value of the c-th crop category for the j-th sampling point in the study area calculated and output by the Segmentation task head module, is the predicted probability value of the c-th crop category for the j-th sample point in the study area by the Expansion task head module, and H and W are the height and width of the input monthly composite temporal reflectivity data.
[0039] As described above, the dense pseudo-label is obtained based on the following formula:
[0040]
[0041] E j,c = 0 else
[0042] where τ a and τ b are two adjustable thresholds with a value range of 0 to 1, is the predicted probability value of the c-th crop category for the j-th sampling point in the study area calculated and output by the Segmentation task head module, Calculate the maximum value of the predicted probability values of various crop categories at the j-th sampling point in the study area for the Segmentation task head module. For the Segmentation task head module, calculate the crop category corresponding to the maximum value of the predicted probability values of various crop categories at the j-th sampling point in the study area, b j Is the actual image edge intensity of the j-th sample point of the image edge label; E j,c Is the dense pseudo-label of the c-th crop category at the j-th sampling point in the study area.
[0043] The present invention has the following beneficial effects compared with the prior art:
[0044] 1. Constructed a new temporal semantic segmentation model U-TempoNet that combines a convolutional neural network and a long short-term memory network. This model can deeply explore the complex spatial and temporal features of crops in temporal remote sensing data, thereby achieving more efficient and accurate crop category recognition.
[0045] 2. Introduced an edge detection task to obtain the image edge intensity, enhancing the perception ability of the spatial structure information of crops, enabling it to more accurately restore the boundaries of crop fields, thereby improving the accuracy of the segmentation results.
[0046] 3. By introducing a pseudo-label supervision task, constructing dense pseudo-labels based on the image edge label and the predicted probability values in the Segmentation task head module whose predicted probability values are greater than a set adjustable threshold, and performing supervised training, the richness of the training samples of the model is improved, further deepening the understanding and discrimination ability of crop categories, and enhancing its generalization ability in practical applications.
[0047] Through the integration of the above innovative technologies, the present invention effectively improves the accuracy of crop category recognition based on sparse point-like crop labels and provides more reliable technical support for large-scale remote sensing agricultural monitoring. BRIEF DESCRIPTION OF THE DRAWINGS
[0048] Figure 1 Is the flowchart of the present invention;
[0049] Figure 2 Is the distribution map of the study area of the present invention; among them, (A) is the distribution map of study area A; (B) is the distribution map of study area B; (C) is the distribution map of study area C;
[0050] Figure 3 Is the structural diagram of the temporal semantic segmentation model of the present invention;
[0051] Figure 4 Is the multi-task weak supervision learning framework diagram of the present invention;
[0052] Figure 5 These are the evaluation results of the crop classification accuracy of the present invention. Among them, (a) represents the evaluation result of the crop classification accuracy of the present invention in the Jianghan Plain test area; (b) is the evaluation result of the crop classification accuracy of the present invention in the Songnen Plain test area; (c) is the evaluation result of the crop classification accuracy of the present invention in the Sanjiang Plain test area. Detailed implementation manners
[0053] To facilitate the understanding and implementation of the present invention by those of ordinary skill in the art, the present invention will be further described in detail below in conjunction with embodiments. It should be understood that the embodiments described herein are only used to illustrate and explain the present invention, and are not used to limit the present invention.
[0054] As Figure 1 shown, a sparse sample crop mapping method based on weakly supervised semantic segmentation includes the following steps:
[0055] Step S1: Obtain Sentinel-2 remote sensing images in the study area based on the GEE (Google Earth Engine) platform and perform cloud removal processing to obtain monthly composite temporal reflectance data:
[0056] Step S1.1: Sentinel-2 remote sensing images of study areas A, B, and C were collected respectively based on the GEE platform (the distribution of study areas and field sampling points is shown in Figure 1 ). The collection time range was from January 1, 2021 to December 31, 2021; in order to capture the complex spatio-temporal characteristics of areas with strong heterogeneity, the present invention selected four spectral bands (i.e., blue, green, red, and near-infrared bands) with a spatial resolution of 10 meters of Sentinel-2 remote sensing images as the research data;
[0057] Step S1.2: For the collected Sentinel-2 remote sensing images, use the Sentinel-2 cloud probability product provided by the GEE platform to remove the pixels with a corresponding cloud probability higher than 65%, and obtain the cloud-removed Sentinel-2 reflectance data. The monthly median synthesis method is adopted, and the monthly composite temporal reflectance data is obtained by performing median synthesis on the cloud-removed Sentinel-2 reflectance data for each month during the period from January 1, 2021 to December 31, 2021.
[0058] Step S2: Obtain sparse point crop labels in the study area (which can be obtained through on-site sampling by ground personnel or known sparse point crop labels). Process the monthly composite temporal reflectance data based on the Sobel operator to obtain image edge labels:
[0059] Step S2.1: Ground personnel went to the study area for on-site sampling in late April 2021 and late August 2021 respectively. Through a handheld GPS instrument, they recorded the geographical coordinates of the categories of crops in the study area and obtained sparse dot-like crop labels of the pixels corresponding to the geographical coordinates.
[0060] Step S2.2: Process the monthly composite temporal reflectance data obtained in Step S1 based on Sobel filtering to generate edge gradient maps for each corresponding month one by one. Then, normalize these edge gradient maps to the range of 0 to 1 through 2% linear stretching to generate edge intensity maps for each month. Finally, perform mean synthesis on these edge intensity maps to generate high-quality image edge labels. By aggregating the means of multi-temporal information, the influence of cloud and rain is effectively reduced, thus ensuring the high quality of the image edge labels.
[0061] Step S3: The present invention couples a convolutional neural network (CNN) and a long short-term memory network (LSTM) to construct a temporal semantic segmentation model (U-TempoNet) specifically for monthly composite temporal reflectance data. The overall structure of the model is as Figure 3 shown, which includes a multi-scale spatial feature encoder, a temporal feature encoder, and a multi-stage feature decoder. This model can not only capture the spatial features in remote sensing data but also fully exploit the temporal variation features therein, thereby obtaining robust and representative "spatio-temporal spectrum" features:
[0062] S3.1 Multi-scale spatial feature encoder: For the monthly composite temporal reflectance data obtained in Step S1, different numbers of cascaded DownConv modules are used to extract spatial feature maps of different scales corresponding to the composite reflectance data of each month (such as Figure 3As shown). The DownConv module consists of three Conv2DN layers. The first Conv2DN layer uses a convolutional operation with a stride of 2 to reduce the resolution of the input feature map to 1 / 2 of the original. The feature map output by the first Conv2DN layer is input into the second Conv2DN layer, and the output feature map of the second Conv2DN layer is input into the third Conv2DN layer. The result of adding the output feature map of the third Conv2DN layer and the output feature map of the second Conv2DN layer is used as the output spatial feature map of the DownConv module. Through this additive skip connection, the DownConv module can extract deeper features while avoiding the vanishing gradient caused by an overly deep model. Each DownConv module generates spatial feature maps of a corresponding scale. Compared with the previous scale, both the height and width are reduced by half, while the number of channels is tripled. Different numbers of cascaded DownConv modules are used to process the monthly synthetic temporal reflectivity data for each month to obtain spatial feature maps of different scales for each month. The spatial feature maps obtained at the same scale but different months are concatenated along the time dimension to obtain the temporal-spatial feature map at this scale. Generating the temporal-spatial feature maps at each scale results in multi-scale temporal-spatial feature maps. The dimension of the temporal-spatial feature map is T×C×H×W, where T is the number of months, C is the number of channels of the spatial feature map, and H and W are the height and width of the temporal-spatial feature map.
[0063] S3.2 Temporal Feature Encoder: For the temporal-spatial feature maps of each scale obtained in step S3.1, the AtBiLSTM module is used to model the temporal dependencies between the feature maps of different months. The AtBiLSTM wraps two layers of bidirectional LSTM (BiLSTM) modules and a 1×1 convolutional layer with a softmax activation function. Given a temporal-spatial feature map with dimensions T×C×H×W, where T is the number of months, C is the number of channels of the spatial feature map, and H and W are the height and width of the temporal-spatial feature map. First, the temporal-spatial feature map is converted into a three-dimensional vector with dimensions (HW)×T×C, which can be regarded as converting the H×W temporal-spatial features in the temporal-spatial feature map into one-dimensional vectors, with the number of months for each temporal-spatial feature being T and the number of feature channels being C. Subsequently, two layers of bidirectional BiLSTM modules are used to explore the temporal dependencies in these temporal-spatial features, and the three-dimensional vector generates the corresponding spatio-temporal feature sequence after passing through two layers of bidirectional BiLSTM modules. The corresponding spatio-temporal feature sequence of the scale is input into a 1×1 convolutional layer with a softmax activation function to learn the attention weights for different months. The probability values output by the softmax activation function are used as the attention weights for each month. Finally, the features of different months in the corresponding spatio-temporal feature sequence of the scale are multiplied by the corresponding attention weights and then weighted and summed to generate the spatio-temporal feature map of the scale, and the spatio-temporal feature maps of each scale generate the final multi-scale spatio-temporal feature map.
[0064] S3.3 Multi-Stage Feature Decoder: For the spatio-temporal feature maps of each scale obtained in step S3.2, since the spatio-temporal feature maps of different scales can capture different aspects of the target crop. The deeper spatio-temporal feature maps, although with lower resolution, have richer contextual semantic information. On the other hand, the lower-level spatio-temporal feature maps contain more spatial details, which helps for more refined analysis. To effectively fuse information with different resolutions and semantic richness, a multi-stage feature decoder is used to perform deep fusion on the spatio-temporal feature maps of different scales, so as to obtain robust and representative spatio-temporal spectral features.
[0065] The multi-stage feature decoder includes an Atrous Spatial Pyramid Pooling (ASPP) module and multiple Joint Attention mechanism modules (JA). The Atrous Spatial Pyramid Pooling module includes three atrous convolution modules and one pooling module. The dilation rates of the three atrous convolution layers are set to 1, 2, and 4 in sequence to expand the receptive field and capture more context information. The Joint Attention mechanism module includes a channel attention module and a spatial attention module. The channel attention module and the spatial attention module process the input features of the Joint Attention mechanism module respectively. The output features of the channel attention module are used as the feature channel weight vector to multiply with the input features of the Joint Attention mechanism module to obtain the first intermediate feature. The output features of the spatial attention module are used as the feature spatial weight vector to multiply with the input features of the Joint Attention mechanism module to obtain the second intermediate feature. After the first intermediate feature and the second intermediate feature are multiplied, they are used as the output features of the Joint Attention mechanism module. Channel attention re-adjusts the contributions of features in different channels by assigning attention weights, while spatial attention highlights significant regions by assigning attention weights to different positions.
[0066] n is the serial number of the scale, and the value range of n is from 1 to N, where N is the total number of scales. The scales from scale 1 to scale N become deeper in sequence. The spatio-temporal feature map corresponding to scale N is input into the Atrous Spatial Pyramid Pooling module. After the feature map output by the Atrous Spatial Pyramid Pooling module passes through the upsampling module corresponding to scale N, the decoded feature map corresponding to scale N is obtained.
[0067] For scales 1 to (N - 1): The spatio-temporal feature map corresponding to scale n is concatenated with the decoded feature map corresponding to scale n + 1, and then undergoes feature fusion through the corresponding Joint Attention mechanism module, and then passes through the upsampling module corresponding to scale n to obtain the decoded feature map corresponding to scale n.
[0068] Finally, the decoded feature map corresponding to scale 1 is used as the spatio-temporal spectral feature output by the temporal semantic segmentation model.
[0069] Step S4. In order to improve the performance of the temporal semantic segmentation model constructed in step S3 under the condition of sparse samples, the present invention designs a multi-task weakly supervised learning network based on multi-task collaborative training ( Figure 4 ). The multi-task weakly supervised learning network includes a Segmentation task head module, an Edge task head module, and an Expansion task head module. Each task head module includes two layers of 3×3 convolutional layers.
[0070] Among them, the Segmentation task head module performs semantic segmentation tasks. It is trained using the sparse dot-like crop labels obtained in step S2.1 and the robust and representative spatio-temporal spectral features obtained in step S3.3 to identify different crop categories and generate the predicted probability values of different crop categories for all sampling points (pixel points) in the study area. The Edge task head module uses the image edge labels obtained in step S2.2 as labels and also combines the spatio-temporal spectral features obtained in step S3.3 for training to perform image edge detection tasks, assisting the temporal semantic segmentation model to learn the spatial structure information of crops and obtaining the predicted edge intensities of all sampling points in the study area. The Expansion task head module constructs dense pseudo-labels using the image edge labels obtained in step S2.2 and the predicted probability values with higher values (greater than the set adjustable threshold) in the Segmentation task head module, and performs supervised training by combining spatio-temporal spectral features and dense pseudo-labels, and outputs the predicted probability values of different crop categories for all sample points in the study area to improve the crop recognition accuracy of the model under sparse few-sample conditions.
[0071] S4.1: For the Segmentation task head module. The present invention uses Focal loss to calculate the loss function L seg (Formula 1) to reduce the influence of class imbalance problems. In Formula (1), y and p s respectively represent the sparse dot-like crop labels obtained in step S2.1 and the predicted probability values of the Segmentation module. n g and k respectively represent the total number of sample points of the sparse dot-like crop labels and the number of crop categories. i and c are respectively the serial numbers of the sample points of the sparse dot-like crop labels and the serial numbers of the crop categories.
[0072]
[0073] Among them, p s i,c is the predicted probability value of the c-th crop category of the i-th sample point of the sparse dot-like label obtained in step S2.1 by the Segmentation task head module, y i,c is the actual probability value of the c-th crop category of the i-th sample point of the sparse dot-like crop label obtained in step S2.1, and γ is an adjustable coefficient with a value range greater than zero in a focal loss. The Segmentation task head module calculates and outputs the predicted probability values of all sampling points in the study area. The sample points of the sparse dot-like labels are only part of the sampling points in the study area. The serial number of the sample points of the sparse dot-like labels is i, and the serial number of the sampling points in the study area is j.
[0074] S4.2: For the Edge task head module, construct the edge loss function \(L\) based on the Tanimoto loss (Equation 2). edge (Equation 4). Where, \(b\) and respectively represent the image edge label obtained in step S2.2 and the predicted edge intensity of the Edge task head module. \(L\) t represents the loss function of the Edge task head module (Tanimoto loss function), which can effectively address the problem of the imbalance in the number of edge pixels and non-edge pixels in the edge detection task.
[0075]
[0076] Where, \(j\) represents the serial number of the sample point in the study area, is the predicted edge intensity of the \(j\)-th sample point by the Edge task head module, and \(b\) j is the actual image edge intensity of the \(j\)-th sample point of the image edge label obtained in step S2.2. The sampling points of the image edge label are all the sampling points in the study area, is the first type of edge loss function, is the second type of edge loss function.
[0077] S4.3: For the Expansion task head module, use the result with a higher predicted probability value in the Segmentation task head module as the label amplification pool. On this basis, amplify the sparse punctate crop labels, and the generated dense pseudo-labels are used for the training of the Expansion task head module. In order to improve the amplification accuracy of the dense pseudo-labels and remove the influence of mixed pixels in the image boundary part, the image edge label obtained in S2.2 is introduced to constrain the amplification range of generating dense pseudo-labels( Figure 4 ). The generation process of the dense pseudo-labels \(E\in[0,k]\) (where, 1 to \(k\) represent different types of crops, and 0 represents unlabeled pixels) constrained by the image edge label is shown in Equation (5).
[0078]
[0079] Where, \(\tau\) a and \(\tau\) b are two adjustable thresholds with a value range of 0 to 1, which are used to constrain the amplification range of the dense pseudo-labels. Their values are set to 0.99 and 0.15 respectively. The Expansion task head module also constructs the loss function \(L\) exp (Equation 6). is the predicted probability value of the \(c\)-th crop category of the \(j\)-th sampling point in the study area calculated and output by the Segmentation task head module, Calculate the maximum value of the predicted probability values of various crop categories at the j-th sampling point in the study area for the Segmentation task head module. Calculate the crop category corresponding to the maximum value of the predicted probability values of various crop categories at the j-th sampling point in the study area for the Segmentation task head module, b j Is the actual image edge intensity of the j-th sample point of the image edge label obtained in step S2.2; E j,c Is the dense pseudo-label of the c-th crop category at the j-th sampling point in the study area.
[0080] In order to enforce consistency constraints between the Segmentation task head module and the Expansion task head module, before inputting the "spatiotemporal spectrum" features extracted in step S3.3 into the Expansion task head module, the present invention applies perturbations to them using the dropout method, and after applying the perturbations, the "spatiotemporal spectrum" features are input into the Expansion task head module for training. The consistency loss function L between the predicted probability values of different crop categories by the Expansion task head module and the Segmentation task head module con Can be expressed by formula (7).
[0081]
[0082] In formula (6), |||| 2 Represents the calculation of the squared value, n e Represents the number of non-zero sample points in the dense pseudo-label, m is the serial number of the non-zero sample points in the dense pseudo-label, Is the predicted probability value of the c-th crop category for the m-th non-zero sample point in the dense pseudo-label by the Expansion task head module, E m,c Is the probability value of the c-th crop category for the m-th non-zero sample point in the dense pseudo-label, Is the predicted probability value of the c-th crop category at the j-th sampling point in the study area calculated by the Segmentation task head module, Is the predicted probability value of the c-th crop category for the j-th sample point in the study area by the Expansion task head module. In formula (7), H and W are the height and width of the input monthly composite temporal reflectance data. The final total loss function L is as shown in formula (8), λ e And λ cThey are the weight coefficients of the edge detection loss and the consistency loss respectively. By jointly training multiple groups of tasks, the temporal semantic segmentation model can effectively utilize the crop spatial structure information and the rich context semantic information of dense pseudo-labels, making up for the shortage of sparse sample data, thereby improving the crop mapping accuracy.
[0083] L = L seg + L exp + λ e L edge + λ c L con (8)
[0084] Step S5: Under the condition of minimizing the total loss function, supervise and train the temporal semantic segmentation model in step S3 and the multi-task weakly supervised learning network in step S4.
[0085] Step S6: Obtain the monthly composite temporal reflectance data and image edge labels of the area to be predicted, obtain the sparse point-like crop labels in the study area, use the temporal semantic segmentation model and the multi-task weakly supervised learning network for prediction, obtain the final prediction probability values of different crop categories for all sample points in the area to be predicted, and then process them through the argmax function to obtain the final crop categories of all sample points in the area to be predicted.
[0086] The final crop categories of all sample points in the area to be predicted are compared with the crop recognition results of two traditional machine learning models (Random Forest RF and Support Vector Machine SVM), and three deep learning models (DCM, Transformer, and Patch-based CNN). Metrics such as F1-score, User’s Accuracy, and Producer’s Accuracy are used to evaluate the accuracy of the results, and the evaluation results are as Figure 5 shown. The method of the present invention has higher accuracy and better recognition effect compared with these methods.
[0087] The above is only the preferred embodiment of the present invention, and the protection scope of the present invention is not limited to the above embodiments. All technical solutions falling within the idea of the present invention belong to the protection scope of the present invention. It should be noted that for those of ordinary skill in the art, several improvements and refinements made without departing from the principle of the present invention should be regarded as within the protection scope of the present invention.
Claims
1. A sparse sample crop mapping method based on weakly supervised semantic segmentation, characterized in that: The following steps are involved: Step S1: Obtain monthly synthetic time series reflectivity data in the study area; Step S2: Obtain sparse point crop labels in the study area and process the monthly synthetic time series reflectivity data to obtain image edge labels; Step S3: construct a time series semantic segmentation model, which generates spatiotemporal spectral features based on monthly synthetic time series reflectivity data; Step S4: Construct a multi-task weakly supervised learning network, which includes a Segmentation task head module, an Edge task head module, and an Expansion task head module. In the Segmentation task head module: based on sparse point crop labels and spatiotemporal spectral features, the predicted probability values of different crop categories at all sampling points in the study area are generated; In the Edge task head module: based on the image edge labels and spatiotemporal spectral features, the predicted edge strength of all sampling points in the study area is generated; In the Expansion task head module: based on the image edge labels and the dense pseudo labels constructed by the predicted probability values in the Segmentation task head module whose predicted probability values are greater than the set adjustable threshold, the final predicted probability values of different crop categories for all sample points in the study area are generated; Step S5: construct a total loss function to train the temporal semantic segmentation model and the multi-task weakly supervised learning network; Step S6: Obtain monthly synthetic time series reflectance data, image edge labels, and sparse point crop labels of the area to be predicted, and use the time series semantic segmentation model and multi-task weakly supervised learning network to obtain the final crop category of all sample points in the area to be predicted.
2. The sparse sample crop mapping method based on weakly supervised semantic segmentation according to claim 1, characterized in that: The monthly synthetic time series reflectivity data in step S1 is obtained based on the following steps: Sentinel-2 remote sensing images of the study area were collected, and the Sentinel-2 cloud probability products provided by the GEE platform were used to remove pixels with corresponding cloud probabilities higher than 65% to obtain the de-clouded Sentinel-2 reflectivity data. The de-clouded Sentinel-2 reflectivity data were median-synthesized to obtain monthly synthetic time-series reflectivity data.
3. The sparse sample crop mapping method based on weakly supervised semantic segmentation according to claim 1, characterized in that: The image edge label in step S2 is obtained based on the following steps: The monthly synthetic time series reflectivity data is processed based on Sobel filtering to generate the edge gradient map of the corresponding month. The edge gradient map is normalized to generate the edge intensity map of each month. The edge intensity map is mean-synthesized to generate the image edge label.
4. The sparse sample crop mapping method based on weakly supervised semantic segmentation according to claim 1, characterized in that: The temporal semantic segmentation model in step S3 includes a multi-scale spatial feature encoder, a temporal feature encoder, and a multi-stage feature decoder. In the multi-scale spatial feature encoder: DownConv modules with different numbers of series connections are used to extract spatial feature maps of different scales corresponding to the synthetic reflectance data of each month. The spatial feature maps of the same scale and obtained in different months are connected along the time dimension to obtain the temporal spatial feature map of the scale. The temporal spatial feature maps of each scale are generated as multi-scale temporal spatial feature maps. In the temporal feature encoder: the temporal spatial feature map is converted into a three-dimensional vector of dimension (HW)×T×C, where T is the number of months of the temporal spatial feature, C is the number of feature channels, H and W are the height and width of the temporal spatial feature map, and the three-dimensional vector is passed through two layers of bidirectional BiLSTM modules to generate a spatiotemporal feature sequence corresponding to the scale. The spatiotemporal feature sequence corresponding to the scale is input into a 1×1 convolutional layer with a softmax activation function to obtain the attention weights of different months. The features of different months in the spatiotemporal feature sequence corresponding to the scale are multiplied by the corresponding attention weights and then weighted summed to generate the spatiotemporal feature map corresponding to the scale. The multi-stage feature decoder includes a dilated spatial pyramid pooling module and multiple joint attention mechanism modules. n is the serial number of the scale, and the value range of n is 1 to N. N is the total number of scales. Scales 1 to N are successively deeper. The spatiotemporal feature map corresponding to scale N is input into the dilated spatial pyramid pooling module. The feature map output by the dilated spatial pyramid pooling module passes through the upsampling module corresponding to scale N to obtain the decoding feature map corresponding to scale N. For scale 1-scale (N-1): the spatiotemporal feature map corresponding to scale n is concatenated with the decoding feature map corresponding to scale n+1, and then the feature fusion is performed through the corresponding joint attention mechanism module, and then the decoding feature map corresponding to scale n is obtained through the upsampling module corresponding to scale n; The decoded feature map corresponding to scale 1 is used as the spatiotemporal spectral feature output by the temporal semantic segmentation model.
5. The sparse sample crop mapping method based on weakly supervised semantic segmentation according to claim 4, characterized in that: The atrous spatial pyramid pooling module includes three atrous convolution modules and one pooling module, and the joint attention mechanism module includes a channel attention module and a spatial attention module.
6. The sparse sample crop mapping method based on weakly supervised semantic segmentation according to claim 1, characterized in that: The total loss function L is based on the following formula: L=L seg +L exp +λ e L edge +λ c L con Among them, L seg is the loss function of the Segmentation task head module, L edge is the edge loss function of the Edge task head module, L exp is the loss function of the Expansion task head module, L con is the consistency loss function between the predicted probability values of different crop categories by the Expansion task head module and the Segmentation task head module, λ e and λ c are the weight coefficients of edge detection loss and consistency loss respectively.
7. The sparse sample crop mapping method based on weakly supervised semantic segmentation according to claim 6, characterized in that: The loss function L of the Segmentation task head module seg Based on the following formula: Among them, n g and k represent the total number of sample points and the number of crop categories of sparse point crop labels, respectively. (is the predicted probability value of the cth crop category of the ith sample point of the sparse point label by the Segmentation task head module, y i,c is the actual probability value of the cth crop category of the ith sample point of the sparse point crop label, and γ is an adjustable coefficient with a value range greater than zero.
8. The sparse sample crop mapping method based on weakly supervised semantic segmentation according to claim 6, characterized in that: The edge loss function L of the Edge task head module is edge Based on the following formula: Among them, j represents the serial number of the sample point in the study area, is the predicted edge strength of the jth sample point by the Edge task head module, b j is the actual image edge strength of the jth sample point of the image edge label, is the first type marginal loss function, It is the second type of marginal loss function.
9. The sparse sample crop mapping method based on weakly supervised semantic segmentation according to claim 6, characterized in that: The loss function L of the expansion task head module exp and consistency loss L con Based on the following formula: Among them, |||| 2 Indicates square value calculation, n e represents the number of non-zero sample points of dense pseudo-labels, m is the sequence number of non-zero sample points of dense pseudo-labels, is the predicted probability value of the cth crop category of the mth sample point with dense pseudo labels that is non-zero, E m,c is the probability value of the cth crop category of the mth sample point with dense pseudo labels, The Segmentation task head module calculates and outputs the predicted probability value of the cth crop category at the jth sampling point in the study area. is the predicted probability value of the cth crop category at the jth sample point in the study area by the Expansion task head module, and H and W are the height and width of the input monthly synthetic time series reflectivity data.
10. The sparse sample crop mapping method based on weakly supervised semantic segmentation according to claim 1, characterized in that: The dense pseudo-labels are obtained based on the following formula: AND j,c =0else Among them, τ a and τ b are two adjustable thresholds ranging from 0 to 1. The Segmentation task head module calculates and outputs the predicted probability value of the cth crop category at the jth sampling point in the study area. The Segmentation task head module calculates and outputs the maximum predicted probability value of various crop categories at the j-th sampling point in the study area. The Segmentation task head module calculates and outputs the crop category corresponding to the maximum predicted probability value of various crop categories at the jth sampling point in the study area, b j is the actual image edge strength of the jth sample point of the image edge label; E j,c is the dense pseudo label of the cth crop category at the jth sampling point in the study area.