Intelligent monitoring method of cultivated land resources based on multi-temporal remote sensing data
Through the spectral attention mechanism and the improved self-attention mechanism, and combined with the mask-distillation feature reconstruction strategy, the U_TAE network architecture is constructed, which solves the problem of insufficient spectral feature extraction and timing feature extraction in the farmland monitoring of multi-time phase remote sensing data, and improves the monitoring accuracy and scope of application.
Patent Information
- Application Number
- CN202310534675.3
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2023-05-12
- Publication Date
- 2025-08-15
- Estimated Expiration
- 2043-05-12
AI Technical Summary
The existing multi-time phase remote sensing data farmland monitoring methods have shortcomings in dealing with spectral feature extraction, timing feature extraction and data integrity, resulting in limited monitoring accuracy and scope of application, which cannot meet the needs of local agricultural applications.
The spectral attention mechanism and an improved self-attention mechanism are used to optimize timing feature extraction, and combined with a mask-distillation-based feature reconstruction strategy, the U_TAE network architecture is constructed to enhance the model's attention to farmland sensitive bands and improve the robustness of timing data.
The accuracy and scope of application of multi-time phase remote sensing data farmland monitoring has been improved, the learning ability of crop phenological attributes and artificial agricultural activity modes has been enhanced, and the monitoring effect in complex environments has been improved.
Smart Images

Figure CN116563709B_ABST
Abstract
Description
Technical Field
[0001] The present invention belongs to the technical field of remote sensing intelligent monitoring of cultivated land resources, and in particular relates to an intelligent monitoring method for cultivated land resources oriented to multi-temporal remote sensing data. Background Art
[0002] Compared with on-site farmland surveys, the combination of remote sensing and artificial intelligence, and the use of massive remote sensing data and deep learning models to achieve automated remote sensing image interpretation, provide new solutions for the precise monitoring of farmland resources.
[0003] Since the 21st century, with the continuous development of machine learning technology, a variety of classification methods have been used for cultivated land resource monitoring, such as random forests, maximum likelihood methods, nearest neighbor methods, support vector machines, and artificial neural networks. These methods have been actively applied to global cultivated land mapping at medium and low resolutions at global scales. However, because these methods are designed based on specific domain knowledge, they can only extract low- or medium-level features from raw data under limited circumstances. They are unable to adapt to the dramatic variations in intra-class differences in cultivated land during large-scale cultivated land resource monitoring. Their poor spatiotemporal generalization leads to suboptimal product mapping accuracy. They are also unable to adapt to the dramatic variations in intra-class differences in cultivated land characteristics caused by regional differences in climate, geography, agricultural activities, and other factors. Consequently, during large-scale cultivated land mapping, production personnel must manually adjust parameters or design features based on local characteristics, adapting to changes in geographic location and satellite imaging conditions, thereby obtaining and deploying multiple locally optimal models. This approach significantly increases the time and labor costs of the process. Furthermore, due to practical production efficiency and labeling costs, local model deployment often fails to fully account for the diverse local characteristics of cultivated land, enabling finer-grained segmentation. For example, current regional segmentation is typically based on climate, but factors such as topography and crop types within the same climate region are often not considered. As a result, the accuracy and completeness of cultivated land parcel extraction remain insufficient for specific applications at the local scale. In recent years, deep convolutional neural networks (DCNNs), with their powerful feature extraction capabilities, have garnered widespread attention in cultivated land mapping and have been applied to global-scale cultivated land mapping. DCNNs can automatically learn representative and discriminant features from training data in a hierarchical manner, and the resulting high-level features exhibit a certain degree of spatiotemporal generalization. This significantly improves the accuracy of cultivated land parcel extraction in large-scale cultivated land resource monitoring and enables a single model to adapt to a wider range of cultivated land local characteristics. However, current deep learning-based intelligent cultivated land monitoring methods still struggle to achieve ideal accuracy and meet the needs of local agricultural applications. This is primarily due to three key issues.
[0004] First, with the continuous improvement of sensor spectral resolution, remote sensing Earth observation data contains increasingly rich spectral information. However, most deep learning models currently used in remote sensing are based on or modified network architectures from the computer vision field. These model architectures are primarily designed for natural images containing only RGB bands. Directly inputting all bands into the model for learning not only fails to fully extract the rich and diverse spectral features in remote sensing data, but also fails to fully explore the relationships between multiple bands. To address this issue, one approach is to use data dimensionality reduction techniques such as principal component analysis to map the important information from multiple bands to a lower dimension to fully utilize the spectral features. However, this approach may result in a certain degree of spectral information loss and is extremely sensitive to noise. Another approach is to calculate various spectral indices, such as the Normalized Difference Vegetation Index (NDVI), the Normalized Water Index (NDWI), and the Enhanced Vegetation Index (EVI), and add them as auxiliary bands to the original RGB information. This operation incorporates artificial priors into the spectral feature extraction process, helping the model focus on bands that are highly sensitive to vegetation. However, these indices can only reflect the growth status of vegetation, but not the land use type. If there is no vegetation on the cultivated land or the vegetation coverage is insufficient, these indices cannot extract cultivated land information.
[0005] Secondly, given the unique characteristics of cultivated land, its visual characteristics in satellite imagery vary significantly throughout the year, and the degree of confusion with other land cover types varies over time. For example, during the planting season, cultivated land has minimal visual differences from other land features such as grassland and regular shrubs, but during the fallow season, it is easily confused with bare land. This further increases intra-class variability in cultivated land during large-scale mapping based on single-temporal composite data. With the continuous advancement of remote sensing technology, the availability of high-resolution multi-temporal data has gradually increased, providing new possibilities for further improving the accuracy of large-scale cultivated land monitoring. Furthermore, the successful application of various large-scale cultivated land mapping methods based on multi-temporal data further demonstrates the necessity and feasibility of incorporating multi-temporal data into cultivated land classification. Furthermore, with the rapid development of deep neural networks for time series data in the remote sensing field, cultivated land mapping no longer relies on feature preprocessing results designed by expert experience, but can instead use imagery as input for autonomous learning and modeling of temporal phenological patterns. The application of deep learning-based temporal networks not only significantly reduces the manual and time-consuming preprocessing efforts, but also significantly enhances the classifier's robustness to temporal outliers through its high-dimensional, multi-level feature extraction capabilities. Currently, researchers have achieved end-to-end automated farmland mapping by leveraging temporal network architectures from computer vision or adapting existing single-temporal networks to the characteristics of remote sensing data. The strong generalization of high-level temporal features extracted from multi-temporal data using deep learning networks has been demonstrated in multiple regions around the world, significantly improving the automation and accuracy of farmland mapping. Among these, the use of self-attention mechanisms to extract temporal features is a current mainstream research direction. By further reducing the distance between elements in convolutional neural networks from logarithmic path length to constant path length, self-attention mechanisms can better capture temporal internal correlations within the data, thereby improving model performance. Furthermore, the self-attention mechanism expands the receptive field of the convolutional neural network. By directly comparing features across all spatiotemporal locations, it can be used to capture local and global long-range dependencies. When using the self-attention mechanism, triangular positional encoding is often used to embed temporal information into the data. Triangular fixed positional encoding shows strong similarity between adjacent time phases. However, in agricultural production, monthly changes can represent significant changes. We want the network to aggregate features specific to "arable land," extracting specific months (such as plowing, sowing, transplanting, and harvesting) and combining them to improve feature separability.
[0006] Finally, given that in real-world applications, complete and dense time series are difficult to obtain in some regions due to factors such as climate and satellite revisit cycles, the encoder component of existing models struggles to fully extract time series features, significantly degrading model performance. A common solution to this problem is to interpolate time series data to supplement missing time series or to use generative models to reconstruct time series data. However, interpolation can lead to deviations in the crop phenology and agricultural activity patterns reflected in the time series data. For example, in Hunan Province, the rainy season significantly reduces available data from March to April, but planting season has already begun. Directly interpolating images from February and May would result in the time series failing to reflect these agricultural phenology patterns. Using generative models for cloud removal or time series reconstruction requires training with large amounts of relevant labeled data in specific environments, which significantly increases costs and makes them difficult to apply to all regions. In addition, when the cloud cover area is too large or adjacent temporal phases are missing, the generative model is based on probability distribution modeling, which may cause confusion in the image reconstructed by the model or some land features that do not belong to the area, thereby misleading the subsequent multi-temporal semantic segmentation model for farmland extraction. Summary of the Invention
[0007] In view of this, the purpose of the present invention is to improve the performance of the existing multi-temporal semantic segmentation model and the robustness of the model to incomplete time series, thereby enhancing the accuracy and applicability of intelligent monitoring of cultivated land. To this end, on the one hand, a spectral attention mechanism is added and a temporal self-attention mechanism is improved in the process of constructing the multi-temporal semantic segmentation network, so as to add the modeling of the relationship between multiple bands, and optimize the model's learning process for crop phenological attributes and artificial agricultural activity patterns, ultimately achieving the goal of improving the extraction effect of multi-temporal semantic segmentation on the scope of cultivated land. On the other hand, taking into account the problem that the accuracy of cultivated land extraction is greatly reduced due to incomplete time series in the multi-temporal semantic segmentation network during actual application, the present invention develops a feature reconstruction strategy based on mask-distillation, thereby enhancing the applicability of the intelligent monitoring method of cultivated land. Based on the above two points, the present invention proposes an intelligent monitoring method for cultivated land for multi-temporal data.
[0008] Specifically, the present invention discloses an intelligent monitoring method for cultivated land resources based on multi-temporal remote sensing data, comprising the following steps:
[0009] Step 1: Build a spectral attention module to increase the importance of sensitive bands for cultivated land and reduce interference from other bands;
[0010] Step 2: Optimize the temporal feature extraction module based on the self-attention mechanism, improve the original L_TAE attention mechanism, and replace the original fixed position encoding with a learnable position encoding to mine and model the temporal features influenced by crop phenology and artificial agricultural activity patterns;
[0011] Step 3: Module combination and basic model design: Using the U_TAE network architecture as the basic framework, the spectral attention module and the optimized temporal attention module are inserted into the network framework to form the basic multi-temporal semantic segmentation network framework;
[0012] Step 4: Build a feature reconstruction strategy based on mask-distillation to improve the robustness of the model to incomplete time series data;
[0013] Step 5: Mapping of cultivated land cover.
[0014] Furthermore, in order to increase the importance of sensitive bands for cultivated land and reduce interference from other bands, the spectral attention module is constructed as follows:
[0015] Step 101: Define the time series multispectral data as a vector X of size T×C×H×W, and the time series vector X consists of T elements x of size C×H×W. t Composition, where T represents the time series length, C represents the number of channels, H represents the image length, and W represents the image width;
[0016] In the spectral attention module, each x t Split them out and perform spectral attention weighted operations separately;
[0017] Step 102: First, t Perform maximum pooling and average pooling to provide a general description of each band of the sample from multiple angles. The formula for this step is as follows:
[0018] f Max ,f Avg =AvgPool(x t ),MaxPool(x t );
[0019] AvgPool means average pooling, MaxPool means maximum pooling;
[0020] Step 103: Use the self-attention mechanism to operate on the above features respectively, that is, use three independent 1*1 convolutions conv_K, conv_Q, and conv_V to map the input features to construct key, query, and value respectively, and use the matrix multiplication of key and query to obtain its channel attention, and then multiply it with the value to obtain the weighted feature f Max_outWith f Avg_out In order to obtain the relationship between the various bands, the formula of this step is as follows:
[0021] f Max_out =Softmax(conv_K(f Max )×conv_Q(f Max ) T )×conv_V(f Max );
[0022] f Avg_out =Softmax(conv_K(f Avg )×conv_Q(f Avg ) T )×conv_V(f Avg );
[0023] Step 104: The obtained f Max_out With f Avg_out Add and normalize to get the final eigenvector X c Its size is the same as the original size of X, which is T×C×H×W. The formula for this step is as follows:
[0024] X c =sigmoid(f Max_out +f Avg_out );
[0025] Sigmoid is the activation function.
[0026] Furthermore, the self-attention mechanism is used to explore the mutual relationship between high-density time series images at the same spatial location, and the position encoding part is optimized to better model the time series characteristics of cultivated land.
[0027] The improved lightweight temporal attention encoder L_TAE is used to extract the temporal attention a and modify it to replace the original triangular position encoding with the learnable position encoding L_PE to obtain the temporal attention of the image. t Encoding is performed to inject temporal position information into the feature vector, thereby enhancing the temporal encoder's ability to learn crop phenological characteristics; the temporal attention a is calculated as follows:
[0028] a=L_TAE(e+L_PE(I t ));
[0029] e is the input.
[0030] Furthermore, the spectral attention module is inserted into the U-TAE network framework and replaces the original temporal attention module, ultimately forming a basic multi-temporal semantic segmentation network framework. The module combination and network construction method are as follows:
[0031] Step 301: Insert the spectral attention module into the basic network front end to form a basic multi-temporal semantic segmentation network framework The supervision signal is constructed by label data Y to guide the network optimization process; the final input image data is a vector X of size T×C×H×W, where T represents the length of the time series, C represents the number of channels, H represents the image length, and W represents the image width; the spectral attention module obtains X of the same size as X c ;
[0032] Step 302: X c Input into the spatial encoder for spatial feature extraction, the spatial encoder ε l There are four layers in total, l is the number of layers, and the input of each layer is the output e of the previous layer l-1 , and apply the same operation to each time series, and finally obtain the size of T×C at different scales l ×H l ×W l The spatial characteristics of C l Represents the number of channels at different scales, H l Represents the image length at different scales, W l Represents the image width at different scales. The formula for this step is as follows:
[0033]
[0034] Step 303: Use the improved lightweight temporal attention encoder L_TAE to extract the temporal attention a of the lth layer l ; Among them, the original triangular position code is replaced by a learnable position code L_P to obtain the image time I t Encoding is performed to inject temporal position information into the feature vector, enhancing the learning ability of the temporal encoder for crop phenological characteristics;
[0035] a 4 =L_TAE(e 4 +L_PE(I t ));
[0036] Step 304: Get a l Perform linear interpolation and multiply it with the feature vector obtained by the spatial encoder through 1*1 convolution Processing is performed to obtain a size of C l ×H l ×W lThe eigenvector f l ;
[0037] a l =resize a 4 to H l ×W l ;
[0038]
[0039] resize a 4 to H l ×W l Indicates that a 4 Adjust to H l ×W l size;
[0040] Step 305: The feature vector f learned in the previous step l Input to decoder D l Perform feature decoding to obtain spatiotemporal features at all scales, decoder D l There are four layers in total, l is the number of layers, and transposed convolution is used in each decoder layer. To perform upsampling operation, each layer input is the output of the previous layer, and the size is C l ×H l ×W l The eigenvector d l , and finally correlate the decoder output with the encoder output in a U-net manner; finally, input the obtained feature vector into the convolutional layer Conv_out to obtain the final prediction result probability output, and obtain the final prediction result through softmax; the formula description of this step is as follows:
[0041]
[0042] y=softmax(Conv_out(d 4 )).
[0043] Furthermore, to improve the robustness of the model to incomplete time series data, the feature reconstruction strategy based on mask-distillation is constructed as follows:
[0044] Step 401: Randomly mask the input time series X in the time dimension, with the mask ratio being α to obtain the masked time series X m ,The formula description of this step is as follows:
[0045] X m =random_mask(X);
[0046] Step 402: Use two independent multi-temporal semantic segmentation networks and They serve as student models and teacher models respectively, where the input of the student model is the masked time series X m The input of the teacher model is the unmasked time series X and label Y;
[0047] Step 403: Use offline distillation mode to train the teacher model first, using binary cross entropy loss in the process BCE Back propagation is performed to guide the model parameter update; the cross entropy loss is formulated as follows, where y i is the label value, p(y i ) is the prediction result, N represents the number of pixels:
[0048]
[0049] Step 404: Construct three loss functions, namely binary cross entropy loss BCE , distillation loss KD , feature reconstruction loss loss FR ; The binary cross entropy is calculated using the label and the predicted value, which is exactly the same as the loss calculation in the previous step; the distillation loss is calculated by the teacher model and the student model prediction value Y s With Y t Calculation; Feature reconstruction loss is calculated by the intermediate feature f between the teacher model and the student model s With f t Calculation; the formula description of each loss is as follows:
[0050]
[0051] Step 405: Train the student model. In this process, the three loss functions are weighted to obtain the final loss value loss. All To perform back propagation to guide the model parameter update; the formula description of this step is as follows:
[0052] loss All =αloss BCE +βloss KD +γloss RF .
[0053] Furthermore, the final trained model was used to map cultivated land cover in my country's cloudy and rainy subtropical monsoon climate region. The cultivated land mapping method is as follows:
[0054] Step 501: Divide the study area and select data areas with relatively complete time series to create a full time series dataset; mask the full time series dataset to obtain a time series missing simulation dataset;
[0055] Step 502: Use the strategy of step 4 to train the model obtained in step 3 to obtain the final intelligent farmland monitoring model;
[0056] Step 503: Use the obtained model to predict the entire area and splice the prediction results to obtain the final cultivated land monitoring results.
[0057] The beneficial effects of the present invention are as follows:
[0058] The present invention combines the Squeeze-and-Excitation module with the self-attention mechanism to model the relationships between bands in multispectral data, allowing the network to adaptively adjust the response values of each band. This strengthens the model's focus on sensitive bands and reduces the noise impact caused by other bands during the farmland extraction process.
[0059] To address the problem of biased prior knowledge of triangular position coding in adjacent time phases in the temporal feature module based on the self-attention mechanism, the present invention replaces it with learnable position coding, thereby better learning the periodic change characteristics brought about by the phenological attributes of crops and artificial agricultural activity patterns during the temporal feature modeling process.
[0060] To address the data missing problem caused by cloud obstruction and other reasons during the actual application of intelligent farmland monitoring methods, the present invention develops a feature reconstruction learning strategy based on mask-distillation, thereby improving the robustness of the multi-temporal model to incomplete time series data, allowing the model to better adapt to the complex and changeable farmland monitoring environment. BRIEF DESCRIPTION OF THE DRAWINGS
[0061] Figure 1 is a flow chart of the present invention;
[0062] Figure 2 The spectral attention mechanism proposed in this invention;
[0063] Figure 3 The improved temporal attention mechanism of the present invention;
[0064] Figure 4 A schematic diagram of the overall network architecture after modification used in the present invention;
[0065] Figure 5 This is the mask-reconstruction feature reconstruction learning strategy proposed in this invention. DETAILED DESCRIPTION
[0066] The present invention will be further described below with reference to the accompanying drawings, but the present invention is not limited in any way. Any changes or substitutions made based on the teachings of the present invention fall within the scope of protection of the present invention.
[0067] like Figure 1 As shown, the intelligent farmland monitoring method for multi-temporal data disclosed in the present invention includes the following steps:
[0068] Step 1: Spectral attention module construction;
[0069] Step 2: Basic model design including temporal attention module;
[0070] Step 3: Module combination and network construction;
[0071] Step 4: Construct a feature reconstruction strategy based on mask-distillation;
[0072] Step 5: Mapping of cultivated land cover.
[0073] In the process of intelligent monitoring of cultivated land for multi-temporal data, considering that some bands are more sensitive to the changes in crops in cultivated land and not all spectral bands are beneficial to cultivated land extraction, in order to make the model pay more attention to important bands and alleviate the interference caused by other spectral bands, an attention mechanism that can automatically adjust the importance of different spectral channels is constructed. The specific structure is as follows: Figure 2 The specific steps are as follows:
[0074] Step 101: Define the time series multispectral data as a vector X of size T×C×H×W, and the time series vector X consists of T x vectors of size C×H×W. t In the spectral attention module, each x t Split them out and perform spectral attention weighted operations separately.
[0075] Step 102: First, t Perform maximum pooling and average pooling to provide a general description of each band of the sample from multiple angles. The formula for this step is as follows:
[0076] f Max ,f Avg =AvgPool(x t ),MaxPool(x t );
[0077] Step 103: Use the self-attention mechanism to operate on the above features respectively, that is, use three independent 1*1 convolutions conv_K, conv_Q, and conv_V to map the input features to construct key, query, and value, and use the matrix multiplication of key and query to obtain its channel attention, and then multiply it with the value to obtain the weighted feature f Max_out With f Avg_out Thus, the relationship between each band is obtained. The formula of this step is as follows:
[0078] f Max_out =Softmax(conv_K(f Max )×conv_Q(f Max ) T )×conv_V(f Max );
[0079] f Avg_out =Softmax(conv_K(f Avg )×conv_Q(f Avg ) T )×conv_V(f Avg );
[0080] Step 104: The obtained f Max_out With f Avg_out Add and normalize to get the final eigenvector X c Its size is the same as the original size of X, which is T×C×H×W. The formula for this step is as follows:
[0081] X c =sigmoid(f Max_out +f Avg_out );
[0082] Since the temporal feature extraction model based on the self-attention mechanism uses triangular position coding to inject temporal information into the data, it often shows a strong correlation between adjacent time phases, while ignoring the periodic change characteristics caused by crop phenology and artificial agricultural activities within the cultivated land. Therefore, the present invention replaces the position coding in the temporal feature extraction module based on the self-attention mechanism with a learnable position coding. The improved temporal feature extraction module based on the self-attention mechanism is as follows: Figure 3 The specific operation is to replace the original triangular position code with the learnable position code L_PE to obtain the image time I t Encoding is performed to inject temporal position information into the feature vector, thereby enhancing the learning ability of the temporal encoder for crop phenological characteristics.
[0083] a=L_TAE(e+L_PE(I t ));
[0084] Insert the spectral attention module into the U-TAE network framework and replace the original temporal attention module to form the basic multi-temporal semantic segmentation network framework. The final framework is as follows Figure 4 As shown, the specific steps include:
[0085] Step 301: Insert the spectral attention module into the basic network front end to form a basic multi-temporal semantic segmentation network framework The supervision signal is constructed by label data Y to guide the network optimization process. The final input image data is a vector X of size T×C×H×W, where T represents the length of the time series, C represents the number of channels, H represents the image length, and W represents the image width; the spectral attention module obtains an X of the same size as X. c .
[0086] Step 302: X c Input into the spatial encoder for spatial feature extraction, the spatial encoder ε l There are four layers in total, where the input of each layer is the output e of the previous layer l-1 , and apply the same operation to each time series, and finally obtain the size of T×C at different scales l ×H l ×W l The spatial characteristics of the step are described as follows:
[0087]
[0088] Step 303: Use the improved lightweight temporal attention encoder L_TAE to extract temporal attention a l Among them, the original triangular position code is replaced by a learnable position code L_P to obtain the image time I t Encoding is performed to inject temporal position information into the feature vector, thereby enhancing the learning ability of the temporal encoder for crop phenological characteristics.
[0089] a 4 =L_TAE(e 4 +L_PE(I t ));
[0090] Step 304: Get a l Perform linear interpolation and multiply it with the feature vector obtained by the spatial encoder through 1*1 convolution Processing is performed to obtain a size of C l ×H l ×W l The eigenvector f l .
[0091] a l =resize a 4 to H l ×W l ;
[0092]
[0093] Step 305: The feature vector f learned in the previous step l Input to decoder D l Perform feature decoding to obtain spatiotemporal features at all scales, decoder D l There are four layers in total, and each decoder layer uses transposed convolution To perform upsampling operation, each layer input is the output of the previous layer, and the size is C l ×H l ×W l The eigenvector d l , and finally the decoder output is correlated with the encoder output in a U-net manner. The resulting feature vector is input into the convolutional layer Conv_out to obtain the final prediction probability output, and the final prediction result is obtained through softmax. The formula for this step is as follows:
[0094]
[0095] y=softmax(Conv_out(d 4 ));
[0096] Considering the data missing problem caused by cloud occlusion and other reasons in the actual application of the intelligent farmland monitoring method, a feature reconstruction learning strategy based on mask-distillation is used to improve the robustness of the multi-temporal model to incomplete time series data, so that the model can better adapt to the complex and changeable farmland monitoring environment. The overall strategy is as follows Figure 5 The specific operations are as follows:
[0097] Step 401: Randomly mask the input time series X in the time dimension, with the mask ratio being α to obtain the masked time series X m ,The formula description of this step is as follows:
[0098] X m =random_mask(X);
[0099] Step 402: Use two independent multi-temporal semantic segmentation networks and They serve as student models and teacher models respectively, where the input of the student model is the masked time series X mThe input of the teacher model is the unmasked time series X and label Y;
[0100] Step 403: Use offline distillation mode to train the teacher model first, using binary cross entropy loss in the process BCE Back propagation is performed to guide the model parameter update. The cross entropy loss is formulated as follows, where y i is the label value, p(y i ) is the prediction result, N represents the number of pixels:
[0101]
[0102] Step 404: Construct three loss functions, namely binary cross entropy loss BCE , distillation loss KD , feature reconstruction loss loss FR The binary cross entropy is calculated using the label and the predicted value, which is exactly the same as the loss calculation in the previous step; the distillation loss is calculated by the teacher model and the student model prediction value Y s With Y t Calculation; Feature reconstruction loss is calculated by the intermediate feature f between the teacher model and the student model s With f t Calculation. The formulas of each loss are described as follows:
[0103]
[0104] Step 405: Train the student model. In this process, the three loss functions are weighted to obtain the final loss value loss. All Back propagation is used to guide the model parameter update. The formula for this step is as follows:
[0105] loss All =αloss BCE +βloss KD +γloss RF ;
[0106] The obtained model is used for intelligent farmland monitoring, that is, to complete farmland mapping. The specific steps are as follows:
[0107] Step 501: Divide the study area and select data areas with relatively complete time series to produce a full time series data set; mask the full time series data set to obtain a time series missing simulation data set.
[0108] Step 502: Use the strategy of step 4 to train the model obtained in step 3 to obtain the final intelligent farmland monitoring model;
[0109] Step 503: Use the obtained model to predict the entire area and splice the prediction results to obtain the final cultivated land monitoring results.
[0110] To demonstrate the effectiveness of the proposed method, experiments were conducted in Hunan Province, China, a subtropical monsoon region. Multispectral data from the Sentinel-2 satellite was collected as the image source and used to annotate cultivated land. Fourteen monthly images were collected from April 2019 to May 2020, and ten bands of the Sentinel-2 multispectral data at 10m and 20m resolution were used as the final imagery.
[0111] Hunan Province was evenly divided into 121 data units at 0.5° latitude and longitude intervals. Data was downloaded and processed within each unit. First, data was selected within these units based on the missingness of remote sensing time series data. Considering both the data volume and uniform coverage of the study area, approximately 30 units with missing time series for 0-2 months were selected throughout the study area to serve as the full-temporal dataset mentioned in Step 5. Finally, after grid cropping, 15,488 data samples of 256x256 size were obtained. One-tenth of this data was used for model training, and the remaining data was used to verify model accuracy. Random masks were used to mask 25% to 75% of the data in the temporal dimension to simulate missing time series. Seven units with missing time series ranging from 35% to 92% were selected from the entire region as the real dataset to verify the mask-distillation feature reconstruction. Specific parameters are shown in Table 1.
[0112] Table 1 Dataset description
[0113] Dataset name Number of coverage areas Number of samples Full time series dataset 30 15488 Time series missing simulation data set 30 15488 Time series missing real data set 7 3388
[0114] To demonstrate the effectiveness of this method, experiments were conducted on the three datasets mentioned above. The full time series dataset was compared with other model architectures to demonstrate the improvement of this method in farmland monitoring accuracy when the time series data is complete. The simulated dataset with missing time series and the real dataset with missing time series demonstrate the improvement of this method compared to other methods when the time series is incomplete. In this process, the experiments on the full time series dataset used the following three indicators for accuracy verification, and the mIou indicator, which is commonly used to measure the accuracy of farmland monitoring, was used for accuracy verification in the simulated dataset with missing time series and the real dataset with missing time series:
[0115] mIoU value, intersection over union ratio:
[0116]
[0117] OA represents the overall accuracy of all pixels and is defined as follows:
[0118]
[0119] The F1 value is the weighted average of precision and recall, comparing the two indicators comprehensively:
[0120]
[0121] TP represents the number of pixels whose predicted results are correct and whose true results are positive; FN represents the number of pixels whose predicted results are incorrect and whose true results are positive; FP represents the number of pixels whose predicted results are incorrect and whose true results are negative.
[0122] The comparison accuracy with other networks in the full-time dataset is shown in Table 2:
[0123] Table 2 Comparison of the full-time dataset with other network architectures
[0124] Method Name mIou OA F1-score ConvLSTM 77.63% 90.92% 79.74% UConvLSTM 77.66% 91.17% 79.50% BUConvLSTM 78.08% 91.22% 80.12% FPN-ConvLSTM 78.11% 91.45% 79.89% U_Net_3D 78.11% 91.19% 80.21% U_TAE 78.32% 91.29% 80.09% Ours 78.54% 91.62% 80.39%
[0125] The comparative accuracy of mIou under different ratios in the time series missing simulation dataset is shown in Table 3:
[0126] Table 3 Comparison of other network architectures for the time series missing simulation dataset
[0127]
[0128]
[0129] The comparative accuracy of mIou in different data missingness regions in the real time series missing data set is shown in Table 4, where X / 14 means that there are X months of missing data in a total of 14 months; S_X represents the region number.
[0130] Table 4 Comparative accuracy of mIou in different data missing degree regions in real time series missing data sets
[0131]
[0132] The method proposed in this invention outperforms the existing methods in three or one indicators on three data sets.
[0133] The beneficial effects of the present invention are as follows:
[0134] The present invention combines the Squeeze-and-Excitation module with the self-attention mechanism to model the relationships between bands in multispectral data, allowing the network to adaptively adjust the response values of each band. This strengthens the model's focus on sensitive bands and reduces the noise impact caused by other bands during the farmland extraction process.
[0135] To address the problem of biased prior knowledge of triangular position coding in adjacent time phases in the temporal feature module based on the self-attention mechanism, the present invention replaces it with learnable position coding, thereby better learning the periodic change characteristics brought about by the phenological attributes of crops and artificial agricultural activity patterns during the temporal feature modeling process.
[0136] To address the data missing problem caused by cloud obstruction and other reasons during the actual application of intelligent farmland monitoring methods, the present invention develops a feature reconstruction learning strategy based on mask-distillation, thereby improving the robustness of the multi-temporal model to incomplete time series data, allowing the model to better adapt to the complex and changeable farmland monitoring environment.
[0137] As used herein, the word "preferred" is intended to serve as an example, instance, or illustration. Any aspect or design described herein as "preferred" is not necessarily to be construed as advantageous over other aspects or designs. Rather, the use of the word "preferred" is intended to present concepts in a concrete manner. As used in this application, the term "or" is intended to mean an inclusive "or" rather than an exclusive "or." That is, unless otherwise specified or clear from the context, "X employs A or B" is intended to mean any of the naturally inclusive permutations. That is, if X employs A; X employs B; or X employs both A and B, then "X employs A or B" is satisfied in any of the foregoing examples.
[0138] Moreover, although the present disclosure has been shown and described with respect to one or implementation, those skilled in the art will think of equivalent variations and modifications based on reading and understanding of this specification and the accompanying drawings. The present disclosure includes all such modifications and variations and is limited only by the scope of the appended claims. In particular, with respect to the various functions performed by the above-mentioned components (such as elements, etc.), the terms used to describe such components are intended to correspond to any component (unless otherwise indicated) that performs the specified function of the component (such as it is functionally equivalent), even if structurally different from the disclosed structure that performs the function in the exemplary implementation of the present disclosure shown herein. In addition, although the specific features of the present disclosure have been disclosed with respect to only one of several implementations, such features can be combined with one or other features of other implementations that can be desired and advantageous for a given or specific application. Moreover, insofar as the terms "including", "having", "containing" or their variations are used in specific embodiments or claims, such terms are intended to be included in a manner similar to the term "comprising".
[0139] The functional units in the embodiments of the present invention may be integrated into a single processing module, or each unit may exist physically separately, or multiple or more units may be integrated into a single module. The aforementioned integrated module may be implemented in the form of hardware or in the form of a software functional module. If the integrated module is implemented in the form of a software functional module and sold or used as an independent product, it may also be stored in a computer-readable storage medium. The aforementioned storage medium may be a read-only memory, a magnetic disk, or an optical disk, etc. The aforementioned devices or systems may execute the storage method in the corresponding method embodiment.
[0140] In summary, the above embodiment is one implementation method of the present invention, but the implementation method of the present invention is not limited to the described embodiment. Any other changes, modifications, substitutions, combinations, and simplifications that deviate from the spirit and principles of the present invention should be equivalent replacement methods and are included in the scope of protection of the present invention.
Claims
1. An intelligent monitoring method for cultivated land resources based on multi-temporal remote sensing data, characterized in that: The following steps are involved: Step 1: Build a spectral attention module to increase the importance of sensitive bands for cultivated land and reduce interference from other bands; Step 2: Optimize the temporal feature extraction module based on the self-attention mechanism, improve the original L_TAE attention mechanism, and replace the original fixed position encoding with a learnable position encoding to mine the temporal features influenced by crop phenology and artificial agricultural activity patterns for modeling; Step 3: Module combination and basic model design: Using the U_TAE network architecture as the basic framework, the spectral attention module and the optimized temporal attention module are inserted into the network framework to form the basic multi-temporal semantic segmentation network framework; Step 4: Build a feature reconstruction strategy based on mask-distillation to improve the robustness of the model to incomplete time series data; Step 5: Mapping of cultivated land cover; The construction of the spectral attention module specifically includes: Step 101: defining the time series multispectral data as a vector X of size T×C×H×W, and the time series vector X consists of T elements x of size C×H×W t Composition, where T represents the time series length, C represents the number of channels, H represents the image length, and W represents the image width; Step 102: Place x t Perform maximum pooling and average pooling to provide a general description of each band of the sample from multiple angles; Step 103: Use the self-attention mechanism to operate on the above features respectively, that is, use three independent 1*1 convolutions conv_K, conv_Q, and conv_V to map the input features to construct key, query, and value respectively, and use the matrix multiplication of key and query to obtain its channel attention, and then multiply it with the value to obtain the weighted feature f Max_out With f Avg_out Thus, the relationship between each band is obtained; Step 104: The obtained f Max_out With f Avg_out Add and normalize to get the final eigenvector X c Its size is the same as the original size of X, which is T×C×H×W.
2. The intelligent monitoring method for cultivated land resources based on multi-temporal remote sensing data according to claim 1, characterized in that: The spectral attention module construction method In step 101, each x t Split them out and perform spectral attention weighted operations separately; The formulaic description of step 102 is as follows: f Max ,f Avg =AvgPool(x t ),MaxPool(x t ); AvgPool means average pooling, MaxPool means maximum pooling; The formulaic description of step 103 is as follows: f Max_out =Softmax(conv_K(f Max )×conv_Q(f Max ) T )×conv_V(f Max ); f Avg_out =Softmax(conv_K(f Avg )×conv_Q(f Avg ) T )×conv_V(f Avg ); The formulaic description of step 104 is as follows: X c =sigmoid(f Max_out +f Avg_out ); Sigmoid is the activation function.
3. The intelligent monitoring method for cultivated land resources based on multi-temporal remote sensing data according to claim 1, characterized in that: The self-attention mechanism is used to explore the mutual relationship between high-density time series images at the same spatial location, and the position encoding part is optimized to better model the time series characteristics of cultivated land. The improved lightweight temporal attention encoder L_TAE is used to extract the temporal attention a and modify it to replace the original triangular position encoding with the learnable position encoding L_PE to obtain the temporal attention of the image. t Encoding is performed to inject temporal position information into the feature vector, thereby enhancing the temporal encoder's ability to learn crop phenological characteristics; the temporal attention a is calculated as follows: a=L_TAE(e+L_OR(I t )): e is the input.
4. The intelligent monitoring method for cultivated land resources based on multi-temporal remote sensing data according to claim 1, characterized in that: The spectral attention module is inserted into the U-TAE network framework and replaces the original temporal attention module, ultimately forming the basic multi-temporal semantic segmentation network framework. The module combination and network construction method are as follows: Step 301: Insert the spectral attention module into the basic network front end to form a basic multi-temporal semantic segmentation network framework The supervision signal is constructed by label data Y to guide the network optimization process; the final input image data is a vector X of size T×C×H×W, where T represents the length of the time series, C represents the number of channels, H represents the image length, and W represents the image width; the spectral attention module obtains X of the same size as X c ; Step 302: X c Input into the spatial encoder for spatial feature extraction, the spatial encoder ε l There are four layers in total, l is the number of layers, and the input of each layer is the output e of the previous layer l-1 , and apply the same operation to each time series, and finally obtain the size of T×C at different scales l ×H l ×W l spatial characteristics; C l Represents the number of channels at different scales, H l Represents the image length at different scales, W l Represents the image width at different scales. The formula for this step is as follows: Step 303: Use the improved lightweight temporal attention encoder L_TAE to extract the temporal attention a of the lth layer l ; Among them, the original triangular position code is replaced by a learnable position code L_P to obtain the image time I t Encoding is performed to inject temporal position information into the feature vector, enhancing the learning ability of the temporal encoder for crop phenological characteristics; fence 4 =L_TAE(e 4 +L_OR(I t )): Step 304: Get a l Perform linear interpolation and multiply it with the feature vector obtained by the spatial encoder through 1*1 convolution Processing is performed to obtain a size of C l ×H l ×W l The eigenvector f l ; a l =resize a 4 to H l ×W l ; resize a 4 to H l ×W l Indicates that a 4 Adjust to H l ×W l size; Step 305: The feature vector f learned in the previous step l Input to decoder D l Perform feature decoding to obtain spatiotemporal features at all scales, decoder D l There are four layers in total, l is the number of layers, and transposed convolution is used in each decoder layer. To perform upsampling operation, each layer input is the output of the previous layer, and the size is C l ×H l ×W l The eigenvector d l , and finally correlate the decoder output with the encoder output in a U-net manner; finally, input the obtained feature vector into the convolutional layer Conv_out to obtain the final prediction result probability output, and obtain the final prediction result through softmax; the formula description of this step is as follows: y=softmax(Conv_out(d 4 ))。 5. The intelligent monitoring method for cultivated land resources based on multi-temporal remote sensing data according to claim 1, characterized in that: To improve the robustness of the model to incomplete time series data, the feature reconstruction strategy based on mask-distillation is constructed as follows: Step 401: Randomly mask the input time series X in the time dimension, with the mask ratio being α to obtain the masked time series X m ,The formula description of this step is as follows: X m =random_mask(X); Step 402: Use two independent multi-temporal semantic segmentation networks and They serve as student models and teacher models respectively, where the input of the student model is the masked time series X m The input of the teacher model is the unmasked time series X and label Y; Step 403: Use offline distillation mode to train the teacher model first, using binary cross entropy loss in the process BCE Back propagation is performed to guide the model parameter update; the cross entropy loss is formulated as follows, where y i is the label value, p(y i ) is the prediction result, N represents the number of pixels: Step 404: Construct three loss functions, namely binary cross entropy loss BCE , distillation loss KD , feature reconstruction loss loss FR ; The binary cross entropy is calculated using the label and the predicted value, which is exactly the same as the loss calculation in the previous step; the distillation loss is calculated by the teacher model and the student model prediction value Y s With Y t Calculation; Feature reconstruction loss is calculated by the intermediate feature f between the teacher model and the student model s With f t Calculation; the formula description of each loss is as follows: Step 405: Train the student model. In this process, the three loss functions are weighted to obtain the final loss value loss. All To perform back propagation to guide the model parameter update; the formula description of this step is as follows: loss All =αloss BCE +βloss KD +γloss RF 。 6. The intelligent monitoring method for cultivated land resources based on multi-temporal remote sensing data according to claim 1, characterized in that: The final trained model is used to map cultivated land cover in my country's subtropical monsoon climate region with cloudy and rainy weather. The cultivated land mapping method is as follows: Step 501: Divide the study area and select data areas with relatively complete time series to create a full time series dataset; mask the full time series dataset to obtain a time series missing simulation dataset; Step 502: Use the strategy of step 4 to train the model obtained in step 3 to obtain the final intelligent farmland monitoring model; Step 503: Use the obtained model to predict the entire area and splice the prediction results to obtain the final cultivated land monitoring results.
Citation Information
Patent Citations
Vegetable field monitoring method based on multi-source multi-temporal remote sensing image data
CN105447494A
Cultivated land information extraction method of high-resolution remote sensing image based on edge enhancement
CN114596502A