Multi-temporal remote sensing image crop classification method based on space-time attention U-shaped network

By constructing a spatiotemporal attention U-shaped network and utilizing the spatiotemporal feature information of multi-temporal remote sensing data, the problems of insufficient utilization of multi-temporal remote sensing data and the influence of cloud cover in existing technologies are solved, and high-precision crop classification and good generalization ability are achieved.

CN120808040AActive Publication Date: 2025-10-17DALIAN MARITIME UNIVERSITY

Patent Information

Application Number
CN202511051891.8
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-07-29
Publication Date
2025-10-17
Estimated Expiration
2045-07-29

AI Technical Summary

Technical Problem

Existing technologies find it difficult to fully utilize the temporal characteristics of multi-phase remote sensing data, and temporal remote sensing data are easily affected by factors such as cloud cover, resulting in reduced classification accuracy and requiring complex data preprocessing processes.

Method used

A multi-temporal remote sensing image crop classification method based on a spatiotemporal attention U-shaped network is constructed. The key growth period characteristics of crops are adaptively captured through the spatiotemporal attention mechanism, including a convolutional block attention module, a lightweight temporal attention encoder module, a dynamic upsampling module and an adaptive feature fusion module, reducing the dependence on complex preprocessing.

Benefits of technology

It achieves high-precision crop classification, improves computational efficiency and classification accuracy, has good generalization ability, and can adaptively process time series noise.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120808040A_ABST
    Figure CN120808040A_ABST
Patent Text Reader

Abstract

The invention belongs to the technical field of remote sensing image processing, and relates to a multi-temporal remote sensing image crop classification method based on a space-time attention U-shaped network. A convolutional block attention module, a lightweight time attention encoder module, a dynamic up-sampling module and an adaptive feature fusion module are integrated, crop classification is performed on multi-temporal remote sensing images, and in the training process of a network model, key growth period features of crops are adaptively captured through a space-time attention mechanism, so that cloud coverage interference is effectively suppressed, and the accuracy of crop classification is improved. And time sequence information and spatial feature information are fully utilized. Under the condition that the number of the training sets is sufficient, accurate classification of farmland crops based on multi-temporal remote sensing data is realized, and a good effect is achieved. The method shows high environment adaptability in space generalization and time generalization, and especially has remarkable advantages in the aspect of space generalization.
Need to check novelty before this filing date? Find Prior Art

Description

TECHNICAL FIELD

[0001] The present application relates to the technical field of remote sensing image processing, in particular to a multi-temporal remote sensing image crop classification method based on a spatio-temporal attention mechanism of deep learning. BACKGROUND

[0002] Traditional crop classification methods mainly rely on pixel-based classification algorithms and machine learning methods such as support vector machines, random forests, etc. These methods perform well in single-temporal remote sensing image classification, but have difficulty in fully exploiting the temporal characteristics of multi-temporal remote sensing data and are highly dependent on feature engineering. The development of deep learning technology has brought new breakthroughs to crop classification, but existing convolutional neural networks mainly process spatial and spectral dimensions and fail to fully utilize time series information, which is insufficient in handling long-time series crop phenology change characteristics.

[0003] Multi-temporal remote sensing data has important value in crop classification and can capture the spectral variation characteristics of crops in different growth periods. However, time series remote sensing data is easily affected by factors such as cloud cover, leading to temporal discontinuity and data loss, which seriously affects classification accuracy. In addition, existing methods often require complex data preprocessing procedures, which have high deployment thresholds in practical applications. Therefore, a crop classification method is needed that can fully utilize multi-temporal information and adaptively handle temporal noise. SUMMARY

[0004] The present application aims to provide a multi-temporal remote sensing image crop classification method based on a spatio-temporal attention U-shaped network, which utilizes a spatio-temporal attention mechanism to adaptively capture key growth period characteristics of crops, achieving high-precision crop classification without relying on complex preprocessing, and improving classification accuracy and computational efficiency.

[0005] To achieve the above functions, the present application designs a multi-temporal remote sensing image crop classification method based on a spatio-temporal attention U-shaped network, which performs the following steps for multi-temporal remote sensing images:

[0006] Obtain a multi-temporal remote sensing image dataset of crops, wherein the multi-temporal remote sensing image dataset includes a time series dataset, read remote sensing images of a crop-dominant plot, and label dataset labels and generalization set labels, divide the time series dataset, dataset labels, and generalization set labels into a training set, a validation set, a test set, and a spatial generalization set;

[0007] Construct a spatio-temporal attention U-shaped network model, including an encoder path and a decoder path;

[0008] Construct a convolution block attention module and a lightweight temporal attention encoder module in the encoder path, and integrate a dynamic upsampling module and an adaptive feature fusion module in the decoder path;

[0009] training the spatio-temporal attention U-shaped network model using the training set and the validation set;

[0010] calculating a loss value of a category map output by the spatio-temporal attention U-shaped network model and a real map using a cross-entropy loss function, and updating parameters of the spatio-temporal attention U-shaped network model through back propagation;

[0011] using the trained spatio-temporal attention U-shaped network model to perform crop classification on the test set, and outputting a crop classification result;

[0012] using the spatial generalization set, the data set label and the generalization set label to test the generalization ability of the trained model, and evaluating the spatial generalization ability of the spatio-temporal attention U-shaped network model;

[0013] using the trained spatio-temporal attention U-shaped network model to perform time generalization ability test on the data set of other years, and evaluating the time generalization ability of the network model.

[0014] Further, the convolution block attention module comprises:

[0015] extracting channel statistical information using global average pooling and maximum pooling, generating attention weights M c (X) of the channel dimension through a multi-layer perceptron and a sigmoid function, and the formula is as follows:

[0016] M c (X)=σ(MLP1(AvgPool(X))+MLP2(MaxPool(X)))

[0017] wherein, X is an input time sequence feature map, σ is a sigmoid function, MLP1 and MLP2 are different multi-layer perceptrons, AvgPool and MaxPool are average pooling operation and maximum pooling operation respectively;

[0018] performing element-wise multiplication on the input time sequence feature map using the attention weights M c (X) of the channel dimension, performing global pooling and maximum pooling operations, and obtaining a spatial position importance map M s (X) based on a sigmoid function processing:

[0019] M s (X)=σ(Conv 7×7 ([AvgPool(X⊙M c (X));MaxPool(X⊙M c (X))]))

[0020] wherein, Conv 7×7is a convolutional layer with a 7×7 convolution kernel, and ⊙ is element-wise multiplication;

[0021] Use the importance map of spatial location M s (X) and the attention weight M of the channel dimension c After element-by-element multiplication of (X) and the input time series feature map, the final output feature map X is obtained. CBA , the process is expressed by the following formula:

[0022] X CBA =X⊙M c (X)⊙M s (X⊙M c (X))

[0023] Furthermore, when building a lightweight temporal attention encoder module:

[0024] The time series feature map X CBA After point-by-point convolution and group normalization, the temporal feature map X is obtained d , add sine-cosine position encoding to get the temporal position embedding feature map X pos , the formula is as follows:

[0025] X pos =X d +PE(pos,i)

[0026]

[0027] Among them, PE(pos,i) is the position encoding function, pos is the time position index, i is the channel index, d m is the dimension of the feature map; the time dimension is weighted by the multi-head self-attention mechanism to obtain the adaptive learning attention weight Attention, and the attention weight is combined with the time series feature map X pos Perform dot product operation to obtain feature map X attn , the formula is as follows:

[0028]

[0029] X attn =Attention·X pos

[0030] Among them, n is the number of attention heads, Q (n) is a fixed learnable vector, K (n) is the time series feature graph X pos After the linear layer transformation, d k is the key-value dimension;

[0031] The output feature map X is obtained by processing through a multi-layer perception LTAE , and the formula is as follows:

[0032] X LTAE =MLP2·RELU(MLP1·X attn )

[0033] wherein, RELU is a nonlinear activation function;

[0034] Attention weights Attention will also be applied in different scale skip connection processing, first up-sampling to the corresponding scale of the time sequence feature map X mul , weighted sum to get the skip connection feature map X skip , the formula is as follows:

[0035]

[0036] wherein, T is the time step.

[0037] Further, when constructing the dynamic up-sampling module:

[0038] The feature map of the skip connection feature map X skip after the offset convolution layer and the feature map after the range adjustment convolution layer and then nonlinear processing by 0.5 times Sigmoid function Dot product, and add the original grid position alpha, get the offset Delta, the formula is as follows:

[0039] Delta = PS(Conv offset (X skip )·sigma(Conv scope (X skip ))·0.5)+alpha

[0040] wherein, PS is pixel recombination, Conv offset is offset convolution layer, Conv scope is range adjustment convolution layer;

[0041] The skip connection feature map X skip and the offset Delta are calculated using the bilinear interpolation grid sampling function to obtain the feature map X grid , the formula is as follows:

[0042] X grid =Grid_Sample(X skip ,Delta)

[0043] wherein, Grid_Sample is a grid sampling function.

[0044] Further, when constructing the adaptive feature fusion module: up-sampling feature map X gridand the skip connection feature map X skip The feature map X' is obtained by point-by-point convolution dimension reduction processing grid and Y' skip The element-level product of the two feature maps is calculated, a similarity map is generated through convolution operation and a Sigmoid function, and adaptive fusion is performed based on the similarity map to obtain the feature map X AFF , and the formula is as follows:

[0045] X AFF = (1-sigma (Conv (X' grid * X' skip ))) * X grid + sigma (Conv (X' grid * X' skip )) * X skip .

[0046] Further, the multi-temporal remote sensing image data is Sentinel-2 satellite multispectral image data, including 10 bands of visible light band, near-infrared band, red edge band, narrow wave near-infrared band and short wave infrared band.

[0047] Further, the convolution block attention module includes a channel attention submodule and a spatial attention submodule, the channel attention submodule extracts channel statistical information through global average pooling and maximum pooling, uses a multi-layer perceptron with channel dimension reduction processing to generate channel dimension attention weights, the spatial attention submodule uses a convolution kernel to generate a spatial position importance map, and the channel and spatial attention are applied in sequence and fused with the original feature through a residual connection.

[0048] Further, the lightweight time attention encoder module performs convolution processing on the input time sequence feature map to obtain a feature map with a dimension of d m , rearranges and group normalizes the feature tensor, adds a sine-cosine position encoding enhancement model to enhance the perception of time position, performs weighted calculation on the time dimension through a multi-head self-attention mechanism, and aggregates the time sequence features based on the adaptive learning attention weights.

[0049] Further, the dynamic upsampling module calculates a dynamic offset through an offset convolution layer, provides an initial sampling position for each output pixel through a range adjustment convolution layer, and performs dynamic upsampling under the guidance of the offset to obtain an output feature map, and integrates a post-processing convolution layer and a residual connection mechanism to enhance the feature expression capability.

[0050] Further, the adaptive feature fusion module respectively processes the up-sampling feature map and the skip connection feature map through convolution dimension reduction, calculates the element-level product of the two feature maps, generates a similarity map through convolution operation and a Sigmoid function, performs adaptive fusion based on the similarity map, and further improves the feature extraction capability through a convolution layer and a residual connection.

[0051] By adopting the technical scheme, the application discloses a multi-temporal remote sensing image crop classification method based on a spatio-temporal attention mechanism of deep learning, constructs a spatio-temporal attention U-shaped network model, classifies crops in multi-temporal remote sensing images, integrates a convolution block attention module, a lightweight time attention encoder module, a dynamic up-sampling module and an adaptive feature fusion module in an encoder and a decoder path respectively, fully utilizes the spatio-temporal feature information of multi-temporal remote sensing data, can adaptively process time sequence noise such as cloud coverage, does not need a complex data preprocessing process, realizes high-precision crop classification and good generalization ability, and provides an effective technical scheme for agricultural remote sensing monitoring. BRIEF DESCRIPTION OF DRAWINGS

[0052] In order to more clearly illustrate the technical solutions in the embodiments of the present application or the prior art, the drawings needed to be used in the embodiments or prior art description will be briefly introduced. Obviously, the drawings in the following description are only some embodiments described in the present application, and other drawings can be obtained by those skilled in the art without creative labor.

[0053] Figure 1 A flowchart of a multi-temporal remote sensing image crop classification method based on a spatio-temporal attention U-shaped network provided by the present application is shown in the figure.

[0054] Figure 2 The overall architecture diagram of the spatio-temporal attention U-shaped network model provided by the present application is shown in the figure.

[0055] Figure 3 The structure diagram of the convolution block attention module of the present application is shown in the figure.

[0056] Figure 4 The structure diagram of the lightweight time attention encoder module of the present application is shown in the figure.

[0057] Figure 5 The structure diagram of the dynamic up-sampling module of the present application is shown in the figure.

[0058] Figure 6 The structure diagram of the adaptive feature fusion module of the present application is shown in the figure.

[0059] Figure 7Schematic diagram of the classification results of the classification method of the present invention and the comparative classification method, wherein (ah) corresponds to the satellite image and the prediction map of each method. DETAILED DESCRIPTION

[0060] To make the technical solutions and advantages of the present invention more clear, the technical solutions in the embodiments of the present invention are clearly and completely described below in conjunction with the accompanying drawings in the embodiments of the present invention:

[0061] like Figure 1 The method for crop classification in multi-temporal remote sensing images based on a spatiotemporal attention U-shaped network is shown, and includes the following steps:

[0062] Obtain a multi-temporal remote sensing image dataset of crops, including a time series dataset. Read remote sensing images of plots dominated by crops, annotate them with dataset labels and generalization set labels, and divide the time series dataset, dataset labels, and generalization set labels into a training set, a validation set, a test set, and a spatial generalization set.

[0063] Construct a spatiotemporal attention U-shaped network model, including an encoder path and a decoder path;

[0064] Integrate convolutional block attention module and lightweight temporal attention encoder module in the encoder path;

[0065] Applying a dynamic upsampling module and an adaptive feature fusion module in the decoder path;

[0066] Use the training set and validation set data to train the spatiotemporal attention U-shaped network model;

[0067] Use the cross entropy loss function to calculate the loss value between the category map output by the spatiotemporal attention U-type network model and the real map, and backpropagate to update the parameters of the spatiotemporal attention U-type network model;

[0068] Use the trained model to classify crops on the test set data and output the crop classification results;

[0069] The overall architecture of the spatiotemporal attention U-shaped network model is shown in the figure Figure 2 As shown in the figure, a spatiotemporal attention U-shaped network model is constructed, consisting of an encoder path and a decoder path. The encoder path integrates a convolutional block attention module and a lightweight temporal attention encoder module; the decoder path applies a dynamic upsampling module and an adaptive feature fusion module. This network model is built in Python using the PyTorch deep learning library.

[0070] The structural diagram of the convolution block attention module is as follows Figure 3 Specifically, the following steps are included:

[0071] Use global average pooling and maximum pooling to extract channel statistics, and then generate the attention weight M of the channel dimension after passing it through a multi-layer perceptron and a sigmoid function. c (X), the formula is as follows:

[0072] M c (X)=σ(MLP1(AvgPool(X))+MLP2(MaxPool(X)))

[0073] Use the channel-dimensional attention weight M c (X) is element-wise multiplied with the input temporal feature map, and then global pooling and maximum pooling are performed. Then, a 7×7 convolution kernel is used for convolution, and finally a sigmoid function is added to obtain the spatial position importance map M. s (X), the formula is as follows:

[0074] M s (X)=σ(Conv 7×7 ([AvgPool(X⊙M c (X));MaxPool(X⊙M c (X))]))

[0075] Use the importance map of spatial location M s (X) and the attention weight M of the channel dimension c After element-by-element multiplication of (X) and the input time series feature map, the final output feature map X is obtained. CBA , the whole process is expressed by the following formula:

[0076] X CBA =X⊙M c (X)⊙M s (X⊙M c (X)

[0077] The structural diagram of the lightweight temporal attention encoder module is as follows Figure 4 Specifically, the following steps are included:

[0078] The time series feature map X CBA After point-by-point convolution and group normalization, the temporal feature map X is obtained d , add sine-cosine position encoding to get the temporal position embedding feature map X pos , the formula is as follows:

[0079] X pos =X d +PE(pos,i)

[0080]

[0081] Then, the time dimension is weighted and calculated by the multi-head self-attention mechanism to obtain the attention weight Attention learned adaptively, and the attention weight is combined with the time sequence feature map X pos to perform dot product operation to obtain the feature map X attn , and the formula is as follows:

[0082]

[0083] X attn = Attention X pos

[0084] Finally, the output feature map X LTAE is obtained by processing through the multi-layer perception, and the formula is as follows:

[0085] X LTAE = MLP2 RELU (MLP1 X attn )

[0086] The attention weight Attention will also be applied in the different scale skip connection processing, first up-sampling it to the time sequence feature map X mul of the corresponding scale, and then performing weighted summation to obtain the skip connection feature map X skip , and the formula is as follows:

[0087]

[0088] The structure diagram of the dynamic up-sampling module is shown in Figure 5 , and specifically includes the following steps:

[0089] The feature map of the skip connection feature map X skip after the offset convolution layer is dot multiplied with the feature map after the range adjustment convolution layer and the nonlinear processing of the 0.5 times Sigmoid function, and the original grid position a is added to obtain the offset D, and the formula is as follows:

[0090] D = PS (Conv offset (X skip ) · sigma (Conv scope (X skip )) · 0.5) + a

[0091] The skip connection feature map X skip is calculated with the offset D using the bilinear interpolation grid sampling function to obtain the feature map X grid , and the formula is as follows:

[0092] X grid = Grid_Sample(Xskip ,Δ)

[0093] The structural diagram of the adaptive feature fusion module is shown in Figure 6 , and specifically includes the following steps:

[0094] The up-sampling feature map X grid and the skip connection feature map X skip are respectively processed by point-wise convolution dimension reduction to obtain feature maps X' grid and Y' skip '. Then, the element-level product of the two features is calculated, and a similarity map is generated through convolution operation and Sigmoid function, and adaptive fusion is performed based on the similarity map to obtain the feature map X AFF , as shown in the following formula:

[0095] X AFF = (1-σ(Conv(X' grid ⊙X' skip )))⊙X grid +σ(Conv(X' grid ⊙X' skip ))⊙X skip

[0096] The spatio-temporal attention U-shaped network model is trained for crop classification learning, and multi-temporal remote sensing image data and corresponding crop class labels are used to train the model. The following steps are implemented during the training process:

[0097] S1. For the given multi-temporal remote sensing image data, the spatial features and channel information are extracted through the convolution block attention module.

[0098] S2. Use the lightweight time attention encoder module to perform weighted calculation on the time dimension to obtain the adaptive learning attention weight.

[0099] S3. Apply the dynamic up-sampling module in the decoder path to realize accurate up-sampling and improve the accuracy of feature reconstruction.

[0100] S4. Use the adaptive feature fusion module to integrate feature information at different levels to reduce the semantic gap.

[0101] S5. Use the cross-entropy loss function to calculate the loss value of the model output and the true label, and update the model parameters through back propagation.

[0102] The above steps are repeatedly trained for 80 epochs, and the spatio-temporal attention U-shaped network model learns multi-temporal crop classification features to achieve the best classification effect.

[0103] The spatio-temporal attention U-shaped network model is tested using the data and corresponding labels in the test set, and the classification effect is calculated.

[0104] The multi-temporal remote sensing image dataset used in this embodiment is Sentinel-2 multispectral satellite image data covering three years. The Sentinel-2 satellite is equipped with a multispectral instrument that can collect spectral information of the earth's surface at different spatial resolutions of 10m, 20m and 60m. This embodiment selects 10 bands, including 10m resolution visible and near-infrared bands, and 20m resolution red edge, narrow wave near-infrared and short wave infrared bands. The data of 20m resolution bands will be resampled to 10m resolution using bilinear interpolation. The dataset labels and generalization set labels include four classes of rice, corn, soybeans and background, and the dataset includes 800 training (70%), validation (15%) and test (15%) plots, 700 spatial generalization test plots, each plot size is 128x128 pixels.

[0105] The classification method comparison experiment respectively uses Ms-TTC, UTempoNet, ConvGRU, ConvLSTM, Unet3d, support vector machine (SVM) and random forest (RF) and other methods, and the accuracy analysis of the experimental results uses overall accuracy (Overall Accuracy, OA), mean intersection over union (mIoU), F1 score (F1-score), precision (Precision) and recall (Recall) and other indicators.

[0106] The model training settings are as follows:

[0107] The encoder width of the spatio-temporal attention U-shaped network model is set to [64, 64, 64, 128], the decoder width is set to [32, 64, 64, 128], and the output convolution layer is set to [32, 4], corresponding to four class outputs.

[0108] The number of attention heads in the lightweight temporal attention encoder module is set to 16, the feature dimension is set to 256, and the key-value space dimension is set to 4.

[0109] The convolution kernel size in the model downsampling stage is 4, the step is 2, the padding is 1, and the encoder normalization layer uses GroupNorm.

[0110] The Adam optimizer is used in the training process, the initial learning rate is set to 0.001, the weight decay is set to 5e-4, the batch size is set to 4, the training rounds are set to 80 rounds, and the parameters of the remaining comparison methods are configured according to the original environment.

[0111] Under this condition, three repeated experiments were carried out, and the sample distribution of the three-year data set constructed by the method of the application is shown in Table 1. The classification accuracy of the method of the application and the comparative experiment method on the three-year data set is shown in Table 2. The classification accuracy of the method of the application and the generalization experiment method on the three-year space generalization data set is shown in Table 3. The classification accuracy of the method of the application on the three-year space generalization data set is shown in Table 4. The classification results of the method of the application and the comparative experiment method are shown in Figure 7

[0112] From the classification accuracy table 1, the method proposed in the application shows excellent performance in the crop classification task. On the three-year data set, the overall classification accuracy of the method of the application is greatly improved compared with other deep learning models and traditional algorithms, and the average accuracy is improved by 0.48% compared with the suboptimal model.

[0113] From the space generalization classification accuracy table 2, the method proposed in the application shows excellent performance in the crop classification task. On the three-year data set, the overall space generalization classification accuracy of the method of the application is greatly improved compared with other deep learning models and traditional algorithms, and the average accuracy is improved by 3.5% compared with the suboptimal model.

[0114] From the time generalization classification accuracy table 3, the method proposed in the application shows certain robustness when applied across years.

[0115] Table 1: Sample distribution of three-year data set

[0116]

[0117]

[0118] Table 2: Classification accuracy of each method on three-year data set

[0119]

[0120]

[0121] Table 3: Space generalization classification accuracy of each method on three-year data set

[0122]

[0123]

[0124] Table 4: Time generalization classification accuracy of the method of the application

[0125]

[0126] ​The above merely describes preferred specific embodiments of the present application, but the protection scope of the present application is not limited thereto, and any person skilled in the art, according to the technical solution and inventive concept of the present application, makes equivalent replacement or change within the technical range disclosed by the present application, which should be covered within the protection scope of the present application.

Claims

1. A method for crop classification in multi-temporal remote sensing images based on spatiotemporal attention U-type network, characterized by: include: Obtain a multi-temporal remote sensing image dataset of crops, including a time series dataset. Read remote sensing images of plots dominated by crops, annotate them with dataset labels and generalization set labels, and divide the time series dataset, dataset labels, and generalization set labels into a training set, a validation set, a test set, and a spatial generalization set. Construct a spatiotemporal attention U-shaped network model, including an encoder path and a decoder path; Construct a convolutional block attention module and a lightweight temporal attention encoder module in the encoder path, and integrate a dynamic upsampling module and an adaptive feature fusion module in the decoder path; Use the training set and validation set to train the spatiotemporal attention U-shaped network model; Use the cross entropy loss function to calculate the loss value between the category map output by the spatiotemporal attention U-type network model and the real map, and backpropagate to update the parameters of the spatiotemporal attention U-type network model; Use the trained spatiotemporal attention U-shaped network model to classify crops on the test set and output the crop classification results; Use the spatial generalization set, dataset labels, and generalization set labels to test the generalization ability of the trained model to evaluate the spatial generalization ability of the spatiotemporal attention U-shaped network model; The trained spatiotemporal attention U-shaped network model is used to test its temporal generalization ability on datasets from other years to evaluate the temporal generalization ability of the network model.

2. The method for crop classification from multi-temporal remote sensing images based on a spatiotemporal attention U-network according to claim 1, characterized in that: Constructing the convolutional block attention module includes: Use global average pooling and maximum pooling to extract channel statistics, and generate the attention weight M of the channel dimension through a multi-layer perceptron and a sigmoid function. c (X), the formula is as follows: M c (X)=σ(MLP1(AvgPool(X))+MLP2(MaxPool(X))) Where X is the input time series feature map, σ is the sigmoid function, MLP1 and MLP2 are different multi-layer perceptrons, AvgPool and MaxPool are average pooling operations and maximum pooling operations respectively; Use the channel-dimensional attention weight M c (X) performs element-by-element multiplication with the input temporal feature map, performs global pooling and maximum pooling operations, and obtains the importance map M of the spatial position based on the sigmoid function. s (X): M s (X)=σ(Conv 7×7 ([AvgPool(X⊙M c (X));MaxPool(X⊙M c (X))])) Among them, Conv 7×7 is a convolutional layer with a 7×7 convolution kernel, and ⊙ is element-wise multiplication; Use the importance map of spatial location M s (X) and the attention weight M of the channel dimension c After element-by-element multiplication of (X) and the input time series feature map, the final output feature map X is obtained. CBA , the process is expressed by the following formula: X CBA =X⊙M c (X)⊙M s (X⊙M c (X))。 3. The method for crop classification from multi-temporal remote sensing images based on a spatiotemporal attention U-network according to claim 1, characterized in that: When building a lightweight temporal attention encoder module: The time series feature map X CBA After point-by-point convolution and group normalization, the temporal feature map X is obtained d , add sine-cosine position encoding to get the temporal position embedding feature map X pos , the formula is as follows: X pos =X d +PE(pos,i) Among them, PE(pos,i) is the position encoding function, pos is the time position index, i is the channel index, d m is the dimension of the feature map; the time dimension is weighted by the multi-head self-attention mechanism to obtain the adaptive learning attention weight Attention, and the attention weight is combined with the time series feature map X pos Perform dot product operation to obtain feature map X attn , the formula is as follows: X attn =Attention·X pos Among them, n is the number of attention heads, Q (n) is a fixed learnable vector, K (n) is the time series feature graph X pos After the linear layer transformation, d k is the key-value dimension; The output feature map X is obtained by multi-layer perceptron processing LTAE , the formula is as follows: X LTAE =MLP2·CLOCK(MLP1·X attn ) Among them, RELU is a nonlinear activation function; Attention weights will also be applied to jump connection processing at different scales, first upsampling them to the time series feature map X of the corresponding scale mul , perform weighted summation to obtain the skip connection feature map X skip , the formula is as follows: Where T is the time step.

4. The method for crop classification from multi-temporal remote sensing images based on a spatiotemporal attention U-network according to claim 1, characterized in that: When building a dynamic upsampling module: The skip connection feature map X skip The dot product of the feature map after the offset convolution layer and the feature map after the range adjustment convolution layer and the nonlinear processing of the 0.5 times Sigmoid function is performed, and the original grid position α is added to get the offset Δ. The formula is as follows: Δ=PS(Conv offset (X skip )·σ(Conv scope (X skip ))·0.5)+a Among them, PS is pixel reorganization, Conv offset is the offset convolution layer, Conv scope Adjust the convolutional layer for range; The skip connection feature map X skip The offset Δ is calculated using a bilinear interpolation grid sampling function to obtain the feature map X grid , the formula is as follows: X grid =Grid_Sample(X skip ,Δ) Among them, Grid_Sample is the grid sampling function.

5. The method for crop classification from multi-temporal remote sensing images based on a spatiotemporal attention U-network according to claim 1, characterized in that: When building the adaptive feature fusion module: upsample the feature map X grid and skip connection feature map X skip The feature map X′ is obtained by point-by-point convolution dimensionality reduction. grid and Y′ skip , calculate the element-wise product of the two feature maps, and generate a similarity map through convolution operation and Sigmoid function, and then perform adaptive fusion based on the similarity map to obtain the feature map X AFF , the formula is as follows: X AFF =(1-σ(Conv(X′ grid ⊙X′ skip )))⊙X grid +σ(Conv(X′ grid ⊙X′ skip ))⊙X skip 。 6. The method for crop classification from multi-temporal remote sensing images based on a spatiotemporal attention U-network according to claim 1, characterized in that: The multi-temporal remote sensing image data is the Sentinel-2 satellite multispectral image data, including 10 bands: visible light band, near infrared band, red edge band, narrowband near infrared band and shortwave infrared band.

7. The method for crop classification from multi-temporal remote sensing images based on a spatiotemporal attention U-type network according to claim 1, characterized in that: The convolution block attention module includes a channel attention submodule and a spatial attention submodule. The channel attention submodule extracts channel statistical information through global average pooling and maximum pooling, and uses a multi-layer perceptron with channel dimensionality reduction to generate attention weights of the channel dimension. The spatial attention submodule uses a convolution kernel to generate an importance map of the spatial position. Channel and spatial attention are applied sequentially and fused with the original features through residual connections.

8. The method for crop classification from multi-temporal remote sensing images based on a spatiotemporal attention U-network according to claim 1, characterized in that: The lightweight temporal attention encoder module performs convolution processing on the input temporal feature map to obtain a dimension d m The feature map is constructed, the feature tensor is rearranged and group normalized, sine-cosine position encoding is added to enhance the model's perception of time position, the time dimension is weighted through a multi-head self-attention mechanism, and the time series features are aggregated based on the attention weights learned adaptively.

9. The method for crop classification from multi-temporal remote sensing images based on a spatiotemporal attention U-network according to claim 1, characterized in that: The dynamic upsampling module calculates the dynamic offset through the offset convolution layer, provides the initial sampling position for each output pixel through the range adjustment convolution layer, performs dynamic upsampling under the guidance of the offset to obtain the output feature map, and integrates the post-processing convolution layer and the residual connection mechanism to enhance the feature expression capability.

10. The method for crop classification based on multi-temporal remote sensing images using a spatiotemporal attention U-network according to claim 1, characterized in that: The adaptive feature fusion module performs convolution dimensionality reduction on the upsampled feature map and the skip connection feature map, calculates the element-wise product of the two feature maps, generates a similarity map through convolution operation and Sigmoid function, performs adaptive fusion based on the similarity map, and further improves the feature extraction capability through convolution layers and residual connections.

Citation Information

Patent Citations

  • Crop space-time generalization classification method and system based on deep learning

    CN113469122A

  • Remote sensing image vegetation extraction method and device

    CN118196629A

  • Sequential SAR crop classification method based on lightweight linear attention

    CN119131470A

  • Forest region cross-season change detection method and system and computer program

    CN119180723A

  • Detection method using fusion network based on attention mechanism, and terminal device

    US11222217B1

Cited By

  • Crop type extraction method, device, equipment and medium

    CN122244694A