Deep learning rice planting area extraction method fusing remote sensing image time-space spectrum information
Through the CTH-Net model, the spectral, temporal and spatial characteristics of remote sensing images are integrated, and the problem of low rice extraction accuracy under complex landscapes is solved, achieving high-precision and stable extraction in rice planting areas.
Patent Information
- Application Number
- CN202510349814.4
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-03-24
- Publication Date
- 2025-07-08
- Estimated Expiration
- 2045-03-24
AI Technical Summary
The prior art has low rice extraction accuracy under complex landscapes, making it difficult to effectively utilize the spatiotemporal and spatial spectral information of remote sensing images, especially in diversified planting modes, which lack classification accuracy and robustness.
A CTH-Net hybrid deep learning model was constructed, combined with CNN and Transformer branches, and introduced the space-time spectral fusion TSSF module, residual module and Dropout regularization, and fused spectral, time and space characteristics through a multi-level attention mechanism to improve rice extraction accuracy.
The accuracy and stability of rice extraction have been significantly improved, especially in complex landscapes, the classification accuracy of single-season rice, double-season rice and abandoned land has reached 99.69%, and good regional adaptability has been demonstrated.
Smart Images

Figure CN120279415A_ABST
Abstract
Description
Technical Field
[0001] The present invention belongs to the technical fields of remote sensing technology, image processing, and big data technology, and particularly relates to a deep learning rice extraction method that fuses spatio-temporal spectral information of remote sensing images. Background Art
[0002] As one of the major food crops globally, accurate monitoring of its planting areas is of great significance for ensuring food security, optimizing water resource management, and addressing climate change. However, traditional ground survey methods are inefficient and have limited coverage, making it difficult to meet the dynamic monitoring needs of large-scale rice planting areas. With the development of remote sensing technology, rice distribution mapping based on satellite remote sensing images has become an efficient and reliable method. Sentinel-2 and Landsat series satellite images, as open-access data resources, provide rich spatio-temporal information, which can effectively support rice extraction in complex landscapes and provide data support for precision agriculture and agricultural policy-making.
[0003] Remote sensing crop mapping methods can be classified into two categories according to the time dimension: single-temporal and multi-temporal methods. Single-temporal methods rely on spectral features at a specific time point, with the advantages of simple calculation and fast processing speed. However, they are difficult to capture the dynamic changes in rice growth, resulting in unstable classification results and low accuracy. Although multi-temporal methods can better utilize spectral responses and growth cycle characteristics by analyzing images throughout the growth cycle, there are still deficiencies in fusing time, space, and spectral information. Especially in complex landscapes and diverse planting patterns, the classification accuracy and robustness need to be improved. In recent years, deep learning technology has been widely applied in remote sensing image classification. Convolutional neural networks (CNNs) are one of the most commonly used architectures. In crop mapping, CNNs can process high-dimensional remote sensing data and extract deep features related to crop types, geographical locations, and spectral information, thereby improving classification accuracy. However, CNNs have limitations in capturing global time dependencies in time series data. In contrast, recurrent neural networks (RNNs) and long short-term memory networks (LSTMs) perform well in processing time series data. LSTM overcomes the vanishing gradient problem of standard RNNs by introducing a gating mechanism, making it more stable in modeling long-term dependencies. In recent years, the Transformer model has made breakthrough progress in the field of natural language processing and has gradually been applied to remote sensing image analysis. Compared with RNNs and LSTMs, the self-attention mechanism of the Transformer can more effectively capture long-term time dependencies. In addition, the labor-intensive task of generating training samples poses challenges to large-scale crop mapping. Although some researchers have solved this problem by developing automatic sample generation methods, these methods are not directly applicable to the extraction of more refined single-season and double-season rice in complex landscapes. Therefore, the existing technologies have limitations and need to be improved. Summary of the Invention
[0004] The objective of the present invention is to provide a deep learning method for extracting paddy rice planting areas by integrating the spatio-temporal-spectral information of remote sensing images, so as to improve the accuracy of paddy rice extraction under complex landscapes.
[0005] The technical solution of the present invention is as follows:
[0006] A deep learning method for extracting paddy rice planting areas by integrating the spatio-temporal-spectral information of remote sensing images, comprising the following steps:
[0007] Step 1: Construct a time series of spectral features and texture features of Sentinel-2 images;
[0008] Step 2: Establish an image set integrating spectral-time-space features;
[0009] Step 3: Construct a CTH-Net hybrid deep learning model; the CTH-Net hybrid deep learning model includes: a CNN branch, a Transformer branch, a TSSF module, a residual module, Dropout regularization, and a fully connected layer;
[0010] In CTH-Net, image samples are input into the CNN branch and the Transformer branch; the CNN branch extracts combinations of local spatial and spectral features at each time point; the Transformer branch is used to extract global spatio-temporal-spectral features from the input data;
[0011] CTH-Net introduces a spatio-temporal-spectral fusion TSSF module with a multi-level attention mechanism for integrating the local features extracted by the CNN branch and the global features extracted by the Transformer branch; the TSSF module embeds a spatial attention SA mechanism, the SA mechanism calculates the average pooling feature and the maximum pooling feature, and learns the weights of important regions through a 3×3 convolutional layer, and finally uses the Sigmoid activation function for normalization;
[0012] The residual module effectively alleviates the common gradient vanishing problem in deep networks by promoting the learning of complex feature representations;
[0013] The Dropout regularization module further enhances the generalization ability of the model and reduces the risk of overfitting;
[0014] The features extracted are mapped to the category space through the fully connected layer, and the output is converted into a probability distribution using the softmax function to effectively handle the diversity of land cover types.
[0015] In the method described above, the CNN branch consists of three 3×3 convolutional layers, and each convolutional layer is equipped with Batch Normalization and ReLU activation function to ensure the effectiveness and stability of feature extraction.
[0016] In the method described above, the Transformer branch first expands the channels of the input data through a 1×1 convolutional layer; subsequently, positional encoding is added to the input data so that the Transformer can capture the potential relationships between different time points.
[0017] In the method described above, the Transformer branch consists of three layers of Transformer encoders, and each encoder contains self-attention, multi-layer perceptron (MLP), layer normalization, and residual connection; the Transformer encoder uses the self-attention mechanism and MLP to capture the correlations between different time points, and then extracts the global features of rice.
[0018] In the method described above, the SA mechanism integrates different local information, enhances the ability of detail extraction, and thus reduces the misclassification of land cover caused by fallow land and fragmented rice distribution. Its mathematical expression is as follows:
[0019]
[0020] where represents the global max-pooling feature, represents the global average-pooling feature, σ is the Sigmoid function, and C 3×3 represents the convolution operation with a convolution kernel size of 3×3.
[0021] In the method described above, a channel attention CA module is introduced in the Transformer branch to adaptively adjust the weights of each channel and enhance the feature representation; adaptive max-pooling and adaptive average-pooling are applied to the input features, and then two-stage channel feature transformation is performed through one-dimensional convolution and ReLU activation, which simplifies the model while retaining the cross-channel feature extraction ability; the CA module learns the importance of each channel and selectively enhances or suppresses specific features.
[0022] In the method described above, the CA module enables the model to focus on the most relevant channels, thereby improving the classification accuracy. Its mathematical expression is as follows:
[0023]
[0024] where represents the adaptive max-pooling feature, represents the adaptive average-pooling feature, and C 1×1Denotes a convolution operation with a convolution kernel size of 1×1, and σ is the Sigmoid function.
[0025] The method introduced a multi-head attention mechanism (MHSA) to fuse the spatial, temporal, and spectral features extracted by the CNN and Transformer branches; MHSA enhances the feature expression ability of the model by computing multiple independent attention heads in parallel. Each attention head independently focuses on different aspects of the input data, thereby capturing richer information.
[0026] In the method, in MHSA, the input features are first mapped into query (Query, Q), key (Key, K), and value (Value, V) matrices respectively through learnable parameter matrices. Subsequently, for each attention head, the following formula is used to calculate its output:
[0027]
[0028] MultiHead(Q, K, V) = Concat(head1, head1,..., head h )W 0
[0029] where d k is the feature dimension of each attention head, which is used to scale the dot product result to prevent gradient instability; the attention scores are normalized by the softmax function to highlight key features and suppress noise; then, the outputs of each attention head are concatenated and linearly transformed through the learnable parameter matrix W 0 to generate the final feature representation. The output of the multi-head attention mechanism is further passed through a residual module to strengthen the feature flow and improve the stability of the model.
[0030] In the method, the residual module consists of two 3×3 convolutional layers, and each convolutional layer is equipped with BatchNormalization behind it to stabilize the training process; in addition, a shortcut connection of 1×1 convolution directly connects the input and output of the module, enabling the model to learn the residual mapping rather than the complete transformation.
[0031] The beneficial effects of the present invention are as follows: By optimizing the spectral-time-space multi-dimensional data fusion method, the efficient integration of spectral, time, and space features of remote sensing images is achieved. The CTH-Net introduces a spatio-temporal-spectral fusion (TSSF) module and a residual module, significantly enhancing the model's ability to extract rice in complex landscapes. The results show that CTH-Net reaches 99.69% in terms of accuracy, precision, recall, and F1 score for rice extraction, and performs stably (>96%) in different categories such as single-season rice, double-season rice, and fallow land. Compared with other models such as CNN, Transformer, LSTM, and SVM, CTH-Net performs excellently in dealing with fragmented rice distributions and mixed land types, improving the accuracy of rice extraction. In addition, CTH-Net exhibits good regional transferability and can adapt to different agricultural environments. BRIEF DESCRIPTION OF THE DRAWINGS
[0032] Figure 1 Flowchart of the deep learning method for extracting rice planting areas by fusing spatio-temporal-spectral information of remote sensing images;
[0033] Figure 2 Structure diagram of the CTH-Net model;
[0034] Figure 3 . Classification results of CTH-Net: (a) Classification map of six land cover types; (b) Rice extraction map;
[0035] Figure 4 . Classification results of CTH-Net in the model transfer area: (a) Classification map of six land cover types; (b) Rice extraction map;
[0036] Figure 5 . Evaluation of the regional transfer ability of CTH-Net; DETAILED DESCRIPTION OF THE INVENTION
[0037] The present invention will be described in detail below in conjunction with specific embodiments.
[0038] Step 1: Construct a time series of spectral features and texture features of Sentinel-2 images
[0039] Obtain Sentinel-2 Level-2A (L2A) satellite remote sensing images from the Copernicus Open Access Hub of the European Space Agency (https: / / scihub.copernicus.eu / ). The Multispectral Imager (MSI) carried by the Sentinel-2 satellite consists of 13 bands, with the spectral range covering from visible light to shortwave infrared, and the spatial resolutions are 10m, 20m, and 60m respectively. Resample the band data with spatial resolutions of 20m and 60m to 10m to ensure the consistency of the spatial resolutions of all band data. Use the quality_scene_classification quality flag to mask clouds and shadows, and obtain 19 cloud-free images from March to October 2020.
[0040] Select 10 spectral bands (B2, B3, B4, B5, B6, B7, B8, B8A, B11, and B12) to capture rice growth characteristics. To better reflect vegetation vitality, water content, and soil conditions and enhance the rice extraction ability, use the spectral bands to calculate 5 spectral indices: The Normalized Difference Vegetation Index (NDVI) and the Enhanced Vegetation Index (EVI) are effective in detecting vegetation changes and are widely used in crop mapping; the Normalized Difference Water Index (NDWI) is used to identify changes in soil and vegetation water content and is crucial for identifying paddy fields that require a large amount of water during the growth stage; the Modified Soil Adjusted Vegetation Index (MSAVI) and the Brightness Index (BI2) are used to consider the soil background effect and the brightness differences between different land cover types, which are crucial for accurately extracting rice in cases of complex soil conditions or low vegetation coverage. These 10 spectral bands and 5 spectral indices together form 15 spectral features.
[0041] Image texture features can describe the local variations and patterns of pixel intensities, effectively distinguish different land cover types, and identify structural details in the image. The Gray-Level Co-Occurrence Matrix (GLCM) can capture the gray-level distribution patterns in different directions, distances, and intensities. By calculating the frequencies of pixel value combinations through GLCM, 10 texture features of the image are described, including Homogeneity, Angular Second Moment (ASM), Contrast, Energy, Maximum Probability, Dissimilarity, Variance, Entropy, Mean, and Correlation. To optimize the extraction effect, the first principal component is extracted using Principal Component Analysis (PCA), the image is quantized into 32 gray levels, the pixel displacement is set to 1, and the directions are 0°, 45°, 90°, and 135°. The texture feature values in these directions are averaged to obtain a comprehensive set of texture features for analysis.
[0042] Considering the dimensional differences between spectral features and texture features, a data mapping strategy is used to standardize these feature values into a unified feature space. This not only ensures the comparability of features but also enhances the consistency of data processing, thereby improving the reliability of analysis. The data mapping strategy is based on the confidence interval and percentile truncation stretching technique, which determines a confidence interval covering the data distribution, captures the main distribution, and excludes extreme values at the same time. By selecting the interval from 1% to 99%, the gray-level values corresponding to the 1st percentile and the 99th percentile are calculated. These values serve as reference points for linear stretching to map the original gray-level values into a unified interval. This process enhances the pixel contrast and allows different feature values to be represented on the same scale.
[0043] By arranging 25 feature images (including 15 spectral features and 10 texture features) at each Sentinel-2 image time point, a time series of Sentinel-2 image spectral features and texture features is constructed.
[0044] Step 2: Establish an image set integrating spectral-time-spatial features
[0045] Use an active learning strategy to select high-quality single-season and double-season rice samples. First, generate NDVI time series curves for each rice sample and distinguish between single-season and double-season rice through manual interpretation. The NDVI curve of double-season rice shows two distinct peaks, corresponding to the growth seasons of early rice (March to July) and late rice (July to October) respectively; the single-season rice shows only one peak, and its growth period is from late May to October. These manually labeled samples are used to train a support vector machine (SVM) classifier to predict the rice types of a wider range of samples. Samples with a prediction probability exceeding 85% are classified as high-confidence samples and directly incorporated into the sample dataset; samples with lower confidence are subject to further manual inspection to ensure accuracy. The combination of manually labeled samples and high-confidence samples is used to iteratively train the SVM model to obtain a rice sample dataset suitable for complex landscapes.
[0046] Based on the high-quality sample dataset, construct an image set that integrates spectral, temporal, and spatial features to comprehensively capture the synergistic effects between these features. First, starting from the starting time point of the time series, extract the values of each pixel in all bands and form these band values into a one-dimensional feature vector. As the time series progresses, repeat this process at each subsequent time point to ensure complete recording of the temporal information of each pixel. Finally, integrate the one-dimensional feature vectors obtained from 19 time points into a two-dimensional matrix and convert it into a 19×25 grayscale image, where each pixel value represents the spectral or spatial features at a specific time point. For any remotely sensed image of size W×L, this method is applied to each pixel to generate W×L grayscale images, ultimately constituting a multi-dimensional feature dataset. Through this method, a detailed spectral and spatial response map is established for each pixel at different time points, comprehensively integrating temporal, spatial, and spectral features.
[0047] Step 3: Construct a CTH-Net hybrid deep learning model
[0048] CTH-Net is a hybrid architecture that combines the convolutional neural network (CNN) and the Transformer mechanism, aiming to enhance the ability to extract rice spectral, spatial, and temporal features for accurate rice extraction. Rice exhibits not only seasonal spectral dynamic changes but also unique spatial features. Therefore, efficiently characterizing these features is crucial for accurate rice identification. To enhance the ability to extract high-level abstract features of the target rice, CTH-Net utilizes CNN and Transformer to mine deep representative features. In CTH-Net, image samples are input into the CNN branch to extract combinations of local spatial and spectral features at each time point. The CNN branch consists of three 3×3 convolutional layers, each followed by batch normalization and the ReLU activation function to ensure the effectiveness and stability of feature extraction. Meanwhile, the Transformer branch is used to extract global spatio-temporal-spectral features from the input data. This branch first expands the channels of the input data through a 1×1 convolutional layer to improve the feature expression ability. Subsequently, to preserve the time series information, positional encoding is added to the input data, enabling the Transformer to capture the potential relationships between different time points. The Transformer branch consists of three layers of Transformer encoders, each encoder containing self-attention, a multi-layer perceptron (MLP), layer normalization, and residual connection. The Transformer encoder uses the self-attention mechanism and MLP to capture the correlations between different time points and then extracts the global features of rice. Compared with traditional convolutional networks, the Transformer has significant advantages in processing long time series and can more effectively identify the temporal variation features between single-season and double-season rice.
[0049] To effectively integrate the local features extracted by the CNN branch and the global features extracted by the Transformer branch, CTH-Net introduces a spatio-temporal-spectral fusion (TSSF) module with a multi-level attention mechanism. This module embeds the spatial attention (SA) mechanism to emphasize the responses of key regions in the feature maps extracted by convolutional operations. This mechanism calculates the average pooling features and the maximum pooling features, learns the weights of important regions through a 3×3 convolutional layer, and finally normalizes them using the Sigmoid activation function. The SA mechanism integrates different local information and enhances the ability to extract details, thus reducing the land cover misclassification caused by fallow land and fragmented rice distribution. Its mathematical expression is as follows:
[0050]
[0051] Among them, represents the global maximum pooling feature, represents the global average pooling feature, and σ is the Sigmoid function, C 3×3 represents a convolutional operation with a convolution kernel size of 3×3. In the Transformer branch, a channel attention (CA) mechanism is introduced to adaptively adjust the weights of each channel and enhance the feature representation. Adaptive max pooling and adaptive average pooling are applied to the input features, followed by two-stage channel feature transformation through one-dimensional convolution and ReLU activation, which simplifies the model while retaining the cross-channel feature extraction ability. Different from traditional convolution, the CA module learns the importance of each channel and selectively enhances or suppresses specific features. This mechanism enables the model to focus on the most relevant channels, thereby improving the classification accuracy. Its mathematical expression is as follows:
[0052]
[0053] Among them, represents the adaptive max pooling feature, represents the adaptive average pooling feature, C 1×1 represents a convolutional operation with a convolution kernel size of 1×1, and σ is the Sigmoid function. Different from traditional convolution, the CA module learns the importance of each channel and selectively enhances or suppresses specific features. This mechanism enables the model to focus on the most relevant channels, thereby improving the classification accuracy.
[0054] In addition, a multi-head attention mechanism (MHSA) is introduced to fuse the spatial, temporal, and spectral features extracted by the CNN and Transformer branches. Different from the traditional attention mechanism (using a single weight matrix), MHSA enhances the feature expression ability of the model by calculating multiple independent attention heads in parallel. Each attention head independently focuses on different aspects of the input data, thereby capturing richer information. In MHSA, the input features are first mapped to query (Q), key (K), and value (V) matrices through learnable parameter matrices respectively. Subsequently, for each attention head, its output is calculated using the following formula:
[0055]
[0056] MultiHead(Q,K,V)=Concat(head1,head1,...,head h )W 0
[0057] Among them, d k is the feature dimension of each attention head, Used to scale the dot product result to prevent gradient instability. The attention scores are normalized by the softmax function to highlight key features and suppress noise. Then, the outputs of each attention head are concatenated and linearly transformed through a learnable parameter matrix W 0 to generate the final feature representation. The output of the multi-head attention mechanism is further passed through a residual module to strengthen the feature flow and improve the stability of the model.
[0058] The residual module effectively alleviates the vanishing gradient problem commonly found in deep networks by facilitating the learning of complex feature representations. The module consists of two 3×3 convolutional layers, each followed by Batch Normalization to stabilize the training process. Additionally, a 1×1 convolutional shortcut connection directly links the input and output of the module, enabling the model to learn the residual mapping rather than the complete transformation. This design not only helps to preserve temporal, spatial, and spectral features during feature fusion but also allows the model to capture complex relationships in the input data through effective gradient flow, thereby improving the accuracy of rice extraction under complex surface conditions. To further enhance the generalization ability of the model and reduce the risk of overfitting, Dropout regularization is introduced. Finally, the extracted features are mapped to the class space through a fully connected layer, and the output is converted into a probability distribution using the softmax function to effectively handle the diversity of land cover types.
[0059] Step 4: Evaluate the accuracy of rice extraction
[0060] To evaluate the accuracy of the rice extraction results, a confusion matrix is introduced to quantify the correct and incorrect predictions of the model. As shown in Table 1, based on the confusion matrix, several accuracy evaluation metrics are further calculated, including accuracy, precision, recall, and F1 score. The value of accuracy ranges from 0 to 1 and is used to measure the overall performance of the model when classifying six land cover types (single-season rice, double-season rice, water body, forest, fallow land, and others). A high value indicates good classification performance of the model. Precision also ranges from 0 to 1 and reflects the ability of the model to accurately identify the rice planting area while minimizing the misclassification of other land cover types as rice. High precision values in the extraction of single-season rice and double-season rice indicate a low error rate of the model. The value of recall also ranges from 0 to 1 and is used to measure the comprehensiveness of the model in identifying single-season rice and double-season rice. A high recall value means that the model can better cover the actual rice planting area and reduce the occurrence of omission cases. The F1 score is the balanced value of precision and recall, ranging from 0 to 1. A score close to 1 indicates that the model has achieved a good balance between accurate identification and comprehensive coverage, showing strong overall performance.
[0061] Table 1. Performance Evaluation of Different Models
[0062]
[0063] As Figure 3 shown, the study area is mainly forested, and water bodies are mainly distributed in the main stream of the Xiangjiang River and surrounding lakes and irrigation canals. Paddy fields are concentrated in plains and valleys, where single-season rice is widely distributed, and double-season rice is concentrated in the middle and near water sources. Buildings are mainly concentrated near the central river valley. Abandoned farmland is scattered in the farmland due to human and natural factors. These results are consistent with the field survey.
[0064] As Figure 4 shown, in the model transfer area, forest is the main land use type, farmland is widely distributed in plains and valleys, and water bodies are mainly the Lushui River and surrounding lakes and irrigation canals. The proportion of single-season rice is higher than that of double-season rice and is widely distributed, while double-season rice is mainly concentrated in the middle and south and is mostly near water sources, conforming to the local irrigation pattern.
[0065] As Figure 5 shown, the evaluation results of the model's transfer ability on six land use types show that it performs excellently in all four indicators. Each indicator of single-season rice is about 0.97, and each indicator of double-season rice is close to 0.96. The overall classification accuracy of rice reaches 0.9960, indicating that the model can effectively identify rice planting areas. The precision and recall of abandoned farmland are 0.9814 and 0.9875 respectively. Each indicator of land use types such as forest and water body exceeds 0.99, verifying the good regional adaptability of the model.
[0066] It should be understood that those of ordinary skill in the art can make improvements or transformations according to the above description, and all such improvements and transformations shall fall within the protection scope of the appended claims of the present invention.
Claims
1. A deep learning method for extracting rice planting areas by integrating spatio-spectral information of remote sensing images, characterized in that, It includes the following steps: Step 1: Construct the time series of Sentinel-2 image spectral features and texture features; Step 2: Establish an image set that fuses spectral-time-space features; Step 3: Construct a CTH-Net hybrid deep learning model; the CTH-Net hybrid deep learning model includes: a CNN branch, a Transformer branch, a TSSF module, a residual module, Dropout regularization, and a fully connected layer; In CTH-Net, image samples are input into the CNN branch and the Transformer branch; the CNN branch extracts the local spatial and spectral feature combinations at each time point; the Transformer branch is used to extract global spatio-temporal-spectral features from the input data; CTH-Net introduces a spatio-temporal-spectral fusion TSSF module with a multi-level attention mechanism to integrate the local features extracted by the CNN branch and the global features extracted by the Transformer branch; the TSSF module embeds a spatial attention SA mechanism, and the SA mechanism calculates the average pooling feature and the max pooling feature, and learns the weights of important regions through a 3×3 convolutional layer, and finally uses the Sigmoid activation function for normalization; The residual module effectively alleviates the common gradient vanishing problem in deep networks by promoting the learning of complex feature representations; Dropout regularization further enhances the generalization ability of the model and reduces the risk of overfitting; The features extracted are mapped to the category space through the fully connected layer, and the output is converted into a probability distribution using the softmax function to effectively handle the diversity of land cover types.
2. The method according to claim 1, characterized in that, The CNN branch consists of three 3×3 convolutional layers, and each convolutional layer is equipped with batch normalization and ReLU activation functions to ensure the effectiveness and stability of feature extraction.
3. The method according to claim 1, wherein The Transformer branch first expands the channels of the input data through a 1×1 convolutional layer; subsequently, positional encoding is added to the input data to enable the Transformer to capture the potential relationships between different time points.
4. The method according to claim 1, characterized in that, The Transformer branch consists of three layers of Transformer encoders, and each encoder contains self-attention, a multi-layer perceptron, layer normalization, and residual connections; the Transformer encoder uses the self-attention mechanism and the MLP to capture the correlations between different time points, and then extracts the global features of rice.
5. The method according to claim 1, wherein The SA mechanism integrates different local information, enhances the detail extraction ability, and thus reduces the land cover misclassification phenomenon caused by fallow land and fragmented rice distribution. Its mathematical expression is as follows: Among them, represents the global maximum pooling feature, represents the global average pooling feature, σ is the Sigmoid function, C 3×3 represents a convolution operation with a convolution kernel size of 3×3.
6. The method according to claim 4, characterized in that A channel attention CA module is introduced in the Transformer branch to adaptively adjust the weights of each channel and enhance the feature representation; adaptive max pooling and adaptive average pooling are applied to the input features, and then two-stage channel feature transformation is performed through one-dimensional convolution and ReLU activation, simplifying the model while retaining the cross-channel feature extraction ability; the CA module learns the importance of each channel and selectively enhances or suppresses specific features.
7. The method according to claim 6, wherein The CA module enables the model to focus on the most relevant channels, thereby improving the classification accuracy. Its mathematical expression is as follows: Among them, represents the adaptive maximum pooling feature, represents the adaptive average pooling feature, C 1×1 represents a convolution operation with a convolution kernel size of 1×1, and σ is the Sigmoid function.
8. The method according to claim 6, characterized in that The multi-head attention mechanism (MHSA) is introduced to fuse the spatial, temporal, and spectral features extracted by the CNN and Transformer branches; MHSA enhances the feature expression ability of the model by computing multiple independent attention heads in parallel. Each attention head independently focuses on different aspects of the input data, thereby capturing richer information.
9. The method according to claim 7, wherein In MHSA, the input features are first mapped to query (Q), key (K), and value (V) matrices respectively through learnable parameter matrices. Subsequently, for each attention head, its output is calculated using the following formula: MultiHead(Q,K,V)=Concat(head1,head1,...,head h )W 0 where d k is the feature dimension of each attention head, which is used to scale the dot product result to prevent gradient instability; the attention scores are normalized by the softmax function to highlight key features and suppress noise; then, the outputs of each attention head are concatenated and linearly transformed by the learnable parameter matrix W 0 to generate the final feature representation. The output of the multi-head attention mechanism is further passed through a residual module to strengthen the feature flow and improve the stability of the model.
10. The method according to claim 1, wherein The residual module consists of two 3×3 convolutional layers, and each convolutional layer is equipped with Batch Normalization behind it to stabilize the training process; In addition, a shortcut connection of 1×1 convolution directly connects the input and output of the module, enabling the model to learn the residual mapping rather than the complete transformation.
Citation Information
Patent Citations
Large-scale regional rice remote sensing classification method under complex conditions
CN116109943A
Hyperspectral image classification method based on hybrid transformation
CN117788875A
Deep learning method for remote sensing parameter space-time spectrum fusion
CN118351456A
Method and device for generating satellite remote-sensing image with high spatial, temporal and spectral resolutions
US12315104B1