CNN-ViT-based time sequence remote sensing crop classification method and device
By combining technical means of CNN and ViT, a spatiotemporal encoder is constructed to capture the spatiotemporal characteristics of crops, solving the problem of insufficient spatial and temporal characteristics modeling in the existing technology, and significantly improving the accuracy and efficiency of crop classification.
Patent Information
- Application Number
- CN202510100229.0
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-01-22
- Publication Date
- 2025-05-02
- Estimated Expiration
- 2045-01-22
AI Technical Summary
The prior art has problems in crop classification with insufficient spatial and temporal feature modeling, limited classification accuracy and low computational efficiency.
Using a time series remote sensing crop classification method based on CNN and ViT, a spatiotemporal encoder is constructed to capture the long-range dependence of the time dimension and the characteristic interaction relationship of the spatial dimension by combining the spatial local feature extraction of the convolutional neural network and the spatial interaction modeling capabilities of Vision Transformer.
It significantly improves the accuracy and efficiency of crop classification, can more comprehensively reflect the growth cycle and spatial distribution characteristics of crops, and provides a novel technical solution for crop monitoring and management.
Smart Images

Figure CN119919734A_ABST
Abstract
Description
Technical Field
[0001] The present invention belongs to the technical field of remote sensing image classification and deep learning, and proposes a time series remote sensing crop classification method and device based on convolutional neural network (CNN) and Vision Transformer (ViT). Background Art
[0002] Crop classification is an important task in modern agricultural research, and is of great significance for crop planting area statistics, yield prediction and precision agricultural management. With the intensification of climate change, population growth and increasingly scarce land resources, efficient and accurate monitoring of crop planting conditions has become a key issue that needs to be addressed. Remote sensing imaging technology has been widely used in crop classification because of its ability to quickly and widely obtain surface information. In particular, the introduction of time series remote sensing data can effectively utilize the spectral change information of crops at different growth stages to improve classification accuracy.
[0003] At present, intelligent crop classification technology is mainly divided into two categories: traditional machine learning methods and deep learning methods. Traditional machine learning methods are represented by support vector machines and random forests. They manually extract spectrum, texture, shape and other features from remote sensing images and input them into classifiers to achieve crop classification. However, these methods are highly dependent on manual features and are difficult to cope with complex and diverse scenarios. In recent years, with the rapid development of deep learning, technologies such as convolutional neural networks (CNN), recurrent neural networks (RNN) and vision transformers (ViT) have been widely used in crop classification tasks. CNN, with its powerful spatial feature extraction capabilities, can efficiently model single-phase remote sensing images; RNN has outstanding performance in modeling time series data and is suitable for modeling multi-phase remote sensing images; ViT effectively models spatial long-distance dependencies in images through the self-attention mechanism, showing great potential in image classification tasks.
[0004] Although the above methods have achieved certain results in crop classification tasks, the following problems still exist: (1) Traditional machine learning methods rely too much on manual features, and their classification performance is limited by the quality of feature extraction, making it difficult to maintain high accuracy in complex scenes; (2) Although deep learning methods can automatically learn features, when CNN or RNN models are used alone, it is difficult to achieve interactive modeling of spatial and temporal features, which limits the modeling ability of time series remote sensing images; (3) Although the Transformer model has the ability to model global features, its high computational complexity makes it face significant computational bottlenecks when processing high-resolution multi-temporal remote sensing images. (4) Existing methods usually fail to fully integrate spatial and temporal information, ignoring the dynamic characteristics of crops at different spatiotemporal scales, resulting in limited classification accuracy. Summary of the invention
[0005] The present invention aims to overcome the above-mentioned shortcomings of the prior art and provide a time series remote sensing crop classification method and device based on CNN and ViT.
[0006] This paper combines the powerful spatial local features of convolutional neural networks with the spatiotemporal interaction modeling capabilities of Vision Transformer to effectively solve the problems of insufficient spatial and temporal feature modeling, limited classification accuracy, and low computational efficiency in current methods. When processing time series remote sensing image crop classification tasks, this method not only significantly improves the classification accuracy and efficiency, but also provides a novel technical solution for crop monitoring and management.
[0007] To achieve the above object, the first aspect of the present invention relates to a time series remote sensing crop classification method based on CNN and ViT, comprising the following steps:
[0008] Step 1: Dataset construction;
[0009] Step 1.1: Data acquisition: select remote sensing image data with a spatial resolution of H×W and a channel number of C at each observation time point T to form a time series image frame set X i :
[0010]
[0011] in, is the t-th frame image data of the i-th sample.
[0012] Then select and X i The corresponding pixel-level category label Y i Composition of image dataset D:
[0013] D={(X i ,Y i )|i=1,2,…,N} (2)
[0014] Among them, X i is the i-th time series remote sensing image sample, Y i is the pixel-level category label corresponding to the i-th sample, and N is the total number of samples in the image dataset.
[0015] Step 1.2: Data cropping. To reduce computational complexity and adapt to model input requirements, and to ensure that the sub-image contains rich spatial information, each input image data X i Cut into subgraphs of size H'×W' to form a subgraph set X i :
[0016]
[0017] in, represents the kth cropped sub-image in the tth frame of the i-th sample, M is the number of sub-images after each frame image is divided into blocks, and is defined as:
[0018]
[0019] Among them, H′ and W′ are the height and width of the cropped sub-image.
[0020] The pruned dataset D′ can be represented as a set of pruned subgraphs and their corresponding labels:
[0021]
[0022] in, represents the kth sub-image after cropping, Indicates the pixel-level label corresponding to the cropped sub-image.
[0023] Step 1.3: Dataset division: divide the processed dataset D′ into the following three subsets in proportion: training set D train , validation set D val , test set D test , in order to improve the independence and generalization performance of model training and avoid overfitting.
[0024] Step 2: Convolution feature extraction;
[0025] Step 2.1: Convolutional feature extraction module construction. Remote sensing images usually have significant spatial structural information (such as land boundaries and textures) and local feature distribution (such as crop shape and arrangement). Although Transformer is good at global feature modeling, its direct modeling ability for local spatial features is limited. In order to enhance the extraction of local spatial information, a convolutional feature extraction module is constructed to improve local perception capabilities. For each frame of image X in the input data X, (t) Through the lightweight convolution module f CNN Perform feature extraction and generate a convolution feature map. Specifically, the following steps are included:
[0026] First, the input image is feature extracted through the two-dimensional convolution layer Conv2D to extract local spatial structure information; then the feature map generated by the convolution operation is normalized to balance the data distribution and eliminate the magnitude difference between different features, while accelerating the model training process and stabilizing the gradient flow; finally, the ReLU activation function is introduced to make the activated features sparse, improve the feature discrimination ability, and effectively avoid the gradient disappearance problem. The final convolution feature map It can provide input support for subsequent spatiotemporal feature modeling, as shown below:
[0027]
[0028] Among them, H′×W′ is the spatial resolution after convolution, and D is the channel dimension after the convolution operation.
[0029] Step 2.2: Patch segmentation and linear projection, after convolution feature extraction is completed, the generated feature map Still maintain a high resolution and number of channels. To further process the feature map, it is divided into non-overlapping patches of size h×w, each of which contains local spatial features. The total number of patches N after segmentation can be expressed as:
[0030]
[0031] For the feature map of the tth frame The i-th Patch is represented by P (t,i) , each P (t,i) Convolutional features of local areas are included, preserving high-dimensional channel information in the area:
[0032] P (t,i) ∈R h×w×D ,i=1,2,…,N (8)
[0033] In order to convert each Patch into a Token suitable for the input format of the Transformer encoder, each P (t ,i) Do the following: First, change P (t,i) Expand to a one-dimensional vector; then, through a linear projection operation, the one-dimensional vector is mapped to P (t,i) Corresponding TokenZ (t,i) , expressed as:
[0034]
[0035] Among them, W is the weight matrix of linear projection, b is the bias vector of linear projection, and D Z It is the characteristic dimension of Token.
[0036] Finally, using the same method, all patches of time frame t are converted into Token sequence Z (t) :
[0037]
[0038] Where N is the number of split patches, D Z It is the characteristic dimension of Token.
[0039] Step 3: Encoder stage;
[0040] Step 3.1: Construction of spatiotemporal position coding. In order to retain the time information of each frame and the spatial position information of each patch in the time series remote sensing image, the generated Token sequence Z (t) Add time position coding and spatial position coding.
[0041] Temporal position coding is used to represent the temporal position information of different frames in a time series. Since the acquisition time of remote sensing images is often irregular, the traditional fixed temporal position coding method is difficult to effectively model the time interval. Therefore, the present invention introduces dynamic temporal position coding by constructing a time lookup table P T ∈R T′×D , where all observation time information is embedded in it, where T′ is the number of all observation times in the data, and D is the feature dimension of the Token. For the feature representation Z of time frame t (t) , by finding the corresponding time position code P T [t], add it to each Token to get the time-enhanced feature It is expressed as:
[0042]
[0043] Spatial position coding is used to identify the spatial position information of each patch within the image, ensuring that the spatial structural characteristics of the image are retained during the time series modeling process. The grid coordinate coding method is used to identify the spatial position information of each patch (p x ,p y ) generates two separate vectors encoding the row and column coordinates respectively, calculated as follows:
[0044]
[0045] in, is the sinusoidal component of the row coordinate position encoding, Encode the cosine component for the row coordinate position, is the sinusoidal component in the column coordinate position encoding, is the cosine component in the column coordinate position encoding.
[0046] The row coordinate encoding and column coordinate encoding are concatenated to obtain the position encoding vector P of each Patch. S [i]:
[0047]
[0048] Stack the position encoding vectors of all patches to get the complete spatial position encoding matrix P S :
[0049]
[0050] And add the spatial position coding to the Token after the temporal position coding to get the representation of the Token after the temporal and spatial position coding:
[0051]
[0052] For the T frames of the entire time series, the Token sequence after all spatiotemporal position encoding is expressed as:
[0053]
[0054] Step 3.2: Construct a spatiotemporal encoder. The spatiotemporal encoder uses a Transformer module, which consists of a multi-head self-attention mechanism (MSA) and a feed-forward network (FFN) to capture long-range dependencies in the time dimension and model feature interactions in the spatial dimension. In each layer, the Token after the input spatiotemporal position encoding is first Normalize, then use the multi-head self-attention mechanism to capture the long-range dependencies in the time dimension and the feature interactions in the modeling space dimension, update the time series and spatial features, and the calculation process is as follows:
[0055]
[0056]
[0057] After the processing of the temporal encoder and the spatial encoder, the final extracted feature is represented as Z final :
[0058]
[0059] Step 4: Decoder stage;
[0060] Step 4.1: Pixel-level segmentation head construction, by final A linear projection operation is performed to convert the spatial features of each category into pixel-level prediction results. For category k and spatial position i, the features are mapped through a multi-layer perception (MLP) to map the high-dimensional features to a low-dimensional space, and the pixel-level probability distribution p of each category in the current spatial block is obtained. k,i .
[0061] Step 4.2: Spatial reorganization and merging, after completing the linear projection of all blocks, the pixel-level probability results of each block are reassembled back to the resolution of the original input image. k,i Rearrange them into the global image space to generate a complete pixel-level prediction result map, thus achieving pixel-level classification of the entire remote sensing image.
[0062] Step 5: Model training and use;
[0063] In the model training and use phase, the training set D defined in step 1.3 is used. train , validation set D val , test set D test Perform the following tasks:
[0064] Model training: Use the training set to optimize the model and gradually update the model parameters through block feature extraction, encoder processing, and decoder pixel-level semantic segmentation operations.
[0065] Model Validation: Use the validation set to tune hyperparameters and evaluate the performance of the model on unseen data.
[0066] Model testing: Use the test set to perform final performance evaluation on the model.
[0067] Finally, through the full-process processing of the model, the category probability distribution corresponding to each pixel is generated for the input time series remote sensing image data to achieve classification prediction of crops.
[0068] The second aspect of the present invention relates to a time series remote sensing crop classification device based on CNN-ViT, comprising a memory and one or more processors, wherein the memory stores executable code, and when the one or more processors execute the executable code, they are used to implement the time series remote sensing crop classification method based on CNN-ViT of the present invention.
[0069] The third aspect of the present invention relates to a computer-readable storage medium having a program stored thereon, which, when executed by a processor, implements the CNN-ViT-based time series remote sensing crop classification method of the present invention.
[0070] The advantages of the present invention are:
[0071] 1) The present invention adopts a small-scale block method in the block feature extraction stage, combined with the CNN module to further capture local spatial details and texture information, while retaining high-resolution spatial features. By introducing the spatiotemporal encoder, the model can simultaneously focus on the interaction of local and global spatiotemporal features, thereby significantly improving the accuracy of crop classification.
[0072] 2) The present invention introduces dynamic time position coding and static space position coding to effectively deal with the modeling problem of irregular time intervals and spatial position information in time series data. Dynamic time position coding enhances the modeling ability of the model for spatiotemporal correlation information, while static space position coding retains the spatial distribution characteristics, reduces the loss of position information, and can more comprehensively reflect the growth cycle and spatial distribution characteristics of crops. BRIEF DESCRIPTION OF THE DRAWINGS
[0073] Figure 1 is a flow chart of the method of the present invention.
[0074] Figure 2 It is a network structure diagram of the present invention.
[0075] Figure 3(a) is a remote sensing image, Figure 3(b) is a true value image of crop labels, and Figure 3(c) is a prediction result image of the CNN-ViT model.
[0076] Figure 4 It is a device diagram of the present invention. DETAILED DESCRIPTION
[0077] In order to further clarify the content of the present invention, the following embodiments of the present invention are further described in detail so that those skilled in the art can understand and practice. It should be noted that the following embodiments are only used to demonstrate and explain the technical solutions of the present invention and do not constitute a limitation on the scope of the present invention. It should be understood that not all features of all actual implementation methods are the same as in this embodiment. In actual engineering projects, some implementation details may be adjusted according to specific conditions and goals.
[0078] Example 1
[0079] This embodiment proposes a time series remote sensing crop classification method based on CNN and ViT. In view of the problems that traditional deep learning models are difficult to capture spatiotemporal features at the same time when processing time series remote sensing image data, the computational complexity is high, and it is difficult to ensure classification accuracy and efficiency when facing large-scale crop classification tasks, a hybrid architecture combining convolutional neural network CNN and ViT is proposed. This method achieves accurate classification of crops in time series remote sensing images through efficient block feature extraction, spatiotemporal encoding and decoding strategies. The specific implementation mode of the present invention is illustrated in conjunction with the accompanying drawings, and the process is as follows:
[0080] This embodiment uses the remote sensing image datasets XJDS and PASTIS of some areas of Xinjiang Uygur Autonomous Region, where the XJDS dataset contains 2066 time series satellite image samples (SITS), and the PASTIS dataset contains 2433 time series satellite image samples, each sample is a 128×128 pixel remote sensing image, the image has 10 spectral bands, and the spatial resolution is 10m. The XJDS dataset uses 5 types of crop labels, and the PASTIS dataset uses 18 types of crop labels. Each sample of these two datasets has 33 to 61 time acquisition frames.
[0081] Step 1: Dataset construction;
[0082] Step 1.1: Select 2066 XJDS dataset samples and 2400 remote sensing image samples with a spatial resolution of 128×128 and 10 channels from the PASTIS dataset to form a time series image frame set X i , and then extract X i Corresponding crop category label Y i Composition of image dataset D XJDS and D PASTIS .
[0083] Step 1.2: In order to reduce the computational overhead during the experiment, each sample of the XJDS dataset and the PASTIS dataset is cropped into a small block of 24×24, while retaining all time series frames. This cropping method can reduce the computational complexity without losing the original image information, while ensuring that the model output can be reassembled back to the original image size, and finally obtain 60,500 24×24 time series image samples.
[0084] Step 1.3: Dataset division: divide the 60,500 samples into: 40,500 training sample sets D train , 10,000 validation sample sets D val , 10,000 test sample sets D test .
[0085] Step 2: Block feature extraction;
[0086] Step 2.1: Convolutional feature extraction module is constructed, which is based on the input T frame 24×24 image sample data X∈R T×24×24×10 , where each frame of image has 10 bands. First, for each frame of image X in the input data X (t) Input to 16 3×3 convolutional layers, and the output is a 24×24×16 feature map Then the feature map Input to 32 3×3 convolutional layers to obtain a 24×24×32 feature map Then the feature map is obtained through a 2×2 maximum pooling operation
[0087] Step 2.2: Linear mapping, after obtaining the feature map After that, the output feature map is segmented into patches. The size of the feature map is 12×12×32, which is divided into non-overlapping patches of size 4×4. Each patch contains the convolutional features of the local area, and the number of patches after division is 9. The dimension of each patch is 4×4×32, which contains the spatial features and high-dimensional channel information in the area. Then each patch is flattened into a one-dimensional vector with a dimension of 512, and then the flattened vector is mapped to 128 dimensions through linear projection. Through linear projection, the TokenZ corresponding to each patch is obtained. (t,i) Finally, for each frame t, concatenate the tokens corresponding to all patches in order to obtain the token sequence Z (t) = {Z (t,1) ,Z (t ,2) ,...,Z (t,9)}, as the input of the Transformer encoder.
[0088] Step 3: Encoder stage;
[0089] Step 3.1: Time position encoding construction, first in the generated Token sequence Z (t) According to the collection time of all samples in the XJDS dataset and the PASTIS dataset, a time lookup table P is constructed. T , retrieve the corresponding time position encoding vector P through the lookup table T [t]. Then add it to the Token, as follows:
[0090]
[0091] Spatial position coding is constructed using grid coordinate coding to encode the block coordinate position (p x ,p y ) generates two separate vectors, encoding row and column coordinates respectively, calculated as follows:
[0092]
[0093] in is the sinusoidal component of the row coordinate position encoding, Encode the cosine component for the row coordinate position, is the sinusoidal component in the column coordinate position encoding, is the cosine component in the column coordinate position encoding.
[0094] The row coordinate code and column coordinate code are concatenated to obtain the position code vector P of each block. S[i], stack the position encoding vectors of all blocks to obtain the complete spatial position encoding matrix P S .
[0095] For the input Token sequence Z of the encoder module (t) , combining the temporal position encoding and the spatial position encoding to form the final input feature
[0096]
[0097] Step 3.2: Construct a spatiotemporal encoder. The spatiotemporal encoder uses the Transformer module for calculation. In each layer, the Token after the input spatiotemporal position encoding is first Normalize it, then use the MSA module to calculate the multi-head attention mechanism to get the output Y TS Then, the two-layer fully connected network and activation function of MLP are used to obtain the final extracted feature representation Z final .
[0098] Step 4: Decoder stage;
[0099] Pixel-level segmentation head construction, Z final The features are mapped to crop categories through multi-layer perception MLP to obtain the crop classification probability distribution of each pixel. All category probabilities of each pixel are normalized by Softmax to calculate the classification probability of each pixel.
[0100] After completing the linear projection of all blocks, the pixel-level probability results of each block are re-stitched back to the resolution of the original input image 24×24 to obtain the complete crop classification results of the remote sensing image samples.
[0101] Step 5: Model training and use;
[0102] In the model training and use phase, use the training set D created in step 1.3 train , validation set D val , test set D test Train, validate and test the model, set the hyperparameters: Epochs to 100, Batch Size to 16, LearningRate to 0.001, and the network structure is as follows: Figure 2 As shown. By performing block feature extraction, encoder stage processing and pixel-level semantic segmentation operations on the input time series remote sensing image data, the category probability distribution corresponding to each pixel is finally obtained to achieve crop classification. The results obtained by this method are shown in Figure 3, where Figure 3(a) is a remote sensing image, Figure 3(b) is a crop label true value map, and Figure 3(c) is a CNN-ViT model prediction result map.
[0103] Example 2
[0104] Reference Figure 4 This embodiment relates to a time series remote sensing crop classification device based on CNN-ViT, including a memory and one or more processors, wherein the memory stores executable code, and when the one or more processors execute the executable code, they are used to implement the time series remote sensing crop classification method based on CNN-ViT of Example 1.
[0105] Example 3
[0106] This embodiment relates to a computer-readable storage medium having a program stored thereon. When the program is executed by a processor, the time series remote sensing crop classification method based on CNN-ViT of the present invention is implemented.
[0107] The above is only a description of an embodiment of the present invention, but the protection scope of the present invention should not be regarded as limited to the specific forms described in the embodiment. The protection scope of the present invention also extends to equivalent technical means that can be conceived by those skilled in the art based on the concept of the present invention.
Claims
1. A time series remote sensing crop classification method based on CNN-ViT, characterized in that: include: Step 1: Dataset construction: preprocessing and standardizing time series remote sensing image data to form a temporally and spatially consistent training dataset; Step 2: Convolutional feature extraction: using the local feature extraction module to obtain the spatial structure information in the image and enhance the model's ability to perceive local details; Step 3: In the encoder stage, temporal and spatial position encoding are combined to generate a spatiotemporal joint feature representation that integrates the dynamic changes between frames and the interaction of spatial features; Step 4: In the decoder stage, the joint features are decoded layer by layer, mapped into category probability distribution, and the classification results are reconstructed; Step 5: Model training and use to achieve prediction of crop planting categories and generate classification results.
2. The time series remote sensing crop classification method based on CNN-ViT according to claim 1 is characterized in that: The data set construction described in step 1 specifically includes: Step 1.1: Data acquisition: select remote sensing image data with a spatial resolution of H×W and a channel number of C at each observation time point T to form a time series image frame set X i : in, is the t-th frame image data of the i-th sample; Then select and X i The corresponding pixel-level category label Y i Composition of image dataset D: D={(X i ,Y i )∣i=1,2,…,N} (2) Among them, X i is the i-th time series remote sensing image sample, Y i is the pixel-level category label corresponding to the i-th sample, and N is the total number of samples in the image dataset; Step 1.2: Data cropping. To reduce computational complexity and adapt to model input requirements, and to ensure that the sub-image contains rich spatial information, each input image data X i Cut into subgraphs of size H'×W' to form a subgraph set X i : in, represents the kth cropped sub-image in the tth frame of the i-th sample, M is the number of sub-images after each frame image is divided into blocks, and is defined as: Among them, H′ and W′ are the height and width of the cropped sub-image; The pruned dataset D′ can be represented as a set of pruned subgraphs and their corresponding labels: in, represents the kth sub-image after cropping, Indicates the pixel-level label corresponding to the cropped sub-image; Step 1.3: Dataset division: divide the processed dataset D′ into the following three subsets in proportion: training set D train , validation set D val , test set D test , in order to improve the independence and generalization performance of model training and avoid overfitting.
3. The time series remote sensing crop classification method based on CNN-ViT according to claim 2 is characterized in that: The convolution feature extraction described in step 2 specifically includes: Step 2.1: Convolutional feature extraction module construction. Remote sensing images usually have significant spatial structure information and local feature distribution. Although Transformer is good at global feature modeling, its direct modeling ability for local spatial features is limited. In order to enhance the extraction of local spatial information, a convolutional feature extraction module is constructed to improve local perception capabilities. For each frame of image X in the input data X, (t) Through the lightweight convolution module f CNN Perform feature extraction and generate a convolution feature map; specifically, the following steps are included: First, the input image is feature extracted through the two-dimensional convolution layer Conv2D to extract local spatial structure information; then the feature map generated by the convolution operation is normalized to balance the data distribution and eliminate the magnitude difference between different features, while accelerating the model training process and stabilizing the gradient flow; finally, the ReLU activation function is introduced to make the activated features sparse, improve the feature discrimination ability, and effectively avoid the gradient disappearance problem; the final convolution feature map is generated It can provide input support for subsequent spatiotemporal feature modeling, as shown below: Among them, H′×W′ is the spatial resolution after convolution, and D is the channel dimension after convolution operation; Step 2.2: Patch segmentation and linear projection, after convolution feature extraction is completed, the generated feature map Still maintain a high resolution and number of channels; to further process the feature map, it is divided into non-overlapping patches of size h×w, each patch contains local spatial features; the total number of patches N after segmentation can be expressed as: For the feature map of the tth frame The i-th Patch is represented by P (t,i) , each P (t,i) Convolutional features of local areas are included, preserving high-dimensional channel information in the area: P (t,i) ∈R h×w×D ,i=1,2,…,N (8) In order to convert each Patch into a Token suitable for the input format of the Transformer encoder, each P (t,i) Do the following: First, change P (t,i) Expand to a one-dimensional vector; then, through a linear projection operation, the one-dimensional vector is mapped to P (t,i) Corresponding TokenZ (t,i) , expressed as: Among them, W is the weight matrix of linear projection, b is the bias vector of linear projection, and D Z It is the characteristic dimension of Token; Finally, using the same method, all patches of time frame t are converted into Token sequence Z (t) : Where N is the number of split patches, D Z It is the characteristic dimension of Token.
4. The time series remote sensing crop classification method based on CNN-ViT according to claim 3 is characterized in that: The encoder stage described in step 3 specifically includes: Step 3.1: Construction of spatiotemporal position coding. In order to retain the time information of each frame and the spatial position information of each patch in the time series remote sensing image, the generated Token sequence Z (t) Add time position coding and space position coding; Temporal position coding is used to represent the temporal position information of different frames in a time series. Since the acquisition time of remote sensing images is often irregular, it is difficult for traditional fixed temporal position coding methods to effectively model time intervals. Therefore, the present invention introduces dynamic temporal position coding by constructing a time lookup table P T ∈R T′×D , embed all the observation time information into it, where T′ is the number of all observation times in the data, and D is the feature dimension of Token; for the feature representation Z of time frame t (t) , by finding the corresponding time position code P T [t], add it to each Token to get the time-enhanced feature It is expressed as: Spatial position coding is used to identify the spatial position information of each patch within the image, ensuring that the spatial structural characteristics of the image are retained during the time series modeling process; the grid coordinate coding method is used to identify the coordinate position (p x ,p y ) generates two separate vectors encoding the row and column coordinates respectively, calculated as follows: in, is the sinusoidal component of the row coordinate position encoding, Encode the cosine component for the row coordinate position, is the sinusoidal component in the column coordinate position encoding, is the cosine component in the column coordinate position encoding; The row coordinate encoding and column coordinate encoding are concatenated to obtain the position encoding vector P of each Patch. S [i]: Stack the position encoding vectors of all patches to get the complete spatial position encoding matrix P S : And add the spatial position coding to the Token after the temporal position coding to get the representation of the Token after the temporal and spatial position coding: For the T frames of the entire time series, the Token sequence after all spatiotemporal position encoding is expressed as: Step 3.2: Construct a spatiotemporal encoder. The spatiotemporal encoder uses a Transformer module, which consists of a multi-head self-attention mechanism (MSA) and a feed-forward network (FFN) to capture long-range dependencies in the time dimension and model feature interactions in the spatial dimension. In each layer, the Token after the input spatiotemporal position encoding is first Normalize, then use the multi-head self-attention mechanism to capture the long-range dependencies in the time dimension and the feature interactions in the modeling space dimension, update the time series and spatial features, and the calculation process is as follows: After the processing of the temporal encoder and the spatial encoder, the final extracted feature is expressed as 5. The time series remote sensing crop classification method based on CNN-ViT according to claim 4 is characterized in that: The decoder stage described in step 4 specifically includes: Step 4.1: Pixel-level segmentation head construction, by final Perform a linear projection operation to convert the spatial features of each category into pixel-level prediction results; for category k and spatial position i, map the features through a multi-layer perception MLP to map the high-dimensional features to the low-dimensional space, and obtain the pixel-level probability distribution p of each category in the current spatial block. k,i ; Step 4.2: Spatial reorganization and merging. After completing the linear projection of all blocks, the pixel-level probability results of each block are reassembled back to the resolution of the original input image; by transforming the predicted probability p of each block into k,i Rearrange them into the global image space to generate a complete pixel-level prediction result map, thus achieving pixel-level classification of the entire remote sensing image.
6. The time series remote sensing crop classification method based on CNN-ViT according to claim 5 is characterized in that: The model training and use described in step 5 specifically include: In the model training and use phase, the training set D defined in step 1.3 is used. train , validation set D val , test set D test Perform the following tasks: Model training: Use the training set to optimize the model and gradually update the model parameters through block feature extraction, encoder processing, and decoder pixel-level semantic segmentation operations; Model validation: Use the validation set to adjust hyperparameters and evaluate the performance of the model on unseen data; Model testing: Use the test set to perform final performance evaluation on the model; Finally, through the full-process processing of the model, the category probability distribution corresponding to each pixel is generated for the input time series remote sensing image data to achieve classification prediction of crops.
7. A time series remote sensing crop classification device based on CNN-ViT, characterized in that: It comprises a memory and one or more processors, wherein the memory stores executable code, and when the one or more processors execute the executable code, it is used to implement the time series remote sensing crop classification method based on CNN-ViT according to any one of claims 1 to 5.
8. A computer-readable storage medium, characterized in that: A program is stored thereon, and when the program is executed by a processor, the time series remote sensing crop classification method based on CNN-ViT described in any one of claims 1 to 5 is implemented.
Citation Information
Patent Citations
Remote sensing image change detection method based on improved Transform twin network
CN115984700A
Vision Transform-LSTM-based multi-time-sequence remote sensing image crop classification method
CN118429715A
Multi-resolution transformer for video quality assessment
WO2023182987A1
Cited By
Crop classification method based on time sequence remote sensing image
CN115661554A
A crop classification method based on time-series remote sensing images
CN115661554B
Pressure sensor data shape recognition method
CN121564501A