An intelligent remote sensing extraction method for crops based on plot scale

Through the dual-branch feature extractor of the integrated convolutional neural network and Transformer network, combining attention mechanism and full convolutional neural network and long and short-term memory network, the problem of low crop classification accuracy in complex terrain areas is solved, and high-precision extraction of cultivated land plots and crop planting structures is achieved.

CN117132884BActive Publication Date: 2025-07-25FUZHOU UNIV +1
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202310849650.2
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2023-07-12
Publication Date
2025-07-25
Estimated Expiration
2043-07-12

AI Technical Summary

Technical Problem

The existing crop classification methods are difficult to effectively identify cultivated land plots and their crop planting structures in complex terrain areas. The traditional methods have poor generalization ability and insufficient feature extraction, resulting in low classification accuracy.

Method used

A two-branch feature extractor with integrated convolutional neural network and Transformer network is adopted, combining attention mechanism modules, full convolutional neural networks and long and short-term memory networks, a deep learning model is built to extract spatial and temporal features of plots and crops, and to improve the model's feature extraction capabilities.

Benefits of technology

The accuracy of crop classification has been improved, especially the accuracy of plot extraction and crop classification accuracy in complex terrain areas, and the construction of local precise agriculture and high-standard farmland databases has been supported.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN117132884B_ABST
    Figure CN117132884B_ABST
Patent Text Reader

Abstract

The present invention relates to a method for remotely sensing and intelligently extracting crops at the plot scale. First, the extraction of high-resolution remote sensing image plots is regarded as an image semantic segmentation problem, integrating a convolutional neural network and a Transformer network to construct a neural network model that takes into account spatial details and long-distance contextual semantics; based on the extracted fine plots, a temporal remote sensing dataset is constructed using medium-resolution images, and combining the representational advantages of the long short-term memory model and the fully convolutional neural network model in terms of temporal features and spatial features, a new temporal deep learning model is constructed for precise crop classification. The present invention makes full use of the advantages of convolutional neural networks, Transformer models, and long short-term memory models in terms of spatial detail, long-distance contextual information modeling, and temporal feature extraction. By integrating advanced deep network models, it effectively extracts plot image features and crop phenological features at different levels, and realizes precise classification of the crop planting structure at the plot scale.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the research field of cultivated land plot extraction and crop planting structure extraction, and particularly to a method for intelligent remote sensing extraction of crops based on the plot scale. Background Art

[0002] Crop planting information is important basic data for crop growth monitoring and yield estimation. Timely and accurately obtaining crop planting types and their spatio-temporal change information is of great significance for optimizing and adjusting crop planting structures and rationally allocating water and soil resources. Most of the existing crop classification methods take pixels as the analysis unit, and relatively few studies are carried out using plots as the basic analysis unit, especially for the complex terrain areas in the south. Currently, most of the methods for plot and crop extraction rely on a single convolutional neural network or object-oriented segmentation technology, but these methods are still difficult to effectively and accurately identify cultivated land plots and their crop planting structures in different regions.

[0003] Convolutional neural networks are good at extracting local semantic features but are insufficient in modeling long-distance context information, while the Transformer model is the opposite. Based on this, for the design of a plot extraction model, a dual-branch feature extractor can be constructed by integrating a convolutional neural network and a Transformer network, and the information extracted by both can be further utilized through a feature fusion module, thereby improving the extraction ability of thematic information in remote sensing images. In addition, long short-term memory network models are good at extracting temporal features, and fully convolutional neural networks are good at extracting spatial features. Therefore, for the design of a crop classification model, a fully convolutional neural network and a long short-term memory network model can be integrated to effectively and fully mine the phenological and spatial features of crops, thereby improving the feature extraction ability of the model and further enhancing the accuracy of crop classification.

[0004] Convolutional neural networks, Transformer networks, and long short-term memory neural networks have been widely used in natural image segmentation, instance segmentation, and object detection. The present invention will integrate these network models into the research on the extraction of crop planting structure information at the plot scale of remote sensing images. Summary of the Invention

[0005] The purpose of the present invention is to provide a method for intelligent remote sensing extraction of crops based on the plot scale, which overcomes the problems of weak generalization ability and insufficient feature extraction of traditional methods, and effectively improves the accuracy of crop classification.

[0006] To achieve the above purpose, the technical solution of the present invention is: a method for intelligent remote sensing extraction of crops based on the plot scale, including the following steps:

[0007] Step S1, obtaining high spatial resolution remote sensing images of the study area, and performing preprocessing operations on the images, including orthorectification, radiation correction, panchromatic fusion, image cropping, and image resampling;

[0008] Step S2: construct a deep network model integrating spatial details and long-distance context semantics, namely, the land parcel extraction model CLCFormer; specifically, based on a dual-branch network framework, different types of feature extraction networks are introduced to extract image semantic features at different levels; wherein the feature extraction network is composed of a convolutional neural network that is good at local feature extraction and a Transformer network that is good at long-distance global context semantic modeling;

[0009] Step S3: construct an attention mechanism module based on expert knowledge to integrate semantic features of different levels and differences;

[0010] Step S4, constructing farmland plot dataset;

[0011] Step S5, land parcel extraction model training and testing;

[0012] Step S6, obtaining a long-time series medium-resolution remote sensing image dataset of the study area, and performing preprocessing operations on the remote sensing images, including radiation correction, atmospheric correction, cloud removal, and fusion mosaicking;

[0013] Step S7, introducing the attention long short-term memory network model AttentionLSTM that is good at temporal feature extraction and the fully convolutional neural network FCN that is good at spatial information expression, and constructing a multivariate attention long short-term memory network model, namely the crop classification model ALSTM-FCN;

[0014] Step S8: Taking the plot as the basic analysis unit, constructing a plot-level crop time series feature dataset;

[0015] Step S9, training and testing the crop classification model ALSTM-FCN;

[0016] Step S10: Integrate the trained and tested plot extraction model CLCFormer and crop classification model ALSTM-FCN to complete the plot-level crop classification in the study area.

[0017] In one embodiment of the present invention, in step S2, the convolutional neural network EfficientNet-B3 and the Transformer network SwinV2-transformer are respectively selected as the feature extractors of the dual-branch architecture to capture the different levels M1, M2...M n (n≤4), discriminative and robust image semantic features.

[0018] In an embodiment of the present invention, step S3 specifically includes the following steps:

[0019] Step S31: Based on the multi-level features obtained in step S2 and the concept of the attention mechanism, construct a bidirectional feature fusion module BiFFM to fuse different and robust image semantics from different branches;

[0020] Step S32: Based on the features fused in step S31 and the attention mechanism, construct an attention gate module ATG with dilated convolution to improve the features;

[0021] Step S33: Based on the features improved in step S32, use an attention residual block ATR to enhance the feature expression ability of the model.

[0022] In an embodiment of the present invention, in step S31, the bidirectional feature fusion module BiFFM is formed as follows:

[0023] First, for the features generated by the convolutional neural network CNN branch, the processing method is as follows:

[0024] P c = δ(DSCscSEDFCC n ) × C n

[0025] In the formula, δ is the Sigmoid function, DSC is the depthwise separable convolution, scSE is the convolution module integrating spatial and channel attention, C n is the feature from the nth layer in the CNN branch, and P c is the refined feature;

[0026] For the features generated by the Transformer branch, the processing method is as follows:

[0027] S m = DSC(AvgPoolT n + DSC(MaxPool(T n )))

[0028] S F = δ(DSCReLUS m ) × T n

[0029] In the formula, T n is the feature from the nth layer in the Transformer branch, S m is the enhanced feature, ReLU is the rectified linear unit, AvgPool is the average pooling layer, MaxPool is the max pooling layer, and S F is the refined feature;

[0030] Finally, the feature expression ability of the model is enhanced by using concatenation operations and the attention residual block ATR. The specific implementation is as follows:

[0031] S F = Drop(ATRConcatP c , S F )

[0032] In the formula, Drop is a dropout layer used to prevent the model from overfitting.

[0033] In step S32, the attention gate module ATG is defined as follows:

[0034] S ED = δ(BNDSCBR(DSCE F + DSCD F )) × E F

[0035] In the formula, BR is a combination of the batch normalization layer BN and the rectified linear unit ReLU, E F comes from the features after BiFFM fusion, and D F comes from the features generated by the decoder;

[0036] In step S33, the attention residual block ATR is defined as follows:

[0037] O F = cSeσBR(σBRS ED + BN(σS ED ))

[0038] In the formula, σ is a 3×3 convolution, S ED is the feature after passing through the attention gate module ATG, and cSe is the spatial attention mechanism module.

[0039] In an embodiment of the present invention, step S4 is specifically implemented as follows: Based on the image data preprocessed in step S1, a label production method is selected, the plots in the remote sensing image are vectorized, the pixel values of the plots are assigned as A, and the pixel values of non-plots are assigned as B. Finally, a vector-to-raster algorithm is used to obtain the plot sample label dataset in the study area.

[0040] In an embodiment of the present invention, step S5 specifically includes the following steps:

[0041] Step S51: Based on the plot sample label dataset constructed in S41, the image and label data are cropped into a dataset of N×N pixel size, and the dataset is divided into a training set, a validation set, and a test set according to C:D:E. Data augmentation operations such as flipping, rotation, and color jittering are used to expand the dataset;

[0042] Step S52: Use the AdamW optimizer with an initial learning rate of Lr as the optimizer for the plot extraction model, and use the weighted cross-entropy loss WBCE and the weighted intersection over union loss WIoU as the loss functions of the plot extraction model. Set the training batch of the model to N and the number of model iterations to R;

[0043] Step S53: Use the pre-trained EfficientNet and Swin-transformer models on the large dataset ImageNet to initialize the network model weights to accelerate the model convergence speed and improve the model training efficiency, and obtain the optimal model weight file;

[0044] Step S54: Based on the model weight file obtained in Step S53, conduct model accuracy evaluation on the test dataset, and use the test-time augmentation method TTA to further improve the prediction effect of the model on the test set.

[0045] In an embodiment of the present invention, Step S7 specifically includes the following steps:

[0046] Step S71: Select the fully convolutional neural network FCN and the attention long short-term memory network AttentionLSTM model as the backbones of the crop classification model respectively, and capture strong spatial features and phenological features by constructing a dual-branch feature extractor;

[0047] Step S72: Integrate the convolutional block attention module CBAM in the FCN branch. Specifically, CBAM consists of a spatial attention module and a channel attention module:

[0048] O i = I SA (I CA (x F ))

[0049] I CA = δ(MLP(P max (x F )) + MLP(P avg (x F )))) × x F

[0050] I SA = δ(θ(Concat(S max (x F ), S avg (x F )))) × x F

[0051] In the formula, I CA is the result of the feature x F after passing through the channel attention module, and I SAIt is feature x F The result after passing through the spatial attention module. MLP is a multi-layer perceptron. Concat represents the concatenation operation. θ is a 7×7 convolution. P max is P avg They are the adaptive max pooling and adaptive average pooling operations respectively. S max and S avg are the maximum and average values of the features along the channel dimension.

[0052] In an embodiment of the present invention, the specific implementation manner of step S8 is as follows: Based on the image data preprocessed in step S6, with the plot boundary as the constraint, the average value of the pixels within the boundary range is taken as the feature value of the plot, the phenological feature calculation at the plot scale is carried out, different crop types are assigned to the plots according to the features, and the S-G filtering algorithm is used to smooth the phenological sequence to obtain the plot-level crop time series feature dataset.

[0053] In an embodiment of the present invention, step S9 specifically includes the following steps:

[0054] Step S91: Construct training sample data. Assign the plot sample labels of the target crops as 1, 2…n, where n is the number of crop categories to be classified, and assign the plot labels of non-target crops as 0. The dimension of the time series dataset is (N, Q T ×M), where N is the total number of samples, Q T is the time step of the time series dataset, and M is the number of variables to be processed at each time step;

[0055] Step S92: Model training: Use the Adam optimizer with an initial learning rate of Lr as the optimizer for the crop classification model ALSTM-FCN. Use the cross-entropy loss as the loss function of the model. Set the training batch of the model to N and the number of model iterations to R to obtain the optimal weight file of the model;

[0056] Step S93: Model prediction: Based on the weight file obtained in step S92, conduct the test and accuracy evaluation of the crop classification model ALSTM-FCN.

[0057] In an embodiment of the present invention, step S10 specifically includes the following steps:

[0058] Step S101: Use the already trained and tested plot extraction model CLCFormer to complete the extraction of the study cultivated land plots;

[0059] Step S102: Based on the already extracted cultivated land plots, use the already trained and tested crop classification model ALSTM-FCN to complete the plot-level crop classification in the study area.

[0060] Compared with the prior art, the present invention has the following beneficial effects: The present invention overcomes both the problem of low accuracy in plot extraction in complex terrain areas by traditional convolutional neural networks and the problem of insufficient extraction of crop spatial features by traditional long short-term memory neural networks. The method of the present invention further improves the accuracy of plot extraction in complex terrain areas by integrating convolutional neural networks and Transformer networks, and further improves the accuracy of crop classification by integrating long short-term memory networks and fully convolutional networks, providing a technical reference for local development of precision agriculture and the construction of high-standard farmland databases. BRIEF DESCRIPTION OF THE DRAWINGS

[0061] Figure 1 It is a schematic flowchart of the method of the embodiment of the present invention.

[0062] Figure 2 It is a structural diagram of the plot extraction model CLCFormer of the embodiment of the present invention.

[0063] Figure 3 It is a structural diagram of relevant modules in CLCFormer of the embodiment of the present invention.

[0064] Figure 4 It is a structural diagram of the crop classification model ALSTM-FCN of the embodiment of the present invention.

[0065] Figure 5 It is a detailed diagram of partial extraction results of the embodiment of the present invention. DETAILED DESCRIPTION OF THE EMBODIMENTS

[0066] The technical solutions of the present invention will be specifically described below with reference to the drawings.

[0067] It should be noted that the following detailed description is exemplary and is intended to provide further illustration of the present application. Unless otherwise specified, all technical and scientific terms used herein have the same meaning as commonly understood by those of ordinary skill in the technical field to which the present application belongs.

[0068] It should be noted that the terms used herein are only for describing specific embodiments and are not intended to limit the exemplary embodiments according to the present application. As used herein, unless the context clearly indicates otherwise, the singular form is also intended to include the plural form. In addition, it should be understood that when the terms "comprising" and / or "including" are used in this specification, they indicate the presence of features, steps, operations, devices, components, and / or combinations thereof.

[0069] As Figure 1 shown, this embodiment provides a remote sensing intelligent extraction method for crop planting structure based on plot scale, including the following steps:

[0070] Step S1, obtaining high spatial resolution remote sensing images of the study area, and performing preprocessing operations on the images, including orthorectification, radiation correction, full color fusion, image cropping, image resampling, etc.;

[0071] Step S2: Construction of a deep network model integrating spatial details and long-range contextual semantics, namely the CLCFormer model (e.g. Figure 3 Specifically, based on the dual-branch network framework, different types of feature extraction networks are introduced to extract image semantic features at different levels. The feature extraction network consists of a convolutional neural network that is good at local feature extraction and a Transformer network that is good at long-range global context semantic modeling.

[0072] Step S3: Based on step S2, an attention mechanism module is constructed based on expert knowledge to fully integrate semantic features of different levels and differences, so as to improve the accuracy of land extraction under complex terrain; specifically, by introducing an attention fusion module and a feature optimization module, the feature expression ability of the model is improved and the generalization ability of the model is enhanced;

[0073] Step S4, constructing farmland plot dataset;

[0074] Step S5, land parcel extraction model training and testing;

[0075] Step S6, obtaining a long-time series medium-resolution remote sensing image dataset of the study area, and performing preprocessing operations on the remote sensing images, including radiation correction, atmospheric correction, cloud removal, fusion mosaicking, etc.;

[0076] Step S7, introducing the attention long short-term memory network model AttentionLSTM that is good at temporal feature extraction and the fully convolutional neural network FCN that is good at spatial information expression, constructing a multivariate attention long short-term memory network model ALSTM-FCN, to improve the accuracy of crop classification in different regions and different seasons;

[0077] Step S8: Taking the plot as the basic analysis unit, constructing a plot-level crop time series feature dataset;

[0078] Step S9, training and testing the crop classification model ALSTM-FCN;

[0079] Step S10: Integrate the trained and tested plot extraction model CLCFormer and crop classification model ALSTM-FCN to complete the plot-level crop classification in the study area.

[0080] In this embodiment, step S2 specifically includes the following steps:

[0081] Step S21: Select the convolutional neural network EfficientNet-B3 and the Transformer network SwinV2-transformer as the feature extractors of the two-branch architecture respectively to capture different levels of M1, M2... M n (n ≤ 4), different, and robust image semantic features.

[0082] In this embodiment, step S3 specifically includes the following steps:

[0083] Step S31: Based on the multi-level features obtained in step S2, construct a bidirectional feature fusion module BiFFM based on the concept of the attention mechanism, fully fuse the different and robust image semantics from different branches, and improve the model's recognition of plots with high inter-class similarity;

[0084] Step S32: Based on the features fused in step S31, construct an attention gate module ATG with dilated convolution based on the attention mechanism to enhance the model's feature learning ability for plots of different sizes and shapes;

[0085] Step S33: Based on the features improved in step S32, use an attention residual block ATR to further improve the model's feature expression ability and alleviate the problem of blurred plot boundaries caused by crop occlusion.

[0086] In this embodiment, in step S31, the bidirectional feature fusion module BiFFM is formed as follows:

[0087] First, for the features generated by the convolutional neural network CNN branch, the processing method is as follows:

[0088] P c = δ(DSCscSEDFCC n )×C n

[0089] In the formula, δ is the Sigmoid function, DSC is the depthwise separable convolution, scSE is the convolutional module integrating spatial and channel attention, C n is the feature from the nth level in the CNN branch, and P c is the refined feature.

[0090] For the features generated by the Transformer branch, the processing method is as follows:

[0091] S m = DSC(AvgPoolT n + DSC(MaxPool(T n ))

[0092] S F=δ(DSCReLUS m )×T n

[0093] where T n is the feature from the nth layer in the Transformer branch, S m is the enhanced feature, ReLU is the rectified linear unit, AvgPool is the average pooling layer, MaxPool is the max pooling layer, and S F is the refined feature;

[0094] Finally, the feature expression ability of the model is enhanced by using concatenation operations and the attention residual block ATR, and the specific implementation method is as follows:

[0095] S F =Drop(ATRConcatP c , S F )

[0096] where Drop is the dropout layer used to prevent the model from overfitting.

[0097] In this embodiment, in step S32, the attention gate module ATG is defined as follows:

[0098] S ED =δ(BNDSCBR(DSCE F + DSCD F ))×E F

[0099] where BR is the combination of the batch normalization layer BN and the rectified linear unit ReLU, E F comes from the feature after BiFFM fusion, and D F comes from the feature generated by the decoder.

[0100] In this embodiment, in step S33, the attention residual block ATR is defined as follows:

[0101] O F =cSeσBR(σBRS ED + BN(σS ED ))

[0102] where σ is a 3×3 convolution, S ED is the feature after passing through the attention gate module ATG, and cSe is the spatial attention mechanism module used to enhance the significant features between channels.

[0103] In this embodiment, the specific implementation of step S4 is as follows: Based on the image data preprocessed in step S1, a common label-making method is selected, such as the surface vector construction method built into ArcGIS software, to vectorize the plots in the remote sensing image, assign the pixel value of the plot as A, and assign the non-plot pixel value as B. Finally, the vector-to-raster algorithm is used to obtain the plot sample label dataset in the study area;

[0104] In this embodiment, step S5 specifically includes the following steps:

[0105] Step S51: Based on the sample dataset constructed in S41, for the convenience of the experiment, the image and label data are cropped into a dataset of N×N pixel size, and the dataset is divided into a training set, a validation set, and a test set according to C:D:E. Data augmentation operations such as flipping, rotation, and color jitter are used to expand the training sample dataset and improve the generalization ability of the model;

[0106] Step S52: Use the AdamW optimizer with an initial learning rate of Lr as the optimizer for plot extraction, use the weighted cross-entropy loss WBCE and the weighted intersection over union loss WIoU as the loss functions of the plot extraction model, set the training batch of the model to N, and set the number of model iterations to R;

[0107] Step S53: Use the EfficientNet and Swin-transformer models pre-trained on the large dataset ImageNet to initialize the network model weights to accelerate the model convergence speed and improve the model training efficiency, and obtain the optimal model weight file;

[0108] Step S54: Based on the model weight file obtained in step S53, conduct model accuracy evaluation on the test dataset, and use the test-time augmentation method TTA to further improve the prediction effect of the model on the test dataset.

[0109] In this embodiment, step S7 specifically includes the following steps:

[0110] Step S71: Select the fully convolutional neural network FCN and the attention long short-term memory network Attention LSTM model as the backbones of the crop classification model respectively, and capture strong spatial features and phenological features by constructing a two-branch feature extractor;

[0111] Step S72: Integrate the convolutional block attention module CBAM in the FCN branch to further enhance the model's ability to capture spatial details and reduce the interference of the background area. Specifically, the CBAM module consists of a spatial attention module and a channel attention module:

[0112] O i =I SA (ICA (x F ))

[0113] I CA =δ(MLP(P max (x F )) + MLP(P avg (x F )))) × x F

[0114] I SA =δ(θ(Concat(S max (x F ), S avg (x F )))) × x F

[0115] In the formula, I CA is the result after the feature x F passes through the channel attention module, and I SA is the result after the feature x F passes through the spatial attention module. MLP is a multi-layer perceptron, Concat represents the concatenation operation, θ is a 7×7 convolution, P max is P avg which are adaptive max pooling and adaptive average pooling operations respectively, and S max and S avg are the maximum and average values of the features along the channel dimension.

[0116] In this embodiment, step S8 is specifically implemented as follows:

[0117] Based on the image data preprocessed in step S6, with the plot boundary as the constraint, the average value of the pixels within the boundary is taken as the feature value of the plot, the phenological features at the plot scale are calculated, different crop types are assigned to the plots according to the features, and the S-G filtering algorithm is used to smooth the phenological sequence to obtain the plot-level crop time series feature dataset.

[0118] In this embodiment, step S9 specifically includes the following steps:

[0119] Step S91: Construct training sample data. Assign the plot sample labels of the target crops as 1, 2…n (n is the number of crop categories to be classified), assign the plot labels of non-target crops as 0, and the dimension of the time series dataset is (N, Q T ×M), where N is the total number of samples, Q T is the time step of the time series dataset, and M is the number of variables to be processed at each time step;

[0120] Step S92, Model Training: Use the Adam optimizer with an initial learning rate of Lr as the optimizer for the crop classification model ALSTM-FCN. Use cross-entropy loss as the loss function of the model. Set the training batch of the model to N and the number of model iterations to R to obtain the optimal weight file of the model.

[0121] Step S93, Model Prediction: Based on the weight file obtained in Step S92, conduct the test and accuracy evaluation of the crop classification model ALSTM-FCN.

[0122] In this embodiment, Step S10 specifically includes the following steps:

[0123] Step S101, Use the already trained and tested plot extraction model CLCFormer to complete the extraction of the research cultivated land plots.

[0124] Step S102, Based on the already extracted cultivated land plots, use the already trained and tested crop classification model ALSTM-FCN to complete the plot-level crop classification in the research area.

[0125] In this embodiment, the GF-2 PMS image on September 19, 2019 in a certain town of a certain province is used as the data source for cultivated land plot extraction. The spatial resolution of the preprocessed image is 1m, and the bands are the red, green, and blue bands. This embodiment uses 5000 images with a pixel size of 256×256 and the corresponding plot labels for plot extraction model training. Based on the extracted plots, 15 scenes of Sentinel-1A radar time series remote sensing images from February to July 2019 in a certain town are used. After preprocessing, the VH and VV backscattering coefficient values are obtained, and a time series dataset of plot-level tobacco phenological characteristics of VV and VH is constructed. This embodiment uses 29234 tobacco samples and 41993 non-tobacco samples for model training.

[0126] As Figure 2 shown, it is the structure diagram of the plot extraction model CLCFormer constructed in this embodiment. It can be seen from the figure that CLCFormer consists of two different network branches, namely EfficientNet-B3 and SwinV2-Transformer. These two backbone networks are used to enhance the ability of CNN to capture long-distance global semantics and repair the insufficient extraction of fine-grained spatial details by Transformer. Next, use the bidirectional feature fusion module BiFFM to aggregate the different-level and different semantic features of the two branches, and use the attention gate module ATG with dilated convolution to expand the model receptive field to capture more context semantic information. Finally, use the attention residual block ATR to further improve the feature expression ability of the model and enhance the performance of the model.

[0127] As shown Figure 4 in the figure, it is the structure diagram of the crop extraction model ALSTM-FCN constructed in this embodiment. The model extracts the temporal features of crop growth based on the Attention Long Short-Term Memory Neural Network (AttentionLSTM), integrates the Fully Convolutional Neural Network (FCN) to extract the spatial features of crop types, and finally obtains the final crop spatial distribution map by fusing the features of the dual branches. Specifically, the input of ALSTM-FCN is the time-series remote sensing image data, which passes through the ALSTM and FCN branches respectively. The ALSTM branch first performs a dimension flipping operation, and then passes through the AttentionLSTM layer and the dropout layer to capture the robust crop temporal features. The FCN branch consists of 3 encoders and 1 global pooling layer. Each encoder includes a spatio-temporal convolutional layer, a normalization layer, and an activation function. The first two encoders end with CBAM attention. The time-series image data obtains the spatial features of crop types through the FCN module. Finally, the temporal and spatial features of the crop are fused through the concatenation layer to complete the extraction of the target crop.

[0128] As shown Figure 5 in the figure, it is a partial experimental result diagram of this embodiment in a certain town of a certain province. It can be seen from the figure that the plots extracted by using the proposed plot extraction model have clear boundaries and good shapes. The proposed method can accurately identify plots of different shapes and sizes. Secondly, the proposed method ignores the interference of irrelevant regions and reduces the missed extraction and mis-extraction. In addition, under the constraint of the plot boundary, the tobacco extracted by using the proposed crop classification model has regular boundaries and good shapes. The experimental results demonstrate the effectiveness of the proposed method for extracting the plot-level crop planting structure information.

[0129] The cultivated land plot extraction model of the present invention fully integrates the advantages of CNN and Transformer, overcomes the problems of insufficient feature extraction ability and low plot extraction accuracy of traditional methods, and improves the effect of plot extraction in complex terrain areas; in addition, the crop classification model of the present invention effectively utilizes the advantages of FCN and LSTM in spatial and temporal feature extraction, and improves the accuracy of crop classification. The method of the present invention fully considers the influence of different neural network models and different features on the cultivated land plot extraction result and the crop extraction result.

[0130] The above are the preferred embodiments of the present invention. All changes made according to the technical solution of the present invention, when the functions and effects produced do not exceed the scope of the technical solution of the present invention, shall fall within the protection scope of the present invention.

Claims

1. A method for remotely sensing and intelligently extracting crops based on the plot scale, characterized in that, It includes the following steps: Step S1: Obtain the high-spatial-resolution remote sensing images of the study area, and perform preprocessing operations on the images, including orthorectification, radiometric correction, panchromatic fusion, image cropping, and image resampling; Step S2: Construct a deep network model that integrates spatial details and long-distance context semantics, namely, the land parcel extraction model CLCFormer. Specifically, based on the dual-branch network framework, different types of feature extraction networks are introduced to extract image semantic features at different levels. The feature extraction networks are composed of a convolutional neural network that is good at local feature extraction and a Transformer network that is good at long-distance global context semantic modeling. The convolutional neural network EfficientNet-B3 and the Transformer network SwinV2-transformer are selected as feature extractors of the dual-branch architecture to capture different levels of M1, M2…M n (n≤4), differentiated, and robust image semantic features; Step S3: Construct an attention mechanism module based on expert knowledge to fuse different-level and different semantic features; specifically, it includes the following steps: Step S31: Based on the multi-level features obtained in Step S2, construct a bidirectional feature fusion module BiFFM based on the attention mechanism concept to fuse different and robust image semantics from different branches; Step S32: Based on the features fused in Step S31, construct an attention gate module ATG with dilated convolution based on the attention mechanism to improve the features; Step S33: Based on the features improved in Step S32, use an attention residual block ATR to enhance the feature expression ability of the model; In Step S31, the bidirectional feature fusion module BiFFM is formed as follows: First, for the features generated by the convolutional neural network CNN branch, the processing method is as follows: P c = δ(DSC(scSE(DFC(C n )))) × C n where δ is the Sigmoid function, DSC is the depthwise separable convolution, scSE is the convolutional module integrating spatial and channel attention, C n is the feature from the n-th level in the CNN branch, and P c is the refined feature; For the features generated by the Transformer branch, the processing method is as follows: S m = DSC(AvgPool(T n )) + DSC(MaxPool(T n )) S F = δ(DSC(ReLU(S m ))) × T n where T n is the feature from the n-th layer in the Transformer branch, S m is the enhanced feature, ReLU is the rectified linear unit, AvgPool is the average pooling layer, MaxPool is the max pooling layer, and S F is the refined feature; Finally, the feature expression ability of the model is enhanced by using concatenation operation and attention residual block ATR. The specific implementation method is as follows: S F = Drop(ATR(Concat(P c ,S F ))) In the formula, Drop is a dropout layer used to prevent the model from overfitting; In Step S32, the attention gate module ATG is defined as follows: S ED = δ(BN(DSC(BR(DSC(E F )) + DSC(D F )))) × E F where BR is a combination of the batch normalization layer BN and the rectified linear unit ReLU, E F is from the features after BiFFM fusion, D F is from the features generated by the decoder; In Step S33, the definition of the attention residual block ATR is as follows: O F = cSe(σ(BR(σ(BR(S ED )))))+ BN(σ(S ED )) where σ is a 3×3 convolution, S ED is the feature after passing through the attention gate module ATG, and cSe is the spatial attention mechanism module; Step S4: Construct a farmland plot dataset; Step S5: Train and test the plot extraction model; Step S6: Obtain the long-term medium-resolution remote sensing image dataset of the study area, and perform preprocessing operations on the remote sensing images, including radiometric correction, atmospheric correction, cloud removal, and fusion mosaicking; Step S7: Introduce an attention long short-term memory network model Attention LSTM that is good at extracting temporal features and a fully convolutional neural network FCN that is good at expressing spatial information, and construct a multi-variable attention long short-term memory network model, namely the crop classification model ALSTM-FCN; specifically, it includes the following steps: Step S71: Select the fully convolutional neural network FCN and the attention long short-term memory network Attention LSTM model as the backbones of the crop classification model respectively, and capture strong spatial features and phenological features by constructing a two-branch feature extractor; Step S72: Integrate a convolutional attention module CBAM in the FCN branch. Specifically, CBAM consists of a spatial attention module and a channel attention module: O i = I SA (I CA (x F )) I CA = δ(MLP(P max (x F )) + MLP(P avg (x F ))) × x F I SA = δ(θ(Concat(S max (x F ),S avg (x F ))))×x F Where, I CA is the result after the feature x F passes through the channel attention module, and I SA is the result after the feature x F passes through the spatial attention module. MLP is a multi-layer perceptron, Concat represents the concatenation operation, θ is a 7×7 convolution, and P max is P avg are the adaptive max pooling and adaptive average pooling operations respectively, and S max and S avg are the maximum and average values of the features along the channel dimension; Step S9: Construct a plot-level crop temporal feature dataset with plots as the basic analysis unit; Step S10: Integrate the trained and tested plot extraction model CLCFormer and the crop classification model ALSTM-FCN to complete the plot-level crop classification in the study area. ​ 2. The method for remotely sensing and intelligently extracting crops based on plot scale according to claim 1, wherein The specific implementation of step S4 is as follows: Based on the image data preprocessed in step S1, select a label-making method, vectorize the plots in the remote sensing image, assign the pixel value of the plot as A, and assign the non-plot pixel value as B. Finally, use the vector-to-raster algorithm to obtain the plot sample label dataset in the study area.

3. The method for remotely sensing and intelligently extracting crops based on plot scale according to claim 2, wherein, Step S5 specifically includes the following steps: Step S51: Based on the constructed plot sample label dataset, crop the image and label data into a dataset of N×N pixel size, and divide the dataset into a training set, a validation set, and a test set according to C:D:E. Use data augmentation operations such as flipping, rotation, and color jittering to expand the dataset. Step S52: Use the AdamW optimizer with an initial learning rate of Lr as the optimizer for the plot extraction model, and use the weighted cross-entropy loss WBCE and the weighted intersection over union loss WIoU as the loss functions of the plot extraction model. Set the training batch of the model to N and the number of model iterations to R. Step S53: Use the pre-trained EfficientNet and Swin-transformer models on the large dataset ImageNet to initialize the network model weights to accelerate the model convergence speed and improve the model training efficiency, and obtain the optimal model weight file. Step S54: Based on the model weight file obtained in step S53, conduct model accuracy evaluation on the test dataset, and use the test-time augmentation method TTA to further improve the prediction effect of the model on the test set.

4. A method for remotely sensing and intelligently extracting crops based on plot scale according to claim 1, characterized in that, The specific implementation of step S8 is as follows: Based on the image data preprocessed in step S6, with the plot boundary as the constraint, take the average value of the pixels within the boundary as the characteristic value of the plot, calculate the phenological characteristics at the plot scale, assign different crop types to the plots according to the characteristics, and use the S-G filtering algorithm to smooth the phenological sequence to obtain the plot-level crop time series characteristic dataset.

5. A method for remotely sensing and intelligently extracting crops based on plot scale according to claim 1, characterized in that, Step S9 specifically includes the following steps: Step S91: Construct training sample data. Assign the plot sample labels of the target crops as 1, 2... n, where n is the number of crop categories to be classified, and assign the plot labels of non-target crops as 0. The dimension of the time series dataset is (N, Q T ×M), where N is the total number of samples, and Q T is the time step of the time series dataset, and M is the number of variables to be processed at each time step; Step S92: Model training: Use the Adam optimizer with an initial learning rate of Lr as the optimizer for the crop classification model ALSTM-FCN, and use the cross-entropy loss as the loss function of the model. Set the training batch of the model to N and the number of model iterations to R to obtain the optimal weight file of the model. Step S93: Model prediction: Based on the weight file obtained in step S92, conduct the test and accuracy evaluation of the crop classification model ALSTM-FCN.

6. The method for remotely sensing and intelligently extracting crops based on plot scale according to claim 1, characterized in that, Step S10 specifically includes the following steps: Step S101: Use the trained and tested plot extraction model CLCFormer to complete the extraction of cultivated land plots in the study area. Step S102: Based on the extracted cultivated land plots, use the trained and tested crop classification model ALSTM-FCN to complete the plot-level crop classification in the study area.