Semantic and boundary joint learning-based time sequence remote sensing image crop classification method
By constructing a multi-level spatiotemporal feature extractor and semantic boundary collaborative network, the problem of spectral similar crop distinction difficulties and environmental interference in traditional crop classification methods is solved, high-precision crop classification is achieved, and the development of precision agriculture is supported.
Patent Information
- Application Number
- CN202510343859.0
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-03-22
- Publication Date
- 2025-07-04
AI Technical Summary
Traditional crop classification methods are difficult to distinguish crops with similar spectra, ignore morphological and texture changes during crop growth, and environmental factors interfere with reducing classification accuracy, making them inefficient in processing large-scale images.
A multi-level spatiotemporal feature extractor is built, combining ResNet and lightweight time attention encoder LTAE to capture the spatiotemporal semantic features of crops and plot boundary features, and fuse spatiotemporal information through semantic boundary collaborative network to optimize the loss function of multi-task joint learning module.
Significantly improve the accuracy of crop classification, improve the accuracy of crop category distinction and plot boundary definition, reduce missed inspections and missed inspections, and support the intelligent development of precision agriculture.
Smart Images

Figure CN120259880A_ABST
Abstract
Description
Technical Field
[0001] The present invention proposes a method for classifying crops in time - series remote - sensing images based on joint learning of semantics and boundaries, which relates to the research technical field of crop classification and extraction. Background Art
[0002] In the field of precision agriculture, accurately grasping the information of crop planting distribution is the key to realizing scientific planting management, yield prediction, and rational resource allocation. The crop classification technology based on time - series remote - sensing images is crucial. Traditional crop classification methods usually split the extraction of plot boundaries and crop semantic classification into independent steps. First, the boundaries of farmland plots are obtained by using image segmentation technology or manual digitization means, and then the crop types are judged according to the spectral characteristics (such as NDVI) of the crops within the plots. Although this method has certain effects, there are many problems. On the one hand, it is difficult to distinguish crops with similar spectra using a single spectral feature. On the other hand, traditional methods do not deeply mine the dynamic information of crop growth in time - series remote - sensing images. They only simply analyze the changes of spectral characteristics in the time series, ignoring the dynamic changes in aspects such as morphology and texture during the crop growth process. In addition, in practical applications, environmental factors such as clouds and the atmosphere interfere with the spectral information of the images, reducing the classification accuracy, and it is also difficult to meet the requirements in terms of efficiency when processing large - scale images. Summary of the Invention
[0003] To solve these problems, the present invention innovatively proposes a multi - scale spatio - temporal feature joint optimization mechanism. By constructing a multi - level spatio - temporal feature extractor, the spatio - temporal semantic features of crop types and fine plot boundary features are captured respectively, enabling the model to identify crop categories at the macroscopic level and pay attention to the details of plot boundaries during the feature learning process. At the same time, a semantic - boundary collaborative network is used to fuse these two types of features, deeply mining the spatio - temporal information in the images, and thus significantly improving the accuracy of crop classification. This innovative mechanism can effectively overcome the limitations of traditional methods that are difficult to carry out fine crop classification in an end - to - end manner, provide a more reliable technical guarantee for precision agriculture, contribute to the development of agricultural intelligence, improve the refined management level of agricultural production, and promote the sustainable development of agriculture.
[0004] The present invention proposes a method for classifying crops in time - series remote - sensing images based on joint learning of semantics and boundaries, including multi - scale spatio - temporal feature extraction and a multi - task semantic - boundary collaborative module. The method for classifying crops in time - series remote - sensing images based on joint learning of semantics and boundaries includes the following steps:
[0005] Step S1: Obtain medium - resolution time - series remote - sensing images and crop sample data of the study area, and pre - process the images;
[0006] Step S2: Based on the time-series remote sensing images in Step S1, construct a multi-level spatio-temporal feature extraction network to extract temporal and spatial features, including: using ResNet as the backbone for spatial feature extraction, extracting multi-level spatial texture details through the residual network, and capturing dynamic features in the long time series through the lightweight temporal attention encoder LTAE, and then obtaining spatio-temporal features at different levels through the fusion module;
[0007] Step S3: Based on the multi-level spatio-temporal feature extraction network in Step S2, construct a multi-task semantic and boundary joint learning network, including a boundary prior guidance module and a semantic and boundary enhancement module; among them, the boundary prior guidance module uses boundary information from low-level fine-grained to enhance the boundaries of high-level semantic feature representations, and the semantic and boundary enhancement module highlights semantic information and boundary details more, which can further improve the geometric accuracy of crop extraction results;
[0008] Step S4: Establish a multi-task joint loss function, select the optimizer of the model and perform parameter fine-tuning;
[0009] Step S5: Train the spatio-temporal feature extraction model in Step S3, and use the trained model for the remote sensing images and crop samples obtained in Step S1 to perform crop classification and extraction.
[0010] Furthermore, Step S1 also includes the following content:
[0011] Obtain medium-resolution time-series remote sensing images and crop sample data of the study area, and the content of image preprocessing includes: radiometric correction, geometric correction, georegistration, and image cropping.
[0012] Furthermore, Step S2 includes the following content:
[0013] Step S21: Use the pre-trained ResNet module as the backbone network for spatial feature extraction, make full use of its deep feature expression ability, so as to provide multi-level feature maps with rich semantics; among them, the core of the ResNet module is to solve the gradient disappearance and degradation problems in the training of deep networks through residual blocks (Residual Blocks), so that the network can effectively learn multi-level spatial representations;
[0014] Step S22: Construct a temporal attention encoder to process the temporal dimension of the feature map and extract temporal information; in order to efficiently capture dynamic features in the long time series, introduce the lightweight temporal attention encoder LTAE; the lightweight temporal attention encoder LTAE aims to make full use of the temporal dependencies in the lightweight temporal attention encoder LTAE to perform multi-step processing on the time series data on the premise of ensuring computational efficiency, so as to identify patterns and trends in temporal changes;
[0015] Among them, the lightweight temporal attention encoder (LTAE) is a variant of the lightweight self-attention mechanism in remote sensing time series classification; the lightweight temporal attention encoder (LTAE) calculates multiple triples from an input sequence of size to obtain multi-head attention T×d in , and produces T×n head ×d out (n head the number of output heads), d in refers to the feature dimension of each time step of the input data, d out refers to the dimension of the feature vector output after the attention mechanism.
[0016] The multi-head attention of the lightweight temporal attention encoder (LTAE) is applied along the channel dimension, and then the outputs of all heads are concatenated into a vector. The output of each head is defined as the sum of the corresponding inputs weighted by the attention mask in the time dimension; the attention mechanism of the lightweight temporal attention encoder (LTAE) can be summarized by the following equation for the h-th head, as follows:
[0017]
[0018] where x h refers to the h-th channel of the x-th, Q h refers to the h-th learnable main query, refers to the h-th transformation matrix of the key,
[0019] n head refers to the number of independent attention output heads calculated in parallel in the multi-head attention mechanism, d out refers to the dimension of the feature vector output after the attention mechanism.
[0020] Furthermore, step S2 also includes the following:
[0021] Step S23: Construct a multi-level spatio-temporal feature fusion module, and splice the multi-level spatial feature maps and multi-level temporal feature vectors on the channel through a feature-level fusion strategy; among them, constructing a multi-level spatio-temporal feature fusion module and splicing the multi-level spatial feature maps and multi-level temporal feature vectors on the channel through a feature-level fusion strategy includes fusing each stage of the ResNet module and the lightweight temporal attention encoder (LTAE) to generate spatio-temporal features of the corresponding level; this multi-level fusion can capture spatio-temporal relationships at different scales and levels of abstraction, and improve the expression ability of the model.
[0022] Furthermore, step S3 includes the following:
[0023] Step S31: Construct a semantic and boundary joint learning network; in the crop classification task based on time-series remote sensing images, low-level semantic information is rich in fine-grained boundary information features, which can accurately outline the edges of crop plots, while high-level semantic information has rich semantic data, which is conducive to accurately identifying crop categories. Based on the multi-level spatio-temporal features obtained by the multi-level spatio-temporal feature extraction network in Step S2, we construct a semantic and boundary joint learning network, which includes a boundary prior guidance module, a semantic information enhancement module, a boundary information enhancement module, and a multi-task joint learning network;
[0024] Step S32: Construct a boundary prior guidance module. Specifically, the boundary prior information obtained from low-level semantics is used to further improve the boundary details in high-level semantic features. Specifically, the boundary semantic features are extracted from the low-level semantic feature F B . First, a 1×1 convolutional operation is used to reduce the channel dimension, and then a Sigmoid function is used to generate a boundary mask as the guidance information; the crop semantic features are extracted from the high-level semantic feature F C , and are processed by a 1×1 convolutional operation to obtain semantic features that can be used for fusion; the boundary mask F B and the crop semantic feature F C are multiplied pixel by pixel with weights to achieve feature enhancement based on the boundary prior;
[0025] F R = σ 1×1 (F C ) + σ 1×1 (F C ) × δ(σ 1×1 (F B ))
[0026] where the meaning of F R is: the enhanced feature generated by fusing the original feature; F C and the boundary prior feature; F B .
[0027] Furthermore, Step S3 also includes the following content:
[0028] Step S33: Construct a semantic information enhancement module. Based on the convolutional operation and the normalized channel attention mechanism module, enhance the output farmland semantic information. Specifically, the input features are first processed in two branches: one branch directly passes through a 3×3 convolution, and the other branch extracts deep features through two 3×3 convolutions, as well as normalization and activation operations. Further, these two features are fused by channel-wise weighted summation, and a channel attention weight is generated through an additional 3×3 convolution, finally dynamically adjusting the input features. This module can effectively improve the feature expression ability of the model, enhance the attention to key channels, and at the same time suppress redundant information, thereby improving the overall performance of the model.
[0029] Step S34: Construct a boundary information enhancement module to enhance the low-level boundary detail features and output the plot boundary and plot mask information. Specifically, the boundary information enhancement module mainly combines depthwise separable convolution (DWconv), channel attention mechanism (cSE), and depthwise separable convolution (DSConv). This module can effectively enhance the extraction and expression ability of boundary information, enabling the model to more accurately identify boundaries and details when dealing with complex image segmentation tasks, thereby improving the overall performance.
[0030] Step S35: Construct a multi-task joint learning module, which generates classification maps, boundary maps, and masks Figure 3 as outputs, and optimizes classification accuracy, boundary refinement, and structured feature expression respectively. The classification map corresponds to the classification loss (Loss1), which is used to improve the prediction accuracy of crop categories; the boundary map corresponds to the boundary loss (Loss2), which strengthens the expression ability of the boundary; the mask map corresponds to the mask loss (Loss3), which captures fine-grained structural features to ensure regional integrity. The design of multi-task learning significantly improves the comprehensive performance of the model by jointly optimizing these loss functions.
[0031] Furthermore, Step S4 includes the following contents:
[0032] Step S41: Based on the image data preprocessed in Step S1, perform label making. The common label making method is manual interpretation. Collect crop classification sample labels, and divide the data into training set, validation set, and test set according to a certain proportion.
[0033] Step S42: Set hyperparameters, set the initial learning rate Lr, the optimizer type Adam, the training batch B of the model, and the number of model iterations E.
[0034] Furthermore, Step S4 also includes the following contents:
[0035] Step S43: Based on the multi-task semantic boundary collaboration module constructed in S3, for the output plot boundaries, plot masks, and classification feature maps, design a multi-task loss function. The calculation of the crop classification extraction loss can use the multi-class cross-entropy loss function, and the calculation formula is:
[0036]
[0037] where y[i] represents the true label of the i-th class, p[i] represents the predicted probability of the i-th class, n is the number of classes, and l Seg is the loss of the classification task;
[0038] For the boundary and mask tasks, use the binary cross-entropy loss function to evaluate the difference between the model prediction and the true label. The calculation formula is:
[0039]
[0040] where y[i] and p[i] represent the true label and predicted probability of the i-th pixel point respectively, m is the number of pixel points, and the calculations of the boundary and mask losses are similar. l Mask and l Bod represent the losses of the mask and boundary respectively.
[0041] Step S44: Aggregate the losses of different tasks with weights to obtain the total loss l Total , and the formula is:
[0042] l Total = θ1l Mask + θ2l Bod + θ3l Seg
[0043] In the formula, θ1, θ2, and θ3 represent the weighted coefficients of the losses of each task, indicating the contribution ratio of different tasks to the total loss. Gradually adjust the parameter sizes of the multi-task through cross-validation, and determine that the coefficients θ1 and θ2 of the mask and boundary of the auxiliary tasks are 0.15, and the coefficient θ3 of the main task classification is 0.7. Finally, obtain the total loss l Total .
[0044] Furthermore, step S5 includes the following contents:
[0045] Step S51: After the parameter setting is completed, use the dataset obtained in step S1 to train the multi-task time series remote sensing image crop classification model and save the optimal model weights;
[0046] Step S52: Use the weight file obtained in step S51 to classify the time series remote sensing images obtained in step S1 to obtain the crop classification results of the region.
[0047] The present invention has the following advantages:
[0048] 1. Comprehensive and efficient spatio-temporal feature extraction: The constructed multi-level spatio-temporal feature extraction network uses the ResNet module as the backbone for spatial feature extraction, effectively solving the problems of gradient disappearance and degradation in the training of deep networks and being able to accurately extract multi-level spatial texture details. At the same time, the lightweight temporal attention encoder LTAE captures the dynamic features in the long time series, and through the fusion module, different levels of spatio-temporal features are obtained, comprehensively and efficiently integrating the information of the time and space dimensions, providing a rich and high-quality feature basis for subsequent classification.
[0049] 2. Joint semantic and boundary learning to improve accuracy: The semantic and boundary joint learning network constructed based on multi-tasks, the boundary prior guidance module uses low-level fine-grained boundary information to enhance the boundaries of high-level semantic feature representations; the semantic and boundary enhancement module highlights semantic information and boundary details, significantly improving the geometric accuracy of the crop extraction results, playing a key role in crop category differentiation and boundary definition, and greatly improving the classification accuracy.
[0050] The present invention has the following beneficial effects: A method for classifying crops in time-series remote sensing images by jointly learning semantics and boundaries is constructed, fully mining the multi-level spatio-temporal features of the time-series images, and through the introduction of a semantic boundary collaboration module, end-to-end one-step crop classification is achieved. BRIEF DESCRIPTION OF THE DRAWINGS
[0051] Figure 1 It is a schematic flow chart of the method of the present invention.
[0052] Figure 2 It is a flow chart of the network model extraction of the present invention.
[0053] Figure 3 It is a model diagram of the boundary prior guidance module of the present invention.
[0054] Figure 4 It is a classification result diagram of some crops in a foreign region in a preferred embodiment of the present invention.
[0055] Figure 5 It is a classification result diagram of some crops in a domestic region in a preferred embodiment of the present invention. DETAILED DESCRIPTION OF THE INVENTION
[0056] The present invention will be further described below in conjunction with the drawings and embodiments.
[0057] It should be noted that the following detailed description is illustrative and is intended to provide further explanation of the present application. Unless otherwise specified, all technical and scientific terms used herein have the same meaning as commonly understood by those of ordinary skill in the technical field to which the present application belongs.
[0058] It should be noted that the terms used herein are only for describing specific embodiments and are not intended to limit the exemplary embodiments according to the present application; as used herein, unless the context clearly indicates otherwise, the singular form is also intended to include the plural form. In addition, it should also be understood that when the terms "comprising" and / or "including" are used in this specification, they indicate the presence of features, steps, operations, devices, components, and / or combinations thereof.
[0059] As Figure 1 shown, the present invention proposes a method for classifying crops in time-series remote sensing images based on joint learning of semantics and boundaries, including multi-scale spatio-temporal feature extraction and a multi-task semantic boundary collaboration module. The method for classifying crops in time-series remote sensing images based on joint learning of semantics and boundaries includes the following steps:
[0060] Step S1: Obtain medium-resolution time-series remote sensing images and crop sample data of the study area, and preprocess the images;
[0061] Step S2: Based on the time-series remote sensing images in Step S1, construct a multi-level spatio-temporal feature extraction network to extract temporal and spatial features, including: using ResNet as the backbone for spatial feature extraction, extracting multi-level spatial texture details through the residual network, and capturing dynamic features in the long time series through the lightweight temporal attention encoder LTAE, and then obtaining spatio-temporal features at different levels through the fusion module;
[0062] Step S3: Based on the multi-level spatio-temporal feature extraction network in Step S2, construct a multi-task semantic and boundary joint learning network, including a boundary prior guidance module and a semantic and boundary enhancement module; among them, the boundary prior guidance module uses boundary information from low-level fine-grained to enhance the boundaries of high-level semantic feature representations, and the semantic and boundary enhancement module highlights semantic information and boundary details, which can further improve the geometric accuracy of crop extraction results;
[0063] Step S4: Establish a multi-task joint loss function, select an optimizer for the model, and perform parameter fine-tuning;
[0064] Step S5: Train the spatio-temporal feature extraction model in Step S3, and use the trained model for the remote sensing images and crop samples obtained in Step S1 for crop classification and extraction.
[0065] Furthermore, Step S1 also includes the following content:
[0066] The content of obtaining medium-resolution time-series remote sensing images and crop sample data of the study area and preprocessing the images includes: radiometric correction, geometric correction, georegistration, and image cropping.
[0067] Further, step S2 includes the following content:
[0068] Step S21: Use a pre-trained ResNet module as the backbone network for spatial feature extraction, making full use of its deep feature expression ability to provide a semantically rich multi-level feature map. Among them, the core of the ResNet module is to solve the gradient disappearance and degradation problems in the training of deep networks through residual blocks (Residual Blocks), enabling the network to effectively learn multi-level spatial representations;
[0069] Step S22: Construct a temporal attention encoder to process the temporal dimension of the feature map and extract temporal information. To efficiently capture dynamic features in long time series, a lightweight temporal attention encoder LTAE is introduced. The lightweight temporal attention encoder LTAE aims to make full use of the temporal dependencies in the lightweight temporal attention encoder LTAE to perform multi-step processing on time series data on the premise of ensuring computational efficiency, so as to identify patterns and trends in temporal changes;
[0070] Among them, the lightweight temporal attention encoder LTAE is a variant of the lightweight self-attention mechanism in remote sensing time series classification. The lightweight temporal attention encoder LTAE calculates multiple triples from an input sequence of size to obtain multi-head attention T×d in and produces T×n head ×d out (n head the number of output heads), d in refers to the feature dimension of each time step of the input data, d out refers to the dimension of the feature vector output after the attention mechanism.
[0071] The multi-head attention of the lightweight temporal attention encoder LTAE is applied along the channel dimension, and then the outputs of all heads are concatenated into a vector. The output of each head is defined as the sum of the corresponding inputs weighted by the attention mask in the temporal dimension. The attention mechanism of the lightweight temporal attention encoder LTAE can be summarized by the following equation for the h-th head, as follows:
[0072]
[0073] where x h refers to the h-th channel of the x-th, Q h refers to the h-th learnable main query, refers to the h-th transformation matrix of the key,
[0074] n head refers to the number of independent attention output heads calculated in parallel in the multi-head attention mechanism, d outRefers to the dimension of the feature vector output after the attention mechanism.
[0075] Furthermore, step S2 further includes the following:
[0076] Step S23: Construct a multi-level spatio-temporal feature fusion module, and splice the multi-level spatial feature maps and the multi-level temporal feature vectors on the channel through a feature-level fusion strategy; among them, constructing a multi-level spatio-temporal feature fusion module and splicing the multi-level spatial feature maps and the multi-level temporal feature vectors on the channel through a feature-level fusion strategy includes fusing each stage of the ResNet module and the lightweight temporal attention encoder LTAE to generate spatio-temporal features at the corresponding level; this multi-level fusion can capture spatio-temporal relationships at different scales and abstraction levels, and improve the expression ability of the model.
[0077] Furthermore, step S3 includes the following:
[0078] Step S31: Construct a semantic and boundary joint learning network; in the crop classification task based on time-series remote sensing images, low-level semantic information is rich in fine-grained boundary information features, which can accurately outline the edges of crop plots, while high-level semantic information has rich semantic data, which is beneficial to accurately identify crop categories. Based on the multi-level spatio-temporal features obtained by the multi-level spatio-temporal feature extraction network in step S2, we construct a semantic and boundary joint learning network, which includes a boundary prior guidance module, a semantic information enhancement module, a boundary information enhancement module, and a multi-task joint learning network;
[0079] Step S32: Construct a boundary prior guidance module. Specifically, the boundary prior information obtained from low-level semantics is used to further improve the boundary details in the high-level semantic features; specifically, boundary semantic features are extracted from the low-level semantic feature F B , first reduce the channel dimension through a 1×1 convolution operation, and then generate a boundary mask through the Sigmoid function as the guiding information; crop semantic features are extracted from the high-level semantic feature F C , and processed through a 1×1 convolution operation to obtain semantic features available for fusion; the boundary mask F B and the crop semantic feature F C are multiplied pixel by pixel with weights to achieve feature enhancement based on the boundary prior;
[0080] F R = σ 1×1 (F C ) + σ 1×1 (F C ) × δ(σ 1×1 (F B ))
[0081] Among them, FR means: By fusing the original feature F C and the boundary prior feature F B the enhanced feature generated.
[0082] Furthermore, step S3 also includes the following:
[0083] Step S33: Construct a semantic information enhancement module, based on a convolutional operation and a normalized channel attention mechanism module, to enhance the output farmland semantic information; where the input feature is first divided into two branches for processing: one branch directly passes through a 3×3 convolution, and the other branch extracts deep features through two 3×3 convolutions as well as normalization and activation operations; further, these two features are fused by channel-wise weighted addition, and a channel attention weight is generated through an additional 3×3 convolution, and finally the input feature is dynamically adjusted; this module can effectively improve the feature expression ability of the model, enhance the attention to key channels, and at the same time suppress redundant information, thereby improving the overall performance of the model;
[0084] Step S34: Construct a boundary information enhancement module to enhance the low-level boundary detail features and output plot boundary and plot mask information; specifically, the boundary information enhancement module mainly combines depthwise separable convolution (DWconv), channel attention mechanism (cSE), and depthwise separable convolution (DSConv); this module can effectively enhance the extraction and expression ability of boundary information, enabling the model to more accurately identify boundaries and details when dealing with complex image segmentation tasks, thereby improving the overall performance;
[0085] Step S35: Construct a multi-task joint learning module, by generating a classification map, a boundary map, and a mask Figure 3 outputs, respectively optimizing classification accuracy, boundary refinement, and structured feature expression. The classification map corresponds to the classification loss (Loss1), which is used to improve the prediction accuracy of crop categories; the boundary map corresponds to the boundary loss (Loss2), which strengthens the expression ability of the boundary; the mask map corresponds to the mask loss (Loss3), which captures fine-grained structural features and ensures regional integrity. The design of multi-task learning significantly improves the comprehensive performance of the model by jointly optimizing these loss functions.
[0086] Furthermore, step S4 includes the following:
[0087] Step S41: Based on the image data preprocessed in step S1, perform label making, and the common label making is manual interpretation; collect crop classification sample labels, and divide the data into a training set, a validation set, and a test set according to a ratio;
[0088] Step S42: Hyperparameter setting, set the initial learning rate Lr, optimizer type Adam, training batch B of the model, and model iteration times E.
[0089] Further, step S4 also includes the following content:
[0090] Step S43: Based on the multi-task semantic boundary collaboration module constructed in S3, for the output plot boundary, plot mask, and classification feature map, design a multi-task loss function. The calculation of the crop classification extraction loss can use the multi-class cross-entropy loss function, and the calculation formula is:
[0091]
[0092] where y[i] represents the true label of the i-th category, p[i] represents the predicted probability of the i-th category, n is the number of categories, and l Seg is the loss for the classification task;
[0093] For the boundary and mask tasks, use the binary cross-entropy loss function to evaluate the difference between the model prediction and the true label. The calculation formula is:
[0094]
[0095]
[0096] where y[i] and p[i] represent the true label and predicted probability of the i-th pixel point respectively, m is the number of pixel points, and the calculation of the boundary and mask losses is similar. l Mask and l Bod represent the losses of the mask and boundary respectively.
[0097] Step S44: Aggregate the losses of different tasks with weights to obtain the total loss l Total , and the formula is:
[0098] l Total = θ1l Mask + θ2l Bod + θ3l Seg
[0099] In the formula, θ1, θ2, and θ3 represent the weighted coefficients of the losses of each task, indicating the contribution ratio of different tasks to the total loss. By cross-validation, gradually adjust the parameter sizes of the multi-task, and determine that the coefficients θ1 and θ2 of the mask and boundary of the auxiliary task are 0.15, and the coefficient θ3 of the main task classification is 0.7. Finally, obtain the total loss l Total .
[0100] Further, step S5 includes the following content:
[0101] Step S51: After the parameter setting is completed, use the dataset obtained in Step S1 to train a multi-task time series remote sensing image crop classification model and save the optimal model weights;
[0102] Step S52: Use the weight file obtained in Step S51 to classify the time series remote sensing images obtained in Step S1 to obtain the crop classification results of the region.
[0103] In this embodiment, all available Sentinel-2 remote sensing image sequences in a foreign region in 2021 are collected. After preprocessing in Step 1, there are 10 spectral bands, the spatial resolution is 10 meters, and the bands are the red, green, and blue bands. A total of 2433 images with a pixel size of 128×128 and the corresponding crop type labels are used for model training. The dataset collected in a certain place covers a domestic region, with a longitude range from 112°0′ to 113°10′ and a latitude range from 30°40′ to 31°40′. The remote sensing images of this dataset are taken by the Sentinel-2 satellite and contain 13 bands with different spatial resolutions (10 meters, 20 meters, and 60 meters). In the experiment, three bands with the lowest resolutions, namely Band 1 (coastal aerosol), Band 9 (water vapor), and Band 10 (SWIR - cirrus), are excluded. The JM dataset mainly contains plots with regular shapes and clear boundaries. These plots are usually large, but at the same time, they also contain a small number of small, scattered, and irregularly shaped farmlands.
[0104] Figure 2 Shows the crop classification method for time series remote sensing images based on the joint learning of semantics and boundaries constructed in this embodiment. This model is mainly composed of a spatio-temporal feature extraction network and a multi-task semantic boundary collaboration module, with time series remote sensing images as the input. First, the ResNet module is used to extract spatial features. Its residual structure effectively solves the gradient problem in the training of deep networks through multi-layer convolution, activation functions, and downsampling operations. Then, the lightweight time attention encoder LTAE is used to extract time features, which can adaptively focus on the importance of different time steps through the attention mechanism and capture the dynamic changes in crop growth. Subsequently, the feature fusion module fuses the multi-level time and spatial features, and the features are concatenated in the channel dimension through Concat. Thus, spatio-temporal features at different levels are gradually obtained, where the low levels retain the details of the image, and the high levels focus on abstract semantic information, as shown in the specific high-level and low-level feature maps. Based on the multi-level spatio-temporal features, the accuracy is further improved through the semantic and boundary joint learning network. Among them, the boundary prior guidance module uses the low-level fine-grained boundary information to enhance the boundaries of the high-level semantic features; the semantic information enhancement module highlights the semantic information to make the crop category recognition more accurate; the boundary information enhancement module enhances the low-level boundary details and outputs clear plot boundaries and mask information to improve the geometric accuracy of the crop extraction results.
[0105] As shown Figure 3 in the figure, the outputs of the low- and high-level spatio-temporal feature maps serve as the inputs to the semantic boundary collaboration module. The boundary semantics are generated from the low-level features. First, through a 1×1 convolutional layer, and then through the Sigmoid activation function to obtain enhanced boundary features; the crop type semantics come from the high-level features and are also processed through a 1×1 convolutional layer. The feature maps of both are fused through a multiplication operation. The fused boundary and semantic features are further processed through a 1×1 convolutional layer and enter the self-attention mechanism module, where Q, K, and V are Query, Key, and Value respectively. The self-attention mechanism further enhances the features by capturing global context information. The post-processed feature map is output through a 1×1 convolutional layer to obtain a refined crop classification map.
[0106] Figure 4 The figure shows the partial crop classification result map of a certain foreign region in this embodiment. The remote sensing images, labeled samples, and predicted classification results of two regions are respectively shown. It can be seen from the figure that whether it is small-scale crop planting in urban areas or large-scale crop planting in rural areas, the crop classification results obtained by the proposed method have a high consistency with the ground truth, and there are few missed detections and misdetections. The experimental results demonstrate the effectiveness of the proposed method for crop classification.
[0107] Figure 5 The figure shows the classification result map of a certain domestic dataset. Although this dataset only contains three classifications: winter wheat, winter rapeseed, and background areas, the degree of plot fragmentation is high, and there are many small, scattered, and irregularly shaped farmlands. Even in such a complex situation, the model still has good performance and can accurately extract information about smaller plots. The predicted output is relatively close to the labeled samples in terms of the distribution and boundary definition of small plots, indicating that the model has strong capabilities in capturing the features and identifying the boundaries of small plots and has certain advantages in dealing with small-scale and irregular plots.
[0108] The present invention has the following advantages:
[0109] 1. Comprehensive and efficient spatio-temporal feature extraction: The constructed multi-level spatio-temporal feature extraction network uses the ResNet module as the backbone for spatial feature extraction, effectively solving the problems of gradient disappearance and degradation in the training of deep networks and being able to accurately extract multi-level spatial texture details. At the same time, the lightweight temporal attention encoder LTAE captures the dynamic features in long time series, and through the fusion module, different-level spatio-temporal features are obtained, comprehensively and efficiently integrating the information in the time and space dimensions, providing a rich and high-quality feature basis for subsequent classification.
[0110] 2. Joint learning of semantics and boundaries improves accuracy: Based on the multi-task constructed semantic and boundary joint learning network, the boundary prior guidance module uses low-level fine-grained boundary information to enhance the boundaries of high-level semantic feature representations; the semantic and boundary enhancement module highlights semantic information and boundary details, significantly improving the geometric accuracy of crop extraction results, playing a key role in crop category discrimination and boundary definition, and greatly improving classification accuracy.
[0111] The present invention provides a method for classifying crops in time-series remote sensing images based on joint learning of semantics and boundaries, which fully excavates the multi-level spatio-temporal features of time-series images, and realizes end-to-end one-step crop classification by introducing a semantic boundary collaboration module; in addition, it integrates semantic information enhancement, boundary information enhancement and multi-task learning strategies, while improving semantic feature extraction, further supplementing fine-grained boundary information, thereby significantly reducing the phenomena of missed detection and misdetection, and providing technical support for efficiently extracting the regular boundaries of crops.
[0112] The above are the preferred embodiments of the present invention. All changes made according to the technical solutions of the present invention, when the functions and effects produced do not exceed the scope of the technical solutions of the present invention, shall fall within the protection scope of the present invention.
Claims
1. A method for classifying crops in time-series remote sensing images based on joint learning of semantics and boundaries, characterized in that, It includes a multi-scale spatio-temporal feature extraction and multi-task semantic boundary collaboration module. Among them, a method for classifying crops in time series remote sensing images based on joint learning of semantics and boundaries includes the following steps: Step S1: Obtain medium-resolution time series remote sensing images and crop sample data of the study area, and preprocess the images; Step S2: Based on the time series remote sensing images in Step S1, construct a multi-level spatio-temporal feature extraction network to extract temporal and spatial features, including: using the ResNet module as the backbone for spatial feature extraction, extracting multi-level spatial texture details through the residual network, and capturing dynamic features in the long time series through the lightweight time attention encoder LTAE, and then obtaining spatio-temporal features at different levels through the fusion module; Step S3: Based on the multi-level spatio-temporal feature extraction network in Step S2, construct a multi-task semantic and boundary joint learning network, including a boundary prior guidance module and a semantic and boundary enhancement module; among them, the boundary prior guidance module uses boundary information from low-level fine-grained to enhance the boundaries of high-level semantic feature representations, and the semantic and boundary enhancement module highlights semantic information and boundary details more; Step S4: Establish a multi-task joint loss function, select the optimizer of the model and perform parameter fine-tuning; Step S5: Train the spatio-temporal feature extraction model in Step S3, and use the trained model for the remote sensing images and crop samples obtained in Step S1 to perform crop classification and extraction.
2. The crop classification method for time series remote sensing images based on joint learning of semantics and boundaries according to claim 1, wherein, Step S1 also includes the following content: The content of obtaining medium-resolution time series remote sensing images and crop sample data of the study area and preprocessing the images includes: radiometric correction, geometric correction, georegistration, and image cropping.
3. A method for classifying crops in time series remote sensing images based on joint learning of semantics and boundaries according to claim 1, characterized in that, Step S2 includes the following content: Step S21: Use the pre-trained ResNet module as the backbone network for spatial feature extraction to obtain multi-level feature maps; among them, the ResNet module is used to solve the problems of gradient disappearance and degradation in the training of deep networks; Step S22: Construct a time attention encoder to process the time dimension of the feature maps and extract temporal information; among them, constructing the time attention encoder includes: the lightweight time attention encoder LTAE; the lightweight time attention encoder LTAE uses the time dependence relationship in the time series data to perform multi-step processing on the time series data and identify the patterns and trends in the temporal changes; Among them, the lightweight temporal attention encoder (LTAE) is a variant of the lightweight self-attention mechanism in remote sensing time series classification; the LTAE calculates multiple triples from the input sequence of size to obtain multi-head attention T×d in , and generates T×n head ×d out (n head number of output heads), where d in refers to the feature dimension of each time step of the input data, and d out refers to the dimension of the feature vector output after the attention mechanism; Among them, the multi-head attention of the lightweight time attention encoder LTAE is applied along the channel dimension, and the outputs of all heads are concatenated into a vector. The output of each head is defined as the sum of the corresponding inputs weighted by the attention mask in the time dimension; the attention mechanism of the lightweight time attention encoder LTAE is summarized by the following equation for the h-th head, as follows: where x h refers to the h-th channel of the x-th, Q h refers to the h-th learnable main query, refers to the h-th transformation matrix of the key, n head refers to the number of independent attention output heads for parallel computation in the multi-head attention mechanism, d out refers to the dimensionality of the feature vector output after passing through the attention mechanism.
4. A method for classifying crops in time - series remote - sensing images based on joint learning of semantics and boundaries according to claim 3, characterized in that, Step S2 also includes the following content: Step S23: Construct a multi-level spatio-temporal feature fusion module, and splice the multi-level spatial feature maps and the multi-level temporal feature vectors on the channels through a feature-level fusion strategy; wherein, constructing a multi-level spatio-temporal feature fusion module and splicing the multi-level spatial feature maps and the multi-level temporal feature vectors on the channels through a feature-level fusion strategy includes fusing each stage output by the lightweight temporal attention encoder (LTAE) and the lightweight temporal attention encoder (LTAE) to generate spatio-temporal features at the corresponding level.
5. A method for classifying crops in time series remote sensing images based on joint learning of semantics and boundaries according to claim 1, characterized in that, Step S3 includes the following: Step S31: Construct a semantic and boundary joint learning network; based on the multi-level spatio-temporal features obtained by the multi-level spatio-temporal feature extraction network in Step S2, construct a semantic and boundary joint learning network, which includes a boundary prior guidance module, a semantic information enhancement module, a boundary information enhancement module, and a multi-task joint learning network; Step S32: Construct a boundary prior guidance module, where boundary semantic features are extracted from low-level semantic features F B . First, the channel dimension is reduced through a 1×1 convolution operation, and then a boundary mask is generated through the Sigmoid function as guidance information; crop semantic features are extracted from high-level semantic features F C , processed through a 1×1 convolution operation to obtain semantic features that can be used for fusion; the boundary mask F B and the crop semantic features F C are multiplied pixel by pixel with weights to achieve feature enhancement based on boundary prior; F R = σ 1×1 (F C ) + σ 1×1 (F C ) × δ(σ 1×1 (F B )) Among them, F R means: by fusing the original feature; F C and the boundary prior feature; F B The enhanced feature generated.
6. The crop classification method for time series remote sensing images based on joint learning of semantics and boundaries according to claim 5, characterized in that Step S3 also includes the following: Step S33: Construct a semantic information enhancement module, including a channel attention mechanism module based on convolution operations and normalization, to enhance the output farmland semantic information; wherein, the input features are first divided into two branches for processing: one branch directly passes through a 3×3 convolution, and the other branch extracts deep features through two 3×3 convolutions, as well as normalization and activation operations; further, these two features are fused by channel-wise weighted summation, and a channel attention weight is generated through an additional 3×3 convolution, and finally the input features are dynamically adjusted; Step S34: Construct a boundary information enhancement module, including enhancing low-level boundary detail features and outputting plot boundary and plot mask information; wherein, the boundary information enhancement module mainly combines depthwise separable convolution (DWconv), channel attention mechanism (cSE), and depthwise separable convolution (DSConv); Step S35: Construct a multi-task joint learning module, including generating three outputs of a classification map, a boundary map, and a mask map to optimize classification accuracy, boundary refinement, and structured feature expression respectively; wherein the classification map corresponds to a classification loss, which is used to improve the prediction accuracy of crop categories; the boundary map corresponds to a boundary loss to strengthen the expression ability of the boundary; the mask map corresponds to a mask loss to capture fine-grained structural features and ensure regional integrity.
7. A method for classifying crops in time-series remote sensing images based on joint learning of semantics and boundaries according to claim 1, characterized in that Step S4 includes the following: Step S41: Perform label making based on the image data preprocessed in Step S1. The image data for label making includes collecting crop classification sample labels and dividing the data into a training set, a validation set, and a test set according to a ratio; Step S42: Hyperparameter setting, which includes setting the initial learning rate Lr, the optimizer type Adam, the training batch B of the model, and the number of model iterations E.
8. A method for classifying crops in time - series remote - sensing images based on joint learning of semantics and boundaries according to claim 7, characterized in that, Step S4 also includes the following: Step S43: Based on the multi-task semantic boundary collaboration module constructed in S3, for the output plot boundary, plot mask, and classification feature map, design a multi-task loss function. The calculation of the crop classification extraction loss uses a multi-class cross-entropy loss function, and the calculation formula is: Among them, y[i] represents the true label of the i-th category, p[i] represents the predicted probability of the i-th category, n is the number of categories, and l Seg is the loss for the classification task; For boundary and masking tasks, a binary cross-entropy loss function is used to evaluate the difference between the model prediction and the true label, and the calculation formula is as follows: Among them, y[i] and p[i] respectively represent the true label and predicted probability of the i-th pixel point, m is the number of pixel points, and the calculation of the boundary and the mask loss is similar. l Mask and l Bod respectively represent the losses of the mask and the boundary; Step S44: Aggregate the losses of different tasks with weights to obtain the total loss l Total , and the formula is: l Total = θ1l Mask + θ2l Bod + θ3l Seg In the formula, θ1, θ2, and θ3 represent the weighted coefficients of the losses of each task, indicating the contribution ratio of different tasks to the total loss.
9. A method for classifying crops in time - series remote sensing images based on joint learning of semantics and boundaries according to claim 1, characterized in that, Step S5 includes the following contents: Step S51: After the parameter setting is completed, use the dataset obtained in step S1 to train the multi-task time series remote sensing image crop classification model and save the optimal model weights; Step S52: Use the weight file obtained in step S51 to classify the time series remote sensing images obtained in step S1 to obtain the crop classification results of the region.
Citation Information
Cited By
Farmland segmentation and crop classification method and system based on multi-source remote sensing data deep learning
CN120808168A
Farmland intelligent identification method of interactive double-branch network based on deep learning
CN121353780A
Multi-scale space-time fusion method and system for realizing phenological recognition based on remote sensing image
CN122313340A
Multi-scale spatio-temporal fusion method and system for realizing phenology recognition based on remote sensing image
CN122313340B