Time sequence remote sensing image land coverage classification method based on multi-prototype cooperation mechanism

By using the multi-prototype collaboration mechanism and phenological invariant feature learning module in time-series remote sensing data, the classification misjudgment problem caused by crop growth diversity in large-scale land mapping is solved, and high-precision and stable crop mapping across phenological areas is achieved.

CN120219952APending Publication Date: 2025-06-27XIDIAN UNIV
View PDF 1 Cites 0 Cited by

Patent Information

Application Number
CN202510254239.X
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-03-05
Publication Date
2025-06-27

AI Technical Summary

Technical Problem

In large-scale land mapping tasks, due to the high classification misjudgment rate caused by crop growth diversity, deep models are prone to overconfidence during the training stage, but their generalization ability on the test dataset is relatively insufficient.

Method used

Through the feature mapping of time-series remote sensing data and the precise modeling of class-level feature spaces, a multi-prototype collaboration mechanism is adopted to generate fine-grained prototypes, optimize the spatial characteristics of the prototypes, and extract the adversarial learning strategy to align the feature distributions of different subclass spaces through the phenological invariant feature learning module.

Benefits of technology

It significantly improves the discriminant ability and interpretability of feature representation, enhances the generalization ability of the model, enables it to adapt to the complex environments of different phenological regions, and improves the classification accuracy and stability of large-scale cross-phenological regions crop mapping.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120219952A_ABST
    Figure CN120219952A_ABST
Patent Text Reader

Abstract

The invention discloses a time sequence remote sensing image land cover classification method based on a multi-prototype cooperation mechanism, and mainly solves the problem that an existing large-scale remote sensing charting method neglects crop growth diversity in a region, and consequently the error classification rate is too high, and the scheme comprises the steps that 1, feature mapping is conducted on an input time sequence image; 2) learning a plurality of prototypes for each category by using a method based on multiple prototypes, and capturing subtle differences under different growth distributions through a time sequence measurement perception mechanism; 3) based on a distance measurement penalty strategy, realizing accurate modeling of a feature space through multiple optimization targets; and 4) by using an adversarial learning strategy, through constructing a game optimization process between the feature extractor and the domain discriminator, realizing alignment of different subclass spatial features, and obtaining a final classification result. According to the method, the problem of inconsistent distribution caused by time sequence change can be effectively relieved, the model generalization ability is improved, and the method is easy to deploy and popularize in an agricultural actual scene with limited resources.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention belongs to the technical field of remote sensing image crop mapping, and further relates to a remote sensing image land classification method, specifically a method for classifying land cover in time-series remote sensing images based on a multi-prototype collaboration mechanism. It can be used for land planning and management. By identifying the crop types in the plot area, it provides a scientific basis for agricultural production, resource management, and environmental protection. Background Art

[0002] Land cover classification refers to using remote sensing technology and related calculation methods to identify and classify the crop types and their growth conditions in a specific area by analyzing the time-series image data obtained from satellite or aircraft sensors, also known as land crop mapping. This technology is widely used in fields such as agricultural monitoring, food production prediction, land use research, and environmental protection. With the continuous progress of remote sensing technology, remote sensing crop mapping has been significantly improved in terms of accuracy, efficiency, and application scope. However, due to the variability of crop growth, the growth state is easily affected by factors such as climate conditions, soil characteristics, and photoperiod changes, which easily causes misclassification phenomena and brings great difficulties and challenges to the accurate land mapping task.

[0003] Han et al. proposed a spatio-temporal multi-level attention crop mapping method based on time-series SAR imagery in their published paper "Spatio-temporal multi-level attention crop mapping method using time-series SAR imagery" (published in the ISPRS Journal of Photogrammetry and Remote Sensing, pages 293 to 310 in 2023). This method first extracts image features at different scales based on a feature pyramid, then uses the self-attention mechanism to extract temporal information and spatial information from multi-scale features respectively, and finally uses a feature fusion algorithm to fuse spatio-temporal information and trains a classifier to achieve the crop mapping process. However, this method of stacking Transformers will increase the complexity of the model, consume a large amount of computing resources, and at the same time, its generalization ability may lead to overlearning due to insufficient training data and it is difficult to directly generalize to other regions.

[0004] Jilin University proposed a large-scale cross-phenological zone crop mapping method based on time-series remote sensing images in its patent application "A large-scale cross-phenological zone crop mapping method based on time-series remote sensing images" (application number CN 202210919289.1 application publication number: CN 115439754A). This method extracts generalized deep phenological features by constructing a time-series spectral network, and then constructs a phenological alignment network based on the pre-trained network to achieve feature alignment between the source domain and the target domain. According to the multi-level deep phenological alignment loss, the distribution distance of different regions is constrained, and the knowledge transfer from the source domain to the target domain is realized, completing the cross-phenological region crop mapping. However, the growth diversity within the region is ignored, resulting in partial information loss during feature alignment, which reduces the classification accuracy in the crop mapping task. Summary of the invention

[0005] The purpose of the present invention is to address the deficiencies of the above-mentioned prior art and propose a time-series remote sensing image land cover classification method based on a multi-prototype collaborative mechanism to solve the problems of high classification misjudgment rate due to crop growth diversity in large-scale land mapping tasks, and deep models are prone to overconfidence in the training stage and relatively insufficient generalization ability on the test data set.

[0006] The present invention aims to realize large-scale regional crop mapping through feature mapping of time series remote sensing data and accurate modeling of class-level feature space. The specific implementation ideas are as follows: First, feature mapping is performed on the read-in time series remote sensing data, and multiple prototypes are learned for each category to fully capture the diversity characteristics of crop growth. Then, through the time series metric perception mechanism, the subtle differences under different growth distributions are captured, so as to more accurately characterize the growth dynamics of crops. On this basis, a prototype optimization module is designed, and the prototype is used to guide the feature optimization direction, forcing the distribution distance of heterogeneous samples in the feature space to be distanced, and at the same time, the distance between different distributions of the same type of samples is distanced, thereby significantly improving the clarity of the classification decision boundary and realizing accurate modeling of the class-level feature space. Finally, by aligning the sample distribution of crops under different growth characteristics, the phenological invariant characteristics of crops are extracted, and while effectively alleviating the distribution inconsistency problem caused by time series changes, the generalization ability of the model is further improved, so that it can adapt to the complex environment of different phenological zones. Through the above steps, the present invention realizes the accurate modeling of crop growth characteristics and the alignment of multi-distribution phenological characteristics, providing reliable technical support for large-scale cross-phenological zone crop mapping.

[0007] The present invention achieves the above-mentioned purpose by the following specific steps:

[0008] 1) Obtain land cover data from public datasets to construct a sample set and divide it into a training set and a test set;

[0009] (2) Construct a feature extractor F based on the ResNet network architecture. Input the time-series image data in the training set, and use F to obtain the pixel-level feature maps of all pixels in the time-series image X, and form a time-series feature map; Denote the pixel-level feature map of the i-th pixel x i as z i , where i = 1, 2,..., n, and n represents the number of pixels in X;

[0010] (3) Generate fine-grained prototypes:

[0011] (3a) Assign multiple prototypes to each category, establish a prototype set and use these prototypes as cluster centers for clustering, where E represents the crop category, K represents the predetermined number of potential distributions, and p e,k is the prototype of the k-th distribution in crop category e; Initialize all prototypes, that is, randomly select K samples from crop category E as the initial prototypes;

[0012] (3b) Design a time-series metric loss dist to calculate the shape similarity and distance similarity of time-series vectors:

[0013]

[0014] where α is a weighting factor, and d cos represents the cosine distance, and SCS represents the shape context similarity;

[0015] (3c) According to the time-series feature map obtained in step (2), use the prototype as the cluster center, identify the differences between features and prototypes through an unsupervised clustering method, and retrieve the prototype p i ' that is closest to z e,k , and obtain the distribution pseudo-label K' of each pixel;

[0016] (3d) Update p e,k ' using momentum according to the following formula to obtain the updated prototype p e,k ":

[0017] p e,k " = βp e,k '+ (1 - β)z e,k ,

[0018] where β ∈ (0, 1) represents the coefficient of momentum update, and z e,k represents the feature vector belonging to the k-th distribution in category e; Through multiple iterative updates, obtain the fine-grained prototype p e,k "';

[0019] (4) Optimize the spatial features of the fine-grained prototypes:

[0020] (4a) Design three feature optimization losses, including the between-class separability loss L cpe , the within-class separability loss L spe , and the clustering compactness loss L pe ; where L cpe is used to clarify the between-class classification boundary, L spe is used to clarify the classification boundaries of multiple distributions within a class, and L pe is used to improve the clustering compactness;

[0021] (4b) Use the three feature optimization losses designed in step (4a) to construct the total loss function L p :

[0022] L p = L pe + θ1L cpe + θ2L spe ,

[0023] where θ1 and θ2 respectively represent the weight coefficients of the between-class and within-class separability losses;

[0024] (4c) Use the loss function L p to optimize the parameters of the feature extractor through backpropagation and generate the optimized fine-grained prototype space features;

[0025] 5) Construct a classifier and achieve optimization through phenology-invariant feature learning:

[0026] (5a) Freeze the weights of the feature generator F and map the optimized fine-grained prototype space features through a multi-layer perceptron; design a classifier for crop classification and a discriminator for distinguishing the distribution labels to which the crops belong;

[0027] (5b) Construct an adversarial loss function L adv based on the gradient reversal layer technique, which includes the classifier loss and the discriminator loss;

[0028] (5c) Use the loss function L adv to optimize the parameters of the multi-layer perceptron, classifier, and discriminator through backpropagation to obtain the optimized classifier;

[0029] 6) Use the test set data as the input of the feature generator F and obtain the land cover classification results using the optimized multi-layer perceptron and classifier.

[0030] Compared with the prior art, the present invention has the following advantages:

[0031] First, since the present invention generates multiple prototype representations for each crop category through the fine-grained prototype generation module, it can accurately capture the phenological feature changes of crops at different growth stages. These prototypes not only encode the typical spectral features of crops but also reflect their spatio-temporal evolution laws, thus significantly enhancing the discriminative ability and interpretability of feature representations. Compared with traditional methods, this fine-grained modeling approach can better handle the problem of intra-class diversity and provide richer feature support for crop classification.

[0032] Second, the present invention introduces a class-invariant feature learning module that adopts an adversarial learning strategy to align the feature distributions in different sub-class spaces, which is one of the core innovations of the present invention. This module can extract class-invariant features that are robust to scene changes and effectively alleviate the problem of feature distribution deviation caused by factors such as environmental differences and phenological changes. This design significantly enhances the generalization ability of the model in cross-region tasks and enables it to maintain high accuracy and stability in complex scenarios.

[0033] Third, the present invention adopts a phased optimization strategy of first feature space modeling and adversarial learning, which significantly improves the training stability and convergence speed. Through the distance penalty strategy in the prototype optimization module, while expanding the range of intra-class feature distributions, the inter-class feature interval is maximized, realizing the precise modeling of the crop feature space. This optimization method can significantly enhance the adaptability and stability of the model in unseen scenarios and provide a reliable guarantee for the generalization ability of large-scale crop mapping. BRIEF DESCRIPTION OF THE DRAWINGS

[0034] Figure 1 is the implementation flowchart of the present invention;

[0035] Figure 2 is the schematic diagram of the prototype optimization method in the present invention;

[0036] Figure 3 is the comparison chart of the recognition performance simulation results of classification using the present invention and existing methods. DETAILED DESCRIPTION OF THE EMBODIMENTS

[0037] The following further describes the present invention with reference to the drawings.

[0038] Example 1: Referring to the attached Figure 1 , a method for land cover classification of time-series remote sensing images based on a multi-prototype cooperation mechanism proposed by the present invention specifically includes the following steps:

[0039] Step 1) Obtain land cover data from a public dataset to construct a sample set, and divide it into a training set and a test set;

[0040] Step 2) Construct a feature extractor F based on the ResNet network architecture. In this embodiment, the preferably constructed feature extractor includes a standard 7×7 convolutional block, a max pooling layer, and four residual blocks. Each residual block contains two 3×3 convolutional operations, a batch normalization layer BN, a rectified linear unit ReLU, a standard 1×1 convolutional operation, and a residual connection represented by a batch normalization layer BN. The extractor extracts features from the last three layers of the ResNet network, and uses the attention technology to extract temporal dimension features and spatial dimension features. Finally, the multi-scale information is fused to obtain a feature map. Input the temporal image data in the training set into F to obtain the pixel-level feature map of all pixels in the temporal image X, and form a temporal feature map; denote the pixel-level feature map of the i-th pixel x i as z i , where i = 1, 2,..., n, and n represents the number of pixels in X.

[0041] In this embodiment, the temporal feature map is specifically obtained according to the following steps:

[0042] (2a) Input the temporal image data in the training set, and set the number of temporal images input in each batch to B. Each temporal image X is organized into a four-dimensional tensor T×C×H×W containing a spatial structure and a time series, where T represents the time series length, C represents the number of polarization channels, and H×W represents the spatial image size 64×64;

[0043] (2b) Use ResNet50 as the basic feature extractor, and design a feature extractor F based on the ResNet network architecture;

[0044] (2c) Obtain the pixel-level feature map z i of the i-th pixel x i in the temporal image X through the feature extractor F:

[0045] z i = F(x i ), where i ∈ (1, n).

[0046] Among them, z i ∈R D , and D represents the updated temporal feature dimension; take i = 1, 2,..., n to obtain the pixel-level feature maps of all pixels in the temporal image X according to the above formula, and construct a temporal feature map Z ∈ R D×H×W .

[0047] Step 3) Generate fine-grained prototypes:

[0048] (3a) Assign multiple prototypes to each category and establish a prototype set And use these prototypes as cluster centers for clustering, where E represents the crop category, K represents the predetermined number of potential distributions, and p e,k is the prototype of the k-th distribution in crop category e; initialize all prototypes, that is, randomly select K samples from crop category E as the initial prototypes;

[0049] (3b) Design a temporal metric loss dist to calculate the shape similarity and distance similarity of temporal vectors:

[0050]

[0051] where α is a weighting factor, and d cos represents the cosine distance, and SCS represents the shape context similarity:

[0052]

[0053] where SCS(p e,k ,z i ) represents the shape similarity of the feature vectors of p e,k and z i , and represent the means of the feature vectors of p e,k and z i respectively, and represent the variances of the feature vectors of p e,k and z i respectively; j = 1, 2,..., t represents the j-th time point, and t is the length of the time series.

[0054] (3c) According to the temporal feature mapping graph obtained in step 2), use the prototypes as cluster centers, identify the differences between features and prototypes through an unsupervised clustering method, and retrieve the prototype p i ' that is closest to z e,k , and obtain the distribution pseudo-label K' of each pixel; in this embodiment, identify the differences between features and prototypes through an unsupervised clustering method, which is implemented according to the following formula:

[0055]

[0056] (3d) Update p e,k ' with momentum according to the following formula to obtain the updated prototype p e,k ":

[0057] p e,k " = βp e,k '+ (1 - β)z e,k ,

[0058] Among them, β∈(0,1) represents the coefficient of momentum update, z e,k represents the feature vector belonging to the kth distribution in class e; through multiple iterative updates, the fine-grained prototype p that characterizes the diverse distribution of crop phenology is obtained e,k ”';

[0059] Step 4) Optimize the spatial features of fine-grained prototypes:

[0060] (4a) Design three feature optimization losses, including the inter-class separability loss L cpe , intra-class separable loss L spe , clustering compactness loss L pe ; where L cpe Used to clarify the classification boundary between classes, L spe Used to clarify the classification boundaries of multiple distributions within a class, L pe Used to improve clustering density; the specific expression is as follows:

[0061]

[0062] in, represents the distance between feature z and the most dissimilar prototype among the fine-grained prototypes corresponding to its category; It represents the distance between the most similar prototype among the fine-grained prototypes corresponding to the feature z and its category; e' represents other categories outside the e category, and k' represents other distributions outside the kth distribution.

[0063] (4b) Use the three features designed in step (4a) to optimize the loss and construct the total loss function L p :

[0064] L p =L pe +θ1L cpe +θ2L spe ,

[0065] Among them, θ1 and θ2 represent the weight coefficients of inter-class and intra-class separable losses respectively;

[0066] (4c) Using the loss function L p The feature extractor parameters are optimized by back-propagation to generate optimized fine-grained prototype space features. This embodiment uses the Adam optimizer and sets the learning rate to 0.001 for optimization.

[0067] Step 5) Build a classifier and optimize it through phenological invariant feature learning:

[0068] (5a) Freeze the weights of the feature generator F, and map the optimized fine-grained prototype space features through a multi-layer perceptron; design a classifier for crop classification and a discriminator for distinguishing the distribution labels to which the crops belong. In this embodiment, the multi-layer perceptron consists of three linear layers, uses the GELU activation function, and maintains the original feature dimension; both the classifier and the discriminator use two standard 3×3 convolutional blocks, use the PRELU activation function, and then use a 1×1 convolutional layer to post-process the mapped features, and output the final mapped results and distribution label predictions.

[0069] (5b) Construct an adversarial loss function L based on the gradient reversal layer technique adv , including the classifier loss and the discriminator loss; the formula is as follows:

[0070] L adv =l ce (C e (F mlp ),e)+l ce (C k (R λ (F mlp ))),k),

[0071] where, F mlp represents the mapped features of the multi-layer perceptron, l ce represents the standard cross-entropy loss, C e represents the class classifier, C k represents the distribution classifier, R λ is the gradient reversal layer controlled by the hyperparameter λ.

[0072] (5c) Use the loss function L adv to optimize the parameters of the multi-layer perceptron, the classifier and the discriminator through backpropagation, and obtain the optimized classifier; in this embodiment, the Adam optimizer is used, and the learning rate is set to 0.001 for optimization.

[0073] Step 6) Use the test set data as the input of the feature generator F, and use the optimized multi-layer perceptron and classifier to obtain the land cover classification results.

[0074] Embodiment 2: Refer to the appendix Figure 1-2 , the overall implementation steps of the classification method proposed in this embodiment are the same as those in Embodiment 1. Now, a specific example is given to further describe the implementation process of the present invention in detail:

[0075] Refer to Figure 1, for the input sequence of temporal images, the present invention first obtains a feature map through a feature embedding module. In the fine-grained classification module, using the initialized prototype set, intra-class clustering is performed on the extracted features based on the class representation to capture the potential intra-class distribution. However, simply relying on distribution clustering cannot effectively achieve the comprehensive decoupling and optimization of the feature space. This is mainly because there is a lack of an effective gradient constraint mechanism to guide the feature update direction during the feature learning process, resulting in limited decoupling effect. Therefore, the present invention introduces a prototype separation module, which effectively realizes the separation and decoupling of the fine-grained feature space through an optimization method based on prototype learning, thereby enhancing the discriminability between features. At the same time, at the class-level feature level, a class-invariant feature learning module is designed. Through the alignment operation between multiple sub-distributions within a class, the module realizes the consistent generalization of features under different data distribution scenarios and largely eliminates the problem of distribution shift. The specific implementation process is as follows:

[0076] Step 1, construct a temporal feature map.

[0077] 1.1) Set the number of input temporal images in each batch. Set the number of input temporal images to 16 time series images in each batch, where the temporal dimension size of the time series image is set to T, which needs to be specified according to the actual data length. In this embodiment, it is set to 41 here. For all data, interpolation or slicing is required to ensure that the lengths of all input sequences are the same.

[0078] 1.2) Divide the temporal images into a training set, a validation set, and a test set. To ensure the reliability of the model prediction accuracy, in this embodiment, it is preferably allocated in a ratio of 6:4 for the training set and the test set;

[0079] 1.3) Organize the input data X into a four-dimensional tensor T×C×H×W containing a spatial structure and a time series, where T represents the sequence length, C represents the number of polarization channels, and H×W represents the spatial image size 64×64. Pixel-level feature mapping is achieved through the feature extractor F. For any x i There is where z i ∈R D , x i is any pixel in the space, and D represents the updated temporal feature dimension.

[0080] 1.4) In the setting of the feature extractor F in 1.3), ResNet50 is used as the basic feature extractor. The designed ResNet network architecture consists of a standard 7×7 convolutional block, a max pooling layer, and four residual blocks. Each residual block contains two 3×3 convolutional operations, batch normalization (BN), a rectified linear unit (ReLU), and a residual connection represented by a standard 1×1 convolutional operation and BN. After different convolutional and pooling operations, the spatial size of the input patch cube is reduced by 1 / 2, 1 / 4, and 1 / 8 respectively, and the channel dimension gradually increases. Then, the features of the last three layers of ResNet are extracted, and attention technology is used to extract the temporal dimension features and spatial dimension features, and the multi-scale information is fused. The technology based on multi-scale feature fusion can obtain richer and more discriminative feature representations.

[0081] Step 2, fine-grained prototype generation.

[0082] 2.1) As Figure 1 shown, first, multiple prototypes are assigned to each category to establish a prototype set, and these prototypes are used as clustering centers for clustering, where C represents the crop category, K represents the predetermined number of potential distributions, and p c,k is the prototype of the k-th distribution in crop category c. For any prototype, in this embodiment, one of the simplest settings is adopted for initialization, and K samples are randomly selected from the crop category C as the initial prototypes.

[0083] 2.2) To accurately perceive the difference between the features of different temporal pixels and the features of the generated prototypes, a temporal metric loss is designed to calculate the shape similarity and distance similarity in the temporal vector. The specific formula is designed as follows:

[0084]

[0085] where α is a weighting factor that needs to be manually set, which controls the weights of the cosine distance and shape similarity in dist, d represents the cosine distance, and SCS represents the shape context similarity.

[0086] 2.3) According to 1.3), for any pixel x i to obtain the pixel-level feature z i , through an unsupervised clustering method, using the prototype as the clustering center to identify the difference between the feature and the prototype, so as to capture the potential changes inside the crop. The specific design is as follows

[0087]

[0088] By retrieving z i and p c,kInfer the distribution label of each pixel based on the nearest prototype among them.

[0089] 2.4) To enable the prototype to fully capture the feature information in all samples while reducing the huge resource burden, the update of the prototype does not require introducing additional parameters to be optimized. Instead, momentum update is performed considering the online clustering results. The update process is as follows:

[0090] p c,k = βp c,k + (1 - β)z c,k ,

[0091] where β ∈ (0, 1) represents the coefficient of momentum update, and z c,k represents the feature vector belonging to the k-th distribution in class c. This strategy ensures that the prototype can timely reflect the changes in data distribution and effectively integrate the information of various samples. In this embodiment, preferably through three iterations, fine-grained prototypes characterizing the diverse distribution of crop phenology are obtained.

[0092] Step 3, Prototype space feature optimization;

[0093] 3.1) To better construct the complex class-level relationships in the remote sensing time domain, this embodiment conducts analysis from the following three aspects. First, prototypes of different classes should be separated from each other. This separation can reduce the confusion between different crop classes, thereby more effectively distinguishing the instances of each crop class. Second, multiple prototypes within the same class should also maintain an appropriate distance to depict rich intra-class relationships while expanding the feature space. Finally, the clustering centered on each prototype should be as compact as possible, thereby further enhancing the distinguishability between different prototypes, as Figure 2 shown.

[0094] 3.2) Referring to Figure 2 , considering that first, prototypes of different classes should be well separated to minimize the confusion between crop classes and promote more effective discrimination; second, multiple prototypes within the same class should maintain an appropriate distance to maximize the boundary between intra-class distributions while being clearly defined; finally, the clusters around each prototype should be as compact as possible to enhance the discrimination ability between different prototypes. The present invention designs three optimization strategies,

[0095]

[0096] where represents the distance between the prototype that is least similar to z within the class c to which z belongs. As Figure 2As shown in (a), the main purpose is to reduce the distance between the feature and the nearest prototype while increasing the distance to the prototypes that are not relevant to it. Here, the distance of the least similar prototype within the class is used to characterize the worst distribution of the class, which can effectively model the inter-class relationship and maintain clear inter-class separability.

[0097]

[0098] Among them represents the distance between the prototype that is most similar to z within the class c to which z belongs. As Figure 2 shown in (b), the feature should be further close to its corresponding prototype and relatively far from other prototypes, regardless of whether this prototype belongs to the same crop class as the feature. This can better describe the rich intra-class relationship and avoid the problem of blurring the classification boundary between different crop classes.

[0099]

[0100] To improve the compactness of each cluster with the prototype as the clustering center so that it can be distinguished, the features should be closely aggregated near their respective prototypes, by minimizing the distance between each feature and its designated prototype, thereby effectively reducing the intra-cluster difference.

[0101] 3.3) Total loss representation:

[0102] L p = L pe + θ1L cpe + θ2L spe ,

[0103] where θ1 and θ2 represent the weight coefficients respectively, and the loss L p includes L pe , L cpe , L spe respectively optimize the overall learning process from three different levels of distribution compactness, intra-class relationship, and inter-class relationship. Therefore, by constructing class-level prototype optimization, the complex class-level relationships in the temporal feature space can be fully described, the expression ability of the feature space is effectively improved, the separation degree of crop classes is significantly improved, and the risk of misclassification is reduced.

[0104] 3.4) In the loss L pAfter the calculation is completed, the model is optimized based on backpropagation. Using the Adam optimizer with a learning rate set to 0.001, the Adam optimizer combines the advantages of the momentum method and adaptive learning rate, enabling efficient adjustment of the parameter update step size, thus accelerating model convergence and avoiding getting stuck in local optima. Through backpropagation optimization, the feature extractor can gradually adjust its parameters to generate prototype features with stronger distribution representation capabilities. This process not only enhances the discriminability of the features but also enables the model to better capture the internal structure and semantic information of the data, thereby improving the overall performance.

[0105] Step 4: Phenology-invariant feature learning.

[0106] 4.1) First, freeze the weights of the feature generator F to prevent changing the model parameters during this process, and obtain the features optimized in 3.4). Subsequently, these features are mapped through the MLP block. A classifier for crop classification and a discriminator for distinguishing the distribution labels to which the crops belong are designed. The MLP block consists of three linear layers, uses the GELU activation function, and maintains the original feature dimension. Both the classifier and the discriminator use two standard 3×3 convolutional blocks, and then use a 1×1 convolutional layer to post-process the mapped features, thereby generating the final mapped results and distribution label predictions.

[0107] 4.2) Use the gradient reversal layer technique to update the loss of the distribution discriminator. This method enables adversarial training, making it impossible for the discriminator to distinguish the differences between distributions. This forces the sub-distribution features from different distributions to align and achieves phenology-invariant feature learning. The specific formula is as follows:

[0108] L adv =l ce (C c (F mlp ),c)+l ce (C k (R λ (F mlp )),k),

[0109] where F mlp represents the mapped features of the multi-layer perceptron, l ce represents the standard cross-entropy loss, C c represents the class classifier, C k represents the distribution classifier, and R λ is the gradient reversal layer controlled by the hyperparameter λ.

[0110] 4.3) In the loss L advAfter the calculation is completed, the model is optimized based on backpropagation. Using the Adam optimizer with a learning rate set to 0.0001, this process only updates the parameters of the MLP module, classifier, and discriminator. This design is based on adversarial learning. By introducing a discriminator to distinguish features from different distributions, the model is forced to learn more generalizable feature representations. In this way, the model can better adapt to diverse input distributions and enhance its robustness and generalization ability in complex scenarios.

[0111] Embodiment 3: Refer to the appendix Figure 1-2 , the overall implementation steps of the classification method proposed in this embodiment are the same as those in Embodiment 1 or 2. Now, several functional modules involved in the present invention will be further described:

[0112] Refer to Figure 1 , the present invention first obtains a pixel-level feature map from time series data through a feature encoder; then, through a fine-grained prototype generation module, calculates the distance between features and the initialized prototype set according to a temporal metric that combines distance similarity and shape similarity, and performs a clustering process; then captures the potential differences inside the data through the prototype to generate fine-grained prototypes. After that, through a prototype optimization module, using three loss functions, optimizes the features from three directions: between classes, within classes, and feature aggregation degree, so as to realize the construction of the class-level space; secondly, through a phenology-invariant feature learning module, freezes the previously trained feature encoder, and learns phenology-invariant features through adversarial training of the classifier and discriminator to enhance the generalization ability of the model.

[0113] Fine-grained prototype generation module: In order to effectively capture complex spatio-temporal scenarios and potential phenological changes inside crops, a fine-grained prototype generation module is adopted. That is, first, multiple prototypes are assigned to each category to establish a fine-grained prototype set, and these prototypes are used as clustering centers for clustering. Then, an unsupervised clustering method is used, with the prototype as the clustering center to identify the differences between features and prototypes, so as to capture the potential changes inside the crop.

[0114] Prototype optimization module: In fine-grained prototype classification, the feature labels and their potential distributions of each category are obtained through prototype clustering, so as to construct a basic spatial structure. However, it should be noted that the main purpose of the clustering algorithm is to aggregate data and extract feature labels, rather than optimize the features themselves. Since this process is unsupervised, there may still be couplings between similar features, which may affect the accuracy of the clustering result. Therefore, the present invention designs three loss strategies from three perspectives, as Figure 2As shown below; first, prototypes of different categories should be separated from each other, and this separation can reduce the confusion between different crop categories, thereby more effectively distinguishing instances of each crop category. Second, multiple prototypes within the same category should also maintain an appropriate distance to depict rich intra-class relationships and expand the feature space at the same time. Finally, the clustering centered on each prototype should be as compact as possible, thereby further enhancing the distinguishability between different prototypes.

[0115] Invariant Feature Learning Module: In large-scale mapping tasks, there may be some unknown potential distributions that have not appeared in the data used for training. To address this issue, the phenological invariant features of crops can be extracted by learning invariant features, thereby improving the generalization ability of the model and adapting to more complex distribution scenarios. Specifically, by using a gradient reversal layer (a commonly used technique that promotes adversarial training by reversing the gradient), the distribution classifier loss is updated to form adversarial training, which forces the alignment of sub-distribution features under different distributions and realizes class-invariant feature learning. It should be noted that although the present invention also employs adversarial learning techniques, compared with traditional direct adversarial learning methods, the present method has significant advantages. Specifically, before implementing adversarial learning, the present invention first constructs a structured feature representation space through feature space modeling. This prior feature space modeling provides good initialization conditions for subsequent adversarial learning, enabling the model to be optimized in a relatively stable and semantically clear feature space, thereby effectively alleviating the common training instability problem in traditional adversarial learning methods.

[0116] The effects of the present invention will be further described below in conjunction with simulation experiments.

[0117] 1. Simulation Conditions:

[0118] The hardware platform used in the present invention is as follows: The CPU uses an eight-core and eight-thread Intel Core i7-9700k with a main frequency of 3.6 GHz and a memory of 64 GB; the GPU uses two Nvidia RTX 3090s with a video memory of 24 GB each. The software platform used is as follows: The operating system uses Ubuntu16.04LTS, the deep learning computing framework uses PyTorch 1.10, and the programming language uses Python 3.9.

[0119] In this simulation experiment, mF1 and mIOU are used as evaluation indicators for crop mapping. Among them, mF1 is used to measure the comprehensive performance of the model in the classification task, and the classification accuracy and balance of the model for various crops are evaluated by calculating the average value of the F1-score for each category; mIOU is used to measure the regional overlap accuracy of the model in the semantic segmentation task, and the ratio of the intersection to the union of the predicted region and the true region is calculated, and the average value of the IOU for each category is taken to evaluate the positioning accuracy and segmentation effect of the model on the crop boundary.

[0120] 2. Simulation Content and Result Analysis

[0121] The proposed method and five existing methods, namely ConvLSTM, 3D Unet, UTAE, TPE-RNN, and STMA, are tested on the temporal remote sensing datasets PASTIS-R and Brandenburg under the above simulation conditions, and the tracking results are evaluated using the above evaluation indicators. The results are shown in Table 1.

[0122] Table 1 Mapping Comparison Results between the Existing Technology and the Proposed Method on PASTIS-R and Bransenburgs

[0123]

[0124] As shown in Table 1, the performance of different methods is presented. By observing the results, it is easy to find that the proposed method outperforms other models on both datasets, demonstrating its superiority in the field of crop mapping. On the PASTIS-R dataset, MDPL achieved an mIOU of 65.78% and an mF1 of 76.82%, which are higher than all other comparison methods. In particular, the improvement in the mF1 score indicates that the MDPL method can effectively capture positive samples while maintaining a high accuracy rate, with better precision and recall. In the Brandenburg dataset, compared with other methods, the average performance improvement of MDPL shows that the mIOU increased by 6.95% and the mF1 increased by 10.24%. It can be found that mf1 has a significant increase, and at the same time, it has strong adaptability to remote sensing scenarios with uneven class distributions.

[0125] Through Figure 3 the results shown, it can be seen that under different training sample ratios, the performance of the proposed method reaches the optimal effect. Even in the case of a low training sample ratio, the model still exhibits excellent performance, which fully demonstrates the significant advantage of the proposed method in generalization. This characteristic makes the proposed method of great value in practical applications, especially in scenarios where training data is scarce or the annotation cost is high, as it can significantly reduce the dependence on the amount of data while maintaining a high performance level.

[0126] In summary, the method for land cover classification of temporal remote sensing images based on the multi-prototype collaborative mechanism has important potential and value in cartographic methods. From the experimental results, this method performs excellently on both the temporal remote sensing datasets PASTIS-R and Brandenburg. By comparing the performance metrics of different methods on these datasets, this method achieves the best results. It shows that this framework can significantly distinguish similar crop categories in complex scenarios, reduce the misrecognition rate, and exhibits excellent generalization performance. The method of the present invention is easy to deploy and promote in the actual agricultural scenarios with limited resources, saving a large amount of time and cost for agricultural management departments and research institutions, and can provide reliable technical support for global agricultural monitoring and large-scale crop mapping.

[0127] The above simulation analysis proves the correctness and effectiveness of the method proposed by the present invention.

[0128] The parts not detailed in the present invention belong to the common general knowledge of those skilled in the art.

[0129] The above are only the preferred embodiments of the present invention, and are not intended to limit the present invention. Obviously, for professionals in the field, after understanding the content and principle of the present invention, various modifications and changes in form and details may be made without departing from the principle and structure of the present invention. However, these modifications and changes based on the idea of the present invention are still within the scope of protection of the claims of the present invention.

Claims

1. A land cover classification method for time series remote sensing images based on a multi-prototype collaborative mechanism, characterized in that: The steps include: 1) Obtain land cover data from public datasets to construct a sample set and divide it into a training set and a test set; 2) Construct a feature extractor F based on the ResNet network architecture, input the time series image data in the training set, and use F to obtain the pixel-level feature map of all pixels in the time series image X to form a time series feature map; The i-th pixel x i The pixel-level feature map is denoted as z i , i = 1, 2, ..., n, n represents the number of pixels in X; 3) Generate fine-grained prototypes: (3a) Assign multiple prototypes to each category and establish a prototype set These prototypes are clustered as cluster centers, where E represents the crop category, K represents the potential distribution of the predetermined number, and p e,k is the prototype of the kth distribution in crop category e; all prototypes are initialized, that is, K samples are randomly selected from crop category E as the initial prototypes; (3b) Design the time series metric loss dist to calculate the shape similarity and distance similarity of the time series vector: Among them, α is the weighting factor, d cos represents cosine distance, SCS represents shape context similarity; (3c) Based on the time series feature map obtained in step 2), the prototype is used as the cluster center, and the difference between the feature and the prototype is identified through unsupervised clustering method, and the distance z is obtained by retrieval. i The most recent prototype e,k ', and obtain the distribution pseudo label K' of each pixel; (3d) According to the following formula, p e,k 'Perform momentum update and obtain the updated prototype p e,k ”: p e,k ”=βp e,k '+(1-β)z e,k , Among them, β∈(0,1) represents the coefficient of momentum update, z e,k represents the feature vector belonging to the kth distribution in class e; through multiple iterative updates, the fine-grained prototype p that characterizes the diverse distribution of crop phenology is obtained e,k ”'; 4) Optimizing the spatial features of fine-grained prototypes: (4a) Design three feature optimization losses, including the inter-class separability loss L cpe , intra-class separable loss L spe , clustering compactness loss L pe ; where L cpe Used to clarify the classification boundary between classes, L spe Used to clarify the classification boundaries of multiple distributions within a class, L pe Used to improve clustering density; (4b) Use the three features designed in step (4a) to optimize the loss and construct the total loss function L p : L p =L pe +θ1L cpe +θ2L spe , Among them, θ1 and θ2 represent the weight coefficients of inter-class and intra-class separable losses respectively; (4c) Using the loss function L p Optimize the feature extractor parameters through back-propagation to generate optimized fine-grained prototype space features; 5) Build a classifier and optimize it through phenological invariant feature learning: (5a) Freeze the weights of the feature generator F and map the optimized fine-grained prototype space features through a multi-layer perceptron; design a classifier for crop classification and a discriminator for distinguishing the distribution labels to which the crops belong; (5b) Constructing adversarial loss function L based on gradient reversal layer technology adv , including classifier loss and discriminator loss; (5c) Using the loss function L adv Optimize the parameters of the multilayer perceptron, classifier and discriminator through back propagation to obtain the optimized classifier; 6) Use the test set data as the input of the feature generator F, and use the optimized multi-layer perceptron and classifier to obtain the land cover classification results.

2. The method according to claim 1, characterized in that: The time series feature map in step 2) is obtained according to the following steps: (2a) Input the time series image data in the training set, and set the number of input time series images in each batch to B, where each time series image X is organized into a four-dimensional tensor T×C×H×W containing spatial structure and time series, where T represents the length of the time series, C represents the number of polarization channels, and H×W represents the spatial image size of 64×64; (2b) Use ResNet50 as the basic feature extractor and design a feature extractor F based on the ResNet network architecture; (2c) Obtain the i-th pixel x in the time series image X through the feature extractor F i The pixel-level feature map z i : z i =F(x i ),i∈(1,n), Among them, z i ∈R D , D represents the updated temporal feature dimension; take i = 1, 2, ..., n and obtain the pixel-level feature map of all pixels in the temporal image X according to the above formula, and construct the temporal feature map Z∈R D×H×W .

3. The method according to claim 2, characterized in that: The feature extractor F in step (2b) includes a standard 7×7 convolution block, a maximum pooling layer and four residual blocks, wherein each residual block contains two 3×3 convolution operations, a batch normalization layer BN, a rectified linear unit ReLU, a standard 1×1 convolution operation and a residual connection represented by a batch normalization layer BN; the extractor extracts the features of the last three layers of the ResNet network, and uses the attention technology to extract the time dimension features and the space dimension features, and finally fuses the multi-scale information to obtain the feature map.

4. The method according to claim 1, characterized in that: The shape context similarity SCS in step (3b) is obtained according to the following formula: Among them, SCS (p e,k ,z i ) indicates p e,k and z i The shape similarity of the feature vectors, and Respectively represent p e,k and z i The mean of the eigenvectors, and Respectively represent p e,k and z i Variance of the eigenvector; j=1,2,...,t represents the jth time point, and t is the length of the time series.

5. The method according to claim 1, characterized in that: In step (3c), the difference between features and prototypes is identified by unsupervised clustering method, which is implemented as follows:

6. The method according to claim 1, characterized in that: The inter-class separability loss L in step (4a) cpe , intra-class separable loss L spe , clustering compactness loss L pe , as follows: in, represents the distance between feature z and the most dissimilar prototype among the fine-grained prototypes corresponding to its category; It represents the distance between the most similar prototype among the fine-grained prototypes corresponding to the feature z and its category; e' represents other categories outside the e category, and k' represents other distributions outside the kth distribution.

7. The method according to claim 1, characterized in that: The back-propagation optimization described in steps (4c) and (5c) both use the Adam optimizer and set the learning rate to 0.

001.

8. The method according to claim 1, characterized in that: The multilayer perceptron in step (5a) consists of three linear layers, uses a GELU activation function, and maintains the original feature dimension.

9. The method according to claim 1, characterized in that: The classifier and discriminator described in step (5a) both use two standard 3×3 convolution blocks, adopt the PRELU activation function, and then use a 1×1 convolution layer to post-process the mapped features and output the final mapping results and distribution label predictions.

10. The method according to claim 1, characterized in that: In step (5b), the gradient reversal layer technique is used to update the loss of the discriminator, and the formula is as follows: L adv =l ce (C e (F mlp ),e)+l ce (C k (R λ (F mlp )),k), Among them, F mlp Represents the mapping features of the multi-layer perceptron, l ce represents the standard cross entropy loss, C e represents the class classifier, C k represents the distribution classifier, R λ is a gradient reversal layer controlled by the hyperparameter λ.

Citation Information

Patent Citations

  • Large-range cross-phenological-area crop drawing method based on time sequence remote sensing image

    CN115439754A