Remote sensing image crop classification method based on self-supervised learning

Through self-supervised learning networks and agent task networks, unlabeled remote sensing image data is used for feature extraction and classification, which solves the problem of traditional methods' dependence on labeled data and realizes efficient and adaptable remote sensing image crop classification.

CN118537652BActive Publication Date: 2025-09-12NORTH CHINA INST OF AEROSPACE ENG
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202410740434.9
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2024-06-08
Publication Date
2025-09-12
Estimated Expiration
2044-06-08

AI Technical Summary

Technical Problem

Traditional remote sensing image crop classification methods require a large amount of labeled data, and it is difficult to obtain and label remote sensing image data, resulting in reduced classification accuracy and difficulty in adapting to the complexity and variability of remote sensing images.

Method used

A self-supervised learning method is adopted to extract and classify features from unlabeled remote sensing image data through a self-supervised learning network and an agent task network. This includes data preprocessing, construction of a coarse classification sample dataset, construction and training of a self-supervised learning network, feature extraction using the normalized vegetation index and principal component analysis, and feature clustering using a convolutional attention module and a clustering network.

Benefits of technology

It reduces the dependence on labeled data, improves the generalization and adaptability of the model, and enables efficient crop classification in different remote sensing image scenes, saving time and cost.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN118537652B_ABST
    Figure CN118537652B_ABST
Patent Text Reader

Abstract

This invention discloses a remote sensing image crop classification method based on self-supervised learning. This method extracts rule-compliant crop sample data through a fully automated sample selection method. This crop sample data is processed into a coarse training sample set. Representation learning is performed using proxy tasks combined with an attention mechanism. Similar images of each remote sensing image are determined based on feature similarity. The proxy features are then used as prior conditions for semantic clustering, and refined classification is performed using the maximized dot product after softmax as the loss function. This method performs unsupervised training on the image's own feature similarity learning method, completing crop classification through clustering, avoiding the tedious and expensive process of manually annotating data. Compared with traditional manual feature extraction methods, this method is more adaptable to the complexity and variability of remote sensing images and can better utilize unlabeled data for training, improving the accuracy of crop classification.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to a remote sensing image crop classification method, in particular to a remote sensing image crop classification method based on self-supervised learning, and belongs to the technical field of agricultural remote sensing. Background Art

[0002] Currently, remote sensing image crop classification refers to the task of assigning a semantic label of a predefined crop category to the input remote sensing image, which is an important research direction in the field of remote sensing image processing. Traditional image classification methods are mainly based on manually extracted features, such as local binary patterns (LBP) and gray-level co-occurrence matrix (GLCM). For different remote sensing images, the features they extract may vary greatly, resulting in a decrease in classification accuracy. This is because traditional methods require sufficient professional knowledge and domain experience for feature extraction and selection, especially for special classification targets such as crops. It is difficult to fully adapt to the complexity and changes of remote sensing images.

[0003] In recent years, the development of deep learning techniques has greatly advanced research in crop classification from remote sensing images. These methods can automatically learn features from raw data and improve classification accuracy through multi-layer neural networks. However, traditional deep learning methods require large amounts of labeled data, which is difficult to implement in the field of remote sensing imagery. Acquiring and labeling large amounts of remote sensing image data is difficult and expensive, and obtaining sufficient and representative samples is even more challenging. Therefore, a new method that can better utilize unlabeled data for training is needed to improve crop classification accuracy. Summary of the Invention

[0004] The technical problem to be solved by the present invention is to provide a remote sensing image crop classification method based on self-supervised learning.

[0005] In order to solve the above technical problems, the technical solution adopted by the present invention is:

[0006] A remote sensing image crop classification method based on self-supervised learning includes the following steps:

[0007] Step 1: Preprocessing of remote sensing image data: After preprocessing the remote sensing image data, a fused image is obtained. The normalized vegetation index (NDVI) of each pixel is calculated from the fused image to obtain a normalized vegetation index distribution map. The calculation method of the normalized vegetation index (NDVI) is as follows:

[0008]

[0009] In the formula, B4 represents the red band, and B5 represents the near-red band;

[0010] Step 2: Construction of coarse classification sample dataset: consists of the following specific steps:

[0011] Step 2-1: Use the OTSU segmentation algorithm to segment the normalized vegetation index distribution map into wheat areas and non-wheat areas;

[0012] Step 2-2: The B6 and B7 short red bands of the fused image are subjected to dimensionality reduction processing using the principal component analysis method (PCA) to obtain a short red band SWIR image retaining only the first principal component;

[0013] Step 2-3: Use the OTSU segmentation algorithm to segment the short red band SWIR image into corn areas and non-corn areas;

[0014] Step 2-4: Find the intersection of the non-wheat area in the NDVI distribution map and the non-corn area in the short red band SWIR image, and mark it as the corn area;

[0015] Step 2-5: Label the corresponding areas in the fused image using corn, soybean, and wheat areas. Areas with more than one category label will not be labeled. Areas with an area smaller than the preset area threshold will not be labeled.

[0016] Step 2-6: Cut the fused image according to the preset cutting size to obtain the segmented image, and mark the labels as wheat area, soybean area and corn area. Select the coarse sample training data according to the preset ratio of the segmented image of the corn area;

[0017] Step 3: Divide the training set and test set: adjust the segmented images in the coarse classification sample dataset to the preset resolution, normalize the pixel values ​​to between 0 and 1, and select the training set and test set according to the preset ratio and preset number;

[0018] Step 4: Create a similar dataset: Perform two RandAugment enhancements on the segmented images in the training set to obtain an expanded training set.

[0019] Step 5: Construct a self-supervised learning network: The self-supervised learning network consists of a cascaded proxy task network and a clustering network. The proxy task network is composed of a cascaded backbone network, an average pooling layer Avgpool, a fully connected layer Linear1, an activation function ReLU1, a fully connected layer Linear2, and a normalization layer Normalization.

[0020] The backbone network includes a DBR1 component cascaded in sequence, two 2DS-2CBAM components with the same structure, and the first to third combination components with the same structure; the first combination component is composed of a cascaded 3DS-3CBAM component and a 2DS-2CBAM component; the 2DS-2CBAM component is composed of a DBR2 component, a convolutional attention module CBAM1, a DB1 component, a convolutional attention module CBAM2, and an SR1 component cascaded in sequence, and the input end of the 2DS-2CBAM component is connected to the input end of the DBR2 component and the SR1 component respectively; the 3DS-3CBAM component includes a DBR3 component, a DB2 component, a DB3 component, a convolutional attention module CBAM3- The input of CBAM5, SR2 components, and 3DS-3CBAM components are connected to the input of SR2 components through DB2 components, convolutional attention module CBAM3, and the other input is connected to the input of SR2 components through DBR3 components, convolutional attention module CBAM4, DB3 components, and convolutional attention module CBAM5. The DBR1 component, DBR2 component, and DBR3 component have the same structure. The DBR1 component consists of a cascaded convolutional layer conv2d, a batch normalization BN layer, and an activation function ReLU2. The SR1 component and SR2 component have the same structure. The SR1 component consists of a cross-layer connection layer shortcut and an activation function ReLU3.

[0021] Step 6: Train the self-supervised learning network: Use the training set to train the self-supervised learning network;

[0022] Step 7: Use self-supervised learning network model for crop classification:

[0023] The remote sensing data to be classified is input into the self-supervised learning network to obtain the crop classification results.

[0024] Furthermore, the data preprocessing in step 1 includes radiometric calibration, atmospheric correction, orthorectification and radiometric correction.

[0025] Furthermore, in the data preprocessing step 1, the multispectral data of the remote sensing satellite image is subjected to radiometric calibration, atmospheric correction, and orthorectification, and then the panchromatic image is subjected to radiometric correction and orthorectification. Finally, the corrected multispectral data and the panchromatic image are resampled according to a preset resolution and fused to obtain a fused image.

[0026] Furthermore, in step 1, the multispectral data is resampled according to the resolution of the panchromatic image and then fused with the panchromatic image to obtain a fused image.

[0027] Furthermore, the DBR1 component parameters are: convolution kernel size of 3x 3, padding of 1, stride of 2, and dilation of 1.

[0028] Furthermore, the DB1 component parameters are: convolution kernel size 3x 3, padding 0, stride 2, and dilation 1.

[0029] The beneficial effects of adopting the above technical solution are:

[0030] (1) The present invention does not require a large amount of labeled data, and the model can learn by itself from unlabeled data, thus greatly reducing the workload of manual labeling.

[0031] (2) The present invention learns similarities from unlabeled data and can learn more features from it, thereby improving the generalization ability of the model.

[0032] (3) The present invention has strong adaptability and can be used in different remote sensing image scenes and crop classification tasks to expand the training model in different self-supervised tasks.

[0033] (4) The present invention does not require data labeling and manual intervention, which greatly saves time and cost, especially for the production of large-scale remote sensing image data. BRIEF DESCRIPTION OF THE DRAWINGS

[0034] Figure 1 is a flow chart of the present invention;

[0035] Figure 2 is a flow chart of the fully automatic sample selection of the present invention;

[0036] Figure 3 is a flowchart of the self-supervised task learning of the present invention;

[0037] Figure 4 This is a diagram of the proxy task network structure of Example 1 of the present invention;

[0038] Figure 5 is a structural diagram of a 2DS-2CBAM assembly according to Example 1 of the present invention;

[0039] Figure 6 3DS-3CBAM component structure diagram of Example 1 of the present invention;

[0040] Figure 7 is a structural diagram of a DBR component according to embodiment 1 of the present invention;

[0041] Figure 8 This is a structural diagram of the SR component of Example 1 of the present invention. DETAILED DESCRIPTION

[0042] Example 1:

[0043] A remote sensing image crop classification method based on self-supervised learning includes the following steps:

[0044] Step 1: Preprocessing of remote sensing image data: After preprocessing the remote sensing image data, a fused image is obtained. The normalized vegetation index (NDVI) of each pixel is calculated from the fused image to obtain a normalized vegetation index distribution map. The calculation method of the normalized vegetation index (NDVI) is as follows:

[0045]

[0046] In the formula, B4 represents the red band, and B5 represents the near-red band;

[0047] Data preprocessing includes radiometric calibration, atmospheric correction, orthorectification, and radiometric correction. The remote sensing image data used in this example is high-resolution Landset 8 satellite imagery. This example uses ENVI software to perform radiometric calibration, atmospheric correction, and orthorectification on the multispectral data of the Landset 8 remote sensing satellite imagery. Radiometric and orthorectification are then performed on the panchromatic imagery.

[0048] Radiometric calibration converts the digital quantization value DN of multispectral data and panchromatic images into physical quantities such as radiometric brightness and reflectivity. This embodiment uses the radiometric calibration tool Radiation Calibration in ENVI to read parameters from the metadata file to complete the calibration.

[0049] Atmospheric correction uses the complete ENVI fast atmospheric correction tool FLASSH to eliminate the radiation error in multispectral data and invert the true surface reflectance of the ground object.

[0050] Orthorectification uses ENVI's RPC correction to eliminate the image point displacement caused by projection errors caused by terrain undulations and errors such as the internal and external states of the sensor in multispectral data and panchromatic images, thereby obtaining orthophotos that can accurately and objectively represent the shape and spatial position of the objects.

[0051] The corrected multispectral data is resampled to generate a multispectral remote sensing image with the same resolution as the panchromatic image. ENVI's GS fusion method fuses the multispectral remote sensing image and the panchromatic image into a fused image. In this embodiment, the resolution is 15 meters.

[0052] Step 2: Construction of coarse classification sample dataset: consists of the following specific steps:

[0053] Step 2-1: Use the OTSU segmentation algorithm to segment the normalized vegetation index distribution map into wheat areas and non-wheat areas;

[0054] NDVI, short for Normalized Difference Vegetation Index, is an index used in remote sensing image analysis. NDVI values ​​can be used to distinguish wheat from soybeans. It reflects the growth and coverage of vegetation by comparing the reflectance of red and near-infrared light. Generally, wheat has a higher NDVI value due to its high chlorophyll content and denser leaves, while soybeans have a lower NDVI value due to their low chlorophyll content and sparser leaves, making it an effective way to distinguish wheat from soybeans.

[0055] Step 2-2: The B6 and B7 short red bands of the fused image are subjected to dimensionality reduction processing using the principal component analysis method PCA to obtain a short red band SWIR image with only the first principal component retained. The purpose of PCA processing of the B6 and B7 short red bands is to map high-dimensional data to a low-dimensional space. The first principal component retains the maximum data variance information and minimizes data redundancy and noise as much as possible. The first principal component of soybeans is higher than that of corn, and they can be distinguished.

[0056] Step 2-3: Use the OTSU segmentation algorithm to segment the short red band SWIR image into corn areas and non-corn areas;

[0057] Step 2-4: Mark the intersection of the non-wheat area in the wheat distribution map and the non-corn area in the short red band SWIR image as the corn area;

[0058] Step 2-5: Label the corresponding areas in the fused image using corn, soybean, and wheat areas. Areas with more than one category label will not be labeled. Areas with an area smaller than the preset area threshold will not be labeled.

[0059] In this example, a histogram-based statistical method is used to automatically determine the threshold ranges for crop extraction based on confidence intervals to obtain the first and second thresholds. The OTSU algorithm calculates the inter-class variance of the image based on the first and second thresholds, thereby separating the image into foreground and background. Through three segmentations, segmentation maps for wheat, soybean, and corn regions are obtained.

[0060] The NDVI threshold for wheat is 51 to 102, the NDVI threshold for soybean is 103 to 153, the short red band reflectance threshold for soybean is 102 to 128, and the short red band reflectance threshold for corn is 51 to 102. In this embodiment, the area threshold is a 10*10 pixel block.

[0061] Step 2-6: Chop the fused image into a pre-set size to obtain a segmented image. Label the image as wheat, soybean, and corn regions, and select coarse sample training data according to a pre-set ratio. In this embodiment, the pre-set size is 64*64, and the coarse sample training data is selected at a ratio of 1:1:1 for wheat, soybean, and corn regions.

[0062] Different label annotations are distinguished by file naming. In the file name 0_*, 1_*, 2_*, 0, 1, and 2 represent wheat, soybean, and corn respectively, and a coarse classification sample data set can be obtained.

[0063] The present invention utilizes threshold-based image segmentation to complete the segmentation task, ignores image outliers, and automatically selects appropriate area thresholds of various types to prevent interference in the training part.

[0064] The present invention removes the training samples of excess crops in the image and sets the samples of each category evenly, so as to obtain pure data samples.

[0065] Step 3: Divide the coarse classification sample dataset into training and test sets: Save the files in the coarse classification sample dataset as jpg images and resize them to 32*32. Normalize the pixel values ​​to between 0 and 1 and store them in an 8:2 ratio, with 80% for the training set and 20% for the test set. Ensure that the number of samples in each category is no less than 2000. The training set is used to train the self-supervised learning network, and the prediction set is used to test the performance of the self-supervised learning network.

[0066] This embodiment processes the above data into a collection text file for the training set, a collection text file for the test set containing category information, and a statistical information file for storing the total number of samples, sample type names, and the number of pixels contained, in order to perform category classification on the subsequent collection test set.

[0067] Step 4: Create a similar dataset: The RandAugment algorithm is applied twice to the files in the coarse classification sample dataset to generate an augmented dataset. This randomly increases sample diversity and helps better learn similar features for the subsequent proxy task. The two augmented results obtained from the two transformations serve as the data input to the proxy task learning network. The RandAugment algorithm applies n augmentation operations to each image, where n is an integer greater than or equal to 1, and the parameters of each operation are randomly selected. Common augmentation operations include random horizontal flips, rotations, and changes to image brightness and contrast.

[0068] Step 5: Construct a self-supervised learning network: The self-supervised learning network is composed of the agent task network and the clustering network. The agent task network is composed of the cascaded backbone network, average pooling layer Avgpool, fully connected layer Linear1, activation function ReLU1, fully connected layer Linear2 and normalization layer Normalization.

[0069] The backbone network includes a DBR1 component cascaded in sequence, two 2DS-2CBAM components with the same structure, and the first to third combination components with the same structure; the first combination component is composed of a cascaded 3DS-3CBAM component and a 2DS-2CBAM component; the 2DS-2CBAM component is composed of a DBR2 component, a convolutional attention module CBAM1, a DB1 component, a convolutional attention module CBAM2, and an SR1 component cascaded in sequence, and the input end of the 2DS-2CBAM component is connected to the input end of the DBR2 component and the SR1 component respectively; the 3DS-3CBAM component includes a DBR3 component, a DB2 component, a DB3 component, a convolutional attention module CBAM3 -CBAM5, SR2 components, the input of the 3DS-3CBAM component is connected to the input of the SR2 component through the DB2 component and the convolutional attention module CBAM3 in sequence, and the other path is connected to the input of the SR2 component through the DBR3 component, the convolutional attention module CBAM4, the DB3 component, and the convolutional attention module CBAM5 in sequence; the DBR1 component, the DBR2 component and the DBR3 component have the same structure, and the DBR1 component consists of a cascaded convolution layer conv2d, a batch normalization BN layer and an activation function ReLU2; the SR1 component and the SR2 component have the same structure, and the SR1 component consists of a cross-layer connection shortcut and an activation function ReLU3.

[0070] In this embodiment, five convolutional attention modules (CBAMs) are inserted into the deep convolutional neural network Restnet-18 network as the backbone network for feature extraction. The network is normalized to the L2 norm, that is, the Euclidean distance. The entire network learns by narrowing the difference between two eigenvalues ​​to cope with the complexity of remote sensing images themselves and the problem of small and inconspicuous features. The number of nearest neighbors is then determined through an instance discrimination task based on noise contrast estimation.

[0071] The Convolutional Attention Module (CBAM) is an attention mechanism for computer vision tasks, designed to improve the performance of convolutional neural networks (CNNs) in image processing. It adaptively captures important information from different channels using channel attention, mapping each channel to a smaller vector, and then mapping this vector back to a vector with the original number of channels. These vectors represent the importance of each channel in a given feature map. The vectors are weighted to obtain crop features.

[0072] The Euclidean distance measurement method calculates each image, including minimizing the feature distance between enhanced images, that is, the nearest neighbor image corresponding to each image, clustering the neighbor images together, and separating images of different categories.

[0073] After two enhancements, images 1 and 2 are obtained and fed into the backbone network at their original 3*32*32 size. A DBR component and two 2DS-2CBAM components are used to extract 64*32*32 feature channel feature maps from the original images. Each channel corresponds to some abstract features in the image. The feature channel feature map is then operated three times through three identically connected combination components, becoming sizes of 128*16*16, 256*8*8 and 512*4*4 respectively. The convolution process parameters of the DBR component and DB component in the 2DS-2CBAM component are as follows: the convolution kernel size is 3x 3, padding is 1, stride is 1, and dilation is 1. The parameters of the DB component process are as follows: the convolution kernel size is 1*1, padding is 0, stride is 2, and dilation is 1. The parameters of the DBR component are as follows: the convolution kernel size is 3*3, padding is 1, stride is 2, and dilation is 1. The parameters of the DB component convolution process are as follows: the convolution kernel size is 3x 3, padding is 1, stride is 1, and dilation is 1.

[0074] In the DBR component, the Conv2d layer is used to perform convolution operations on the input features, the BN layer is used to normalize the output features, and the ReLU activation function layer is used to increase the nonlinear expression ability of the network. The function of the DBR module is to convert the input image into a high-dimensional feature representation, providing richer and more effective features for the network's subsequent classification or detection tasks.

[0075] The DBR component is similar to the DB component, except that the DB component uses the Relu activation function, while the DBR component uses a residual connection. A residual connection directly adds the input to the output before performing a Relu activation. This allows for deeper networks while ensuring the stability of gradient propagation and mitigating the vanishing gradient problem, thereby improving model performance and training speed.

[0076] The cross-layer connection layer shortcut in the SR component connects the input directly to the output, avoiding the vanishing or exploding gradient problems caused by excessively deep network layers. This cross-layer connection allows information to flow more smoothly within the neural network, improving network performance and effectiveness. ReLU, or rectified linear unit, is a commonly used activation function. It converts negative numbers to 0 while leaving positive numbers unchanged, thereby enhancing the nonlinear fitting capabilities of the neural network. It also prevents the vanishing gradient problem and improves network training speed and performance. The SR component, through the combination of cross-layer connections and the ReLU activation function, increases the depth and width of the neural network, improving its expressiveness and performance.

[0077] Use the training set to train the self-supervised learning network, and use the prediction set to test the performance of the self-supervised learning network.

[0078] In the feature extraction process of the above network, a channel attention mechanism (CBAM) is added after the convolution layer of each DBR component and DB component to adaptively focus on the feature information extraction of farmland and reduce the interference of other information. Because the remote sensing farmland information itself has obvious spatial characteristics, it can be effectively combined with the spatial attention relationship of CBAM to learn the spatial position relationship.

[0079] The CBAM attention mechanism is a key mechanism that enhances the expressiveness and performance of neural networks. It consists of two main components: channel attention and spatial attention. Channel attention is used to learn relationships between channels, while spatial attention is used to learn relationships between spatial locations. Specifically, channel attention learns the importance weight of each channel by performing global pooling and full connectivity operations on each channel; spatial attention learns relationships between different locations by performing global pooling and full connectivity operations on the feature maps of each channel.

[0080] The 512*4*4 feature map is downsampled and features extracted to 512 through average pooling. Average pooling can reduce the size of the feature map and smooth the input features, making the features more stable and smooth. The result is then passed through a fully connected layer to output features of dimension 512. After a combination of a ReLU layer and a fully connected layer, the 512 dimension is reduced to 128, thereby refining more valuable features and reducing the computational complexity and parameter count of the model, thereby improving the model's generalization ability and training efficiency.

[0081] The feature maps obtained by the network are used to calculate the similarity of the two feature maps through the Euclidean distance. For the n-dimensional feature vectors A and B, their Euclidean distance D(A,B) is expressed as:

[0082]

[0083] That is, through the above formula, we can get the distance between two feature vectors in space. The smaller the Euclidean distance, the higher the similarity between the two feature maps. Conversely, the smaller the Euclidean distance, the lower the similarity between the two feature maps. We then perform normalization to ensure that the distance value is within a certain range. This is done by dividing the distance value by the dimension of the feature vector, i.e.:

[0084]

[0085] The distance value D′(A, B) obtained in this way ranges from 0 to 1, which is convenient for subsequent calculations and comparisons.

[0086] The normalized results above are constrained using a proxy loss function (Noise Contrastive Estimation (NCE)) to construct the model for this module. NCE is a method for training contrastive loss. Contrastive loss aims to make samples of the same category closer in feature space, while samples of different categories are more dispersed. Specifically, given a sample and a noise sample, their feature vectors are used as input to the model to predict whether they belong to the same category.

[0087] Based on the preliminary similarity classification results obtained above, the K-means method was combined with the clustering process to produce three clusters. The core idea of ​​this algorithm is to assign a data point to the cluster center closest to it. The cluster center is obtained by calculating the average of all data points within the cluster. Each cluster is then updated with the new cluster center until the change in cluster center is less than a certain threshold or a predetermined number of iterations is reached.

[0088] After using k-means to cluster and get 3 cluster centers, we can assign each sample x i Assign to the nearest cluster center j and get a category label y for each sample i , where y i ∈{1, 2, 3}. At this point, we can incorporate the cross entropy loss function to constrain the learning of the classification model by calculating the cross entropy between the negative logarithmic softmax value of the distance between each sample and the cluster center and the one-hot encoded category label.

[0089] The negative logarithmic softmax function is a commonly used probability distribution function. Its function is to convert a vector into a probability distribution vector, in which each component is non-negative and the sum is 1.

[0090] Specifically, for each sample x i , and its distance from each cluster center can be expressed as d ij =||xi-cj||2, where c jDenotes the jth cluster center. Then we can define the negative logarithmic softmax function:

[0091]

[0092] in Represents the probability estimate that sample xi belongs to category j, and we use the one-hot encoded category label y i To define the weighted cross entropy loss function:

[0093]

[0094] Where N represents the number of samples, w i is the weight of category j, because the samples are uniform, the weight y ij Represents sample x i Whether it belongs to category j, cross entropy is a loss function used to measure the difference between the model prediction result and the actual result. The smaller the cross entropy value, the closer the model prediction result is to the actual result.

[0095] The virtual labels are made based on the prediction results above, and the real labels and virtual labels are mixed together in a certain proportion. They are input into the cross entropy loss function as new labels to calculate the loss, increase the amount of data, optimize the classification results, and improve the generalization ability of the model.

[0096] The accuracy of different crops is calculated through the confusion matrix, and the label content of the test set itself is combined to assign wheat, soybean or corn category names to different clustering results, and they are stored separately to complete the final classification task of naming.

[0097] The confusion matrix is ​​a table used to evaluate the performance of a classification model, in which columns represent predictions and rows represent true labels. Each element in the matrix represents the number of samples whose predictions are for the category corresponding to the column when the true label is the category corresponding to the row.

[0098] Step 6: Train the self-supervised learning network: Use the training set to train the self-supervised learning network;

[0099] In the hyperparameter setting, the agent task selected a batch size of 512 to perform the representation learning process for 500 epochs, where the learning rate was 0.4 and gradually decreased during training. The temperature parameter was 0.1, where the temperature hyperparameter is used to control the scale of the normalized activation in the contrastive loss function. The weight decay was 10-4, combined with L2 regularization to complete the network task.

[0100] To speed up training, we transfer the weights obtained from the clustering task to the initial clustering step, which we perform for 100 epochs using a batch size of 128. The loss function weight is set to λ = 5. The higher weight avoids grouping samples too early during training.

[0101] After adding the virtual labels, we trained for an additional 200 epochs using a self-labeling process with a threshold of 0.99. A weighted cross-entropy loss was used to compensate for the imbalance in confidence samples across clusters. The learning rate was set to 10⁻¹⁴ and the weight decay was set to 10⁻¹⁴. Throughout the self-supervised learning process, the network weights were updated using the Adam algorithm, which offers advantages such as faster convergence and better generalization performance, to complete the model's final optimization for the classification task.

[0102] Step 7: Use self-supervised learning network model for crop classification:

[0103] The remote sensing data to be classified is input into the self-supervised learning network to obtain the crop classification results.

[0104] The present invention extracts crop sample data that meets the rules through a fully automated sample selection method, combines data processing to produce a training sample set, uses proxy tasks for representation learning, and mines the nearest neighbors of each image based on feature similarity, that is, similar images of each remote sensing image. The attention mechanism is combined to better focus on the characteristic differences of the crops themselves, and then uses the proxy features as the prior conditions for semantic clustering. The maximized dot product after softmax is used as the loss function to classify each image and its mined neighbors into a group of categories. Finally, virtual labels are created to remove noise problems, optimize and reclassify, and then combine the test set and information file in the dataset to complete the final category name assignment.

[0105] The present invention can explore similarities from the image itself through self-supervised learning, reducing a series of cost issues. By combining the semantic clustering analysis method of the proxy task with the attention mechanism module, it can effectively focus on the characteristics of the crop itself and can be better applied to crop classification. Secondly, it can also handle the complexity and variability of the remote sensing image itself well.

Claims

1. A remote sensing image crop classification method based on self-supervised learning, characterized by: The following steps are involved: Step 1: Preprocessing of remote sensing image data: After preprocessing the remote sensing image data, a fused image is obtained, and the normalized vegetation index (NDVI) of each pixel is calculated from the fused image to obtain a normalized vegetation index distribution map; The calculation method of normalized vegetation index NDVI is: In the formula, B4 represents the red band, and B5 represents the near-red band; Step 2: Construction of coarse classification sample dataset: consists of the following specific steps: Step 2-1: Use the OTSU segmentation algorithm to segment the normalized vegetation index distribution map into wheat areas and non-wheat areas; Step 2-2: The B6 and B7 short red bands of the fused image are subjected to dimensionality reduction processing using the principal component analysis method (PCA) to obtain a short red band SWIR image retaining only the first principal component; Step 2-3: Use the OTSU segmentation algorithm to segment the short red band SWIR image into corn areas and non-corn areas; Step 2-4: Find the intersection of the non-wheat area in the NDVI distribution map and the non-corn area in the short red band SWIR image, and mark it as the soybean area; Step 2-5: Label the corresponding areas in the fused image using corn, soybean, and wheat areas. Areas with more than one category label will not be labeled. Areas with an area smaller than the preset area threshold will not be labeled. Step 2-6: Cut the fused image according to the preset cutting size to obtain the segmented image, and mark the labels as wheat area, soybean area and corn area. Select the coarse sample training data according to the preset ratio of the segmented image of the corn area; Step 3: Divide the training set and test set: adjust the segmented images in the coarse classification sample dataset to the preset resolution, normalize the pixel values ​​to between 0 and 1, and select the training set and test set according to the preset ratio and preset number; Step 4: Create a similar dataset: Perform two RandAugment enhancements on the segmented images in the training set to obtain an expanded training set. Step 5: Construct a self-supervised learning network: The self-supervised learning network is composed of the agent task network and the clustering network. The agent task network is composed of the cascaded backbone network, average pooling layer Avgpool, fully connected layer Linear1, activation function ReLU1, fully connected layer Linear2 and normalization layer Normalization. The backbone network includes a DBR1 component cascaded in sequence, two 2DS-2CBAM components with the same structure, and the first to third combination components with the same structure; the first combination component is composed of a cascaded 3DS-3CBAM component and a 2DS-2CBAM component; the 2DS-2CBAM component is composed of a DBR2 component, a convolutional attention module CBAM1, a DB1 component, a convolutional attention module CBAM2, and an SR1 component cascaded in sequence, and the input end of the 2DS-2CBAM component is connected to the input end of the DBR2 component and the SR1 component respectively; the 3DS-3CBAM component includes a DBR3 component, a DB2 component, a DB3 component, a convolutional attention module CBAM3 -CBAM5, SR2 components, the input of the 3DS-3CBAM component is connected to the input of the SR2 component through the DB2 component and the convolutional attention module CBAM3 in sequence, and the other path is connected to the input of the SR2 component through the DBR3 component, the convolutional attention module CBAM4, the DB3 component, and the convolutional attention module CBAM5 in sequence; the DBR1 component, DBR2 component and DBR3 component have the same structure. The DBR1 component consists of a cascaded convolutional layer conv2d, a batch normalization BN layer and an activation function ReLU2; the SR1 component and SR2 component have the same structure. The SR1 component consists of a cross-layer connection shortcut and an activation function ReLU3; Step 6: Train the self-supervised learning network: Use the training set to train the self-supervised learning network; Step 7: Use self-supervised learning network model for crop classification: The remote sensing data to be classified is input into the self-supervised learning network to obtain the crop classification results.

2. The method for crop classification based on remote sensing images based on self-supervised learning according to claim 1, characterized in that: The data preprocessing in step 1 includes radiometric calibration, atmospheric correction, orthorectification and radiometric correction.

3. The method for crop classification based on remote sensing images based on self-supervised learning according to claim 1, characterized in that: In the step 1, the data preprocessing includes performing radiometric calibration, atmospheric correction, and orthorectification on the multispectral data of the remote sensing satellite image, and then performing radiometric correction and orthorectification on the panchromatic image. Finally, the corrected multispectral data and the panchromatic image are resampled according to a preset resolution and fused to obtain a fused image.

4. The method for crop classification based on remote sensing images based on self-supervised learning according to claim 1, characterized in that: In step 1, the multispectral data is resampled according to the resolution of the panchromatic image and then fused with the panchromatic image to obtain a fused image.

5. The method for crop classification based on remote sensing images based on self-supervised learning according to claim 1, characterized in that: The parameters of the DBR1 component in step 5 are: convolution kernel size 3 x 3, padding 1, stride 2, and dilation 1.

6. The method for crop classification based on remote sensing images based on self-supervised learning according to claim 1, characterized in that: The DB1 component parameters in step 5 are: convolution kernel size 3 x 3, padding 0, stride 2, and dilation 1.

Citation Information

Patent Citations

  • High-resolution remote sensing image crop classification method based on deep learning

    CN110287869A

  • Remote sensing image semantic segmentation method based on supervised long-range correlation

    CN115953577A