Multi-class Feature Recognition Method for Remote Sensing Images Based on Self-Supervised Learning
By using a combination of self-supervised learning and gradient-enhanced decision tree in the multi-category land object recognition in remote sensing images, deep features are extracted and classified, and the existing methods rely on labeled data, low computing efficiency, and insufficient category distinction capabilities are solved, and high-precision and low-cost multi-category land object recognition is achieved.
Patent Information
- Application Number
- CN202510272076.8
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-03-10
- Publication Date
- 2025-06-20
- Estimated Expiration
- 2045-03-10
AI Technical Summary
The existing multi-category land object recognition methods for remote sensing images rely on a large amount of manual labeling data, which is expensive and difficult to adapt to changes in different land object categories and environments. It is difficult to efficiently model the spatial and spectral characteristics of complex remote sensing data, especially in small sample environments, and the classification performance deteriorates when dealing with multi-category imbalance problems.
The self-supervised learning method is used to pre-train features of remote sensing image data, and deep features are extracted through multi-view comparison tasks, spatial consistency tasks and spectral domain transformation tasks. Combined with the gradient enhancement decision tree classification module and feature distillation module, high-quality feature extraction and classification of multiple categories of land objects are achieved.
It significantly improves the classification accuracy and category distinction ability of complex geographic categories, reduces dependence on manual annotation data, improves the adaptability and computing efficiency of the model, and maintains a high resolution ability under different geographic categories, lighting conditions and space-time scales.
Smart Images

Figure CN119810672B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of remote sensing images, and particularly to a method for multi-class land object recognition in remote sensing images based on self-supervised learning. Background Art
[0002] With the rapid development of remote sensing technology and artificial intelligence, the automated processing of remote sensing images plays an increasingly important role in the fields of land and resources management, ecological environment monitoring, agricultural mapping, and urban planning. Among them, multi-class land object recognition is one of the important tasks of remote sensing image analysis, aiming to automatically extract information on different types of land objects from high-resolution, multi-spectral, or hyperspectral images to achieve accurate classification of land objects.
[0003] Currently, the multi-class land object recognition of remote sensing images mainly relies on traditional supervised learning methods, that is, by manually annotating a large amount of remote sensing image data to train deep learning models or machine learning classifiers. Existing methods have improved the automation level of land object classification to a certain extent, but there are still the following main problems:
[0004] First, the high dependence on labeled data limits the adaptability of the model. Remote sensing images usually cover vast geographical areas and various land object types, and there are significant differences in the spectral characteristics of land objects in different regions. Therefore, in order to train a classification model with generalization ability, large-scale and diverse high-quality labeled data are required. However, the manual annotation cost of remote sensing image data is extremely high, and the precise annotation of multi-class complex land objects requires the participation of domain experts, which leads to high data acquisition costs and difficulty in covering all possible land object categories and environmental changes. In addition, manually annotated data is easily affected by subjective factors and has certain noise, further reducing the reliability of the classification model.
[0005] Second, existing classification algorithms are difficult to efficiently model the spatial and spectral characteristics of complex remote sensing data. Remote sensing images have multi-dimensional characteristics, including spatial distribution, spectral information, and texture features. Traditional classification methods, such as random forests and support vector machines, often have difficulty fully modeling the non-linear relationships between these complex features. Although deep learning methods have been widely applied to remote sensing image analysis in recent years, existing methods usually require a large amount of training data and are prone to overfitting in small-sample environments, resulting in poor adaptability of the model on different data sets.
[0006] In addition, the current classification methods still face great challenges in dealing with the problem of multi-class imbalance. The distribution of ground object categories in remote sensing images is often extremely unbalanced. For example, the categories of buildings and roads in urban areas may occupy a large proportion, while the data volume of vegetation and wetland categories is small. Traditional classification methods usually tend to focus on high-frequency categories, resulting in low recognition accuracy for frequency categories below the threshold. Although existing methods such as class resampling and cost-sensitive learning can partially alleviate this problem, there are still problems of decreased classification performance in the case of highly skewed data distributions, and the misclassification rate of the classifier is relatively high when the boundaries of ground object categories are blurred and the spectral characteristics between categories are similar.
[0007] Therefore, there is an urgent need for a new method to break through the dependence on labeled data in traditional supervised learning and combine an efficient classification mechanism to improve the recognition accuracy and adaptability of multi-class ground objects in remote sensing images. Summary of the Invention
[0008] An object of the present invention is to propose a multi-class ground object recognition method for remote sensing images based on self-supervised learning, which significantly improves the classification accuracy and category discrimination ability of complex ground object categories.
[0009] A multi-class ground object recognition method for remote sensing images based on self-supervised learning according to an embodiment of the present invention includes the following steps:
[0010] S1. Obtain remote sensing image data and preprocess the remote sensing image data, including radiometric correction, geometric correction, noise removal, cloud occlusion area filling, and image enhancement processing, to generate standardized remote sensing image data;
[0011] S2. Perform feature pre-training on the standardized remote sensing image data based on self-supervised learning. The self-supervised learning module automatically extracts discriminative deep feature representations from the standardized remote sensing image data by constructing multi-view contrast tasks, spatial consistency tasks, and spectral domain transformation tasks, and outputs deep feature representations;
[0012] S3. Input the deep feature representations into the gradient-boosted decision tree classification module. The gradient-boosted decision tree classification module consists of multiple weak classifiers, and performs preliminary multi-class ground object classification on the deep feature representations through iterative optimization, and outputs a preliminary classification result;
[0013] S4. Calculate the classification error according to the preliminary classification result and dynamically adjust the weights of the weak classifiers based on the gradient feedback mechanism, optimize the decision boundaries of each weak classifier, and update the gradient-boosted decision tree classification module;
[0014] S5. Convert the deep feature representation into a low-dimensional feature representation via a feature distillation module, and fuse it with the intermediate classification features after weight adjustment by the gradient-boosted decision tree classification module to generate a fused optimized feature representation. The optimized feature representation is further input into the gradient-boosted decision tree classification module for fine classification training, and an optimized classification result is output;
[0015] S6. Implement spatial consistency optimization, classification boundary refinement, and class reassignment on the optimized classification result to generate a final multi-class object recognition result for the remote sensing image.
[0016] Optionally, S1 includes the following steps:
[0017] S11. Collect remote sensing image data, where the resolution of the remote sensing image data is R, the number of bands is B, and the image size is ;
[0018] S12. Perform radiometric correction, geometric correction, and noise removal on the remote sensing image data;
[0019] S13. Fill in the cloud-covered areas of the remote sensing image data after radiometric correction, geometric correction, and noise removal. Based on the principle of local similarity, estimate the pixel values of the cloud-covered areas using the surrounding pixel information, calculate the nearest neighbor unoccluded pixel set, and fill it using the weighted interpolation method;
[0020] S14. Optimize the contrast of the remote sensing image data using the adaptive histogram equalization method and output the standardized remote sensing image data D:
[0021] ;
[0022] where is the Nth standardized remote sensing image, and N is the total number of remote sensing image data.
[0023] Optionally, S2 includes the following steps:
[0024] S21. Use the self-supervised learning method to perform feature pre-training on the standardized remote sensing image data D and construct a cross-scale spatial-spectral fusion self-supervised feature extraction model F:
[0025] ;
[0026] where is the feature representation of the standardized remote sensing image after feature extraction, , d is the feature dimension;
[0027] S22. Construct a cross-scale contrast learning mechanism, perform multi-scale decomposition on the standardized remote sensing image data, define the scale transformation set T, and construct multi-scale enhanced samples and , calculate their feature vectors ;
[0028] Define the scale contrast loss:
[0029] ;
[0030] Among them, is the cosine similarity of the feature vectors, is the temperature parameter, and j represents the j-th scale transformation sample in the scale contrast learning mechanism;
[0031] S23. Design a ground object-background segmentation constraint, construct a ground object mask , define the ground object pixel set and the background pixel set , and calculate the mean feature of the ground object and the background:
[0032]
[0033] ;
[0034] Among them, M is the number of ground object pixels, K is the number of background pixels, and represent the representation vectors of the ground object pixel or the background pixel in the deep feature space;
[0035] Define the ground object discrimination loss to make the features of the ground object and the background distinguishable:
[0036] ;
[0037] S24. Adopt a spectral-spatial fusion enhancement mechanism to construct spectral perturbation samples:
[0038] ;
[0039] Among them, is Gaussian noise, is the perturbation coefficient, is the remote sensing image after spectral perturbation, and calculate the spectral invariance loss:
[0040] ;
[0041] Among them, is the remote sensing image after spectral perturbation The features processed by the cross-scale spatio-spectral fusion self-supervised feature extraction model, are for remote sensing images The features processed by the cross-scale spatio-spectral fusion self-supervised feature extraction model;
[0042] And by fusing spatial information, a local consistency constraint is defined:
[0043] ;
[0044] Wherein, is the m-th pixel of the normalized remote sensing image and is the pixel The surrounding neighborhood window contains pixel points with similar spatial positions, represents the feature representation of the pixel point in the remote sensing image after being processed by the cross-scale spatio-spectral fusion self-supervised feature extraction model, represents a certain pixel point in the normalized remote sensing image , is the number of pixel points within the neighborhood window;
[0045] S25. Calculate the final improved self-supervised learning loss, combining cross-scale contrast learning, object-background segmentation constraint, and spectral-spatial fusion mechanism:
[0046] ;
[0047] Wherein, is the weight parameter;
[0048] And output the set of deep feature representations extracted by self-supervised learning:
[0049] ;
[0050] Wherein, is the deep feature representation of the N-th remote sensing image.
[0051] Optionally, the S3 includes the following steps:
[0052] S31. Use the gradient boosting decision tree classification module to perform preliminary classification on the deep feature representation, construct the gradient boosting decision tree classification module G, and define the mapping:
[0053] ;
[0054] Wherein, is the predicted class label, is the deep feature representation of the i-th remote sensing image, and the value range is , C is the total number of ground object categories;
[0055] S32. Use a decision tree model based on gradient boosting for classification training and construct a loss function:
[0056] ;
[0057] Among them, is the true class label, is the true probability of the c-th class, is the class probability predicted by the gradient-boosted decision tree classification module;
[0058] S33. Adopt a weighted gradient update strategy to calculate the negative gradient in the t-th iteration:
[0059] ;
[0060] Based on the negative gradient construct the t-th decision tree , and update the classification model:
[0061] ;
[0062] Among them, is the learning rate, which controls the contribution of each decision tree;
[0063] S34. Use a deep feature importance calculation mechanism to evaluate the contributions of different feature dimensions, and define the feature importance score:
[0064] ;
[0065] Among them, is the j-th feature, T is the total number of decision trees, is the number of nodes of the t-th decision tree, is the loss reduction brought by node splitting. Sort the feature importance scores and select the most important feature dimensions;
[0066] S35. Adopt a class balance adjustment strategy to calculate the balance coefficient of the class distribution:
[0067] ;
[0068] Among them, is the number of samples of the c-th class. Adjust the loss function based on the balance coefficient :
[0069] ;
[0070] Iteratively optimize the gradient-boosted decision tree classification module until the loss function converges or reaches the preset number of iteration rounds , output the preliminary multi-class land cover classification results:
[0071] ;
[0072] Among them, is the set of preliminary classification labels for all remote sensing images, is the predicted class label for the Nth remote sensing image.
[0073] Optionally, the S4 includes the following steps:
[0074] S41. By taking the gradient of the loss function with respect to the preliminary classification results, obtain the gradient feedback value:
[0075] ;
[0076] S42. For the multiple weak classifiers of the gradient-boosted decision tree classification module, let the weight of the t-th weak classifier be , where , T is the total number of weak classifiers. According to the gradient feedback value of each sample i, calculate the weight adjustment amount of this weak classifier, and define the adjustment amount as:
[0077] ;
[0078] Among them, represents the weight adjustment of the t-th weak classifier under the gradient of the i-th sample;
[0079] S43. Implement dynamic weight update for each weak classifier:
[0080] ;
[0081] Among them, is the updated weight of the t-th weak classifier, is the learning rate parameter, controlling the update step size;
[0082] S44. Use the updated weight to re-optimize the decision boundaries of each weak classifier, and combine the deep feature representation to perform overall fine-tuning or retraining on the gradient-boosted decision tree classification module to construct the updated classification module G';
[0083] S45. Output the classification results obtained by predicting the deep feature representation Z using the updated gradient-boosted decision tree classification module G':
[0084] ;
[0085] Among them, is the predicted class probability vector after dynamic adjustment of the Nth remote sensing image.
[0086] Optionally, S5 includes the following steps:
[0087] S51. Use a feature distillation module to perform dimensionality reduction transformation on the deep feature representation. The feature distillation module optimizes redundant features while maintaining the original feature space information by constructing a lightweight representation learning network, and generates a low-dimensional feature representation after dimensionality reduction.
[0088] S52. Use an information entropy optimization strategy to screen the low-dimensional feature representation after dimensionality reduction, calculate the information contribution degree of each feature dimension based on the class discrimination ability, retain the most discriminative features, and eliminate low-correlation features.
[0089] S53. Use a feature fusion strategy to fuse the low-dimensional feature representation output by the feature distillation module with the intermediate classification features after weight adjustment of the optimized gradient-enhanced decision tree classification module. In the fusion process, dynamic weighting parameters are set for different feature sources, so that the deep features extracted by self-supervised learning and the highly discriminative features generated by the gradient-enhanced decision tree classification module work together to form an optimized feature representation after fusion.
[0090] S54. Use the fused feature representation to retrain the gradient-enhanced decision tree classification module. In the training process, introduce a hierarchical learning strategy, first train the easy-to-classify samples, and gradually transition to the boundary-fuzzy and easily confused categories.
[0091] S55. Use a fine-grained classification training strategy to perform local iterative optimization on the fused feature representation. Dynamically adjust the class discrimination boundary during the training process, increase the decision tree depth for the confused categories above the threshold, and adopt a sample balancing strategy for the frequency categories below the threshold.
[0092] S56. Output the optimized multi-class ground object recognition classification result. The optimized multi-class ground object recognition classification result is generated based on the fused feature representation and has undergone multi-level optimization and adjustment.
[0093] The beneficial effects of the present invention are:
[0094] (1) The present invention adopts a self-supervised learning strategy for cross-scale spatial-spectral fusion, combines multi-view contrast learning, spatial consistency learning, and spectral transformation enhancement, and realizes high-quality feature extraction of multi-category ground objects in remote sensing images without a large amount of manually labeled data. By introducing ground object-background segmentation constraints and spectral invariance regularization, it can extract robust representative features under different ground object categories, lighting conditions, and spatio-temporal scales, and can fully utilize the inherent correlation of spectral information when processing hyperspectral and multispectral data, enabling the ground object classification to still maintain a high resolution ability in the face of spectral mixing effects.
[0095] (2) During the remote sensing image classification process, the present invention combines a feature distillation and gradient-boosted decision tree classification module. The feature distillation module reduces the dimension of high-dimensional deep features, removes redundant information, extracts the most discriminative low-dimensional feature representations, and fuses them with the intermediate features of the gradient-boosted decision tree classifier, thereby improving the classification accuracy and reducing the computational complexity. By constructing a lightweight feature distillation network, the input dimension of the features is reduced by 35% - 50%, but the information retention rate can still reach more than 95%. In addition, the feature fusion mechanism enables the GBDT to utilize the discriminative features extracted by deep learning while maintaining the adaptability to small sample data.
[0096] (3) During the training process of the gradient-boosted decision tree classification module, the present invention introduces a multi-stage dynamic weight adjustment strategy based on gradient feedback, conducts gradient analysis on the classification errors of different categories, and adjusts the decision weights of weak classifiers in each iteration. By constructing an adaptive class balance loss function, the classification accuracy of frequency classes below the threshold is improved, and the local feature reweighting mechanism is used to enable the classifier to dynamically adjust the decision boundary for confused classes above the threshold, improving the discrimination ability of the classification model in complex ground object environments. Description of the Drawings
[0097] The drawings are used to provide a further understanding of the present invention and constitute a part of the specification. They are used together with the embodiments of the present invention to explain the present invention and do not constitute a limitation to the present invention. In the drawings:
[0098] Figure 1 is a flowchart of a method for identifying multi-category ground objects in remote sensing images based on self-supervised learning proposed by the present invention. Detailed Embodiments
[0099] Now, the present invention will be further described in detail with reference to the drawings. These drawings are all simplified schematic diagrams, only showing the basic structure of the present invention in a schematic manner, so they only show the components related to the present invention.
[0100] Refer to Figure 1, A method for multi-class land object recognition in remote sensing images based on self-supervised learning, comprising the following steps:
[0101] S1. Obtain remote sensing image data and preprocess the remote sensing image data, including radiometric correction, geometric correction, noise removal, filling of cloud-covered areas, and image enhancement processing, to generate standardized remote sensing image data;
[0102] S2. Perform feature pre-training on the standardized remote sensing image data based on self-supervised learning. The self-supervised learning module automatically extracts discriminative deep feature representations from the standardized remote sensing image data by constructing multi-view contrast tasks, spatial consistency tasks, and spectral domain transformation tasks, and outputs the deep feature representations;
[0103] S3. Input the deep feature representations into the gradient-boosted decision tree classification module. The gradient-boosted decision tree classification module consists of multiple weak classifiers, and performs preliminary multi-class land object classification on the deep feature representations through iterative optimization, and outputs the preliminary classification results;
[0104] S4. Calculate the classification error according to the preliminary classification results and dynamically adjust the weights of the weak classifiers based on the gradient feedback mechanism, optimize the decision boundaries of each weak classifier, and update the gradient-boosted decision tree classification module;
[0105] S5. Convert the deep feature representations into low-dimensional feature representations via the feature distillation module, and fuse them with the intermediate classification features of the gradient-boosted decision tree classification module after weight adjustment to generate the fused optimized feature representations. The optimized feature representations are further input into the gradient-boosted decision tree classification module for fine classification training, and the optimized classification results are output;
[0106] S6. Implement spatial consistency optimization, classification boundary refinement, and class reassignment on the optimized classification results to generate the final multi-class land object recognition results of the remote sensing images.
[0107] In this embodiment, S1 includes the following steps:
[0108] S11. Collect remote sensing image data. The resolution of the remote sensing image data is R, the number of bands is B, and the image size is ;
[0109] S12. Perform radiometric correction, geometric correction, and noise removal on the remote sensing image data;
[0110] S13. Fill the cloud-covered areas of the remote sensing image data after radiometric correction, geometric correction, and noise removal. Based on the principle of local similarity, estimate the pixel values of the cloud-covered areas using the surrounding pixel information, calculate the nearest neighbor unoccluded pixel set, and use the weighted interpolation method for filling;
[0111] S14. Optimize the contrast of remote sensing image data using the adaptive histogram equalization method and output the standardized remote sensing image data D:
[0112] ;
[0113] Among them, is the Nth standardized remote sensing image, and N is the total number of remote sensing image data.
[0114] In this embodiment, S2 includes the following steps:
[0115] S21. Use the self-supervised learning method to perform feature pre-training on the standardized remote sensing image data D and construct a cross-scale spatial-spectral fusion self-supervised feature extraction model F:
[0116] ;
[0117] Among them, is the feature representation after feature extraction of the standardized remote sensing image , and d is the feature dimension; , d is the feature dimension;
[0118] S22. Construct a cross-scale contrast learning mechanism, perform multi-scale decomposition on the standardized remote sensing image data, define the scale transformation set T, and construct multi-scale enhanced samples and , and calculate their feature vectors ;
[0119] Define the scale contrast loss:
[0120] ;
[0121] Among them, is the cosine similarity of the feature vectors, is the temperature parameter, and j represents the jth scale transformation sample in the scale contrast learning mechanism;
[0122] S23. Design the object-background segmentation constraint, construct the object mask , define the object pixel set and the background pixel set , and calculate the mean of the object and background features:
[0123]
[0124] ;
[0125] Among them, M is the number of object pixels, K is the number of background pixels, and Represents the ground object pixel or the background pixel representation vector in the depth feature space;
[0126] Define the ground object discrimination loss to make the ground object and background features distinguishable:
[0127] ;
[0128] S24. Adopt the spectral-spatial fusion enhancement mechanism to construct spectral perturbation samples:
[0129] ;
[0130] wherein, is Gaussian noise, is the perturbation coefficient, is the remotely sensed image after spectral perturbation, and calculate the spectral invariance loss:
[0131] ;
[0132] wherein, is the remotely sensed image after spectral perturbation is the feature processed by the cross-scale spatial-spectral fusion self-supervised feature extraction model, is the remotely sensed image is the feature processed by the cross-scale spatial-spectral fusion self-supervised feature extraction model;
[0133] And fuse the spatial information to define the local consistency constraint:
[0134] ;
[0135] wherein, is the m-th pixel of the normalized remotely sensed image , is the pixel the neighborhood window around, including the pixel points close to its spatial position, represents the feature representation of the pixel point in the remotely sensed image after being processed by the cross-scale spatial-spectral fusion self-supervised feature extraction model, represents a certain pixel point in the normalized remotely sensed image , is the number of pixel points in the neighborhood window;
[0136] S25. Calculate the final improved self-supervised learning loss, combining cross-scale contrast learning, ground object-background segmentation constraint and spectral-spatial fusion mechanism:
[0137] ;
[0138] wherein, is a weight parameter;
[0139] and output a set of deep feature representations extracted by self-supervised learning:
[0140] ;
[0141] wherein, is the deep feature representation of the Nth remote sensing image.
[0142] In this embodiment, S3 includes the following steps:
[0143] S31. Use the gradient-boosted decision tree classification module to perform preliminary classification on the deep feature representation, construct the gradient-boosted decision tree classification module G, and define the mapping:
[0144] ;
[0145] wherein, is the predicted class label, is the deep feature representation of the ith remote sensing image, and the value range is , and C is the total number of ground object classes;
[0146] S32. Use the decision tree model based on gradient boosting for classification training and construct the loss function:
[0147] ;
[0148] wherein, is the true class label, is the true probability of the cth class, is the class probability predicted by the gradient-boosted decision tree classification module;
[0149] S33. Use the weighted gradient update strategy to calculate the negative gradient of the tth iteration:
[0150] ;
[0151] Based on the negative gradient construct the tth decision tree , and update the classification model:
[0152] ;
[0153] wherein, is the learning rate, which controls the contribution of each decision tree;
[0154] S34. Use the deep feature importance calculation mechanism to evaluate the contributions of different feature dimensions and define the feature importance score:
[0155] ;
[0156] Among them, is the j-th feature, T is the total number of decision trees, is the number of nodes of the t-th decision tree, is the reduction in loss brought by node splitting. Sort the feature importance scores and select the most important feature dimensions;
[0157] S35. Calculate the balance coefficient of the class distribution using the class balance adjustment strategy:
[0158] ;
[0159] Among them, is the number of samples of the c-th class. Based on the balance coefficient adjust the loss function:
[0160] ;
[0161] S36. Iteratively optimize the gradient boosting decision tree classification module until the loss function converges or reaches the preset number of iteration rounds , and output the preliminary multi-class land cover classification result:
[0162] ;
[0163] Among them, is the set of preliminary classification labels for all remote sensing images, is the predicted class label of the N-th remote sensing image.
[0164] In this embodiment, S4 includes the following steps:
[0165] S41. Obtain the gradient feedback value by taking the gradient of the loss function with respect to the preliminary classification result:
[0166] ;
[0167] S42. For multiple weak classifiers in the gradient boosting decision tree classification module, let the weight of the t-th weak classifier be , among which, , T is the total number of weak classifiers. According to the gradient feedback value of each sample i, calculate the weight adjustment amount of this weak classifier, and define the adjustment amount as:
[0168] ;
[0169] Among them, Denote the weight adjustment of the $t$-th weak classifier under the gradient guidance for the $i$-th sample;
[0170] S43. Implement dynamic weight update for each weak classifier:
[0171] ;
[0172] where is the weight of the updated $t$-th weak classifier, is the learning rate parameter, controlling the update step size;
[0173] S44. Use the updated weight to re-optimize the decision boundaries of each weak classifier, and combine the deep feature representation to perform overall fine-tuning or retraining on the gradient-boosted decision tree classification module to construct the updated classification module $G'$;
[0174] S45. Output the classification result obtained by predicting the deep feature representation $Z$ using the updated gradient-boosted decision tree classification module $G'$:
[0175] ;
[0176] where is the predicted class probability vector of the $N$-th remote sensing image after dynamic adjustment.
[0177] In this embodiment, S5 includes the following steps:
[0178] S51. Use the feature distillation module to perform dimensionality reduction transformation on the deep feature representation. The feature distillation module optimizes redundant features by constructing a lightweight representation learning network while maintaining the information of the original feature space, and generates a low-dimensional feature representation after dimensionality reduction;
[0179] S52. Use the information entropy optimization strategy to screen the low-dimensional feature representation after dimensionality reduction, calculate the information contribution degree of each feature dimension based on the class discrimination ability, retain the most discriminative features and eliminate low-correlation features;
[0180] S53. Use the feature fusion strategy to fuse the low-dimensional feature representation output by the feature distillation module with the intermediate classification features after weight adjustment of the optimized gradient-boosted decision tree classification module. During the fusion process, dynamic weighting parameters are set for different feature sources, so that the deep features extracted by self-supervised learning and the highly discriminative features generated by the gradient-boosted decision tree classification module act synergistically to form a fused optimized feature representation;
[0181] S54. Retrain the gradient-boosted decision tree classification module using the fused feature representation, and introduce a hierarchical learning strategy during training. First, train on the easily classifiable samples and gradually transition to the categories with blurred boundaries and easy confusion.
[0182] S55. Adopt a fine-grained classification training strategy to locally iteratively optimize the fused feature representation. Dynamically adjust the class discrimination boundary during training. Increase the decision tree depth for the confused categories above the threshold, and adopt a sample balancing strategy for the frequency categories below the threshold.
[0183] S56. Output the optimized multi-class land cover recognition classification results. The optimized multi-class land cover recognition classification results are generated based on the fused feature representation and have undergone multi-level optimization and adjustment.
[0184] Example 1: To verify the effectiveness of the remote sensing image multi-class land cover recognition method based on gradient-boosted decision tree and self-supervised learning in complex land cover classification tasks, this example selects the multi-spectral remote sensing image data of City A and its surrounding areas, and conducts a land cover classification experiment in combination with the geographic information system. The area contains multi-class land covers such as urban buildings, roads, farmlands, forests, water bodies, and wetlands, with complex spectral characteristics and highly mixed spatial distributions, providing a real experimental environment for remote sensing image classification tasks.
[0185] Traditional remote sensing land cover classification methods usually rely on a large amount of manually labeled data and use random forests, support vector machines, or deep neural networks for classification. However, due to high labeling costs, data imbalance, and insufficient classification accuracy in complex environments, traditional methods have great limitations in dealing with multi-class, cross-scale, and hyperspectral land cover classification tasks. This invention combines self-supervised learning with gradient-boosted decision trees to automatically learn the deep features of remote sensing images without the need for large-scale manually labeled data, and combines the decision tree model for efficient classification.
[0186] This experiment uses the Sentinel-2 multi-spectral remote sensing images of City A and its surrounding areas from January 2023 to December 2023, with a spatial resolution of 10m, including 12 spectral bands, and the image coverage area is about 5000 square kilometers. The experimental data includes 1200 high-resolution images, of which 800 are used for training and 400 are used for testing. The land cover distribution data provided by the geographic database is used for land cover category labeling to evaluate the classification accuracy.
[0187] The specific land cover categories and the number of samples are shown in Table 1 below:
[0188] Table 1 Statistics of Land Cover Category Training and Test Samples
[0189] Feature class Number of training samples Number of test samples Percentage (%) Urban buildings 50,000 20,000 18.3 Roads 40,000 15,000 14.6 Farmland 70,000 25,000 21.1 Forests 65,000 22,000 19.8 Water bodies 35,000 12,000 10.5 Wetlands 25,000 10,000 7.9 Others 20,000 8,000 7.8
[0190] During the experiment, we used traditional methods (CNN, random forest, SVM) and the method of the present invention (self-supervised learning + gradient boosting decision tree) to perform multi-category land object recognition on remote sensing images, and made a comprehensive comparison of classification accuracy, computational efficiency, and model stability.
[0191] Firstly, the remote sensing images are subjected to radiation correction, geometric correction, noise removal, cloud-occluded area filling and image enhancement to generate standardized remote sensing image data. The traditional method directly uses the data for training, while the method of the present invention further adopts self-supervised learning for feature pre-training to mine the spectral and spatial characteristics of the data.
[0192] The present invention adopts cross-scale spatial-spectral fusion self-supervised learning to perform multi-perspective comparative learning, object-background segmentation constraints, and spectral enhancement tasks on the training data, automatically learns the deep features of different categories of objects, and constructs an efficient feature extraction network through unsupervised feature learning, which reduces the dependence on manually labeled data and improves the separability and generalization ability of classification features.
[0193] In the classification stage, traditional methods directly train CNN, random forest and SVM based on labeled data, while the method of the present invention uses gradient enhanced decision trees for classification and combines feature distillation to reduce the dimensionality of high-dimensional features, reduce computational overhead, and improve the decision-making ability of the model. During the training process, we use multi-stage gradient feedback optimization to adjust the decision boundary to improve the ability to distinguish confused categories.
[0194] During the testing phase, we used 400 test images to compare the performance of the proposed method with traditional methods in terms of classification accuracy, computational efficiency, and category recognition ability.
[0195] The experimental results shown in Table 2 below show that the method of the present invention is significantly superior to the traditional method in terms of classification accuracy, and has better discrimination ability in categories with high spectral similarity (forest and wetland, road and building).
[0196] Table 2 Comparison of classification accuracy between the method of the present invention and the traditional method
[0197] Method Overall classification accuracy (OA, %) Kappa coefficient Urban buildings (%) Roads (%) Farmland (%) Forests (%) Water bodies (%) Wetlands (%) Random forest 79.6 0.75 85.2 78.1 80.4 76.3 82.5 68.7 SVM 82.3 0.78 86.7 80.5 82.8 78.9 84.2 71.2 CNN 85.4 0.81 89.1 84.2 86.7 82.1 87.5 74.5 Method of the present invention 92.6 0.89 94.3 90.8 93.1 89.7 92.4 85.2
[0198] The method of the present invention is superior to traditional methods in both training time and inference time, and can achieve efficient processing on large-scale data sets. The method of the present invention reduces the training time by 53.5% and increases the inference speed by 59.8%, and has a significant advantage in computing efficiency.
[0199] Traditional methods have large errors in classifying frequency categories below the threshold. The method of the present invention improves the recognition ability of frequency categories below the threshold through feature distillation and gradient optimization, increasing the wetland recognition accuracy by 10.7% and significantly improving the classification accuracy of farmland and forest categories.
[0200] This embodiment shows that the method of the present invention can effectively solve the problems of traditional remote sensing ground object classification, such as high dependence on labeled data, low computational efficiency, and insufficient category discrimination ability. It is superior to traditional methods in terms of classification accuracy, computational speed, and generalization ability, and can be widely applied to the fields of remote sensing data automatic classification, land use monitoring, and environmental assessment, providing an efficient and stable solution for intelligent analysis of remote sensing images.
[0201] The present invention adopts a self-supervised learning strategy of cross-scale spatial-spectral fusion, combined with multi-view contrast learning, spatial consistency learning, and spectral transformation enhancement, to achieve high-quality feature extraction of multi-category ground objects in remote sensing images without a large amount of manually labeled data. By introducing ground object-background segmentation constraints and spectral invariance regularization, it can extract robust representative features under different ground object categories, lighting conditions, and spatio-temporal scales, and can make full use of the inherent correlation of spectral information when processing hyperspectral and multispectral data, enabling ground object classification to maintain a high resolution ability in the face of spectral mixing effects.
[0202] The present invention combines feature distillation and gradient-enhanced decision tree classification module in the process of remote sensing image classification. The feature distillation module reduces the dimension of high-dimensional deep features, removes redundant information, extracts the most discriminative low-dimensional feature representation, and fuses it with the intermediate features of the gradient-enhanced decision tree classifier, thereby improving the classification accuracy and reducing the computational complexity. By constructing a lightweight feature distillation network, the input dimension of features is reduced by 35%-50%, but the information retention rate can still reach more than 95%. In addition, the feature fusion mechanism enables GBDT to utilize the discriminative features extracted by deep learning while maintaining adaptability to small sample data.
[0203] The present invention introduces a multi-stage dynamic weight adjustment strategy based on gradient feedback in the training process of the gradient-enhanced decision tree classification module, conducts gradient analysis on the classification errors of different categories, and adjusts the decision weights of weak classifiers in each round of iteration. By constructing an adaptive category balance loss function, the classification accuracy of frequency categories below the threshold is improved, and the local feature reweighting mechanism is used to enable the classifier to dynamically adjust the decision boundary for confused categories above the threshold, improving the discrimination ability of the classification model in complex ground object environments.
[0204] As described above, it is only the preferred specific implementation manner of the present invention, but the protection scope of the present invention is not limited thereto. Any person skilled in the art within the technical scope disclosed by the present invention, according to the technical solution and inventive concept of the present invention, making equivalent replacements or changes should be covered within the protection scope of the present invention.
Claims
1. A multi-category object recognition method for remote sensing images based on self-supervised learning, characterized in that: The steps include: S1. Acquire remote sensing image data and preprocess the remote sensing image data, including radiation correction, geometric correction, noise removal, cloud occlusion area filling and image enhancement processing, to generate standardized remote sensing image data; S2. performing feature pre-training on the standardized remote sensing image data based on self-supervised learning, wherein the self-supervised learning module automatically extracts deep-level feature representations with discriminative capabilities from the standardized remote sensing image data by constructing multi-view comparison tasks, spatial consistency tasks, and spectral domain transformation tasks, and outputs the deep feature representations; S21. Use self-supervised learning method to pre-train the features of standardized remote sensing image data D, and construct a cross-scale spatial-spectral fusion self-supervised feature extraction model F; S22. Construct a cross-scale contrast learning mechanism, perform multi-scale decomposition on standardized remote sensing image data, define the scale transformation set T, and construct multi-scale enhanced samples and Calculate their eigenvectors respectively And define scale contrast loss; S23. Design object-background segmentation constraints and construct object mask M i , define the ground object pixel set P i and background pixel set B i , and calculate the mean of ground object and background features, define ground object distinction loss, and make ground object and background features distinguishable; S24. Using the spectral-spatial fusion enhancement mechanism, construct the spectral perturbation sample I' i , calculate the spectral invariance loss, fuse the spatial information, and define the local consistency constraint; S25. Calculate the final improved self-supervised learning loss, combining cross-scale contrastive learning, object-background segmentation constraints and spectral-spatial fusion mechanism, and output the deep feature representation set Z extracted by self-supervised learning; S3. Input the deep feature representation into a gradient enhanced decision tree classification module, which is composed of a plurality of weak classifiers, performs preliminary multi-category classification of the deep feature representation through iterative optimization, and outputs preliminary classification results; S4. Calculate the classification error based on the preliminary classification results and dynamically adjust the weights of the weak classifiers based on the gradient feedback mechanism, optimize the decision boundaries of each weak classifier, and update the gradient enhanced decision tree classification module; S41. Through the loss function Find the gradient of the preliminary classification results and obtain the gradient feedback value g i ,in, is the predicted category label, y i is the true category label; S42. For multiple weak classifiers of the gradient enhanced decision tree classification module, let the weight of the tth weak classifier be w (t) , where t = 1, 2, ..., T, T is the total number of weak classifiers, and the gradient feedback value g of each sample i is i Calculate the weight adjustment of the weak classifier and define the adjustment as S43. Dynamic weight update is implemented for each weak classifier; S44. Using the updated weights Reoptimize the decision boundaries of each weak classifier and combine the deep feature representation Z i Fine-tune or retrain the gradient boosted decision tree classification module as a whole to construct an updated classification module G'; S45. Output the updated gradient boosted decision tree classification module G' to predict the classification result of the deep feature representation Z S5. The deep feature representation is converted into a low-dimensional feature representation through the feature distillation module, and is fused with the intermediate classification features after weight adjustment of the gradient boosting decision tree classification module to generate a fused optimized feature representation, which is further input into the gradient boosting decision tree classification module for fine classification training, and outputs the optimized classification result; S6. Implement spatial consistency optimization, classification boundary refinement and category reallocation on the optimized classification results to generate the final remote sensing image multi-category object recognition results.
2. The method for multi-category object recognition in remote sensing images based on self-supervised learning according to claim 1, characterized in that: The S1 comprises the following steps: S11. Collecting remote sensing image data, wherein the resolution of the remote sensing image data is R, the number of bands is B, and the image size is H×W; S12. Performing radiation correction, geometric correction and noise removal on the remote sensing image data; S13. Filling the cloud-blocked area of the remote sensing image data after radiation correction, geometric correction and noise removal, based on the principle of local similarity, using the surrounding pixel information to estimate the pixel value of the cloud-blocked area, calculating the nearest neighbor unblocked pixel set and filling it using the weighted interpolation method; S14. Adopting the adaptive histogram equalization method to optimize the contrast of the remote sensing image data, and outputting the standardized remote sensing image data D: D={I1,I2,...,I N }; Among them, I N is the Nth standardized remote sensing image, and N is the total number of remote sensing image data.
3. The method for multi-category object recognition in remote sensing images based on self-supervised learning according to claim 1, characterized in that: The S2 comprises the following steps: S21. Use self-supervised learning methods to pre-train the features of standardized remote sensing image data D and construct a cross-scale spatial-spectral fusion self-supervised feature extraction model F: F:I i →Z i ; Among them, Z i is the standardized remote sensing image I i After feature extraction, the feature representation is d is the feature dimension; S22. Construct a cross-scale contrast learning mechanism, perform multi-scale decomposition on standardized remote sensing image data, define the scale transformation set T, and construct multi-scale enhanced samples and Calculate their eigenvectors respectively Define scale contrast loss: in, is the cosine similarity of the feature vector, τ is the temperature parameter; S23. Design object-background segmentation constraints and construct object mask M i , define the ground object pixel set P i and background pixel set B i , and calculate the mean of ground objects and background features: Among them, M is the number of ground object pixels, K is the number of background pixels, F(p im ) and F(b ik ) represents the ground feature pixel p im or background pixel b ik Representation vector in deep feature space; Define the object discrimination loss to make the objects distinguishable from the background features: S24. Use the spectral-spatial fusion enhancement mechanism to construct spectral perturbation samples: I' i =I i +α·N(0,σ 2 ); Among them, N(0,σ 2 ) is Gaussian noise, α is the disturbance coefficient, I' i Calculate the spectral invariance loss for the remote sensing image after spectral perturbation: Among them, F(I' i ) is the remote sensing image I' after spectral perturbation i Features processed by the self-supervised feature extraction model through cross-scale spatial-spectral fusion; And integrate spatial information to define local consistency constraints: Among them, p im is the standardized remote sensing image I i The mth pixel, W(p im ) is pixel p im The surrounding neighborhood window contains pixels close to its spatial position, F(p k ) represents the pixel point p in the remote sensing image k The feature representation after processing by the cross-scale spatial-spectral fusion self-supervised feature extraction model, p k Represents the standardized remote sensing image I i A pixel point in |W(p im )| is the number of pixels in the neighborhood window; S25. Calculate the final improved self-supervised learning loss, combining cross-scale contrastive learning, object-background segmentation constraints and spectral-spatial fusion mechanism: Among them, λ1,λ2,λ3,λ4 are weight parameters; And output the set of deep feature representations extracted by self-supervised learning: Z={Z1,Z2,...,Z N }; Among them, Z N is the deep feature representation of the Nth remote sensing image.
4. The method for multi-category object recognition in remote sensing images based on self-supervised learning according to claim 1, characterized in that: The S3 comprises the following steps: S31. Use the gradient enhanced decision tree classification module to perform preliminary classification on the deep feature representation, build the gradient enhanced decision tree classification module G, and define the mapping: in, To predict the class label, Z i is the deep feature representation of the i-th remote sensing image, with a value range of {1,2,...,C}, where C is the total number of ground object categories; S32. Use the decision tree model based on gradient boosting for classification training and construct the loss function: Among them, y i is the true category label, y i,c is the true probability of the cth class, P(y i,c ∣Z i ) is the category probability predicted by the gradient boosted decision tree classification module; S33. Using the weighted gradient update strategy, calculate the negative gradient of the tth iteration: Based on negative gradient Construct the tth decision tree h (t) , update the classification model: Among them, η1 is the learning rate, which controls the contribution of each decision tree; S34. Use a deep feature importance calculation mechanism to evaluate the contribution of different feature dimensions and define the feature importance score: Among them, f j is the jth feature, T is the total number of decision trees, N t is the number of nodes in the t-th decision tree, The loss reduction caused by node splitting is calculated, the feature importance scores are sorted, and the most important feature dimensions are selected; S35. Calculate the balance coefficient of category distribution using the category balance adjustment strategy: Among them, N c is the number of samples in the cth class, based on the equalization coefficient w c Adjust the loss function: S36. Iteratively optimize the gradient enhanced decision tree classification module until the loss function converges or reaches the preset number of iterations T max , output preliminary multi-category classification results: in, is the preliminary classification label set for all remote sensing images, is the predicted category label of the Nth remote sensing image.
5. The method for multi-category object recognition in remote sensing images based on self-supervised learning according to claim 1, characterized in that: The S4 comprises the following steps: S41. Through the loss function Find the gradient of the preliminary classification results and obtain the gradient feedback value: S42. For multiple weak classifiers of the gradient enhanced decision tree classification module, let the weight of the tth weak classifier be w (t) , where t = 1, 2, ..., T, T is the total number of weak classifiers, and the gradient feedback value g of each sample i is i Calculate the weight adjustment of the weak classifier and define the adjustment as: in, represents the weight adjustment of the t-th weak classifier under the guidance of the gradient for the i-th sample; S43. Implement dynamic weight update for each weak classifier: in, is the updated weight of the tth weak classifier, η2 is the learning rate parameter, which controls the update step size; S44. Using the updated weights Reoptimize the decision boundaries of each weak classifier and combine the deep feature representation Z i Fine-tune or retrain the gradient boosted decision tree classification module as a whole to construct an updated classification module G'; S45. Output the classification result obtained after the updated gradient enhanced decision tree classification module G' predicts the deep feature representation Z: in, is the predicted category probability vector of the Nth remote sensing image after dynamic adjustment.
6. The method for multi-category object recognition in remote sensing images based on self-supervised learning according to claim 1, characterized in that: The S5 comprises the following steps: S51. A feature distillation module is used to reduce the dimension of the deep feature representation. The feature distillation module optimizes redundant features while maintaining the original feature space information by constructing a lightweight representation learning network, and generates a low-dimensional feature representation after dimensionality reduction; S52. Use information entropy optimization strategy to screen the low-dimensional feature representation after dimensionality reduction, calculate the information contribution of each feature dimension based on the category discrimination ability, retain the most discriminative features and eliminate low-relevance features; S53. A feature fusion strategy is used to fuse the low-dimensional feature representation output by the feature distillation module with the intermediate classification features after weight adjustment by the optimized gradient boosting decision tree classification module. During the fusion process, dynamic weighting parameters are set for different feature sources so that the deep features extracted by self-supervised learning and the high-discrimination features generated by the gradient boosting decision tree classification module work together to form a fused optimized feature representation; S54. Use fusion feature representation to retrain the gradient boosting decision tree classification module, introduce a hierarchical learning strategy in the training process, first train the easy-to-classify samples, and gradually transition to the boundary fuzzy and easily confused categories; S55. Use fine classification training strategy to perform local iterative optimization on fusion feature representation, dynamically adjust category discrimination boundary during training, increase decision tree depth for high confusion categories, and use sample balancing strategy for low frequency categories; S56. Outputting optimized multi-category ground object recognition and classification results, wherein the optimized multi-category ground object recognition and classification results are generated based on the fusion feature representation and undergo multi-level optimization and adjustment.
Citation Information
Patent Citations
Long-tail image recognition method based on self-supervision and self-distillation
CN113837238A
Remote sensing image building roof contour identification method, device, equipment and medium
CN116645595A
All-electric drive aircraft tractor control system and method
CN119087822A
Remote sensing visual system based on supervised training
CN119762835A