NSGA-II-XGBoost machine learning-based high-precision identification method and equipment for cultivated land in plateau mountain area, and medium

By adopting the NSGA-II-XGBoost model in the identification of cultivated land in the plateau mountainous areas, combining multi-source remote sensing data and texture features, the problems of low accuracy and overfitting of cultivated land under complex terrain are solved, and high-precision cultivated land recognition and model optimization are achieved.

CN120217048APending Publication Date: 2025-06-27KUNMING UNIV OF SCI & TECH
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510302210.4
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-03-14
Publication Date
2025-06-27

AI Technical Summary

Technical Problem

Under the complex terrain of the plateau mountainous areas, existing farmland identification and classification models show lower classification accuracy and higher risk of overfitting, especially in areas with similar spectral characteristics or blurred boundaries.

Method used

Using a machine learning method based on NSGA-II-XGBoost, a multi-source remote sensing data (spectral, radar, and terrain data), pre-processing and preliminary classification, combining texture feature extraction and terrain feature fusion, a multi-source feature data set is generated. The model is interpreted using SHAP, snatched irrelevant features, and optimized the hyperparameters and feature combinations of the model through NSGA-II.

Benefits of technology

It significantly improves the accuracy of cultivated land classification in complex terrain areas, avoids overfitting, improves the interpretability and generalization capabilities of the model, and can effectively deal with the challenges of similar spectral characteristics or blurred boundaries.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120217048A_ABST
    Figure CN120217048A_ABST
Patent Text Reader

Abstract

The invention relates to the technical field of cultivated land recognition, in particular to a highland mountain area cultivated land high-precision recognition method and device based on NSGA-II-XGBoost machine learning and a medium, and the method specifically comprises the following steps: S1, obtaining multi-source remote sensing data, and carrying out the preprocessing of the multi-source remote sensing data; s2, performing preliminary classification on the data obtained in the step S1 through a GEE platform, and enhancing classification precision by using a texture feature extraction technology; s3, fusing the data obtained in the step S2 according to the topographic features and the texture features to generate a multi-source feature data set; s4, explaining the XGBoost model by using the SHAP, evaluating the contribution of each feature in the topographic features and the texture features to the cultivated land identification result, and eliminating irrelevant or redundant features so as to optimize the input features of the model; and S5, a non-dominated sorting genetic algorithm II is adopted to optimize the hyper-parameter and feature combination of the XGBoost classification model, and an NSGA-II-XGBoost model is obtained. The method can effectively improve classification precision, can cope with challenges of similar spectral features or fuzzy boundaries, and has high application potential.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of cultivated land identification, and particularly to a high-precision identification method, device and medium for cultivated land in plateau mountainous areas based on NSGA-II-XGBoost machine learning. Background Art

[0002] Due to the complex terrain and drastic undulation in plateau mountainous areas, the distribution of cultivated land shows the characteristics of fragmentation and small scale, making it a severe challenge to achieve high-precision identification of cultivated land in remote sensing images. Accurate identification of cultivated land can not only significantly improve the implementation effect of precision agriculture (such as precision fertilization, irrigation and pest management), thereby effectively improving crop yield and quality, but also provide important technical support for the scientific planning and refined management of agricultural resources.

[0003] Existing cultivated land identification and classification models have great difficulty in distinguishing cultivated land from other land types with similar spectral features or fuzzy boundaries in complex terrain areas such as plateau mountainous areas. Traditional remote sensing image classification methods, such as support vector machine (SVM), random forest (RF) and conventional XGBoost models, although can effectively handle conventional cultivated land identification problems, show low classification accuracy and high overfitting risk in areas with high altitude, complex terrain and similar spectral features.

[0004] In recent years, with the rapid development of deep learning and machine learning technologies, remote sensing image analysis methods based on neural networks have been widely used. However, in the complex terrain of plateau mountainous areas, how to overcome the influence of terrain factors on the identification accuracy is still an urgent problem to be solved. Most of the existing multi-objective optimization frameworks lack the comprehensive application of spectral, radar and terrain data, and most models fail to fully explore the correlation and complementarity between various features when processing multi-source data, which affects the improvement of classification accuracy. Summary of the Invention

[0005] The features and advantages of the present invention are partially stated in the following description, or can be obvious from the description, or can be learned by practicing the present invention.

[0006] To overcome the problems of the prior art, the present invention provides a high-precision identification method for cultivated land in plateau mountainous areas based on NSGA-II-XGBoost machine learning, which specifically includes the following steps:

[0007] S1. Obtain multi-source remote sensing data and preprocess the multi-source remote sensing data;

[0008] S2. Perform preliminary classification on the data obtained in step S1 through the GEE platform, and use texture feature extraction technology to enhance the classification accuracy;

[0009] S3. Integrate the data obtained in step S2 based on topographic features and texture features to generate a multi-source feature dataset;

[0010] S4. Use SHAP to interpret the XGBoost model, evaluate the contribution of each feature in topographic features and texture features to the cultivated land recognition result, and eliminate irrelevant or redundant features, thereby optimizing the input features of the model;

[0011] S5. Optimize the hyperparameters and feature combinations of the XGBoost classification model using the Non-dominated Sorting Genetic Algorithm II to obtain the NSGA-II-XGBoost model.

[0012] Preferably, the multi-source remote sensing data includes spectral data, radar data, and elevation data.

[0013] Preferably, the preprocessing of the spectral data includes cloud removal and resampling through the Cloud Score+ algorithm of GEE, and also includes radiometric correction, geometric correction, and atmospheric correction.

[0014] Preferably, the preprocessing of the radar data includes noise removal, geometric correction, radiometric calibration, and terrain correction, and at least one of using the Lee despeckling algorithm to remove speckle noise.

[0015] Preferably, the preprocessing of the elevation data includes calculating slope and aspect based on DEM.

[0016] Preferably, the spectral indices of the spectral data include the Normalized Difference Vegetation Index, Enhanced Vegetation Index, Normalized Difference Water Index, and Bare Soil Index.

[0017] Preferably, the texture feature extraction technique includes the Gray Level Co-occurrence Matrix.

[0018] Preferably, the objective function optimized by the Non-dominated Sorting Genetic Algorithm II includes classification accuracy, F1 score, recall rate, and Kappa coefficient.

[0019] The present invention also provides a computer device, including: a memory, a processor, and a computer program stored on the memory and executable on the processor, where the processor executes the computer program to implement the steps of the high-precision cultivated land recognition method for plateau mountainous areas based on NSGA-II-XGBoost machine learning as described above.

[0020] The present invention also provides a computer-readable storage medium, on which a computer program is stored, and when the computer program is executed by a processor, it implements the steps of the high-precision cultivated land recognition method for plateau mountainous areas based on NSGA-II-XGBoost machine learning as described above.

[0021] Advantages of the present invention: The present invention combines spectral, radar, and terrain data, fully explores the influence of terrain factors on cultivated land classification, and significantly improves the classification accuracy in complex terrain areas. By adopting the NSGA-II-XGBoost model optimization framework, the hyperparameters and feature combinations are optimized through genetic algorithms, improving the model accuracy and generalization ability. Particularly in complex terrains such as plateau mountainous areas, overfitting is effectively avoided. At the same time, the SHAP value analysis is used to quantitatively evaluate the feature contributions, enhancing the interpretability and performance of the model. This optimization method can effectively improve the classification accuracy for the problem of cultivated land identification in high-altitude and complex terrains, and can address the challenges of similar spectral features or blurred boundaries, having strong practical application potential. Brief Description of the Drawings

[0022] The present invention will be specifically described below with reference to the accompanying drawings and examples. The advantages and implementation manners of the present invention will become more apparent. The content shown in the accompanying drawings is only used for the explanation of the present invention and does not constitute any limitation to the present invention. In the drawings:

[0023] Figure 1 is a flowchart of a high-precision identification method for cultivated land in plateau mountainous areas based on NSGA-II-XGBoost machine learning in a specific embodiment of the present invention;

[0024] Figure 2 is a visualization diagram of SHAP values of some features of a high-precision identification method for cultivated land in plateau mountainous areas based on NSGA-II-XGBoost machine learning in a specific embodiment of the present invention;

[0025] Figure 3 is a flowchart of NSGA-II optimizing XGBoost of a high-precision identification method for cultivated land in plateau mountainous areas based on NSGA-II-XGBoost machine learning in a specific embodiment of the present invention;

[0026] Figure 4 are cultivated land distribution maps processed by different models. Among them, Figure A is the cultivated land distribution map based on the XGBoost model; Figure B is the cultivated land distribution map based on the NSGA-II-XGBoost model; Figure C is the cultivated land distribution map of the third national land survey (the third survey); Figure D is the comparison result map of different cultivated land plots processed by the XGBoost and NSGA-II-XGBoost models respectively. Detailed Embodiment

[0027] The embodiments of the present invention will be further described in detail below with reference to the accompanying drawings and examples. The following examples are used to illustrate the present invention, but cannot be used to limit the scope of the present invention.

[0028] As Figure 1As shown in the figure, the present invention provides a high-precision identification method for cultivated land in plateau mountainous areas based on NSGA-II-XGBoost machine learning, which specifically includes the following steps:

[0029] S1. Obtain multi-source remote sensing data and preprocess the multi-source remote sensing data; the multi-source remote sensing data includes spectral data, radar data, and elevation data; the spectral (Sentinel-2) data is sourced from the Google Earth Engine (GEE) platform, and the preprocessing of the spectral data includes cloud removal and resampling through the Cloud Score+ algorithm of GEE, and also includes radiometric calibration, geometric calibration, and atmospheric calibration. The influence of the atmosphere on the remote sensing data is removed through atmospheric calibration to ensure that the reflectance values in the image can truly reflect the features of the ground objects.

[0030] The radar (Sentinel-1) data is sourced from the GEE platform, and the preprocessing steps include noise removal, geometric calibration, radiometric calibration, and terrain correction, and the Lee despeckling algorithm is used to remove speckle noise, thereby improving the image quality. The structural features of the ground surface (such as terrain undulation, vegetation coverage, etc.) can be extracted through radar data, and these data are particularly important in plateau mountainous areas.

[0031] The elevation (SRTM DEM) data is sourced from the GEE platform, and the preprocessing includes calculating the slope and aspect based on the DEM (Digital Elevation Model) and providing them as input features to the subsequent classification algorithm.

[0032] S2. Initially classify the data obtained in step S1 through the GEE platform and use texture feature extraction technology to enhance the classification accuracy; initially classify the spectral data through the percentage method of the GEE platform. This method classifies the remote sensing image into different land use types according to the matching degree between the spectral values of the pixels and the training samples. This step can effectively distinguish land use types such as cultivated land, forest, grassland, and bare soil. The texture feature extraction technology includes methods such as the gray-level co-occurrence matrix (GLCM) to further enhance the classification accuracy. Texture features can help the model distinguish the subtle differences existing in different land use types. Especially when identifying cultivated land types, it can distinguish different land management and use methods.

[0033] S3. Fuse the data obtained in step S2 based on the terrain features and texture features to generate a multi-source feature dataset; the terrain features include DEM, slope, aspect, altitude, etc. This process not only retains the spectral information of the remote sensing image but also adds terrain information, which helps to improve the performance of the classification model in complex terrains. Data fusion can effectively enhance the feature space of the classification model, help the model better understand the complex ground object distribution pattern in plateau mountainous areas, and thus improve the accuracy of cultivated land identification.

[0034] S4. Use SHAP (Shapley Additive Explanations) to interpret the XGBoost model, evaluate the contribution of each feature in terrain features and texture features to the cultivated land recognition result, and eliminate irrelevant or redundant features, thereby optimizing the input features of the model; SHAP values can quantify the impact of each feature on the classification result and reveal the positive or negative impact of each feature. According to the SHAP analysis results, select the features with the greatest contribution to the classification, and eliminate irrelevant or redundant features, thereby optimizing the input features of the model. This process can improve the computational efficiency of the model and enhance its classification accuracy.

[0035] S5. Use the non-dominated sorting genetic algorithm II to optimize the hyperparameters and feature combinations of the XGBoost classification model to obtain the NSGA-II-XGBoost model; the goal is to improve the performance of the model on the training set and test set by optimizing the hyperparameters and feature combinations of the model, avoid overfitting and ensure the generalization ability of the model. The optimized model can still show strong stability and accuracy when facing unknown data.

[0036] This embodiment obtains data based on the Google Earth Engine (GEE) platform, calls the Sentinel-1 and Sentinel-2 remote sensing images of the study area and performs preprocessing. Specifically, when preprocessing the remote sensing images, it includes performing atmospheric correction, radiometric correction on the Sentinel-2 remote sensing images, and using the CloudScore+ cloud removal algorithm to perform cloud removal processing on the Sentinel-2 data; using the Lee algorithm to perform speckle removal processing on the Sentinel-1 data. In this embodiment, the remote sensing images used include the visible light band, near-infrared band, short-wave infrared band, etc. of Sentinel-2; in order to better reflect different ground object information, spectral indices such as the Normalized Difference Vegetation Index (NDVI), Enhanced Vegetation Index (EVI), Normalized Difference Water Index (NDWI), and Bare Soil Index (BSI) are selected, and the Gray Level Co-occurrence Matrix (GLCM) is calculated using the B8 band data in Sentinel-2. This is a statistical-based texture feature extraction method. The Gray Level Co-occurrence Matrix (GLCM) includes the Angular Second Moment (ASM), Entropy (ENT), Sum Entropy (SEMT), Difference Entropy (DENT), Average Gray Value (SAVG), Inverse Difference Moment (IDM), Contrast, Correlation, Sum of Squares Variance (SVar), and Variance (Var). These statistics can quantitatively describe the texture characteristics of the image and reveal the texture feature differences between cultivated land and non-cultivated land. To extract the texture information in the remote sensing image, we set a window size of 3×3 and combined spectral and texture features to deeply analyze the image. These spectral indices can effectively distinguish ground objects such as cultivated land, grassland, water bodies, and urban buildings, improving the classification accuracy. The data types and sources are specifically shown in Table 1 as follows:

[0037] Table 1 Data Types and Sources

[0038]

[0039] In terms of radar data, this embodiment uses parameters such as the VV polarization backscattering coefficient, VH polarization backscattering coefficient, Polarization Ratio (PR), Total Polarization Ratio (TP), Non-uniform Polarization Ratio (NTPD), and Entropy of Sentinel-1. These parameters are particularly suitable for the identification of cultivated land in the complex terrain of plateau mountainous areas.

[0040] The classification of the spectral data is processed using the percentile method of GEE. Specifically, by constructing a histogram of the feature set and calculating the specified percentile values (5%, 25%, 50%, 75%, 95%) of the feature distribution, this method can effectively capture the spectral and radar information of Sentinel-2 and is widely used in land cover classification. This method extracts the key information of each input feature set by constructing a histogram of the feature set and calculating the specified percentile values of the feature distribution.

[0041] Fuse the topographic features extracted from the SRTM DEM data with the remote sensing data to generate a multi-source feature dataset. Cultivated land in mountainous areas of the plateau is usually distributed in areas with gentle slopes, so topographic features play an important role in cultivated land identification. Fuse spectral data, radar data, topographic data, and texture features to generate a multi-dimensional feature dataset, providing rich input data for subsequent classification models.

[0042] Use the SHAP (Shapley Additive Explanations) method to quantify the contribution of each feature in cultivated land identification. Through SHAP analysis, identify the most important features for cultivated land identification, and then optimize the selection of features. As Figure 2 shown, according to the SHAP analysis results, remove the features with low SHAP values, retain the features that contribute more to cultivated land identification, reduce data redundancy, and improve the model training efficiency.

[0043] Use the Non-dominated Sorting Genetic Algorithm II (NSGA-II) to optimize the hyperparameters and feature combinations of the XGBoost classification model, and then obtain the NSGA-II-XGBoost model. The specific optimization process is as Figure 3 shown, the optimization process mainly includes the following steps:

[0044] 1) Set the optimization objectives: Set four key performance indicators: accuracy, F1 score, recall, and Kappa coefficient. These indicators comprehensively reflect the classification ability of the model.

[0045] 2) Optimization of key hyperparameter selection: The key hyperparameters include maximum depth (max_depth), learning rate (learning_rate), column sampling ratio (threshold), L2 regularization (lambda), and L1 regularization (alpha). These hyperparameters have a significant impact on the model performance. The optimization process of key hyperparameter selection is as follows:

[0046] a. Initial population generation: Randomly generate multiple combinations of hyperparameters. Each combination represents an individual. Set the population size to N, and each individual i represents a combination of hyperparameters. Each hyperparameter value is randomly selected from a predefined range, as shown in Equation (1):

[0047] P0 = {p1, p2,... p N} (1)

[0048] Then evaluate these individuals. Use 5-fold cross-validation to calculate the performance metrics of the model, and train the XGBoost model using each combination of hyperparameters to calculate its fitness score to measure the performance of the model.

[0049] b. Fitness evaluation: Evaluate the individuals, calculate the fitness value of each individual. Calculate the performance metrics of the hyperparameters through 5-fold cross-validation, and train the XGBoost model using each combination of hyperparameters to calculate its fitness score to measure the performance of the model. Evaluate four objective functions: Accuracy, F1-Score, Recall, and Kappa coefficient, as shown in Equation (2):

[0050] f j (p i ) = Accuracy(p i ) (2)

[0051] Evaluate the relative position of individuals in the objective space through Crowding Distance, as shown in Equation (3):

[0052]

[0053] where and are the maximum and minimum values of objective j, and are the objective function values of adjacent individuals in the objective space.

[0054] Based on these scores, perform non-dominated sorting and crowding degree calculation on the individuals. The Pareto rank is used to evaluate the superiority and inferiority in multi-objective optimization, and the crowding degree is used to maintain the diversity of the population. Then perform selection, crossover, and mutation operations to generate new offspring individuals, continuously explore better combinations of hyperparameters. Finally, merge the parent and offspring individuals and perform screening according to the Pareto rank and crowding degree to keep the population size unchanged.

[0055] c. Selection, Crossover, and Mutation: Use Roulette Wheel Selection or Tournament Selection to generate a new generation of population. The function is shown in Equation (4):

[0056] P selected = Select(P) (4)

[0057] And improve diversity through the Crossover operation and the Mutation operation.

[0058] d. Termination Condition: The optimization process continues until the maximum number of generations is reached or the fitness no longer improves significantly. The function is shown in Equation (5):

[0059] Terminate if Generation ≥ Max Generation or Fitness change ≤ ε# (5)

[0060] During the optimization process, the termination condition is continuously checked, that is, it is judged whether the maximum number of iterations or the fitness threshold is reached. If the termination condition is met, the best combination of hyperparameters is output. Otherwise, return to the previous steps and continue iterative optimization until the optimal solution is found.

[0061] Finally, NSGA-II returns multiple optimized solutions that balance multiple objectives, forming the optimal combination of hyperparameters and features, and then obtaining the NSGA-II-XGBoost model

[0062] Model Validation

[0063] To verify the feasibility of the model, a certain plateau characteristic agricultural demonstration area is selected for model validation. The terrain of this area is high in the northwest and low in the southeast, showing a stepped shape sloping towards the southeast. Mountainous areas and alpine mountainous areas account for 87.5% of the total area. The altitude ranges from 1445m to 3294m, and the forest coverage rate is 58.3%. Agriculture develops well in most areas, and the cultivated land area is 6087.9 km 2First, a sample library is constructed by combining visual interpretation and field sampling to ensure the representativeness and accuracy of the samples, providing high-quality sample data for model training. Specifically, different processing methods provided by the present invention are adopted for different types of multi-source remote sensing data. Specifically: 1) Sentinel-2 data: To eliminate the influence of clouds on the image, the Cloud Score+ algorithm is used for cloud removal. In addition, the bands with a resolution of 20m are uniformly resampled to 10m to improve the spatial resolution. 2) Radar data: First, radiometric calibration is performed to make the data closer to the true surface scattering characteristics. Then, the Lee despeckling algorithm is used to remove speckle noise, thereby improving the image quality. In addition, geometric correction is performed on the image in combination with terrain information to reduce the distortion caused by terrain undulation and ensure the spatial accuracy of the image in mountainous and hilly areas. 3) Slope and Aspect data: These data are calculated by using DEM through the GEE platform. 4) Land use / cover data: To ensure the consistency of the classification system, the present invention reclassifies three datasets of ESA, ESRI, and CRLC, and combines field data. By using the spatial overlay analysis method and visual interpretation method, based on the spatial distribution consistency of various data, sample points are randomly and evenly selected. Finally, a total of 4000 sample points are selected by this method, including 2000 cultivated land samples, 500 forest samples, 500 grassland samples, 500 water body samples, and 500 impervious surface samples. 70% of all sample points are used for model training, and the remaining 30% of the sample points are used for model verification.

[0064] The non-dominated sorting genetic algorithm II (NSGA-II) is used to optimize the hyperparameters and feature combinations of the XGBoost classification model to improve the classification accuracy and generalization ability of the model. During the optimization process, NSGA-II optimizes multiple objective functions, including classification accuracy (Accuracy), F1-score (F1-Score), recall (Recall), and Kappa coefficient (Kappa), ensuring that the model can achieve optimal performance in multiple aspects. In addition, through the multi-objective optimization framework, while ensuring the classification accuracy, overfitting is avoided and the practical application ability of the model is improved. Specifically, the types and search ranges of the optimized hyperparameters are shown in Table 2:

[0065] Table 2 Types and search ranges of hyperparameters

[0066] Learning rate 0.01-0.3 Depth of tree 3-15 Column sampling ratio 0.01-0.99 Regularization parameter L1(0 - 1); L2(0 - 10) Number of trees 50-500

[0067] Based on the optimized XGBoost model, high-precision identification of cultivated land in plateau mountainous areas is carried out. The optimized XGBoost model is optimized by NSGA-II, which can accurately identify the types of cultivated land and distinguish them from other types (such as grasslands, forests, etc.). Evaluation indicators such as Kappa coefficient, overall accuracy (OA), recall rate (Recall), and F1 score are used to evaluate the accuracy of the cultivated land identification results. Through these indicators, the classification effect of the model under the complex terrain of plateau mountainous areas can be comprehensively evaluated. The specific results are shown in Table 3 as follows:

[0068] Table 3 Classification Results

[0069]

[0070] As can be seen from Table 3, the NSGA-I-XGBoost model is superior to other models in terms of classification accuracy, Kappa coefficient, F1 score, and recall rate. Especially in areas with similar spectral features and blurred boundaries, the classification accuracy of the optimization method has been significantly improved, verifying the advantages of the NSGA-11-XGBoost model under complex terrain. At the same time, in order to further verify the feasibility of the optimization algorithm, the present invention also uses XGBoost for cultivated land identification. Finally, the results of the two algorithms are compared with the data of the third national land survey, as Figure 4 shown in A, B, and C. The cultivated land identification effect of XGBoost optimized by the model of the present invention is closer to the data of the third national land survey, which shows the feasibility of the optimization algorithm. Specifically, the number of cultivated land grid cells identified by XGBoost is 1,518,535, while the number of cultivated land grid cells identified by NSGA-II-XGBoost is 14,819,179. Although the number of cultivated land identified by XGBoost is larger, as can be seen from Figure 4 Figure D, when XGBoost identifies large continuous cultivated land, it often appears fragmented, resulting in discontinuous cultivated land. In contrast, the identification effect of NSGA-II-XGBoost is more coherent and can better identify continuous cultivated land areas. In addition, XGBoost will also identify more non-cultivated land as cultivated land, while NSGA-I1-XGBoost performs more accurately in this regard and can effectively reduce misidentification. This shows that the optimization algorithm not only improves the classification accuracy but also can better handle complex terrain and the identification problems of different types of land use.

[0071] The present invention is not only applicable to plateau mountainous areas, but also can be extended to other areas with complex terrains. Since the method combines multiple remote sensing data and machine learning techniques, it can cope with the monitoring of land use changes and the identification of arable land under different terrains and climate conditions. Whether it is mountainous areas, hilly areas, or other areas with complex terrains, relatively accurate arable land identification results can be obtained through this method. In addition, this technology also has important significance for the protection and management of arable land. Through high-precision land use classification and long-term monitoring, it can provide reliable data support for land management departments to help them better carry out land use planning, arable land protection, and resource management. Especially in the identification of abandoned arable land and the assessment of arable land quality, the present invention provides a more scientific analysis method.

[0072] The preferred embodiments of the present invention have been described above with reference to the accompanying drawings. Those skilled in the art can implement the present invention in various variant schemes without departing from the scope and essence of the present invention. For example, the features shown or described as part of one embodiment can be used in another embodiment to obtain another embodiment. The above are only the preferred and feasible embodiments of the present invention, and thus do not limit the scope of the rights of the present invention. Any equivalent changes made by using the content of the specification and drawings of the present invention are included within the scope of the rights of the present invention.

Claims

1. A high-precision identification method for cultivated land in plateau mountainous areas based on NSGA-II-XGBoost machine learning, characterized by: The specific steps include: S1. Acquire multi-source remote sensing data and pre-process the multi-source remote sensing data; S2. Preliminary classification of the data obtained in step S1 is performed through the GEE platform, and texture feature extraction technology is used to enhance classification accuracy; S3. Fusion of the data obtained in step S2 based on terrain features and texture features to generate a multi-source feature data set; S4. Use SHAP to interpret the XGBoost model, evaluate the contribution of each feature of terrain and texture features to the results of cultivated land identification, eliminate irrelevant or redundant features, and thus optimize the input features of the model; S5. The non-dominated sorting genetic algorithm II is used to optimize the hyperparameters and feature combinations of the XGBoost classification model to obtain the NSGA-II-XGBoost model.

2. The high-precision identification method of cultivated land in plateau mountainous areas based on NSGA-II-XGBoost machine learning according to claim 1 is characterized in that: The multi-source remote sensing data includes spectral data, radar data, and elevation data.

3. The high-precision identification method of cultivated land in plateau mountainous areas based on NSGA-II-XGBoost machine learning according to claim 2 is characterized in that: The spectral data preprocessing includes declouding and resampling by using the Cloud Score+ algorithm of GEE, as well as radiation correction, geometric correction and atmospheric correction.

4. The high-precision identification method of cultivated land in plateau mountainous areas based on NSGA-II-XGBoost machine learning according to claim 2 is characterized in that: The radar data preprocessing includes at least one of denoising, geometric correction, radiation calibration and terrain correction, and removing speckle noise using a Lee despeckle algorithm.

5. The high-precision identification method of cultivated land in plateau mountainous areas based on NSGA-II-XGBoost machine learning according to claim 2 is characterized in that: The preprocessing of the elevation data includes calculating the slope and the slope direction according to the DEM.

6. The high-precision identification method of cultivated land in plateau mountainous areas based on NSGA-II-XGBoost machine learning according to claim 2 is characterized in that: The spectral indexes of the spectral data include a normalized vegetation index, an enhanced vegetation index, a normalized water index and a bare soil index.

7. The high-precision identification method of cultivated land in plateau mountainous areas based on NSGA-II-XGBoost machine learning according to claim 1 is characterized in that: The texture feature extraction technique includes a gray level co-occurrence matrix.

8. The high-precision identification method of cultivated land in plateau mountainous areas based on NSGA-II-XGBoost machine learning according to claim 1 is characterized in that: The objective functions optimized by the non-dominated sorting genetic algorithm II include classification accuracy, F1 score, recall rate and Kappa coefficient.

9. A computer device comprising: A memory, a processor, and a computer program stored in the memory and executable on the processor, wherein the processor executes the computer program to implement the steps of a high-precision identification method for cultivated land in plateau mountainous areas based on NSGA-II-XGBoost machine learning as described in any one of claims 1 to 8.

10. A computer-readable storage medium having a computer program stored thereon, characterized in that: When the computer program is executed by a processor, the steps of the high-precision identification method of cultivated land in plateau mountainous areas based on NSGA-II-XGBoost machine learning are implemented as described in any one of claims 1 to 8.