A scientific greening vegetation species recommendation method based on random forest machine learning

By constructing a vegetation species recommendation model based on random forest machine learning, the problem of unstable vegetation species recommendation was solved, and accurate learning and transparent decision-making on the relationship between vegetation species and site environment were achieved, thereby improving the stability of afforestation survival rate and effectiveness.

CN122309864APending Publication Date: 2026-06-30RES INST OF FOREST RESOURCE INFORMATION TECHN CHINESE ACADEMY OF FORESTRY
View PDF 0 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
RES INST OF FOREST RESOURCE INFORMATION TECHN CHINESE ACADEMY OF FORESTRY
Filing Date
2026-04-02
Publication Date
2026-06-30

AI Technical Summary

Technical Problem

Existing methods for recommending vegetation species rely on expert experience, making it difficult to guarantee the consistency and objectivity of the judgment results. Furthermore, it is difficult to quantify the interaction between different environmental factors, resulting in low survival rates of afforestation in some areas and unstable afforestation effectiveness.

Method used

A random forest-based machine learning approach was adopted. By acquiring field survey data of afforestation patches, a patch attribute database was constructed, feature parameterization and model training were performed, and a vegetation species recommendation model was built using the random forest algorithm. Multi-scale spatial independence verification and interpretability analysis were conducted to optimize the recommendation model.

Benefits of technology

It achieves accurate learning of the complex nonlinear relationship between vegetation species and site environment, provides transparent and verifiable decision-making basis, improves the stability and effectiveness of afforestation survival rate, and ensures the applicability of the model in different geographical regions and the transparency of recommendation results.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN122309864A_ABST
    Figure CN122309864A_ABST
Patent Text Reader

Abstract

This invention relates to the field of scientific afforestation and discloses a method for recommending vegetation species based on random forest machine learning. The method includes the following steps: acquiring field survey data of afforestation plots; parameterizing the survey data in the plot attribute database based on plant ecology principles; constructing a vegetation species recommendation model using a random forest algorithm; performing multi-scale spatial independence verification and interpretability analysis on the vegetation species recommendation model; optimizing the vegetation species recommendation model based on the analysis results; and outputting recommended vegetation species and corresponding decision-making criteria. By parameterizing the field survey data of afforestation plots to construct a feature set, constructing a vegetation species recommendation model using a random forest algorithm, and performing multi-scale spatial independence verification and interpretability analysis on the model, the recommendation model is optimized based on the analysis results. This achieves accurate learning of the complex nonlinear relationship between vegetation species and the site environment.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of scientific greening technology, specifically to a method for recommending scientific greening vegetation species based on random forest machine learning. Background Technology

[0002] Scientific afforestation is an important way to promote ecological civilization and achieve carbon peaking and carbon neutrality, and it is also a core component of national land space ecological restoration. The principle of "planting the right tree in the right place" is fundamental to scientific afforestation. It requires selecting suitable vegetation species based on site conditions to ensure afforestation survival rates and ecological benefits. Vegetation species recommendations involve a comprehensive assessment of multi-dimensional environmental factors such as landforms, topography, and soil. These factors have complex non-linear relationships with vegetation growth, and different factors interact and couple with each other.

[0003] Currently, vegetation species recommendations are mainly based on expert experience and manual judgment. Forestry technicians select suitable vegetation species through on-site surveys and personal experience. However, different technicians may reach different recommendations, making it difficult to guarantee the consistency and objectivity of the judgment results. Furthermore, it is difficult to quantify the interaction between different environmental factors. When multiple factors act simultaneously, judgments are often made based on intuition rather than quantitative analysis. This makes it difficult to comprehensively consider the complex coupling relationships of multidimensional factors such as landforms, soil, and topography. For plots with special environmental conditions or complex combinations of multiple factors, judgment biases are prone to occur, which can easily lead to low afforestation survival rates and unstable afforestation results in some plots. Summary of the Invention

[0004] To address the shortcomings of existing technologies, this invention provides a scientific method for recommending vegetation species for afforestation based on random forest machine learning. This method solves the problem that existing vegetation species recommendations often result in low survival rates and unstable afforestation outcomes in some areas.

[0005] To achieve the above objectives, the present invention provides the following technical solution: a scientific method for recommending vegetation types for afforestation based on random forest machine learning, comprising the following steps:

[0006] Obtain field survey data of afforestation plots, standardize the survey data, and construct a plot attribute database. The survey data includes geomorphic factors, soil factors, and vegetation factors.

[0007] Based on the principles of plant ecology, the survey data in the image patch attribute database are parameterized to construct a feature set for model training. The feature set includes original features and derived features.

[0008] Using the feature set as input variables and vegetation type as target variables, a vegetation type recommendation model is constructed using the random forest algorithm. The vegetation type recommendation model is then subjected to multi-scale spatial independence verification and interpretability analysis. Based on the analysis results, the vegetation type recommendation model is optimized to obtain the recommendation model.

[0009] The survey data of the plot to be predicted is input into the recommendation model, which outputs recommended vegetation types and corresponding decision-making criteria.

[0010] By adopting the above technical solution, feature sets are constructed by parameterizing the field survey data of afforestation plots. A vegetation type recommendation model is built using the random forest algorithm. The model is then subjected to multi-scale spatial independence verification and interpretability analysis. Based on the analysis results, the recommendation model is optimized, thereby achieving accurate learning of the complex nonlinear relationship between vegetation types and site environment. It also provides transparent and verifiable decision-making basis, solving the problem that existing vegetation type recommendations easily lead to low survival rates and unstable afforestation results in some plots.

[0011] Preferably, the construction of the map patch attribute database includes the following steps:

[0012] Conduct field surveys of suitable afforestation sites, already afforested sites, and degraded forest sites;

[0013] Record the landform type, topographic location, elevation, slope, aspect, and slope position of each map patch as landform and topographic factors;

[0014] Record the soil type, soil texture, soil layer thickness, and soil pH value for each patch as soil factors;

[0015] The main vegetation species, vegetation density, tree age, afforestation method, and afforestation effectiveness assessment of each patch are recorded as vegetation factors. The afforestation effectiveness assessment includes first, second, third, and fourth levels.

[0016] Add time stamps and spatial location information to each map feature, and enter all data into the database in a unified format to obtain the map feature attribute database.

[0017] Preferably, constructing the feature set for model training includes the following steps:

[0018] The landform type, topographic location, altitude, slope, aspect, slope position, soil type, soil texture, soil layer thickness, and soil pH value in the survey data are used as the original features.

[0019] Based on the characteristics of light factor derived from slope aspect, samples corresponding to sunny slopes, semi-sunny slopes, and semi-shaded slopes are labeled as light-loving, while samples corresponding to shady slopes are labeled as shade-loving.

[0020] Based on the soil pH value, acid-base adaptability characteristics are derived, and pH value is divided into a preset number of acid-base adaptability levels. Each level corresponds to a preset pH value range and is assigned a corresponding level code.

[0021] Based on soil type, soil fertility derivative characteristics are generated, and red soil, brown soil, brown soil, black soil, alluvial soil, and paddy soil are marked as fertilizer-loving type, while chestnut soil, desert soil, irrigated alluvial soil, wet soil, saline-alkali soil, lithological soil, and alpine soil are marked as barren-tolerant type.

[0022] Based on the characteristics of root adaptation derived from soil texture, sandy loam, loam, sandy clay, light clay, and medium clay are marked as suitable for deep roots, while extremely heavy sandy soil, heavy sandy soil, medium sandy soil, light sandy soil, sandy silt, silt, heavy clay, and extremely heavy clay are marked as suitable for shallow roots.

[0023] The original features and the derived features are combined to form a feature set for model training.

[0024] Preferably, the step of constructing a vegetation species recommendation model using the random forest algorithm includes the following steps:

[0025] Using the feature set as input variables, the vegetation species corresponding to the first or second afforestation effectiveness evaluation samples are used as positive samples, and the remaining samples are used as negative samples to construct a supervised learning dataset.

[0026] Set the number of trees in the random forest to a preset number, and the number of node split features to the square root of the total number of input features;

[0027] Using the Gini index as the node splitting criterion, a random forest model is trained on the supervised learning dataset to obtain a vegetation type recommendation model.

[0028] Preferably, in the multi-scale spatial independence verification and interpretability analysis of the vegetation species recommendation model, the multi-scale spatial independence verification includes the following steps:

[0029] The study area was divided into regular grids according to various preset scales;

[0030] For each scale, the samples are divided into training and validation sets according to grid cells, so that samples within the same regular grid do not appear in the training and validation sets at the same time.

[0031] Record the sample partitioning results of the training and validation sets at each scale.

[0032] Preferably, in the multi-scale spatial independence verification and interpretability analysis of the vegetation species recommendation model, the interpretability analysis includes the following steps:

[0033] The TreeSHAP algorithm is used to calculate the Shapley value of each feature in each sample;

[0034] The mean absolute Shapley value of each feature is calculated based on the Shapley values ​​of all samples, and is used as the global feature importance.

[0035] For a single sample, a waterfall plot of Shapley values ​​for each feature is generated, showing the direction and magnitude of each feature's contribution to the prediction result.

[0036] Preferably, the interpretability analysis further includes a multi-scale stability test, specifically comprising the following steps:

[0037] For each vegetation type recommendation model trained at each preset scale, the average Shapley value of each feature at each scale is calculated.

[0038] Calculate the coefficient of variation of the average Shapley value for each feature at different scales;

[0039] Features with a coefficient of variation less than a preset threshold are selected as cross-scale stable features.

[0040] Preferably, optimizing the vegetation species recommendation model based on the analysis results includes the following steps:

[0041] Based on the analysis results, the cross-scale stable features are used as the current feature set;

[0042] Select the spatial scale with the highest validation accuracy, and redivide the training set and validation set according to that spatial scale;

[0043] Based on the newly partitioned training and validation sets, a random forest model is trained on the current feature set, and the global importance of each feature is calculated.

[0044] Features with global importance below a preset importance threshold are removed to obtain the pruned feature set;

[0045] Repeat training, importance calculation and feature pruning on the pruned feature set until the verification accuracy no longer improves or the feature set remains unchanged for a set number of consecutive iterations;

[0046] The model with the highest validation accuracy is selected as the recommended model.

[0047] Preferably, the output of recommended vegetation types and corresponding decision-making criteria includes the following steps:

[0048] The survey data of the plot to be predicted is parameterized to obtain the feature vector to be predicted;

[0049] Input the feature vector to be predicted into the recommendation model, and output the recommended vegetation type and the corresponding prediction confidence.

[0050] Calculate the local Shapley value of each feature in the feature vector to be predicted, and output the key driving factors in order of absolute value.

[0051] The contribution direction of the local Shapley value is matched with the ecological preference of the corresponding vegetation type in the built-in plant ecological characteristic database, and the recommended vegetation type and the corresponding decision basis are output. When the contribution direction contradicts the ecological preference, an early warning prompt is output.

[0052] This invention provides a scientific method for recommending vegetation types for greening based on random forest machine learning. It has the following beneficial effects:

[0053] 1. This invention constructs a feature set by parameterizing the field survey data of afforestation plots, builds a vegetation type recommendation model using the random forest algorithm, and performs multi-scale spatial independence verification and interpretability analysis on the model. Based on the analysis results, the recommendation model is optimized, thereby realizing the accurate learning of the complex nonlinear relationship between vegetation type and site environment, and providing a transparent and verifiable decision basis. This solves the problem that existing vegetation type recommendations are prone to causing low survival rates of afforestation in some plots and unstable afforestation results.

[0054] 2. This invention verifies the spatial independence of the vegetation species recommendation model at multiple scales by dividing the training and validation sets at different spatial scales. This ensures that the model learns real ecological relationships rather than spurious correlations caused by spatial proximity. Furthermore, it combines multi-scale stability tests to screen cross-scale stable features, thereby improving the model's generalization ability and applicability in different geographical regions, and ensuring the stability of afforestation results.

[0055] 3. This invention parameterizes survey data based on plant ecology principles, constructing a feature set containing original and derived features. This enables the model to more accurately learn the ecological relationship between vegetation species and the site environment. Interpretability analysis then outputs the decision-making basis and key driving factors for each recommendation result, achieving a shift from experience-based judgment to data-driven approaches. Furthermore, by iteratively optimizing the vegetation species recommendation model based on the analysis results, using cross-scale stable features as a foundation, and combining feature importance pruning and validation accuracy optimization, a high-precision and concise recommendation model is obtained. Simultaneously, the contribution direction of local Shapley values ​​is matched and validated with the built-in plant ecological characteristic database. When outputting recommendation results, prediction confidence and ecological validation conclusions are provided, achieving transparency and verifiability in the recommendation process. Attached Figure Description

[0056] Figure 1 This is a diagram illustrating the survey situation of the Tuban proposed in an embodiment of the present invention.

[0057] Figure 2 This is a database map of afforestation plot attributes proposed in an embodiment of the present invention.

[0058] Figure 3 This is a database map of suitable afforestation plots proposed in an embodiment of the present invention.

[0059] Figure 4 This is a ranking diagram of the importance of land parcel attribute features proposed in an embodiment of the present invention.

[0060] Figure 5 This is a precision graph of the machine learning model proposed in an embodiment of the present invention;

[0061] Figure 6 This is a flowchart of a scientific vegetation species recommendation method based on random forest machine learning proposed in this invention;

[0062] Figure 7 This is an architecture diagram of a scientific greening vegetation type recommendation system based on random forest machine learning proposed in an embodiment of the present invention. Detailed Implementation

[0063] The technical solution of the present invention will now be clearly and completely described with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of the present invention, and not all embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of the present invention.

[0064] Example 1:

[0065] In a first embodiment of the present invention, the present invention provides a scientific method for recommending vegetation species for greening based on random forest machine learning, such as... Figure 6 As shown, it includes the following steps:

[0066] Obtain field survey data of afforestation plots, standardize the survey data, and construct a plot attribute database. The survey data includes geomorphological factors, soil factors, and vegetation factors.

[0067] Furthermore, constructing a map feature attribute database includes the following steps:

[0068] Conduct field surveys of suitable afforestation sites, already afforested sites, and degraded forest sites;

[0069] Record the landform type, topographic location, elevation, slope, aspect, and slope position of each map patch as landform and topographic factors;

[0070] Record the soil type, soil texture, soil layer thickness, and soil pH value for each patch as soil factors;

[0071] Record the main vegetation species, vegetation density, tree age, afforestation method, and afforestation effectiveness assessment for each patch as vegetation factors. The afforestation effectiveness assessment includes first, second, third, and fourth levels.

[0072] Add time stamps and spatial location information to each map feature, and enter all data into the database in a unified format to obtain the map feature attribute database.

[0073] Specifically, firstly, in step one, obtaining field survey data of afforestation patches is the foundation for building a vegetation type recommendation model. The purpose is to collect sufficient and high-quality sample data to provide reliable data support for subsequent feature parameterization and model training. Through systematic surveys of patches of different afforestation types, a patch attribute database containing multi-dimensional environmental factors and vegetation response information is established, thereby transforming the raw data obtained from the field survey into a structured data set that can be used for model learning.

[0074] In this embodiment, the field survey of afforestation sites covers three types: suitable afforestation sites, already afforested sites, and degraded forest sites. Data collection for each type of site must be conducted according to unified survey specifications to ensure data comparability and consistency. For each site, the following three types of factors must be recorded as survey data: geomorphological factors, including geomorphological type, topographic location, altitude, slope, aspect, and slope position; soil factors, including soil type, soil texture, soil layer thickness, and soil pH; and vegetation factors, including main vegetation species, vegetation density, tree age, afforestation method, and afforestation effectiveness assessment. The afforestation effectiveness assessment is divided into four levels according to preset standards: Level 1, Level 2, Level 3, and Level 4, corresponding to the degree of excellence or inferiority of the afforestation effect.

[0075] To facilitate subsequent data processing and model training, the collected survey data needs to be standardized. Time stamps and spatial location information are added to each map patch, and all survey data are entered into a database system according to a unified data format, thus constructing a map patch attribute database. Each map patch sample in this database corresponds to a complete feature vector, containing the specific values ​​of all the aforementioned survey factors, providing fundamental data support for subsequent feature parameterization and model construction.

[0076] Based on the principles of plant ecology, the survey data in the map attribute database is parameterized to construct a feature set for model training. The feature set includes original features and derived features.

[0077] Furthermore, constructing a feature set for model training includes the following steps:

[0078] The landform type, topographic location, altitude, slope, aspect, slope position, soil type, soil texture, soil layer thickness, and soil pH value in the survey data are used as the original features.

[0079] Based on the characteristics of light factor derived from slope aspect, samples corresponding to sunny slopes, semi-sunny slopes, and semi-shaded slopes are labeled as light-loving, while samples corresponding to shady slopes are labeled as shade-loving.

[0080] Based on the soil pH value, acid-base adaptability characteristics are derived, and pH value is divided into a preset number of acid-base adaptability levels. Each level corresponds to a preset pH value range and is assigned a corresponding level code.

[0081] Based on soil type, soil fertility derivative characteristics are generated, and red soil, brown soil, brown soil, black soil, alluvial soil, and paddy soil are marked as fertilizer-loving type, while chestnut soil, desert soil, irrigated alluvial soil, wet soil, saline-alkali soil, lithological soil, and alpine soil are marked as barren-tolerant type.

[0082] Based on the characteristics of root adaptation derived from soil texture, sandy loam, loam, sandy clay, light clay, and medium clay are marked as suitable for deep roots, while extremely heavy sandy soil, heavy sandy soil, medium sandy soil, light sandy soil, sandy silt, silt, heavy clay, and extremely heavy clay are marked as suitable for shallow roots.

[0083] The original features and derived features are combined to form a feature set for model training.

[0084] Specifically, in step two, the survey data in the map attribute database is parameterized based on the principles of plant ecology. The aim is to transform the original survey data into numerical features that can characterize the ecological relationship between vegetation growth and environmental factors, thereby constructing a feature set suitable for training the random forest model. By introducing domain knowledge to encode and derive the original features, the model can learn the inherent laws between vegetation types and site conditions more effectively.

[0085] In this embodiment, the construction of the feature set for model training first requires determining the original features. The landform type, topographic location, altitude, slope, aspect, slope position, soil type, soil texture, soil layer thickness, and soil pH value in the survey data obtained in step one are directly used as the original features. These features constitute the basic variables describing the site conditions.

[0086] Next, based on the principles of plant ecology, derived features are generated to enhance the model's learning ability to adapt to ecological conditions. Based on slope aspect, light factor derived features are generated: samples corresponding to sunny slopes, semi-sunny slopes, and semi-shaded slopes are marked as light-loving, and samples corresponding to shady slopes are marked as shade-loving. This treatment is based on the differences in solar radiation received by different slope aspects, reflecting the vegetation's preference for light conditions.

[0087] Based on the acid-base adaptability characteristics derived from soil pH, pH values ​​are divided into a preset number of acid-base adaptability levels. Each level corresponds to a preset pH range and is assigned a corresponding level code. For example, samples with a pH value less than 6.5 are classified as acidic, samples with a pH value between 6.5 and 7.5 are classified as neutral, and so on, thus quantifying the restrictive effect of soil acidity and alkalinity on vegetation growth.

[0088] Based on soil type, soil fertility characteristics are derived. Soil types such as red soil, brown soil, brown soil, black soil, alluvial soil, and paddy soil have high natural fertility and are suitable for the growth of fertilizer-loving vegetation. Therefore, the samples corresponding to these soil types are marked as fertilizer-loving. On the other hand, soil types such as chestnut soil, desert soil, irrigated alluvial soil, wet soil, saline-alkali soil, lithological soil, and alpine soil have low fertility or have limiting factors and are suitable for the growth of barren-tolerant vegetation. Therefore, they are marked as barren-tolerant.

[0089] Based on the characteristics of root adaptation derived from soil texture, sandy loam, loam, sandy clay, light clay, and medium clay are relatively loose textures, which are conducive to deep root growth, and are therefore marked as suitable for deep roots; while very heavy sandy soil, heavy sandy soil, medium sandy soil, light sandy soil, sandy silt, silt, heavy clay, and very heavy clay are too sandy or too clayey, which restricts root growth, and are therefore marked as suitable for shallow roots.

[0090] After generating the above-mentioned derived features, the original features are merged with all derived features to form a feature set for model training. This feature set not only retains the original survey information but also incorporates plant ecology knowledge, which can more comprehensively depict the intrinsic relationship between vegetation and site conditions, providing input data for the subsequent learning of the random forest model. This transforms the originally scattered original survey data into numerical features with clear ecological significance, thereby improving the model's ability to learn the laws of vegetation suitability.

[0091] Using feature set as input variable and vegetation type as target variable, a vegetation type recommendation model is constructed using random forest algorithm. The vegetation type recommendation model is then subjected to multi-scale spatial independence verification and interpretability analysis. Based on the analysis results, the vegetation type recommendation model is optimized to obtain the recommendation model.

[0092] Furthermore, a vegetation type recommendation model is constructed using the random forest algorithm, including the following steps:

[0093] Using the feature set as the input variable, the vegetation species corresponding to the first or second sample in the afforestation effectiveness assessment are used as positive samples, and the remaining samples are used as negative samples to construct a supervised learning dataset.

[0094] Set the number of trees in the random forest to a preset number, and the number of node split features to the square root of the total number of input features;

[0095] Using the Gini index as the node splitting criterion, a random forest model was trained on a supervised learning dataset to obtain a vegetation type recommendation model.

[0096] Furthermore, in the multi-scale spatial independence verification and interpretability analysis of the vegetation species recommendation model, the multi-scale spatial independence verification includes the following steps:

[0097] The study area was divided into regular grids according to various preset scales;

[0098] For each scale, the samples are divided into training and validation sets according to grid cells, so that samples within the same regular grid do not appear in the training and validation sets at the same time.

[0099] Record the sample partitioning results of the training and validation sets at each scale.

[0100] Furthermore, in the multi-scale spatial independence verification and interpretability analysis of the vegetation species recommendation model, the interpretability analysis includes the following steps:

[0101] The TreeSHAP algorithm is used to calculate the Shapley value of each feature in each sample;

[0102] The mean absolute Shapley value of each feature is calculated based on the Shapley values ​​of all samples, and is used as the global feature importance.

[0103] For a single sample, a waterfall plot of Shapley values ​​for each feature is generated, showing the direction and magnitude of each feature's contribution to the prediction result.

[0104] Furthermore, interpretability analysis also includes multi-scale stability testing, specifically comprising the following steps:

[0105] For each vegetation type recommendation model trained at each preset scale, the average Shapley value of each feature at each scale is calculated.

[0106] Calculate the coefficient of variation of the average Shapley value for each feature at different scales;

[0107] Features with a coefficient of variation less than a preset threshold are selected as cross-scale stable features.

[0108] Furthermore, the vegetation species recommendation model is optimized based on the analysis results, including the following steps:

[0109] Based on the analysis results, cross-scale stable features are used as the current feature set;

[0110] Select the spatial scale with the highest validation accuracy, and redivide the training set and validation set according to that spatial scale;

[0111] Based on the newly partitioned training and validation sets, a random forest model is trained on the current feature set, and the global importance of each feature is calculated.

[0112] Features with global importance below a preset importance threshold are removed to obtain the pruned feature set;

[0113] Repeat training, importance calculation and feature pruning on the pruned feature set until the verification accuracy no longer improves or the feature set remains unchanged for a set number of consecutive iterations;

[0114] The model with the highest validation accuracy is selected as the recommended model.

[0115] Specifically, in step three, the feature set constructed in step two is used as the input variable, and vegetation type is used as the target variable. A vegetation type recommendation model is constructed using the random forest algorithm. The model is then subjected to multi-scale spatial independence verification and interpretability analysis. Finally, the model is optimized based on the analysis results to obtain the recommendation model. The aim is to learn the nonlinear relationship between vegetation type and site environment from the data through machine learning methods, while ensuring that the model has good generalization ability and interpretability, thereby providing a scientific and reliable decision-making basis for subsequent vegetation recommendations.

[0116] In this embodiment, the first step in constructing a vegetation type recommendation model using the random forest algorithm is to build a supervised learning dataset. The feature set obtained in step two is used as the input variable X, and the vegetation type is used as the target variable Y. However, not all samples are suitable for model training. Only samples with good afforestation results can reflect the correct vegetation adaptability relationship. Therefore, the vegetation types corresponding to the samples whose afforestation results are evaluated as the first or second level are used as positive samples, and the remaining samples are used as negative samples, thus forming a supervised learning dataset for model training.

[0117] During model training, the number of trees in the random forest is set to a preset number, such as 500. The number of split features for each node is typically the square root of the total number of input features. The Gini coefficient is used as the node splitting criterion for each decision tree. For any node... If it contains For samples of each category, the formula for calculating the Gini index is: ,in, Represents a node The middle sample belongs to the first The proportion of vegetation types reflects the impurity of nodes. When a node splits, the feature that causes the largest decrease in the Gini index after the split and its split point are selected for partitioning. By integrating the voting results of all decision trees, the final vegetation type recommendation model is obtained.

[0118] To evaluate the model's generalization ability and avoid overfitting caused by spatial autocorrelation, multi-scale spatial independence validation of the model is necessary. Specifically, the study area is divided into regular grids according to several preset scales, such as 100 meters, 500 meters, and 1000 meters. For each scale, samples are divided into training and validation sets according to grid cells, ensuring that samples within the same regular grid do not appear in both the training and validation sets simultaneously. The sample partitioning results for the training and validation sets at each scale are recorded, and vegetation species recommendation models at each scale are trained based on the data from each scale, with their validation accuracy recorded.

[0119] Regarding model interpretability, the TreeSHAP algorithm is used to calculate the Shapley value of each feature in each sample. The Shapley value, derived from game theory, is used to fairly allocate the contribution of each feature to the prediction result. For the , A sample, its characteristics Shapley value satisfy: ,in, This is the model prediction value for this sample. The mean of the predicted values ​​for all samples. The formula decomposes the predicted value into the sum of the baseline value and the contributions of each feature, thus explaining the marginal impact of each feature on the prediction result. Based on the Shapley values ​​of all samples, the average absolute Shapley value of each feature is calculated as the global feature importance, measuring the average contribution of the feature in the entire model. For a single sample, a waterfall plot of the Shapley values ​​of each feature is generated, visually showing the direction and magnitude in which each feature pushes the predicted value from the baseline value to the final prediction result.

[0120] To identify stable features insensitive to spatial scale, a multi-scale stability test was further conducted. For the vegetation species recommendation model trained at each preset scale, the average Shapley value of each feature at each scale was calculated, denoted as . ,in The scale is then used to represent the feature. The coefficient of variation of the average Shapley value for each feature at different scales is then calculated. : ,in The coefficient of variation reflects the degree of fluctuation in feature contributions at different spatial scales. The smaller the fluctuation, the more stable the feature. Features with a coefficient of variation less than a preset threshold are selected as cross-scale stable features. These features are considered to be the core ecological driving factors of vegetation distribution in the study area.

[0121] Based on the above analysis results, the vegetation type recommendation model is optimized. First, stable features across scales are used as the current feature set. Then, the spatial scale with the highest validation accuracy is selected, and the training and validation sets are re-divided according to this scale. A random forest model is trained on the current feature set based on the re-divided data, and the global importance of each feature is calculated. Features with global importance lower than a preset importance threshold are removed to obtain a pruned feature set. Training, importance calculation, and feature pruning are repeated on the pruned feature set until the validation accuracy no longer improves or the feature set remains unchanged for a preset number of iterations. Finally, the model with the highest validation accuracy is used as the recommendation model.

[0122] Through the above multi-scale verification, interpretability analysis and iterative optimization, the final recommendation model not only has high prediction accuracy, but also provides decision-making basis for each prediction result. At the same time, it selects stable ecological characteristics that are not sensitive to spatial scale, providing technical support for scientific greening and vegetation configuration.

[0123] After the random forest model is trained, for an input feature vector X, the predicted vegetation type output by the model is determined by the voting results of all decision trees. The voting rule is as follows: ,in Indicates the first Decision trees for input The prediction results Indicates taking the mode. The total number of decision trees in the random forest is represented by the voting rule, which integrates the judgment results of each decision tree, reduces the risk of overfitting of a single tree, and improves the generalization ability of the model.

[0124] Input the survey data of the plot to be predicted into the recommendation model, and output the recommended vegetation types and corresponding decision-making basis.

[0125] Furthermore, the recommended vegetation types and corresponding decision-making criteria are output through the following steps:

[0126] The survey data of the plot to be predicted is parameterized to obtain the feature vector to be predicted;

[0127] Input the feature vector to be predicted into the recommendation model, and output the recommended vegetation type and the corresponding prediction confidence.

[0128] Calculate the local Shapley value of each feature in the feature vector to be predicted, and output the key driving factors in order of absolute value.

[0129] The contribution direction of local Shapley values ​​is matched with the ecological preferences of corresponding vegetation types in the built-in plant ecological characteristic database. Recommended vegetation types and corresponding decision-making basis are output, and a warning is output when the contribution direction contradicts the ecological preference.

[0130] Specifically, in step four, the survey data of the plot to be predicted is input into the recommendation model, and the recommended vegetation types and corresponding decision-making basis are output. The aim is to apply the trained model to the vegetation recommendation of actual afforestation plots and provide explainable reasons for the recommendations, thereby enhancing the credibility and practicality of the recommendation results. Through feature parameterization, model prediction, local interpretability analysis and ecological knowledge matching of the plot to be predicted, the final output is a comprehensive decision-making information including the recommended types, confidence levels, key driving factors and ecological verification results.

[0131] First, the survey data of the plots to be predicted needs to be processed by feature parameterization according to the method in step two to obtain the feature vector to be predicted. The feature vector to be predicted contains the same original features and derived features as the training data, and its dimension is consistent with the feature set used during model training to ensure that the model can correctly identify and process the input data.

[0132] Let the feature vector to be predicted be denoted as . Input the recommendation model obtained from step three optimization, and the model outputs recommended vegetation types. and the corresponding prediction confidence level Prediction confidence reflects the model's certainty about the recommendation outcome, and is typically determined by the proportion of decision trees in the random forest that vote for that vegetation type.

[0133] To understand why the model recommends this vegetation type, it is necessary to calculate the local Shapley values ​​of each feature in the feature vector to be predicted. Based on the TreeSHAP interpreter established in step three, for the sample to be predicted... Its characteristics Local Shapley values The following relationship must be satisfied: ,in, The predicted values ​​of the model for the sample to be predicted are represented by class probabilities or class labels. The baseline value is the mean of the predicted values ​​of all samples in the training set. Features The Shapley value of the prediction result for the sample to be predicted. The formula decomposes the prediction result of the sample to be predicted into the sum of the baseline value and the contribution of each feature, thereby enabling a quantitative assessment of the marginal impact of each feature on the prediction result of the sample.

[0134] For the sample to be predicted, the Shapley value of each feature is calculated. The absolute value of a feature represents its contribution to the prediction result, and the positive or negative sign indicates whether the contribution is positive or negative. Features are sorted in descending order of absolute value, and the top few features are output as key driving factors, along with their respective contribution directions and specific values.

[0135] To further ensure the ecological rationality of the recommendations, the contribution direction of local Shapley values ​​is matched with the ecological preferences of corresponding vegetation species in the built-in plant ecological characteristic database. This database pre-stores the suitable ranges or preference attributes of each vegetation species in terms of light, pH, fertility, and root system. Taking spruce as an example, its ecological preference is to prefer cool and moist conditions and tolerate poor soil. If the model shows that soil pH contributes positively to the recommendations and the pH value of the plot is high, but spruce actually prefers slightly acidic soil, a contradiction may arise. In this case, if the contribution direction of a certain feature contradicts the ecological preference of that vegetation species recorded in the database, an early warning is issued, suggesting that the data or model be verified in conjunction with field conditions. Through this linkage verification mechanism, spurious correlations learned by the model can be effectively avoided, improving the scientificity and reliability of the recommendations.

[0136] Example 2:

[0137] In a second embodiment of the present invention, the present invention provides a scientific vegetation species recommendation system based on random forest machine learning, such as... Figure 7 As shown, it includes the following modules:

[0138] The data acquisition module is used to acquire field survey data of afforestation plots, standardize the survey data, and construct a plot attribute database. The survey data includes geomorphological factors, soil factors, and vegetation factors.

[0139] The feature construction module is used to parameterize the survey data in the map attribute database based on the principles of plant ecology, and to construct a feature set for model training. The feature set includes original features and derived features.

[0140] The model building module is used to construct a vegetation type recommendation model using a feature set as the input variable and vegetation type as the target variable, and to perform multi-scale spatial independence verification and interpretability analysis on the vegetation type recommendation model. Based on the analysis results, the vegetation type recommendation model is optimized to obtain the recommendation model.

[0141] The recommendation output module is used to input the survey data of the plot to be predicted into the recommendation model and output the recommended vegetation types and corresponding decision-making basis.

[0142] In a pilot national land greening project, over 300 afforestation plots were planned. Traditional tree species recommendations relied on expert on-site surveys and experience-based judgments, which struggled to comprehensively consider the complex interactions of multiple factors such as topography, soil, and terrain. This resulted in low survival rates and unstable afforestation outcomes in some areas. To address these issues, a scientific vegetation species recommendation system based on random forest machine learning, as provided in this invention, was adopted. Its architecture is as follows: Figure 7 As shown. The specific implementation process of this system is as follows:

[0143] First, the data acquisition module conducted field surveys on suitable afforestation patches, already afforested patches, and degraded forest patches, recording information such as landform type, topographic location, altitude, slope, aspect, slope position, soil type, soil texture, soil layer thickness, soil pH value, main vegetation species, and afforestation effectiveness for each patch. After adding time and spatial markers to all data, the data was uniformly entered into the database to construct a patch attribute database, providing high-quality basic data support for subsequent modeling.

[0144] Subsequently, the feature construction module parameterizes the original data in the map attribute database based on the principles of plant ecology. It transforms the slope aspect into a light-loving or shade-loving attribute, classifies the soil pH value into acid-base adaptability levels, and transforms the soil type and texture into fertility and root adaptability derived features, forming a feature set containing original features and derived features, enabling the model to learn the ecological relationship between vegetation and the environment more effectively.

[0145] Next, the model building module takes the feature set as input and the vegetation type as the target variable, and uses the random forest algorithm to build the initial model. The training set and validation set are divided through multi-scale spatial independence verification to eliminate the influence of spatial autocorrelation. The TreeSHAP algorithm is used for interpretability analysis and screening of cross-scale stable features. After iterative optimization, features with low importance are removed. Finally, a high-precision and interpretable recommendation model is obtained, which ensures the generalization ability of the model at different spatial scales and the reliability of the recommendation results.

[0146] Finally, the recommendation output module inputs the characteristic data of the plot to be predicted into the recommendation model, outputs the recommended vegetation types, prediction confidence, key driving factors and their contribution directions, and matches and verifies them with the built-in plant ecological characteristic database. When contradictions occur, it automatically issues an early warning, providing a scientific, transparent and verifiable decision-making basis for afforestation planning, and improving the survival rate of afforestation and the effectiveness of ecological restoration.

[0147] Example 3: Please refer to the appendix. Figure 1 -Appendix Figure 5Step 1: Conduct field surveys of afforestation factors for suitable afforestation sites, degraded forest sites, and existing afforestation sites. Referring to the relevant requirements in the "Guidelines for Land and Sea Use Classification in Territorial Spatial Survey, Planning, and Land Use Control," investigate and analyze the geomorphological, topographic, and soil attributes of the sites, and establish a site attribute database. Table 1 shows the attributes and corresponding field information that need to be recorded during the survey of sites; for suitable afforestation sites, items 1-19 are mandatory.

[0148] Table 1: Attributes and Field Information of Surveyed Map Patches

[0149]

[0150]

[0151]

[0152]

[0153] Step Two: When using surveyed map features to conduct AI-based scientific afforestation recommendations, the first step is to establish the correspondence between the surveyed map feature content and various factors to facilitate tree (shrub, grass) species recommendations for machine learning models. Based on plant ecology principles, expert knowledge is used to transform survey data into quantifiable features, constructing a labeled dataset for model training. Parameterized features include: landform type (DIMAOLEIXING), topographic location (DIXINGBUWEI), elevation (HAIBA), slope (PODU), aspect (POXIANG), slope position (POWEI), soil type (TU_LEIXING), soil texture (TU_ZHIDI), soil thickness (TU_HOUDU), and soil pH (TU_PH).

[0154] (1) Regarding the light factor, "prefers light" corresponds to "sunny slope, semi-sunny slope and semi-shady slope" in the "slope aspect" section of the survey form, and "prefers shade" corresponds to "shady slope";

[0155] (2) Based on the soil acidity and alkalinity characteristics, the survey content was classified according to pH level and divided into acidic (pH = 4.0-6.5), neutral (pH = 6.5-7.5), slightly alkaline (pH = 7.5-8.0), moderately alkaline (pH = 8.0-8.5) and severely alkaline (pH > 8.5);

[0156] (3) Soil fertility is determined based on the "soil type" surveyed. Soils that "prefer fertilizer" include: red soil (code 1), brown soil (code 2), brown soil (code 3), black soil (code 4), alluvial soil (including sandy ginger black soil) (code 7), and paddy soil (code 9); Soils that "tolerate poor soil" include: chestnut calcareous soil (code 5), desert soil (code 6), irrigated alluvial soil (code 8), wet soil (meadow, marsh soil) (code 10), saline-alkali soil (code 11), lithological soil (code 12), and alpine soil (code 13).

[0157] (4) For soils suitable for root depth, the determination is based on the surveyed "soil texture". Soil textures suitable for "deep-rooted" plants include: sandy loam (code 7), loam (code 8), sandy clay (code 9), light clay (code 10), and medium clay (code 11); Soil textures suitable for "shallow-rooted" plants include: extremely heavy sandy soil (code 1), heavy sandy soil (code 2), medium sandy soil (code 3), light sandy soil (code 4), sandy silt (code 5), silt (code 6), heavy clay (code 12), and extremely heavy clay (code 13).

[0158] Step 3: The topographic type, terrain location, altitude, slope, aspect, slope position, soil type, soil texture, soil thickness, and soil pH of the surveyed afforestation plots are used as feature variables, and planting effectiveness is used as the target variable. Plots with better afforestation effectiveness are used as positive samples, and the remaining plots are used as negative samples, forming a supervised learning dataset. An optimal vegetation recommendation model is constructed using the random forest machine learning algorithm. Overfitting risk is reduced through multi-decision tree voting, and the model outputs a ranking of feature importance. 70% of the survey samples are used for model training, and 30% are used for validation, ensuring that the training and test sets do not overlap geographically.

[0159] Step 4: Apply the trained model to recommend vegetation species for suitable afforestation plots; by evaluating the accuracy of the machine learning model, use the model with the best generalization ability to make predictions and recommend suitable vegetation species for suitable afforestation plots (unafforested). The optimal machine learning model has a training accuracy better than 0.8 and a validation accuracy better than 0.60. Use this model to carry out machine learning vegetation recommendation to obtain the most suitable vegetation for planting.

[0160] Taking the field survey conducted in Longde County, Guyuan City, Ningxia Hui Autonomous Region as an example, a total of 1,203 afforestation sites, both suitable for afforestation and already afforested, were surveyed from April to May 2024. These included:

[0161] (1) 302 suitable afforestation plots were surveyed, including landform type, topographic location, slope position, soil type, and soil layer thickness;

[0162] (2) 901 afforestation plots have been established. The survey content includes landform type, topographic location, slope position, soil type, soil layer thickness, main vegetation, vegetation density, tree age, afforestation method, afforestation density, degraded forest restoration method, and afforestation effectiveness assessment. Among them, 405 plots have been rated as "good", 91 plots as "medium" and 405 plots as "poor".

[0163] The survey data from 901 afforested plots were parameterized, and the features were ranked by importance. The influence on vegetation type was ranked as follows: slope, altitude, topographic location, soil pH, and soil thickness. 70% of the survey samples were used for model training, and 30% for validation. A machine learning training model for optimal vegetation recommendation was constructed using the above five features. The training model achieved an accuracy of 0.9399, and the validation accuracy was 0.6023, indicating that the model has good generalization ability for vegetation recommendation.

[0164] Taking a suitable afforestation plot (unafforested area) in Longde County as an example, the plot has a longitude of 106.198° and a latitude of 35.659°. Surveys show its landform is mountainous, with a sunny slope and brown soil (20mm thick). Based on a trained optimal vegetation recommendation model, spruce was chosen as the suitable vegetation for this plot. Follow-up investigations and effectiveness verification show that in recent years, the area has optimized its afforestation tree species configuration, abandoning the previous model of using pure poplar forests as the sole afforestation subject. Instead, it has adopted pure spruce forests or mixed plantings of spruce and Pinus sylvestris, significantly improving the survival rate and ecological restoration effectiveness, achieving good ecological and practical benefits. The recommended spruce species for this plot aligns with regional afforestation practices, verifying the reliability and applicability of the vegetation recommendation model.

[0165] Although embodiments of the invention have been shown and described, it will be understood by those skilled in the art that various changes, modifications, substitutions and alterations can be made to these embodiments without departing from the principles and spirit of the invention, the scope of which is defined by the appended claims and their equivalents.

Claims

1. A scientific greening vegetation species recommendation method based on random forest machine learning, characterized by, Includes the following steps: Obtain field survey data of afforestation plots, standardize the survey data, and construct a plot attribute database. The survey data includes geomorphic factors, soil factors, and vegetation factors. Based on the principles of plant ecology, the survey data in the image patch attribute database are parameterized to construct a feature set for model training. The feature set includes original features and derived features. Using the feature set as input variables and vegetation type as target variables, a vegetation type recommendation model is constructed using the random forest algorithm. The vegetation type recommendation model is then subjected to multi-scale spatial independence verification and interpretability analysis. Based on the analysis results, the vegetation type recommendation model is optimized to obtain the recommendation model. The survey data of the plot to be predicted is input into the recommendation model, which outputs recommended vegetation types and corresponding decision-making criteria. 2.The scientific green vegetation species recommendation method based on random forest machine learning according to claim 1, characterized in that: The construction of the map feature attribute database includes the following steps: Conduct field surveys of suitable afforestation sites, already afforested sites, and degraded forest sites; Record the landform type, topographic location, elevation, slope, aspect, and slope position of each map patch as landform and topographic factors; Record the soil type, soil texture, soil layer thickness, and soil pH value for each patch as soil factors; The main vegetation species, vegetation density, tree age, afforestation method, and afforestation effectiveness assessment of each patch are recorded as vegetation factors. The afforestation effectiveness assessment includes first, second, third, and fourth levels. Add time stamps and spatial location information to each map feature, and enter all data into the database in a unified format to obtain the map feature attribute database. 3.The scientific green vegetation species recommendation method based on random forest machine learning according to claim 1, characterized in that: The construction of the feature set for model training includes the following steps: The landform type, topographic location, altitude, slope, aspect, slope position, soil type, soil texture, soil layer thickness, and soil pH value in the survey data are used as the original features. Based on the characteristics of light factor derived from slope aspect, samples corresponding to sunny slopes, semi-sunny slopes, and semi-shaded slopes are labeled as light-loving, while samples corresponding to shady slopes are labeled as shade-loving. Based on the soil pH value, acid-base adaptability characteristics are derived, and pH value is divided into a preset number of acid-base adaptability levels. Each level corresponds to a preset pH value range and is assigned a corresponding level code. Based on soil type, soil fertility derivative characteristics are generated, and red soil, brown soil, brown soil, black soil, alluvial soil, and paddy soil are marked as fertilizer-loving type, while chestnut soil, desert soil, irrigated alluvial soil, wet soil, saline-alkali soil, lithological soil, and alpine soil are marked as barren-tolerant type. Based on the characteristics of root adaptation derived from soil texture, sandy loam, loam, sandy clay, light clay, and medium clay are marked as suitable for deep roots, while extremely heavy sandy soil, heavy sandy soil, medium sandy soil, light sandy soil, sandy silt, silt, heavy clay, and extremely heavy clay are marked as suitable for shallow roots. The original features and the derived features are combined to form a feature set for model training.

4. The scientific green vegetation species recommendation method based on random forest machine learning according to claim 1, characterized in that: The method of constructing a vegetation type recommendation model using the random forest algorithm includes the following steps: Using the feature set as input variables, the vegetation species corresponding to the first or second afforestation effectiveness evaluation samples are used as positive samples, and the remaining samples are used as negative samples to construct a supervised learning dataset. Set the number of trees in the random forest to a preset number, and the number of node split features to the square root of the total number of input features; Using the Gini index as the node splitting criterion, a random forest model is trained on the supervised learning dataset to obtain a vegetation type recommendation model.

5. The method for recommending scientific vegetation species for afforestation based on random forest machine learning according to claim 1, characterized in that: In the multi-scale spatial independence verification and interpretability analysis of the vegetation type recommendation model, the multi-scale spatial independence verification includes the following steps: The study area was divided into regular grids according to various preset scales; For each scale, the samples are divided into training and validation sets according to grid cells, so that samples within the same regular grid do not appear in the training and validation sets at the same time. Record the sample partitioning results of the training and validation sets at each scale.

6. The method for recommending scientific vegetation species for afforestation based on random forest machine learning according to claim 1, characterized in that: In the multi-scale spatial independence verification and interpretability analysis of the vegetation species recommendation model, the interpretability analysis includes the following steps: The TreeSHAP algorithm is used to calculate the Shapley value of each feature in each sample; The mean absolute Shapley value of each feature is calculated based on the Shapley values ​​of all samples, and is used as the global feature importance. For a single sample, a waterfall plot of Shapley values ​​for each feature is generated, showing the direction and magnitude of each feature's contribution to the prediction result.

7. The method for recommending scientific vegetation species for afforestation based on random forest machine learning according to claim 6, characterized in that: The interpretability analysis also includes a multi-scale stability test, specifically comprising the following steps: For each vegetation type recommendation model trained at each preset scale, the average Shapley value of each feature at each scale is calculated. Calculate the coefficient of variation of the average Shapley value for each feature at different scales; Features with a coefficient of variation less than a preset threshold are selected as cross-scale stable features.

8. The method for recommending scientific vegetation species for afforestation based on random forest machine learning according to claim 1, characterized in that: The optimization of the vegetation type recommendation model based on the analysis results includes the following steps: Based on the analysis results, the cross-scale stable features are used as the current feature set; Select the spatial scale with the highest validation accuracy, and redivide the training set and validation set according to that spatial scale; Based on the newly partitioned training and validation sets, a random forest model is trained on the current feature set, and the global importance of each feature is calculated. Features with global importance below a preset importance threshold are removed to obtain the pruned feature set; Repeat training, importance calculation and feature pruning on the pruned feature set until the verification accuracy no longer improves or the feature set remains unchanged for a set number of consecutive iterations; The model with the highest validation accuracy is selected as the recommended model.

9. The method for recommending scientific vegetation species for afforestation based on random forest machine learning according to claim 1, characterized in that: The output of recommended vegetation types and corresponding decision-making criteria includes the following steps: The survey data of the plot to be predicted is parameterized to obtain the feature vector to be predicted; Input the feature vector to be predicted into the recommendation model, and output the recommended vegetation type and the corresponding prediction confidence. Calculate the local Shapley value of each feature in the feature vector to be predicted, and output the key driving factors in order of absolute value. The contribution direction of the local Shapley value is matched with the ecological preference of the corresponding vegetation type in the built-in plant ecological characteristic database, and the recommended vegetation type and the corresponding decision basis are output. When the contribution direction contradicts the ecological preference, an early warning prompt is output.