A method and system for predicting tree planting selection based on big data

By integrating the genomic, phenotype and environmental data of the forest planting area, establishing prediction models and optimizing the forest planting structure, the error problem in the prediction of forest planting product selection is solved, accuracy and resource utilization efficiency are improved, planting risks are reduced, and personalized suggestions and pest warnings are provided.

CN119358821BActive Publication Date: 2025-08-08LANZHOU UNIV
View PDF 1 Cites 0 Cited by

Patent Information

Application Number
CN202411393045.X
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2024-10-08
Publication Date
2025-08-08
Estimated Expiration
2044-10-08

AI Technical Summary

Technical Problem

In the prior art, due to the different evaluation criteria for forest planting of different tree species and different uses, there are errors in the prediction results of forest planting product selection, which makes it difficult to improve the evaluation accuracy and accuracy.

Method used

By collecting and integrating genomic, phenotype and growth environment data of the forest planting area, pre-processing and feature extraction, establishing prediction models, combining a comprehensive evaluation index system, optimizing the planting structure, and selecting the forest varieties that are most suitable for the local environment.

Benefits of technology

It improves the accuracy of forest planting and product selection and resource utilization efficiency, reduces production costs and planting risks, provides personalized product selection suggestions, optimizes planting structure, reduces resource waste, improves economic benefits, and provides pest and disease warnings.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119358821B_ABST
    Figure CN119358821B_ABST
Patent Text Reader

Abstract

The present invention discloses a method and system for predicting tree planting selection based on big data, which relates to the technical field of tree planting selection, and includes the following steps: determining a prediction target and a prediction time range, collecting genomic data, phenotypic data, planting data, and growth environment data of the predicted target planting area, and integrating them into a data set; preprocessing the data set, extracting characteristics related to the tree planting varieties in the data set, including uses and key traits, and screening out key characteristics that affect the tree planting varieties through statistical methods; establishing a prediction model based on data characteristics and prediction targets. The present invention integrates and analyzes cross-regional and cross-variety tree planting data through the use of big data technology, identifies key factors that affect tree growth, yield, and quality, and constructs a comprehensive evaluation system based on these common factors to reduce errors caused by a single standard and improve the accuracy of prediction results.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of forest tree planting and selection technology, and in particular to a forest tree planting and selection prediction method and system based on big data. Background Art

[0002] Forestry planting selection refers to the process of selecting the most suitable forestry varieties or tree species for planting in a region based on multiple factors such as specific environmental conditions, market demand, economic benefits and ecological goals during the forestry planting process. At present, with the global attention to environmental protection and sustainable development, forestry, as an indispensable part of the ecosystem, its sustainable development is particularly important. With the continuous development of information technology, forestry informatization and intelligence have also become a convenient way in the process of forestry planting selection. Forestry planting selection methods based on big data have also emerged. They use massive data and advanced analysis technology to optimize the process of forestry planting variety selection. For example, for the three trees of poplar, willow and catalpa, growth rate, stress resistance, wood quality and ecological adaptability are important quantitative traits for selection. Through big data analysis technology, the performance of different varieties in these aspects can be evaluated more scientifically. At the same time, in order to further improve the efficiency of forestry planting, the use of big data for forestry planting selection will be more widely applied and developed.

[0003] In the existing technology, since different tree species and trees with different uses have different evaluation standards, there will be certain errors in the selection prediction results. Therefore, we need to establish a unified evaluation mechanism to improve the evaluation accuracy and the accuracy of the prediction results. To this end, we propose a forest planting selection prediction method and system based on big data. Summary of the Invention

[0004] The purpose of the present invention is to provide a method and system for predicting tree planting selection based on big data to solve the problems raised in the above background technology.

[0005] In order to solve the above technical problems, the technical solution adopted by the present invention is:

[0006] In the first aspect, a method for predicting tree planting selection based on big data includes the following steps:

[0007] Step 1: Determine the prediction target and prediction timeframe, collect genomic data, phenotypic data, planting data, and growth environment data for the target planting area, and integrate them into a dataset; genomic data includes genome sequences or single nucleotide polymorphism data for different tree varieties within the target planting area; growth environment data includes soil quality, climate temperature, humidity, rainfall, light, topography, and other data; planting data includes information on tree varieties, growth conditions, yield, and the occurrence of pests and diseases;

[0008] Step 2: Preprocess the dataset to improve data quality. Preprocessing includes data cleaning, data standardization, and handling missing values. Features related to tree species in the dataset are extracted, including uses and key traits such as growth rate, adaptability, economic value, and ecological benefits. Statistical methods are used to screen out key features that affect tree species. If the number of features is too large, dimensionality reduction techniques can be used to reduce the number of features and improve model training efficiency.

[0009] Step 3: Establish a prediction model based on data characteristics and prediction targets, and use the preprocessed data set to train the prediction model;

[0010] Step 4: Establish a comprehensive evaluation index system based on the purpose and key characteristics of the trees to ensure the comprehensiveness and accuracy of the prediction results. Use the trained prediction model to predict the tree varieties in the target planting area and output the prediction results for each tree variety.

[0011] A further improvement of the technical solution of the present invention is that the data set integration process is:

[0012] Step 101, determine the predicted tree species, such as poplar, willow, catalpa, fast-growing forest, economic forest, ornamental forest, etc., and the specific prediction targets, such as growth rate, wood quality, disease and pest resistance, etc., and determine the time range of the prediction target; such as short-term, medium-term, long-term,

[0013] Step 102: Obtain the genome sequences of different tree species through a database, use drones and sensors to automatically collect visible characteristics of different tree species, such as growth rate, wood quality, and pest and disease resistance, and collect climate data of the target planting area, such as temperature, humidity, rainfall, and light intensity; soil data, such as soil type, pH value, and nutrient content; and topographic data, such as altitude, slope, and terrain characteristics;

[0014] Step 103 : Integrate and associate the genomic data, phenotypic data, planting data, and growth environment data in a unified format to form a complete data set, and analyze the data set to understand the distribution, correlation, and potential problems of the data.

[0015] A further improvement of the technical solution of the present invention is that the step of screening the key features is:

[0016] Step 201: Check and fill missing values in the data set, identify and process outliers in the data set, and perform data standardization;

[0017] Step 202: extracting the genomic, phenotypic, and environmental characteristics of the tree species, and calculating the correlation coefficients and characteristic evaluation coefficients between the predicted target and the genomic, phenotypic, and environmental characteristics, and outputting a correlation analysis report and a characteristic evaluation report.

[0018] Step 203: extracting key features that have an impact on the prediction of forest tree varieties based on the correlation analysis report and the characteristic evaluation report;

[0019] Step 204: extract the features that have the greatest impact on the prediction target based on the key features.

[0020] A further improvement of the technical solution of the present invention is that the calculation formula of the correlation coefficient is:

[0021]

[0022] in, represents the correlation coefficient, Represents the genome feature G i and phenotypic characteristics P j The Pearson correlation coefficient between and Represents the genome features G i and phenotypic characteristics P j and environmental characteristics E k The Pearson correlation coefficient between Represents environmental characteristics E k The correlation coefficient with itself is the variance of the feature.

[0023] A further improvement of the technical solution of the present invention is that the calculation formula of the characteristic evaluation coefficient is:

[0024]

[0025] Among them I F represents the characteristic evaluation coefficient, N represents the total number of features in the feature set, and w i,f Indicates that in the i-th tree in the random forest, λ i represents the singular value in partial least squares regression PLS, w fi and v fi represents the weight of feature f and the i-th component of the response weight in PLS, 1+λ i Indicates the adjustment factor used to balance w fi and v fi The square term, log(1+F f ) indicates that the F statistic is logarithmically transformed.

[0026] A further improvement of the technical solution of the present invention is that the process of training the prediction model is:

[0027] Step 301: Divide the preprocessed data set into a training set, a validation set, and a test set, and use the training set data to train the prediction model;

[0028] Step 302: Use the validation set to validate the prediction model, evaluate the model performance, and select key features for retraining based on the validation results and the characteristic evaluation report, and adjust the prediction model;

[0029] Step 303: Use the test set to test the adjusted prediction model and output the test results. The test set should be kept independent from the training set and the validation set to ensure the objectivity and reliability of the evaluation results.

[0030] A further improvement of the technical solution of the present invention is that: the tree species to be planted in the predicted target planting area are:

[0031] Step 401: Determine the purpose of the trees, and based on the purpose of the trees, identify the key traits that affect the performance of the trees and obtain the evaluation values, such as biomass, trunk morphology, leaf area index, disease resistance, etc.

[0032] Step 402: Use the trained prediction model to predict the tree species in the target planting area and output the prediction results in the form of charts, such as bar charts, line charts, radar charts, etc., to intuitively display the scores or rankings of each tree species in terms of growth rate, wood quality, stress resistance, etc. The prediction results may include the scores or rankings of each species in terms of growth rate, wood quality, stress resistance, etc.

[0033] Step 403: associate each evaluation value with the prediction result to obtain a comprehensive score, analyze the comprehensive score of each variety, identify the best variety, and output a detailed report.

[0034] A further improvement of the technical solution of the present invention is that the output process of the comprehensive evaluation result is:

[0035] Step 501: Based on the predicted ecological and environmental conditions of the target planting area, the adaptability of each tree variety in the target planting area is evaluated, and its growth habits, drought tolerance, cold tolerance, disease and pest resistance, etc. are considered to determine whether they match the regional conditions. Based on the market demand for different tree varieties, the market potential, competitiveness, and market demand trends of each tree variety in the predicted results are analyzed.

[0036] Step 502: Based on the comprehensive score, score each tree species and calculate the comprehensive score of each tree species to obtain a comprehensive score result;

[0037] Step 503: Based on the comprehensive scoring results, the highest-scoring tree species are selected as the tree species to be planted. The final decision is made by considering the following factors: ecological risk: assessing the potential ecological risks of the selected species and taking appropriate measures to prevent and mitigate them; technical feasibility: considering the availability and feasibility of planting technology to ensure that the selected species can be successfully planted and managed; and policy support: understanding and considering the support and restrictions of relevant policies on the planted species to obtain more policy benefits and support.

[0038] Step 504 outputs the comprehensive scoring results in the form of a chart to clarify the comprehensive score and ranking of each tree species. At the same time, planting suggestions are put forward, including variety selection, planting techniques, management measures, etc., to provide scientific basis and guidance for actual planting.

[0039] In a second aspect, a big data-based forestry planting selection prediction system includes a data management center, wherein the data management center is communicatively connected to a data acquisition module, a data preprocessing module, a data analysis and mining module, a prediction model construction module, an evaluation module, a user interaction and feedback module, and a privacy protection module, wherein the modules are electrically connected;

[0040] The data acquisition module is used to collect planting data of forest tree species, and organize, classify and store the collected planting data to ensure the accuracy and completeness of the data, providing a basis for subsequent data analysis, and using cloud storage or distributed database systems to support large-scale data storage and efficient retrieval;

[0041] The data preprocessing module is used to clean, integrate and standardize the collected planting data to ensure the quality and consistency of the planting data, improve the availability and accuracy of the data, provide a reliable data source for subsequent analysis, and extract feature data from the preprocessed data. These features have predictive value for forest tree planting selection;

[0042] The data analysis and mining module is used to use big data technologies such as data mining and machine learning to analyze the collected planting data, and to mine the correlations, trends and patterns between the data, and output data analysis and mining reports to provide a basis for forest tree planting selection prediction;

[0043] The prediction model building module: Based on the data analysis and mining report, it builds a forest tree planting variety prediction model, uses the feature data to train the prediction model, uses the trained prediction model to perform prediction analysis on new planting data, outputs prediction results, provides recommendations for forest tree planting variety selection, and adjusts model parameters according to the evaluation results to optimize model performance;

[0044] The evaluation module is used to evaluate the effect of the prediction model and adjust the model parameters based on the feedback of actual planting results;

[0045] The user interaction and feedback module is used to provide a user-friendly interface to enable growers to obtain product selection prediction results, collect user feedback on the prediction results, and optimize and improve the prediction model, including online query, report generation, early warning prompts, user guides, etc.

[0046] The privacy protection module is used to protect the security of planting data and prevent data leakage and unauthorized access.

[0047] Due to the adoption of the above technical solution, the present invention has the following technical advancements compared to the prior art:

[0048] 1. The present invention provides a method and system for predicting tree planting varieties based on big data. By collecting and analyzing the genome, phenotype and environmental big data of tree varieties, the adaptability and potential performance of each variety in different ecological environments are comprehensively evaluated. According to environmental conditions and variety characteristics, the planting structure is optimized, and the varieties most suitable for the local environment are selected for planting. By rationally allocating planting resources, planters can reduce unnecessary investment, improve resource utilization efficiency, reduce production costs, significantly reduce planting risks, and increase planting success rates.

[0049] 2. The present invention provides a method and system for predicting tree planting selection based on big data. By applying big data technology, it integrates and analyzes cross-regional and cross-variety tree planting data, identifies key factors affecting tree growth, yield and quality, and constructs a comprehensive evaluation system based on these common factors to reduce errors caused by a single standard and improve the accuracy of prediction results.

[0050] 3. The present invention provides a method and system for predicting tree planting selection based on big data. By collecting and analyzing the genome, phenotype and environmental big data of tree varieties, the adaptability and potential performance of each variety in different ecological environments are comprehensively evaluated. According to environmental conditions and variety characteristics, the planting structure is optimized, and the varieties most suitable for the local environment are selected for planting. By rationally allocating planting resources, planters can reduce unnecessary investment, improve resource utilization efficiency, reduce production costs, significantly reduce planting risks, and increase planting success rates.

[0051] 4. This invention provides a big data-based prediction method and system for tree selection for planting. This method predicts the growth performance of different tree varieties under varying conditions, provides personalized selection recommendations, and optimizes planting structures. Furthermore, by accurately predicting tree growth cycles and yields, production plans can be more rationally arranged, reducing resource waste and improving economic efficiency. Furthermore, the system can provide early warnings of pests and diseases, helping growers take timely measures to minimize losses and ensure profitable planting. BRIEF DESCRIPTION OF THE DRAWINGS

[0052] In order to more clearly illustrate the embodiments of the present application or the technical solutions in the prior art, the following briefly introduces the drawings required for use in the embodiments. Obviously, the drawings described below are only some embodiments described in the present invention. For ordinary technicians in this field, other drawings can also be obtained based on these drawings.

[0053] Figure 1 is a flow chart of the method of the present invention;

[0054] Figure 2 A flow chart for screening key features of the present invention;

[0055] Figure 3 It is a system module diagram of the present invention. DETAILED DESCRIPTION

[0056] To make the objectives, technical solutions, and advantages of the embodiments of the present invention more clear, the technical solutions in the embodiments of the present invention will be clearly and completely described below in conjunction with the accompanying drawings of the embodiments of the present invention. Obviously, the described embodiments are only part of the embodiments of the present invention, not all of the embodiments. All other embodiments obtained by ordinary technicians in this field based on the embodiments of the present invention without making any creative efforts shall fall within the scope of protection of the present invention.

[0057] Example 1, as Figure 1 、 2As shown, the present invention provides a method for predicting tree planting selection based on big data, comprising the following steps: Step 1, determining the prediction target and the prediction time range, collecting genomic data, phenotypic data, planting data and growth environment data of the predicted target planting area, and integrating them into a data set; the genomic data is the genome sequence or single nucleotide polymorphism data of different tree planting varieties in the target planting area, the growth environment data is the soil quality, climate temperature, humidity, rainfall, light, etc., as well as topography and landform data, and the planting data is the information on the tree planting variety, growth conditions, yield, disease and pest occurrence, etc.; the data set integration process is: determining the predicted tree planting variety, such as poplar, willow, catalpa, fast-growing forest, economic forest, ornamental forest, etc., and the specific prediction targets, such as growth rate, wood quality, disease and pest resistance, etc., and determining the time range of the prediction target; such as short-term, medium-term, long-term, The genome sequences of different tree species are obtained through databases. UAVs and sensors are used to automatically collect visible characteristics of different tree species, such as growth rate, wood quality, and pest and disease resistance. Climate data of the target planting area, such as temperature, humidity, rainfall, and light intensity, soil data, such as soil type, pH value, and nutrient content, and topographic data, such as altitude, slope, and terrain characteristics, are collected. The genome data, phenotypic data, planting data, and growth environment data are integrated and correlated in a unified format to form a complete data set. The data set is then analyzed to understand the data distribution, correlation, and potential problems. Statistical quantities such as the mean, standard deviation, maximum value, and minimum value of each data indicator are calculated to understand the data distribution. The correlation coefficient is used to analyze the correlation between different data indicators and identify key factors affecting tree growth rate, wood quality, and pest and disease resistance.

[0058] Step 2: Preprocess the data set to improve data quality. Preprocessing includes data cleaning, data standardization, and processing of missing values. Features related to tree species in the data set are extracted, including uses and key traits, such as growth rate, adaptability, economic value, and ecological benefits. Statistical methods are used to screen out key features that affect tree species. If the number of features is too large, the number of features can be reduced through dimensionality reduction technology to improve model training efficiency. The steps for screening key features are: checking and filling missing values in the data set, identifying and processing outliers in the data set, and performing data standardization; extracting genomic, phenotypic, and environmental characteristics of tree species, and calculating the correlation coefficients and characteristic evaluation coefficients between the predicted targets and the genomic, phenotypic, and environmental characteristics, and outputting a correlation analysis report and a characteristic evaluation report at the same time; extracting key features that affect the prediction of tree species based on the correlation analysis report and the characteristic evaluation report; extracting the features that have the greatest impact on the predicted target based on the key features; the calculation formula for the correlation coefficient is:

[0059]

[0060] in, represents the correlation coefficient, Represents the genome feature G i and phenotypic characteristics P j The Pearson correlation coefficient between and Represents the genome features G i and phenotypic characteristics P j and environmental characteristics E k The Pearson correlation coefficient between Represents environmental characteristics E k The correlation coefficient with itself is the variance of the feature;

[0061] The calculation formula of the characteristic evaluation coefficient is:

[0062]

[0063] Among them I F represents the characteristic evaluation coefficient, N represents the total number of features in the feature set, and w i,f Indicates that in the i-th tree in the random forest, λ i represents the singular value in partial least squares regression PLS, w fi and v fi represents the weight of feature f and the i-th component of the response weight in PLS, 1+λ i Indicates the adjustment factor used to balance w fi and v fi The square term, log(1+F f ) indicates logarithmic transformation of F statistic;

[0064] Step 3: Establish a prediction model based on data characteristics and prediction objectives, and use the preprocessed dataset to train the prediction model. The process of training the prediction model is as follows: divide the preprocessed dataset into a training set, a validation set, and a test set, and use the training set data to train the prediction model; use the validation set to verify the prediction model and evaluate the model performance. Based on the verification results and characteristic evaluation report, select key features for retraining and adjust the prediction model; use the test set to test the adjusted prediction model and output the test results. The test set should be kept independent of the training set and validation set to ensure the objectivity and reliability of the evaluation results.

[0065] Step 4: Establish a comprehensive evaluation index system based on the purpose and key traits of trees to ensure the comprehensiveness and accuracy of the prediction results. Use the trained prediction model to predict the tree varieties in the target planting area and output the prediction results of each tree variety. Predicting the tree varieties in the target planting area is as follows: determine the purpose of the trees, and based on the purpose of the trees, identify the key traits that affect the performance of the trees and obtain various evaluation values; such as biomass, trunk morphology, leaf area index, disease resistance, etc. Use the trained prediction model to predict the tree varieties in the target planting area and output the prediction results in the form of charts; such as bar charts, line charts, radar charts, etc., to intuitively display the scores or rankings of each tree variety in terms of growth rate, wood quality, stress resistance, etc. The prediction results may include the scores or rankings of each variety in terms of growth rate, wood quality, stress resistance, etc. Associate each evaluation value with the prediction result to obtain a comprehensive score, analyze the comprehensive score of each variety, identify the best variety, and output a detailed report.

[0066] Step 5: Combined with the ecological environment requirements and market demand, the forecast results are comprehensively evaluated, and the comprehensive evaluation results are output. Based on the comprehensive evaluation results, the most suitable forest tree varieties are selected for planting. The output process of the comprehensive evaluation results is as follows: based on the ecological environment conditions of the predicted target planting area, the adaptability of each forest tree variety in the target planting area is evaluated, and its growth habits, drought resistance, cold resistance, disease and pest resistance, etc. are considered to match the regional conditions. Based on the market demand for different forest tree varieties, the potential, competitiveness and market demand trend of each forest tree variety in the forecast results are analyzed. Based on the comprehensive score, each forest tree variety is scored and the comprehensive score of each forest tree variety is calculated to obtain the comprehensive score. Comprehensive scoring results; based on the comprehensive scoring results, select the tree species with the highest score as the tree planting variety. At the same time, consider the following factors to make the final decision: Ecological risk: Assess the ecological risks that may be brought about by the selected varieties, and take corresponding measures to prevent and alleviate them; Technical feasibility: Consider the availability and feasibility of planting technology to ensure that the selected varieties can be smoothly planted and managed; Policy support: Understand and consider the support and restrictions of relevant policies on planting varieties to obtain more policy preferences and support; Output the comprehensive scoring results in the form of charts to clarify the comprehensive score and ranking of each tree variety. At the same time, put forward planting recommendations, including variety selection, planting technology, management measures, etc., to provide a scientific basis and guidance for actual planting.

[0067] Example 2, as Figure 3As shown, based on Example 1, the present invention further provides a forest tree planting selection prediction system based on big data, including a data management center, the data management center is communicatively connected to a data acquisition module, a data preprocessing module, a data analysis and mining module, a prediction model construction module, an evaluation module, a user interaction and feedback module, and a privacy protection module, wherein the modules are electrically connected;

[0068] Data collection module: used to collect planting data of forest tree varieties, including climate conditions, soil types, water source conditions, historical planting records, market demand, tree growth cycle, pest and disease occurrence, etc., and organize, classify and store the collected planting data to ensure the accuracy and completeness of the data, providing a basis for subsequent data analysis. Cloud storage or distributed database systems are used to support large-scale data storage and efficient retrieval;

[0069] Data preprocessing module: used to clean, integrate and standardize the collected planting data to ensure its quality and consistency, improve its availability and accuracy, provide a reliable data source for subsequent analysis, and extract feature data from the preprocessed data. These features have predictive value for tree planting selection;

[0070] Data analysis and mining module: This module uses big data technologies such as data mining and machine learning to analyze the collected planting data and identify correlations, trends, and patterns between the data. It also outputs data analysis and mining reports to provide a basis for tree planting selection predictions.

[0071] Prediction model building module: Based on data analysis and mining reports, a forest tree variety prediction model is constructed, and the prediction model is trained using feature data. The trained prediction model is used to perform prediction analysis on new planting data, output prediction results, provide recommendations for forest tree variety selection, and adjust model parameters based on the evaluation results to optimize model performance.

[0072] Evaluation module: used to evaluate the effectiveness of the prediction model and adjust the model parameters based on feedback from actual planting results;

[0073] User interaction and feedback module: This module provides a user-friendly interface, enabling growers to obtain prediction results, collect user feedback on the prediction results, and optimize and improve the prediction model. It includes online query, report generation, early warning prompts, user guides, etc.

[0074] Privacy protection module: used to protect the security of planting data and prevent data leakage and unauthorized access.

[0075] The above are only specific embodiments of the present application, but the scope of protection of this application is not limited thereto. Any changes or substitutions that can be easily conceived by a person skilled in the art within the technical scope disclosed in this application should be included in the scope of protection of this application. Therefore, the scope of protection of this application should be based on the scope of protection of the claims.

Claims

1. A method for predicting tree planting selection based on big data, characterized by: The following steps are involved: Step 1: Determine the prediction target and prediction timeframe, collect genomic data, phenotypic data, planting data, and growth environment data for the target planting area, and integrate them into a dataset. Prediction targets include growth rate, wood quality, and pest and disease resistance, and prediction timeframes include short-term, medium-term, and long-term. Step 2: Preprocess the dataset to extract features related to tree species, including uses and key traits, and use statistical methods to screen out key features that affect tree species. Step 3: Establish a prediction model based on data characteristics and prediction targets, and use the preprocessed data set to train the prediction model; Step 4: Establish a comprehensive evaluation index system based on the uses and key characteristics of the trees, use the trained prediction model to predict the tree species in the target planting area, and output the prediction results for each tree species; The tree species planted in the predicted target planting area are: Step 401, determining the purpose of the trees, and based on the purpose of the trees, identifying the key traits that affect the performance of the trees, and obtaining respective evaluation values; Step 402: Use the trained prediction model to predict the tree species in the target planting area and output the prediction results in a graphical form. Step 403: Correlate each evaluation value with the prediction result to obtain a comprehensive score, analyze the comprehensive score of each variety, identify the best variety, and output a detailed report; Step 5: Combined with ecological environment requirements and market demand, conduct a comprehensive evaluation of the forecast results, output the comprehensive evaluation results, and select the most suitable forest planting varieties for planting based on the comprehensive evaluation results.

2. The method for predicting tree planting selection based on big data according to claim 1, characterized in that: The integration process of the dataset is as follows: Step 101, determining the predicted tree species and specific prediction targets, as well as the time range of the prediction targets; Step 102: Obtain genome sequences of different tree species through a database, automatically collect visible features of different tree species using drones and sensors, and collect climate data, soil data, and topographic data of the target planting area; Step 103 : Integrate and associate the genomic data, phenotypic data, planting data, and growth environment data in a unified format to form a complete data set, and analyze the data set to understand the distribution, correlation, and potential problems of the data.

3. The method for predicting tree planting selection based on big data according to claim 2, characterized in that: The steps for screening the key features are: Step 201: Check and fill missing values in the data set, identify and process outliers in the data set, and perform data standardization; Step 202: extracting the genomic, phenotypic, and environmental characteristics of the tree species, and calculating the correlation coefficients and characteristic evaluation coefficients between the predicted target and the genomic, phenotypic, and environmental characteristics, and outputting a correlation analysis report and a characteristic evaluation report. Step 203: extracting key features that have an impact on the prediction of forest tree varieties based on the correlation analysis report and the characteristic evaluation report; Step 204: extract the features that have the greatest impact on the prediction target based on the key features.

4. The method for predicting tree planting selection based on big data according to claim 3, characterized in that: The calculation formula of the correlation coefficient is: ; in, represents the correlation coefficient, Representing genomic features and phenotypic characteristics The Pearson correlation coefficient between and Represents genomic features and phenotypic characteristics and environmental characteristics The Pearson correlation coefficient between Representing environmental characteristics The correlation coefficient with itself is the variance of the feature.

5. The method for predicting tree planting selection based on big data according to claim 4, characterized in that: The calculation formula of the characteristic evaluation coefficient is: ; in represents the characteristic evaluation coefficient, N represents the total number of features in the feature set, Indicates that in the i-th tree in the random forest, represents the singular value in partial least squares regression PLS, and represents the weight of feature f and the i-th component of the response weight in PLS, Represents an adjustment factor for balance and The square term of Indicates that the F statistic is logarithmically transformed.

6. The method for predicting tree planting selection based on big data according to claim 3, characterized in that: The process of training the prediction model is as follows: Step 301: Divide the preprocessed data set into a training set, a validation set, and a test set, and use the training set data to train the prediction model; Step 302: Use the validation set to validate the prediction model, evaluate the model performance, and select key features for retraining based on the validation results and the characteristic evaluation report, and adjust the prediction model; Step 303: Use the test set to test the adjusted prediction model and output the test results.

7. The method for predicting tree planting selection based on big data according to claim 6, characterized in that: The output process of the comprehensive evaluation results is as follows: Step 501: Based on the predicted ecological and environmental conditions of the target planting area, the adaptability of each forest and each tree species in the target planting area is evaluated, and based on the market demand for different tree species, the market potential, competitiveness, and market demand trend of each tree species in the predicted results are analyzed; Step 502: Based on the comprehensive score, score each tree species and calculate the comprehensive score of each tree species to obtain a comprehensive score result; Step 503: Based on the comprehensive scoring results, select the tree species with the highest score as the tree species to be planted; Step 504: Output the comprehensive scoring results in the form of a chart to clarify the comprehensive score and ranking of each tree species.

8. A big data-based tree planting selection prediction system, used to implement a big data-based tree planting selection prediction method according to any one of claims 1 to 7, characterized in that: It includes a data management center, which is communicatively connected to a data acquisition module, a data preprocessing module, a data analysis and mining module, a prediction model building module, an evaluation module, a user interaction and feedback module, and a privacy protection module, wherein the modules are connected by electrical signals; The data collection module is used to collect the planting data of forest tree species and to organize, classify and store the collected planting data; The data preprocessing module is used to clean, integrate and standardize the collected planting data and extract feature data from the preprocessed data; The data analysis and mining module is used to use big data technology of data mining and machine learning to analyze the collected planting data, and to mine the correlation, trends and patterns between the data, and output a data analysis and mining report; The prediction model building module: based on the data analysis and mining report, builds a forest tree planting variety prediction model, uses the feature data to train the prediction model, uses the trained prediction model to perform prediction analysis on new planting data, outputs the prediction results, adjusts the model parameters according to the evaluation results, and optimizes the model performance; The evaluation module is used to evaluate the effect of the prediction model and adjust the model parameters based on the feedback of actual planting results; The user interaction and feedback module is used to provide a user-friendly interface, allowing growers to obtain product selection prediction results, and collect user feedback on the prediction results to optimize and improve the prediction model; The privacy protection module is used to protect the security of planting data and prevent data leakage and unauthorized access.

Citation Information

Patent Citations

  • Rice new variety value evaluation method and system based on generative model

    CN118469155A