A method for predicting soil microbial functions based on machine learning

By collecting and analyzing soil microbial 16S rRNA sequencing data, combining FAPROTAX database and machine learning modeling, the distribution pattern of soil microbial functions is predicted worldwide, solving the problem of difficulty in understanding the impact of climate change on soil microbial functions in the existing technology, and achieving rapid and accurate soil function prediction, providing a theoretical basis for land management.

CN117497053BActive Publication Date: 2025-06-27NANKAI UNIV +1
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202310785544.2
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2023-06-29
Publication Date
2025-06-27
Estimated Expiration
2043-06-29

AI Technical Summary

Technical Problem

It is difficult for the existing technology to deeply understand the impact of climate change on soil microbial community structure and function, and high-throughput sequencing technology has problems such as high sequencing costs, long cycles, large data volumes, and high analysis difficulty.

Method used

By collecting high-throughput sequencing data of soil microorganisms in the literature, using QIIME2 bioinformatic analysis software for analysis, and combining with the FAPROTAX database to predict soil microbial functional genes, constructing a database of global soil carbon-nitrogen cycle-related functions, using machine learning modeling to reveal the regulation of microbial community structure changes by environmental factors, and then predicting the distribution pattern of global soil microbial functions.

Benefits of technology

It has achieved rapid and accurate prediction of the carbon-nitrogen circulation function of global soil microorganisms, provided a theoretical basis for the formulation of land management policies, and overcome the problems of insufficient data and difficulty in analysis in the existing technology.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN117497053B_ABST
    Figure CN117497053B_ABST
Patent Text Reader

Abstract

The present invention relates to a method for predicting soil microbial functions based on machine learning, comprising the following steps: S1 Collect data such as soil microbial gene sequencing serial numbers in the literature, as well as geographical information, climate variables, soil physical and chemical properties, vegetation indices, land cover types, etc. at the corresponding locations; S2 Download global raster files of climate variables, soil physical and chemical properties, vegetation indices, and land cover types from the database; S3 According to the obtained gene sequencing data serial numbers, download microbial gene sequencing data and construct a soil surface microbial function database; S4 Select functional genes reflecting the carbon cycling ability of the soil environment and calculate the soil carbon cycling function index; select functional genes reflecting the nitrogen cycling ability in the soil environment and calculate the soil nitrogen cycling function index; S5 Construct machine learning models for soil carbon cycling functions and soil nitrogen cycling functions respectively.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The invention belongs to the field of environmental technology, and in particular relates to a method for predicting soil microbial functions based on machine learning. Background Art

[0002] Soil microorganisms are diverse and are important components in the plant-soil material cycle. They can adapt to environmental changes by adjusting their number, function and population structure, and are indicative of the health of the ecosystem. A large number of studies have shown that the structure of soil microbial communities is usually directly affected by soil physical and chemical properties (soil temperature, moisture, pH, nutrient content) and environmental factors (temperature, precipitation). At the same time, the structure of soil microbial communities is closely linked to their functional activity. Therefore, environmental changes will change the abundance of microbial genes involved in soil carbon and nitrogen cycles, thereby affecting soil respiration, leading to changes in soil carbon and nitrogen cycles.

[0003] In the context of global warming, rising temperatures may change the soil organic carbon decomposition rate by changing the soil microbial community structure and metabolic pathways, so it is crucial to understand how climate change affects the soil microbial community structure and function. However, due to the limitations of research technology, the understanding of soil microbial community structure and function was not in-depth enough. The development of high-throughput sequencing technology provides us with an opportunity to conduct in-depth research on soil organic carbon and microbial characteristics.

[0004] High-throughput sequencing is currently a widely recognized means for in-depth research on environmental microorganisms. Although the information of metagenomics and metatranscriptomes is more comprehensive and can be used for in-depth analysis of environmental microorganisms, there are problems such as high sequencing cost, long cycle, large amount of data, and high difficulty in analysis. Amplicon sequencing (such as 16S, 18S, ITS, etc.) is the sequencing of PCR products or captured fragments of a specific length. 16S rRNA sequencing is a research method that uses bacterial universal primers to amplify the variable region of DNA in environmental samples to analyze the composition and abundance of bacterial species, population structure, and systematic evolution in the environment. Compared with meta-omics sequencing, although 16S rRNA sequencing cannot go deep into functional annotation analysis, its sequencing cycle is shorter, the cost is lower, and the data volume is small, which greatly reduces the difficulty of analysis. The FAPROTAX database is a database manually constructed by Louca et al. based on published literature. It currently collects more than 7,600 functional annotation information from more than 80 functional groups of more than 4,600 prokaryotic microorganisms, and is still being updated. The FAPROTAX database has, to a certain extent, solved the limitation of 16S rRNA sequencing that cannot be used for functional level research, and can quickly and cost-effectively obtain soil microbial species and functional information.

[0005] In recent years, a large number of soil microorganism studies have provided a large amount of sequencing data for analyzing the global soil microorganism functions. Machine learning has natural advantages in analyzing big data. Machine learning can handle the relationships between multiple variables and labels and make predictions quickly. By screening the data set and adjusting the model parameters, good prediction effects can be achieved. Machine learning methods such as random forest, XGBoost, and neural network are used to test the environmental-related hypotheses of soil microorganism functional abundances at a large scale, reveal how environmental factors affect soil microbial communities and the changes in their carbon and nitrogen cycle functional genes, and further reveal how the carbon and nitrogen cycle functions of ecosystems change. Summary of the Invention

[0006] The technical problem to be solved by the present invention is to overcome the limitations of the prior art and insufficient data. A method for predicting soil microorganism functions based on machine learning is to analyze the existing data of bacterial 16S rRNA high-throughput sequencing technology in the literature using the QIIME2 bioinformatics analysis software, further predict the functional genes of soil microorganisms through the FAPROTAX database, and construct a database related to the global soil carbon and nitrogen cycles. Combining machine learning modeling to reveal how environmental factors regulate the changes in microbial community structure and further how to change the microbial functions related to the carbon and nitrogen cycles, and further predict the distribution patterns of global soil microorganism functions, providing a basic basis for future land management. To achieve the above objectives, the technical solutions of the present invention are as follows.

[0007] A method for predicting soil microorganism functions based on machine learning includes the following steps:

[0008] S1 Collect relevant literature on soil microorganisms - climate change, extract literature information, and collect data such as soil microorganism gene sequencing serial numbers and corresponding geographical information, climate variables, soil physical and chemical properties, vegetation indices, land cover types, etc. in the literature. The method is as follows:

[0009] S2 Download global raster files of climate variables, soil physical and chemical properties, vegetation indices, and land cover types from the database, and use global data to fill in the missing variable values;

[0010] S3 According to the obtained gene sequencing data serial numbers, download the microbial gene sequencing data, use bioinformatics analysis software to process the microbial gene sequencing data, classify and annotate the species, and predict the relative abundances of microbial functional genes; integrate climate variables, soil physical and chemical properties, land cover types, altitude, and longitude and latitude information to construct a soil surface microbial function database;

[0011] S4 Select functional genes reflecting the carbon cycle ability of the soil environment and calculate the soil carbon cycle function index; select functional genes reflecting the nitrogen cycle ability in the soil environment and calculate the soil nitrogen cycle function index;

[0012] S5 constructs machine learning models for soil carbon cycle function and soil nitrogen cycle function respectively: the predictive variables are the soil carbon cycle function index and the soil nitrogen cycle function index; based on the soil surface microbial function database, characteristic variables are determined, divided into a training set and a test set, and based on the machine learning random forest training model, the parameters of the random forest are adjusted to optimize the model effect.

[0013] Further, in S1, the gene sequencing serial number is the 16S rRNA sequencing data serial number.

[0014] Further, S1 is executed according to the following steps:

[0015] S11 retrieves literature topics including soil microorganisms, high-throughput sequencing, climate change, and land cover types, marks the literature containing soil microbial gene sequencing serial numbers, and obtains microbial sequencing data points;

[0016] S12 collects the characteristic variables of the microbial sequencing data points, including longitude and latitude, altitude, sampling year, sampling depth, sequencing instrument, amplification region, climate variables, soil physical and chemical properties, soil carbon and nitrogen content, vegetation index, and land cover types.

[0017] For the method of predicting soil microbial function, in S12, the climate variables include annual average temperature, annual precipitation, precipitation seasonality, temperature seasonality, average daily temperature range, drought index, potential evapotranspiration; the soil physical and chemical properties include soil pH, soil water content, soil bulk density, soil sand content, soil clay content, cation exchange capacity, and soil organic carbon, total nitrogen, and total phosphorus; the vegetation index refers to the normalized difference vegetation index NDVI; the land cover types include grassland, forest, shrub, bare land, and farmland.

[0018] Further, in S4, z-scores are calculated for the selected different functional genes respectively; after calculating the z-scores of each functional gene, the z-scores of all functional genes are summed and averaged to obtain the soil carbon cycle function index and the soil nitrogen cycle function index.

[0019] Furthermore, the carbon cycle functional genes selected in S4 include eight functions: chitinolysis, cellulolysis, xylanolysis, ureolysis, hydrocarbon degradation, fermentation, methanol oxidation, and methylotrophy. The nitrogen cycle functional genes selected include six functions: nitrification, denitrification, nitrogen fixation, nitrate ammonification, nitrogen respiration, and nitrate reduction.

[0020] Furthermore, in S5, the determined characteristic variables include: altitude, annual average temperature, annual precipitation, precipitation seasonality, temperature seasonality, diurnal temperature range of average temperature, aridity index, potential evapotranspiration, soil temperature, soil pH, soil water content, soil bulk density, soil sand content, soil clay content, cation exchange capacity, soil organic carbon, total nitrogen, total phosphorus, normalized difference vegetation index NDVI, land cover type.

[0021] Furthermore, in S3, the relative abundances of microbial functional genes are predicted using the FAPROTAX database.

[0022] The present invention is realized based on a large amount of high-throughput sequencing data. By collecting 16S rRNA sequencing data, climate variables, soil physical and chemical properties, and land cover types in the literature, more than 2,800 bacterial sequencing data from all over the world are collected to construct a global environmental-soil microbial function database, providing a data basis for future big data research on global soil. Through the random forest model of machine learning, the carbon and nitrogen cycle functional genes of global soil microorganisms are accurately evaluated and predicted, revealing the impact of environmental changes on the soil ecosystem, and can be used to predict the functions of soil microorganisms in unknown regions. BRIEF DESCRIPTION OF THE DRAWINGS

[0023] Figure 1 It is a flowchart of a method for predicting soil microbial functions based on machine learning according to the present invention;

[0024] Figure 2 It is the R2 distribution obtained by ten-fold cross-validation of the random forest of the present invention;

[0025] Figure 3This is the graph of the feature importance ranking after the random forest calculation of the present invention. Detailed implementation manners

[0026] The present invention will be further described in detail below in conjunction with the accompanying drawings and specific implementations.

[0027] A method for predicting soil microbial functions based on machine learning predicts the carbon and nitrogen cycling functional genes of soil surface microorganisms according to high-throughput sequencing data obtained from a large number of studies, obtains the soil microbial functions of more than 2,800 global sample points, and constructs a database. Using artificial intelligence methods such as machine learning, climate variables, soil physical and chemical properties, vegetation indices, and land cover data are used as feature variables, and the soil carbon cycling function index and the soil nitrogen cycling index are used as prediction variables to construct a model, and the model parameters are adjusted to obtain an optimal training model. According to the adjusted model, the soil carbon and nitrogen cycling functions of unknown regions can be accurately predicted, and the soil carbon and nitrogen cycling functions worldwide can be efficiently and rapidly evaluated, providing a theoretical basis for the formulation of land management policies.

[0028] As Figure 1 shown, the main steps of the present invention are as follows:

[0029] S1 Collect relevant literatures on global soil microorganisms - climate change, extract the literature information of the literatures, and collect data such as the 16S rRNA sequencing data serial numbers in the literatures and the geographical information, climate variables, soil physical and chemical properties, vegetation indices, land cover types, etc. of the corresponding locations;

[0030] S2 Download the global raster files of climate variables, soil physical and chemical properties, vegetation indices, and land cover types from the database, and use global data to fill in the missing variable values;

[0031] S3 Download the original 16S rRNA sequencing data from NCBI according to the obtained 16S rRNA sequencing data serial numbers, process the collected sequencing data using the QIIME2 bioinformatics analysis software, obtain the relative abundances of microbial functional genes using the FAPROTAX database, and construct a global soil surface microbial function database;

[0032] S4 Calculate the carbon cycling function index and nitrogen cycling function index of the soil respectively;

[0033] S5 Determine the prediction variables and feature variables, divide the feature variables into a training set and a test set, and based on the machine learning random forest training model, observe and adjust the parameters of the random forest to make the model effect reach the optimal.

[0034] Based on the above steps, it will be further introduced in combination with the data processing process and model construction according to the specific implementation manners, including but not limited to: literature retrieval, missing data supplementation, global raster data download, bioinformatics analysis process, definition of carbon and nitrogen cycle function index, z-score calculation, and random forest construction.

[0035] Further, in step S1, 2815 soil bacterial sequencing data from all over the world are collected, and the sampling points are distributed in 7 continents and more than 40 countries around the world. The specific steps are as follows:

[0036] S1.1 Literature retrieval

[0037] Retrieve in the literature database. The retrieval topics include relevant literatures such as soil microbial diversity, climate change, and soil bacterial 16S rRNA. Store them according to the topics and relevant contents for subsequent data extraction; retrieve in web of science with keywords "soil microbial", "climate change", "microbial function", and "16S", mark the literatures containing sequencing data, and record their literature names, sampling regions and times, sequencing regions, sequencing instruments, and sampling depths.

[0038] S1.2 Collection of characteristic variables

[0039] The collected characteristic variables mainly include altitude, annual average temperature, annual precipitation, precipitation seasonality, temperature seasonality, daily temperature range of average temperature, aridity index, potential evapotranspiration, soil temperature, soil pH, soil water content, soil bulk density, soil sand content, soil clay content, cation exchange capacity, and soil organic carbon, total nitrogen, total phosphorus, normalized difference vegetation index NDVI, and land cover type.

[0040] Further, the specific steps for downloading global data in step S2 are as follows:

[0041] S2.1 Download dataset

[0042] The data of annual average temperature, annual precipitation, precipitation seasonality, temperature seasonality, daily temperature range of average temperature, and altitude are from WorldClim global climate dataset (https: / / www.worldclim.org / ), with a spatial resolution of 1000m; the aridity index and potential evapotranspiration are from Global Aridity Index and Potential Evapotranspiration Climate Database v3 ( https: / / cgiarcsi.community / ) The spatial resolution is 1000 m; the soil temperature data is from the dataset NASA / GLDAS / V021 / NOAH / G025 / T3H uploaded on the Google Earth Engine website, and the downloaded spatial resolution is 1000 m; the physical and chemical properties of soil pH and water content are from Soilgrids (https: / / www.isric.org / ), with a spatial resolution of 1000 m; the land cover data is from the ESA 10 m resolution land cover data of the European Space Agency (https: / / cds.climate.copernicus.eu / ), with a spatial resolution of 10 m; the normalized difference vegetation index (NDVI) is from MOD13Q1 (https: / / modis.gsfc.nasa.gov / ), the NDVI spatial resolution is 250 m, the temporal resolution is 16 days, and the annual average value of NDVI during the sampling period is used in this example.

[0043] S2.2 Filling missing values

[0044] Extract the longitude and latitude of the sampling points with missing data, create a CSV file, select the raster data of the variables to be supplemented, import it using the raster package, use the extract function to extract the pixel values corresponding to the longitude and latitude, and perform supplementation to construct a soil microbial carbon and nitrogen cycle function database.

[0045] Furthermore, the specific steps for analyzing 16S rRNA data and predicting microbial functions using FAPROTAX in step S3 are as follows:

[0046] S3.1 Bioinformatics analysis process

[0047] Prepare the metadata and manifest files and import the 16S rRNA sequencing data in fastq format;

[0048] Use DADA2 to remove sequencing noise, errors, and chimeras from the amplicon sequences, select amplicon sequence variants (ASVs), and generate a feature table;

[0049] Import the species annotation classifier for species composition analysis and export the CSV file of species annotation;

[0050] The annotation results were imported into the FAPROTAX plugin to predict the microbial functions. More than 80 microbial functions were predicted. We mainly selected eight carbon cycle functional genes in the soil: Chitinolysis, cellulolysis, xylanolysis, ureolysis, hydrocarbon_degradation, fermentation, methanol_oxidation, and methylotrophy; six nitrogen cycle functional genes in the soil: nitrification, denitrification, nitrogen_fixation, nitrate_ammonification, nitrogen_respiration, and nitrate_reduction.

[0051] S4.1 Define the soil carbon cycle function index and the soil nitrogen cycle function index: Calculate the z-score for each of the carbon cycle functional genes selected in S3.1. After calculating the z-score for each functional gene, sum and average the z-scores of all selected functional genes to obtain the soil carbon cycle function index; Calculate the z-score for each of the nitrogen cycle functional genes selected in S3.1. After calculating the z-score for each functional gene, sum and average the z-scores of all selected functional genes to obtain the soil nitrogen cycle function index.

[0052] The specific formula for the z-score is as follows:

[0053] z = (X - μ) / σ

[0054] Where X is the measured value of the sample feature, μ is the average value of a certain sample feature, and σ is the standard deviation of a certain sample feature;

[0055] Furthermore, the process of using the feature dataset for random forest modeling in step S5 is as follows:

[0056] S5.1 Read the dataset, determine the feature variables and prediction variables. The feature variables include: altitude, annual average temperature, annual precipitation, precipitation seasonality, temperature seasonality, diurnal temperature range, aridity index, potential evapotranspiration, soil temperature, soil pH, soil water content, soil bulk density, soil sand content, soil clay content, cation exchange capacity, soil organic carbon, total nitrogen, total phosphorus, normalized difference vegetation index NDVI, land cover type; The prediction variables are the soil carbon cycle function index and the soil nitrogen cycle function index respectively.

[0057] S5.2 Data splitting: The sample data is divided into a test set and a training set. The training set is used for modeling, and the test set is used to verify the model performance. In this method, 90% of the test samples are randomly selected from the samples to construct the training set.

[0058] S5.3 Cross-validation method: The ten-fold cross-validation method is used. The test set samples are randomly divided into 10 subsets of comparable sizes, and then the model is evaluated and trained 10 times.

[0059] S5.4 Data modeling: The random forest model regression method is selected to model the training set data. The number of decision trees n_estimators and the maximum depth max_depth are specified, and the cross-validation method is specified as ten-fold cross-validation to train the model.

[0060] S5.5 Model optimization: The root mean square error RMSE and the correlation coefficient R 2 are used to evaluate the prediction effects of different models, and the optimal model is selected from the test models according to the model effects.

[0061] where R 2 = ρ X,Y , X represents the actual value, and Y represents the predicted value;

[0062] Figure 2 As shown, the average R 2 of the ten-fold cross-validation of the soil carbon cycle function of the present invention is (0.815 + 0.813 + 0.814 + 0.814 + 0.811 + 0.834 + 0.808 + 0.815 + 0.819 + 0.825) / 10 = 0.817; the average R 2 of the ten-fold cross-validation of the soil nitrogen cycle function is (0.837 + 0.831 + 0.862 + 0.828 + 0.833 + 0.83 + 0.83 + 0.829 + 0.824 + 0.833) / 10 = 0.834

[0063] S5.6 Variable importance analysis based on SHAP values

[0064] The SHAP value is the value assigned to each feature in the predicted value of the prediction sample, and it can also show the positive and negative of the influence. The average of the absolute values of the SHAP values of each feature is taken as the importance of the feature and sorted in descending order, as Figure 3 shown. The influences of variables on the soil carbon cycle function index and the soil nitrogen cycle function index are not completely consistent. Altitude and soil pH are the most important variables affecting the soil carbon cycle function index and the soil carbon cycle function index respectively, while climate variables such as precipitation seasonality, drought index, and annual precipitation are all important factors affecting the soil carbon and nitrogen cycles, indicating that the impact of climate change on the soil and even the entire ecosystem is still crucial.

[0065] In summary, the present invention provides a method for predicting the functions of soil microorganisms. Based on high-throughput sequencing data obtained from a large number of studies, the carbon and nitrogen cycling functional genes of surface soil microorganisms are predicted to obtain the abundances of the functional genes of the soil microbial carbon and nitrogen cycles. The z-scores are calculated for them to obtain the soil carbon cycle index and the soil nitrogen cycle index respectively. The climate conditions, physical and chemical properties of the soil, vegetation index, and land cover type are used as features to input into the model. Through training and learning with the machine learning random forest model, a model that can be used to predict the functions of soil microorganisms is obtained. The soil carbon cycle function index R 2 is 0.817, and the soil nitrogen cycle function index R 2 is 0.834.

[0066] The method for predicting the functions of soil carbon and nitrogen cycles based on machine learning in the present invention can quickly and accurately predict the soil functions in some areas, and then be used to evaluate the ecosystem functions of specific locations and guide the formulation of land management policies.

[0067] The present invention is not limited to the above example methods, and the protection scope of the present invention cannot be limited thereby. Any changes made on the basis of the technical solution according to the technical idea proposed by the present invention fall within the protection scope of the present invention.

Claims

1. A method for predicting soil microbial functions based on machine learning, comprising the following steps: S1 Collect relevant literature on soil microorganisms - climate change, extract literature information, and collect soil microbial gene sequencing serial numbers in the literature, as well as geographical information, climate variables, soil physical and chemical properties, vegetation indices, and land cover type data at the corresponding locations; S2 Download global raster files of climate variables, soil physical and chemical properties, vegetation indices, and land cover types from the database, and use global data to fill in relevant missing variable values; S3 According to the obtained gene sequencing data serial numbers, download microbial gene sequencing data, process the microbial gene sequencing data using bioinformatics analysis software, classify and annotate species, and use the FAPROTAX database to predict the relative abundances of microbial functional genes; integrate climate variables, soil physical and chemical properties, land cover types, altitude, and longitude and latitude information to construct a soil surface microbial function database; S4 Select functional genes reflecting the soil environmental carbon cycling ability and calculate the soil carbon cycling function index; select functional genes reflecting the nitrogen cycling ability in the soil environment and calculate the soil nitrogen cycling function index; Define the soil carbon cycling function index and the soil nitrogen cycling function index: perform z - score calculations on the selected functional genes for carbon cycling ability respectively. After calculating the z - scores of each functional gene, sum and average the z - scores of all selected functional genes to obtain the soil carbon cycling function index; Perform z - score calculations on the selected functional genes for nitrogen cycling ability respectively. After calculating the z - scores of each functional gene, sum and average the z - scores of all selected functional genes to obtain the soil nitrogen cycling function index; The specific formula for the z - score is as follows: z=(X - μ) / σ where X is the measured value of the sample feature, μ is the average value of a certain sample feature, and σ is the standard deviation of a certain sample feature; S5 Construct machine learning models for soil carbon cycling function and soil nitrogen cycling function respectively: The predictive variables are the soil carbon cycling function index and the soil nitrogen cycling function index; based on the soil surface microbial function database, determine the characteristic variables, divide them into a training set and a test set, and based on the machine learning random forest training model, adjust the parameters of the random forest to optimize the model effect.

2. The method for predicting soil microbial functions according to claim 1, wherein S1 is executed according to the following steps: S11 Retrieve literature topics including soil microorganisms, high - throughput sequencing, climate change, and land cover types, mark the literature containing soil microbial gene sequencing serial numbers, and obtain microbial sequencing data points; S12 Collect the characteristic variables of the microbial sequencing data points, including longitude and latitude, altitude, sampling year, sampling depth, sequencing instrument, amplification region, climate variables, soil physical and chemical properties, soil carbon and nitrogen content, vegetation index, and land cover type.

3. The method for predicting soil microbial functions according to claim 2, wherein, In S12, the climate variables include annual average temperature, annual precipitation, precipitation seasonality, temperature seasonality, diurnal temperature range of average temperature, aridity index, and potential evapotranspiration; the soil physical and chemical properties include soil pH, soil water content, soil bulk density, soil sand content, soil clay content, cation exchange capacity, and soil organic carbon, total nitrogen, and total phosphorus; the vegetation index refers to the normalized difference vegetation index NDVI; the land cover types include grassland, forest, shrub, bare land, and farmland.

4. The method for predicting soil microbial functions according to claim 1, wherein, In S5, the determined characteristic variables include: altitude, annual average temperature, annual precipitation, precipitation seasonality, temperature seasonality, diurnal temperature range of average temperature, aridity index, potential evapotranspiration, soil temperature, soil pH, soil water content, soil bulk density, soil sand content, soil clay content, cation exchange capacity, soil organic carbon, total nitrogen, total phosphorus, normalized difference vegetation index NDVI, and land cover type.

5. The method for predicting soil microbial functions according to claim 1, wherein, In S1, the gene sequencing serial number is the serial number of 16S rRNA sequencing data.

6. The method for predicting soil microbial functions according to claim 1, characterized in that, The functional genes for carbon cycling ability selected in S4 include Chitinolysis, cellulolysis, xylanolysis, ureolysis, hydrocarbon_degradation, fermentation, methanol_oxidation, and methylotrophy, a total of eight kinds. The functional genes for nitrogen cycling ability include nitrification, denitrification, nitrogen_fixation, nitrate_ammonification, nitrogen_respiration, and nitrate_reduction, a total of six kinds.