A soil health evaluation method based on soil microorganisms

By using soil microbial metagenomic sequencing and machine learning, a rapid and accurate method for soil health assessment was established, which solves the problems of long testing cycles and high costs in existing technologies and achieves efficient soil health assessment.

CN122493978APending Publication Date: 2026-07-31RES CENT FOR ECO ENVIRONMENTAL SCI THE CHINESE ACAD OF SCI
View PDF 0 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
RES CENT FOR ECO ENVIRONMENTAL SCI THE CHINESE ACAD OF SCI
Filing Date
2026-04-30
Publication Date
2026-07-31

AI Technical Summary

Technical Problem

Existing methods for assessing soil health rely on numerous physicochemical and biological indicators, have long testing cycles and high costs, and fail to effectively utilize soil microorganisms for high-throughput, low-cost assessment.

Method used

By employing soil microbial metagenomic sequencing technology and machine learning methods, this study identifies the species that contribute most to soil health by collecting soil samples, calculating soil health indices, obtaining relative microbial abundance, constructing machine learning models, and performing SHAP analysis. This leads to the establishment of a rapid and accurate method for assessing soil health.

Benefits of technology

It enables rapid, accurate, comprehensive, and high-throughput soil health assessment, reducing testing time and costs, and providing a more efficient assessment method for soil health monitoring.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN122493978A_ABST
    Figure CN122493978A_ABST
Patent Text Reader

Abstract

This invention discloses a soil health assessment method based on soil microorganisms, comprising: S1: collecting soil samples and measuring their physicochemical and biological indicators; S2: calculating the soil health index based on the soil's physicochemical and biological indicators; S3: obtaining the relative abundance of soil microorganisms using metagenomic sequencing technology; S4: establishing a machine learning model for predicting soil health, with the relative abundance of soil microorganisms as the input feature and the soil health index as the output feature; S5: identifying the top 40 species that contribute the most to the prediction of the soil health index using SHAP analysis; S6: reconstructing the soil health index prediction model using the relative abundance of the top 40 species that contribute the most to the prediction of the soil health index as explanatory variables.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of soil quality assessment, and more specifically, to a method for soil health assessment based on soil microorganisms. Background Technology

[0002] Traditional soil health assessment methods rely on physicochemical and biological indicators, such as soil organic matter content, available nutrient content, and soil enzyme activity. However, these methods typically involve numerous testing indicators, long testing cycles, and high costs. Current research indicates that soil microorganisms play a crucial role in maintaining and assessing soil health. However, developing high-throughput, low-cost, and accurate soil health assessment methods using soil microorganisms remains a challenge. Summary of the Invention

[0003] This invention aims to establish a rapid and efficient method for soil health assessment by utilizing soil microbial metagenomic sequencing technology and machine learning, so as to monitor changes in soil health in a timely manner and quickly assess soil health.

[0004] To achieve the above objectives, the present invention provides a soil health assessment method based on soil microorganisms, comprising:

[0005] S1: Collect soil samples and measure their physicochemical and biological indicators;

[0006] S2: Calculate the soil health index based on the soil's physicochemical and biological indicators;

[0007] S3: The relative abundance of soil microorganisms was obtained using metagenomic sequencing technology;

[0008] S4: Establish a machine learning model to predict soil health, with the relative abundance of soil microorganisms as the input feature and the soil health index as the output feature;

[0009] S5: The SHAP analysis method was used to identify the top 40 species that contributed the most to the prediction of soil health index;

[0010] S6: The relative abundance of the top 40 species that contribute the most to the prediction of soil health index will be identified as explanatory variables, and the prediction model of soil health index will be reconstructed.

[0011] In one embodiment of the present invention, optionally, in step S1, the soil physicochemical and biological indicators include pH, electrical conductivity, soil organic matter, total nitrogen, soluble organic carbon, soluble total nitrogen, ammonia nitrogen, nitrate nitrogen, available phosphorus, available potassium, β-glucosidase activity, urease activity, and phosphatase activity.

[0012] In one embodiment of the present invention, optionally, in step S1, at each sampling point, a field with an area of ​​more than 1 mu is selected, and a five-point sampling method is used for sampling, with an interval of 20 m between adjacent sampling points and the center point; and

[0013] The sampling depth was 0-15 cm. A soil auger was used for sampling to ensure that the amount of soil sample taken from the upper and lower layers was consistent, and the surface straw residue was removed.

[0014] In one embodiment of the present invention, optionally, in step S1, at least two samples are collected, wherein:

[0015] The first sample was stored at -80℃ for soil DNA extraction.

[0016] The second sample was stored at 4°C and used to analyze soil β-glucosidase activity, phosphatase activity, urease activity, ammonia nitrogen, nitrate nitrogen, soluble organic carbon, and soluble total nitrogen.

[0017] The remaining samples were air-dried and used to determine pH, electrical conductivity, soil organic matter, total nitrogen, available phosphorus, and available potassium.

[0018] In one embodiment of the present invention, step S2 may optionally include:

[0019] S21: For each physicochemical and biological indicator, the cumulative normal distribution function is used. Calculate the score for each indicator:

[0020] ,

[0021] μ and σ represent the measured value, mean, and standard deviation of each physicochemical and biological indicator, respectively;

[0022] S22: Calculate the soil health index, which is the average of the scores of all indicators, where:

[0023] Indicators beneficial to soil health include soil organic matter, total nitrogen, soluble organic carbon, soluble total nitrogen, ammonia nitrogen, nitrate nitrogen, available phosphorus, available potassium, β-glucosidase activity, urease activity, and phosphatase activity. The score for these indicators is calculated as: score = 100 × CND.

[0024] Indicators detrimental to soil health include electrical conductivity. The score for this type of indicator is calculated as follows: score = 100 × (1 – CND).

[0025] The indicators most favorable to the soil when the logarithmic value is in the middle include pH. When the pH value is between 6.3 and 7.2, the score is 100 points. When the pH value is ≤ 5.4 or ≥ 7.7, the score is 0 points. When the pH value is between the optimal range and the boundary, the score is calculated by linear interpolation.

[0026] In one embodiment of the present invention, step S3 may optionally include:

[0027] S31: All raw data obtained from metagenomic sequencing are quality controlled through SOAPnuke to obtain high-quality, clean data;

[0028] S32: Use MEGAHIT to assemble high-quality clean data to obtain fragment contigs, and select fragment contigs with a length ≥500 bp as the final assembly result;

[0029] S33: Use Prodigal to predict open reading frames from the assembled fragment contigs, cluster the open reading frames using CD-HIT, and select the longest gene in each cluster as the representative sequence to construct a non-redundant gene set;

[0030] S34: Use bowtie2 to map high-quality clean data to a non-redundant gene set, statistically match sequence data, and convert them into abundance quantification per million transcripts;

[0031] S35: Use DIAMOND to compare the non-redundant gene set with the NCBI-nr database for microbial classification annotation.

[0032] In one embodiment of the present invention, optionally, in step S33, when clustering open reading frames using CD-HIT, the parameters are set to sequence similarity of 95% and coverage of 90%.

[0033] In one embodiment of the present invention, optionally, in step S4, three machine learning models—random forest, support vector regression, and extreme gradient boosting—are used to construct a model for predicting soil health index based on the relative abundance of soil microbial species. The model uses the relative abundance of microbial species obtained from metagenomic sequencing as input data and the soil health index as the target variable. 80% of the input data is used as the training set, and the remainder is used as the test set. Hyperparameters are tuned using GridSearchCV, and the results are calculated based on the coefficient of determination R of the test set. 2 The root mean square error (RMSE) and mean absolute error (MAE) are used to evaluate model performance.

[0034] Currently, soil health assessment mainly relies on soil physicochemical and biological indicators, neglecting the role of soil microorganisms. This invention proposes a microorganism-based soil health assessment method, reducing testing time and costs, and establishing a rapid, accurate, comprehensive, and high-throughput method for soil health assessment. This provides a more efficient soil health assessment method for soil health monitoring. Attached Figure Description

[0035] To more clearly illustrate the technical solutions in the embodiments of the present invention or the prior art, the drawings used in the description of the embodiments or the prior art will be briefly introduced below. Obviously, the drawings described below are only some embodiments of the present invention. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.

[0036] Figure 1 This is a flowchart of a soil health assessment method based on soil microorganisms according to an embodiment of the present invention;

[0037] Figure 2 This is a schematic diagram of the spatial distribution of soil health index in 243 paddy field soils in Northeast China, according to an embodiment of the present invention.

[0038] Figure 3 The above are the prediction results of three machine learning models—random forest, support vector regression, and extreme gradient boosting—in one embodiment of the present invention.

[0039] Figure 4 This is a SHAP bar chart of 40 most important microorganisms for predicting soil health index based on a random forest model, according to an embodiment of the present invention.

[0040] Figure 5 This is a scatter plot of measured and predicted values ​​of the soil health index using a random forest model according to an embodiment of the present invention.

[0041] Figure 6 Scatter plot of measured and predicted values ​​of soil health index for bare land and paddy fields with different planting years;

[0042] Figure 7 This is a scatter plot showing the measured and predicted values ​​of the soil health index for paddy fields in a typical black soil region. Detailed Implementation

[0043] The technical solutions of the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of the present invention, and not all embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of the present invention.

[0044] Figure 1 This is a flowchart of a soil health assessment method based on soil microorganisms according to an embodiment of the present invention, as shown below. Figure 1 As shown, the soil health assessment method based on soil microorganisms provided by the present invention includes:

[0045] S1: Collect soil samples and measure their physicochemical and biological indicators;

[0046] S2: Calculate the soil health index based on the soil's physicochemical and biological indicators;

[0047] S3: The relative abundance of soil microorganisms was obtained using metagenomic sequencing technology;

[0048] S4: Establish a machine learning model to predict soil health, with the relative abundance of soil microorganisms as the input feature and the soil health index as the output feature;

[0049] S5: The SHAP analysis method was used to identify the top 40 species that contributed the most to the prediction of soil health index;

[0050] S6: The relative abundance of the top 40 species that contribute the most to the prediction of soil health index will be identified as explanatory variables, and the prediction model of soil health index will be reconstructed.

[0051] In one embodiment of the present invention, optionally, in step S1, the soil physicochemical and biological indicators include pH, electrical conductivity, soil organic matter, total nitrogen, soluble organic carbon, soluble total nitrogen, ammonia nitrogen, nitrate nitrogen, available phosphorus, available potassium, β-glucosidase activity, urease activity, and phosphatase activity.

[0052] In one embodiment of the present invention, optionally, in step S1, at each sampling point, a field with an area of ​​more than 1 mu is selected, and a five-point sampling method is used for sampling, with an interval of 20 m between adjacent sampling points and the center point; and

[0053] The sampling depth was 0-15 cm. A soil auger was used for sampling to ensure that the amount of soil sample taken from the upper and lower layers was consistent, and the surface straw residue was removed.

[0054] In one embodiment of the present invention, optionally, in step S1, at least two samples are collected, wherein:

[0055] The first sample was stored at -80℃ for soil DNA extraction.

[0056] The second sample was stored at 4°C and used to analyze soil β-glucosidase activity, phosphatase activity, urease activity, ammonia nitrogen, nitrate nitrogen, soluble organic carbon, and soluble total nitrogen.

[0057] The remaining samples were air-dried and used to determine pH, electrical conductivity, soil organic matter, total nitrogen, available phosphorus, and available potassium.

[0058] In one embodiment of the present invention, step S2 may optionally include:

[0059] S21: For each physicochemical and biological indicator, the cumulative normal distribution function is used. Calculate the score for each indicator:

[0060] ,

[0061] μ and σ represent the measured value, mean, and standard deviation of each physicochemical and biological indicator, respectively;

[0062] S22: Calculate the soil health index, which is the average of the scores of all indicators, where:

[0063] Indicators beneficial to soil health include soil organic matter, total nitrogen, soluble organic carbon, soluble total nitrogen, ammonia nitrogen, nitrate nitrogen, available phosphorus, available potassium, β-glucosidase activity, urease activity, and phosphatase activity. The score for these indicators is calculated as: score = 100 × CND.

[0064] Indicators detrimental to soil health include electrical conductivity. The score for this type of indicator is calculated as follows: score = 100 × (1 – CND).

[0065] The indicators most favorable to the soil when the logarithmic value is in the middle include pH. When the pH value is between 6.3 and 7.2, the score is 100 points. When the pH value is ≤ 5.4 or ≥ 7.7, the score is 0 points. When the pH value is between the optimal range and the boundary, the score is calculated by linear interpolation.

[0066] In one embodiment of the present invention, step S3 may optionally include:

[0067] S31: All raw data obtained from metagenomic sequencing are quality controlled through SOAPnuke to obtain high-quality, clean data;

[0068] S32: Use MEGAHIT to assemble high-quality clean data to obtain fragment contigs, and select fragment contigs with a length ≥500 bp as the final assembly result;

[0069] S33: Use Prodigal to predict open reading frames from the assembled fragment contigs, cluster the open reading frames using CD-HIT, and select the longest gene in each cluster as the representative sequence to construct a non-redundant gene set;

[0070] S34: Use bowtie2 to map high-quality clean data to a non-redundant gene set, statistically match sequence data, and convert them into abundance quantification per million transcripts;

[0071] S35: Use DIAMOND to compare the non-redundant gene set with the NCBI-nr database for microbial classification annotation.

[0072] In one embodiment of the present invention, optionally, in step S33, when clustering open reading frames using CD-HIT, the parameters are set to sequence similarity of 95% and coverage of 90%.

[0073] In one embodiment of the present invention, optionally, in step S4, three machine learning models—random forest, support vector regression, and extreme gradient boosting—are used to construct a model for predicting soil health index based on the relative abundance of soil microbial species. The model uses the relative abundance of microbial species obtained from metagenomic sequencing as input data and the soil health index as the target variable. 80% of the input data is used as the training set, and the remainder is used as the test set. Hyperparameters are tuned using GridSearchCV, and the results are calculated based on the coefficient of determination R of the test set. 2 The root mean square error (RMSE) and mean absolute error (MAE) are used to evaluate model performance.

[0074] Another specific embodiment of the present invention:

[0075] Soil samples from paddy fields were collected annually from June to September 2021 to 2024. This example collected soil samples from 243 paddy fields in Northeast China. At each sampling point, fields larger than one mu (approximately 0.16 acres) were selected, and a five-point sampling method was used, with an interval of approximately 20 m between adjacent sampling points and the center point. Sampling was conducted during the flooded period, avoiding sampling at compost sites, field ridges, ditches, and areas with unusual terrain. The sampling depth was 0-15 cm, using a soil auger to ensure consistent soil sample volume from all layers, and removing surface straw residue. Each soil sample consisted of five equal subsamples, thoroughly mixed while wearing disposable gloves, and roots and large stones were removed. The mixed samples were divided into three equal portions, placed in sterile self-sealing bags, and immediately transported to the laboratory on ice packs. One sample was stored at -80℃ for soil DNA extraction; the second sample was stored at 4℃ for analyzing soil β-glucosidase (BG) activity, phosphatase (PHOS) activity, urease activity, and ammonia nitrogen (NH4) content. +-N), nitrate nitrogen (NO3) − Soil samples were analyzed for soluble organic carbon (DOC) and soluble total nitrogen (DTN). The remaining samples were air-dried and used to determine pH, electrical conductivity (EC), soil organic matter (SOM), total nitrogen (TN), available phosphorus (AP), and available potassium (AK). A total of 243 paddy field soil samples were collected in this study.

[0076] The indicators included in the soil health assessment are: pH, electrical conductivity, organic matter, total nitrogen, soluble organic carbon, soluble total nitrogen, ammonia nitrogen, nitrate nitrogen, available phosphorus, available potassium, β-glucosidase activity, phosphatase activity, and urease activity. Soil pH was determined using the potentiometric method (HJ 962-2018) for soil pH determination; soil electrical conductivity was determined using the electrode method (HJ 802-2016) for soil electrical conductivity determination; soil organic matter was determined using the standard (NY / T 1121.6-2006) for soil testing, part 6: determination of soil organic matter; total nitrogen was determined using an elemental analyzer; soluble organic carbon and soluble total nitrogen were determined using K2SO4 extract; ammonia nitrogen and nitrate nitrogen were determined using the potassium chloride solution extraction-spectrophotometric method (HJ 634—2012) for soil ammonia nitrogen, nitrite nitrogen, and nitrate nitrogen determination; available phosphorus was determined using the standard (NY / T 1121.7-2014) for soil testing, part 7: determination of available phosphorus in soil; and soil organic matter was determined using the standard (NY / T 889-2004). The available potassium content in soil was determined by measuring available potassium. β-glucosidase and phosphatase activities were measured using a fluorescence method. Urease activity was measured using a soil urease (S-UE) activity assay kit (Beijing Solarbio Science & Technology Co., Ltd.). DNA was extracted from 0.5 g soil samples using the FastDNA Spin Kit for Soil (MP Biomedicals, CA, USA). Metagenomic sequencing was performed using the DNBSEQ-T7 sequencing platform.

[0077] For each physicochemical and biological indicator, the cumulative normal distribution function is used. Calculate the score for each indicator:

[0078] ,

[0079] μ and σ represent the measured value, mean, and standard deviation of each physicochemical and biological indicator, respectively;

[0080] Calculate the soil health index, which is the average of the scores of all indicators, where:

[0081] Indicators beneficial to soil health include soil organic matter, total nitrogen, soluble organic carbon, soluble total nitrogen, ammonia nitrogen, nitrate nitrogen, available phosphorus, available potassium, β-glucosidase activity, urease activity, and phosphatase activity. The score for these indicators is calculated as: score = 100 × CND.

[0082] Indicators detrimental to soil health include electrical conductivity. The score for this type of indicator is calculated as follows: score = 100 × (1 – CND).

[0083] The indicators most favorable to the soil when the logarithmic value is in the middle include pH. When the pH value is between 6.3 and 7.2, the score is 100 points. When the pH value is ≤ 5.4 or ≥ 7.7, the score is 0 points. When the pH value is between the optimal range and the boundary, the score is calculated by linear interpolation.

[0084] All raw data from metagenomic sequencing were quality-controlled using SOAPnuke (v1.5.2) to obtain high-quality clean data. MEGAHIT (v1.2.9) was then used to assemble the high-quality clean data into contigs, with contigs ≥ 500 bp selected as the final assembly. Prodigal (v2.6.3) was used to predict open reading frames (ORFs) from the assembled contigs. CD-HIT (v4.8.1) was used to cluster the ORFs (with parameters set to 95% sequence similarity and 90% coverage), and the longest gene in each cluster was selected as the representative sequence to construct a non-redundant gene set. bowtie2 (v2.5.4) was used to map the high-quality clean data to the non-redundant gene set, and the matched sequence data were statistically analyzed and converted to abundance quantification per million transcripts (TPM). Finally, DIAMOND (v2.1.11) was used to compare the non-redundant gene set with the NCBI-nr database for microbial classification annotation.

[0085] In this embodiment, the spatial distribution of the paddy field soil health index is as follows: Figure 2 As shown, Figure 2 The spatial distribution of soil health index in 243 paddy soil samples in Northeast China is shown.

[0086] Three machine learning models—Random Forest, Support Vector Regression, and Extreme Gradient Boosting (XGBoost)—were employed to construct a model for predicting soil health index based on the relative abundance of soil microbial species. The model uses the relative abundance of microbial species obtained from metagenomic sequencing as input data and the soil health index as the target variable. The dataset was divided into a training set (80%) and a test set (20%). Hyperparameters were tuned using GridSearchCV, and the results were calculated based on the coefficient of determination (R²) of the test set. 2 The model performance is evaluated using the root mean square error (RMSE) and mean absolute error (MAE).

[0087] This study compares the performance of three machine learning models—random forest, support vector regression, and extreme gradient boosting—in predicting soil health indices based on soil microorganisms. The results show that the random forest model has the best predictive performance, with a test set determination coefficient (R²) of 0.65, and root mean square error (RMSE) and mean absolute error (MAE) of 7.71 and 6.23, respectively. The support vector regression model performs second best, with a test set determination coefficient (R²) of 0.62, and RMSE and MAE of 8.02 and 6.51, respectively. The extreme gradient boosting model has the lowest performance, with a test set determination coefficient (R²) of only 0.60, and RMSE and MAE of 8.21 and 6.60, respectively. Figure 3 (As shown). Based on the above comparison, the random forest model was selected for the subsequent identification of indicator species of paddy field soil health.

[0088] Figure 3 In the middle, R 2 Test RMSE Test , and MAE Test These represent the coefficients of determination (R²) between the soil health index measured in the test set and the predicted soil health index, respectively. 2 ), root mean square error (RMSE) and mean absolute error (MAE).

[0089] The SHAP analysis method based on a random forest model was used to identify the top 40 species that contributed most to the prediction of soil health index. To further verify whether these species alone can accurately predict soil health index, the random forest model was reconstructed using the training and test sets mentioned above. If the reconstructed model still maintains good predictive performance, it indicates that these 40 species can effectively assess soil health status and are therefore identified as indicator species for paddy field soil health.

[0090] Based on the SHAP value of the random forest model, a total of 40 microbial species most important for predicting soil health index were identified. Figure 4 To evaluate the predictive performance of these 40 species on soil health index, a random forest model was reconstructed using their relative abundance as explanatory variables. The results show that the model's test set R... 2 The value is 0.62 ( Figure 5 ), and R obtained based on data at the full species level 2 The value (0.65) is comparable, indicating that these 40 species can accurately predict the soil health index of paddy fields, and therefore can be identified as indicator species of paddy field soil health.

[0091] To further validate the general applicability of the constructed random forest model based on the 40 identified indicator microorganisms of paddy field soil health, external datasets from two independent studies from the NCBI database were used for further validation.

[0092] The first dataset included bare soil (i.e., in its original state) and paddy field soils (n=63) that had been planted for 2, 4, 6, 8, 11, 12, 20, and 23 years. The model demonstrated acceptable accuracy in predicting soil health indices, with R0. 2 The value is 0.312 (p < 0.001). Figure 6 ).

[0093] The second dataset contained 120 samples from a typical black soil region. The random forest model performed well in predicting soil health indices, with an R² value of 0.605 (p < 0.001). Figure 7 ).

[0094] This embodiment identified 40 indicator microorganisms that predict soil health in paddy fields. Furthermore, the reliability of these indicator microorganisms in predicting soil health across both temporal and spatial scales was validated using independent datasets. Therefore, a method for rapidly assessing soil health based on soil microorganisms was established.

[0095] Currently, soil health assessment mainly relies on soil physicochemical and biological indicators, neglecting the role of soil microorganisms. This invention proposes a microorganism-based soil health assessment method, reducing testing time and costs, and establishing a rapid, accurate, comprehensive, and high-throughput method for soil health assessment. This provides a more efficient soil health assessment method for soil health monitoring.

[0096] Those skilled in the art will understand that the accompanying drawings are merely schematic diagrams of one embodiment, and the modules or processes shown in the drawings are not necessarily essential for implementing the present invention.

[0097] Those skilled in the art will understand that the modules in the apparatus of the embodiments can be distributed in the apparatus of the embodiments as described in the embodiments, or they can be located in one or more devices different from this embodiment with corresponding changes. The modules of the above embodiments can be combined into one module, or they can be further divided into multiple sub-modules.

[0098] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention, and not to limit them; although the present invention has been described in detail with reference to the foregoing embodiments, those skilled in the art should understand that modifications can still be made to the technical solutions described in the foregoing embodiments, or equivalent substitutions can be made to some of the technical features; and these modifications or substitutions do not cause the essence of the corresponding technical solutions to deviate from the spirit and scope of the technical solutions of the embodiments of the present invention.

Claims

1. A soil health evaluation method based on soil microorganisms, characterized by, include: S1: Collect soil samples and measure their physicochemical and biological indicators; S2: Calculate the soil health index based on the soil's physicochemical and biological indicators; S3: The relative abundance of soil microorganisms was obtained using metagenomic sequencing technology; S4: Establish a machine learning model to predict soil health, with the relative abundance of soil microorganisms as the input feature and the soil health index as the output feature; S5: The SHAP analysis method was used to identify the top 40 species that contributed the most to the prediction of soil health index; S6: The relative abundance of the top 40 species that contribute the most to the prediction of soil health index will be identified as explanatory variables, and the prediction model of soil health index will be reconstructed.

2. The soil-microbe-based soil health evaluation method according to claim 1, characterized by, In step S1, soil physicochemical and biological indicators include pH, electrical conductivity, soil organic matter, total nitrogen, soluble organic carbon, soluble total nitrogen, ammonia nitrogen, nitrate nitrogen, available phosphorus, available potassium, β-glucosidase activity, urease activity, and phosphatase activity.

3. The soil-microbe-based soil health evaluation method according to claim 1, characterized by, In step S1, at each sampling point, a field with an area of ​​more than 1 mu is selected, and a five-point sampling method is used for sampling. The interval between adjacent sampling points and the center point is 20 m. as well as The sampling depth was 0-15 cm. A soil auger was used for sampling to ensure that the amount of soil sample taken from the upper and lower layers was consistent, and the surface straw residue was removed.

4. The soil-microbe-based soil health evaluation method according to claim 1, characterized by, In step S1, at least two samples are collected, wherein: The first sample was stored at -80℃ for soil DNA extraction. The second sample was stored at 4°C and used to analyze soil β-glucosidase activity, phosphatase activity, urease activity, ammonia nitrogen, nitrate nitrogen, soluble organic carbon, and soluble total nitrogen. The remaining samples were air-dried and used to determine pH, electrical conductivity, soil organic matter, total nitrogen, available phosphorus, and available potassium.

5. The soil health assessment method based on soil microorganisms according to claim 1, characterized in that, Step S2 includes: S21: For each physicochemical and biological indicator, the cumulative normal distribution function is used. Calculate the score for each indicator: , μ and σ represent the measured value, mean, and standard deviation of each physicochemical and biological indicator, respectively; S22: Calculate the soil health index, which is the average of the scores of all indicators, where: Indicators beneficial to soil health include soil organic matter, total nitrogen, soluble organic carbon, soluble total nitrogen, ammonia nitrogen, nitrate nitrogen, available phosphorus, available potassium, β-glucosidase activity, urease activity, and phosphatase activity. The score for these indicators is calculated as: score = 100 × CND. Indicators detrimental to soil health include electrical conductivity. The score for this type of indicator is calculated as follows: score = 100 × (1 – CND). The indicators most favorable to the soil when the logarithmic value is in the middle include pH. When the pH value is between 6.3 and 7.2, the score is 100 points. When the pH value is ≤ 5.4 or ≥ 7.7, the score is 0 points. When the pH value is between the optimal range and the boundary, the score is calculated by linear interpolation.

6. The soil health assessment method based on soil microorganisms according to claim 1, characterized in that, Step S3 includes: S31: All raw data obtained from metagenomic sequencing are quality controlled using SOAPnuke to obtain high-quality, clean data; S32: Use MEGAHIT to assemble high-quality clean data to obtain fragment contigs, and select fragment contigs with a length ≥ 500bp as the final assembly result; S33: Use Prodigal to predict open reading frames from the assembled fragment contigs, cluster the open reading frames using CD-HIT, and select the longest gene in each cluster as the representative sequence to construct a non-redundant gene set; S34: Use bowtie2 to map high-quality clean data to a non-redundant gene set, statistically match sequence data, and convert them into abundance quantification per million transcripts; S35: Use DIAMOND to compare the non-redundant gene set with the NCBI-nr database for microbial classification annotation.

7. The soil health assessment method based on soil microorganisms according to claim 6, characterized in that, In step S33, when clustering open reading frames using CD-HIT, the parameters are set to sequence similarity of 95% and coverage of 90%.

8. The soil health assessment method based on soil microorganisms according to claim 1, characterized in that, In step S4, three machine learning models, random forest, support vector regression and extreme gradient boosting, are used to build a model for predicting soil health index based on soil microbial species relative abundance. The model takes the microbial species relative abundance obtained by metagenomic sequencing as input data and the soil health index as the target variable. 80% of the input data is used as the training set and the rest is used as the test set. GridSearchCV is used for hyperparameter tuning, and the model performance is evaluated according to the determination coefficient R 2 , root mean square error RMSE and mean absolute error MAE of the test set.