Evaluation method and system based on soil buffer capability

Through high-throughput sequencing and neural network models, combined with genetic algorithms, analyzing soil microbial community composition and diversity, the problem of traditional evaluation methods ignoring soil microorganisms is solved, and more accurate soil buffering capacity assessment and dynamic monitoring are achieved.

CN119990918AActive Publication Date: 2025-05-13SHANGHAI ACAD OF AGRI SCI

Patent Information

Application Number
CN202510458408.1
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-04-14
Publication Date
2025-05-13
Estimated Expiration
2045-04-14

AI Technical Summary

Technical Problem

Traditional soil buffering capacity assessment methods ignore the key factor of soil microorganisms, which affects the accuracy of the assessment.

Method used

By performing DNA extraction and high-throughput sequencing of soil microbial samples, the microbial community composition and diversity are analyzed, and the soil buffering capacity is predicted in combination with genetic algorithms and neural network models.

Benefits of technology

This method can more accurately evaluate soil buffering capacity, improve the accuracy and comprehensiveness of the assessment, dynamically monitor and predict changes in soil buffering capacity, and assist in the formulation of soil management measures.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119990918A_ABST
    Figure CN119990918A_ABST
Patent Text Reader

Abstract

The invention provides an evaluation method and system based on soil buffer capacity, and relates to the technical field of soil evaluation.The method comprises the steps that a microbial community diversity index is coded to serve as a genotype of a genetic algorithm; randomly generating an initial population, wherein each individual represents one environment factor combination; calculating the fitness value of each individual, carrying out selection, crossover and mutation operations, and repeating genetic operations until a termination condition is met so as to obtain a final key environment factor combination; obtaining microorganism activity data according to the soil respiration rate in the microorganism sample; and according to the final key environment factor combination, the microbial activity data and the trained neural network model, predicting a soil buffer capability value. According to the method, the buffer capacity of the soil can be evaluated more accurately.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of soil evaluation, and in particular to an evaluation method and system based on soil buffering capacity. Background Art

[0002] Some traditional methods for assessing soil buffering capacity rely on laboratory chemical analysis, which usually involves the measurement of multiple chemical indicators of the soil, such as pH, organic matter content and cation exchange capacity. These indicators can indirectly reflect certain physical and chemical properties of the soil and are therefore used to infer the buffering capacity of the soil.

[0003] However, although the above chemical analysis method can provide useful information about some physical and chemical properties of soil, it ignores the key factor of soil microorganisms. Soil microorganisms play an important role in soil ecosystems, including nutrient cycling, organic matter decomposition, soil structure formation, etc. These microbial activities directly affect the buffering capacity of soil because they are involved in many key biochemical processes. Therefore, traditional chemical analysis methods may not be able to capture these dynamic changes, thus affecting the accuracy of the assessment. Summary of the invention

[0004] The technical problem to be solved by the present invention is to provide an evaluation method and system based on soil buffering capacity, which can more accurately evaluate the soil buffering capacity.

[0005] In order to solve the above technical problems, the technical solution of the present invention is as follows: In a first aspect, a soil buffer capacity assessment method is provided, the method comprising: Step 1, extracting DNA from each microbial sample to obtain DNA data; sequencing the DNA data using high-throughput sequencing technology to obtain sequencing data; Step 2, analyzing the community composition of soil microorganisms based on the sequencing data; and constructing a microbial community diversity index based on the community composition of soil microorganisms; Step 3, encode the diversity index of the microbial community as the genotype of the genetic algorithm; randomly generate the initial population, each individual represents a combination of environmental factors; calculate the fitness value of each individual, and perform selection, crossover and mutation operations, and repeat the genetic operation until the termination condition is met to obtain the final key environmental factor combination; Step 4, obtaining microbial activity data based on soil respiration rate in microbial samples; Step 5, predicting the soil buffering capacity value based on the final combination of key environmental factors, microbial activity data and the trained neural network model; Step 6: Classify the soil into different buffering capacity levels according to the soil buffering capacity value.

[0006] Furthermore, before DNA is extracted from each microbial sample to obtain DNA data, the following steps are also included: Divide the target acquisition area into circular grids; obtain remote sensing image data of the target area from remote sensing satellites; Preprocess the remote sensing image data, use the spectral characteristics, texture characteristics and shape characteristics of the preprocessed remote sensing image, automatically identify the vegetation type through the image classification algorithm, and mark and classify the vegetation type in each grid to obtain the vegetation type distribution; According to the distribution of vegetation types, the vegetation coverage is estimated by calculating the difference in spectral characteristics between vegetation and bare soil; According to remote sensing image data, the density of vegetation in the grid is obtained by calculating the number of vegetation pixels per unit area; In each circular grid, a corresponding collection radius is determined for each grid according to the distribution of vegetation types, coverage, and density of vegetation; In each grid, a specific collection point is determined according to the collection radius, and microbial samples are obtained through the collection point.

[0007] Furthermore, environmental factors include soil structure, density, porosity, soil pH, type, quantity, activity of soil microorganisms, topography and climatic conditions.

[0008] Furthermore, the community composition of soil microorganisms is analyzed based on the sequencing data; based on the community composition of soil microorganisms, a microbial community diversity index is constructed, including: Acquire paired sequencing data, and perform preprocessing to obtain paired sequencing data after preprocessing, namely, Read 1 and Read 2; Based on all the sequences in Read 1 and Read 2, a de Bruijn graph is constructed. Each node in the de Bruijn graph represents a k-mer of fixed length, i.e., k consecutive nucleotides, and the edge represents the connection relationship between k-mers. Using the characteristics of the de Bruijn graph, search for k-mers shared in Read 1 and Read 2; traverse the de Bruijn graph and compare the k-mer paths in different reads to identify overlapping nucleotide sequences, i.e., overlapping regions; match nucleotides at corresponding positions in Read 1 and Read 2 based on the overlapping regions; After a successful match, the two reads were merged in the overlapping region to be assembled into a single sequence; Single sequences were clustered to form operational taxonomic units, representative sequences of each OTU were selected, and the corresponding representative sequences were compared with known microbial reference databases to determine the species classification information of each OTU; According to the species classification information, the number of sequences of different species in each sample is calculated, and the number of sequences of different species is converted into relative abundance to generate a species abundance table; The species abundance table was used to calculate species richness index and diversity index, including Shannon diversity index and Simpson diversity index.

[0009] Furthermore, the fitness value of each individual is calculated, including: The Shannon diversity index and Simpson diversity index are determined according to the richness of microbial species and microbial communities in the soil; the contribution value of soil microbial diversity is determined according to the Shannon diversity index and Simpson diversity index; Determine the Pearson correlation coefficient between each environmental factor and soil buffering capacity; determine the correlation contribution value based on each environmental factor and its value in the current individual; For each microbial activity data sample, determine its corresponding expected value and the standard deviation of the microbial activity data; according to each microbial activity data sample and the corresponding expected value and the standard deviation of the microbial activity data, obtain the standardized deviation of the sample; According to the standardized deviation of the sample and the soil respiration rate of the sample, the contribution value of microbial activity data is obtained; The correlation contribution value, the standardized deviation of the sample, and the contribution value of the microbial activity data are fused to obtain the fitness value of each individual.

[0010] Furthermore, the fitness value of each individual is calculated, and selection, crossover and mutation operations are performed, and the genetic operation is repeated until the termination condition is met to obtain the final key environmental factor combination, including: An initial population is randomly generated, and each individual in the population represents a combination of a set of key environmental factors; Calculate the fitness value of each individual, and select the corresponding individual for crossover operation according to the individual's fitness value to generate new offspring individuals; The offspring individuals are mutated, and the new individuals generated by selection, crossover and mutation operations form a new generation of population until the preset number of evolutionary generations is reached to obtain the final individual in the current population as the final solution to the problem; based on the final solution, the final combination of key environmental factors is obtained.

[0011] Furthermore, based on the final combination of key environmental factors, microbial activity data and trained neural network model, the soil buffering capacity value is predicted, including: Select features related to soil buffering capacity from the key environmental factor combination and microbial activity data, and combine the selected features into a feature vector; The feature vector is input into the trained neural network model to obtain the soil buffering capacity value.

[0012] In a second aspect, an evaluation system based on soil buffering capacity includes: An extraction module is used to extract DNA from each microbial sample to obtain DNA data; and sequence the DNA data using high-throughput sequencing technology to obtain sequencing data; Constructing a module for analyzing the community composition of soil microorganisms based on sequencing data; constructing a microbial community diversity index based on the community composition of soil microorganisms; The optimization module is used to encode the diversity index of the microbial community as the genotype of the genetic algorithm; randomly generate the initial population, each individual represents a combination of environmental factors; calculate the fitness value of each individual, and perform selection, crossover and mutation operations, and repeat the genetic operations until the termination condition is met to obtain the final key environmental factor combination; An acquisition module is used to obtain microbial activity data based on soil respiration rate in microbial samples; A prediction module is used to predict the soil buffering capacity value based on the final combination of key environmental factors, microbial activity data, and the trained neural network model; The classification module is used to classify the soil into different buffering capacity levels according to the soil buffering capacity value.

[0013] According to a third aspect, a computing device includes: one or more processors; The storage device is used to store one or more programs. When the one or more programs are executed by the one or more processors, the one or more processors implement the method described.

[0014] In a fourth aspect, a computer-readable storage medium stores a program, and when the program is executed by a processor, the method described is implemented.

[0015] The above solution of the present invention includes at least the following beneficial effects: By analyzing the community composition and diversity of soil microorganisms, the buffering capacity of soil can be evaluated more accurately. Microorganisms are key components of soil ecosystems, and their community structure and diversity have an important impact on soil buffering capacity. Therefore, incorporating microbial data into the evaluation system can significantly improve the accuracy and comprehensiveness of the evaluation.

[0016] By determining the combination of key environmental factors through genetic algorithms and combining them with microbial activity data, this method can dynamically monitor and predict changes in soil buffering capacity.

[0017] By utilizing high-throughput sequencing technology and neural network models, this method realizes the intelligent and automated assessment of soil buffering capacity, which can not only improve the assessment efficiency but also reduce errors caused by human factors.

[0018] Dividing soil into different buffering capacity levels according to its buffering capacity value will help to formulate targeted soil management measures. Different management strategies can be adopted for soils of different levels. BRIEF DESCRIPTION OF THE DRAWINGS

[0019] Figure 1 It is a flow chart of an evaluation method based on soil buffering capacity provided by an embodiment of the present invention.

[0020] Figure 2 It is a schematic diagram of an evaluation system based on soil buffering capacity provided by an embodiment of the present invention. DETAILED DESCRIPTION

[0021] The exemplary embodiments of the present disclosure will be described in more detail below with reference to the accompanying drawings. Although the exemplary embodiments of the present disclosure are shown in the accompanying drawings, it should be understood that the present disclosure can be implemented in various forms and should not be limited by the embodiments set forth herein. On the contrary, these embodiments are provided to enable a more thorough understanding of the present disclosure and to fully convey the scope of the present disclosure to those skilled in the art.

[0022] like Figure 1 As shown, an embodiment of the present invention provides an evaluation method based on soil buffering capacity, the method comprising the following steps: Step 1, extracting DNA from each microbial sample to obtain DNA data; sequencing the DNA data using high-throughput sequencing technology to obtain sequencing data; Step 2, analyzing the community composition of soil microorganisms based on the sequencing data; and constructing a microbial community diversity index based on the community composition of soil microorganisms; Step 3, encode the microbial community diversity index as the genotype of the genetic algorithm; randomly generate the initial population, each individual represents a combination of environmental factors, including soil structure, density, porosity, soil pH, soil microbial species, quantity, activity, topography and climate conditions; calculate the fitness value of each individual, and perform selection, crossover and mutation operations, repeat the genetic operation until the termination condition is met to obtain the final key environmental factor combination; Step 4, obtaining microbial activity data based on soil respiration rate in microbial samples; Step 5, predicting the soil buffering capacity value based on the final combination of key environmental factors, microbial activity data and the trained neural network model; Step 6: Classify the soil into different buffering capacity levels according to the soil buffering capacity value.

[0023] In the embodiment of the present invention, the genetic information of soil microorganisms can be accurately obtained through DNA extraction and high-throughput sequencing technology. This technology has the characteristics of high throughput and high efficiency, and can process a large number of samples at one time, thereby improving the efficiency and accuracy of the evaluation method. By analyzing the sequencing data, the community structure and diversity of soil microorganisms can be fully understood. The construction of the microbial community diversity index helps to quantitatively evaluate the richness and uniformity of the microbial community, thereby more accurately reflecting the health status and buffering capacity of the soil ecosystem. The microbial community diversity index is encoded using a genetic algorithm, and the fitness function value is calculated to optimize the combination of environmental factors, which can efficiently find the key environmental factors that affect the soil buffering capacity. This method not only takes into account the comprehensive effects of multiple environmental factors, but also can quickly find the best factor combination through the optimization algorithm. Obtaining microbial activity data by measuring the soil respiration rate in the microbial sample helps to understand the metabolic activity and vitality of soil microorganisms. These data are one of the important indicators for evaluating the soil buffering capacity and can reflect the dynamic changes and recovery capacity of the soil ecosystem. Combining the final combination of key environmental factors, microbial activity data and the trained neural network model, the soil buffering capacity value can be accurately predicted. This method comprehensively considers multiple influencing factors and improves the accuracy and reliability of the prediction results. Dividing the soil into different buffering capacity levels according to the soil buffering capacity value is helpful to formulate targeted soil management measures and strategies. This grading method can intuitively reflect the strength of the soil buffering capacity.

[0024] In a preferred embodiment of the present invention, before extracting DNA from each microbial sample to obtain DNA data, the process further includes: Divide the target collection area into circular grids; obtain remote sensing image data of the target area from remote sensing satellites, including: using GIS (geographic information system) software to open the map of the target area; setting the radius of the circular grid according to research needs and area size, such as 500 meters, 1 kilometer, etc.; using the grid generation tool of GIS software to create equally spaced circular grids on the map to ensure that the grid covers the entire target area; selecting appropriate remote sensing satellite data sources, such as Landsat, Sentinel-2, etc.; querying and downloading the corresponding remote sensing image data according to the geographical location of the target area and the required time range, and decompressing and converting the downloaded remote sensing images.

[0025] The remote sensing image data is preprocessed, and the spectral characteristics, texture characteristics and shape characteristics of the preprocessed remote sensing images are used to automatically identify the vegetation types through image classification algorithms. The vegetation types in each grid are labeled and classified to obtain the vegetation type distribution. Specifically, the remote sensing image is opened using remote sensing image processing software (such as ENVI, ERDASImagine, etc.) and preprocessing operations are performed, including atmospheric correction, geometric correction, radiation calibration, etc., to improve image quality; spectral features are extracted from the preprocessed remote sensing images. These features include reflectance values ​​of each band, vegetation indexes (such as NDVI, EVI, etc.), water body indexes (such as NDWI), etc., which can reflect the spectral reflectance and absorption characteristics of the objects; the texture features of the image are extracted using the gray level co-occurrence matrix (GLCM), which can describe the surface structure and roughness of the objects.

[0026] The shape features of the objects are extracted through edge detection methods, such as area, perimeter, aspect ratio, complexity, etc. These features help to distinguish the types of objects with different geometric shapes; based on prior knowledge or field survey data, representative object sample points (such as forests, grasslands, water bodies, etc.) are selected and assigned corresponding category labels; a random forest classifier is trained using the extracted spectral, texture and shape features as input and the corresponding category labels as output. During the training process, the model performance is optimized by adjusting parameters such as the number of decision trees and the maximum depth; the trained random forest model is applied to the entire remote sensing image, and for each pixel Carry out classification prediction to obtain preliminary vegetation type classification results; by setting the area threshold, remove small patches in the classification results that are too small, may be noise or misclassified, so as to improve the accuracy of the classification results; use the morphological filtering algorithm to smooth the boundaries of the classification results to reduce the appearance of jagged boundaries and fine patches; evaluate the accuracy of the classification results by constructing a confusion matrix and calculating indicators such as overall accuracy (OA), producer accuracy (PA), and user accuracy (UA); convert the classification results from pixel level to vector data (such as Shapefile format) or raster data (such as GeoTIFF format).

[0027] According to the distribution of vegetation types, the vegetation coverage is estimated by calculating the difference in spectral characteristics between vegetation and bare soil. Specifically, the following steps are performed: determine the size of the grid, divide the study area into grids of specified size, assign an identifier to each pixel based on the previous vegetation type classification results, for example, vegetation pixels are marked as 1 and non-vegetation pixels are marked as 0; for each grid, traverse all pixels in it and count the number of pixels marked as vegetation (1) and non-vegetation (0); compare the vegetation and bare soil in specific bands (such as red light, near-infrared bands, etc.); ) reflectance differences. Based on these differences, set an appropriate threshold to further distinguish vegetation and non-vegetation pixels. For example, the NDVI (normalized difference vegetation index) value can be used. Usually, the NDVI value of vegetation is higher than that of bare soil. Apply the set threshold to reclassify the pixels in the grid and update the number of vegetation and non-vegetation pixels. For each grid, use the formula "vegetation coverage = number of vegetation pixels / total number of pixels" to calculate the vegetation coverage, and save the vegetation coverage value of each grid to a new data layer or data table.

[0028] According to the remote sensing image data, the density of vegetation in the grid is obtained by calculating the number of vegetation pixels per unit area, which includes: superimposing the divided grid data with the loaded remote sensing image data; for each grid, traverse all the pixels in it and count the number of pixels identified as vegetation, which can be done by checking the classification label of each pixel (such as vegetation marked as 1 and non-vegetation as 0 in the previous step). For standardized comparison, the number of vegetation pixels can be converted into density per unit area, such as the number of vegetation pixels per square kilometer or per hectare. According to historical data, different thresholds are set to divide the density of vegetation. For example, three thresholds of high, medium and low can be set to represent dense, medium and sparse vegetation states respectively. For each grid, it is classified as dense, medium or sparse according to the density of vegetation pixels per unit area, which can be done by comparing the density of vegetation pixels with the set threshold. Create a new data layer or attribute table to store the vegetation density assessment results of each grid, and fill the vegetation density classification results of each grid into the newly created data layer or attribute table.

[0029] In each circular grid, a corresponding collection radius is determined for each grid according to the distribution of vegetation types, coverage, and density of vegetation. The calculation formula is: ; in, Indicates the basic radius, which is the initial radius set according to the total area of ​​the study area and the budget (the default value can be set to 100m), for example, R b=50m~200m (need to be adjusted according to the actual area size); Represents the vegetation type weight coefficient, which adjusts the influence of vegetation type on the radius. a =0.4~0.6 (default 0.5); Indicates the value of vegetation type, the impact of different vegetation types on soil heterogeneity (needs to be defined in advance); forest = 1.2; shrub = 1.0; grassland = 0.8; bare soil = 0.5 (higher soil homogeneity); β represents the coverage adjustment coefficient, which controls the correction effect of vegetation coverage on radius, β = 0.2~0.3 (default 0.25); Indicates vegetation coverage, the proportion of vegetation pixels in the grid (unit: percentage), ranging from 0% (bare soil) to 100% (complete coverage); Indicates the density adjustment coefficient, reflecting the adjustment weight of vegetation density to radius, γ = 0.1 ~ 0.2 (default 0.15); Indicates the density of vegetation, which is divided into levels by the density of vegetation pixels per unit area: dense = 3; medium = 2; sparse = 1 (the higher the value, the higher the complexity of the microbial habitat).

[0030] In each grid, a specific collection point is determined according to the collection radius, and microbial samples are obtained through the collection point, which specifically includes: loading the grid data with the collection radius, for each grid, generating a buffer zone according to its collection radius, randomly selecting a specific collection point in the buffer zone or according to a specific rule (such as a center point), recording the geographic location information (latitude and longitude or coordinates) of each collection point, finding the collection point according to the recorded geographic location information, and collecting microbial samples.

[0031] In the embodiment of the present invention, by dividing the target area into circular grids, the distribution of sampling points can be systematically planned to ensure the representativeness and uniformity of the sampling points. This gridding method helps to avoid the randomness and overlap of sampling points, and improves the sampling efficiency and accuracy. Remote sensing image data provides large-scale, high-precision surface information, which can quickly obtain the vegetation distribution and land use conditions of the target area; preprocessing can eliminate noise and interference factors in remote sensing images, improve image quality, and automatically identify vegetation types through image classification algorithms, so that vegetation information in each grid can be quickly and accurately obtained, avoiding the tediousness and time-consumingness of traditional manual surveys, and improving work efficiency and accuracy. Vegetation coverage is one of the important indicators reflecting the status of surface vegetation. By calculating the difference in spectral characteristics between vegetation and bare soil to estimate vegetation coverage, the vegetation density and ecological status of the target area can be quantitatively evaluated. By calculating the number of vegetation pixels per unit area to evaluate the density of vegetation, the description of vegetation status can be further refined. This quantitative method helps to understand the ecological and environmental characteristics within each grid more accurately. Determining the collection radius for each grid based on the distribution of vegetation types, coverage and density of vegetation can ensure that the selection of sampling points is more reasonable. This method fully considers the impact of different vegetation types and densities on the distribution of soil microorganisms, and improves the pertinence and representativeness of the sampling points.

[0032] The above step 1, extracting DNA from each microbial sample to obtain DNA data; sequencing the DNA data using high-throughput sequencing technology to obtain sequencing data, may include: Collect microbial samples from the target area, which usually involves collecting soil or other samples containing microorganisms at the selected sampling point using sterile tools (such as sterile spatulas or cotton swabs) and placing them in sterile containers to avoid external contamination. The collected microbial samples are initially processed to remove impurities and enrich microbial cells, which can include washing with saline or other buffers, and separating microbial cells by physical or chemical methods (such as centrifugation or filtration); in the processed microbial samples, specific lysis buffers or enzymes are used to destroy the cell walls and cell membranes of microorganisms to release DNA. This step is the key to DNA extraction because it ensures that DNA can be effectively released from the cells. After lysis, the sample will contain DNA and other cellular components (such as proteins, RNA, etc.). In order to obtain pure DNA, a series of purification steps are used, such as phenol-chloroform extraction, ethanol precipitation, or the use of specific DNA purification kits. These steps are designed to remove impurities while retaining DNA; after extraction, the quality and quantity of DNA are tested, which is usually done by UV-visible spectrophotometer, gel electrophoresis or other quantitative methods. This step is to ensure that the DNA sample is suitable for subsequent sequencing analysis. After DNA extraction and quality confirmation, preparation for high-throughput sequencing includes constructing sequencing libraries, which involves steps such as DNA fragmentation, connecting adapters, and performing PCR amplification to generate DNA templates suitable for sequencer reading. Finally, the prepared DNA library is loaded onto a high-throughput sequencer (such as an Illumina sequencer) for sequencing. The sequencer reads the sequence information of each DNA fragment and generates a large amount of sequencing data (usually millions of short sequence reads); after sequencing, the generated sequencing data is quality controlled, which includes removing low-quality reads, trimming sequencing adapter sequences, and checking the overall quality of the data. This step is to ensure that subsequent data analysis is based on accurate and reliable sequencing data. Through the above detailed steps, DNA can be extracted from microbial samples and its genetic information can be obtained using high-throughput sequencing technology. These data will then be used to analyze the community composition and diversity of soil microorganisms and to assess the buffering capacity of the soil.

[0033] In a preferred embodiment of the present invention, the community composition of soil microorganisms is analyzed according to the sequencing data; and a microbial community diversity index is constructed according to the community composition of soil microorganisms, including: Acquire paired sequencing data and perform preprocessing to obtain preprocessed paired sequencing data, i.e., Read 1 and Read 2, specifically including: acquiring original sequencing data from a sequencer, which is usually stored in FASTQ format; preprocessing the data, including removing low-quality reads, removing sequencing adapters, trimming low-quality sequence ends, etc.; after preprocessing, higher-quality paired sequencing data, i.e., Read 1 and Read 2, will be obtained.

[0034] Based on all the sequences in Read 1 and Read 2, a de Bruijn graph is constructed. Each node in the de Bruijn graph represents a fixed-length k-mer, that is, k consecutive nucleotides, and the edge represents the connection relationship between k-mers. Specifically, it includes: selecting a suitable k value (the length of the k-mer), which is usually determined according to the characteristics of the sequencing data and analysis requirements, traversing all the sequences in Read 1 and Read 2, and cutting each sequence into segments (k-mers) of length k; in the de Bruijn graph, each k-mer is a node. If two k-mers have k-1 identical nucleotides, an edge is added between the corresponding nodes in the graph.

[0035] The characteristics of the de Bruijn graph are used to search for k-mers shared by Read 1 and Read 2; the overlapping nucleotide sequences, i.e., the overlapping regions, are identified by traversing the de Bruijn graph and comparing the k-mer paths in different reads; nucleotide matching is performed at corresponding positions of Read 1 and Read 2 according to the overlapping regions, specifically including: searching the de Bruijn graph for k-mer nodes that exist in common in Read 1 and Read 2, and these shared k-mers are key points for subsequent sequence splicing; the de Bruijn graph is traversed to compare the k-mer paths in Read 1 and Read 2, and continuous overlapping k-mers in the paths are found, and these overlapping regions are the overlapping regions between the two reads, and the nucleotides of Read 1 and Read 2 are aligned and matched bit by bit in the overlapping regions.

[0036] After a successful match, the two reads are merged in the overlapping region to form a single sequence. Specifically, once the nucleotides are successfully matched in the overlapping region, Read 1 and Read 2 can be merged in this area. The merged sequence is a longer single sequence spliced ​​from the two original reads.

[0037] The single sequences are clustered to form operational taxonomic units, and representative sequences of each OTU are selected. The corresponding representative sequences are compared with known microbial reference databases to determine the species classification information of each OTU, including: clustering the spliced ​​single sequences using clustering algorithms (such as UCLUST, CD-HIT, etc.). The purpose of clustering is to group similar sequences together to form an operational taxonomic unit (OTU). Each OTU represents a microbial species or subspecies. A representative sequence is selected from each OTU, and these representative sequences are compared with known microbial reference databases using alignment tools (such as BLAST, BOWTIE, etc.). The alignment results will give the most likely species classification information for each OTU.

[0038] According to the species classification information, the number of sequences of different species in each sample is calculated, and the number of sequences of different species is converted into relative abundance to generate a species abundance table, which specifically includes: according to the comparison results, the number of sequences of different species (or OTU) in each sample is counted, and these numbers are converted into relative abundance, that is, the proportion of each species in the sample, and these data are sorted to generate a species abundance table, listing the relative abundance of each species in each sample.

[0039] The species abundance table is used to calculate species richness indicators and diversity indices. The diversity indices include the Shannon diversity index and the Simpson diversity index. Specifically, the data in the species abundance table are used to calculate species richness indicators, such as the number of OTUs and the Chao1 index, to evaluate the richness of species in the sample. At the same time, diversity indices, such as the Shannon Index and the Simpson Index, are calculated to quantify the diversity level of species in the sample.

[0040] In the embodiment of the present invention, low-quality reads, sequencing errors and impurity data can be removed through preprocessing, thereby improving the accuracy and reliability of subsequent analysis; the de Bruijn graph can effectively represent the relationship between k-mers in sequencing data, which helps to quickly identify and align common patterns in sequences, and provides a basis for subsequent overlapping region identification and sequence splicing; by searching Read 1 and Read The k-mer shared in 2 can accurately find the overlapping area between two reads; accurate identification of overlapping areas and nucleotide matching are the prerequisites for sequence splicing, which helps to recover longer, original DNA sequences from paired sequencing data; by merging overlapping reads, a more complete DNA sequence can be obtained; through cluster analysis, similar sequences can be grouped together to form OTUs, which helps to simplify the data and reduce the complexity of analysis while retaining species diversity information; by comparing with known microbial reference databases, the species classification information of each OTU can be accurately determined. The species abundance table provides information on the relative number of different species in each sample, which is the basis for calculating species richness and diversity indices. Species richness indicators and diversity indices (such as the Shannon diversity index and the Simpson diversity index) can quantitatively describe the complexity and diversity of microbial communities.

[0041] In a preferred embodiment of the present invention, calculating the fitness value of each individual includes: The Shannon diversity index and Simpson diversity index are determined according to the richness of microbial species and microbial communities in the soil; the contribution value of soil microbial diversity is determined according to the Shannon diversity index and Simpson diversity index; Determine the Pearson correlation coefficient between each environmental factor and soil buffering capacity; determine the correlation contribution value based on each environmental factor and its value in the current individual; For each microbial activity data sample, determine its corresponding expected value and the standard deviation of the microbial activity data; according to each microbial activity data sample and the corresponding expected value and the standard deviation of the microbial activity data, obtain the standardized deviation of the sample; According to the standardized deviation of the sample and the soil respiration rate of the sample, the contribution value of microbial activity data is obtained; The correlation contribution value, the standardized deviation of the sample, and the contribution value of the microbial activity data are fused to obtain the fitness value of each individual.

[0042] In an embodiment of the present invention, the Shannon diversity index S and the Simpson diversity index m are calculated according to the richness of microbial species in the soil and the structure of the microbial community; the Shannon diversity index S reflects the richness of microbial species in the soil, while the Simpson diversity index m reflects the proportional relationship between dominant species and rare species in the microbial community. Combining these two indexes, the contribution of soil microbial diversity to the fitness value, i.e., the soil microbial diversity contribution value, can be determined; the Pearson correlation coefficient between each environmental factor and the soil buffering capacity is calculated to measure the degree of linear correlation between them, and the correlation contribution value of the environmental factor to the fitness value is determined according to the value of each environmental factor in the current individual and its correlation coefficient with the soil buffering capacity. For each microbial activity data sample, the corresponding expected value (i.e., the average value) and the standard deviation of the microbial activity data are calculated, and the standardized deviation of the sample is calculated according to the soil respiration rate, expected value and standard deviation of each sample to reflect the deviation of the sample from the average value, and the contribution of the microbial activity data to the fitness value, i.e., the microbial activity data contribution value, is determined by combining the standardized deviation of the sample and the soil respiration rate. The contribution value of soil microbial diversity, the contribution value of the correlation between environmental factors and soil buffering capacity, and the contribution value of microbial activity data are weighted and fused to obtain the fitness value of each individual.

[0043] By comprehensively considering the soil microbial diversity, the correlation between environmental factors and soil buffering capacity, and microbial activity data, the soil buffering capacity can be evaluated more comprehensively and accurately, and the fitness value can serve as an important basis for soil management and improvement. By comparing the fitness values ​​of different individuals, a combination of environmental factors with higher soil buffering capacity can be selected, thereby optimizing soil management strategies and improving soil quality. In ecological restoration projects, the fitness value can be used to screen out the combination of environmental factors that best suits the current soil conditions, thereby accelerating the recovery and reconstruction of the soil ecosystem and improving the efficiency of ecological restoration.

[0044] In a preferred embodiment of the present invention, when applied in a specific application, the calculation formula of the fitness function value can be: ; in, represents the expected value of soil respiration rate; represents the standard deviation of microbial activity data (soil respiration rate); Indicates Pearson correlation coefficient between environmental factors and soil buffering capacity; Indicates the current individual Environmental factor values; represents the absolute value of the maximum correlation coefficient; represents the Shannon diversity index; represents the Simpson diversity index; Indicates Soil respiration rate of each sample; , and represents the weight coefficient; Indicates the index of the environmental factor, from 1 to ; Represents the total number of environmental factors; Represents the index of the microbial activity data sample, from 1 to ; Represents the total number of microbial activity data samples.

[0045] In an embodiment of the present invention, the calculation of the fitness function value takes into account multiple ecological and environmental factors, and can more comprehensively and accurately evaluate the health of the soil ecosystem. By introducing the Shannon diversity index and the Simpson diversity index, the function can reflect soil biodiversity, which is an important indicator of soil ecosystem stability and functional diversity. At the same time, by considering the correlation between environmental factors and soil buffering capacity, the function can evaluate the adaptability of soil to environmental changes. In addition, by introducing the expected value and fluctuation range of soil respiration rate, the activity state of soil microorganisms can be further understood, thereby more comprehensively evaluating soil quality.

[0046] The Shannon diversity index (S) and Simpson diversity index (m) were calculated based on soil sample data. Both indices are calculated based on species richness and evenness and are used to measure soil biodiversity.

[0047] Calculate the correlation between environmental factors and soil buffering capacity: For each environmental factor (such as temperature, humidity, pH, etc.), calculate the Pearson correlation coefficient between it and the soil buffering capacity and find the maximum absolute value of the correlation coefficient .

[0048] Calculate the expected value and range of soil respiration rate: Calculate the expected value of soil respiration rate based on the microbial activity data sample and fluctuation range , which can be done by performing statistical analysis on the sample data.

[0049] The components of the calculation to calculate the fitness function value: Part I: , reflecting the contribution of soil biodiversity to fitness.

[0050] Part II: , reflecting the contribution of the correlation between environmental factors and soil buffering capacity to fitness.

[0051] Part III: , reflecting the contribution of soil microbial activity to fitness.

[0052] Adding the above three parts together, we can get the total fitness value of each individual, which can be used to evaluate the comprehensive status of the soil ecosystem.

[0053] For example, in order to evaluate the health status and microbial activity of the soil, the above-mentioned fitness function value calculation is used to collect relevant environmental factor data, soil diversity index and microbial activity data, and set weight coefficients.

[0054] Environmental factor data, measuring soil temperature, humidity, pH value, organic matter content and other n environmental factors, and obtaining the value of each environmental factor .

[0055] Soil diversity index, by sampling and analyzing the microbial community in the soil, the Shannon diversity index S and the Simpson diversity index m were calculated.

[0056] Microbial activity data, respiration rate data of N soil samples collected , and calculated the expected value of soil respiration rate and the standard deviation of soil respiration rate .

[0057] Correlation coefficient calculation, calculate the Pearson correlation coefficient between each environmental factor and soil buffering capacity , and found the maximum absolute value of the correlation coefficient .

[0058] Using the data collected above and the set weight coefficients, various data were calculated, and the fitness value F of the soil ecosystem was obtained through calculation. According to the calculated fitness value F, the comprehensive condition of the farmland soil ecosystem can be evaluated. A higher fitness value indicates that the soil ecosystem has higher biodiversity, good adaptability to environmental factors and active microbial communities.

[0059] In a preferred embodiment of the present invention, the fitness value of each individual is calculated, and selection, crossover and mutation operations are performed, and the genetic operation is repeated until the termination condition is met to obtain the final key environmental factor combination, including: An initial population is randomly generated, and each individual in the population represents a combination of a set of key environmental factors, specifically including: determining the number and range of key environmental factors, using a random number generator to generate random values ​​for each environmental factor within this range to form an individual, and repeating the above process until a sufficient number of individuals are generated to form an initial population.

[0060] Calculate the fitness value of each individual, and select the corresponding individual for crossover operation based on the individual fitness value to generate new offspring individuals. Specifically, for each individual in the population, extract the key environmental factor value it represents, calculate based on these values, and record the fitness value of each individual, which will be used for subsequent selection operations; use roulette selection, tournament selection and other methods to select which individuals will participate in the crossover operation based on the individual fitness value. Generally, individuals with high fitness are more likely to be selected. Once a pair of parent individuals are selected, the crossover operation can be performed.

[0061] The offspring individuals are mutated, and the new individuals generated by selection, crossover and mutation operations form a new generation of population until the preset evolutionary generation is reached to obtain the final individual in the current population as the final solution to the problem; according to the final solution, the final key environmental factor combination is obtained, specifically including: randomly selecting one or more points in the gene string of the parent individual as crossover points, exchanging the genes of the two parent individuals at the crossover points, generating two new offspring individuals, and for each offspring individual, deciding whether to mutate with a certain mutation probability. If mutation is decided, one or more gene positions are randomly selected for mutation. The mutation can be a change in the value of the gene, or in some cases, it can be the insertion, deletion or Inversion and other operations; the newly generated offspring individuals together with the parent individuals form a temporary mixed population; according to the fitness value, a new generation of population is selected from this mixed population, which is usually achieved by selecting individuals with higher fitness. The new generation of population will be used for the next round of genetic operations. For multiple rounds of genetic operations, the evolutionary algebra counter is increased after each round. When the evolutionary algebra reaches the preset value, the genetic operation is stopped; after reaching the preset evolutionary algebra, the individual with the highest fitness is selected from the last generation of the population. The combination of key environmental factors represented by this individual is the optimal or approximately optimal solution found by the algorithm. The genetic information of this individual, that is, the value of the key environmental factor, is extracted as the final solution to the problem.

[0062] In an embodiment of the present invention, by randomly generating an initial population, the algorithm can start exploring in a wide search space, which helps to avoid falling into a local optimal solution, thereby making it more likely to find a global optimal solution; the calculation of the fitness function value quantifies the individual's adaptability to the problem, and by evaluating the fitness value of each individual, the algorithm can distinguish which individuals are more likely to be close to the optimal solution. The selection operation ensures that better individuals have a greater chance to participate in reproduction, thereby passing on their excellent genes to the next generation, which helps the algorithm quickly converge to an excellent solution space; the crossover operation simulates the biological hybridization process, combining the excellent genes of the parent individuals to produce potentially better offspring individuals, which helps the algorithm to effectively explore the search space to find better solutions; the mutation operation simulates gene mutation, providing the algorithm with the ability to jump out of the local optimal solution. By randomly changing certain gene values, the algorithm may find new and better search areas; the new generation of population inherits the excellent genes of the previous generation and introduces new mutations, which increases the diversity of the population and helps the algorithm search for the optimal solution globally; through multiple iterations, the algorithm can gradually approach the optimal solution. Each generation will eliminate unfit individuals, retain and reproduce excellent individuals, thereby continuously improving the overall fitness of the population; after multiple generations of genetic operations, the algorithm finally converges to one or several excellent individuals, which represent the optimal or approximately optimal solutions to the problem, and these solutions represent the optimal combination of key environmental factors.

[0063] For example, due to long-term use of chemical fertilizers and improper farmland management in a certain area, the soil buffering capacity has gradually decreased, resulting in the impact on crop yield and quality. In order to improve this situation, genetic algorithms are used to find the best soil improvement plan.

[0064] Identify key environmental factors: soil organic matter content (range: 1%-5%); Soil pH (range: 5.5-8.5); soil microbial population diversity index (range: 1-10); soil cation exchange capacity (cmol / kg, range: 5–30); Use a random number generator to generate random values ​​for each environmental factor within a specified range. Combine these random values ​​to form an "individual" that represents a soil improvement plan. Repeat this process to generate 100 individuals to form the initial population.

[0065] The calculation of the fitness function value takes into account multiple dimensions of soil buffering capacity, such as acid-base balance, nutrient retention, microbial activity, etc. The roulette wheel selection method is used to select the parent generation participating in the crossover according to the fitness value of the individual. Individuals with high fitness have a higher probability of being selected. Two parent individuals are randomly selected, and a crossover point is randomly selected in the gene string of each parent. The gene segments of the two parents after the crossover point are exchanged to generate two new offspring individuals. For each offspring individual, it is decided with a probability of 0.05 whether to mutate. If it is decided to mutate, a gene position (i.e., an environmental factor) is randomly selected and the gene is modified. Make small random adjustments, such as increasing or decreasing its value by 1%; the newly generated offspring individuals together with the parent individuals form a mixed population, and according to the fitness value, select the top 100 individuals with the highest fitness from the mixed population to form a new generation of population, and repeat the selection, crossover and mutation operations on the new generation of population. With each round of operation, the evolutionary generation increases by 1. When the evolutionary generation reaches 50 generations, stop the genetic operation. After the termination condition is reached, select the individual with the highest fitness from the last generation of population as the optimal solution. The combination of environmental factors represented by this individual is the recommended soil improvement plan.

[0066] The above step 4, obtaining microbial activity data based on soil respiration rate in microbial samples, may include: Select representative sampling points that reflect the overall situation of the study area. Use appropriate tools (such as soil samplers) to collect soil samples, ensure the integrity and non-contamination of the samples, place the collected soil samples in appropriate containers, and mark the sampling time, location and other information; bring the soil samples back to the laboratory and process them as soon as possible to avoid sample deterioration, remove impurities such as stones and roots from the samples, ensure the purity of the samples, and grind or sieve the soil if necessary to obtain uniform soil particles; prepare measurement equipment, such as soil respiration measurement system or infrared gas analyzer, place the soil sample in the measurement equipment, ensure good sealing, and open the container. The amount of carbon dioxide released by the soil is measured and recorded, which is usually achieved by measuring the change in carbon dioxide concentration per unit time. The measurement process lasts for a period of time (such as several hours or a day) to obtain accurate respiration rate data; the measured soil respiration rate data is sorted out, including information such as time and carbon dioxide release, and the data is analyzed, such as calculating statistical indicators such as mean and standard deviation; soil respiration rate is one of the important indicators reflecting soil microbial activity. A higher respiration rate usually means more active microbial activity. By comparing soil respiration rate data of different samples or at different time points, the changes and differences in microbial activity can be evaluated.

[0067] In a preferred embodiment of the present invention, the soil buffering capacity value is predicted based on the final key environmental factor combination, microbial activity data and the trained neural network model, including: Select features related to soil buffering capacity from the combination of key environmental factors and microbial activity data, and combine the selected features into a feature vector, including: collecting data on the combination of key environmental factors, including soil organic matter content, soil pH, soil microbial population diversity index, and soil cation exchange capacity. At the same time, microbial activity data should also be collected, including indicators such as the type, quantity, and activity of microorganisms; calculate the Pearson correlation coefficient between each feature and soil buffering capacity, and record the correlation coefficient value between each feature and the target variable; determine a threshold value (such as Pearson correlation coefficient > 0.5) based on the calculated correlation coefficient, and screen out features that are significantly related to soil buffering capacity; standardize the screened features that are most relevant to soil buffering capacity; Confirm that all selected features (such as soil organic matter content, pH, microbial diversity index, and cation exchange capacity) have been standardized; collect all standardized feature data together and store them in a data frame (DataFrame) or array, with each feature corresponding to one column or row; determine the order of features based on needs or the expected input format of the model; the order can be determined by the order of columns in the data set; extract the corresponding values ​​from each feature in the determined order, and for each data point (for example, each soil sample), combine these values ​​into a one-dimensional array, which is a feature vector, which contains the standardized values ​​of all selected features for that data point; The feature vector is input into the trained neural network model to obtain the soil buffering capacity value, specifically including: dividing the data set containing the feature vector and the corresponding soil buffering capacity value into a training set, a validation set and a test set, usually using a ratio of 70% for the training set, 15% for the validation set and 15% for the test set. Design the neural network architecture, including the input layer, hidden layer, and output layer, set initial values ​​for the weights and biases of the neural network, which are usually random; determine a suitable loss function to measure the difference between the model prediction and the actual value, such as mean squared error (MSE); select an optimization algorithm, such as gradient descent (SGD), to update the weights and biases of the model during training; determine training parameters such as training rounds (epochs), batch size (batchsize), and learning rate (learning rate); load the training set data into memory and perform necessary preprocessing, such as feature scaling; in each training round, input the training data into the neural network in batches, calculate the loss function value, and update the model parameters through back propagation and optimizer. After each round, use the validation set to evaluate the performance of the model, and adjust the model architecture or training parameters as needed to optimize the performance; use the test set to evaluate the performance of the trained neural network model to ensure that the model has good generalization ability, and save the trained neural network model as a file; input the combined feature vector into the loaded neural network model, and the neural network model is used to predict the soil buffering capacity value.

[0068] In an embodiment of the present invention, by combining key environmental factors and microbial activity data, various factors affecting soil buffering capacity can be considered more comprehensively. This multi-dimensional data integration enables the neural network model to learn more complex nonlinear relationships, thereby improving the accuracy of soil buffering capacity prediction. Accurate soil buffering capacity prediction provides an important basis for agricultural, environmental protection and land management decisions. For example, in agricultural production, understanding the buffering capacity of the soil helps to formulate reasonable fertilization and irrigation plans and improve crop yield and quality. Through accurate prediction of soil buffering capacity, resources such as water resources, fertilizers and pesticides can be allocated and utilized more effectively, which not only helps to save costs, but also reduces environmental pollution; by predicting and monitoring soil buffering capacity, soil degradation or pollution problems can be discovered in a timely manner, so that corresponding measures can be taken for protection and restoration.

[0069] like Figure 2 As shown, an embodiment of the present invention further provides an evaluation system based on soil buffering capacity, comprising: An extraction module is used to extract DNA from each microbial sample to obtain DNA data; and sequence the DNA data using high-throughput sequencing technology to obtain sequencing data; Constructing a module for analyzing the community composition of soil microorganisms based on sequencing data; constructing a microbial community diversity index based on the community composition of soil microorganisms; The optimization module is used to encode the diversity index of the microbial community as the genotype of the genetic algorithm; randomly generate the initial population, each individual represents a combination of environmental factors; calculate the fitness value of each individual, and perform selection, crossover and mutation operations, and repeat the genetic operations until the termination condition is met to obtain the final key environmental factor combination; An acquisition module is used to obtain microbial activity data based on soil respiration rate in microbial samples; A prediction module is used to predict the soil buffering capacity value based on the final combination of key environmental factors, microbial activity data, and the trained neural network model; The classification module is used to classify the soil into different buffering capacity levels according to the soil buffering capacity value.

[0070] It should be noted that the system is a system corresponding to the above method, and all implementation methods in the above method embodiment are applicable to this embodiment and can achieve the same technical effect.

[0071] The embodiment of the present invention further provides a computing device, comprising: a processor, a memory storing a computer program, wherein when the computer program is executed by the processor, the method described above is executed. All implementations in the above method embodiment are applicable to this embodiment and can achieve the same technical effect.

[0072] The embodiment of the present invention also provides a computer-readable storage medium storing instructions, which, when executed on a computer, enable the computer to execute the method described above. All implementations in the above method embodiment are applicable to this embodiment and can achieve the same technical effect.

Claims

1. A method for evaluating soil buffering capacity, characterized in that: The method comprises: Step 1, extracting DNA from each microbial sample to obtain DNA data; sequencing the DNA data using high-throughput sequencing technology to obtain sequencing data; Step 2, analyzing the community composition of soil microorganisms based on the sequencing data; and constructing a microbial community diversity index based on the community composition of soil microorganisms; Step 3, encode the diversity index of the microbial community as the genotype of the genetic algorithm; randomly generate the initial population, each individual represents a combination of environmental factors; calculate the fitness value of each individual, and perform selection, crossover and mutation operations, and repeat the genetic operation until the termination condition is met to obtain the final key environmental factor combination; Step 4, obtaining microbial activity data based on soil respiration rate in microbial samples; Step 5, predicting the soil buffering capacity value based on the final combination of key environmental factors, microbial activity data and the trained neural network model; Step 6: Classify the soil into different buffering capacity levels according to the soil buffering capacity value.

2. The soil buffer capacity assessment method according to claim 1, characterized in that: Before DNA extraction is performed on each microbial sample to obtain DNA data, it also includes: Divide the target acquisition area into circular grids; obtain remote sensing image data of the target area from remote sensing satellites; Preprocess the remote sensing image data, use the spectral characteristics, texture characteristics and shape characteristics of the preprocessed remote sensing image, automatically identify the vegetation type through the image classification algorithm, and mark and classify the vegetation type in each grid to obtain the vegetation type distribution; According to the distribution of vegetation types, the vegetation coverage is estimated by calculating the difference in spectral characteristics between vegetation and bare soil; According to remote sensing image data, the density of vegetation in the grid is obtained by calculating the number of vegetation pixels per unit area; In each circular grid, a corresponding collection radius is determined for each grid according to the distribution of vegetation types, coverage, and density of vegetation; In each grid, a specific collection point is determined according to the collection radius, and microbial samples are obtained through the collection point.

3. The soil buffer capacity assessment method according to claim 2, characterized in that: Environmental factors include soil structure, density, porosity, soil pH, type, quantity, activity of soil microorganisms, topography and climatic conditions.

4. The soil buffer capacity assessment method according to claim 3, characterized in that: Analyze the community composition of soil microorganisms based on sequencing data; construct a microbial community diversity index based on the community composition of soil microorganisms, including: Acquire paired sequencing data, and perform preprocessing to obtain paired sequencing data after preprocessing, namely, Read 1 and Read 2; Based on all the sequences in Read 1 and Read 2, a de Bruijn graph is constructed. Each node in the de Bruijn graph represents a k-mer of fixed length, i.e., k consecutive nucleotides, and the edge represents the connection relationship between k-mers. Using the characteristics of the de Bruijn graph, search for k-mers shared in Read 1 and Read 2; traverse the de Bruijn graph and compare the k-mer paths in different reads to identify overlapping nucleotide sequences, i.e., overlapping regions; match nucleotides at corresponding positions in Read 1 and Read 2 based on the overlapping regions; After a successful match, the two reads were merged in the overlapping region to be assembled into a single sequence; Single sequences were clustered to form operational taxonomic units, representative sequences of each OTU were selected, and the corresponding representative sequences were compared with known microbial reference databases to determine the species classification information of each OTU; According to the species classification information, the number of sequences of different species in each sample is calculated, and the number of sequences of different species is converted into relative abundance to generate a species abundance table; The species abundance table was used to calculate species richness index and diversity index, including Shannon diversity index and Simpson diversity index.

5. The soil buffer capacity assessment method according to claim 4, characterized in that: Calculate the fitness value of each individual, including: The Shannon diversity index and Simpson diversity index are determined according to the richness of microbial species and microbial communities in the soil; the contribution value of soil microbial diversity is determined according to the Shannon diversity index and Simpson diversity index; Determine the Pearson correlation coefficient between each environmental factor and soil buffering capacity; determine the correlation contribution value based on each environmental factor and its value in the current individual; For each microbial activity data sample, determine its corresponding expected value and the standard deviation of the microbial activity data; according to each microbial activity data sample and the corresponding expected value and the standard deviation of the microbial activity data, obtain the standardized deviation of the sample; According to the standardized deviation of the sample and the soil respiration rate of the sample, the contribution value of microbial activity data is obtained; The correlation contribution value, the standardized deviation of the sample, and the contribution value of the microbial activity data are fused to obtain the fitness value of each individual.

6. The soil buffer capacity assessment method according to claim 5, characterized in that: Calculate the fitness value of each individual, and perform selection, crossover and mutation operations, and repeat the genetic operation until the termination condition is met to obtain the final key environmental factor combination, including: An initial population is randomly generated, and each individual in the population represents a combination of a set of key environmental factors; Calculate the fitness value of each individual, and select the corresponding individual for crossover operation according to the individual's fitness value to generate new offspring individuals; The offspring individuals are mutated, and the new individuals generated by selection, crossover and mutation operations form a new generation of population until the preset number of evolutionary generations is reached to obtain the final individual in the current population as the final solution to the problem; based on the final solution, the final combination of key environmental factors is obtained.

7. The soil buffer capacity assessment method according to claim 6, characterized in that: Based on the final combination of key environmental factors, microbial activity data and trained neural network model, the soil buffering capacity value is predicted, including: Select features related to soil buffering capacity from the key environmental factor combination and microbial activity data, and combine the selected features into a feature vector; The feature vector is input into the trained neural network model to obtain the soil buffering capacity value.

8. An evaluation system based on soil buffering capacity, characterized in that: The system is used to perform the method according to any one of claims 1 to 7, comprising: An extraction module is used to extract DNA from each microbial sample to obtain DNA data; and sequence the DNA data using high-throughput sequencing technology to obtain sequencing data; Constructing a module for analyzing the community composition of soil microorganisms based on sequencing data; constructing a microbial community diversity index based on the community composition of soil microorganisms; The optimization module is used to encode the diversity index of the microbial community as the genotype of the genetic algorithm; randomly generate the initial population, each individual represents a combination of environmental factors; calculate the fitness value of each individual, and perform selection, crossover and mutation operations, and repeat the genetic operations until the termination condition is met to obtain the final key environmental factor combination; An acquisition module is used to obtain microbial activity data based on soil respiration rate in microbial samples; A prediction module is used to predict the soil buffering capacity value based on the final combination of key environmental factors, microbial activity data, and the trained neural network model; The classification module is used to classify the soil into different buffering capacity levels according to the soil buffering capacity value.

9. A computing device, characterized in that include: one or more processors; A storage device for storing one or more programs, when the one or more programs are executed by the one or more processors, the one or more processors implement the method as claimed in any one of claims 1 to 7.

10. A computer-readable storage medium, characterized in that: The computer-readable storage medium stores a program, which, when executed by a processor, implements the method according to any one of claims 1 to 7.

Citation Information

Patent Citations

  • Comprehensive evaluation method for remediation effects of heavy metal contaminated soil based on plants, soil and microorganisms

    CN107066823A

  • Grassland soil degradation evaluation method

    CN107220967A

  • Land space ecological restoration project multi-scale soil environment and biological diversity cooperative monitoring and comprehensive evaluation method

    CN118313677A

  • Microbial functional factor evaluation system and method based on data analysis

    CN119229979A

Cited By

  • Green plant selection method, system and equipment and storage medium

    CN120823912A

  • Formula construction method of acid soil conditioner

    CN121256564A

  • Remote sensing image data correction method

    CN122222879A