An evaluation method and system based on soil buffering capacity
Through high-throughput sequencing and neural network model combined with genetic algorithms, the composition of soil microbial communities is analyzed, and the problem of neglected microbial factors in traditional soil buffering capacity assessment methods is solved, achieving more accurate and efficient soil buffering capacity assessment and management.
Patent Information
- Application Number
- CN202510458408.1
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-04-14
- Publication Date
- 2025-07-08
- Estimated Expiration
- 2045-04-14
AI Technical Summary
Traditional soil buffering capacity assessment methods ignore the important role of soil microorganisms, resulting in inaccurate assessment.
Through high-throughput sequencing technology, the composition of soil microbial communities is analyzed, and the soil buffering capacity is predicted, microbial activity and multiple environmental factors are considered, the microbial community diversity index is constructed, and the sampling point distribution is optimized using remote sensing image data.
It achieves a more accurate assessment of soil buffering capacity, improves the accuracy and efficiency of the assessment, can dynamically monitor changes in soil buffering capacity, and formulate targeted soil management strategies based on the evaluation results.
Smart Images

Figure CN119990918B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of soil assessment, and particularly to an assessment method and system based on soil buffering capacity. Background Art
[0002] Traditional methods for assessing soil buffering capacity rely on laboratory chemical analysis, which typically involves the determination of multiple chemical indicators of the soil, such as pH value, organic matter content, and cation exchange capacity. These indicators can indirectly reflect certain physical and chemical properties of the soil and are thus used to infer the buffering capacity of the soil.
[0003] However, although the above chemical analysis methods can provide useful information about certain physical and chemical properties of the soil, they ignore the key factor of soil microorganisms. Soil microorganisms play an important role in the soil ecosystem, including nutrient cycling, organic matter decomposition, and soil structure formation. These microbial activities directly affect the buffering capacity of the soil because they are involved in many key biochemical processes. Therefore, traditional chemical analysis methods may not be able to capture these dynamic changes, thereby affecting the accuracy of the assessment. Summary of the Invention
[0004] The technical problem to be solved by the present invention is to provide an assessment method and system based on soil buffering capacity, which can more accurately assess the buffering capacity of the soil.
[0005] To solve the above technical problem, the technical solution of the present invention is as follows:
[0006] In the first aspect, an assessment method based on soil buffering capacity, the method includes:
[0007] Step 1, extract DNA from each microbial sample to obtain DNA data; sequence the DNA data using high-throughput sequencing technology to obtain sequencing data;
[0008] Step 2, analyze the community composition of soil microorganisms according to the sequencing data; construct a microbial community diversity index according to the community composition of soil microorganisms;
[0009] Step 3, encode the microbial community diversity index as the genotype of the genetic algorithm; randomly generate an initial population, where each individual represents a combination of environmental factors; calculate the fitness value of each individual, and perform selection, crossover, and mutation operations, and repeat the genetic operations until the termination condition is met to obtain the final key environmental factor combination;
[0010] Step 4, obtain microbial activity data based on the soil respiration rate in the microbial sample;
[0011] Step 5, predict the soil buffering capacity value based on the final combination of key environmental factors, microbial activity data, and the trained neural network model;
[0012] Step 6, classify the soil into different buffering capacity levels according to the soil buffering capacity value.
[0013] Further, before DNA extraction is performed on each microbial sample to obtain DNA data, it also includes:
[0014] Divide the target collection area into circular grids; obtain remote sensing image data of the target area from a remote sensing satellite;
[0015] Preprocess the remote sensing image data, and use the spectral features, texture features, and shape features of the preprocessed remote sensing image to automatically identify the vegetation type through an image classification algorithm, label and classify the vegetation type within each grid to obtain the vegetation type distribution;
[0016] According to the vegetation type distribution, estimate the vegetation coverage by calculating the difference in spectral features between vegetation and bare soil;
[0017] According to the remote sensing image data, calculate the number of vegetation pixels per unit area to obtain the density of vegetation within the grid;
[0018] Within each circular grid, determine a corresponding collection radius for each grid according to the vegetation type distribution, coverage, and density of vegetation;
[0019] Within each grid, determine a specific collection point according to the collection radius, and obtain microbial samples through the collection point.
[0020] Further, the environmental factors include soil structure, density, porosity, soil pH value, types, quantities, activities of soil microorganisms, topography, and climate conditions.
[0021] Further, analyze the community composition of soil microorganisms based on sequencing data; construct a microbial community diversity index according to the community composition of soil microorganisms, including:
[0022] Obtain paired sequencing data and perform preprocessing to obtain preprocessed paired sequencing data, namely Read 1 and Read 2;
[0023] Based on all the sequences in Read 1 and Read 2, construct a de Bruijn graph, where each node in the de Bruijn graph represents a fixed-length k-mer, that is, a continuous k nucleotides, and the edges represent the connection relationship between k-mers;
[0024] Utilize the characteristics of the de Bruijn graph to search for shared k-mers in Read 1 and Read 2; identify overlapping nucleotide sequences, i.e., the overlapping regions, by traversing the de Bruijn graph and comparing k-mer paths in different reads; perform nucleotide matching at the corresponding positions in Read 1 and Read 2 according to the overlapping regions;
[0025] After successful matching, merge the two reads in the overlapping region to splice into a single sequence;
[0026] Perform clustering analysis on the single sequence to form operational taxonomic units (OTUs), select the representative sequences of each OTU, and align the corresponding representative sequences with a known microbial reference database to determine the species classification information of each OTU;
[0027] According to the species classification information, calculate the number of sequences of different species in each sample, and convert the number of sequences of different species into relative abundances to generate a species abundance table;
[0028] Utilize the species abundance table to calculate species richness indices and diversity indices, where the diversity indices include Shannon diversity index and Simpson diversity index.
[0029] Furthermore, calculate the fitness value of each individual, including:
[0030] Determine the Shannon diversity index and Simpson diversity index according to the richness of microbial species in the soil and the microbial community; determine the contribution value of soil microbial diversity according to the Shannon diversity index and Simpson diversity index;
[0031] Determine the Pearson correlation coefficient between each environmental factor and the soil buffering capacity; determine the correlation contribution value according to each environmental factor and its value in the current individual;
[0032] For each microbial activity data sample, determine its corresponding expected value and the standard deviation of the microbial activity data; obtain the standardized deviation of the sample according to each microbial activity data sample and the corresponding expected value and the standard deviation of the microbial activity data;
[0033] According to the standardized deviation of the sample and the soil respiration rate of the sample, obtain the contribution value of the microbial activity data;
[0034] Fuse the correlation contribution value, the standardized deviation of the sample, and the contribution value of the microbial activity data to obtain the fitness value of each individual.
[0035] Furthermore, calculate the fitness value of each individual, and perform selection, crossover, and mutation operations. Repeat the genetic operations until the termination condition is met to obtain the final combination of key environmental factors, including:
[0036] Randomly generate an initial population, where each individual in the population represents a combination of a set of key environmental factors;
[0037] Calculate the fitness value of each individual. According to the fitness value of the individual, select the corresponding individual for crossover operation to generate new offspring individuals;
[0038] Perform mutation operation on the offspring individuals. The new individuals generated through selection, crossover, and mutation operations form a new generation of population until the preset number of evolutionary generations is reached, and the final individual in the current population is used as the final solution to the problem; According to the final solution, the final combination of key environmental factors is obtained.
[0039] Furthermore, according to the final combination of key environmental factors, microbial activity data, and the trained neural network model, predict the soil buffering capacity value, including:
[0040] Select the features related to the soil buffering capacity from the combination of key environmental factors and microbial activity data, and combine the selected features into a feature vector;
[0041] Input the feature vector into the trained neural network model to obtain the soil buffering capacity value.
[0042] In a second aspect, an evaluation system based on soil buffering capacity includes:
[0043] An extraction module for extracting DNA from each microbial sample to obtain DNA data; sequencing the DNA data using high-throughput sequencing technology to obtain sequencing data;
[0044] A construction module for analyzing the community composition of soil microorganisms based on the sequencing data; constructing a microbial community diversity index according to the community composition of soil microorganisms;
[0045] An optimization module for encoding the microbial community diversity index as the genotype of a genetic algorithm; randomly generating an initial population, where each individual represents a combination of environmental factors; calculating the fitness value of each individual, and performing selection, crossover, and mutation operations, repeating the genetic operations until the termination condition is met to obtain the final combination of key environmental factors;
[0046] An acquisition module for obtaining microbial activity data according to the soil respiration rate in the microbial sample;
[0047] A prediction module for predicting the soil buffering capacity value according to the final combination of key environmental factors, microbial activity data, and the trained neural network model;
[0048] A partitioning module, configured to partition soil into different buffer capacity levels according to the soil buffer capacity value.
[0049] In a third aspect, a computing device includes:
[0050] One or more processors;
[0051] A storage device, configured to store one or more programs, which when executed by the one or more processors cause the one or more processors to implement the method described above.
[0052] In a fourth aspect, a computer-readable storage medium stores a program that, when executed by a processor, implements the method described above.
[0053] The above solution of the present invention has at least the following beneficial effects:
[0054] By analyzing the community composition and diversity of soil microorganisms, the buffer capacity of soil can be evaluated more precisely. Microorganisms are key components of the soil ecosystem, and their community structure and diversity have an important impact on the soil buffer capacity. Therefore, incorporating microbial data into the evaluation system can significantly improve the accuracy and comprehensiveness of the evaluation.
[0055] By determining the combination of key environmental factors through a genetic algorithm and combining microbial activity data, this method can dynamically monitor and predict changes in soil buffer capacity.
[0056] Using high-throughput sequencing technology and a neural network model, this method realizes the intelligentization and automation of soil buffer capacity evaluation, which can not only improve the evaluation efficiency but also reduce errors caused by human factors.
[0057] Partitioning soil into different buffer capacity levels according to the soil buffer capacity value helps to formulate targeted soil management measures, and different management strategies can be adopted for soils of different levels. BRIEF DESCRIPTION OF THE DRAWINGS
[0058] Figure 1 is a schematic flowchart of an evaluation method based on soil buffer capacity provided by an embodiment of the present invention.
[0059] Figure 2 is a schematic diagram of an evaluation system based on soil buffer capacity provided by an embodiment of the present invention. DETAILED DESCRIPTION OF THE EMBODIMENTS
[0060] Exemplary embodiments of the present disclosure will now be described in more detail with reference to the accompanying drawings. Although the exemplary embodiments of the present disclosure are shown in the drawings, it should be understood that the present disclosure can be implemented in various forms and should not be limited by the embodiments set forth herein. On the contrary, these embodiments are provided so that the present disclosure can be more thoroughly understood and the scope of the present disclosure can be fully conveyed to those skilled in the art.
[0061] As Figure 1 shown, an embodiment of the present invention provides an evaluation method based on soil buffering capacity, and the method includes the following steps:
[0062] Step 1, extract DNA from each microbial sample to obtain DNA data; sequence the DNA data using high-throughput sequencing technology to obtain sequencing data;
[0063] Step 2, analyze the community composition of soil microorganisms based on the sequencing data; construct a microbial community diversity index according to the community composition of soil microorganisms;
[0064] Step 3, encode the microbial community diversity index as the genotype of the genetic algorithm; randomly generate an initial population, where each individual represents a combination of environmental factors, and the environmental factors include soil structure, density, porosity, soil pH value, types, quantities, activities of soil microorganisms, topography, and climate conditions; calculate the fitness value of each individual, and perform selection, crossover, and mutation operations, and repeat the genetic operations until the termination condition is met to obtain the final combination of key environmental factors;
[0065] Step 4, obtain microbial activity data based on the soil respiration rate in the microbial sample;
[0066] Step 5, predict the soil buffering capacity value according to the final combination of key environmental factors, microbial activity data, and the trained neural network model;
[0067] Step 6, divide the soil into different buffering capacity levels according to the soil buffering capacity value.
[0068] In the embodiments of the present invention, through DNA extraction and high-throughput sequencing technologies, the genetic information of soil microorganisms can be accurately obtained. This technology has the characteristics of high throughput and high efficiency, can process a large number of samples at one time, and improves the efficiency and accuracy of the evaluation method. By analyzing the sequencing data, the community structure and diversity of soil microorganisms can be comprehensively understood. The construction of the microbial community diversity index helps to quantitatively evaluate the richness and evenness of the microbial community, thus more accurately reflecting the health status and buffering capacity of the soil ecosystem. Using a genetic algorithm to encode the microbial community diversity index and calculating the fitness function value to optimize the combination of environmental factors can efficiently find the key environmental factors affecting the soil buffering capacity. This method not only considers the comprehensive effects of multiple environmental factors but also can quickly find the optimal factor combination through an optimization algorithm. By measuring the soil respiration rate in microbial samples to obtain microbial activity data, it helps to understand the metabolic activities and vitality of soil microorganisms. These data are one of the important indicators for evaluating the soil buffering capacity and can reflect the dynamic changes and recovery ability of the soil ecosystem. Combining the final key environmental factor combination, microbial activity data, and the trained neural network model can accurately predict the soil buffering capacity value. This method comprehensively considers various influencing factors and improves the accuracy and reliability of the prediction results; dividing the soil into different buffering capacity levels according to the soil buffering capacity value helps to formulate targeted soil management measures and strategies. This level division method can intuitively reflect the strength of the soil buffering capacity.
[0069] In a preferred embodiment of the present invention, before DNA extraction is performed on each microbial sample to obtain DNA data, it further includes:
[0070] Dividing the target collection area into circular grids; obtaining remote sensing image data of the target area from a remote sensing satellite, specifically including: using GIS (Geographic Information System) software to open the map of the target area; setting the radius of the circular grids according to research requirements and the size of the area, such as 500 meters, 1 kilometer, etc.; using the grid generation tool of GIS software to create equally spaced circular grids on the map to ensure that the grids cover the entire target area; selecting a suitable remote sensing satellite data source, such as Landsat, Sentinel-2, etc.; querying and downloading the corresponding remote sensing image data according to the geographical location of the target area and the required time range, and decompressing and converting the format of the downloaded remote sensing images.
[0071] Preprocess the remote sensing image data. Utilize the spectral features, texture features, and shape features of the preprocessed remote sensing image, and automatically identify the vegetation types through an image classification algorithm. Label and classify the vegetation types within each grid to obtain the vegetation type distribution. Specifically, it includes: opening the remote sensing image using remote sensing image processing software (such as ENVI, ERDAS Imagine, etc.) and performing preprocessing operations, including atmospheric correction, geometric correction, radiometric calibration, etc., to improve the image quality; extracting spectral features from the preprocessed remote sensing image, which include the reflectance values of each band, vegetation indices (such as NDVI, EVI, etc.), water body indices (such as NDWI), etc., and they can reflect the spectral reflection and absorption characteristics of ground objects; using the gray-level co-occurrence matrix (GLCM) to extract the texture features of the image, and these features can describe the surface structure and roughness of ground objects.
[0072] Extract the shape features of ground objects through edge detection methods, such as area, perimeter, aspect ratio, complexity, etc., and these features help to distinguish ground object types with different geometric shapes; according to prior knowledge or field survey data, select representative ground object sample points (such as forests, grasslands, water bodies, etc.) and assign corresponding class labels; use the extracted spectral, texture, and shape features as inputs, and the corresponding class labels as outputs to train a random forest classifier. During the training process, optimize the model performance by adjusting parameters such as the number of decision trees and the maximum depth; apply the trained random forest model to the entire remote sensing image, classify and predict each pixel to obtain a preliminary vegetation type classification result; by setting an area threshold, remove small patches in the classification result that are too small in area and may be noise or misclassified to improve the accuracy of the classification result; use a morphological filtering algorithm to smooth the boundaries of the classification result to reduce the appearance of jagged boundaries and fragmented patches; evaluate the accuracy of the classification result by constructing a confusion matrix and calculating metrics such as overall accuracy (OA), producer accuracy (PA), and user accuracy (UA); convert the classification result from pixel level to vector data (such as Shapefile format) or raster data (such as GeoTIFF format).
[0073] According to the distribution of vegetation types, by calculating the differences in spectral characteristics between vegetation and bare soil, estimate the vegetation coverage, specifically including: determining the size of the divided grid, dividing the study area into grids of a specified size, and based on the previous vegetation type classification results, assigning an identifier to each pixel. For example, vegetation pixels are marked as 1 and non-vegetation pixels are marked as 0; for each grid, traverse all the pixels therein and count the number of pixels marked as vegetation (1) and non-vegetation (0); compare the reflectance differences between vegetation and bare soil in specific bands (such as the red band and the near-infrared band), and based on these differences, set an appropriate threshold to further distinguish vegetation and non-vegetation pixels. For example, the NDVI (Normalized Difference Vegetation Index) value can be used. Generally, the NDVI value of vegetation is higher than that of bare soil; apply the set threshold to reclassify the pixels within the grid and update the number of vegetation and non-vegetation pixels; for each grid, use the formula "vegetation coverage = number of vegetation pixels / total number of pixels" to calculate the vegetation coverage, and save the vegetation coverage value of each grid to a new data layer or data table.
[0074] According to the remote sensing image data, by calculating the number of vegetation pixels per unit area, obtain the density of vegetation within the grid, specifically including: overlaying the divided grid data with the loaded remote sensing image data; for each grid, traverse all the pixels therein and count the number of pixels identified as vegetation, which can be done by checking the classification label of each pixel (such as vegetation marked as 1 and non-vegetation marked as 0 in the previous steps). To standardize the comparison, the number of vegetation pixels can be converted into a density per unit area, such as the number of vegetation pixels per square kilometer or per hectare. According to historical data, set different thresholds to classify the density of vegetation. For example, three thresholds of high, medium, and low can be set, representing dense, medium, and sparse vegetation states respectively. For each grid, classify it as dense, medium, or sparse according to the density of vegetation pixels per unit area, which can be done by comparing the density of vegetation pixels with the set thresholds. Create a new data layer or attribute table to store the evaluation results of the density of vegetation for each grid, and fill the classification results of the density of vegetation for each grid into the newly created data layer or attribute table.
[0075] Within each circular grid, according to the distribution of vegetation types, coverage, and the density of vegetation, determine a corresponding collection radius for each grid, where the collection radius The calculation formula is:
[0076] ;
[0077] where represents the base radius, the initial radius set according to the total area of the study area and the budget (the default value can be set to 100m). For example, Rb = 50 m to 200 m (to be adjusted according to the actual area size); Represents the vegetation type weight coefficient, which adjusts the influence intensity of the vegetation type on the radius. a = 0.4 to 0.6 (default 0.5); Represents the value of the vegetation type, the influence of different vegetation types on soil heterogeneity (to be predefined); Forest = 1.2; Shrub = 1.0; Grassland = 0.8; Bare soil = 0.5 (higher soil homogeneity); β represents the coverage adjustment coefficient, which controls the correction effect of vegetation coverage on the radius, β = 0.2 to 0.3 (default 0.25); Represents the vegetation coverage, the proportion of vegetation pixels in the grid (unit: percentage), with a value range of 0% (bare soil) to 100% (fully covered); Represents the density adjustment coefficient, which reflects the adjustment weight of vegetation density on the radius, γ = 0.1 to 0.2 (default 0.15); Represents the vegetation density, a graded value divided by the vegetation pixel density per unit area, dense = 3; medium = 2; sparse = 1 (the higher the value, the higher the complexity of the microbial habitat).
[0078] In each grid, a specific sampling point is determined according to the sampling radius, and microbial samples are obtained through the sampling point. Specifically, it includes: loading the grid data with the sampling radius, for each grid, generating a buffer according to its sampling radius, randomly or according to specific rules (such as the center point) selecting a specific sampling point in the buffer, recording the geographical location information (latitude and longitude or coordinates) of each sampling point, finding the sampling point according to the recorded geographical location information, and collecting microbial samples.
[0079] In the embodiments of the present invention, by dividing the target area into circular grids, the distribution of sampling points can be systematically planned to ensure the representativeness and uniformity of the sampling points. This grid method helps to avoid the randomness and overlap of sampling points, and improves the sampling efficiency and accuracy. Remote sensing image data provides large-scale and high-precision surface information, and can quickly obtain the vegetation distribution and land use conditions of the target area; preprocessing can eliminate the noise and interference factors in the remote sensing image, improve the image quality, and automatically identify the vegetation type through the image classification algorithm, so as to quickly and accurately obtain the vegetation information in each grid, avoiding the cumbersome and time-consuming traditional manual survey, and improving the work efficiency and accuracy. Vegetation coverage is one of the important indicators reflecting the surface vegetation conditions. By calculating the difference in spectral characteristics between vegetation and bare soil to estimate the vegetation coverage, the vegetation density and ecological conditions of the target area can be quantitatively evaluated; by calculating the number of vegetation pixels per unit area to evaluate the vegetation density, the description of the vegetation conditions can be further refined. This quantification method helps to more accurately understand the ecological environment characteristics in each grid; determining the sampling radius for each grid according to the vegetation type distribution, coverage and vegetation density can ensure that the selection of sampling points is more reasonable. This method fully considers the influence of different vegetation types and densities on the distribution of soil microorganisms, and improves the pertinence and representativeness of the sampling points.
[0080] In step 1 above, DNA extraction is performed on each microbial sample to obtain DNA data; high-throughput sequencing technology is used to sequence the DNA data to obtain sequencing data, which may include:
[0081] Collect microbial samples from the target area, which usually involves collecting soil or other samples containing microorganisms at the selected sampling point using sterile tools (such as sterile spatulas or cotton swabs) and placing them in sterile containers to avoid external contamination. The collected microbial samples are initially processed to remove impurities and enrich microbial cells, which can include washing with saline or other buffers, and separating microbial cells by physical or chemical methods (such as centrifugation or filtration); in the processed microbial samples, specific lysis buffers or enzymes are used to destroy the cell walls and cell membranes of microorganisms to release DNA. This step is the key to DNA extraction because it ensures that DNA can be effectively released from the cells. After lysis, the sample will contain DNA and other cellular components (such as proteins, RNA, etc.). In order to obtain pure DNA, a series of purification steps are used, such as phenol-chloroform extraction, ethanol precipitation, or the use of specific DNA purification kits. These steps are designed to remove impurities while retaining DNA; after extraction, the quality and quantity of DNA are tested, which is usually done by UV-visible spectrophotometer, gel electrophoresis or other quantitative methods. This step is to ensure that the DNA sample is suitable for subsequent sequencing analysis. After DNA extraction and quality confirmation, preparation for high-throughput sequencing includes constructing sequencing libraries, which involves steps such as DNA fragmentation, connecting adapters, and performing PCR amplification to generate DNA templates suitable for sequencer reading. Finally, the prepared DNA library is loaded onto a high-throughput sequencer (such as an Illumina sequencer) for sequencing. The sequencer reads the sequence information of each DNA fragment and generates a large amount of sequencing data (usually millions of short sequence reads); after sequencing, the generated sequencing data is quality controlled, which includes removing low-quality reads, trimming sequencing adapter sequences, and checking the overall quality of the data. This step is to ensure that subsequent data analysis is based on accurate and reliable sequencing data. Through the above detailed steps, DNA can be extracted from microbial samples and its genetic information can be obtained using high-throughput sequencing technology. These data will then be used to analyze the community composition and diversity of soil microorganisms and to assess the buffering capacity of the soil.
[0082] In a preferred embodiment of the present invention, the community composition of soil microorganisms is analyzed according to the sequencing data; and a microbial community diversity index is constructed according to the community composition of soil microorganisms, including:
[0083] Obtain paired sequencing data and perform preprocessing to obtain preprocessed paired sequencing data, namely Read 1 and Read 2. Specifically, it includes: obtaining the original sequencing data from the sequencer, which are usually stored in FASTQ format; performing preprocessing on the data, including removing low-quality reads, removing sequencing adapters, trimming the ends of low-quality sequences, etc.; after preprocessing, high-quality paired sequencing data, namely Read 1 and Read 2, will be obtained.
[0084] Based on all the sequences in Read 1 and Read 2, construct a de Bruijn graph. Each node in the de Bruijn graph represents a k-mer of a fixed length, that is, a continuous k nucleotides, and the edges represent the connection relationships between k-mers. Specifically, it includes: selecting an appropriate k value (the length of the k-mer), usually determined according to the characteristics of the sequencing data and the analysis requirements, traversing all the sequences in Read 1 and Read 2, and cutting each sequence into segments of length k (k-mers); in the de Bruijn graph, each k-mer serves as a node, and if two k-mers have k - 1 identical nucleotides, an edge is added between the corresponding nodes in the graph.
[0085] Utilize the characteristics of the de Bruijn graph to search for shared k-mers in Read 1 and Read 2; identify overlapping nucleotide sequences, that is, overlapping regions, by traversing the de Bruijn graph and comparing the k-mer paths in different reads; based on the overlapping regions, perform nucleotide matching at the corresponding positions in Read 1 and Read 2. Specifically, it includes: searching for the k-mer nodes that coexist in Read 1 and Read 2 in the de Bruijn graph. These shared k-mers are the key points for subsequent sequence assembly; by traversing the de Bruijn graph, comparing the k-mer paths in Read 1 and Read 2, finding consecutive overlapping k-mers in the paths. These overlapping regions are the overlapping regions between the two reads. In the overlapping regions, perform pairwise alignment and matching of the nucleotides in Read 1 and Read 2.
[0086] After successful matching, merge the two reads in the overlapping region to splice them into a single sequence. Specifically, once the nucleotides are successfully matched in the overlapping region, Read 1 and Read 2 can be merged in this region. The merged sequence is a longer single sequence spliced from the two original reads.
[0087] Cluster analysis is performed on the single sequences to form operational taxonomic units (OTUs). Representative sequences of each OTU are selected and compared with a known microbial reference database to determine the species classification information of each OTU. Specifically, it includes: using clustering algorithms (such as UCLUST, CD-HIT, etc.) to cluster the spliced single sequences. The purpose of clustering is to group similar sequences into one set to form an operational taxonomic unit (OTU). Each OTU represents a microbial species or subspecies. Select a representative sequence from each OTU and use alignment tools (such as BLAST, BOWTIE, etc.) to compare these representative sequences with the known microbial reference database. The comparison results will give the most likely species classification information for each OTU.
[0088] Based on the species classification information, calculate the number of sequences of different species in each sample, and convert the number of sequences of different species into relative abundances to generate a species abundance table. Specifically, it includes: according to the comparison results, count the number of sequences of different species (or OTUs) in each sample, convert these numbers into relative abundances, that is, the proportion of each species in the sample, and organize these data to generate a species abundance table listing the relative abundances of each species in each sample.
[0089] Using the species abundance table, calculate species richness indices and diversity indices. The diversity indices include Shannon diversity index and Simpson diversity index. Specifically, it includes: using the data in the species abundance table, calculate species richness indices, such as the number of OTUs, Chao1 index, etc., to evaluate the richness of species in the sample. At the same time, calculate diversity indices, such as Shannon diversity index and Simpson diversity index, to quantify the diversity level of species in the sample.
[0090] In the embodiments of the present invention, through preprocessing, low-quality reads, sequencing errors, and impurity data can be removed, thereby improving the accuracy and reliability of subsequent analysis; the de Bruijn graph can effectively represent the relationships between k-mers in sequencing data, helping to quickly identify and align common patterns in sequences, providing a basis for subsequent overlap region identification and sequence assembly; by searching for shared k-mers in Read 1 and Read 2, the overlapping region between the two reads can be accurately found; accurately identifying the overlapping region and performing nucleotide matching are the prerequisites for sequence assembly, which helps to recover longer and original DNA sequences from paired sequencing data; by merging overlapping reads, a more complete DNA sequence can be obtained; through clustering analysis, similar sequences can be grouped into one set to form OTUs, which helps to simplify the data and reduce the complexity of analysis while retaining the species diversity information; by comparing with known microbial reference databases, the species classification information of each OTU can be accurately determined, and the species abundance table provides the relative quantity information of different species in each sample, which is the basis for calculating species richness and diversity indices. The species richness index and diversity indices (such as the Shannon diversity index and Simpson diversity index) can quantitatively describe the complexity and diversity of the microbial community.
[0091] In a preferred embodiment of the present invention, calculating the fitness value of each individual includes:
[0092] Determining the Shannon diversity index and Simpson diversity index according to the richness of microbial species in the soil and the microbial community; determining the contribution value of soil microbial diversity according to the Shannon diversity index and Simpson diversity index;
[0093] Determining the Pearson correlation coefficient between each environmental factor and the soil buffering capacity; determining the correlation contribution value according to each environmental factor and its value in the current individual;
[0094] For each microbial activity data sample, determining its corresponding expected value and the standard deviation of the microbial activity data; obtaining the standardized deviation of the sample according to each microbial activity data sample and the corresponding expected value and the standard deviation of the microbial activity data;
[0095] According to the standardized deviation of the sample and the soil respiration rate of the sample, to obtain the contribution value of the microbial activity data;
[0096] Fusing the correlation contribution value, the standardized deviation of the sample, and the contribution value of the microbial activity data to obtain the fitness value of each individual.
[0097] In an embodiment of the present invention, according to the richness of microbial species and the microbial community structure in the soil, the Shannon diversity index S and the Simpson diversity index m are calculated; the Shannon diversity index S reflects the richness of microbial species in the soil, while the Simpson diversity index m reflects the proportional relationship between the dominant species and rare species in the microbial community. By combining these two indices, the contribution of soil microbial diversity to the fitness value can be determined, that is, the soil microbial diversity contribution value; the Pearson correlation coefficient between each environmental factor and the soil buffering capacity is calculated to measure the degree of linear correlation between them. According to the value of each environmental factor in the current individual and its correlation coefficient with the soil buffering capacity, the correlation contribution value of the environmental factor to the fitness value is determined. For each microbial activity data sample, its corresponding expected value (i.e., the average value) and the standard deviation of the microbial activity data are calculated. According to the soil respiration rate, the expected value, and the standard deviation of each sample, the standardized deviation of the sample is calculated to reflect the degree of deviation of the sample from the average value. By combining the standardized deviation of the sample and the soil respiration rate, the contribution of the microbial activity data to the fitness value is determined, that is, the microbial activity data contribution value. The soil microbial diversity contribution value, the correlation contribution value between the environmental factor and the soil buffering capacity, and the microbial activity data contribution value are weighted and fused to obtain the fitness value of each individual.
[0098] By comprehensively considering the soil microbial diversity, the correlation between the environmental factor and the soil buffering capacity, and the microbial activity data, the buffering capacity of the soil can be evaluated more comprehensively and accurately. The fitness value can be used as an important basis for soil management and improvement. By comparing the fitness values of different individuals, an environmental factor combination with a higher soil buffering capacity can be selected, thereby optimizing the soil management strategy and improving the soil quality. In ecological restoration projects, the fitness value can be used to screen out the environmental factor combination most suitable for the current soil conditions, thereby accelerating the restoration and reconstruction of the soil ecosystem and improving the ecological restoration efficiency.
[0099] In a preferred embodiment of the present invention, when specifically applied, the calculation formula of the fitness function value can be:
[0100] ;
[0101] Wherein, represents the expected value of the soil respiration rate; represents the standard deviation of the microbial activity data (soil respiration rate); represents the Pearson correlation coefficient between the th environmental factor and the soil buffering capacity; represents the value of the th environmental factor of the current individual; represents the absolute value of the maximum correlation coefficient; represents the Shannon diversity index; represents the Simpson diversity index; represents the soil respiration rate of the th sample; and represent the weight coefficients; represents the index of the environmental factor, ranging from 1 to ; represents the total number of environmental factors; represents the index of the microbial activity data sample, ranging from 1 to ;
[0102] In the embodiments of the present invention, the calculation of the fitness function value takes into account multiple ecological and environmental factors, and can more comprehensively and accurately evaluate the health status of the soil ecosystem. By introducing the Shannon diversity index and the Simpson diversity index, the function can reflect the soil biodiversity, which is an important indicator of the stability and functional diversity of the soil ecosystem. At the same time, by considering the correlation between environmental factors and soil buffering capacity, the function can evaluate the adaptability of the soil to environmental changes. In addition, by introducing the expected value and fluctuation range of the soil respiration rate, the active state of soil microorganisms can be further understood, so as to more comprehensively evaluate the soil quality.
[0103] Calculate the Shannon diversity index (S) and the Simpson diversity index (m) based on the soil sample data. Both of these indices are calculated based on species richness and evenness, and are used to measure soil biodiversity.
[0104] Calculate the correlation between environmental factors and soil buffering capacity:
[0105] For each environmental factor (such as temperature, humidity, pH, etc.), calculate its Pearson correlation coefficient with the soil buffering capacity, and find the absolute value of the maximum correlation coefficient .
[0106] Calculate the expected value and fluctuation range of the soil respiration rate:
[0107] Based on the microbial activity data samples, calculate the expected value and the fluctuation range of the soil respiration rate, which can be completed by statistical analysis of the sample data.
[0108] Calculate the components of the calculation of the fitness function value:
[0109] The first part: , reflecting the contribution of soil biodiversity to fitness.
[0110] Part II: , reflecting the contribution of the correlation between environmental factors and soil buffering capacity to fitness.
[0111] Part III: , reflecting the contribution of soil microbial activity to fitness.
[0112] Adding the above three parts together gives the total fitness value for each individual, which can be used to evaluate the comprehensive status of the soil ecosystem.
[0113] For example, to evaluate the health status of the soil and microbial activity, using the above calculation of the fitness function value, relevant environmental factor data, soil diversity index, and microbial activity data were collected, and weight coefficients were set.
[0114] For environmental factor data, n environmental factors such as soil temperature, humidity, pH value, and organic matter content were measured, and the value of each environmental factor was obtained .
[0115] For the soil diversity index, by sampling and analyzing the microbial community in the soil, the Shannon diversity index S and Simpson diversity index m were calculated.
[0116] For microbial activity data, the respiration rate data of N soil samples were collected , and the expected value of the soil respiration rate was calculated and the standard deviation of the soil respiration rate .
[0117] Correlation coefficient calculation, calculating the Pearson correlation coefficient between each environmental factor and soil buffering capacity , and finding the absolute value of the maximum correlation coefficient .
[0118] Using the above - collected data and set weight coefficients, each item of data was calculated. Through calculation, the fitness value F of the soil ecosystem was obtained; according to the calculated fitness value F, the comprehensive status of the farmland soil ecosystem can be evaluated. A higher fitness value indicates that the soil ecosystem has higher biodiversity, good adaptability to environmental factors, and an active microbial community.
[0119] In a preferred embodiment of the present invention, the fitness value of each individual is calculated, and selection, crossover, and mutation operations are performed. The genetic operations are repeated until the termination condition is met to obtain the final key environmental factor combination, including:
[0120] Randomly generate an initial population. Each individual in the population represents a combination of key environmental factors, specifically including: determining the number and range of key environmental factors, using a random number generator to generate random values for each environmental factor within this range to form an individual, and repeating the above process until a sufficient number of individuals are generated to form the initial population.
[0121] Calculate the fitness value of each individual. Based on the fitness value of the individual, select the corresponding individual for crossover operation to generate new offspring individuals, specifically including: for each individual in the population, extract the key environmental factor values it represents, calculate based on these values, and record the fitness value of each individual, which will be used for subsequent selection operations; adopt methods such as roulette wheel selection and tournament selection to select which individuals will participate in the crossover operation according to the fitness value of the individual. Generally, individuals with higher fitness values have a greater probability of being selected. Once a pair of parent individuals is selected, the crossover operation can be carried out.
[0122] Perform mutation operations on the offspring individuals. The new individuals generated through selection, crossover, and mutation operations form a new generation of population until the preset number of generations of evolution is reached, and the final individual in the current population is taken as the final solution to the problem; according to the final solution, the final combination of key environmental factors is obtained, specifically including: randomly select one or more points in the gene string of the parent individual as crossover points, exchange the genes of the two parent individuals at the crossover points to generate two new offspring individuals. For each offspring individual, determine whether to mutate with a certain mutation probability. If it is decided to mutate, randomly select one or more gene positions for mutation. The mutation can be to change the value of the gene, or in some cases, it can be operations such as gene insertion, deletion, or inversion; the newly generated offspring individuals and the parent individuals together form a temporary mixed population; according to the fitness value, select a new generation of population from this mixed population, which is usually achieved by selecting individuals with higher fitness values. The new generation of population will be used for the next round of genetic operations. Perform multiple rounds of genetic operations. Each time a round is carried out, the generation counter is incremented. When the number of generations of evolution reaches the preset value, stop the genetic operation; after reaching the preset number of generations of evolution, select the individual with the highest fitness from the last generation of the population. The combination of key environmental factors represented by this individual is the optimal or approximate optimal solution found by the algorithm. Extract the gene information of this individual, that is, the values of the key environmental factors, as the final solution to the problem.
[0123] In the embodiments of the present invention, by randomly generating an initial population, the algorithm can start exploring in a wide search space, which helps to avoid falling into local optimal solutions and thus is more likely to find the global optimal solution. The calculation of the fitness function value quantifies the degree of adaptation of an individual to the problem. By evaluating the fitness values of each individual, the algorithm can distinguish which individuals are more likely to approach the optimal solution. The selection operation ensures that more excellent individuals have a greater chance to participate in reproduction, thereby passing on their excellent genes to the next generation, which helps the algorithm to quickly converge to an excellent solution space. The crossover operation simulates the biological hybridization process and produces potentially more excellent offspring individuals by combining the excellent genes of parental individuals, which helps the algorithm to effectively explore the search space to find better solutions. The mutation operation simulates gene mutations and provides the algorithm with the ability to jump out of local optimal solutions. By randomly changing some gene values, the algorithm may discover new and better search regions. The new generation of the population inherits the excellent genes of the previous generation and introduces new mutations, which increases the diversity of the population and helps the algorithm to search for the optimal solution globally. Through multiple iterations, the algorithm can gradually approach the optimal solution. Each generation eliminates unfit individuals, retains and reproduces excellent individuals, thereby continuously improving the overall fitness of the population. After multiple generations of genetic operations, the algorithm finally converges to one or several excellent individuals, which represent the optimal or approximate optimal solutions to the problem, and these solutions represent the optimal combination of key environmental factors.
[0124] For example, due to the long-term use of chemical fertilizers and improper farmland management in a certain area, the soil buffering capacity has gradually decreased, resulting in the impact on crop yield and quality. To improve this situation, a genetic algorithm is used to find the best soil improvement plan.
[0125] Determine the key environmental factors:
[0126] Soil organic matter content (range: 1% - 5%);
[0127] Soil pH value (range: 5.5 - 8.5);
[0128] Soil microbial population diversity index (range: 1 - 10);
[0129] Soil cation exchange capacity (cmol / kg, range: 5 - 30);
[0130] Use a random number generator to generate random values for each environmental factor within the specified range, combine these random values to form an "individual", representing a soil improvement plan. Repeat this process to generate 100 individuals to form the initial population.
[0131] The calculation of the fitness function value takes into account multiple dimensions of soil buffering capacity, such as acid-base balance, nutrient retention, microbial activity, etc.; the roulette wheel selection method is used to select the parent individuals participating in crossover according to the fitness values of the individuals. Individuals with higher fitness have a higher probability of being selected. Two parent individuals are randomly selected, and a crossover point is randomly selected in the gene string of each parent, and the gene segments after the crossover point of the two parents are exchanged to generate two new offspring individuals. For each offspring individual, with a probability of 0.05, it is determined whether to mutate. If it is determined to mutate, a gene position (i.e., a certain environmental factor) is randomly selected, and a small random adjustment is made to this gene, such as increasing or decreasing its value by 1%; the newly generated offspring individuals and the parent individuals together form a mixed population. According to the fitness values, the top 100 individuals with the highest fitness are selected from the mixed population to form a new generation of population. The selection, crossover, and mutation operations are repeated for the new generation of population. Each time a round of operations is performed, the generation number of evolution increases by 1. When the generation number of evolution reaches 50 generations, the genetic operation is stopped. After reaching the termination condition, the individual with the highest fitness is selected from the last generation of population as the optimal solution, and the combination of environmental factors represented by this individual is the recommended soil improvement plan.
[0132] In step 4 above, based on the soil respiration rate in the microbial sample to obtain the microbial activity data, it may include:
[0133] Select representative sampling points that can reflect the overall situation of the study area. Use appropriate tools (such as soil samplers) to collect soil samples to ensure the integrity and pollution-free of the samples. Put the collected soil samples into appropriate containers and mark information such as sampling time and location; bring the soil samples back to the laboratory and process them as soon as possible to avoid sample deterioration. Remove impurities such as stones and roots in the samples to ensure the purity of the samples. If necessary, grind or sieve the soil to obtain uniform soil particles; prepare measurement equipment, such as a soil respiration measurement system or an infrared gas analyzer. Place the soil samples in the measurement equipment to ensure good sealing. Start measuring and record the amount of carbon dioxide released by the soil, which is usually achieved by measuring the change in carbon dioxide concentration per unit time. The measurement process lasts for a period of time (such as several hours or one day) to obtain accurate respiration rate data; organize the measured soil respiration rate data, including information such as time and carbon dioxide release amount, and analyze the data, such as calculating statistical indicators such as average value and standard deviation; soil respiration rate is one of the important indicators reflecting soil microbial activity. A higher respiration rate usually means more active microbial activity. By comparing the soil respiration rate data of different samples or different time points, the changes and differences in microbial activity can be evaluated.
[0134] In a preferred embodiment of the present invention, according to the final key environmental factor combination, microbial activity data, and the trained neural network model, the soil buffering capacity value is predicted, including:
[0135] Select features related to soil buffering capacity from the key environmental factor combination and microbial activity data, and combine the selected features into a feature vector. Specifically, it includes: collecting data on the key environmental factor combination, including soil organic matter content, soil pH value, soil microbial population diversity index, and soil cation exchange capacity, etc. At the same time, data on microbial activity also need to be collected, including indicators such as the types, quantities, and activities of microorganisms; calculating the Pearson correlation coefficient between each feature and soil buffering capacity, and recording the correlation coefficient value between each feature and the target variable; according to the calculated correlation coefficient, determining a threshold (such as Pearson correlation coefficient > 0.5), and screening out features significantly related to soil buffering capacity; for the features most related to soil buffering capacity screened out, perform standardization processing;
[0136] Confirm that all selected features (such as soil organic matter content, pH value, microbial diversity index, and cation exchange capacity) have undergone standardization processing; collect all the standardized feature data together, and these data are stored in a data frame (DataFrame) or an array, with each feature corresponding to a column or a row; determine the order of the features according to the needs or the expected input format of the model; the order can be determined according to the column order in the dataset; according to the determined order, extract the corresponding values from each feature, and for each data point (for example, each soil sample), combine these values into a one-dimensional array, and this one-dimensional array is a feature vector, which contains the standardized values of all selected features of this data point;
[0137] Input the feature vector into the trained neural network model to obtain the soil buffering capacity value, which specifically includes: dividing the dataset containing the feature vector and the corresponding soil buffering capacity value into a training set, a validation set, and a test set. Usually, the proportions of the training set, validation set, and test set are 70%, 15%, and 15% respectively. Design the neural network architecture, including an input layer, hidden layers, and an output layer, and set initial values for the weights and biases of the neural network, which are usually random; determine a suitable loss function to measure the difference between the model prediction and the actual value, such as the mean squared error (MSE); select an optimization algorithm, such as stochastic gradient descent (SGD), to update the weights and biases of the model during training; determine training parameters such as the number of training epochs, batch size, and learning rate; load the training set data into memory and perform necessary preprocessing, such as feature scaling; in each training epoch, input the training data in batches into the neural network, calculate the loss function value, and update the model parameters through backpropagation and the optimizer. After each epoch, evaluate the performance of the model using the validation set, and adjust the model architecture or training parameters as needed to optimize the performance; use the test set to evaluate the performance of the trained neural network model to ensure that the model has good generalization ability, and save the trained neural network model as a file; input the combined feature vector into the loaded neural network model, and the neural network model predicts the soil buffering capacity value.
[0138] In the embodiments of the present invention, by combining key environmental factors and microbial activity data, various factors affecting soil buffering capacity can be considered more comprehensively. This multi-dimensional data integration enables the neural network model to learn more complex non-linear relationships, thereby improving the accuracy of predicting soil buffering capacity. Accurate prediction of soil buffering capacity provides an important basis for agricultural, environmental protection, and land management decisions. For example, in agricultural production, understanding the buffering capacity of the soil helps to formulate reasonable fertilization and irrigation plans, improving crop yield and quality. Through accurate prediction of soil buffering capacity, resources such as water resources, fertilizers, and pesticides can be allocated and utilized more effectively, which not only helps to save costs but also reduces environmental pollution; by predicting and monitoring soil buffering capacity, soil degradation or pollution problems can be detected in a timely manner, and corresponding measures can be taken for protection and restoration.
[0139] As Figure 2 shown, the embodiments of the present invention also provide an evaluation system based on soil buffering capacity, including:
[0140] An extraction module for extracting DNA from each microbial sample to obtain DNA data; sequencing the DNA data using high-throughput sequencing technology to obtain sequencing data;
[0141] A construction module for analyzing the community composition of soil microorganisms based on sequencing data; constructing a microbial community diversity index according to the community composition of soil microorganisms;
[0142] An optimization module for encoding the microbial community diversity index as the genotype of a genetic algorithm; randomly generating an initial population, where each individual represents a combination of environmental factors; calculating the fitness value of each individual, and performing selection, crossover, and mutation operations, repeating the genetic operations until the termination condition is met to obtain the final combination of key environmental factors;
[0143] An acquisition module for obtaining microbial activity data based on the soil respiration rate in a microbial sample;
[0144] A prediction module for predicting the soil buffering capacity value based on the final combination of key environmental factors, microbial activity data, and a trained neural network model;
[0145] A division module for dividing the soil into different buffering capacity levels according to the soil buffering capacity value.
[0146] It should be noted that this system corresponds to the above method. All implementation manners in the above method embodiments are applicable to this embodiment and can achieve the same technical effects.
[0147] An embodiment of the present invention also provides a computing device, including: a processor and a memory storing a computer program. When the computer program is run by the processor, it executes the method as described above. All implementation manners in the above method embodiments are applicable to this embodiment and can achieve the same technical effects.
[0148] An embodiment of the present invention also provides a computer-readable storage medium storing instructions. When the instructions are run on a computer, the computer is made to execute the method as described above. All implementation manners in the above method embodiments are applicable to this embodiment and can achieve the same technical effects.
Claims
1. An evaluation method based on soil buffering capacity, characterized in that The method includes: Step 1, extracting DNA from each microbial sample to obtain DNA data; sequencing the DNA data using high-throughput sequencing technology to obtain sequencing data; Step 2, analyzing the community composition of soil microorganisms based on the sequencing data; constructing a microbial community diversity index according to the community composition of soil microorganisms; Step 3, encoding the microbial community diversity index as the genotype of the genetic algorithm; randomly generating an initial population, where each individual represents a combination of environmental factors; calculating the fitness value of each individual, and performing selection, crossover, and mutation operations, repeating the genetic operations until the termination condition is met to obtain the final key environmental factor combination, where calculating the fitness value of each individual includes: Determining the Shannon diversity index and Simpson diversity index according to the richness of microbial species and the microbial community in the soil; determining the contribution value of soil microbial diversity according to the Shannon diversity index and Simpson diversity index; determining the Pearson correlation coefficient between each environmental factor and the soil buffering capacity; determining the correlation contribution value according to each environmental factor and its value in the current individual; for each microbial activity data sample, determining its corresponding expected value and the standard deviation of the microbial activity data; obtaining the standardized deviation of the sample according to each microbial activity data sample and the corresponding expected value and the standard deviation of the microbial activity data; obtaining the contribution value of the microbial activity data according to the standardized deviation of the sample and the soil respiration rate of the sample; fusing the correlation contribution value, the standardized deviation of the sample, and the contribution value of the microbial activity data to obtain the fitness value of each individual; Step 5, obtaining microbial activity data according to the soil respiration rate in the microbial sample; Step 6, predicting the soil buffering capacity value according to the final key environmental factor combination, microbial activity data, and the trained neural network model; Step 7, dividing the soil into different buffering capacity levels according to the soil buffering capacity value.
2. The evaluation method based on soil buffering capacity according to claim 1, wherein, Before extracting DNA from each microbial sample to obtain DNA data, it further includes: Dividing the target collection area into circular grids; obtaining remote sensing image data of the target area from a remote sensing satellite; Preprocessing the remote sensing image data, and automatically identifying the vegetation type through an image classification algorithm using the spectral features, texture features, and shape features of the preprocessed remote sensing image, and labeling and classifying the vegetation type in each grid to obtain the vegetation type distribution; Estimating the vegetation coverage according to the vegetation type distribution by calculating the difference in spectral features between vegetation and bare soil; Obtaining the density of vegetation in the grid by calculating the number of vegetation pixels per unit area according to the remote sensing image data; In each circular grid, determining a corresponding collection radius for each grid according to the vegetation type distribution, coverage, and density of the vegetation; In each grid, determining a specific collection point according to the collection radius, and obtaining microbial samples through the collection point.
3. The evaluation method based on soil buffering capacity according to claim 2, wherein Environmental factors include soil structure, density, porosity, soil pH value, types, quantities, activities of soil microorganisms, topography and climate conditions.
4. The evaluation method based on soil buffering capacity according to claim 3, wherein Analyze the community composition of soil microorganisms based on sequencing data; construct microbial community diversity indices according to the community composition of soil microorganisms, including: Obtain paired sequencing data and perform preprocessing to obtain preprocessed paired sequencing data, namely Read 1 and Read 2; Construct a de Bruijn graph based on all sequences in Read 1 and Read 2. Each node in the de Bruijn graph represents a k-mer of a fixed length, that is, a continuous k nucleotides, and the edges represent the connection relationships between k-mers; Utilize the characteristics of the de Bruijn graph to search for shared k-mers in Read 1 and Read 2; identify overlapping nucleotide sequences, that is, overlapping regions, by traversing the de Bruijn graph and comparing k-mer paths in different reads; perform nucleotide matching at the corresponding positions of Read 1 and Read 2 according to the overlapping regions; After successful matching, merge the two reads in the overlapping region to splice them into a single sequence; Perform clustering analysis on the single sequence to form operational taxonomic units, select the representative sequences of each OTU, and compare the corresponding representative sequences with a known microbial reference database to determine the species classification information of each OTU; According to the species classification information, calculate the sequence numbers of different species in each sample, and convert the sequence numbers of different species into relative abundances to generate a species abundance table; Utilize the species abundance table to calculate species richness indices and diversity indices. The diversity indices include Shannon diversity index and Simpson diversity index.
5. The evaluation method based on soil buffering capacity according to claim 4, characterized in that Calculate the fitness value of each individual and perform selection, crossover, and mutation operations. Repeat the genetic operations until the termination condition is met to obtain the final combination of key environmental factors, including: Randomly generate an initial population, and each individual in the population represents a combination of a set of key environmental factors; Calculate the fitness value of each individual, and select the corresponding individual for crossover operation according to the fitness value of the individual to generate new offspring individuals; Perform mutation operations on the offspring individuals. The new individuals generated through selection, crossover, and mutation operations form a new generation of population until the preset number of generations of evolution is reached to obtain the final individual in the current population as the final solution to the problem; according to the final solution, obtain the final combination of key environmental factors.
6. The evaluation method based on soil buffering capacity according to claim 5, wherein Predict the soil buffering capacity value based on the final combination of key environmental factors, microbial activity data, and the trained neural network model, including: Select the characteristics related to soil buffering capacity from the combination of key environmental factors and microbial activity data, and combine the selected characteristics into a feature vector; Input the feature vector into the trained neural network model to obtain the soil buffering capacity value.
7. An evaluation system based on soil buffering capacity, characterized in that, The system is used to execute the method described in any one of claims 1 to 6, including: An extraction module for extracting DNA from each microbial sample to obtain DNA data; sequencing the DNA data using high-throughput sequencing technology to obtain sequencing data; A construction module for analyzing the community composition of soil microorganisms based on the sequencing data; constructing a microbial community diversity index according to the community composition of soil microorganisms; An optimization module for encoding the microbial community diversity index as the genotype of a genetic algorithm; randomly generating an initial population, where each individual represents a combination of environmental factors; calculating the fitness value of each individual, and performing selection, crossover, and mutation operations, repeating the genetic operations until the termination condition is met to obtain the final key environmental factor combination; An acquisition module for obtaining microbial activity data based on the soil respiration rate in the microbial sample; A prediction module for predicting the soil buffering capacity value based on the final key environmental factor combination, the microbial activity data, and the trained neural network model; A division module for dividing the soil into different buffering capacity levels according to the soil buffering capacity value.
8. A computing device, characterized in that, Comprising: One or more processors; A storage device for storing one or more programs, which when executed by the one or more processors cause the one or more processors to implement the method according to any one of claims 1 to 6.
9. A computer-readable storage medium, characterized in that, A program is stored in the computer-readable storage medium, and when the program is executed by a processor, the method according to any one of claims 1 to 6 is implemented.
Citation Information
Patent Citations
Comprehensive evaluation method for remediation effects of heavy metal contaminated soil based on plants, soil and microorganisms
CN107066823A
Microbial functional factor evaluation system and method based on data analysis
CN119229979A