Method for analyzing microbial DNA sequencing and its application in formation division
By using shale debris endogenous microbial DNA sequencing and machine learning models, the problem of low resolution in traditional stratigraphic division was solved, achieving high-resolution fine stratigraphic division, meeting the needs of sub-layer division, and improving the accuracy of geological reservoir interpretation.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- CHINA UNIV OF GEOSCIENCES (BEIJING)
- Filing Date
- 2025-12-26
- Publication Date
- 2026-07-07
AI Technical Summary
Traditional stratigraphic methods have low resolution, cannot meet the needs of subdivision, and suffer from uncertainty and high cost.
A method based on DNA sequencing of endogenous microorganisms in shale cuttings was adopted. By collecting cuttings samples during drilling, DNA sequencing, microbial screening and cluster analysis were performed. Machine learning models were used to finely divide the formation, screen out characteristic microbial species and refine them at high resolution.
It achieves a stratigraphic resolution of 1-2m, meeting the requirements for sub-layer division, with high accuracy, simple process, short cycle, and the ability to exclude contaminating bacteria, thus improving the interpretation accuracy of geological reservoirs.
Smart Images

Figure CN121737323B_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of petroleum geological engineering technology, specifically to an analytical method for microbial DNA sequencing and its application in stratigraphic division. Background Technology
[0002] Shale reservoirs typically exhibit significant heterogeneity. To better understand and predict the heterogeneity of shale oil reservoirs, it is necessary to employ appropriate technical means and engineering methods to conduct detailed reservoir delineation in order to achieve effective oil and gas exploration and development.
[0003] Stratigraphic methods are mainly classified into three categories: well logging methods, seismic methods, and geochemical methods. Well logging methods determine formation properties and interfaces by analyzing well logging curves. Real-time formation information is obtained during drilling, facilitating timely adjustments to drilling plans and formation evaluation. However, for heterogeneous and complex formations, the interpretation of well logging methods may involve a degree of subjectivity and uncertainty. Furthermore, mud filtrate intrusion during drilling can contaminate the formation and affect the logging response.
[0004] Seismic methods utilize the propagation characteristics of seismic waves in the subsurface medium. By analyzing reflected and refracted waves in seismic data, they determine the properties and interfaces of strata. Seismic methods can provide information on subsurface structures, stratum thickness, and lithological variations. The advantage of seismic stratigraphy lies in its ability to interpret stratigraphic layers between wells. However, the accuracy of seismic stratigraphy in stratigraphic interpretation is relatively low, and it contains a certain degree of uncertainty.
[0005] Moreover, the vertical or horizontal resolution of traditional stratigraphic division and description methods is generally 2-5m, which is low and cannot meet the needs of small-layer division. Furthermore, it is limited by instrument accuracy and data acquisition, and its reliability and resolution are limited by sampling point density and sample quality. It also suffers from problems such as low accuracy, high cost, long cycle, complex process, high requirements for the field, and easy contamination of the reservoir. Summary of the Invention
[0006] The purpose of this invention is to provide an analytical method for microbial DNA sequencing and its application in stratigraphy, in order to solve the technical problem that the resolution of traditional stratigraphy and description methods is generally 2-5m, which is low and cannot meet the requirements of sub-layer division.
[0007] To address the aforementioned technical problems, this invention specifically provides a stratigraphic method based on endogenous microbial DNA sequencing of shale debris, comprising the following steps:
[0008] During the drilling process, rock cuttings were sampled from the vibrating screen at the wellhead, and rock cuttings samples were collected every 0.5 meters to collect rock cuttings samples from different formation depths.
[0009] The collected rock fragments were processed and subjected to DNA sequencing to obtain the types and abundance of bacteria contained in the rock fragments;
[0010] Preliminary screening of bacterial strains was conducted, and strains from underground oil reservoirs were retained;
[0011] Cluster analysis was performed on the preserved microbial species, and the total length of the collected rock cuttings was initially divided into multiple strata based on the similarity and correlation of microbial species at different strata depths.
[0012] Each stratum is finely divided into multiple sub-layers based on the characteristic fungal species found in each stratum.
[0013] As a preferred embodiment of the present invention, the method for preliminary screening of bacterial strains is as follows:
[0014] The characteristics of the underground reservoir environment in the target block were analyzed, and the DNA of microorganisms in the underground reservoir was sequenced to analyze the common characteristics of microorganisms in the underground reservoir environment in order to construct parameters for preliminary screening of microbial species.
[0015] The DNA sequences obtained from sequencing are compared with known microbial DNA sequences in the target area to retain strains with underground origin characteristics.
[0016] As a preferred embodiment of the present invention, based on the characteristics of underground oil reservoirs being anaerobic and high-temperature, the preliminary screening conditions are set as anaerobic or facultative anaerobic, and the suitable growth temperature is greater than 20°C.
[0017] As a preferred embodiment of the present invention, the clustering analysis method is as follows:
[0018] Species statistical analysis at the genus level was performed on the preserved bacterial DNA sequences to obtain relative abundance maps of bacterial species in each stratum;
[0019] Clustering was performed based on matching distance, standard deviation, and column standardization direction. A ring heatmap was used to display the correlation between different microbial species in each stratum after clustering. Strata with strong correlations were grouped into the same stratum, so as to preliminarily divide the total length of the collected rock debris strata into multiple strata.
[0020] As a preferred embodiment of the present invention, the characteristic species in each stratum refers to species that are highly correlated and abundant in that stratum, but not highly correlated in other strata.
[0021] As a preferred embodiment of the present invention, the method for screening characteristic bacterial species in each stratum is as follows:
[0022] Correlation analysis was performed on the various microbial species in each stratum after preliminary division to identify the few microbial species with strong correlation in each stratum;
[0023] The highly correlated bacterial species in each stratum were compared with highly correlated bacterial species in other strata, and the highly correlated bacterial species unique to each stratum were screened from several species.
[0024] The abundance of each bacterial species in each stratum was compared to verify whether the selected bacterial species had a prominent abundance in that stratum. Species with higher abundance were identified as characteristic bacterial species.
[0025] As a preferred embodiment of the present invention, the method for correlation analysis of various bacterial species is as follows: constructing a correlation heatmap based on the DNA sequencing data of bacterial species in each stratum, and using the Pearson correlation analysis method to analyze the correlation of microbial species in each stratum in order to screen several bacterial species with strong correlation in each stratum.
[0026] The method for comparing the abundance of different species in different strata is as follows: based on the Brectis distance, standard deviation standardization method, and column standardization direction, a bubble chart is constructed to display the abundance of each species in different strata, so as to screen out the species with higher abundance in each stratum.
[0027] As a preferred embodiment of the present invention, the method for constructing a machine learning model for screening characteristic microbial species in strata is as follows: a dataset is constructed using the relative abundance information of microbial species in the horizontal direction of each stratum in the preliminary division of strata and the screening results of characteristic microbial species in each stratum.
[0028] A machine learning model for screening stratigraphic microbial species is obtained by training a classifier model containing multiple classification heads based on the dataset, wherein the number of classification heads is the same as the number m of category labels initially used to classify the stratigraphy.
[0029] The loss function of the machine learning model In the formula, For binary cross-entropy, For the cosine loss term, , The classifier model outputs the predicted characteristic bacterial species for the i-th and j-th classifier heads, initially classifying the strata into i-th and j-th category labels. The true values of characteristic bacterial species for the i-th category label in the strata are used for preliminary classification.
[0030] As a preferred embodiment of the present invention, the method for finely dividing each stratum into multiple smaller layers is as follows:
[0031] According to the preliminary division of the various strata
[0032] The abundance of characteristic bacterial species, using the abundance of the characteristic bacterial species Preliminary category labels corresponding to the abundance of characteristic bacterial species Constructing a small sample dataset And the small sample dataset is augmented to obtain a large sample dataset;
[0033] A large dataset was divided into training and test sets to train a Random Forest (RF) algorithm. Grid Search and Cross-Validation were used to optimize the hyperparameters of the Random Forest, resulting in a set for optimizing the algorithm based on the abundance of feature species. Obtain stratigraphic category labels The classifier model, where the loss function of the classifier model is... In the formula Cross-entropy;
[0034] Each stratum is divided into stratigraphic points with 1-meter intervals, and the abundance of characteristic bacterial species at each stratigraphic point is input into the classifier model to predict the category label of each stratigraphic point. and the predicted probability value corresponding to the category label. ;
[0035] Category labels are the same and the difference in predicted probabilities is less than the threshold corresponding to that category label. Adjacent equidistant stratigraphic points are merged to form smaller layers, thereby achieving a fine division of the strata.
[0036] As a preferred embodiment of the present invention, the method for verifying the accuracy of the preliminary division is as follows:
[0037] Preliminary formation division is achieved by observing abnormal fluctuations in natural gamma radiation, sealing resistivity, and sonic transit time parameters in well logging curves.
[0038] Compare the preliminary stratigraphic division results with the abnormal fluctuations in the well logging curves. If they are consistent, it indicates that the preliminary stratigraphic division results are accurate.
[0039] As a preferred embodiment of the present invention, the method for verifying whether fine segmentation improves the interpretation accuracy of geological reservoirs is as follows:
[0040] Compare the resolution of fine stratigraphic division with that of well logging methods;
[0041] If the resolution of fine stratigraphic division is higher than that of well logging curves, it indicates improved accuracy in interpreting geological reservoirs.
[0042] Compared with the prior art, the present invention has the following advantages:
[0043] This invention employs a stratigraphic division method based on endogenous microbial DNA sequencing of shale cuttings. The method sequentially performs cuttings DNA sequencing, microbial screening, initial stratigraphic division, characteristic microbial screening, and secondary fine stratigraphic division. The stratigraphic division achieves high resolution, reaching 1-2 meters, meeting the requirements for small-layer stratigraphic division. Furthermore, it boasts high accuracy, a simple process, and a short cycle time, and can exclude contaminating microorganisms. Attached Figure Description
[0044] To more clearly illustrate the embodiments of the present invention or the technical solutions in the prior art, the drawings used in the description of the embodiments or the prior art will be briefly introduced below. Obviously, the drawings described below are merely exemplary, and those skilled in the art can derive other embodiments based on the provided drawings without creative effort.
[0045] Figure 1 This is an architecture diagram of the classifier model with a three-classification head in an embodiment of the present invention;
[0046] Figure 2 This is a flowchart of the stratigraphic division method in an embodiment of the present invention. Detailed Implementation
[0047] The technical solutions of the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of the present invention, and not all embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of the present invention.
[0048] like Figure 2 As shown, this invention specifically provides a stratigraphic method based on endogenous microbial DNA sequencing of shale debris, comprising the following steps:
[0049] During the drilling process, rock cuttings were sampled from the vibrating screen at the wellhead, and rock cuttings samples were collected every 0.5 meters to collect rock cuttings samples from different formation depths.
[0050] The collected rock fragments were processed and subjected to DNA sequencing to obtain the types and abundance of bacteria contained in the rock fragments;
[0051] Preliminary screening of bacterial strains was conducted, and strains from underground oil reservoirs were retained;
[0052] Cluster analysis was performed on the preserved microbial species, and the total length of the collected rock cuttings was initially divided into multiple strata based on the similarity and correlation of microbial species at different strata depths.
[0053] Each stratum is finely divided into multiple sub-layers based on the characteristic fungal species found in each stratum.
[0054] This invention employs a stratigraphic division method based on endogenous microbial DNA sequencing of shale cuttings. The method sequentially performs cuttings DNA sequencing, microbial screening, initial stratigraphic division, characteristic microbial screening, and secondary fine stratigraphic division. The stratigraphic division achieves high resolution, reaching 1-2 meters, meeting the requirements for small-layer stratigraphic division. Furthermore, it boasts high accuracy, a simple process, and a short cycle time, and can exclude contaminating microorganisms.
[0055] The sample collection and pretreatment methods are as follows:
[0056] During the drilling process, rock cuttings were sampled from the vibrating screen at the wellhead, with samples taken approximately every 0.5 meters, for a total sampling length of 22.60 meters.
[0057] During sample collection, contact between the target sample and the sample collector should be minimized to reduce contamination. Gloves should be worn during sampling to reduce skin microbial contamination. Sample containers should be cleaned and sterilized with ultraviolet light before sampling. After sampling, the samples should be immediately placed in primary and secondary containers, then stored in dry ice or liquid nitrogen at 4°C–80°C, and transported to the laboratory immediately.
[0058] After the samples arrived at the laboratory, the rock cuttings were first cleaned to remove impurities, mud, surface chemicals, etc. The rock cutting samples were then ground and filtered to leave microorganisms on the filter membrane surface, making it easier to extract microbial DNA. The filtered samples were collected in centrifuge tubes, and the mixture was centrifuged. The resulting bacterial precipitate was then refrigerated until extraction.
[0059] The DNA extraction and sequencing methods are as follows:
[0060] A magnetic bead milling method was used to extract DNA by breaking down samples through high-speed beating, thereby disrupting microorganisms and individual cells and separating nucleic acids, proteins, and other cellular substances (Gautam 2022). PCR amplification was then performed to replicate DNA in low-biomass samples, significantly increasing the amount of trace DNA (Wages 2005). In this study, the V4 region was selected as the variable region, and fragments were amplified from total genomic DNA in rock debris samples using universal 16S rRNA targeting primers 27f (5'-AGAGTTTGATCATGGCTCAG-3') and 1492R (5'-TAGGGTTACCTTGTTACGACTT-3'). After PCR amplification, agarose gel electrophoresis was performed. The samples were thawed on ice, thoroughly mixed, centrifuged, and a suitable amount was taken for analysis.
[0061] Furthermore, the method for screening bacterial strains is as follows:
[0062] In deep subsurface environments, microbial diversity is influenced by stratigraphic depth and exhibits both vertical and horizontal variations. To investigate these variations, we performed horizontal species visualization statistical analysis of rock debris microbial DNA sequences to create a response chart of microorganisms with stratigraphic depth, showcasing the relative abundance of microbial species in each stratum.
[0063] The tested bacterial strains were initially screened to determine whether they originated from underground oil reservoirs rather than from contamination during the sampling and extraction process.
[0064] The specific screening method involves a detailed analysis of the characteristics of the underground reservoir environment, including factors such as temperature, pressure, and pH. Underground reservoirs typically exhibit anaerobic, high-temperature, high-pressure, and high-mineralization characteristics, which significantly influence the survival and distribution of underground microorganisms.
[0065] At the same time, the DNA sequence of the obtained strain can be compared with the microbial sequences in known underground environments to determine whether the obtained strain has characteristics of underground origin.
[0066] Based on preliminary assessment, aerobic bacteria typically originate above ground, while anaerobic bacteria typically originate underground. Therefore, the selected strains were limited to anaerobic and facultative anaerobic bacteria with a suitable growth temperature >20°C, as previous studies have demonstrated that these characteristics apply to underground bacteria. The screened strains are considered to originate underground and will proceed to downstream analysis.
[0067] Preliminary screening of the tested bacterial strains:
[0068] 1. Conduct characteristic analysis of the underground reservoir environment of the target block to construct parameters for preliminary screening of microbial strains;
[0069] 2. Compare the DNA sequence of the tested bacterial strain with the known microbial DNA sequences of the target block.
[0070] Furthermore, the method for the first clustering is as follows:
[0071] Species statistical analysis at the genus level was performed on the preserved bacterial DNA sequences to obtain relative abundance maps of bacterial species in each stratum;
[0072] Clustering was performed based on matching distance, standard deviation, and column standardization direction. A ring heatmap was used to display the correlation between different microbial species in each stratum after clustering. Strata with strong correlations were grouped into the same stratum, so as to preliminarily divide the total length of the collected rock debris strata into multiple strata.
[0073] Furthermore, characteristic species in each stratum refer to species that are highly correlated and abundant in that stratum, but not highly correlated in other strata.
[0074] Furthermore, the screening method for characteristic bacterial species in each stratum is as follows:
[0075] Correlation analysis was performed on the various microbial species in each stratum after preliminary division to identify the few microbial species with strong correlation in each stratum;
[0076] The highly correlated bacterial species in each stratum were compared with highly correlated bacterial species in other strata, and the highly correlated bacterial species unique to each stratum were screened from several species.
[0077] The abundance of each bacterial species in each stratum was compared to verify whether the selected bacterial species had a prominent abundance in that stratum. Species with higher abundance were identified as characteristic bacterial species.
[0078] Furthermore, the method for correlation analysis of various microbial species is as follows: a correlation heatmap is constructed based on the DNA sequencing data of microbial species in each stratum, and the correlation of microbial species in each stratum is analyzed using the Pearson correlation analysis method to screen out several microbial species with strong correlation in each stratum.
[0079] The method for comparing the abundance of different species in different strata is as follows: based on the Brectis distance, standard deviation standardization method, and column standardization direction, a bubble chart is constructed to display the abundance of each species in different strata, so as to screen out the species with higher abundance in each stratum.
[0080] Furthermore,
[0081] Based on the abundance of characteristic bacterial species in each stratum, using the abundance of the characteristic bacterial species Preliminary category labels corresponding to the abundance of characteristic bacterial species Constructing a small sample dataset And the small sample dataset is augmented to obtain a large sample dataset;
[0082] Since small sample datasets have limited data volume, data augmentation is employed to increase the data size. This can be achieved by using generative adversarial networks to learn and generate different sample data based on the small sample dataset, or by introducing data from other similar scenarios to accumulate data. Bootstrap sampling can also be combined to generate diverse datasets. This approach avoids the overfitting defects that can easily occur due to insufficient sample size and improves model generalization.
[0083] A large dataset was divided into training and test sets to train a Random Forest (RF) algorithm. Grid Search and Cross-Validation were used to optimize the hyperparameters of the Random Forest, resulting in a set for optimizing the algorithm based on the abundance of feature species. Obtain stratigraphic category labels The classifier model, where the loss function of the classifier model is... In the formula Cross-entropy;
[0084] Each stratum is divided into stratigraphic points with 1-meter intervals, and the abundance of characteristic bacterial species at each stratigraphic point is input into the classifier model to predict the category label of each stratigraphic point. and the predicted probability value corresponding to the category label. ;
[0085] Category labels are the same and the difference in predicted probabilities is less than the threshold corresponding to that category label. Adjacent, equally divided stratigraphic points are merged to form smaller layers, thus achieving fine-grained stratigraphic division. Based on predicted probabilities, an adaptive threshold (based on the probability statistics of each category) is used to merge adjacent points of the same category, forming refined smaller layer divisions.
[0086] The method for setting the threshold includes the following steps:
[0087] Calculate the mean of the predicted probabilities among all adjacent stratigraphic pairs in each category label of the preliminary stratigraphic division. and standard deviation ;
[0088] The threshold corresponding to each category label is obtained by combining the mean and standard deviation of the predicted probabilities. In the formula, k is an adjustable parameter.
[0089] Based on the abundance of characteristic bacterial species in each stratum, using the abundance of the characteristic bacterial species Preliminary category labels corresponding to the abundance of characteristic bacterial species Constructing a small sample dataset And the small sample dataset is augmented to obtain a large sample dataset;
[0090] Since small sample datasets have limited data volume, data augmentation is employed to increase the data size. This can be achieved by using generative adversarial networks to learn and generate different sample data based on the small sample dataset, or by introducing data from other similar scenarios to accumulate data. Bootstrap sampling can also be combined to generate diverse datasets. This approach avoids the overfitting defects that can easily occur due to insufficient sample size and improves model generalization.
[0091] A large dataset was divided into training and test sets to train a Random Forest (RF) algorithm. Grid Search and Cross-Validation were used to optimize the hyperparameters of the Random Forest, resulting in a set for optimizing the algorithm based on the abundance of feature species. Obtain stratigraphic category labels The classifier model, where the loss function of the classifier model is... In the formula Cross-entropy;
[0092] Each stratum is divided into stratigraphic points with 1-meter intervals, and the abundance of characteristic bacterial species at each stratigraphic point is input into the classifier model to predict the category label of each stratigraphic point. and the predicted probability value corresponding to the category label. ;
[0093] Category labels are the same and the difference in predicted probabilities is less than the threshold corresponding to that category label. Adjacent, equally divided stratigraphic points are merged to form smaller layers, thus achieving fine-grained stratigraphic division. Based on predicted probabilities, an adaptive threshold (based on the probability statistics of each category) is used to merge adjacent points of the same category, forming refined smaller layer divisions.
[0094] The method for setting the threshold includes the following steps:
[0095] Calculate the mean of the predicted probabilities among all adjacent stratigraphic pairs in each category label of the preliminary stratigraphic division. and standard deviation ;
[0096] The threshold corresponding to each category label is obtained by combining the mean and standard deviation of the predicted probabilities. In the formula, k is an adjustable parameter, which can be set to 2. It can also be set according to actual needs and adjusted in combination with the degree of variation of actual data: if the strata are very homogeneous, k should be set smaller to make the division more refined; if the variation is large, k should be set larger to avoid over-segmentation.
[0097] Unlike traditional hard classification that directly outputs category labels, this invention further utilizes predicted probabilities to merge stratigraphic points, embodying the idea of soft classification. Threshold adaptation employs the calculation of the mean and standard deviation of the predicted probabilities of adjacent points within each category, reflecting the degree of internal consistency of that stratigraphic category. Furthermore, the threshold formula... This means that if the probability fluctuation within a category is small (small standard deviation), the threshold is strict, the merging conditions are stringent, and it is suitable for homogeneous strata; if the fluctuation is large (large standard deviation), the threshold is relatively lenient, allowing for some internal variation and avoiding excessive splitting. Moreover, the parameter k can be used to adjust the sensitivity of merging to adapt to the needs of different geological scenarios.
[0098] Furthermore, only when adjacent points have the same category and their probability difference is less than [a certain value]... Merging is done only when necessary to ensure high similarity in microbial characteristics within the partitioned units, while allowing for gradual transitions.
[0099] This invention employs a probability-driven rather than hard-classification merging strategy, utilizing adaptive thresholds to consider the internal variations of different strata, thereby achieving sub-meter-level high-resolution partitioning.
[0100] This invention utilizes the aforementioned machine learning method for fine division, specifically: the initially divided strata are further divided into small segments with 1-meter intervals;
[0101] Based on the abundance of bacterial species in each segment, the similarity measure is calculated using Euclidean distance. The selected characteristic bacterial species are used as the main clustering features to perform principal coordinate analysis on the similarity matrix, and the coordinates of each bacterial species on the principal coordinate axis are obtained. The principal coordinate analysis can reduce the dimensionality of the high-dimensional similarity matrix to a low-dimensional space. By observing the distribution of bacterial species in the low-dimensional space, the similarity of bacterial species in different segments can be judged.
[0102] The closer the microbial species are, the stronger their correlation. Merging adjacent 1-meter segments with strong correlation forms a small layer, thus achieving fine division of the strata.
[0103] Furthermore, the method to verify whether the preliminary division is accurate is as follows:
[0104] Preliminary formation division is achieved by observing abnormal fluctuations in natural gamma radiation, sealing resistivity, and sonic transit time parameters in well logging curves.
[0105] Compare the preliminary stratigraphic division results with the abnormal fluctuations in the well logging curves. If they are consistent, it indicates that the preliminary stratigraphic division results are accurate.
[0106] Furthermore, the method to verify whether fine segmentation improves the interpretation accuracy of geological reservoirs is as follows:
[0107] Compare the resolution of fine stratigraphic division with that of well logging methods;
[0108] If the resolution of fine stratigraphic division is higher than that of well logging curves, it indicates improved accuracy in interpreting geological reservoirs.
[0109] The method described above will be further illustrated by a specific embodiment below:
[0110] Horizontal species visualization statistical analysis was performed on the DNA sequences of microorganisms in shale debris at depths of 3670–3693 meters to establish a response chart of microorganisms with formation depth, showing the relative abundance of microbial species in each formation.
[0111] In the formation microbial response profile, each colored bar represents a specific microbial species, and its length indicates the abundance of that species.
[0112] This feature based on microbial response to changes in reservoir properties can supplement or replace other rock-based stratigraphic methods to provide more comprehensive geological stratification data and reduce model uncertainty.
[0113] The relative abundance maps of species in different strata show that there are significant differences in bacterial genera among the strata, and the resolution is relatively high. Therefore, the method of constructing a liquid production profile based on microbial DNA sequencing diagnosis has a considerable resolution.
[0114] The strains were screened according to the following principles:
[0115] (1) Anaerobic and facultative anaerobic bacteria;
[0116] (2) Suitable growth temperature > 20℃;
[0117] (3) It has been proven that underground fungi are not subject to the above conditions.
[0118] After screening, ten species that met the above screening criteria were retained as rock debris species derived from underground for further downstream analysis: Methanosarcina, Enterococcus, Bifidobacterium, Methanobacterium, Streptococcus, Lactobacillus, Clostridium XIVa, Pseudomonas, Clostridium IV, and Methanobrevibacter.
[0119] By performing genus-level species statistical analysis on the DNA sequences of rock debris microorganisms, a subsurface stratification method based on microbial species and abundance can be established, thus obtaining a relative abundance map of species in each stratum.
[0120] The relative abundance maps of species in each stratum show the bacterial species and their abundance at different depths, with different species represented by different colors.
[0121] Based on correlation analysis, matching distance, standardization using the Standard Scaler, and standardization direction of the columns, a circular heatmap is obtained.
[0122] The clustering results from the ring heat map can roughly divide the 20-meter strata into three smaller layers.
[0123] The first sub-level is from 3670 to 3679 meters, the second sub-level is from 3680 to 3684 meters, and the third sub-level is from 3685 to 3693 meters.
[0124] The microbial community exhibited high vertical and horizontal specificity, demonstrating the high-resolution characteristics of the stratigraphic characterization method based on the sequencing of endogenous microbial DNA from rock fragments.
[0125] DNA results from rock fragments in each sub-layer were presented at the genus level. Furthermore, comparison of bacterial species in each sub-layer revealed that each sub-stratum unit could be identified with a unique bacterial marker.
[0126] Based on the fact that inter-class distance equals the minimum distance between two classes of objects, Braycurtis distance, the standardization method of the standard scaler, and the standardization direction of the column, a bubble matrix related heatmap for each sub-layer can be constructed.
[0127] In this embodiment, the representative bacterial species of the first sublayer is Streptococcus, the representative bacterial species of the second sublayer is Clostridium IV, and the representative bacterial species of the third sublayer is Methanobrevibacter.
[0128] The fine stratigraphic division results generated by the machine learning model show that 3670-3672 is one layer, 3673-3674 is one layer, 3675-3679 is one layer, 3680-3684 is one layer, 3685-3687 is one layer, and 3688-3693 is one layer. This shows that high-resolution fine stratigraphic division can be achieved based on the sequencing results of endogenous microbial DNA from rock debris, with a resolution of up to 1-2m.
[0129] Based on the stratigraphic division results, all stratigraphic data are divided into four categories: 3670-3674 is represented by pink dots, 3675-3679 by orange dots, 3680-3684 by purple dots, and 3685-3693 by green dots.
[0130] The fine stratigraphic division results are as follows: 3670-3672 meters, 3673-3674 meters, 3675-3679 meters, 3680-3684 meters, 3685-3687 meters, and 3688-3693 meters. At depths of 3673 meters, 3680 meters, and 3685 meters, the natural gamma radiation (GR), containment resistivity (CBL), and sonic transit time (SATT) parameters of the well logging curves all exhibit abnormal fluctuations. Therefore, based on these abnormal fluctuations, the measured stratigraphic interval of more than 20 meters can be divided into three smaller layers, with a resolution of only 2-5 meters.
[0131] According to the stratigraphic microbial response profile shown in the bar chart, significant changes in microorganisms at the genus level were observed at 3673 meters, 3680 meters, and 3685 meters, thus verifying that the fine classification results provided by this invention are consistent with geological laws.
[0132] Based on the new stratigraphic refinement results obtained above, which are based on shale microbial DNA sequencing, we divided the 3670-3672 meter range into one layer, the 3673-3674 meter range into one layer, the 3675-3679 meter range into one layer, the 3680-3684 meter range into one layer, the 3685-3687 meter range into one layer, and the 3688-3693 meter range into one layer. This resulted in the same stratigraphic range being divided into 5 sub-layers, achieving a resolution of up to 1-2 meters.
[0133] This refined stratigraphic division result achieves the same stratigraphic division results as the well logging curve method, while also enabling even finer division. It demonstrates that the microbial stratigraphic division method can be cross-validated with traditional well logging techniques, thus confirming the reliability and practicality of the microbial stratigraphic division technology. Furthermore, this method achieves higher-precision stratigraphic division, reaching a high-resolution stratigraphic division of 1-2 meters, which is more accurate than the well logging curve stratigraphic division method.
[0134] By applying the above-mentioned characteristic bacterial species screening steps to different stratigraphic division scenarios, this invention can obtain a large amount of data accumulation, namely, the relative abundance information of bacterial species in the horizontal direction of each stratigraphic layer in different scenarios (relative abundance map of bacterial species) and the screening results of characteristic bacterial species in each stratigraphic layer, which are recorded and stored as a dataset.
[0135] On this dataset, machine learning algorithms are used to enable the model to autonomously predict characteristic bacterial species in the stratigraphy, replacing the entire screening process mentioned above and improving efficiency.
[0136] Based on the above embodiments, the detailed structure of a classifier model containing multiple classification heads is described, such as... Figure 1 As shown, specifically: the classifier consists of a feature encoder and three classification heads. The three classification heads are used to output the predicted values of the featured bacteria in the first sub-layer (3670 to 3679 meters), the predicted values of the featured bacteria in the second sub-layer (3680 to 3684 meters), and the predicted values of the featured bacteria in the third sub-layer (3685 to 3693 meters), respectively.
[0137] When setting the loss for training the classifier, the cross-entropy loss between the predicted values and the true values output by the three classifier heads is used as the loss term. This ensures that the three classification heads can output predicted values close to the true values, achieving high-precision prediction. These are normalization coefficients used for normalization; the labels output by the classification header can be used directly here. Perform the calculation.
[0138] Since the screening of characteristic microbial species aims to identify the most specific microbial species for each stratum, facilitating subsequent fine stratigraphic division, the characteristic microbial species to be screened in the first sub-layer (3670 to 3679 meters), the second sub-layer (3680 to 3684 meters), and the third sub-layer (3685 to 3693 meters) should be different. This is further supported by the conclusions drawn in the above examples: the representative microbial species of the first sub-layer is Streptococcus, the representative microbial species of the second sub-layer is Clostridium IV, and the representative microbial species of the third sub-layer is Methanobrevibacter.
[0139] In other words, it is necessary to ensure that the predicted values output by the three classification heads are as different as possible. This is reflected in the loss function as follows: , characterizing when The lower the similarity, the better. The closer it is to 0, the more effective it is in actual calculations. The calculation uses the predicted probability vector output by the classification head. The predicted vector is a probability value, and each component is non-negative, so the cosine similarity is also non-negative, effectively ranging from 0 to 1. Thus, when the similarity is 1, this term is 1; when the similarity is 0, this term is 0. This loss term encourages a reduction in similarity, satisfying the expectation of minimizing the loss, which is the expectation that the predicted values output by the three classification heads are as different as possible. These are hyperparameters and should be set according to the actual situation. This is the normalization coefficient, used for normalization.
[0140] Combining the two loss terms to form the total loss term enables the accurate selection of representative bacterial species with specific stratigraphic characteristics, without the need for cumbersome procedures. Applied to the above embodiment, the predicted representative species for the first sublayer is Streptococcus, for the second sublayer it is Clostridium IV, and for the third sublayer it is Methanobrevibacter, demonstrating the effectiveness of this classifier.
[0141] The above embodiments are merely exemplary embodiments of this application and are not intended to limit this application. The scope of protection of this application is defined by the claims. Those skilled in the art can make various modifications or equivalent substitutions to this application within its substance and scope of protection, and such modifications or equivalent substitutions should also be considered to fall within the scope of protection of this application.
Claims
1. A method for analyzing microbial DNA sequencing, characterized in that, Includes the following steps: S100. After initially dividing the formation according to the set sampling depth interval and collecting drilling cuttings, the microbial information of the drilling cuttings from each formation is extracted using DNA sequencing technology. S200. Based on the preliminary stratigraphic and microbial information, data analysis is performed to obtain data on stratigraphy, stratigraphic microbial species, and stratigraphic microbial species abundance. S300. Using cluster analysis, a screening model for strata, stratum microbial species, and stratum microbial species abundance is established. Characteristic microbial species of each stratum are obtained by screening based on differences in stratum microbial species abundance. Similarity relationships between stratum microbial species are established using the characteristic microbial species of each stratum. In step S300, the cluster analysis includes the following steps: S301. Define the matching distance between two clusters by the maximum distance between element pairs, use the standard deviation standardization method, standardize the column direction, and perform cluster analysis on the microbial information of each stratum based on the relative abundance information of species in the horizontal direction of each stratum, establish the relationship between each stratum and the relative abundance of species, and screen the characteristic species of each stratum. S302. Using characteristic microbial species in each stratum as the main feature, establish the relationship between each stratum and the characteristic microbial species using machine learning algorithms; Based on the screening results of characteristic microbial species in each stratum, a machine learning model for screening characteristic microbial species in strata can be constructed, including the following steps: constructing a dataset using the relative abundance information of microbial species in the horizontal direction of each stratum in the preliminary division of strata and the screening results of characteristic microbial species in each stratum; A machine learning model for screening stratigraphic microbial species is obtained by training a classifier model containing multiple classification heads based on the dataset, wherein the number of classification heads is the same as the number m of category labels initially used to classify the stratigraphy. The loss function of the machine learning model In the formula, For binary cross-entropy, For the cosine loss term, , The classifier model outputs the predicted characteristic bacterial species for the i-th and j-th classifier heads, initially classifying the strata into i-th and j-th category labels. True values of characteristic bacterial species for the i-th category label in the strata were initially determined. Among them, the For hyperparameters, the This is the normalization coefficient.
2. The method according to claim 1, characterized in that, In step S301, the method for establishing the relative abundance relationship between different strata and bacterial species includes the following steps: Pearson correlation analysis was used to analyze the correlation of microbial species in different strata, and a bubble chart was used to visualize the constructed matrix.
3. The method according to claim 2, characterized in that, The method for constructing the bubble chart is as follows: Correlation heatmaps were constructed based on high-throughput sequencing data of microbial communities in samples from different strata. Bubble matrix correlation heatmaps for each stratum depth were constructed based on Brecourtis distance, standard deviation normalization, and column normalization direction. In the bubble matrix heatmap, the bacterial composition under different strata depths is displayed in the form of a bubble diagram, showing the abundance distribution of different bacterial species in samples in each pre-divided stratum, so as to screen characteristic bacterial species in each stratum.
4. The method according to claim 1, characterized in that, Step S200 includes the following steps: Genera-level species visualization statistical analysis was performed on the DNA sequences of shale debris microorganisms at a set sampling depth to establish a relative abundance map of microorganisms in different strata with varying depths.
5. An application of the microbial DNA sequencing analysis method according to any one of claims 1 to 4 in stratigraphic division, characterized in that, Based on the microbial DNA sequencing analysis method, the abundance of characteristic bacterial species in adjacent strata is used to refine the preliminary strata division. The refined segmentation includes the following steps: Based on the abundance of characteristic bacterial species in each stratum, using the abundance of the characteristic bacterial species Preliminary category labels corresponding to the abundance of characteristic bacterial species Constructing a small sample dataset And the small sample dataset is augmented to obtain a large sample dataset; A large dataset was divided into training and test sets to train a Random Forest (RF) algorithm. Grid Search and Cross-Validation were used to optimize the hyperparameters of the Random Forest, resulting in a set for optimizing the algorithm based on the abundance of feature species. Obtain stratigraphic category labels The classifier model, where the loss function of the classifier model is... In the formula Cross-entropy; Each stratum is divided into stratigraphic points with 1-meter intervals, and the abundance of characteristic bacterial species at each stratigraphic point is input into the classifier model to predict the category label of each stratigraphic point. and the predicted probability value corresponding to the category label. ; Category labels are the same and the difference in predicted probabilities is less than the threshold corresponding to that category label. Adjacent equidistant stratigraphic points are merged to form smaller layers, thereby achieving a fine division of the strata.
6. The application according to claim 5, characterized in that, The method for setting the threshold includes the following steps: Calculate the mean of the predicted probabilities among all adjacent stratigraphic pairs in each category label of the preliminary stratigraphic division. and standard deviation ; The threshold corresponding to each category label is obtained by combining the mean and standard deviation of the predicted probabilities. In the formula, k is an adjustable parameter.
7. The application according to claim 5, characterized in that, The resolution of the refined division is 1-2m.
Citation Information
Patent Citations
Oil well underground layering method based on geological microflora characteristic analysis
CN114317791A
Fluid production profile dynamic interpretation method based on microorganism DNA sequencing diagnosis
CN119418769A