Marker method based on multi-dimensional intestinal flora characteristics and application thereof

By employing a multidimensional gut microbiota feature labeling method and machine learning algorithms, the problem of accurately identifying differentially expressed bacterial genera in existing technologies has been solved, enabling efficient gut microbiota prediction and health intervention.

CN115472227BActive Publication Date: 2026-05-12AIAGE LIFE SCI CORP LTD
View PDF 2 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
AIAGE LIFE SCI CORP LTD
Filing Date
2022-08-31
Publication Date
2026-05-12

AI Technical Summary

Technical Problem

Current technologies cannot accurately identify differentially expressed bacterial genera in the gut microbiota, resulting in low prediction efficiency.

Method used

The study employed a multidimensional gut microbiota feature labeling method, including converting absolute abundance to relative abundance, calculating the sum of occurrence frequency and abundance, screening for differentially expressed bacterial genera, and constructing a classifier model using machine learning algorithms.

Benefits of technology

It improves prediction efficiency and accuracy, enables rapid establishment of classifier models, accurately identifies differentially expressed bacterial genera, and enhances the accuracy of health intervention selection.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN115472227B_ABST
    Figure CN115472227B_ABST
Patent Text Reader

Abstract

The application relates to a marking method based on multi-dimensional intestinal flora characteristics and application thereof, and belongs to the cross technical field of microbiomics and artificial intelligence. First, the sum of a first occurrence frequency and a relative abundance of a first genus is calculated, all the first genera are screened, and second genera are obtained. Then, the average relative abundance of the second genera is calculated, all the second genera are screened, and third genera are obtained. Then, the average relative abundance difference coefficient of the third genera is calculated, all the third genera are screened, and fourth genera are obtained. Finally, the second occurrence frequency and the third occurrence frequency of the fourth genera are calculated, all the fourth genera are screened, and difference genera are obtained, so that the intestinal flora characteristics are marked, the difference genera can be accurately determined through step-by-step screening, and the prediction efficiency is improved. In addition, through the sample set constructed by the selected difference genera, a classifier model can be quickly established, the classifier model is evaluated, and the prediction accuracy can be greatly improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the interdisciplinary field of microbiome and artificial intelligence, and in particular to a labeling and screening method based on multidimensional gut microbiota characteristics and its application. Background Technology

[0002] Recent epidemiological, pathological, omics, cellular, and animal studies reveal that the gut microbiota significantly mediates metabolic health. Gut flora influences host metabolic homeostasis, and dysbiosis can lead to various common metabolic diseases, including obesity, type 2 diabetes, non-alcoholic liver disease, metabolic heart disease, and malnutrition. Gut microbiology holds promise for developing non-invasive fecal-based tests, dynamic monitoring, and health prediction. By monitoring significant changes in gut microbiota abundance, individuals can gain a deeper understanding of their health status and choose appropriate health interventions. However, current predictions rely on monitoring the abundance of all gut microbiota genera, failing to accurately identify differentially expressed genera, resulting in low predictive efficiency.

[0003] Therefore, there is an urgent need for a labeling method for multidimensional gut microbiota characteristics based on microbial abundance and frequency, and its application. Summary of the Invention

[0004] The purpose of this invention is to provide a labeling method based on multidimensional gut microbiota characteristics and its application, which can accurately identify differentially expressed bacterial genera, thereby improving prediction efficiency.

[0005] To achieve the above objectives, the present invention provides the following solution:

[0006] A labeling method based on multidimensional gut microbiota characteristics, the labeling method comprising:

[0007] Obtain the absolute abundance of each first genus in the gut microbiota of each of the multiple samples; the samples include healthy samples and disease samples;

[0008] For each of the first bacterial genus, the absolute abundance of the first bacterial genus is converted into relative abundance, and the first occurrence frequency of the first bacterial genus in all the samples and the sum of the relative abundance of the first bacterial genus in all the samples are calculated based on the relative abundance; all the first bacterial genus are screened based on the first occurrence frequency and the sum of the relative abundance to obtain the second bacterial genus;

[0009] For each second genus, the average relative abundance of the second genus in all the disease samples is calculated based on the relative abundance of the second genus; all second genera are screened based on the average relative abundance to obtain a third genus;

[0010] For each of the third genera, the average relative abundance difference coefficient of the third genera is calculated based on the average relative abundance of the third genera; all the third genera are screened based on the average relative abundance difference coefficient to obtain the fourth genera;

[0011] For each of the fourth bacterial genus, a second occurrence frequency of the fourth bacterial genus in all the disease samples and a third occurrence frequency of the fourth bacterial genus in all the healthy samples are calculated based on the relative abundance of the fourth bacterial genus; all the fourth bacterial genus are screened based on the average relative abundance difference coefficient, the second occurrence frequency and the third occurrence frequency to obtain differential bacterial genus; the differential bacterial genus is the labeling result of the intestinal flora characteristics.

[0012] A method for classifying and evaluating gut microbiota features, the method comprising:

[0013] Differential bacterial genera are obtained using the labeling method described above, and a sample set is constructed using the relative abundance of differential bacterial genera in each of the multiple samples as sample data.

[0014] The sample set is divided into a training set and a test set; using the training set as input, various machine learning algorithms are used to model the data, resulting in a classifier model for each machine learning algorithm; the machine learning algorithms include random forest algorithm, linear regression algorithm, K-nearest neighbor algorithm, and decision tree algorithm;

[0015] Using the test set as input, the classifier model corresponding to each machine learning algorithm is evaluated according to evaluation metrics, including accuracy, recall, F1-score, and ROC curve.

[0016] A classifier model for gut microbiota features, the classifier model being constructed based on a classifier modeling and evaluation method for gut microbiota features:

[0017] Differential bacterial genera are obtained using the labeling method described above, and a sample set is constructed using the relative abundance of differential bacterial genera in each of the multiple samples as sample data.

[0018] The sample set is divided into a training set and a test set; using the training set as input, various machine learning algorithms are used to model the data, resulting in a classifier model for each machine learning algorithm; the machine learning algorithms include random forest algorithm, linear regression algorithm, K-nearest neighbor algorithm, and decision tree algorithm;

[0019] Using the test set as input, the classifier model corresponding to each machine learning algorithm is evaluated according to evaluation metrics, including accuracy, recall, F1-score, and ROC curve.

[0020] A terminal device includes a processor and a computer-readable storage medium for storing a plurality of instructions, the processor for implementing each of the instructions, the instructions being adapted to be loaded by the processor and perform the following processing:

[0021] Differential bacterial genera are obtained using the labeling method described above, and a sample set is constructed using the relative abundance of differential bacterial genera in each of the multiple samples as sample data.

[0022] The sample set is divided into a training set and a test set; using the training set as input, various machine learning algorithms are used to model the data, resulting in a classifier model for each machine learning algorithm; the machine learning algorithms include random forest algorithm, linear regression algorithm, K-nearest neighbor algorithm, and decision tree algorithm;

[0023] Using the test set as input, the classifier model corresponding to each machine learning algorithm is evaluated according to evaluation metrics, including accuracy, recall, F1-score, and ROC curve.

[0024] According to specific embodiments provided by the present invention, the present invention discloses the following technical effects:

[0025] This invention provides a labeling method based on multidimensional gut microbiota characteristics and its application. First, the absolute abundance of a first genus is converted to relative abundance. Then, the sum of the first occurrence frequency and relative abundance of the first genus is calculated to screen all first genera, resulting in a second genus. Next, the average relative abundance of the second genus is calculated to screen all second genera, resulting in a third genus. Then, the average relative abundance difference coefficient of the third genus is calculated to screen all third genera, resulting in a fourth genus. Finally, the second and third occurrence frequencies of the fourth genus are calculated to screen all fourth genera, resulting in differentially expressed genera. This completes the labeling of multidimensional gut microbiota characteristics. Through stepwise screening, differentially expressed genera can be accurately identified, thereby improving prediction efficiency. Furthermore, by constructing a sample set using the selected differentially expressed genera, a classifier model can be quickly established. Evaluating the classifier model can significantly improve prediction accuracy. Attached Figure Description

[0026] To more clearly illustrate the technical solutions in the embodiments of the present invention or the prior art, the drawings used in the embodiments will be briefly introduced below. Obviously, the drawings described below are only some embodiments of the present invention. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.

[0027] Figure 1This is a flowchart of the marking method provided in Embodiment 1 of the present invention;

[0028] Figure 2 This is a flowchart of the classifier modeling and evaluation method provided in Embodiment 2 of the present invention;

[0029] Figure 3 This is a schematic diagram of the AUC curve provided in Embodiment 2 of the present invention;

[0030] Figure 4 This is a schematic diagram of the classification data summary provided in Embodiment 2 of the present invention;

[0031] Figure 5 This is a schematic diagram illustrating the importance of features provided in Embodiment 2 of the present invention. Detailed Implementation

[0032] The technical solutions of the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of the present invention, and not all embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of the present invention.

[0033] The purpose of this invention is to provide a labeling method based on multidimensional gut microbiota characteristics and its application, which can accurately identify differentially expressed bacterial genera, thereby improving prediction efficiency.

[0034] To make the above-mentioned objects, features and advantages of the present invention more apparent and understandable, the present invention will be further described in detail below with reference to the accompanying drawings and specific embodiments.

[0035] Example 1:

[0036] This embodiment provides a labeling method based on multidimensional gut microbiota characteristics, such as... Figure 1 As shown, the marking method includes:

[0037] S1: Obtain the absolute abundance of each first genus in the gut microbiota of each of the multiple samples; the samples include healthy samples and disease samples;

[0038] In this embodiment, disease samples refer to population samples diagnosed by professional physicians as suffering from a specific disease X, while healthy samples refer to population samples diagnosed by professional physicians as not suffering from the specific disease X and with relatively normal indicators. Population samples generally refer to fecal samples from this population. Specific disease X refers to diseases caused by intestinal flora imbalance, including constipation, diarrhea, recurrent colds, obesity, type 2 diabetes, non-alcoholic liver disease, metabolic heart disease, and malnutrition. The labeling method in this embodiment can be used to accurately identify the differential bacterial genus for any of the above-mentioned diseases, facilitating subsequent prediction. Identifying the differential bacterial genus improves prediction efficiency.

[0039] For each sample, the method for obtaining the absolute abundance of each first genus of bacteria in its gut microbiota is as follows: bacterial DNA is extracted from the sample (i.e., human fecal sample) using 16S rRNA sequencing and then sequenced using Illumina Miseq to obtain sequencing data. The sequencing data is then analyzed using QIIME (2020.2) software to obtain the absolute abundance of each first genus of bacteria in the gut microbiota of that sample.

[0040] It should be noted that the first bacterial genus is the same for each sample. In this case, the data for each sample is equivalent to a row vector in a matrix, and the data composition is: ID + Age + Disease Type + Antibiotic Type + OTU1 + OTU2 + ... + OTU Len OTU is the absolute abundance of a certain bacterial genus in a sample, and each sample has Len OTU values.

[0041] Characteristic difference analysis was performed based on the absolute abundance of each first genus in each sample obtained from S1 to identify differentially expressed genera.

[0042] S2: For each of the first bacterial genus, convert the absolute abundance of the first bacterial genus into relative abundance, calculate the first occurrence frequency of the first bacterial genus in all the samples and the sum of the relative abundance of the first bacterial genus in all the samples based on the relative abundance; screen all the first bacterial genus based on the first occurrence frequency and the sum of the relative abundance to obtain the second bacterial genus;

[0043] In S2, converting the absolute abundance of the first genus to relative abundance may include:

[0044] (1) For each sample, calculate the ratio of the absolute abundance of the first genus to the sum of the absolute abundances of all first genus species in the sample to obtain the median abundance of the first genus.

[0045] For each sample, the formula for calculating the median abundance of each first genus in that sample is:

[0046]

[0047] where \(i\) is the sample serial number and \(j\) is the genus serial number; \(p\) ij is the intermediate abundance of the \(j\)th first genus in the \(i\)th sample; \(n\) ij is the absolute abundance of the \(j\)th first genus in the \(i\)th sample; \(t\) is the total number of first genera.

[0048] (2) Determine whether the intermediate abundance is less than the first preset threshold; if so, set the intermediate abundance to 0, otherwise, keep the intermediate abundance unchanged to obtain the adjusted abundance of the first genus;

[0049] For the intermediate abundance of each first genus, if \(p\) ij \(<Q\), set the adjusted abundance \(p'\) ij \( = 0\), otherwise, \(p'\) ij \( = p\) ij to filter out low-abundance genera in each sample and set the intermediate abundance of the first genus in the sample with an intermediate abundance lower than the first preset threshold \(Q\) (\(Q\) can be \(2\times10\) -5 ) to 0.

[0050] (3) Calculate the ratio of the adjusted abundance of the first genus to the sum of the adjusted abundances of all first genera in the sample to obtain the relative abundance of the first genus.

[0051] The formula for calculating the relative abundance is:

[0052]

[0053] where \(P\) ij is the relative abundance of the \(j\)th first genus in the \(i\)th sample; \(p'\) ij is the adjusted abundance of the \(j\)th first genus in the \(i\)th sample.

[0054] Using the above three steps, the absolute abundance of each first genus in each sample can be converted into a relative abundance.

[0055] In S2, calculating the first occurrence frequency of the first genus in all samples and the sum of the relative abundances of the first genus in all samples can include:

[0056] (1) Calculate the first occurrence frequency:

[0057]

[0058] where \(f\) j is the first occurrence frequency of the \(j\)th first genus, that is, the proportion of samples with a relative abundance of the \(j\)th first genus greater than 0 among all samples; \(n\) j is the number of samples with a relative abundance of the \(j\)th first genus greater than 0; \(N\) Ais the total number of disease samples; N B is the total number of healthy samples.

[0059] (2) Calculate the sum of relative abundances:

[0060]

[0061] where S j is the sum of the relative abundances of the j-th first genus in all samples.

[0062] Using the above formula, the first occurrence frequency and the sum of relative abundances of each first genus can be calculated. Subsequently, these two parameters are used to screen all first genera.

[0063] In S2, all first genera are screened according to the first occurrence frequency and the sum of relative abundances to obtain the second genus, which may include: removing the first genera with a first occurrence frequency less than the second preset threshold and a sum of relative abundances less than the third preset threshold from all first genera to obtain the second genus.

[0064] Specifically, if f j < F and S j < R, then the j-th first genus is filtered; otherwise, it is not filtered. All unfiltered first genera are the second genus. Here, F is the second preset threshold for the frequency of screening qualified condition genera, which can be set to 0.1, 0.15, etc.; R is the third preset threshold for the abundance of screening qualified condition genera, which can be set to 0.001, 0.0001, etc.

[0065] S3: For each of the second genera, calculate the average relative abundance of the second genus in all the disease samples according to the relative abundance of the second genus; screen all the second genera according to the average relative abundance to obtain the third genus;

[0066] In S3, the formula for calculating the average relative abundance of the second genus in all disease samples according to the relative abundance of the second genus is as follows:

[0067]

[0068] where M Aj is the average relative abundance of the j-th second genus; P ij is the relative abundance of the j-th second genus in the i-th disease sample.

[0069] Using the above formula, the average relative abundance of all second genera in disease samples can be calculated.

[0070] In S3, screening all second genera based on average relative abundance to obtain third genera may include selecting second genera with average relative abundance greater than a fourth preset threshold or average relative abundance less than a fifth preset threshold as third genera.

[0071] Specifically, disease samples with an average relative abundance greater than a fourth preset threshold P are selected. j_high Or less than the fifth preset threshold P j_low The second genus of bacteria that meets the aforementioned conditions is the third genus of bacteria. The set of genera H consisting of the third genus of bacteria can be represented as: M Aj >P j_high Or M Aj <P j_low .

[0072] S4: For each of the third genera, calculate the average relative abundance difference coefficient of the third genera based on the average relative abundance of the third genera; screen all the third genera based on the average relative abundance difference coefficient to obtain the fourth genera;

[0073] In S4, calculating the average relative abundance difference coefficient of the third genus based on its average relative abundance may include: calculating the average relative abundance difference coefficient of the third genus based on its average relative abundance using the difference coefficient calculation formula.

[0074] In this embodiment, the formula for calculating the difference coefficient is:

[0075]

[0076] Where, p Acontrol M is the average relative abundance difference coefficient for the j-th third genus; Aj p represents the average relative abundance of the j-th third genus; j_mean The average relative abundance of the j-th third genus in healthy individuals.

[0077] Using the above formula, the average relative abundance difference coefficient of each third genus in the genus set H can be determined.

[0078] It should be noted that in this embodiment, the average relative abundance p of healthy individuals is... j_mean This is a pre-obtained value, obtained as follows:

[0079] (1) Collection and screening of healthy population samples

[0080] Healthy individuals were sampled through self-collection or downloading from public online databases, with samples selected from individuals within a suitable age range (e.g., 20-70 years old). After analyzing the bacterial genus data of the healthy individuals' samples, samples with too few genera (e.g., less than 10 species) were filtered out to eliminate interference from abnormal samples. In this embodiment, the healthy individuals' samples refer to healthy control group samples with multiple diseases, and a reference interval for the microbial community of healthy individuals was established based on a large number of healthy individuals' samples.

[0081] (2) Calculation of reference intervals for fungal genera

[0082] Suppose the absolute abundance of microbial genera in the healthy population sample is known, i is the sample number, j is the genera number, and the total number of samples is N. control .

[0083] 1) For each healthy population sample, calculate the relative abundance of each bacterial genus in that sample:

[0084]

[0085] Where, p ij_control n represents the relative abundance of genus j in sample i of healthy individuals; ij_control t represents the absolute abundance of genus j in sample i of healthy individuals; t represents the total number of genus j.

[0086] 2) Calculate the reference intervals for bacterial genera in all healthy population samples:

[0087] The relative abundance p of genus j in all healthy population samples 1j_control p 2j_control ...sort in ascending order, and take a specific percentile interval as the genus reference interval for the healthy population sample [P] j_low P j_high For example, the 20th percentile is used as the lower limit of the reference interval for bacterial genera in healthy individuals, and the 80th percentile is used as the upper limit of the reference interval for bacterial genera in healthy individuals. These upper and lower limits are the fourth and fifth preset thresholds in S3, respectively. The anchor point is calculated using the reference interval for bacterial genera in healthy individuals, and the differential bacterial genera with the most obvious differential characteristics are selected through calculation.

[0088] (3) For each genus j, calculate the frequency of its occurrence in all healthy population samples:

[0089]

[0090] Among them, f j_control N represents the proportion of healthy individuals whose relative abundance of genus j is greater than 0 among all healthy individuals; that is, the frequency of occurrence of genus j. j This represents the number of healthy individuals whose relative abundance of genus j is greater than 0.

[0091] (4) For each genus j, calculate the average relative abundance:

[0092]

[0093] Where, p j_mean This refers to the average relative abundance of the jth bacterial genus in healthy individuals.

[0094] In S4, all third genera are screened based on the average relative abundance difference coefficient to obtain the fourth genera. Specifically, the third genera whose absolute value of the average relative abundance difference coefficient is greater than the sixth preset threshold are selected as the fourth genera.

[0095] Specifically, select abs(p) Acontrol The third genus of bacteria at the sixth preset threshold θp is taken as the fourth genus, forming the genus set H1. Here, abs represents the absolute value.

[0096] S5: For each of the fourth bacterial genera, calculate the second occurrence frequency of the fourth bacterial genera in all the disease samples and the third occurrence frequency of the fourth bacterial genera in all the healthy samples based on the relative abundance of the fourth bacterial genera; screen all the fourth bacterial genera based on the average relative abundance difference coefficient, the second occurrence frequency and the third occurrence frequency to obtain differential bacterial genera; the differential bacterial genera are the labeling results of the intestinal flora characteristics.

[0097] In S5, calculating the second occurrence frequency of the fourth genus in all disease samples and the third occurrence frequency of the fourth genus in all healthy samples based on the relative abundance of the fourth genus may include:

[0098] (1) Calculate the second frequency of occurrence:

[0099] For the bacterial genus set H1, the formula for calculating the frequency of the j-th fourth bacterial genus in the disease sample, i.e., the second frequency of occurrence, is as follows:

[0100]

[0101] Among them, f Aj The second occurrence frequency of the j-th fourth genus of bacteria; n Aj denoted as the number of disease samples in which the relative abundance of the j-th fourth genus of bacteria is greater than 0.

[0102] (2) Calculate the frequency of the third occurrence:

[0103] For a set of bacterial genera H1, the formula for calculating the frequency of the j-th fourth bacterial genus in healthy samples, i.e., the third frequency, is as follows:

[0104]

[0105] Among them, f Bj The third occurrence frequency of the j-th fourth genus; n Bj The number of healthy samples with a relative abundance of the j-th fourth genus greater than 0.

[0106] In S5, the fourth genera are screened based on the average relative abundance difference coefficient, the second occurrence frequency, and the third occurrence frequency. The differentially occurring genera can include:

[0107] (1) Selecting average relative abundance difference coefficients greater than 0 (i.e., p) Acontrol The fourth genus with a relative abundance difference coefficient >0 was selected as the first dominant genus in the disease sample; genera with an average relative abundance difference coefficient less than 0 (i.e., p) were selected. Acontrol The fourth genus of bacteria with <0 was selected as the second dominant genus of bacteria in healthy samples;

[0108] (2) Select the second frequency of occurrence that is greater than or equal to the third frequency of occurrence (i.e., f). Aj ≥f Bj ), or the difference between the third frequency of occurrence and the second frequency of occurrence is less than the seventh preset threshold (i.e., f). Bj -f Aj <θ f The first dominant genus of bacteria is considered as the differentially expressed genus; θ f This is the seventh preset threshold;

[0109] (3) Select the third occurrence frequency that is greater than or equal to the second occurrence frequency (i.e., f). Bj ≥f Aj ), or the difference between the second and third occurrence frequencies is less than the seventh preset threshold (i.e., f). Aj -f Bj <θ f The second dominant genus of bacteria was identified as the differential genus.

[0110] The differentially observed bacterial genera selected in (2) above form the dominant bacterial genera set H2 of the disease sample, and the differentially observed bacterial genera selected in (3) form the dominant bacterial genera set H3 of the healthy sample. The combination of H2 and H3 yields the bacterial genera set H4, which is the final differentially observed bacterial genera.

[0111] The labeling method provided in this embodiment can accurately identify differentially expressed bacterial genera through a step-by-step screening process of the sample analysis. Subsequent rapid prediction can then be performed using these differentially expressed genera, greatly improving prediction efficiency. Furthermore, the screening process also incorporates the bacterial genera reference range and average relative abundance of healthy individuals to obtain differentially expressed bacterial genera more objectively.

[0112] Example 2:

[0113] This embodiment provides a method for classifying and evaluating gut microbiota features, such as... Figure 2 As shown, the method includes:

[0114] T1: Differential bacterial genera are obtained using the labeling method described in Example 1, and a sample set is constructed using the relative abundance of differential bacterial genera in each of the multiple samples as sample data;

[0115] For each sample in S1 of Example 1, the relative abundance of each differentially expressed bacterial genus in that sample is obtained as the sample data for that sample. The sample data of all samples constitute the sample set.

[0116] T2: Divide the sample set into a training set and a test set; using the training set as input, model the sample set using various machine learning algorithms to obtain a classifier model for each machine learning algorithm; the machine learning algorithms include random forest algorithm, linear regression algorithm, K-nearest neighbor algorithm and decision tree algorithm;

[0117] In this embodiment, the sample set can be divided into a training set and a test set in an 8:2 ratio.

[0118] In this embodiment, the relative abundance of the differentially abundant bacterial genera obtained through screening is used as the sample data. Compared with using the relative abundance of all bacterial genera as the sample data, this can greatly reduce the amount of training and improve the modeling speed of the classifier model.

[0119] T3: Using the test set as input, evaluate the classifier model corresponding to each machine learning algorithm according to the evaluation metrics; the evaluation metrics include accuracy, recall, F1-score and ROC curve.

[0120] After evaluating the classifier model corresponding to each machine learning algorithm according to evaluation metrics, selecting the best-performing classifier model as the target model can greatly improve prediction accuracy. Of course, this embodiment can also optimize various input parameters of the target model.

[0121] The classifier modeling and evaluation method provided in this embodiment can evaluate the classifier models of various machine learning algorithms, effectively improving prediction accuracy and precision.

[0122] Here, this embodiment provides an example to further illustrate the classifier modeling and evaluation method of this embodiment:

[0123] I. Calculation of Reference Intervals for Fungal Genus in Healthy Individuals

[0124] (1) Sample collection and screening

[0125] Healthy individuals were sampled through self-collection or downloading from public online databases, and samples from individuals within a suitable age range (e.g., 20-70 years old) were selected. After analyzing the bacterial genera data of the healthy individuals' samples, samples with too few genera (e.g., less than 10 species) were filtered out to eliminate interference from abnormal samples.

[0126] This example screened a total of 7773 healthy individuals aged 20-70. The age and gender distribution is shown in Table 1 below:

[0127] Table 1

[0128]

[0129]

[0130] (2) Calculation of reference intervals for fungal genera

[0131] Suppose the absolute abundance of microbial genera in the healthy population sample is known, i is the sample number, j is the genera number, and the total number of samples is N. control ;

[0132] 1) For each healthy population sample, calculate the relative abundance of each bacterial genus in that sample:

[0133]

[0134] Where, p ij_control n represents the relative abundance of genus j in sample i of healthy individuals; ij_control t represents the absolute abundance of genus j in sample i of healthy individuals; t represents the total number of genus j.

[0135] 2) Calculate the reference intervals for bacterial genera in all healthy population samples:

[0136] The relative abundance of genus j in all healthy population samples was sorted in ascending order, and a specific percentile interval was taken as the reference interval for genus j in healthy population samples [P]. j_low P j_high For example, the 20th percentile can be used as the lower limit of the reference interval for bacterial genera in healthy individuals, and the 80th percentile can be used as the upper limit of the reference interval for bacterial genera in healthy individuals.

[0137] (3) For each genus j, calculate the frequency of its occurrence in all healthy population samples:

[0138]

[0139] Among them, f j_control N represents the proportion of healthy individuals whose relative abundance of genus j is greater than 0 among all healthy individuals; that is, the frequency of occurrence of genus j. jThis represents the number of healthy individuals whose relative abundance of genus j is greater than 0.

[0140] (4) For each genus j, calculate the average relative abundance:

[0141]

[0142] Where, p j_mean This refers to the average relative abundance of the jth bacterial genus in healthy individuals.

[0143] Based on this, the frequency of occurrence, different quantiles, and average relative abundance of each fungal genus were obtained, and some data are shown in Table 2 below.

[0144] Table 2

[0145]

[0146] II. Feature Difference Analysis

[0147] This example selected constipation for analysis, with a total of 964 samples, including 482 constipation samples and 482 healthy samples.

[0148] (1) Absolute abundance to relative abundance

[0149] 1) For each sample, calculate the median abundance of each first genus in that sample:

[0150]

[0151] Where i is the sample number and j is the genus number; p ij n represents the median abundance of the j-th species of the first genus in the i-th sample; ij t represents the absolute abundance of the j-th species of the first genus in the i-th sample; t represents the total number of species of the first genus.

[0152] The results are shown in Table 3 below:

[0153] Table 3

[0154]

[0155] 2) Filter out low abundance of the first genus, and remove all intermediate abundances in the sample below 2*10. -5 The intermediate abundance of the first genus was changed to 0, and the partial results are shown in Table 4 below:

[0156] Table 4

[0157]

[0158]

[0159] 3) Further calculate the relative abundance

[0160] (2) Filter the first genus of bacteria with low occurrence frequency

[0161] 1) Calculate the first occurrence frequency of the first genus of bacteria:

[0162]

[0163] where f j is the first occurrence frequency of the j-th first genus of bacteria, that is, the proportion of samples with a relative abundance of the j-th first genus of bacteria greater than 0 in all samples; n j is the number of samples with a relative abundance of the j-th first genus of bacteria greater than 0; N A is the total number of constipation samples; N B is the total number of healthy samples.

[0164] 2) Calculate the sum of the relative abundances of each first genus of bacteria:

[0165]

[0166] where S j is the sum of the relative abundances of the j-th first genus of bacteria in all samples.

[0167] 3) Filtering method

[0168] If f j < F and S j < R, then filter the first genus of bacteria j from all samples, otherwise do not filter. Here, F is the second preset threshold for the frequency of screening qualified conditional genera of bacteria, which can be set to 0.1, 0.15, etc.; R is the third preset threshold for the abundance of screening qualified conditional genera of bacteria, which can be set to 0.001, 0.0001, etc.

[0169] The partial results obtained are as shown in Table 5 below:

[0170] Table 5

[0171]

[0172]

[0173] (3) Calculate the differences

[0174] 1) Difference in average relative abundance

[0175] Calculate the average relative abundance of all second genera of bacteria in constipation samples:

[0176]

[0177] where M Aj is the average relative abundance of the j-th second genus of bacteria; P ijLet represent the relative abundance of the j-th second bacterial genus in the i-th constipation sample.

[0178] Among the constipation samples screened, the average relative abundance was greater than P. j_high or less than P j_low The collection of fungi H:M Aj >P j_high Or M Aj <P j_low .

[0179] Calculate the average relative abundance difference coefficient in the fungal genus set H:

[0180]

[0181] Where, p Acontrol M is the average relative abundance difference coefficient for the j-th third genus; Aj p represents the average relative abundance of the j-th third genus; j_mean The average relative abundance of the j-th third genus in healthy individuals.

[0182] Select abs(p) Acontrol The third genus of bacteria in θp is considered as the fourth genus, forming the genus set H1.

[0183] 2) Frequency difference

[0184] For the bacterial genus set H1, calculate the frequency of the fourth bacterial genus j in constipation samples and healthy samples:

[0185]

[0186] Among them, f Aj The second occurrence frequency of the j-th fourth genus of bacteria; n Aj The number of constipation samples with a relative abundance of the j-th fourth genus of bacteria greater than 0.

[0187]

[0188] Among them, f Bj The third occurrence frequency of the j-th fourth genus; n Bj The number of healthy samples with a relative abundance of the j-th fourth genus greater than 0.

[0189] The dominant bacterial genus (p) in constipation samples screened for "relative abundance mean difference" Acontrol >0), if f Aj ≥f Bj or f Bj -f Aj <θ f Then, genus j is denoted as the dominant genus set H2 of the constipation sample;

[0190] For the second dominant bacterial genus (p) in healthy samples selected by screening "relative abundance mean difference" Acontrol <0), if f Bj ≥f Aj or f Aj -f Bj <θ f Then, let genus j be the set of dominant genus in the healthy sample, H3.

[0191] H2 and H3 are merged into the genera set H4, thus obtaining the final differential genera.

[0192] Finally, a total of 25 species of H2 were screened out, in order: 'Aggregatibacter', 'Enterobacter', 'Klebsiella', 'Enterococcus', 'Eggerthella', 'Streptococcus', '[Euba cterium]halliigroup', 'Bifidobacterium', 'Collinsella', 'Adlercreutzia', 'Blautia', 'Lactobacillus', 'Erysipelotrichaceae' UCG-003', 'Turicibacter', 'Anaerostipes', 'Tyzzerella 3', 'Fusicatenibacter', '[Ruminococcus]gnavus group', 'Dorea', '[Ruminococcus]torques group', 'Ruminococcaceae UCG-004', 'Marvinbryantia', 'Agathobacter', 'Ruminococcaceae UCG-013', 'Clostridium sensustricto 1', there are 12 species of H3, in order: 'Ruminococcaceae UCG-005', 'Ruminiclostridium 9', '[Eubacterium]xylanophilum group', 'Alistipes', '[Eubacterium]eligens group', 'Sutterella', 'Barnesiella', 'Escherichia-Shigella', 'Ruminiclostridium 6', 'Bilophila', 'Akkermansia'.

[0193] III. Classifier Model

[0194] In constipation and healthy samples, all bacterial genera in bacterial set H4 were screened out and denoted as sample set H5, as shown in Table 6 below:

[0195] Table 6

[0196]

[0197]

[0198] The sample set was divided into training and test sets in an 8:2 ratio. Conventional machine learning algorithms, such as random forest, linear regression, K-nearest neighbor, or decision tree, were used sequentially for modeling. These models were evaluated based on relevant technical metrics of machine learning algorithms, including accuracy, recall, F1-score, and ROC curve. Based on the evaluation results, the Gradient Boosting Classifier model was selected as the best model, with an accuracy of 92.48%, an AUC of 97.36%, a recall of 92.14%, and an F1-score of 92.38%. This embodiment can also use the `classification.compare_models` method in the Python `pycaret` module to perform 10-fold cross-validation on each model to assist in selecting the best model. Parameter optimization can also be performed on the selected model.

[0199] The evaluation metrics for each model are shown in Table 7 below:

[0200] Table 7

[0201]

[0202]

[0203] Figure 3 This is a schematic diagram of the AUC curve, where AUC = 0.98. Figure 4 This is a schematic diagram of the classification data summary. The precision, recall, and F-score are all greater than 0.9, indicating that the differentially expressed bacterial species screened based on the labeling method of this embodiment can effectively distinguish between the two groups of data. Figure 5 This is a schematic diagram illustrating the importance of features, showing the weight of each fungal genus in the model.

[0204] Example 3:

[0205] This embodiment provides a classifier model for gut microbiota features, which is constructed based on a classifier modeling and evaluation method for gut microbiota features.

[0206] Differential bacterial genera were obtained using the labeling method described in Example 1, and a sample set was constructed using the relative abundance of the differential bacterial genera in each of the multiple samples as sample data.

[0207] The sample set is divided into a training set and a test set; using the training set as input, various machine learning algorithms are used to model the data, resulting in a classifier model for each machine learning algorithm; the machine learning algorithms include random forest algorithm, linear regression algorithm, K-nearest neighbor algorithm, and decision tree algorithm;

[0208] Using the test set as input, the classifier model corresponding to each machine learning algorithm is evaluated according to evaluation metrics, including accuracy, recall, F1-score, and ROC curve.

[0209] Example 4:

[0210] This embodiment provides a terminal device, including a processor and a computer-readable storage medium. The computer-readable storage medium stores a plurality of instructions, and the processor implements each of the instructions. The instructions are adapted to be loaded by the processor and executed for the following processing:

[0211] Differential bacterial genera were obtained using the labeling method described in Example 1, and a sample set was constructed using the relative abundance of the differential bacterial genera in each of the multiple samples as sample data.

[0212] The sample set is divided into a training set and a test set; using the training set as input, various machine learning algorithms are used to model the data, resulting in a classifier model for each machine learning algorithm; the machine learning algorithms include random forest algorithm, linear regression algorithm, K-nearest neighbor algorithm, and decision tree algorithm;

[0213] Using the test set as input, the classifier model corresponding to each machine learning algorithm is evaluated according to evaluation metrics, including accuracy, recall, F1-score, and ROC curve.

[0214] The various embodiments in this specification are described in a progressive manner, with each embodiment focusing on the differences from other embodiments. The same or similar parts between the various embodiments can be referred to each other.

[0215] This document uses specific examples to illustrate the principles and implementation methods of the present invention. The descriptions of the above embodiments are only for the purpose of helping to understand the method and core ideas of the present invention. Furthermore, those skilled in the art will recognize that, based on the ideas of the present invention, there will be changes in the specific implementation methods and application scope. Therefore, the content of this specification should not be construed as a limitation of the present invention.

Claims

1. A labeling method based on multidimensional gut microbiota characteristics, characterized in that, The marking method includes: Obtain the absolute abundance of each first genus in the gut microbiota of each of the multiple samples; the samples include healthy samples and disease samples; For each of the first bacterial genus, the absolute abundance of the first bacterial genus is converted into relative abundance, and the first occurrence frequency of the first bacterial genus in all the samples and the sum of the relative abundance of the first bacterial genus in all the samples are calculated based on the relative abundance; all the first bacterial genus are screened based on the first occurrence frequency and the sum of the relative abundance to obtain the second bacterial genus; For each second genus, the average relative abundance of the second genus in all the disease samples is calculated based on the relative abundance of the second genus; all second genera are screened based on the average relative abundance to obtain a third genus; For each of the third genera, the average relative abundance difference coefficient of the third genera is calculated based on the average relative abundance of the third genera; all the third genera are screened based on the average relative abundance difference coefficient to obtain the fourth genera; For each of the fourth bacterial genus, a second occurrence frequency of the fourth bacterial genus in all the disease samples and a third occurrence frequency of the fourth bacterial genus in all the healthy samples are calculated based on the relative abundance of the fourth bacterial genus; all the fourth bacterial genus are screened based on the average relative abundance difference coefficient, the second occurrence frequency and the third occurrence frequency to obtain differential bacterial genus; the differential bacterial genus is the labeling result of the intestinal flora characteristics; Specifically, the step of calculating the average relative abundance difference coefficient of the third genus based on its average relative abundance includes: calculating the average relative abundance difference coefficient of the third genus using the difference coefficient calculation formula based on its average relative abundance. The formula for calculating the difference coefficient includes: ; in, For the first The average relative abundance difference coefficient of each third genus; For the first Average relative abundance of the third genus; For the first healthy people Average relative abundance of the third genus.

2. The marking method according to claim 1, characterized in that, The conversion of the absolute abundance of the first genus to relative abundance specifically includes: For each sample, the ratio of the absolute abundance of the first genus to the sum of the absolute abundances of all first genus species in the sample is calculated to obtain the median abundance of the first genus. Determine whether the intermediate abundance is less than a first preset threshold; if so, set the intermediate abundance to 0; otherwise, keep the intermediate abundance unchanged to obtain the adjusted abundance of the first genus. The relative abundance of the first genus is obtained by calculating the ratio of the adjusted abundance of the first genus to the sum of the adjusted abundances of all first genus species in the sample.

3. The marking method according to claim 1, characterized in that, The step of screening all the first genera based on the sum of the first occurrence frequency and the relative abundance to obtain the second genera specifically includes: removing the first genera whose first occurrence frequency is less than a second preset threshold and whose relative abundance is less than a third preset threshold from all the first genera to obtain the second genera.

4. The marking method according to claim 1, characterized in that, The step of screening all second genera based on the average relative abundance to obtain a third genera specifically includes: selecting second genera whose average relative abundance is greater than a fourth preset threshold or whose average relative abundance is less than a fifth preset threshold as the third genera.

5. The marking method according to claim 1, characterized in that, The step of screening all the third genera according to the average relative abundance difference coefficient to obtain the fourth genera specifically includes: selecting the third genera whose absolute value of the average relative abundance difference coefficient is greater than a sixth preset threshold as the fourth genera.

6. The marking method according to claim 1, characterized in that, The step of screening all the fourth genera based on the average relative abundance difference coefficient, the second occurrence frequency, and the third occurrence frequency to obtain differentially occurring genera specifically includes: The fourth bacterial genus with an average relative abundance difference coefficient greater than 0 is selected as the first dominant bacterial genus in the disease sample; the fourth bacterial genus with an average relative abundance difference coefficient less than 0 is selected as the second dominant bacterial genus in the healthy sample. The first dominant bacterial genus that has a second occurrence frequency greater than or equal to the third occurrence frequency, or whose difference between the third occurrence frequency and the second occurrence frequency is less than a seventh preset threshold, is selected as the differential bacterial genus. The second dominant bacterial genus, whose third occurrence frequency is greater than or equal to the second occurrence frequency, or whose difference between the second occurrence frequency and the third occurrence frequency is less than the seventh preset threshold, is selected as the differential bacterial genus.

7. A method for modeling and evaluating classifiers based on gut microbiota characteristics, characterized in that, The method includes: Differential bacterial genera are obtained using the labeling method according to any one of claims 1-6, and a sample set is constructed using the relative abundance of differential bacterial genera in each of the multiple samples as sample data; The sample set is divided into a training set and a test set; using the training set as input, various machine learning algorithms are used to model the data, resulting in a classifier model for each machine learning algorithm; the machine learning algorithms include random forest algorithm, linear regression algorithm, K-nearest neighbor algorithm, and decision tree algorithm; Using the test set as input, the classifier model corresponding to each machine learning algorithm is evaluated according to evaluation metrics, including accuracy, recall, F1-score, and ROC curve.

8. A terminal device, comprising a processor and a computer-readable storage medium, the computer-readable storage medium for storing a plurality of instructions, the processor for implementing each of the instructions, characterized in that, The instructions are adapted to be loaded by the processor and executed for the following processing: Differential bacterial genera are obtained using the labeling method according to any one of claims 1-6, and a sample set is constructed using the relative abundance of differential bacterial genera in each of the multiple samples as sample data; The sample set is divided into a training set and a test set; using the training set as input, various machine learning algorithms are used to model the data, resulting in a classifier model for each machine learning algorithm; the machine learning algorithms include random forest algorithm, linear regression algorithm, K-nearest neighbor algorithm, and decision tree algorithm; Using the test set as input, the classifier model corresponding to each machine learning algorithm is evaluated according to evaluation metrics, including accuracy, recall, F1-score, and ROC curve.