Microbial composition based on enterotype typing, method for constructing an allergy prediction model, and model and use thereof
By constructing an allergy prediction model based on intestinal typing and using machine learning to identify microbial markers in fecal samples, the problem of insufficient accuracy in non-invasive detection of childhood allergic diseases in existing technologies has been solved, and a high-accuracy assessment of childhood allergy risk has been achieved.
Patent Information
- Application Number
- CN202310439648.8
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2023-04-23
- Publication Date
- 2026-02-24
- Estimated Expiration
- 2043-04-23
AI Technical Summary
Current technology lacks non-invasive methods for directly diagnosing allergic diseases in children, especially the accuracy of diagnosing allergic diseases in children through specific bacteria in fecal samples needs to be improved.
An allergy prediction model based on intestinal typing was constructed. By screening 16S sequencing data suitable for age groups, the intestinal typing model was trained using machine learning methods to identify 15 microbial biomarkers, including Incertae Sedis, Bombiscardovia, and Eggerthella, for the detection of microbial compositions in fecal samples. The allergy prediction model was then constructed and applied to the kit.
This method enables non-invasive detection of allergy risk in children using fecal samples, with an accuracy rate of 94.62%, providing a new non-invasive method for allergy detection in children.
Smart Images

Figure CN116662798B_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of microbial molecular diagnostics technology, specifically to microbial compositions based on enterotyping, methods for constructing allergy prediction models, and their models and applications. Background Technology
[0002] Allergic diseases are common immune system disorders, characterized by a pathological overreaction of the body's immune system to persistent stimulation by specific antigens or re-stimulation by the same antigen. Common allergic diseases include allergic rhinitis, atopic dermatitis, allergic asthma, and food allergies. Globally, approximately 30% to 40% of the population suffers from allergic diseases, making it a widespread public health issue. Common allergic reactions belong to type I hypersensitivity reactions. The primary mechanism involves the binding of allergens entering the body to IgE antibodies on the surface of mast cells, triggering a series of biochemical reactions in the cell membrane. This leads to mast cell degranulation, releasing allergy mediators such as histamine, leukotrienes, serotonin, prostaglandins (PG), and kinins, thereby causing a series of allergic reaction symptoms.
[0003] Allergies are more common in children. Epidemiological studies have found that in addition to genetic predisposition, lifestyle changes, including increased cesarean sections, antibiotic use, and dietary habits, are significantly contributing to the rise in allergy symptoms. Children often lack sufficient exposure to environmental microorganisms during their growth, resulting in a lack of stimulation from corresponding antigens and a significantly increased probability of developing allergic diseases later in life. These factors directly or indirectly affect the development of gut microbiota, which plays a dominant role in forming immune responses, especially in early life.
[0004] The gut microbiota is an indispensable part of the human body, not only helping the body absorb nutrients from food, but also playing a vital role in functions including metabolism, biological barriers, immune regulation, and host defense. Gut microbiota can indirectly influence an individual's response to immunotherapy. The flora colonizing the intestinal mucosa plays a crucial role in the maturation of the host's immune system, manifested in the maintenance of epithelial cell integrity and the stimulation of immune tolerance by gut microbiota and its metabolites. For example, short-chain fatty acids are metabolites produced by the fermentation and degradation of some dietary fibers by gut microbiota, and these products participate in regulating the body's health and disease development. Butyrate, as an inhibitor of histone deacetylases, promotes the expression of FOXP3, thus enhancing the induced inhibitory function of Treg cells and effectively regulating the immune system. Conversely, the occurrence of allergic diseases and the use of medications can also lead to dysbiosis of the gut microbiota. A study using antihistamines to treat chronic spontaneous urticaria found significant enrichment of Prevotella, Megamonas, and Escherichia in the gut of urticaria patients, while Blautia, Alistipes, Anaerostipes, and Lachnospira were significantly reduced in the gut of antihistamine-resistant patients. The symbiotic and co-evolutionary relationship between the gut microbiota and the human body can promote the development of the host immune system and regulate its balance.
[0005] Current reports on non-invasive methods for diagnosing allergic diseases involve detecting the expression levels of relevant proteins in urine. There are also reports of diagnosing food allergies using gut microbiota. However, these methods are currently only tested on mice and have not been directly validated in humans, let alone in children. Their accuracy needs improvement. There are currently no methods or means to directly diagnose allergic diseases in children using specific bacteria found in fecal samples. Therefore, there is an urgent need for a method that can directly diagnose allergic diseases in children using specific bacteria found in fecal samples. Summary of the Invention:
[0006] Purpose of the invention: The technical problem to be solved by the present invention is to provide a microbial biomarker composition for allergic populations based on intestinal typing.
[0007] Another technical problem to be solved by the present invention is to provide the application of a combination of microbial biomarkers for allergic populations based on intestinal typing in the preparation of reagents or kits for identifying drugs for allergic diseases.
[0008] Another technical problem that this invention aims to solve is to provide a kit for identifying allergic diseases.
[0009] Another technical problem that this invention aims to solve is to provide a method for constructing the allergy prediction model based on intestinal typing.
[0010] Technical Solution: To solve the above-mentioned technical problems, the present invention provides a microbial biomarker composition for allergic individuals based on intestinal typing, comprising one or more of the following: Incertae Sedis, Bombiscardovia, Eggerthella, Erysipelatoclostridium, Lachnospira, [Eubacterium]eligens group, Ligilactobacillus, [Ruminococcus]gauvreauii group, Lachnospiraceae NK4A136 group, Bifidobacterium, Veillonella, Collinsella, Faecalibacterium, [Clostridium]innocuum group, and Enterococcus.
[0011] The present invention also includes the application of the aforementioned microbial biomarker composition for allergic populations based on intestinal typing in the preparation of reagents or kits for identifying allergic diseases.
[0012] The present invention also includes a kit for identifying allergic diseases, the kit comprising a detection reagent for detecting one or more of the following in a sample containing the gut microbiota of the subject: Incertae Sedis, Bombiscardovia, Eggerthella, Erysipelatoclostridium, Lachnospira, [Eubacterium]eligens group, Ligilactobacillus, [Ruminococcus]gauvreauii group, Lachnospiraceae NK4A136 group, Bifidobacterium, Veillonella, Collinsella, Faecalibacterium, [Clostridium]innocuum group, and Enterococcus.
[0013] The samples include, but are not limited to, fecal samples.
[0014] This invention also includes a method for constructing the allergy prediction model based on intestinal typing, comprising the following steps:
[0015] 1) Filter samples suitable for the appropriate age group from publicly available data from NCBI or SRA;
[0016] 2) Download the raw 16S sequencing data and corresponding sample information data selected in step one, including age and whether or not there are allergies;
[0017] 3) Identify the microbial community composition structure data of each sample using the 16S analysis workflow;
[0018] 4) Intestinal type model training: Use the structural topic model to train the microbial community composition structure data of the above-processed samples to obtain the optimal number of intestinal types, the composition structure of each intestinal type, and the parameters of the intestinal type model.
[0019] 5) Use the intestinal type model to predict the intestinal type of the training samples, count the number of allergies in each intestinal type, and determine the intestinal type distribution of each sample;
[0020] 6) Feature selection: The data in step 5) is split into training set and test set. Then, the importance of microbial features in the model is evaluated in the training set by random forest method. Finally, 15 microbial features with the highest importance are selected as markers.
[0021] 7) Model training: Using the 15 microbial biomarkers from step 6) as features, train a binary classification model that can distinguish between normal children and allergic children using logistic regression, support vector machine, or random forest algorithms.
[0022] In step 1), the age range is 1-2 years old.
[0023] Among them, the intestinal type distribution in step 5) includes E1 represented by Blautia, E2 represented by Bacteroides, and E3 represented by Bifidobacterium.
[0024] Among them, the 15 microbial markers in step 6) include Incertae Sedis, Bombiscardovia, Eggerthella, Erysipelatoclostridium, Lachnospira, [Eubacterium]eligens group, Ligilactobacillus, [Ruminococcus]gauvreauii group, Lachnospiraceae NK4A136 group, Bifidobacterium, Veillonella, Collinsella, Faecalibacterium, [Clostridium]innocuum group, and Enterococcus.
[0025] The present invention also includes the allergy prediction model obtained by the construction method described above.
[0026] The present invention also includes the application of the allergy prediction model in the preparation of a system for judging the allergy risk of a sample.
[0027] The present invention also includes a method for allergy diagnosis and prediction based on gut type and machine learning.
[0028] This invention includes an allergy assessment method comprising the following steps:
[0029] 1) Intestinal type classification is determined based on large-scale publicly available 16S sequencing data, and an intestinal type classification model is trained.
[0030] 2) Based on the above gut type classification results, machine learning methods are applied to model and search for microbial biomarkers related to allergic children;
[0031] 3) Extract microbial genomic DNA from stool samples of normal children and children with allergies;
[0032] 4) PCR amplification of the V3-V4 region (or other variable region) of the 16S ribosomal rRNA sequence, constructing a library and sequencing it.
[0033] 4) Perform bioinformatics analysis on sequencing data to identify the structure of the microbial community;
[0034] 6) Based on the above microbial community identification results, the intestinal type of the sample is calculated using the intestinal type classification model from step one;
[0035] 7) Based on the above intestinal type classification results, the corresponding machine model and microbial markers are used to determine the allergy risk of the sample.
[0036] In summary, this invention, based on research using large-scale publicly available data, collects a large amount of publicly available microbial sequencing data to classify the gut microbiota types of children of different ages. Then, it uses machine learning methods to identify allergy biomarkers in different gut types and predicts the risk of allergies in infants and young children using these microbial biomarkers. This invention involves collecting children's feces, extracting genomic DNA from the microorganisms in the feces under laboratory conditions, amplifying it, and using next-generation sequencing technology to detect the structure of the fecal microbiota. The gut type of the sample is determined from the microbiota structure, and machine learning methods are used to predict and assess the allergy risk of the sample.
[0037] Beneficial effects: Compared with the prior art, the present invention has the following advantages:
[0038] 1) This invention proposes for the first time an allergy prediction model based on intestinal type stratification. It first classifies the sample by intestinal type and then uses the corresponding machine learning model to predict the allergy risk of the sample.
[0039] 2) This invention, through training and learning from more than 2,000 cases of data, identified 15 microbial biomarkers in the gut type centered on Bifidibacterium, which can effectively predict the allergy risk of children with this gut type.
[0040] 3) This invention provides a non-invasive method to detect a child's allergy risk using only a child's stool sample. This invention provides a new technological foundation for future research on childhood allergies and offers a new non-invasive method for detecting childhood allergies. Attached Figure Description
[0041] Figure 1 Distribution of intestinal patterns in 2310 samples;
[0042] Figure 2 Composition of each enterotype genera (enterotype 1 represents E1, enterotype 2 represents E2, and enterotype 3 represents E3);
[0043] Figure 3 Performance of the machine learning allergy model on the test set. Detailed Implementation
[0044] Example 1: Learning the Intestinal Type Classification Model in 1-2 Year Old Children
[0045] This embodiment mainly includes the following steps: data search and filtering, data download, data analysis, model training, and model prediction.
[0046] 1. By searching and reading literature on children's fecal gut microbiota, we screened literature suitable for the age group (1-2 years old), and then analyzed the literature to select studies that used 16S sequencing and whose data were publicly available. Specifically, we selected studies from NCBI or ENA, totaling 2311 samples.
[0047] 2. Download the raw sequencing data and corresponding sample information (such as age, allergies, etc.) selected in step 1 from the public database;
[0048] Table 1
[0049]
[0050]
[0051] 3. Sequencing data analysis:
[0052] The sequencing data were processed using the 16S amplicon analysis workflow based on DADA2 (https: / / benjjneb.github.io / dada2 / tutorial.html). The operation steps are as follows:
[0053] (1) First, the 16S rRNA sequence of each sample is obtained by removing the adapter and splitting the sequence. Then, after quality filtering, noise removal, merging and chimera removal, a 100% consistent ASV (Amplicon Sequence Variant) sequence is obtained.
[0054] (2) Align the ASV sequence to the Silva database (v138) to obtain the species name of the sequence.
[0055] (3) Integrate the naming of different reference sequences, merge the abundance of sequences annotated to the same species into the abundance of the same species, and obtain the microbial composition data of the samples used.
[0056] 4. Intestinal type model training: Use the structured topic model (R language stm package) to train the above-processed sample microbial composition data to obtain the optimal number of intestinal types, the composition structure of each intestinal type, and the parameters of the intestinal type model;
[0057] 5. Using the intestinal type model, predict the intestinal type of the training samples, count the number of allergic cases in each intestinal type (see Table 2, composition structure of the three intestinal types), and determine the intestinal type distribution of each sample (see Table 2). Figure 1 It is evident that among infants aged 1-2 years, Bacteroides enterotype has the highest proportion, followed by Bifidobacterium enterotype. This is closely related to breastfeeding and complementary food introduction during the early growth stages of infants. In terms of the incidence of allergies, the allergy rate among infants with Bifidobacterium enterotype is only 8.59%, far lower than the other two enterotypes (both above 14%).
[0058] Table 2 Distribution of allergic and normal infants in different gut types.
[0059]
[0060]
[0061] Example 2: Search for allergy-related microbial biomarkers in 1-2 year old children with Bifidobacterium enterotype
[0062] This embodiment mainly includes the following steps: data preparation, feature selection, model training, and model evaluation.
[0063] 1. Data processing: Samples with the intestinal type Bifidobacterium were selected from the more than 2,310 samples in Example 1;
[0064] 2. Feature selection: The data from the first step was split into training and test sets in a ratio of 8:2. Then, the importance of microbial features in the model was evaluated in the training set using the random forest method. Finally, 15 of the most important microbial features were selected as biomarkers (see Table 3 below, 15 microbial biomarkers found in the intestinal type with Bifidobacterium as the core).
[0065] Table 3 15 Microbial Markers
[0066]
[0067] 3. Model Training: Using the 15 markers from step 2 as features, a binary classification model is trained to distinguish between normal children and allergic children. The random forest algorithm is selected for model training. Based on the probability output by the model, a probability greater than 50% is selected as allergic, and a probability less than or equal to 50% is selected as normal.
[0068] 4. Model Evaluation: The trained model is evaluated on the test set. The model's AUC on the test set is as follows: Figure 3 The confusion matrix predicted by the model on the test set is shown in Table 4. As can be seen from Table 4, 7 out of the 19 allergy cases predicted by the model were actually allergies, while all the normal samples predicted by the model were actually normal, demonstrating the model's good specificity. The performance parameters of the model on the test set are shown in Table 5. As can be seen from the standard, the model's accuracy reaches 94.62%, and its specificity reaches 100%.
[0069] Table 4 Confusion matrix predicted by the model
[0070]
[0071] Table 5 shows the model's parameter performance on the prediction set.
[0072]
[0073] Example 3: Allergy Risk Assessment in Children Aged 1-2 Years
[0074] This embodiment mainly includes the following steps: sample collection, extraction and library construction and sequencing, data processing, intestinal type determination and allergy risk assessment.
[0075] 1. Sample collection: Collect 2g of fecal samples (e.g., samples from children aged 1-2 years) from 223 cases in Table 4 according to the sample collection procedure, and freeze them.
[0076] 2. Extraction, library preparation, and sequencing:
[0077] a) Extraction: Soil microbial DNA extraction kit (OMEGASoil DNA Kit, M5635-02) was used to extract genomic DNA from microorganisms.
[0078] b) PCR amplification: Microbial RNA contains multiple conserved and variable regions. Here, primers 338F (5'-ACTCCTACGGGAGGCAGCA-3') and 806R (5'-GGACTACHVGGGTWTCTAAT-3') were used to amplify the V3-V4 region of the 16S rRNA gene in the sample by PCR.
[0079] PCR was performed using NEB Q5 DNA high-fidelity polymerase, and the system is shown in Table 6:
[0080] Table 6
[0081]
[0082] The operation process is as follows:
[0083] After preparing all the necessary components for the PCR reaction, pre-denature the template DNA at 98°C for 30 seconds on a PCR instrument to ensure complete denaturation. Then, proceed with the amplification cycle. In each cycle, first, denature the template at 98°C for 15 seconds, then lower the temperature to 50°C and hold for 30 seconds to allow the primers to fully anneal to the template; then, hold at 72°C for 30 seconds to allow the primers to extend on the template and synthesize DNA, completing one cycle. Repeat this cycle 25–27 times to accumulate a large amount of amplified DNA fragments. Finally, hold at 72°C for 5 minutes to ensure complete product extension, and store at 4°C.
[0084] The amplification results were subjected to 2% agarose gel electrophoresis, and the target fragment was recovered using the Axygen gel recovery kit. The fragment size was approximately 480 bp.
[0085] c) Library construction: Library construction was performed using the Illumina TruSeq Nano DNA LT Library Prep Kit. The first step was end repair, which involved using the End Repair Mix 2 from the kit to remove the protruding bases at the 5' end of the DNA, fill in the missing bases at the 3' end, and add a phosphate group at the 5' end.
[0086] The specific steps are as follows:
[0087] The first step is the excision of the protruding bases at the 5' end of the DNA:
[0088] (1) Take 30 ng of the DNA fragment obtained in the above steps and add water to 60 μL, then add 40 μL of End RepairMix2;
[0089] (2) Mix well by pipetting and place on a PCR instrument at 30°C for 30 min;
[0090] (3) The end repair system was purified using BECKMAN AMPure XP beads (purchased from BECKMAN), and finally eluted with 17.5 μL of Resuspension buffer.
[0091] The second step is adding an A base to the 3' end. In this process, an A base is added separately to the 3' end of the DNA to prevent the DNA fragment from self-ligating, and at the same time to ensure that the DNA is connected to a sequencing adapter with a protruding T base at the 3' end. The specific steps are as follows:
[0092] (1) Add 12.5 μL of A-Tailing Mix to the DNA after fragment selection;
[0093] (2) Mix well by blowing with a pipette and incubate on a PCR instrument. The program is as follows: 37℃, 30min; 70℃, 5min; 4℃, 5min; 4℃, ∞.
[0094] The third step is to add a linker with a specific tag. This process is to allow the DNA to ultimately hybridize into the flow cell. The specific steps are as follows:
[0095] (1) Add 2.5 μL of Resuspension buffer, 2.5 μL of Igation Mix and 2.5 μL of DNA adapter Index to the system of the product obtained in the second step.
[0096] (2) Mix well by pipetting and place on a PCR instrument and incubate at 30°C for 10 min;
[0097] (3) Add 5 μL of Stop Ligation buffer;
[0098] (4) The system with added linkers was purified using BECKMAN AMPure XP beads.
[0099] The fourth step is to amplify the DNA fragment with the adapter added by PCR, and then purify the PCR system using BECKMAN AMPure XPbeads.
[0100] The fifth step is to perform final fragment selection and purification of the library using 2% agarose gel electrophoresis.
[0101] d) Sequencing: First, the library undergoes quality control. After passing quality control, sequencing is performed. The library to be sequenced (with non-reproducible indices) is first serially diluted to 2 nM, then mixed according to the required data volume ratio. The mixed library is denatured into single strands using 0.1N NaOH for sequencing. Specifically, paired-end sequencing of 2 × 250 bp is performed on an Illumina NovaSeq machine using the NovaSeq6000SP Reagent Kit (500 cycles).
[0102] 3. Data Processing: The sequencing data was processed using the DADA2 (https: / / benjjneb.github.io / dada2 / tutorial.html) 16S amplicon analysis workflow. The operation steps are as follows:
[0103] a) First, the 16S rRNA sequence of each sample is obtained by removing the adapter and splitting the sequence. Then, after quality filtering, noise reduction, merging and chimera removal, a 100% identical ASV sequence (Amplicon Sequence Variant) is obtained.
[0104] b) Align the ASV sequence to the Silva database (v138) to obtain the species name of the sequence.
[0105] c) Integrate the nomenclature of different reference sequences and merge the abundance of sequences annotated to the same species into the abundance of the same species;
[0106] d) In a single sample, the abundance of species is converted into relative abundance, that is, the number of sequences (reads) of each species divided by the total number of sequences (reads) in the sample; finally, the relative abundance of each bacterium is obtained, that is, the bacterial community composition structure of the sample. 4. Enterotype determination: Based on the bacterial community composition structure of the above sample, the enterotype classification algorithm in Example 1 is used to calculate the enterotype type of the sample;
[0107] 5. Allergy Risk: Different machine learning models and microbial biomarkers were selected for different gut types. The output of the machine model was the allergy risk assessment result for that sample. Ultimately, 7 cases of allergy and 216 cases of normal results were detected from the 223 samples tested. This conclusion is consistent with the model validation results in Example 2.
Claims
1. A microbial biomarker composition for allergic individuals based on gut type classification, characterized in that, It includes Incertae Sedis, Bombiscardovia Egerte (Eggerthella) genus *Clostridium* (Erysipelatoclostridium) genus *Trichophyton* (Lachnospira) Picky Eubacterium (Eubacterium eligens group), Lactobacillus (Ligilactobacillus), Ruminococcus gauvreauii group Family: Trichophyceae, Genus: NK4A136 (Lachnospiraceae NK4A136 group), Bifidobacterium (Bifidobacterium) Veillonella (Veillonella) Collins (Collinsella) spp. (Faecalibacterium), Clostridium innocuum group and Enterococcus spp. Enterococcus) .
2. The use of the microbial biomarker composition for allergic populations based on intestinal typing as described in claim 1 in the preparation of reagents or kits for identifying allergic diseases.
3. A kit for identifying allergic diseases, characterized in that, The kit contains detection reagents, which are used to detect samples of the subject's gut microbiota containing... Incertae Sedis, Bombiscardovia, Egerte (Eggerthella) genus *Clostridium* (Erysipelatoclostridium) genus *Trichophyton* (Lachnospira) Picky Eubacterium (Eubacterium eligens group), Lactobacillus (Ligilactobacillus), Ruminococcus gauvreauii group, Family: Trichophyceae, Genus: NK4A136 (Lachnospiraceae NK4A136 group), Bifidobacterium (Bifidobacterium) Veillonella (Veillonella) Collins (Collinsella) spp. (Faecalibacterium), Clostridium innocuum group and Enterococcus spp. Enterococcus) .
4. The kit for identifying allergic diseases according to claim 3, characterized in that, The samples include fecal samples.
5. A method for constructing an allergy prediction model based on intestinal typing, characterized in that, Includes the following steps: 1) Screen for children's samples from publicly available data in NCBI or SRA; 2) Download the raw 16S sequencing data and corresponding sample information data selected in step one, including age and whether or not there are allergies; 3) Identify the microbial community composition structure data of each sample using the 16S analysis workflow; 4) Intestinal type model training: Use the structural topic model to train the microbial community composition structure data of the above-processed samples to obtain the optimal number of intestinal types, the composition structure of each intestinal type, and the parameters of the intestinal type model. 5) Use the intestinal type model to predict the intestinal type of the training samples, count the number of allergies in each intestinal type, and determine the intestinal type distribution of each sample; 6) Feature selection: The data in step 5) is split into training set and test set. Then, the importance of microbial features in the model is evaluated in the training set by random forest method. Finally, 15 microbial features with the highest importance are selected as markers. The markers include: Incertae Sedis, Bombiscardovia, Egerte (Eggerthella) genus *Clostridium* (Erysipelatoclostridium) genus *Trichophyton* (Lachnospira) Picky Eubacterium (Eubacterium eligens group), Lactobacillus (Ligilactobacillus), Ruminococcus gauvreauii group Family: Trichophyceae, Genus: NK4A136 (Lachnospiraceae NK4A136 group), Bifidobacterium (Bifidobacterium) Veillonella (Veillonella) Collins (Collinsella) spp. (Faecalibacterium), Clostridium innocuum group and Enterococcus spp. Enterococcus); 7) Model training: Using the 15 microbial biomarkers from step 6) as features, train a binary classification model that can distinguish between normal children and allergic children using logistic regression, support vector machine, or random forest algorithms.
6. The method for constructing an allergy prediction model according to claim 5, characterized in that, The age range in step 1) is 1-2 years old.
7. The method for constructing an allergy prediction model according to claim 5, characterized in that, The intestinal pattern distribution in step 5) includes: Blautia E1, represented by [the specific brand], Bacteroides E2, represented by [the specific name], Bifidobacterium E3 is represented by this.
8. The allergy prediction model obtained by the construction method according to any one of claims 5 to 7.
9. The application of the allergy prediction model according to claim 8 in constructing a system for judging the allergy risk of a sample.
Citation Information
Patent Citations
Methylation marker for diagnosing pneumointestinal adenocarcinoma
CN115094142A
Intestine type typing method and device based on intestinal microbial flora structure and medium
CN115938484A