Microbial marker combination for detecting obstructive sleep apnea in children and use thereof
The predictive model constructed by combining oral microbial biomarkers and the random forest algorithm solves the problems of non-invasiveness and accuracy in the diagnosis of childhood OSA, and realizes non-invasive and objective risk assessment and screening of childhood OSA.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2026-01-06
- Publication Date
- 2026-03-20
AI Technical Summary
Existing diagnostic methods for childhood obstructive sleep apnea (OSA), such as nocturnal polysomnography, are complex and expensive, parent questionnaires have limited accuracy, and there is a lack of non-invasive and effective microbial biomarkers for risk assessment.
Using a combination of oral microbial biomarkers, including nine genera such as Streptococcus and Socrates, a predictive model was constructed using a random forest algorithm to assess the risk of OSA in children based on the abundance of microorganisms in saliva samples.
It provides a non-invasive and objective tool for assessing the risk of OSA in children, with high accuracy and specificity, and is suitable for large-scale preliminary screening and monitoring the effectiveness of sleep interventions.
Smart Images

Figure CN121459950B_ABST
Abstract
Description
TECHNICAL FIELD
[0001] The present application relates to the field of biotechnology and health informatics, in particular, to a microorganism marker combination and its application in risk assessment and auxiliary diagnosis of obstructive sleep apnea (OSA) in children. BACKGROUND
[0002] Obstructive sleep apnea (OSA) is a common sleep disorder that repeatedly occurs during sleep, and its incidence in children is increasing. Long-term OSA can lead to impaired neurocognitive development, cardiovascular complications, and growth retardation in children, and a series of serious problems.
[0003] Currently, the gold standard for diagnosing OSA in children is overnight polysomnography. However, the process of overnight polysomnography is complex and expensive, and requires wearing multiple sensors overnight in a specialized sleep laboratory, which causes significant psychological burden and discomfort for children, resulting in poor clinical popularity and compliance of children. In addition, existing screening questionnaires (such as the Children's Sleep Habits Questionnaire) mainly rely on parental subjective recall, and have limited accuracy and specificity.
[0004] In recent years, microbiome research has shown that there is a "gut-brain axis" two-way communication between gut microbes and sleep regulation. However, the association between oral saliva microbes, especially OSA, is still in the blank. The collection of saliva samples is completely non-invasive, simple and low-cost, and is an ideal source of health screening samples for children. Currently, existing technologies have not disclosed saliva microorganism markers that can effectively indicate the risk of OSA in children, and there is a lack of objective and non-invasive assessment methods based on such markers. SUMMARY
[0005] To solve the above problems, the present application provides a non-invasive and objective sleep health screening tool for children based on a combination of oral microorganism markers, which outputs a "risk assessment report" and can be applied to health management, early warning and effect monitoring, and has practicality.
[0006] The specific technical solutions adopted by the present application are as follows:
[0007] In a first aspect, the present application provides a microorganism marker combination for detecting obstructive sleep apnea in children, which is a combination of oral microorganism markers, and specifically includes the following 9 microorganism genera:
[0008] Streptococcus (g__Streptococcus), Solobacterium (g__Solobacterium), Schaalia (g__Schaalia), Capnocytophaga (g__Capnocytophaga), Stomatobaculum (g__Stomatobaculum), Mogibacterium (g__Mogibacterium), Clostridia UCG-014 (g__Clostridia_UCG-014), Selenomonas (g__Selenomonas), Pseudoleptotrichia (g__Pseudoleptotrichia).
[0009] In a second aspect, the present application provides an application of the microorganism marker combination in constructing a prediction model of obstructive sleep apnea in children. The prediction model of obstructive sleep apnea in children is a classification model constructed by using a random forest algorithm based on the relative abundance of the 9 microorganism genera in a saliva sample of a subject child to determine whether the subject child has OSA.
[0010] In a third aspect, the present application provides a method for constructing a prediction model of obstructive sleep apnea in children, comprising the following steps:
[0011] S1. Data collection: collecting saliva samples of children aged 6-12 years old, the training set containing samples of the healthy group and the OSA group, and detecting the relative abundance of microorganisms at the genus level in the samples;
[0012] S2. Feature matrix construction: constructing a feature matrix X with the relative abundance of all microorganism genera in the samples and a label vector y with the group labels confirmed by polysomnography;
[0013] S3. Training a random forest model containing T decision trees using the feature matrix X and the label y, the key parameters of the random forest model and their typical setting values being as follows:
[0014] Number of decision trees: T = 500; number of candidate feature for node splitting: M is the total number of features; maximum depth of decision tree: not limited, until the stop condition is met; minimum number of samples for node splitting: 2; random seed: 42;
[0015] S4. Feature selection: calculating the feature importance score of each microorganism genus by using the average precision drop method, ranking the importance scores in descending order, and retraining the random forest model with the top N important features as input, and evaluating the AUC value on the training set. When N = 9, the model AUC reaches a peak value, and the 9 microorganism genera are determined as the final marker combination.
[0016] S5. Model solidification: using the determined 9 microbial genera features and all training samples, retrain the random forest classifier according to the parameters in step S3 and save the model.
[0017] Further, in step S1, the saliva sample is collected after the child wakes up, brushes teeth, rinses mouth, eats and drinks water.
[0018] Further, in step S1, the relative abundance is the proportion of the sequence number of a single microbial genus to the total sequence number of the sample at the genus level. The detection of the relative abundance is well known to those skilled in the art and can be obtained by extracting sample genomic DNA and performing 16S rRNA gene sequencing.
[0019] In a fourth aspect, the present application provides a prediction model for obstructive sleep apnea in children constructed by the construction method.
[0020] In a fifth aspect, the present application provides a system for predicting obstructive sleep apnea in children, comprising the following modules:
[0021] a data input module for inputting the abundance data of the microbial marker combination in the obtained biological sample of the subject;
[0022] a database storage module for storing the abundance data of the microbial marker combination in the biological samples of a population, the population comprising OSA patients and healthy subjects;
[0023] a disease prediction module connected to the data input module and the database storage module, respectively, for constructing a prediction model using the abundance data of the microbial marker combination in the biological samples of the population, and diagnosing whether the subject has OSA based on the abundance data of the microbial marker combination of the subject obtained from the data input module;
[0024] an output module for outputting a risk assessment report, the report content including the probability of the subject belonging to the OSA group and the healthy group category, and the final determination result of the subject belonging to the OSA group or the healthy group.
[0025] In a sixth aspect, the present application provides a computer readable storage medium storing a computer program, the computer program being executed by a processor to realize the functions of the system for predicting obstructive sleep apnea in children.
[0026] In a seventh aspect, the present application provides a computer device comprising a memory and a processor, the memory storing a computer program, and the processor being configured to execute the computer program to realize the functions of the system for predicting obstructive sleep apnea in children.
[0027] In an eighth aspect, the present application provides use of a reagent for detecting the abundance of the combination of the microbial markers in the preparation of a kit for screening children with obstructive sleep apnea.
[0028] Further, the kit for screening children with obstructive sleep apnea can comprise a genomic DNA extraction reagent and a reagent for detecting the abundance of the combination of the microbial markers; the reagent for detecting the abundance of the combination of the microbial markers comprises primers and / or probes capable of amplifying the microbial markers. The genomic DNA extraction reagent and the abundance detection reagent are well known to those skilled in the art and can be obtained by conventional means, and therefore are not particularly limited.
[0029] Compared with the prior art, the present application has the beneficial effects and significant progress in that:
[0030] 1. The present application first discovers a combination of oral saliva microbial markers significantly related to children with obstructive sleep apnea, which comprises the following 9 microbial genera: g__Streptococcus, g__Solobacterium, g__Schaalia, g__Capnocytophaga, g__Stomatobaculum, g__Mogibacterium, g__Clostridia_UCG-014, g__Selenomonas, and g__Pseudoleptotrichia. By detecting the abundance characteristics of the combination of the microbial markers, the risk of children with obstructive sleep apnea can be effectively predicted. Experiments have proved that the prediction model constructed based on the combination of the markers has high accuracy.
[0031] 2. The obstructive sleep apnea related microbial markers of the present application are detected based on saliva microbial sequencing data, and the results are accurate and safe after strict data screening and verification.
[0032] 3. The present application provides an obstructive sleep apnea prediction model construction method, which is based on a group of auxiliary diagnostic microbial markers screened out, and constructs a random forest model with higher specificity, better screening efficiency and accuracy.
[0033] 4. The combination of the obstructive sleep apnea related microbial markers of the present application can be used for preparing an obstructive sleep apnea diagnosis reagent or kit, which is used for the non-invasive auxiliary diagnosis of children with OSA.
[0034] 5. The application scene of the present application is clear, and the sampling method is non-invasive; it is particularly suitable for large-scale preliminary screening, and long-term and dynamic biological monitoring of the effect of sleep intervention measures. BRIEF DESCRIPTION OF DRAWINGS
[0035] Figure 1 Figure 1: Venn diagram of microbial species in saliva of children with OSA and healthy children in Example 1.
[0036] Figure 2 Figure 2: Heatmap of microbial species in saliva of children with OSA and healthy children in Example 1.
[0037] Figure 3 Figure 3: ROC curve of training set in Example 2.
[0038] Figure 4 Figure 4: ROC curve of independent validation set in Example 3.
[0039] Figure 5 Figure 5: OSA screening result of a validation sample in Example 4. DETAILED DESCRIPTION
[0040] The present application is further described below in conjunction with the accompanying drawings and specific examples. In the following examples, the reagents involved are commercially available conventional reagents unless otherwise specified, and the experimental operations involved are conventional operations in the art unless otherwise specified.
[0041] Example 1: Preliminary comparative analysis of oral microbial differences between children with OSA and healthy controls
[0042] 1. Data collection
[0043] Inclusion criteria:
[0044] (1) Age 6-12 years old;
[0045] (2) Diagnosed as OSA according to the Guidelines for Diagnosis and Treatment of Obstructive Sleep Apnea in Children by polysomnography;
[0046] (3) Not received OSA-related treatment (such as adenoidectomy, continuous positive airway pressure therapy, etc.).
[0047] Exclusion criteria:
[0048] (1) Used antibiotics, probiotics or prebiotics within 4 weeks;
[0049] (2) Had acute respiratory tract infection or oral inflammation within 2 weeks;
[0050] (3) Had congenital heart disease, nervous system disease, genetic metabolic disease or other serious chronic systemic disease;
[0051] (4) Had malformation of maxillofacial development or diagnosed genetic syndrome (such as Down syndrome);
[0052] (5) Unable to complete saliva sample collection.
[0053] 2. Saliva sample collection and processing
[0054] Saliva sample quantity: A total of 72 samples, including 34 samples from the healthy group and 38 samples from the OSA group.
[0055] All samples were collected after the children brushed their teeth, rinsed their mouths, ate, and drank water in the morning to ensure the originality of the saliva microbial state. The children were asked to spit the non-irritating saliva into the collection tube, and the entire process was ensured to be free of contamination.
[0056] 3. DNA extraction and 16S rRNA gene sequencing
[0057] (1) DNA extraction: After genomic DNA extraction using the E.Z.N.A.® soil DNA kit (Omega Bio-tek, Norcross, GA, U.S.), the extracted genomic DNA was detected using 1% agarose gel electrophoresis.
[0058] (2) PCR amplification: PCR amplification was performed using universal primers for the V3-V4 region of the 16S rRNA gene (upstream primer 338F (5'-ACTCCTACGGGAGGCAGCAG-3'); downstream primer 806R (5'-GGACTACHVGGGTWTCTAAT-3')). The amplification products were quality checked using agarose gel electrophoresis.
[0059] (3) Fluorescence quantification: Based on the preliminary quantitative results of electrophoresis, the PCR products were detected and quantified using the QuantiFluor™ -ST blue fluorescence quantification system (Promega Corporation). Then, according to the sequencing requirements of each sample, different samples were mixed in the corresponding proportion to form the library for sequencing (different samples were distinguished by the index on the 5' end of the primer).
[0060] (4) Sequencing: The constructed library was quantified and uniformly mixed for PE250 double-end sequencing.
[0061] 4. Bioinformatics analysis process
[0062] After sample splitting, the PE reads obtained by sequencing are first quality controlled and filtered according to the sequencing quality, and meanwhile, spliced according to the overlap relationship between the double-end reads to obtain the optimized data after quality control and splicing. Then, the sequence noise reduction method (DADA2 / Deblur) is used to process the optimized data to obtain ASV (Amplicon Sequence Variant) representative sequences and relative abundance information. Based on the ASV representative sequences and relative abundance information, community diversity analysis and species difference analysis are performed.
[0063] 5、Results
[0064] The species Venn diagram can be used to count the number of species shared and unique in the sample, and can directly show the similarity and overlap of the species composition of the two groups of samples. Figure 1 The species Venn diagram shows that there are 93 species shared by the two groups of samples, 15 species unique to the healthy group, and 18 species unique to the OSA group.
[0065] The community Heatmap diagram is a color gradient to represent the data size in a two-dimensional matrix or table, and presents the community species composition information. Usually, clustering is performed according to the similarity of the abundance between species or samples, and the results are presented on the community heatmap diagram, so that high-abundance and low-abundance species can be block clustered. Through color change and similarity, the similarity and difference of the community composition of the two groups of samples at each classification level are reflected. Figure 2 The Heatmap diagram shows the distribution of the top dominant species of the OSA group and the healthy group, and the species change trends of the two groups are inconsistent.
[0066] The above results show that there are obvious differences in oral microorganisms between the OSA group and the healthy group, and there is a basic condition for distinguishing children with OSA from healthy people based on oral microorganisms.
[0067] Example 2: Screening of oral microorganism marker combination and model construction
[0068] 1. Best risk screening model construction and training method
[0069] In this embodiment, a random forest algorithm is used to construct a children's OSA risk classification model. The training of the model is a data-driven process based on clear rules and parameters, and the specific steps are as follows:
[0070] (1) Data preparation and feature definition:
[0071] The feature matrix X is constructed with the genus-level relative abundance data of oral microbiota of all training samples (i.e. 72 samples from Example 1 as the training set). Each row represents a sample and each column represents a genus (feature). The class label y of the sample is "OSA group" or "healthy group" diagnosed by PSG.
[0072] (2) Training process of random forest model:
[0073] A random forest containing T decision trees is trained using the feature matrix X and label y. The construction of a single decision tree h k follows the following reproducible rules:
[0074] ① Bootstrap sampling: N samples (N is the total number of samples in the training set) are randomly sampled with replacement from the original training set to form the training subset of the tree. This process is ensured to be reproducible by setting a fixed random seed (random_state).
[0075] ② Node splitting rule: When splitting at each non-leaf node of the tree: first, randomly select m features from all M microbial features as a candidate feature subset. The m is a key parameter of the model, which is usually set as:
[0076]
[0077] Then, from the m candidate features, select the feature and its split threshold that can reduce the Gini Impurity the most after splitting. The Gini Impurity of node t is calculated as:
[0078]
[0079] where c is the number of classes (C=2 in this invention), and p(i|t) is the proportion of samples belonging to class I in node t.
[0080] ③ Stopping rule for tree growth: when the tree reaches the preset maximum depth (max_depth), or the number of samples contained in the node is less than the minimum split sample number (min_samples_split), or the Gini Impurity of the node is 0, stop splitting and mark the node as a leaf node. The output of the leaf node is the class of the majority of samples in the node.
[0081] (3) Model parameter setting and solidification:
[0082] In this example, the key parameters of the random forest model and their typical values are as follows:
[0083] Number of decision trees (n_estimators): T=500;
[0084] Number of candidate features for node splitting (max_features): ;
[0085] Maximum depth of decision tree (max_depth): No limit, until the stopping condition is met;
[0086] Minimum number of split samples per node (min_samples_split): 2;
[0087] Random seed (random_state): 42 (The purpose of fixing the random seed is to ensure that the Bootstrap sampling and node random feature selection processes are completely repeatable).
[0088] (4) Model prediction (voting) mechanism:
[0089] For a new saliva sample X new, the OSA risk assessment process is as follows:
[0090] The relative abundance feature vector of the microbial genera in this sample is input into each decision tree h in the trained random forest. k .
[0091] Each tree h k Produce an independent prediction category c k (OSA group or healthy group).
[0092] Random forests integrate the predictions of all trees using a majority voting method to obtain the final category. :
[0093]
[0094] in, This is an indicator function; its value is 1 when the condition within the parentheses is true, and 0 otherwise. Final output: This represents the category with the most votes, and can also output the probability of belonging to the OSA category. .
[0095] 2. Feature selection and determination of optimal combination
[0096] To obtain a concise and efficient model, we perform feature importance analysis based on the trained random forest model described above.
[0097] (1) Feature importance assessment: The feature importance score for each microbial genus was calculated using the average precision decline method. This score reflects the degree to which the model prediction accuracy decreases after the value of the feature is randomly shuffled in the model prediction. The greater the decrease, the more important the feature.
[0098] (2) Determine the optimal feature combination: We ranked all the microbial genera by importance score in descending order, and then retrained the random forest model using the top N important features as input, and calculated its AUC value on the training set. The results showed that when the number of variables N = 9 in the top importance ranking, the model performance (AUC) of the random forest constructed using this number of variables reached a peak. Thus, the minimum optimal feature combination for constructing the final screening model was determined, i.e., the following 9 microbial genera:
[0099] g__Streptococcus, g__Solobacterium, g__Schaalia, g__Capnocytophaga, g__Stomatobaculum, g__Mogibacterium, g__Clostridia_UCG-014, g__Selenomonas, g__Pseudoleptotrichia.
[0100] (3) Final model solidification: Using the 9 key features and all training samples, retrain the final random forest classifier according to the process and parameters described in Part 1, and save the model (including tree structure, splitting rule, parameters, etc.) for subsequent clinical sample evaluation.
[0101] The classification accuracy of this group of 9-microbial combination features was evaluated. As shown in Figure 3 , the point marked on the ROC curve is the optimal critical value (specificity Specificity = 0.73 and sensitivity Sensitivity = 0.88); the AUC is the area under the corresponding curve: 0.75. This indicates that the classification prediction model has certain discrimination accuracy.
[0102] Example 3: Microbial marker and model detection of OSA effect verification
[0103] 1. Study subjects and sample collection: Collect 20 children sample populations, including 10 OSA patients and 10 healthy people's saliva samples.
[0104] 2. Sample collection and sequencing: As before (steps 3-4 of Example 1), obtain 16S rRNA gene sequencing data for saliva samples and calculate their relative abundance at the genus level.
[0105] Feature extraction: From the above results, extract the relative abundance values of the 9 key microbial genera to form a 9-dimensional feature vector.
[0106] 3. Diagnostic efficiency analysis: draw the receiver operating characteristic curve, analyze the sensitivity and specificity of the microbial marker combination obtained by screening between the healthy group and the OSA subjects in the validation set, to judge its diagnostic efficiency for OSA disease in children.
[0107] The ROC curve was drawn and the AUC value was calculated with the abundance combination of the microbial marker as the detection variable, and the results are shown in Figure 4 The AUC value is 0.710, the specificity is 0.75, and the sensitivity is 0.786, indicating that the combination of saliva microbial markers can be used as a microbial marker for diagnosing OSA disease in children.
[0108] Example 4: Diagnosis of critical and difficult samples
[0109] For a child to be screened, the OSA risk assessment process is as follows:
[0110] 1. Sample collection and sequencing: as before, obtain 16S rRNA gene sequencing data of saliva samples and calculate their relative abundance at the genus level.
[0111] 2. Feature extraction: from the above results, extract the relative abundance values of the 9 key microbial genera to form a 9-dimensional feature vector.
[0112] 3. Model prediction: input the 9-dimensional feature vector into the final random forest model saved in step (3) of Example 2 (the model file contains the structure, split points and parameters of all decision trees).
[0113] 4. Result output: the model runs all the decision trees inside it and performs majority voting to generate a visual report as shown in Figure 5 The report shows the prediction results and the probability of belonging to each category, and the category with the highest probability is taken as the final report result. The final output of this specific embodiment is that the child belongs to the "OSA group" category and the probability of belonging to this group is "Prob_OSA".
[0114] Figure 5 The prediction sample result table shows that the prediction sample belongs to the OSA group. According to the diagnostic gold standard PSG, the child indeed has chronic intermittent hypoxia and sleep fragmentation during sleep, which meets the diagnostic criteria for obstructive sleep apnea. Moreover, the child is close to the diagnostic boundary and is difficult to assess. The prediction result of this child by the method is consistent with the actual situation, indicating that the OSA screening microbial marker combination and model of the present application have high reliability.
[0115] The specific embodiments are only illustrative of the present application, and are not a limitation of the present application, any change made by those skilled in the art after reading the specification of the present application will be protected by the patent law as long as it is within the scope of the claims of the present application.
Claims
1. A combination of microbial biomarkers for detecting obstructive sleep apnea in children, characterized in that, It is a combination of oral microbial markers, including the following 9 genera: Streptococcus, Sorafenib, Salmonella, Carbonylphage, Oral Bacillus, Diplobacterium, Clostridium UCG-014, Lunatomium, and Pseudomonas.
2. The application of the microbial biomarker combination according to claim 1 in constructing a predictive model for obstructive sleep apnea in children, characterized in that, The childhood obstructive sleep apnea prediction model uses a classification model constructed with a random forest algorithm based on the relative abundance of the nine microbial genera in the saliva samples of the tested children to determine whether the tested children have OSA.
3. A method for constructing a predictive model for obstructive sleep apnea in children, characterized in that, Includes the following steps: S1. Data Collection: Saliva samples were collected from children aged 6-12 years to detect the relative abundance of horizontal microorganisms in the samples. The training set included samples from the healthy group and the OSA group. S2. Feature matrix construction: The feature matrix X is constructed using the relative abundance of all microbial genera in the sample, and the label vector y is constructed using the group labels diagnosed by polysomnography. S3. Train a random forest model containing T decision trees using the feature matrix X and the label y. The key parameters of the random forest model and their typical settings are as follows: Number of decision trees: T=500; Number of candidate features for node splitting: M represents the total number of features; maximum depth of the decision tree: unlimited, until the stopping condition is met; minimum number of split samples per node: 2; random seed: 42; S4. Feature selection: The feature importance score of each microbial genus is calculated using the average precision descent method, and the features are sorted in descending order of importance score. The top N important features are then used as input to retrain the random forest model. The AUC value is evaluated on the training set. When N=9, the model AUC reaches its peak, and the nine microbial genera described in claim 1 are determined as the final biomarker combination. S5. Model Consolidation: Using the identified features of the 9 microbial genera and all training samples, retrain the random forest classifier according to the parameters in step S3 and save the model.
4. The construction method according to claim 3, characterized in that, In step S1, saliva samples were collected from the children after they woke up in the morning, before they brushed their teeth, rinsed their mouths, ate, and drank water; genomic DNA was extracted from the samples and 16S rRNA sequencing was performed to obtain relative abundance data of microorganisms at the genus level.
5. A childhood obstructive sleep apnea prediction model constructed using the construction method described in claim 3.
6. A system for predicting obstructive sleep apnea in children, characterized in that, Includes the following modules: The data input module is used to input the abundance data of the combination of microbial markers as described in claim 1 in the obtained biological samples of the subjects; A database storage module for storing abundance data of the microbial biomarker combinations in biological samples of a population, including OSA patients and healthy subjects; The disease prediction module is connected to the data input module and the database storage module, respectively, and is used to construct a prediction model using the abundance data of the combination of microbial markers in the biological samples of the population, and to diagnose whether the subject has OSA based on the abundance data of the combination of microbial markers of the subject obtained from the data input module. The output module outputs a risk assessment report, which includes the probability that the subject belongs to the OSA group or the healthy group, as well as the final determination of whether the subject belongs to the OSA group or the healthy group.
7. A computer-readable storage medium, characterized in that, It stores a computer program that, when executed by a processor, can perform the functions of the system as described in claim 6.
8. A computer device, characterized in that, It includes a memory and a processor, the memory storing a computer program and the processor executing the computer program to perform the functions of the system as described in claim 6.
9. The use of the reagent for detecting the abundance of the microbial biomarker combination of claim 1 in the preparation of a kit for screening obstructive sleep apnea in children.
Citation Information
Patent Citations
Systems and methods for screening obstructive sleep apnea during wakefulness using anthropometric information and tracheal breathing sounds
CA3089395A1
Oral cavity microorganism for identifying taste sensitivity as well as identification method and application of oral cavity microorganism
CN120442769A