A method for realizing health status monitoring based on an intestinal microbiota statistical model
Through PCA-based statistical models and health index, combined with the comparison of PCA learning and characteristic contribution diagnosis, the problems of low unhealth detection rates and inability to analyze differences and risks of healthy people in the existing technology are solved, and efficient health status monitoring and personalized diagnosis are achieved.
Patent Information
- Application Number
- CN202310332827.1
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2023-03-30
- Publication Date
- 2025-06-24
- Estimated Expiration
- 2043-03-30
AI Technical Summary
The existing gut flora health index method has a low detection rate when distinguishing healthy and unhealthy populations, which cannot analyze the differences and risks in healthy populations, and cannot trace back to the most responsible species associated with the unhealthy phenotype.
Relative abundance data of intestinal microorganisms were obtained through metagenomic sequencing, statistical models and health indexes based on PCA were established, and health patterns in healthy populations were revealed using comparative PCA learning methods, and targets were identified through characteristic contribution diagnosis and identification.
It improves the detection rate of unhealthy people, can detect unhealthy in a timely manner, reduces the economic losses of patients, and produces greater economic benefits; at the same time, analyzes the differences and risks in healthy people, reduces the risk of illness in healthy people, and achieves personalized diagnosis.
Smart Images

Figure CN116364287B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to a method for realizing health status monitoring based on an intestinal microbiota statistical model, and belongs to the field of microbiological technology. Background Art
[0002] The gut microbiota refers to a large number of microbial communities existing in the human gastrointestinal tract. Interpreting the role of these important organs in health has attracted great interest in the health research community. After a decade of efforts, it has now become a global consensus that these microorganisms are crucial for human health because they replace many functions of the host, and any dysregulation of the microbiota can significantly affect the host's immunity, metabolism, and even neurobehavior. In addition, many active joint projects have brought an increasing impetus to large-scale data analysis and personal health understanding, which has generated a large amount of reference datasets that can be used for large-scale retrospective studies. Therefore, highly automated and powerful bioinformatics tools for personalized health status inference are expected to transform the composition of the human microbiome into useful clinical indications for non-invasive health monitoring, diagnosis, and treatment.
[0003] Generally, these high-throughput raw sequencing data reads are clustered and organized into operational taxonomic units (OTUs) for downstream analysis, which are usually high-dimensional matrices with high variability and high sparsity. In population-level health analysis and disease-related feature exploration, statistical monitoring teams are essential to extract appropriate knowledge and intelligence from the composition table and facilitate timely health warnings.
[0004] In the microbiome literature, principal component analysis (PCA) is the most widely adopted statistical method. As a simple and effective model for data inspection, interpretation, and utilization, PCA allows feature extraction and knowledge representation by deconstructing the differences or correlations between samples. In this process, the high-dimensional composition will be significantly reduced, and an elegant sorted visualization can be presented for discrimination between sample groups. However, the research in this field is limited to qualitative evaluation and lacks quantitative analysis. This in turn affects the practical scope of unhealthy phenotype analysis and health understanding.
[0005] The existing Gut Microbiota Health Index (GMHI) method is formulated based on 50 species that simultaneously include healthy prevalent and healthy rare species. This method can distinguish healthy species from unhealthy species with relatively high balanced accuracy. However, when distinguishing between healthy and unhealthy populations, the detection rate for unhealthy populations is relatively low. For clinical applications, this may become implausible and deceptive because the cost of missing an unhealthy alert is likely to be a disaster for an unhealthy diagnosis. Moreover, this method only considers the classification of healthy and unhealthy populations and does not consider the further division of healthy populations and the differences and risks existing among healthy populations. In addition, when formulating this method, the impact of healthy prevalent and healthy rare species on the gut microbiota health index was not considered, so it is impossible to trace back to those most responsible species related to the reported unhealthy phenotypes. These drawbacks limit the model interpretation for further personalized medication and are difficult to support personalized health analysis. Summary of the Invention
[0006] To solve the existing problems of low unhealthy detection rate, inability to analyze the differences and risks existing among healthy populations, and inability to trace back to those most responsible species related to the reported unhealthy phenotypes, the present invention provides a method for realizing health status monitoring based on a gut microbiota statistical model. The method includes:
[0007] Step 1: Obtain the relative abundance data of gut microbiota of population samples through metagenomic sequencing and record the sample labels. Set all sample data labeled as healthy as the training set, and the remaining samples as the test set;
[0008] Step 2: Preprocess the data in the training set and the test set, including feature selection and data standardization;
[0009] Step 3: Establish a statistical model and a health index for the preprocessed training set data based on PCA (Principal Component Analysis), and perform health prediction on the training set data;
[0010] Step 4: Select the training set sample data and test set sample data predicted to be healthy in Step 3, and reveal the healthy patterns among healthy populations through contrastive PCA (Contrastive Principal Component Analysis) learning, and further analyze the differences among healthy patterns;
[0011] Step 5: Select the preprocessed test set samples, establish a health index through the statistical model in Step 3, and perform health prediction on the test set samples. And diagnose and identify the corresponding targets through the contribution of each feature to the health index.
[0012] Optionally, the feature selection in Step 2 is:
[0013] Based on obtaining the relative abundance data of gut microbiota in a population sample, the Kolmogorov-Smirnov test is used for samples labeled as healthy and unhealthy to identify those healthy prevalent and healthy scarce features, where the healthy prevalent features are defined by rejecting the alternative hypothesis that, at a significance level of 0.001, the empirical cumulative density of the feature in the healthy population is less than that in the unhealthy population, and the healthy scarce features are also defined by rejecting the alternative hypothesis that, at a significance level p, the empirical cumulative density of the feature in the unhealthy population is less than that in the healthy population.
[0014] Optionally, the data normalization described in step 2 is as follows:
[0015] For a reasonable analysis using PCA, the data in the training set and test set after the above feature selection needs to be transformed. Relative abundance is considered in the present invention. These values can range from 0 to relatively large actual values, and most magnitudes range from 10 -3 to 10 -1 . Considering that low-abundance features may play an important role in health conditions, the following logarithmic transformation is designed and applied:
[0016] lt(x) = log(2x + σ)
[0017] To avoid numerical problems at the origin, a small σ is added, where σ is designed to be 10 -5 , once the data has been transformed, z-score normalization is performed to adjust the mean to 0 and the standard deviation to 1.
[0018] Optionally, step 3 includes:
[0019] Step 3.1: The preprocessed training set sample data is composed into an OTU matrix. Assuming the OTU matrix X ∈ R N×D consisting of D microbial features of N samples, the PCA model structure is as follows:
[0020] X = TP T + E (9)
[0021] where T ∈ R N×d is the score matrix, P ∈ R D×d is the loading matrix, d is the dimension of the retained principal components, E is the residual matrix, and by performing the eigenvalue decomposition of the covariance matrix S = X T X / N - 1, the following can be obtained:
[0022]
[0023] where is the residual load, Λ and are the eigenvalues of the principal component and the residual subspace;
[0024] Therefore, the principal component subspace PCS and the residual subspace RS of the data are defined as:
[0025]
[0026]
[0027] where C and are the projection matrices of the principal component subspace and the residual subspace;
[0028] To determine the correct number of principal components (PCs), assume that the eigenvalues are arranged in descending order λ1, λ2,..., λ D , and the percentage of explained variance (PEVs) for each eigenvalue is defined as λ i / ∑ i λ i , and the correct number of PCs is determined by the cumulative PEVs meeting the result;
[0029] Step 3.2: According to the determined number of PCs, design the PMI chart, RMI chart, and CMI chart;
[0030] The PMI chart is as follows: Given a sample x, PMI monitors the principal component subspace, which is defined as PMI(x) = x T Dx, where D = PΛ -1 P T , and the control limits at the confidence level (1 - α)100% are determined by the chi-square distribution , and the principal component dimension d is the degree of freedom;
[0031] The RMI chart is as follows: RMI monitors the residual subspace, which is defined as The control limit is where is calculated using the eigenvalues;
[0032] The CMI chart is as follows: The combined index is defined as CMI = x T Φx, where The control limit is where and
[0033] Step 3.3: Transform the PMI index, RMI index, and CMI index into the general quadratic form Ind(x) = x T Mx, and the M of each index is respectively expressed as D, and Φ as described above. If any index exceeds the corresponding threshold, an unhealthy situation is reported.
[0034] Optionally, the process of comparing PCA learning in step 4 includes:
[0035] Performing comparative PCA learning on healthy individuals reported through the health index in the training set and unhealthy individuals in the test set:
[0036] Seeking a comparison direction p with weights between high target health variance and low unhealthy variance through comparative PCA * , where p * = argmax(S h - αS uh ). S h and S uh represent the covariance matrices of the above-mentioned healthy and unhealthy individuals respectively, and α represents a contrast parameter used to quantify the weight between high target health variance and low unhealthy variance.
[0037] Optionally, the differences between the healthy patterns in step 4 include:
[0038] 1) Differences in the species composition of healthy patterns;
[0039] 2) Differences in the geographical distribution of healthy patterns;
[0040] 3) Differences in the metabolic systems of healthy patterns:
[0041] First, map OTUs at the genus level or lower to the genome-scale metabolic model GSMMs, and standardize the reactions according to the OTU abundances in the samples;
[0042] Then, perform a two-sample t-test on the reaction abundances of each pair of healthy patterns to determine the reactions with significantly different abundances between groups, extract the metabolic subsystems of the reactions from the GSMM, and perform enrichment analysis on the subsystems using Fisher's exact test;
[0043] Finally, calculate the average abundance differences of each subsystem in each pair of healthy patterns for comparative analysis.
[0044] Optionally, the contribution diagnosis process of each feature to the health index in step 5 includes:
[0045] According to the health index, use the BHC (Bacteria-to-health-index contribution) plot for contribution diagnosis, and reconstruct the normal state by adding a correction term to the unhealthy components;
[0046] Given an unhealthy sample x, assuming that feature i has potential abnormal behavior, the reconstructed component unit is:
[0047] z i= x - ξ i f i (13)
[0048] where ξ i is the direction and f i is the magnitude of change;
[0049] Then, a goal can be set to optimize the health index:
[0050] minInd(z i ) = (x - ξ i f i ) T M(x - ξ i f i ) (14)
[0051] By taking the first derivative equal to zero, we get:
[0052]
[0053] where M represents the above-mentioned D, and Φ, φ i represents the contribution of feature i to the health index.
[0054] The second object of the present invention is to provide a computer-readable storage medium storing computer-executable instructions, and when the computer-executable instructions are executed by a processor, the above-mentioned intestinal health status monitoring method is implemented.
[0055] The beneficial effects of the present invention are as follows:
[0056] The technical solution provided by the present invention solves the problems of low unhealthy monitoring rate, inability to analyze the differences and risks existing in healthy populations, and inability to trace back to the most responsible species related to the reported unhealthy phenotypes. By establishing a statistical model and a health index for healthy populations, the detection rate of unhealthy populations is improved, the detection of unhealthy conditions can be timely, the economic losses of patients can be reduced, and greater economic benefits can be generated; by using the comparative PCA learning method, the differences existing in healthy populations are further analyzed, the risks existing in healthy populations are evaluated in a timely manner, and the risk of healthy populations getting sick is reduced; through the BHC diagnosis of the statistical model, the most responsible species related to the reported unhealthy phenotypes can be traced back in a timely manner, so as to achieve the effect of personalized diagnosis. Description of the Drawings
[0057] To more clearly illustrate the technical solutions in the embodiments of the present invention, the following will briefly introduce the accompanying drawings required in the description of the embodiments. Obviously, the accompanying drawings in the following description are only some embodiments of the present invention. For those of ordinary skill in the art, without creative efforts, other drawings can be obtained based on these drawings.
[0058] Figure 1 It is the overall process schematic diagram of the present invention. Specific embodiments
[0059] To make the objectives, technical solutions and advantages of the present invention clearer, the following will further describe the embodiments of the present invention in detail with reference to the accompanying drawings.
[0060] First, the basic knowledge related to the present invention will be introduced:
[0061] I. Metagenomic sequencing technology
[0062] Metagenomic sequencing technology is a high-throughput sequencing technology used to study the microbial community structure and function in the environment. It can simultaneously detect a large number of microbial species and their quantity information existing in the environment. The brief summary of metagenomic sequencing technology is as follows:
[0063] 1. Sample collection and DNA extraction: Samples are collected from the environment, and DNA is extracted from the samples.
[0064] 2. DNA library construction: DNA fragments are amplified by PCR and a library is constructed.
[0065] 3. High-throughput sequencing: The library is sequenced using a high-throughput sequencing platform to generate a large amount of sequence data.
[0066] 4. Bioinformatics analysis: Bioinformatics analysis is performed on metagenomic sequencing data, such as OTU clustering, species annotation, functional annotation, diversity analysis, etc., which can deeply study the microbial community structure and function.
[0067] II. OTU matrix
[0068] The OTU matrix is a matrix used to represent metagenomic sequencing data. Each row represents a sample, and each column represents an OTU (operational taxonomic unit). However, due to the existence of a large number of unknown microbial species in the microbial community, OTUs are usually obtained by clustering the sequencing data. In the OTU matrix, each cell represents the abundance value of an OTU in a sample, and this value usually refers to the frequency or quantity of the OTU appearing in the sample. Since the number of samples and the number of OTUs may be very large, the OTU matrix is usually a very sparse matrix.
[0069] III. PCA Method
[0070] PCA is a commonly used method for data dimensionality reduction and feature extraction. It can transform high-dimensional data into low-dimensional data while retaining the main features and structure of the data. Specifically, PCA projects the original data onto a set of orthogonal principal components. Each principal component is a linear combination, and they are arranged in descending order of variance. That is, the first principal component contains the information with the largest variance in the data, the second principal component contains the second-largest information, and so on.
[0071] Example 1:
[0072] This example provides a method for monitoring the intestinal health status. Refer to Figure 1 The method includes:
[0073] Step 1: Obtain the relative abundance data of intestinal microorganisms in population samples through metagenomic sequencing and record the sample labels. Set the data of all samples labeled as healthy as the training set, and the remaining samples as the test set;
[0074] Step 2: Preprocess the data in the training set and test set, including feature selection and data standardization;
[0075] Step 3: Based on PCA (Principal Component Analysis), establish a statistical model and a health index for the preprocessed training set data, and perform health prediction on the training set data;
[0076] Step 4: Select the training set sample data predicted as healthy and the test set sample data in Step 3, and reveal the health patterns in the healthy population through contrastive PCA (Contrastive Principal Component Analysis) learning, and further analyze the differences between the health patterns;
[0077] Step 5: Select the preprocessed test set samples, establish a health index through the statistical model in Step 3, and perform health prediction on the test set samples. And diagnose and identify the corresponding targets through the contribution of each feature to the health index.
[0078] Example 2:
[0079] This example provides a method for monitoring the intestinal health status. Refer to Figure 1 The method includes:
[0080] Step 1: Obtain the relative abundance data of intestinal microorganisms in population samples through metagenomic sequencing and record the sample labels. Set the data of all samples labeled as healthy as the training set, and the remaining samples as the test set;
[0081] Step 2: Preprocess the data in the training set and the test set, including feature selection and data standardization.
[0082] Step 2.1: Based on the obtained relative abundance data of the gut microbiota in the population samples, perform the Kolmogorov-Smirnov test on the samples labeled as healthy and unhealthy to identify those healthy prevalent and healthy scarce features, where the healthy prevalent features are defined by rejecting the alternative hypothesis that, at a significance level of 0.001, the empirical cumulative density of the feature in the healthy population is less than that in the unhealthy population, and the healthy scarce features are also defined by rejecting the alternative hypothesis that, at a significance level p, the empirical cumulative density of the feature in the unhealthy population is less than that in the healthy population.
[0083] Step 2.2: In order to perform a reasonable analysis using PCA, it is necessary to transform the data in the training set and the test set after the above feature selection. Relative abundance is considered in the present invention. The range of these values can be from 0 to a relatively large actual value, and most magnitudes range from 10 -3 to 10 -1 . Considering that low-abundance features may play an important role in the health condition, the following logarithmic transformation is designed and applied:
[0084] lt(x) = log(2x + σ)
[0085] To avoid numerical problems at the origin, a small σ is added, where σ is designed to be 10 -5 . Once the data has been transformed, z-score standardization is performed to adjust the mean to 0 and the standard deviation to 1.
[0086] Step 3: Based on PCA, establish a statistical model and a health index for the preprocessed training set data, and implement health prediction for the training set data.
[0087] Step 3.1: Compose the preprocessed training set sample data into an OTU matrix. Assume the OTU matrix X ∈ R N×D consisting of D microbial features of N samples. The PCA model structure is as follows:
[0088] X = TP T + E (16)
[0089] where T ∈ R N×d is the score matrix, P ∈ R D×d is the loading matrix, d is the dimension of the retained principal components, E is the residual matrix, and by performing the eigenvalue decomposition of the covariance matrix S = X T X / N - 1, the following can be obtained:
[0090]
[0091] wherein is the residual load, Λ and are the eigenvalues of the principal components and the residual subspace;
[0092] Therefore, the principal component subspace PCS and the residual subspace RS of the data are defined as:
[0093]
[0094]
[0095] where C and are the projection matrices of the principal component subspace and the residual subspace;
[0096] To determine the correct number of principal components (PCs), assume that the eigenvalues are arranged in descending order λ1, λ2,..., λ D , and the percentage of explained variance (PEVs) for each eigenvalue is defined as λ i / Σ i λ i , and the correct number of PCs is determined by the cumulative PEVs meeting the result;
[0097] Step 3.2: According to the determined number of PCs, design PMI charts, RMI charts, and CMI charts;
[0098] The PMI chart is: Given a sample x, PMI monitors the principal component subspace, which is defined as PMI(x) = x T Dx, where D = PΛ -1 P T , and the control limits at the confidence level (1 - α)100% are determined by the chi-square distribution , and the principal component dimension d is the degree of freedom;
[0099] The RMI chart is: RMI monitors the residual subspace, which is defined as The control limits are where is calculated using the eigenvalues;
[0100] The CMI chart is: The combined index is defined as CMI = x T Φx, where The control limits are where and
[0101] Step 3.3: Transform the PMI index, RMI index, and CMI index into the general quadratic form Ind(x) = x T Mx, and the M for each index is respectively expressed as the above D, and Φ. If any exponent exceeds the corresponding threshold, an unhealthy situation is reported.
[0102] Step 4: Select the training set sample data and test set sample data predicted to be healthy in Step 3, and reveal the healthy patterns among healthy people through contrastive PCA (Contrastive Principal Component Analysis) learning, and further analyze the differences between healthy patterns;
[0103] Step 4.1: Conduct contrastive PCA learning on the healthy people reported by the health index in the training set and the unhealthy people in the test set. Contrastive PCA seeks the contrast direction p that balances high target healthy variance and low unhealthy variance * , where p * = argmax(S h - αS uh ), S h and S uh represent the covariance matrices of the above-mentioned healthy and unhealthy people respectively, and α represents the contrast parameter used to quantify the weight between high target healthy variance and low unhealthy variance. By eliminating the unhealthy confounding, the healthy patterns can be highlighted. Once the contrast direction is calculated and the latent projection is completed, a Gaussian mixture model is used for clustering, which can be determined by adjusting the number of clusters and according to the Bayesian information criterion (BIC) fitted for each class.
[0104] Step 4.2: After identifying the healthy patterns, examine the microbial composition under each pattern, and give the comparison details of the differences in microbial composition between healthy patterns. Subsequently, according to the geographical information contained in the samples, compare the differences in the geographical distribution of healthy patterns. After examining the microbial composition and geographical information of each pattern, it is necessary to analyze the metabolic system of each pattern. First, map the OTUs at the genus level or lower to the genome-scale metabolic models (GSMMs), and standardize the reactions according to the OTU abundances in the samples. Then, a two-sample t-test is performed on the reaction abundances of each pair of healthy patterns to determine the reactions with significantly different abundances between groups. Extract the metabolic subsystems of the reactions from the GSMM, and use Fisher's exact test for enrichment analysis of the subsystems. Finally, calculate the average abundance differences of each subsystem in each pair of healthy patterns for comparative analysis.
[0105] Step 5: Select the preprocessed test set samples, establish a health index through the statistical model in Step 3, and achieve health prediction for the test set samples. And identify the corresponding targets through the contribution of each feature to the health index.
[0106] According to the health index, the BHC (Bacteria-to-health-index contribution) graph is used for contribution diagnosis, and the normal state is reconstructed by adding a correction term to the unhealthy components;
[0107] Given an unhealthy sample x, assuming that feature i has potential abnormal behavior, the reconstructed component unit is:
[0108] z i = x - ξ i f i (20)
[0109] where ξ i is the direction and f i is the amplitude of change;
[0110] Then the goal can be formulated to optimize the health index:
[0111] minInd(z i ) = (x - ξ i f i ) T M(x - ξ i f i ) (21)
[0112] By taking the first derivative equal to zero, we get:
[0113]
[0114] where M represents D, and Φ above, and φ i represents the contribution of feature i to the health index.
[0115] A method for realizing personalized health status monitoring based on the statistical model of gut microbiota provided in this embodiment, by establishing a statistical model and a health index for healthy people, improves the detection rate of unhealthy people, is more useful for the health management of unhealthy people, can detect unhealthy conditions in a timely manner, reduce the economic losses of patients, and generate greater economic benefits; adopting the comparative PCA learning method, further analyzes the differences existing in healthy people, evaluates the risks existing in healthy people in a timely manner, and reduces the risk of healthy people getting sick; through the BHC diagnosis of the statistical model, those most responsible species related to the reported unhealthy phenotypes can be traced in a timely manner, so as to achieve the effect of personalized diagnosis.
[0116] Some steps in the embodiments of the present invention can be implemented using software, and the corresponding software program can be stored in a readable storage medium, such as an optical disc or a hard disk, etc.
[0117] The above are only the preferred embodiments of the present invention, and are not intended to limit the present invention. Any modifications, equivalent replacements, improvements, etc. made within the spirit and principles of the present invention shall be included within the protection scope of the present invention.
Claims
1. A method for monitoring the intestinal health status, characterized in that, The method includes the following: Step 1: Obtain the relative abundance data of gut microbiota in population samples through metagenomic sequencing, record the sample labels, and set the data of all samples labeled as healthy as the training set, and the remaining samples as the test set; Step 2: Preprocess the data in the training set and the test set, including feature selection and data standardization; Step 3: Based on PCA, establish a statistical model and a health index for the preprocessed training set data, and perform health prediction on the training set data; Step 4: Select the training set sample data and test set sample data predicted as healthy in Step 3, reveal the health patterns in the healthy population through comparative PCA learning, and further analyze the differences between the health patterns; Step 5: Select the preprocessed test set samples, establish a health index through the statistical model in Step 3, perform health prediction on the test set samples, and diagnose and identify the corresponding targets through the contribution of each feature to the health index; The said Step 3 includes: Step 3.1: Compose the preprocessed training set sample data into an OTU matrix. Assume the OTU matrix consists of samples of microbial features. The PCA model structure is as follows: (2) Among them is the score matrix, is the loading matrix, is the retained principal component dimension, is the residual matrix, obtained by performing eigen-decomposition of the covariance matrix as follows: (3) wherein is the residual load, and are the eigenvalues of the principal component and the residual subspace; Therefore, the principal component subspace PCS and the residual subspace RS of the data are defined as: (4) (5) wherein and are the projection matrices of the main component subspace and the residual subspace; Assume that the eigenvalues are arranged in descending order , the percentage of explained variance (PEV) for each eigenvalue is defined as , and the correct number of principal components (PCs) is determined by the cumulative PEV satisfaction result; Step 3.2: Design charts, graphs, and diagrams according to the determined number of the principal components PC; The chart is: Given a sample , monitor the principal component subspace, which is defined as , where , at the confidence level the control limits are determined by the chi-square distribution , and the principal component dimension d is the degree of freedom; The chart is as follows: monitoring residual subspace, which is defined as , and the control limit is , where , calculated using eigenvalues; The said The chart is as follows: The combined index is defined as , where , and the control limit is , where and ; Step 3.3: Convert exponents, exponent sums into a general quadratic form , where the M of each exponent is respectively represented as the above-mentioned D, and , and if any exponent exceeds the corresponding threshold, report an unhealthy situation; The process of diagnosing the contribution of each feature to the health index in Step 5 includes: According to the health index, use the BHC graph for contribution diagnosis, and reconstruct the normal state by adding a correction term to the unhealthy components; Given an unhealthy sample , assuming that the feature i has potential abnormal behavior, the reconstructed constituent units are: (6) wherein is the direction, is the amplitude of change; Then formulate the objective to optimize the health index: (7) Among them, M represents the above-mentioned D, and ; By taking the first derivative equal to zero, we get: (8) Among them represents the feature i 's contribution to the health index.
2. The intestinal health status monitoring method according to claim 1, characterized in that The feature selection in Step 2 is: Based on the obtained relative abundance data of gut microbiota in population samples, use the Kolmogorov-Smirnov test for samples labeled as healthy and unhealthy to identify health prevalence and health scarcity features. Among them, the health prevalence feature is defined by rejecting another hypothesis, that is, at the significance level p, the empirical cumulative density of the feature in the healthy population is less than that in the unhealthy population, and the health scarcity feature is also defined by rejecting another hypothesis, that is, at the significance level p, the empirical cumulative density of the feature in the unhealthy population is less than that in the healthy population.
3. The method for monitoring the intestinal health status according to claim 1, wherein The data standardization in Step 2 includes: Apply the following logarithmic transformation to the relative abundance data: (1) wherein is a minimum value, and the logarithm-transformed values are z-score standardized to adjust the mean to 0 and the standard deviation to 1.
4. The method for monitoring the intestinal health status according to claim 1, wherein The process of comparative PCA learning in Step 4 includes: Perform comparative PCA learning on the healthy population reported by the health index in the training set and the unhealthy population in the test set; Seek the contrast direction with weights between high target healthy variance and low unhealthy variance by comparing PCA , where , and represent the covariance matrices of the healthy and unhealthy populations described above, respectively, represents the contrast parameter used to quantify the weights between high target healthy variance and low unhealthy variance.
5. The method for monitoring the intestinal health status according to claim 1, wherein The differences between the health patterns in Step 4 include: 1) Differences in the species composition of health patterns; 2) Differences in the geographical distribution of health patterns; 3) Differences in the metabolic systems of health patterns: First, map the OTUs at the genus level or lower levels to the genome-scale metabolic models GSMMs, and standardize the reactions according to the OTU abundances in the samples; Then, perform a two-sample t-test on the reaction abundances of each pair of health patterns to determine the reactions with significantly different abundances between groups, extract the metabolic subsystems of the reactions from the GSMMs, and perform enrichment analysis on the subsystems using the Fisher exact test; Finally, calculate the average abundance differences of each subsystem in each pair of health patterns for comparative analysis.
6. The method for monitoring the intestinal health status according to claim 2, wherein The significance level p = 0.
001.
7. The method for monitoring the intestinal health status according to claim 3, characterized in that, The minimum value Take 10 -5 .
8. A computer-readable storage medium storing computer-executable instructions that, when executed by a processor, implement the method for monitoring the intestinal health status according to any one of claims 1-7.
Citation Information
Patent Citations
Method for predicting healthy aging through relative abundance of intestinal flora
CN113186310A