Macro-genome queue matching method based on microbial metabolic background

By employing a metagenomic cohort matching method based on microbial metabolic background, the challenge of cohort construction in microbial research has been solved, and the accuracy of identifying causal relationships between microorganisms and diseases and the effectiveness of cohort matching have been improved while controlling for the influence of host confounding factors.

CN116864004BActive Publication Date: 2026-04-14THE CHILDRENS HOSPITAL ZHEJIANG UNIV SCHOOL OF MEDICINE
View PDF 0 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2023-06-25
Publication Date
2026-04-14

AI Technical Summary

Technical Problem

In microbial research, existing technologies struggle to effectively construct matched cohorts of case and control samples. In particular, as the number of clinical variables increases, the difficulty of sampling for matched samples rises exponentially, and individual differences in microbial communities mask the true causal relationship between microbes and diseases.

Method used

A metagenomic cohort matching method based on microbial metabolic background was used, including metagenomic sequencing data processing, extraction of major microbial metabolic background, nearest neighbor matching algorithm, and matching effect check, to construct matched disease group and control group cohorts. Principal component analysis and propensity score matching were used to control for the influence of host confounding factors.

Benefits of technology

This approach enables the acquisition of matched disease and control microbial cohorts while controlling for host confounding factors, enhancing causal relationship identification in metagenomics research, reducing false positive rates, and improving the accuracy of differentially expressed bacterial species identification.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN116864004B_ABST
    Figure CN116864004B_ABST
Patent Text Reader

Abstract

The application discloses a method for matching a metagenome queue based on a microbial metabolic background and relates to the technical field of metagenomics.S1: metagenome sequencing data processing; standardizing processing of data from whole metagenomics and manually input meta data;S2: extraction of a main metabolic background of microorganisms;S3: matching of the metabolic background of microorganisms; firstly, a nearest neighbor matching algorithm is used to screen matched samples in a control group without missing any main metabolic components;S4: matching effect checking; balance test is conducted on the mean of covariants of the matched disease group and the control group;S5: difference analysis based on the matched queue; if the data of the matched disease group and the control group meet normal distribution, paired sample t test is conducted for difference analysis, otherwise, group wilcoxon test is used for difference analysis.Through the technical method, a matched queue of case and control samples in microbial research is constructed, and the causal relationship identification capability of metagenomics research is strengthened.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of metagenomics, and more specifically, to a metagenomic cohort matching method based on microbial metabolic background. Background Technology

[0002] It is well known that the human body contains a highly diverse and rich microbiome that coordinates a comprehensive interplay of physiological processes and disease susceptibility, playing a vital role in human health and disease. Extensive evidence suggests that, compared to human genetic factors, the microbiome can explain a higher proportion of phenotypic variation in specific diseases within a population, thus serving as a novel biomarker for disease diagnosis or treatment.

[0003] However, the human microbiome is highly individualized, and the correlations between different microbial communities among individuals can be obscured by the uniqueness of the individual's microbiome. Studies have shown that the composition and function of the microbial community differ among individuals. These differences are influenced by multiple factors, including genetics, environment, lifestyle, diet, and age—i.e., host variables—and may mask the true causal relationship between the microbiome and disease. Currently, the construction of matched cohorts is the main way to explore the causal association between microbes and disease. For example, host confounding factors between cohorts can be controlled by using twins or relatives, or by using long-term sampling from the same individual to form a self-control. However, adhering to the principle of matching clinical variables in actual cohort construction is very challenging because the difficulty of sampling matching samples increases exponentially with the number of clinical variables.

[0004] Previous studies have shown that the composition and abundance of the microbiome are strictly constrained by the entire metabolic network, and the core metabolic functions of the microbiome are stable across individuals. Therefore, the “microbial metabolic background” is a better microbial baseline for individual sample matching. Summary of the Invention

[0005] The purpose of this invention is to provide a metagenomic cohort matching method based on microbial metabolic background, so as to obtain matched case and control samples in microbial research, thereby strengthening the causal relationship in metagenomics research.

[0006] The above-mentioned technical objective of the present invention is achieved through the following technical solution: a metagenomic cohort matching method based on microbial metabolic background, comprising the following steps and methods:

[0007] S1: Metagenomic sequencing data processing; standardization of data from whole metagenomics and manually entered meta-data; including data generated by MetaPhAn3 and data generated by HUMAnN3 from the UniRef90 database;

[0008] S2: Extraction of the main metabolic background of microorganisms; First, input the abundance data of microbial metabolic pathways as the metabolic background, and then use principal component analysis to reduce the dimensionality of the high-dimensional data. Within an acceptable range of information loss, retain the most important features, namely the main metabolic background.

[0009] S3: Microbial metabolic background matching; First, the nearest neighbor matching algorithm is used to screen samples in the control group that match the disease group without missing any major metabolic components;

[0010] S4: Matching effect check; perform a balance test on the data distribution and covariate means of the matched disease group and control group;

[0011] S5: Differential analysis based on matched cohorts; if the data of the matched disease group and control group conform to a normal distribution, then a paired-samples t-test is performed for differential analysis; otherwise, an independent group Wilcoxon test is used to compare the true degree of difference between the disease and control groups.

[0012] The present invention is further configured such that, in step S1, the data generated by MetaPhAn3 includes: ① a species-level taxonomic overview, representing the relative abundance from kingdom to species; ② the existence of unique branch-specific markers; and ③ the abundance of unique branch-specific markers.

[0013] The present invention is further configured such that, in step S1, the data generated by HUMAnN3 from the UniRef90 database includes: ① gene abundance; ② metabolic pathway coverage; ③ metabolic pathway abundance.

[0014] In summary, this invention offers the following advantages: Constructing actual cohorts in microbial research presents significant challenges because the difficulty of sampling matching samples increases exponentially with the number of clinical variables. Since core microbial metabolic functions are relatively stable among individuals, this invention uses microbial metabolic background as a matching benchmark. This allows for the acquisition of matched disease and control group microbial research cohorts while controlling for the influence of host confounding factors, thereby strengthening the understanding of causal relationships in metagenomics research. Attached Figure Description

[0015] Figure 1 This is a detailed flowchart of an embodiment of the present invention;

[0016] Figure 2 This is a schematic diagram of the microbial metabolic background sample matching method according to an embodiment of the present invention;

[0017] Figure 3 This is a simulation of microbial research data that was not matched in the embodiments of the present invention;

[0018] Figure 4This represents the probability of identifying differentially expressed bacterial species in unmatched and matched queues under different disease weights in embodiments of the present invention.

[0019] Figure 5 This is a simulation of the metabolic background in an embodiment of the present invention when it is unaffected by disease;

[0020] Figure 6 This is a simulation of the metabolic background being affected by disease in an embodiment of the present invention;

[0021] Figure 7 These are the data matching results (A) distribution histogram and (B) LOVE plot of inflammatory bowel disease according to embodiments of the present invention.

[0022] Figure 8 These are the results of region, age, and BMI before and after matching for inflammatory bowel disease in this embodiment of the invention. Detailed Implementation

[0023] The following is in conjunction with the appendix Figure 1-8 The present invention will be described in further detail below.

[0024] Example: A metagenomic cohort matching method based on microbial metabolic background. First, principal component analysis (PCA) is used to extract major metabolic components from the microbial metabolic pathways of the original mismatched cohort. Then, propensity score matching (PSM) is performed on the extracted major metabolic components to select the control sample that is closest to the given case sample for matching. Finally, a matching cohort is constructed.

[0025] like Figure 1 , Figure 2 As shown, the process includes the following steps:

[0026] 1. Metagenomic sequencing data processing

[0027] Standardization was performed on data from metagenomics and manually entered meta-data, including:

[0028] Generated by MetaPhAn3: ① Species-level taxonomic overview, representing the relative abundance from kingdom to species (relative_abundance); ② Presence of unique clade-specific markers (marker_presence); ③ Abundance of unique clade-specific markers (marker_abundance);

[0029] Generated from HUMAnN3 in the UniRef90 database: ① gene abundance (gene_families); ② pathway coverage (pathway_coverage); ③ pathway abundance (pathway_abundance).

[0030] 2. Extraction of major microbial metabolic background

[0031] Host variables (meta), relative abundance of microbial species (relative_abundance), and pathway abundance of microbial metabolic pathways (pathway_abundance) files are the most commonly used metagenomic data in microbial research. First, microbial metabolic pathway abundance data is input as the metabolic background. Then, principal component analysis is used to reduce the dimensionality of the high-dimensional data, retaining the most important features—the main metabolic background—within an acceptable range of information loss.

[0032] 3. Microbial metabolic background matching

[0033] First, the nearest neighbor (NN) matching algorithm is used to screen potentially matching samples from the control group without omitting any major metabolic components. The propensity score is calculated using a generalized linear model (GLM), and the maximum distance of the score is set by calipers, which restricts the matching of NNs to a range limited to this range.

[0034] The ratio is used to match the number of control units. The default value for the ratio parameter is 1, which means that the control sample with the propensity score closest to the given disease sample is selected to form a matched disease-control cohort with the given disease sample.

[0035] 4. Matching effect check

[0036] A balance test is performed on the covariate means of the matched disease and control groups. If the covariate data distributions of the matched disease and control groups are similar, it can be considered "data balanced." This typically requires the standardized mean difference (SMD) to be significantly smaller than before matching; otherwise, the caliper and ratio need to be readjusted to obtain an accurate matching cohort.

[0037] 5. Difference analysis based on matching queues

[0038] If the matched disease group and control group data conform to a normal distribution, a paired-samples t-test is performed for difference analysis; otherwise, an independent group Wilcoxon test is used to compare the true degree of difference between the disease and control groups.

[0039] The following detailed description of the invention uses simulated data, and verifies its superiority using real-world data. Simulated microbial research data is used to explore the accuracy of the invention in identifying differentially expressed bacterial species (see...). Figure 3 ).

[0040] To verify the feasibility and scope of application of this invention, simulated data from 100 cases in the disease group and 100 cases in the control group were randomly generated. The main metabolic backgrounds included age-driven, sex-driven, and environment-driven metabolic backgrounds. The abundance of the main metabolic backgrounds all conformed to a normal distribution N(μ,σ). 2 ), where μ is the expected value and σ is the variance.

[0041] Subsequently, by setting weights for the primary metabolic background and the disease, differentially expressed microbial species related to the disease are generated, along with other species unrelated to the disease. The microbial species are a linear combination of the primary metabolic background and the disease weight.

[0042] In this example, when the disease weight is greater than or equal to 0.1, the corresponding bacterial species are defined as disease-related differentially expressed bacterial species, and when the disease weight is less than or equal to 0.001, they are defined as other bacterial species. Fifty differentially expressed bacterial species and 50 other bacterial species were randomly generated.

[0043] like Figure 4 As shown, unmatched queues may incorrectly identify other bacterial species as differentially expressed species, while queues matched by this invention can effectively avoid false positives.

[0044] Since the differentially expressed bacterial species in the simulated data are known, the performance of this invention will be evaluated by studying the accuracy of identifying differentially expressed bacterial species before and after matching. The accuracy is calculated using a confusion matrix: Accuracy = (TP + TN) / (TP + FP + TN + FN), where TP is the true positive, TN is the true negative, FP is the false positive, and FN is the false negative.

[0045] The performance of this invention in identifying differentially expressed bacterial species was compared under two different conditions.

[0046] Assuming the metabolic background is unaffected by disease (see...) Figure 5 The difference in metabolic background μ between the disease group and the control group was gradually increased from 0 to 0.2 in increments of 0.02. When there was no difference or a very small difference between the disease and control groups, the original cohorts were considered to be matched. Therefore, the accuracy rate decreased slightly after matching using this invention, which is believed to be due to the loss of some data during matching. However, when there was a difference in metabolic background between the disease and control groups, the accuracy rate of identifying differentially expressed bacteria in the unmatched cohorts dropped rapidly. As the difference increased, the accuracy rate after matching decreased, but remained relatively stable and significantly higher than that of the unmatched study cohorts.

[0047] Assuming that the metabolic background is not independent and is affected by disease (see...) Figure 6 The metabolic background difference between the disease group and the control group was fixed at 0.1, and the influence coefficient ω of the disease on the metabolic background was gradually increased from 0 to 0.1 in increments of 0.02. The results showed that as the influence of the disease on the metabolic background increased, the accuracy of identifying differentially expressed bacteria in the unmatched group decreased rapidly until all bacteria were misidentified as differentially expressed bacteria; while the accuracy of the matched group decreased, it tended to stabilize overall and still had a high accuracy (over 80%), which was significantly higher than that of the unmatched group.

[0048] To avoid the influence of special datasets during the randomization process, 50 sets of data were randomly generated for each case. The average accuracy of the 50 simulations was calculated to compare the ability of the unmatched group and the matched group to identify differentially expressed bacterial species.

[0049] Examples of microbial research in inflammatory bowel disease (IBD).

[0050] Data Processing. Metagenomic data from five IBD studies were obtained from the CuratedMetagenomicData (version 3.4.2) database, including 1730 IBD cases and 796 control cases. Relative abundance data of microorganisms from kingdom to species were generated using MetaPhAn3, and pathway abundance data were generated using HUMAnN3 from the UniRef90 database.

[0051] The matching queue is obtained using the above technical methods, and the selection of calipers and ratios is shown in the table below:

[0052]

[0053] Obtaining the matching queue: After matching using this invention, 2023 matched IBD case studies and 2023 control cases are generated. The histogram of the matching effect distribution for IjazUZ_2017 is shown below. Figure 7 As shown.

[0054] Results Explanation: For example Figure 8 As shown, after matching using this invention, the impact of regional differences on microorganisms was balanced, the age difference between IBD and the control group was significantly reduced, while the difference in BMI disappeared significantly. This avoids spurious associations between host variables and microorganisms caused by confounding variables.

[0055] This specific embodiment is merely an explanation of the present invention and is not intended to limit the invention. After reading this specification, those skilled in the art can make modifications to this embodiment without contributing any inventive step, but such modifications are protected by patent law as long as they are within the scope of the claims of the present invention.

Claims

1. A metagenomic cohort matching method based on microbial metabolic background, characterized by: The steps and methods include the following: S1: Metagenomic sequencing data processing; standardization of data from whole metagenomics and manually entered meta-data, including data generated by MetaPhlAn3 and data generated by HUMAnN3 from the UniRef90 database; The data generated by MetaPhlAn3 includes: ① species-level taxonomic overview, representing relative abundance from kingdom to species; ② the presence of unique clade-specific markers; ③ the abundance of unique clade-specific markers. The data generated from the HUMAnN 3 database of UniRef90 includes: ① gene abundance; ② metabolic pathway coverage; ③ metabolic pathway abundance; S2: Extraction of the main metabolic background of microorganisms; First, input the abundance data of microbial metabolic pathways as the metabolic background, and then use principal component analysis to reduce the dimensionality of the high-dimensional data. Within an acceptable range of information loss, retain the most important features, namely the main metabolic background. S3: Microbial metabolic background matching; First, the nearest neighbor matching algorithm is used to screen samples in the control group that match the disease group without missing any major metabolic components; S4: Matching effect check; perform a balance test on the data distribution and mean of metabolic background covariates of the matched disease group and control group; S5: Differential analysis based on matched cohorts; if the data of the matched disease group and control group conform to a normal distribution, a paired-samples t-test is performed for differential analysis; otherwise, an independent group Wilcoxon test is used to compare the true degree of difference between the disease and control groups.