Time sequence analysis method and device for intestinal flora abundance

By analyzing the time series of gut microbiota abundance, this study solves the problem of existing technologies being unable to effectively capture dynamic patterns, enabling accurate prediction of microbiota dynamic characteristics and identification of key bacterial interactions, thus improving the predictive ability of microbiota behavior.

CN121838879APending Publication Date: 2026-04-10PEOPLES HOSPITAL PEKING UNIV +1
View PDF 0 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-12-19
Publication Date
2026-04-10

AI Technical Summary

Technical Problem

Existing technologies cannot effectively capture and quantify the dynamic patterns of gut microbiota over time, and have failed to fully explore the temporal interactions between bacterial species, resulting in insufficient ability to predict and measure the dynamic behavior of the microbiota.

Method used

We employed time-series analysis methods targeting gut microbiota abundance, and identified trends and fluctuations in species abundance through temporal trend and volatility analysis, temporal association analysis, and the construction of species co-occurrence networks. We also characterized the interactions between different species.

Benefits of technology

It enables accurate prediction of the dynamic characteristics of gut microbiota, identifies key bacterial species and shared interactions, and improves the understanding and prediction capabilities of microbiota behavior.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121838879A_ABST
    Figure CN121838879A_ABST
Patent Text Reader

Abstract

The invention provides a time sequence analysis method and device for intestinal flora abundance, belongs to the technical field of intestinal microbiological data analysis, and solves the problem that characteristics such as trend and fluctuation, time sequence association and time sequence interaction in an intestinal flora abundance time sequence with limited samples and abnormal distribution are difficult to effectively extract in the prior art. According to the time sequence analysis method and device for the intestinal flora abundance, the trend and fluctuation of flora change can be accurately identified according to dynamic data of the intestinal flora abundance along with time. The flora time sequence characteristics identified based on the method provided by the invention can be used as the basis of clinical practices such as disease activity prediction and treatment response.
Need to check novelty before this filing date? Find Prior Art

Description

TECHNICAL FIELD

[0001] The present application provides a time series analysis method and device for intestinal flora abundance, belonging to the technical field of intestinal microbiology data analysis, and is used for solving the problem of extracting time sequence characteristics of intestinal flora. BACKGROUND

[0002] Intestinal flora abundance is a key biological indicator reflecting human metabolism, immunity and disease state. Traditional flora abundance analysis technology mainly relies on static measurement of single time point samples, and describes the microbial composition at a specific moment through cross-sectional study. However, the human health state and the flora itself are in a continuous evolution dynamic process, and this static snapshot method cannot effectively capture and quantify the dynamic pattern of the evolution of the flora over time.

[0003] In recent years, with the progress of time sequence sampling technology, the research paradigm of analyzing the dynamic changes of flora abundance by collecting flora samples at multiple consecutive time points to construct time series data has gradually been valued. However, in the field of flora abundance prediction and measurement technology, the existing methods still have obvious deficiencies. On the one hand, most technologies can only perform basic descriptive statistics or difference comparison on time series data, and lack special analysis methods that can systematically extract and quantify the time sequence dynamic characteristics (such as long-term change trend, periodic fluctuation, time sequence dependence, etc.) of flora abundance. On the other hand, the existing methods usually isolate the abundance changes of a single species, and fail to effectively model the complex interactive relationships between different species in the time dimension, such as cooperation, competition or co-occurrence, resulting in that a large amount of key information contained in the time series data is not fully explored and utilized.

[0004] Therefore, there is an urgent need in the art to develop an analysis method that can deeply analyze the time series of flora abundance, comprehensively quantify its dynamic characteristics and reveal the time sequence interaction between species, so as to improve the prediction and measurement ability of the dynamic behavior of flora. SUMMARY

[0005] In view of the above problems, the present application provides a time series analysis method and device for intestinal flora abundance, which can accurately identify the trend and volatility of flora changes according to the dynamic data of intestinal flora abundance over time.

[0006] The present application provides a time series analysis method for intestinal flora abundance, comprising the following steps:

[0007] Obtaining the time series of the relative abundance of each species in the fecal sample;

[0008] Performing time sequence trend and volatility analysis on the time series to identify the trend and volatility characteristics of the abundance of the species over time;

[0009] The time series and the index parameter check time series are subjected to time series correlation analysis to examine the time series correlation between the bacterial species abundance time series and the index parameter check time series.

[0010] The bacterial species time series are subjected to bacterial species time series interaction analysis, the interaction between different bacterial species is characterized by constructing a bacterial species co-occurrence network, and the key bacterial species and network similarity in the co-occurrence network are analyzed.

[0011] Optionally, the time series are subjected to time series trend analysis according to the statistical quantity of the bacterial species abundance sequence.

[0012] Optionally, the time series are subjected to volatility analysis according to the root mean square first difference of the bacterial species abundance time series.

[0013] Optionally, the time series correlation analysis adopts Spearman rank correlation test to analyze the time series correlation between the time series and the index parameter check time series.

[0014] Optionally, the bacterial species co-occurrence network is constructed, the proximity centrality of the nodes in the bacterial species co-occurrence network is analyzed, and the key bacterial species are determined by evaluating the proximity centrality of each bacterial species.

[0015] Optionally, the proximity centrality c a of node a is expressed as:

[0016]

[0017] wherein, is the number of nodes in the maximum subgraph G * of the bacterial species co-occurrence network G; b is a node in the maximum subgraph G * , b≠a; d(b,a) is the shortest path between node a and k.

[0018] Optionally, the bacterial species co-occurrence network is constructed, and the network similarity is analyzed according to the bacterial species co-occurrence network.

[0019] Optionally, the network similarity J(G p ,G q ) of the pth co-occurrence network G p and the pth co-occurrence network G q is expressed as:

[0020] J(G p ,G q )=|E p ∩E q | / |E p ∪E q |

[0021] wherein, E p and E qThe pth co-occurrence network G p The pth co-occurrence network G q The edge set of the pth co-occurrence network G

[0022] Optionally, the index parameter inspection time series is a clinical inspection series synchronized with the bacterial abundance time series.

[0023] Another aspect of the present application also discloses a time series analysis device for intestinal flora abundance, characterized in that the time series analysis method described above is used to analyze the time series of intestinal flora abundance, comprising:

[0024] An abundance time series preprocessing module is configured to extract the flora abundance time series data from the fecal sample of the subject;

[0025] A trend / fluctuation feature analysis module is configured to identify the time series trend and fluctuation feature of each bacterial abundance of the subject over time;

[0026] A time series correlation analysis module is configured to verify the time series correlation between the bacterial abundance time series and the clinical inspection time series;

[0027] A bacterial interaction analysis module is configured to identify the interaction between different bacteria in the flora by constructing a bacterial co-occurrence network, and to identify key bacteria.

[0028] Compared with the prior art, the present application has at least the following beneficial effects:

[0029] (1) For the time series of intestinal flora abundance of limited samples, the present application uses a non-parametric detection method to reasonably identify the abundance trend, and uses a first-order difference to measure the difference between samples at different time points to effectively recognize the sequence fluctuation, thereby realizing accurate prediction of dynamic characteristics.

[0030] (2) The present application constructs a bacterial co-occurrence network to represent the time series interaction between different bacteria. Considering the combination and sparsity of bacterial abundance data, the SpiecEasi-MB method is used to construct the network. On this basis, the present application also uses node centrality and network similarity to identify key bacteria and common interactions in the intestinal flora.

[0031] (3) The present application uses a non-parametric Spearman rank correlation test to verify the dynamic collaborative change of bacterial abundance and host state and time without assuming normal distribution of the abundance sequence. BRIEF DESCRIPTION OF DRAWINGS

[0032] Figure 1 The flowchart of the time series analysis method for intestinal flora abundance of the present application;

[0033] Figure 2Species with significant trend for each subject (p<0.05);

[0034] Figure 3 Comparison of the overall abundance of gut microbiota between groups;

[0035] Figure 4 Species with the most significant abundance change in each group of subjects;

[0036] Figure 5 Comparison of the abundance of each species between groups;

[0037] Figure 6 Comparison of the number of gut microbiota samples and CRP samples for each subject;

[0038] Figure 7 Species with significant correlation with CRP level for each subject;

[0039] Figure 8 Species co-occurrence network constructed in R, mA and R2A groups;

[0040] Figure 9 Comparison of the centrality of each species between groups.

[0041] Figure 10 Comparison of network similarity between groups;

[0042] Figure 11 Species interaction shared by ≥75% of subjects in each group.

[0043] It should be noted that, Figure 5 and Figure 9 The horizontal axis of the volcano plot is log2(FC), where FC represents the ratio of the median of the two groups compared. Figure 5 The root mean square first difference in Figure 9 The node closeness centrality in DETAILED DESCRIPTION

[0044] In order to more clearly understand the above-mentioned purposes, features and advantages of the present application, the present application will be further described in detail below in conjunction with the drawings and specific embodiments. It should be noted that the embodiments of the present application and the features in the embodiments can be combined with each other without conflict. In addition, the present application can also be implemented in other ways different from those described herein, therefore, the protection scope of the present application is not limited by the specific embodiments disclosed below.

[0045] One specific embodiment of the present application, such as Figures 1-11 discloses a time series analysis method for gut microbiota abundance, the specific steps are as follows:

[0046] Step 1. Obtain the time series of relative abundance of each bacterial species in the fecal sample.

[0047] Optionally, fecal samples of the subject are first collected over a period of time (stored at -80°C). The sampling interval and total period are adjusted according to factors such as disease and treatment progress (e.g., the sampling interval is not less than 1 month / time, and the total period is not less than 1 year). For the collected samples, the following operations are performed:

[0048] Step 11. Extract DNA from the fecal sample;

[0049] Specifically, a fecal DNA extraction kit based on the purification principle of silica gel membrane is used to extract DNA from the fecal sample. The quality of the extracted DNA is detected, including concentration quantification using a fluorescence meter, and evaluation of its integrity and purity by 1.2% agarose gel electrophoresis.

[0050] Step 12. Sequencing library construction and sequencing to obtain raw sequencing data;

[0051] Specifically, a high-throughput sequencing library construction DNA extraction kit with double-stranded DNA adapter is used to prepare a sequencing library, and a magnetic bead solid-phase reversible immobilization (SPRI) technique is used to purify the sequencing library.

[0052] Further, the purified library is accurately quantified based on the principle of real-time fluorescence quantitative polymerase chain reaction (qPCR). The quantified library is subjected to double-end sequencing on a high-throughput sequencing platform to generate raw sequencing data.

[0053] Step 13. Bioinformatics processing;

[0054] Specifically, for the above raw sequencing data, sample demultiplexing is performed, and the raw data is evaluated for quality using a sequence quality assessment tool. A sequence trimming tool is applied to perform quality control steps, removing bases with a quality score below a set threshold (e.g., Phred quality score Q<20), and trimming the end regions of reads with continuous low-quality (e.g., Q<15) bases longer than 5 bp. Finally, reads with an average quality score Q≥20 and a length ≥50 bp are retained, and a set of high-quality sequencing reads is obtained.

[0055] Step 14. Use a taxonomic annotation tool based on k-mer algorithm to identify species in the set of high-quality sequencing reads to obtain the classification of bacterial species in the fecal sample.

[0056] Step 15. Based on the classification results of bacterial species in the fecal sample, calculate the relative abundance of each classification unit (genus, species level) of bacterial species in the fecal sample. The sampling results at different times in the time series are integrated to finally obtain the time series of the relative abundance of each bacterial species.

[0057] Step 2 time trend analysis and volatility analysis

[0058] Step 21 time trend analysis

[0059] The present application adopts a non-parametric test method to identify the trend of the abundance of bacterial species over time, compares the differences between the first detected sample and the later detected sample, does not rely on the assumptions of stationary process and normal volatility, and has good robustness when analyzing the volatility sequence of limited samples. The problem of extracting the time series characteristics of the intestinal flora abundance time series with non-stationary and non-normal volatility by conventional time series analysis methods such as regression analysis or stochastic differential equation is solved.

[0060] Specifically, the statistical quantity S i of the abundance sequence of the i-th bacterial species is obtained

[0061]

[0062] wherein, x represents the t-th sample in the time sequence of the i-th bacterial species, t = 1, 2, …, n-1, x represents the g-th sample in the time sequence of the i-th bacterial species, g = t+1, 2, …, n, n is the total number of samples; sign(.) represents the indicator function; the abundance sequence of the i-th bacterial species

[0063]

[0064] Specifically, if S i is a larger positive number (greater than a preset threshold), it indicates that the later detected samples are generally greater than the first detected samples, thus representing the trend of the abundance sequence of the i-th bacterial species x i increasing over time; on the contrary, S i is a larger negative number in absolute value, indicating that the abundance sequence of the i-th bacterial species x i decreases over time; if S i is close to 0, it indicates that x i has no obvious time trend.

[0065] Further, based on the statistical quantity S i of the abundance sequence of the i-th bacterial species, the regularized statistical quantity τ i of the i-th bacterial species is obtained, which is used to quantify the direction and strength of the time sequence trend, and the expression is:

[0066] In the formula, τ i ∈[-1, 1].

[0067] Step 22 time trend analysis and volatility analysis

[0068] This invention uses first-order difference as the quantization method for volatility, which improves the ability to consider time sequence when quantifying time-series volatility.

[0069] Specifically, obtain the time series x of the abundance of the i-th bacterial species. i The root mean square first difference Dx i The expression is:

[0070]

[0071] in, and Let represent the t-th and t+1-th samples in the time series of the i-th bacterial species, respectively; For the t+1 sample in the time series of the i-th bacterial species The t-th sample in the time series of the i-th bacterial species The first difference between them; n i Let x be the abundance time series of the i-th bacterial species. i The total number of samples.

[0072] Furthermore, based on the root mean square first difference Dx of each bacterial species i The mean values ​​of each bacterial species are obtained to assess the fluctuation of the overall bacterial community abundance. The expression is as follows:

[0073] Dx = <Dx i > (4)

[0074] This invention uses a combination of first-order differences and the aforementioned trend test results to assess the volatility of species abundance. For example, when a species has a large first-order difference but no significant trend, it indicates that the abundance of that species fluctuates significantly.

[0075] Step 3: Temporal correlation analysis;

[0076] When identifying temporal correlations, this invention uses the nonparametric Spearman rank correlation test to evaluate temporal correlations, overcoming the problem of distorted calculation of significance level p-values ​​and incorrect acceptance / rejection of the null hypothesis caused by the traditional Pearson correlation coefficient method when the abundance data does not follow a normal distribution.

[0077] Specifically, the abundance time series x of the i-th bacterial species i And concurrent clinical testing sequences (such as C-reactive protein (CRP) testing sequences). Converted into the abundance sequence rank vector R(x) of the bacterial species respectively i ) and the rank vector R(y) of the clinical examination sequence i ).

[0078] Furthermore, based on the abundance sequence rank vector R(x) of the bacterial species i) and clinical examination sequence rank vector R(y i ) to calculate rank correlation coefficient r s , the expression is:

[0079]

[0080] Wherein, represents the t-th sample in the clinical examination sequence y i synchronized with the i-th species; represents the rank in the time sequence x i ; represents the rank in the time sequence y i ; and are the mean values of the abundance sequence rank vector R(x i ) and the clinical examination sequence rank vector R(y i ) of the i-th species, respectively.

[0081] Further, the simplified expression of the rank correlation coefficient r s is:

[0082]

[0083] Further, the statistical significance of the rank correlation coefficient r s is judged, and the specific steps are:

[0084] First, the null hypothesis is that the species abundance and the clinical examination synchronized therewith have no time sequence correlation, and the alternative hypothesis is that the two have significant correlation. The statistical quantity approximately obeys the t distribution with degrees of freedom n-2 to calculate the significance level p value.

[0085] Then, if the test obtains a significance level p value less than the significance level requirement (such as 0.05), it indicates that there is a significant time sequence correlation between the i-th species abundance and the clinical examination synchronized therewith; if the significance level p value is greater than or equal to the significance level requirement (such as 0.05), it indicates that there is no significant time sequence correlation between the i-th species abundance and the clinical examination synchronized therewith.

[0086] Step 4: Analysis of species time sequence interaction;

[0087] The present application represents the interaction between different species (i.e. the interaction between different species) by constructing a species co-occurrence network, and analyzes the key species in the flora and the species interaction common in different populations according to the network node centrality and network similarity.

[0088] Step 41: Constructing a species co-occurrence network;

[0089] The present application combines the data conversion and graphical model inference framework of the combined data analysis, and assumes that the underlying strain co-occurrence network is sparse, relies on sparse neighborhood and sparse inverse covariance selection algorithms to reconstruct the strain co-occurrence network, which can avoid the false correlation between nodes caused by the combined data characteristics or the confounding bias between nodes, that is, there is no correlation between two strains, but the combined data characteristics or other strains confounding effects lead to false correlation.

[0090] Further, the co-occurrence network includes nodes, edge connections and the field of nodes, wherein each node corresponds to a strain. If there is a statistically significant correlation between the abundance of two strains, a connection between the corresponding two nodes will be connected by an edge.

[0091] Further, the SpiecEasi method infers the network structure based on conditional independence, that is, to identify which strains are connected based on strain abundance data and determine which nodes have edges, it first performs a centered log-ratio transformation on the strain abundance sequence to address the combined data bias.

[0092] On this basis, the Meinshausen-Bühlman (MB) method is used to determine the neighborhood of each node: for any node a, given the state of all nodes in the neighborhood ne(a) of the node a, the state of the node a is conditionally independent of the state of all other nodes (excluding ne(a)) in the co-occurrence network. The state of node a has an edge connection with all nodes in its neighborhood ne(a).

[0093] It can be understood that the conditional independence: given the state of all nodes in the neighborhood ne(a) of the node a, the state of the node a is independent (uncorrelated) of the state of all other nodes (excluding ne(a)) in the co-occurrence network.

[0094] Further, in order to determine the neighborhood ne(a) of node a, the MB method uses Lasso regression to solve the coefficient vector θ a of node a, and the expression is:

[0095]

[0096] where x a is the sample of node a (i.e., the time series of the abundance of the a-th strain), x k is the time series of the abundance of the k-th strain (k≠a); θ k is the regression coefficient corresponding to the k-th strain; is the distance measure.

[0097] Further, based on the coefficient vector θ a, estimate neighborhood ne(a), use the neighborhood to represent that there is a statistically significant correlation between the abundance of two species, expressed as:

[0098] ne(a) = {k e T \ {a}: 0 k ≠ 0} (8)

[0099] Wherein, T is the set of all nodes.

[0100] Step 42 analyzes the closeness centrality of the node;

[0101] The present application determines the positioning of the key species by evaluating the closeness centrality of each species, and the specific steps are as follows:

[0102] Further, the expression of the closeness centrality c a of the node a is as follows:

[0103]

[0104] Wherein, is the number of nodes in the maximum subgraph G * of the species co-occurrence network G; b is a node (b≠a) in the maximum subgraph G * ; d(b,a) is the shortest path between node a and node b.

[0105] Step 43 network similarity analysis;

[0106] When different species co-occurrence networks are constructed (for example, co-occurrence networks of different subjects, co-occurrence networks of different stages of the same subject, etc.), the similarity of two networks is compared based on network topology. When the network similarity is large, it indicates that the two co-occurrence networks may have more common species interactions and common rules; otherwise, it indicates that there are obvious differences or changes in the intestinal flora.

[0107] Specifically, the expression of the network similarity J(G p ,G q ) of the pth co-occurrence network G p and the pth co-occurrence network G q is as follows:

[0108] J(G p ,G q ) = |E p ∩E q | / |E p ∪E q | (10)

[0109] Wherein, E p and E q are the pth co-occurrence network G p and the pth co-occurrence network Gq The edge set.

[0110] Furthermore, based on network similarity J(G) p G q By comparing the similarity between different co-occurrence networks, it is possible to reasonably measure the differences in gut microbiota among different groups / stages, thereby supporting the prediction of disease progression. For example, if the co-occurrence network of a subject's microbiota is highly similar to that of a group of subjects with inflammatory bowel disease, this may indicate that the subject has a potential predisposition to the disease.

[0111] Another aspect of the present invention discloses a time series analysis device for gut microbiota abundance, which uses the aforementioned time series analysis method for gut microbiota abundance to perform time series analysis of gut microbiota abundance. The device includes an abundance time series preprocessing module for extracting microbiota abundance time series data from the subject's sample; a trend / fluctuation feature analysis module for identifying the time series trends and fluctuation characteristics of the abundance of each microbial species in the subject over time; a time series correlation analysis module for examining the time series correlation between microbial abundance time series and clinical examination time series (e.g., CRP sequence); and a microbial interaction analysis module for identifying interactions between different microbial species within the microbiota and identifying core microbial species and other characteristics.

[0112] The following analysis will focus on real sampling data from Crohn's disease subjects receiving biological agent treatment. Using the time series analysis method for gut microbiota abundance proposed in this invention, we will identify the temporal characteristics of the microbiota in subjects with different treatment responses, thereby providing a basis for predicting disease activity and treatment response.

[0113] Step 1: Preprocessing of gut microbiota abundance time series data;

[0114] This study included 16 Crohn's disease patients treated with biologics at the Department of Gastroenterology, Peking University People's Hospital, between 2020 and 2021. All patients received treatment for 0.5–5 years, and provided monthly stool samples for metagenomic profiling analysis during a 12-month observation period. The protocol was approved by the Institutional Review Committee of Peking University People's Hospital (ethics approval number: 2021PHB255), and all patients signed written informed consent forms before enrollment and sample collection. Based on treatment response assessed using the Simplified Crohn's Disease Endoscopic Score (SES-CD), patients were divided into three groups: persistent remission (R), mild activity (mA), and relapse (R2A). There were 8 patients in the R group, 4 in the mA group, and 4 in the R2A group. Finally, based on step 1 of the invention, time-series data (sequence length approximately 12) of different bacterial species abundance for each patient were obtained.

[0115] Step 2: Time series trend and volatility analysis;

[0116] 1. Temporal trend analysis

[0117] Firstly, for each subject and each bacterial species, Mann-Kendall test was used to identify the trend of the abundance over time. This example was analyzed for the bacterial species with average abundance greater than 0.001. Figure 2 The bacterial species with significant trend (p < 0.05) for each subject were shown.

[0118] From the trend characteristics of Figure 2 , we can know that:

[0119] 1) For the subjects in R group, some probiotic species showed increasing trend over time. For example, F. prausnitzii and Alistipes species for subject 2; Akkermansia muciniphila for subject 5; and Bifidobacterium bifidum, Bifidobacterium longum and Roseburia intestinalis for subject 12. In contrast, some Bacteroides, Lachnoclostridium and Lachnospiraceae species showed decreasing trend;

[0120] 2) For the subjects in mA and R2A groups, some pathogenic bacteria significantly increased over time, for example, Klebsiella pneumoniae for subjects 8 and 15, and Fusobacterium mortiferum for subjects 1 and 15. At the same time, some Bacteroides, Clostridium and Veillonella showed decreasing trend.

[0121] 2. Temporal fluctuation analysis

[0122] First order difference of each bacterial species abundance was calculated to assess the fluctuation. Firstly, the overall abundance change of the gut microbiota for each subject was assessed based on equation (4), the results were shown in Figure 3 . At the species level, the abundance change of each bacterial species was calculated based on equation (3), the bacterial species with top abundance change for each group of subjects were shown in Figure 4 .

[0123] From the fluctuation characteristics of Figure 3 and Figure 4 , we can know that:

[0124] 1) The overall abundance change of the gut microbiota for the subjects in R2A group was significantly greater than that for the subjects in R and mA groups (p < 0.05).

[0125] 0.05), while R did not differ significantly from mA group subjects;

[0126] 2) Some pathogenic bacteria had more significant changes in mA and R2A group subjects compared to R group subjects, including Ruminococcus gnavus, K. pneumoniae, and C. perfringens. Figure 4

[0127] On this basis, the Kruskal-Wallis H non-parametric test and Dunn post-hoc test were used to statistically analyze the group differences in the changes in bacterial species abundance, and the results are shown in Table 2. Figure 5

[0128] 1) Some bacterial species had more significant changes in abundance in mA group subjects compared to R group subjects, such as Sutterella megalosphaeroides, K. pneumoniae, Veillonella atypica, Veillonella dispar, and Clostridium butyricum.

[0129] 2) Similarly, Campylobacter concisus, K. pneumoniae, V. atypica, C. perfringens,

[0130] Klebsiella oxytoca, and Eikenella corrodens showed more significant changes in abundance in R2A group subjects compared to R group subjects.

[0131] 3) In contrast, Alistipes shahii and other bacterial species showed more significant changes in abundance in R group subjects.

[0132] Since the significant changes in the abundance of some bacterial species in some subjects may be mainly due to deterministic temporal trends, such as K. pneumoniae in mA and R2A group subjects and A. shahii in R group subjects. In addition to these bacterial species, significant changes in the abundance of other bacterial species indicate their strong volatility.

[0133] In summary, the method proposed by the present application shows that for R group subjects, many probiotic bacterial species have a growth trend, while in mA and R2A group subjects, some pathogenic bacteria significantly increase or have strong volatility.

[0134] Step 3 temporal correlation analysis;

[0135] ​​In this example, the subjects collected C-reactive protein (CRP) samples at the same time of collecting the fecal samples, which is used in clinic to characterize the degree of inflammation. Higher CRP content usually indicates more severe inflammation response in the subject. This example investigates the temporal correlation between the gut microbiota abundance and the CRP level. Note that there is some missing data in the CRP collection process, resulting in less CRP samples than gut microbiota samples. Figure 6 The number of gut microbiota samples and CRP samples for each subject was compared. Since subjects 6 and 11 have more missing CRP samples, this step was not included for subjects 6 and 11. For other subjects, the temporal correlation was analyzed for the species with average abundance greater than 0.001. Figure 7 The species with significant correlation for each subject were shown. It can be seen that:

[0136] 1) For subjects in R and mA groups, there is generally no correlation between the species abundance and CRP, which may be due to the low and stable CRP level of the subjects;

[0137] 2) For subjects in R2A group, there are some species with significant negative correlation between the species and CRP, which belong to typical probiotic genera, including Bifidobacterium, Roseburia, Blautia, and Megamonas.

[0138] In summary, according to the method proposed in the present application, the above-mentioned temporal correlation characteristics can be used as a clinical marker of inflammation response in Crohn's disease subjects, and can provide an indication for exploring the physiological mechanism between gut microbiota and inflammation response.

[0139] Step 4: Temporal interaction analysis of species;

[0140] 1. Network construction;

[0141] First, based on the species abundance sequence data of each subject, an individualized species co-occurrence network was established for each subject. In constructing the network, only species with average abundance greater than 0.0001 were included. Figure 8 The typical network diagram in the three groups was shown.

[0142] 2. Node centrality analysis;

[0143] According to formula (9), the node centrality of each species for each subject was calculated, and the Kruskal-Wallis H non-parametric test and Dunn post-hoc test were used to evaluate the group difference of the centrality of each species, and the results are shown in Figure 9 It can be seen that: Figure 9

[0144] ​1) Some pathogenic bacteria show significant centrality in the mA group, but weak influence in the R group, including Sutterella faecalis, V. atypica, Clostridium septicum, Enterobacter hormaechei, Shigella dysenteriae and Fusobacterium ulcerans;

[0145] 2) Similarly, some pathogenic bacteria show significantly greater centrality in the R2A group than in the R group, such as V. atypica, Morganella morganii, E. corrodens, Klebsiella michiganensis and K. oxytoca.

[0146] 3. Network similarity analysis;

[0147] Based on equation (10), the similarity between the species co-occurrence networks of two subjects in the same group was first compared, and the Kruskal-Wallis H non-parametric test and Dunn post-hoc test were also used to evaluate the group differences in similarity, as shown in Table 2. Figure 10 The results show that the network similarity of the R2A group is slightly greater than that of the R group (p = 0.074).

[0148] On this basis, the common species interactions of subjects in each group were extracted, as shown in Table 3. Figure 11 The results show that:

[0149] 1) The scale of common interactions in the R2A group is larger than that in the mA and R groups, which corresponds to the greater network similarity in the R2A group;

[0150] 2) The common interactions of Bacteroides species appear in all three groups, indicating that such interactions are prevalent in Crohn's disease subjects;

[0151] 3) The interactions of V. atypica and V. dispar exist in both the mA and R2A groups, so this interaction indicates the development of the disease process;

[0152] 4) The interaction of Alistipes megaguti and Alistipes dispar only appears in the R group, indicating that Alistipes-related interactions imply disease remission.

[0153] The above merely describes preferred specific embodiments of the present application, but the scope of protection of the present application is not limited thereto, and any changes or substitutions within the scope of the technology disclosed by the present application can be easily thought of by those skilled in the art, which should be covered within the scope of protection of the present application.

Claims

1. A method for time series analysis of gut microbiota abundance, characterized by, The method comprises the following steps: obtaining a time series of the relative abundance of each bacterial species in the fecal sample; performing time trend and volatility analysis on the time series to identify the time trend and volatility characteristics of the abundance of the bacterial species over time; performing time correlation analysis on the time series and the index parameter inspection time series to test the time correlation between the time series of the abundance of the bacterial species and the index parameter inspection time series; performing time interaction analysis on the time series of the bacterial species by constructing a bacterial co-occurrence network to represent the interaction between different bacterial species and analyzing the key bacterial species and network similarity in the co-occurrence network.

2. The time series analysis method according to claim 1, characterized by, Performing time trend analysis on the time series according to the statistical quantity of the abundance sequence of the bacterial species.

3. The time series analysis method according to claim 1, characterized by, Performing volatility analysis on the time series according to the root mean square first-order difference of the time series of the abundance of the bacterial species.

4. The time series analysis method according to claim 1, characterized by, The time correlation analysis uses Spearman rank correlation test to perform time correlation analysis on the time series and the index parameter inspection time series.

5. The time series analysis method of claim 1, wherein, Constructing a bacterial co-occurrence network, analyzing the closeness centrality of the nodes in the bacterial co-occurrence network, and determining the key bacterial species by evaluating the closeness centrality of each bacterial species.

6. The time series analysis method according to claim 5, characterized in that, The proximity centrality c of node a is given by the expression: a The proximity centrality c of node a is given by the expression: wherein, is the number of nodes in the largest subgraph G * of the species co-occurrence network G; b is the number of nodes in the largest subgraph G * , b≠a; d(b,a) is the shortest path between nodes a and k.

7. The time series analysis method of claim 1, wherein, Constructing a bacterial co-occurrence network and analyzing the network similarity according to the bacterial co-occurrence network.

8. The time series analysis method according to claim 7, characterized by, The pth co-occurrence network G p The pth co-occurrence network G q The expression of network similarity J(G p , G q ) is: J(G p ,G q ) = |E p ∩E q | / |E p ∪E q | wherein E p with E q is the edge set of the p-th co-occurrence network G p with the p-th co-occurrence network G q .

9. The time series analysis method according to any one of claims 1 to 8, characterized in that, The index parameter inspection time series is a clinical inspection sequence synchronized with the time series of the abundance of the bacterial species.

10. A device for time series analysis of gut microbiota abundance, characterized by, The method for time series analysis of the abundance of the intestinal flora comprises: an abundance time series preprocessing module for extracting the time series data of the abundance of the flora from the fecal sample of the subject; a trend / volatility characteristic analysis module for identifying the time trend and volatility characteristics of the abundance of each bacterial species of the subject over time; a time correlation analysis module for testing the time correlation between the time series of the abundance of the bacterial species and the time series of the clinical inspection; a bacterial interaction analysis module for identifying the interaction between different bacterial species within the flora by constructing a bacterial co-occurrence network and identifying key bacterial species.