A dynamic analysis method for multi-time sequence intestinal flora gene detection data

CN122531469APending Publication Date: 2026-08-07JIANGSU GUOYUAN ADVANCED INSTR TECH RES INST CO LTD
View PDF 2 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
JIANGSU GUOYUAN ADVANCED INSTR TECH RES INST CO LTD
Filing Date
2026-05-13
Publication Date
2026-08-07

AI Technical Summary

Technical Problem

[0004]上述现有技术各自解决了不同层面的技术问题,CN105046094A解决的是如何将个体与群体数据库进行横向比较的问题,CN112151117B解决的是如何在菌种层面捕捉菌群互作网络动态变化的问题,但两者均存在一个共同的且尚未被现有技术认识到并解决的技术缺陷:均无法对同一个体自身的宏基因组功能通路时间序列数据进行三分量(趋势分量、周期分量、余项分量)分解,并在功能通路层面基于个体自身历史基线实现偏离度的动态评估与偏离动力学类型的自动区分

Benefits of technology

[0049] I. This invention acquires metagenomic sequencing data from multiple time points of the same individual and constructs a functional pathway abundance matrix. It identifies the baseline data window in an individual's healthy state and decomposes the time series of each functional pathway into trend, periodic, and residual components. It then establishes individualized time-series baseline models for each component. This allows the reference system for dynamic assessment of gut microbiota function to shift from population statistical standards to the individual's own historical homeostasis. It enables precise quantification of multidimensional deviations of functional pathways and automatic differentiation of abnormal dynamic types. This effectively solves the technical defects of existing technologies that cannot distinguish functional pathway abnormalities of different causes and cannot judge the nature of deviations based on the individual's own homeostasis. It provides a stable and reliable quantitative analysis basis for individualized dynamic monitoring of gut microbiota function.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN122531469A_ABST
    Figure CN122531469A_ABST
Patent Text Reader

Abstract

The application discloses a kind of dynamic analysis methods of multiple time sequence intestinal flora gene detection data, it is related to bioinformatics and microorganism group gene detection data analysis technical field, the method first obtains the metagenome sequencing data of multiple time points stool sample of same individual, constructs functional pathway abundance matrix after completing functional annotation, identifies baseline period data window under individual health steady state, time series decomposition is carried out to the time series of each functional pathway in baseline period and establishes individualized time series baseline model, calculates each dimension deviation degree when new detection data arrives, completes abnormal type classification based on deviation degree combination mode and outputs analysis result.The application can be applied to the longitudinal dynamic analysis of intestinal flora metagenome sequencing data, provides accurate quantifiable analysis scheme that can be landed for personalized intestinal flora health monitoring.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of bioinformatics and microbiome gene detection data analysis technology, specifically a dynamic analysis method for multi-time series gut microbiota gene detection data. Background Technology

[0002] The gut microbiota is closely related to various human health conditions. With the advancement and cost reduction of next-generation sequencing technology, gut microbiota gene testing services based on metagenomic sequencing have gradually entered the commercial application stage. However, a single cross-sectional test can only reflect the static composition of the microbiota at a certain moment, making it difficult to capture the dynamic changes of the gut microbiota over time. Therefore, more and more research and commercial tests are beginning to adopt longitudinal designs with multi-timepoint sampling in order to obtain deeper information on the functional status of the microbiota through time-series data analysis. How to effectively analyze multi-timepoint gut microbiota metagenomic sequencing data and extract dynamic functional information with personalized clinical guidance value has become a technical problem that urgently needs to be solved in this field.

[0003] In the prior art, invention patent publication number CN105046094A discloses a gut microbiota detection system and method, and a dynamic database. Its technical solution compares the individual gut microbiota data collected each time with the total gut microbiota data in a storage device, allowing the examinee to understand the differences between themselves and healthy individuals. This method uses population statistics as the reference benchmark and does not utilize historical detection data from multiple time points within the same individual to construct an individualized reference system. Another invention patent publication number CN112151117B discloses a dynamic observation device and detection method based on time-series metagenomic data. It constructs a microbial interaction network through dynamic time warping and Pearson correlation analysis, and the network can dynamically change as new samples are added. This method focuses on the analysis of covariation relationships between microbial species, without involving the extraction of temporal features at the functional pathway level, nor establishing a temporal baseline model of an individual's own functional pathways.

[0004] The aforementioned existing technologies each address technical problems at different levels. CN105046094A addresses how to conduct horizontal comparisons between individual and population databases, while CN112151117B addresses how to capture dynamic changes in microbial community interaction networks at the species level. However, both share a common technical deficiency that has not yet been recognized and resolved by existing technologies: neither can decompose the time-series data of a single individual's metagenomic functional pathways into three components (trend component, periodic component, and residual component), nor can they dynamically assess the degree of deviation and automatically distinguish the type of deviation dynamics based on the individual's own historical baseline at the functional pathway level. This deficiency prevents existing technologies from distinguishing between trend deviations (functional pathway abundance shows a continuous upward or downward trend), periodic deviations (functional pathway abundance fluctuates normally according to a certain rhythm), and perturbative deviations (temporary abnormalities caused by temporary factors), and also from judging the degree and nature of deviation within the context of an individual's own functional homeostasis.

[0005] In summary, current technologies lack a method that can organically unify time-series signal decomposition and individualized baseline deviation assessment at the functional pathway level, preventing the full exploration and utilization of the rich dynamic functional information in multi-timepoint gut microbiota metagenomic sequencing data. Therefore, it is necessary to provide a new analytical method to fill the aforementioned technological gap. Summary of the Invention

[0006] The purpose of this invention is to overcome the shortcomings of existing technologies and provide a dynamic analysis method for multi-time-series gut microbiota gene detection data. By constructing a time-series matrix of gut microbiota functional pathway abundance at multiple time points, establishing an individualized time-series baseline model, quantifying multi-dimensional deviations, and automatically distinguishing abnormality types, it can achieve individualized dynamic assessment of gut microbiota functional status, accurately identify functional abnormalities with different dynamic causes, and provide a reliable quantitative analysis scheme for gut microbiota health monitoring.

[0007] To solve the above-mentioned technical problems, the present invention provides the following technical solution: a dynamic analysis method for multi-time-series gut microbiota gene detection data, the method comprising the following steps:

[0008] Step S1: Obtain metagenomic sequencing data of fecal samples from multiple time points of the same individual, perform functional annotation on the metagenomic sequencing data at each time point, extract the abundance values ​​of multiple functional pathways at each time point, and construct a functional pathway abundance matrix. The rows of the functional pathway abundance matrix correspond to functional pathways, and the columns correspond to time points.

[0009] Step S2: From all time points of an individual, identify the continuous time points where the individual is in a healthy state and the abundance value of the functional pathways meets the preset stability criteria as the baseline period data window. The preset stability criteria include that the abundance variation coefficient of the preset core functional pathway set among multiple functional pathways is less than a preset threshold.

[0010] Step S3: Perform time series decomposition on the time series of each functional pathway in the baseline data window, decomposing the time series of each functional pathway into trend component series, periodic component series, and residual component series. Based on the decomposed trend component series, periodic component series, and residual component series, establish an individualized time series baseline model for each functional pathway. The individualized time series baseline model includes the statistical characteristics of the trend component baseline, the envelope of the periodic component baseline, and the statistical distribution of the residual component baseline.

[0011] Step S4: When the metagenomic sequencing data of an individual at a new time point arrives, the functional pathway abundance value at the new time point is added to the time series of the corresponding functional pathway. The time series after the addition is decomposed again. Based on the individualized time series baseline model, the deviation of the trend component, the deviation of the periodic component, and the deviation of the residual component at the new time point on each functional pathway are calculated respectively.

[0012] Step S5: Based on the combination pattern of the trend component deviation, periodic component deviation, and residual component deviation of each functional pathway, classify the deviation events of each functional pathway into one of the preset multiple abnormal types, and output the deviation analysis results in units of functional pathways.

[0013] By acquiring metagenomic functional pathway abundance data of the same individual at multiple time points, identifying the individual's own baseline window, decomposing each functional pathway time series into three components (trend, cycle, and remainder) and establishing an individualized baseline model, and calculating the deviation of each component when new data arrives and classifying abnormality types according to combination patterns, the reference system for dynamic assessment of functional pathways can be shifted from population statistics to the individual's own historical steady state, and the deviation dynamics type can be automatically distinguished, providing a precise quantitative basis for personalized gut microbiota function monitoring.

[0014] Furthermore, in step S3, the time series of each functional pathway is decomposed using a seasonal trend decomposition method based on local weighted regression, and the abundance value of functional pathway p at time point t is determined. Decomposed into:

[0015]

[0016] In the formula, The abundance value of functional pathway p at time point t. Let p be the trend component value of the functional pathway at time point t. Let be the periodic component value of functional pathway p at time point t. The residual component value of functional pathway p at time point t;

[0017] The trend component value The periodic component values ​​reflect the long-term evolution direction of functional pathway abundance within the baseline data window. The periodic fluctuation pattern reflecting the abundance of functional pathways, the residual component values It reflects the residual irregular fluctuations after removing trend and periodic components.

[0018] By employing a seasonal trend decomposition method based on local weighted regression, the time series of functional pathway abundance is clearly separated into trend components, periodic components, and residual components. This method can robustly extract the variable components at different time scales from a finite-length time series, reduce the interference of outliers on the decomposition results, and provide high-quality three-component input data for the subsequent establishment of a reliable individualized baseline model.

[0019] Furthermore, the trend component baseline statistical features in step S3 include: extracting the trend component sequence of each functional pathway p within the baseline period data window. The starting and ending values ​​are used to calculate the trend slope, which is the change in the trend component value per unit time. The trend slope is used as the baseline feature of the trend component of the functional pathway p.

[0020] By extracting the starting and ending values ​​of the trend component sequence during the baseline period to calculate the trend slope as a baseline feature, the long-term evolution direction and rate of change of the abundance of each functional pathway during the steady state period can be characterized with a simple quantitative indicator. This allows for sensitive detection of trend reversal or significant rate change when a new time point is reached, thus improving the early perception of chronic functional drift.

[0021] Furthermore, the baseline envelope of the periodic component in step S3 is established in the following manner:

[0022] Extract the periodic component sequence of each functional pathway p within the baseline data window. The peaks and troughs of each period are connected sequentially to form the upper envelope, and the troughs of each period are connected sequentially to form the lower envelope.

[0023] By connecting the peaks and troughs of each periodic component sequence during the baseline period to form upper and lower envelopes as periodic baselines, the normal fluctuation boundary morphology of the abundance of functional pathways in an individual can be characterized in a nonparametric manner, adapting to the differences in periodic rhythms among different individuals, and providing a personalized reference framework for achieving highly specific detection of periodic rhythm abnormalities.

[0024] Furthermore, the baseline statistical distribution of the remaining components in step S3 is established in the following way:

[0025] Calculate the remainder component sequence for each functional pathway p within the baseline data window. Standard deviation Based on standard deviation Establish reference ranges for the residual components of functional pathway p. The reference ranges for the residual components are as follows: k is a preset multiple.

[0026] Individualized models were established by calculating the standard deviation of the baseline residual component series. The reference range can objectively quantify the inherent random fluctuation amplitude of each functional pathway, which is not explained by trends and cycles. This allows the judgment of residual fluctuation anomalies at new time points to be made within the individual's habitual fluctuation level, effectively reducing misjudgments caused by differences in the fluctuation background between individuals.

[0027] Furthermore, the trend component deviation in step S4 is the absolute value of the difference between the update trend slope of the functional path p after the addition of the new time point and the trend slope of the functional path p within the baseline data window.

[0028] The periodic component deviation is the penetration depth of the periodic component value at a new time point relative to the periodic component baseline envelope.

[0029] The deviation of the remainder component is the absolute value of the Z-score of the remainder component at the new time point.

[0030] By defining three types of deviation—the absolute difference in trend slope, the penetration depth of the periodic component envelope, and the absolute value of the residual component Z-score—the deviation assessment can be decoupled from the general abundance difference into three independent dimensions: the amount of trend change, the degree of periodic transgression, and the intensity of instantaneous fluctuations. This makes the characterization of the degree of deviation more dynamically interpretable and more granularly comparable.

[0031] Furthermore, the preset multiple exception types in step S5 include:

[0032] Trend deviation type, that is, the deviation of the trend component exceeds the preset trend deviation threshold, while the deviation of the periodic component and the deviation of the residual component do not exceed their respective preset thresholds;

[0033] The periodic imbalance type of deviation means that the deviation of the periodic component exceeds the preset periodic deviation threshold while the deviation of the trend component does not exceed the preset trend deviation threshold.

[0034] Acute disturbance type deviation, i.e., only the deviation of the remaining component exceeds the preset disturbance deviation threshold;

[0035] Composite deviation type, that is, at least two of the deviations of the trend component deviation, periodic component deviation and residual component deviation exceed their respective preset thresholds at the same time.

[0036] Normal fluctuation type, that is, the deviation of trend component, periodic component and residual component does not exceed their respective preset thresholds.

[0037] By automatically mapping the combination patterns of the three-component deviation to five types—trend deviation, rhythm disorder deviation, acute disturbance deviation, compound deviation, and normal fluctuation—the system can accurately distinguish the dynamic causes of functional pathway deviation behavior. This allows the test report to not only indicate whether an abnormality exists, but also to specify whether the abnormality is a persistent change in direction, a rhythm disorder, or an occasional fluctuation, guiding more targeted follow-up measures.

[0038] Furthermore, after outputting the deviation analysis results in terms of functional pathways in step S5, the method further includes: aggregating the deviation analysis results of all functional pathways across pathways, calculating the percentage of abnormal pathways, where the percentage of abnormal pathways is the proportion of the number of functional pathways with deviation events to the total number of functional pathways, and generating an overall functional status rating of the individual gut microbiota based on the percentage of abnormal pathways.

[0039] By aggregating deviation analysis results across all functional pathways, calculating the proportion of abnormal pathways, and generating an overall functional status rating, the deviation information of hundreds of functional pathways can be compressed into a global health level indicator. This helps users and clinicians quickly grasp the overall stability of gut microbiota function and improves the actual readability and decision support efficiency of dynamic analysis results.

[0040] Furthermore, the cross-pathway aggregation also includes:

[0041] The core stable pathway set is identified from the baseline period data window. The core stable pathway set consists of multiple functional pathways whose trend component baseline statistical characteristics satisfy a preset stability condition and whose residual component baseline statistical distribution satisfies a preset low volatility condition.

[0042] When the number of functional pathways experiencing deviation events in the core stable pathway set exceeds a preset threshold, a warning of the corresponding level is triggered.

[0043] By identifying a set of core stable pathways with stable trends and low fluctuations in other pathways during the baseline period, and triggering an upgraded warning when deviations occur in this set, the monitoring focus can be locked on the most stable functional modules of the individual, avoiding the dilution of attention to the damage signals of the microecological base by fluctuations in secondary pathways, and significantly improving the timeliness and specificity of the warning for systemic imbalances in the microbial community.

[0044] Furthermore, the method also includes:

[0045] When a functional pathway is identified as having a deviation event, the abundance distribution data of the functional pathway in the population health database is queried, the abundance value of the functional pathway at the new time point is compared with the functional pathway abundance distribution data, and the population reference validation result is output.

[0046] When both the deviation event identified by the individualized time-series baseline model and the population reference validation results indicate anomalies, the confidence level of the deviation event is marked as high confidence.

[0047] By querying a population database for secondary validation when individualized baseline deviations are detected, and marking bi-source consistent anomalies as high confidence, we can cross-validate deviation judgments using population statistics while maintaining the priority of individual references. This avoids misjudging individual norms by population standards and provides additional objective support for anomalies discovered through individualization, thereby improving the reliability of the analysis conclusions.

[0048] Compared with existing technologies, this dynamic analysis method for multi-time-series gut microbiota gene detection data has the following advantages:

[0049] I. This invention acquires metagenomic sequencing data from multiple time points of the same individual and constructs a functional pathway abundance matrix. It identifies the baseline data window in an individual's healthy state and decomposes the time series of each functional pathway into trend, periodic, and residual components. It then establishes individualized time-series baseline models for each component. This allows the reference system for dynamic assessment of gut microbiota function to shift from population statistical standards to the individual's own historical homeostasis. It enables precise quantification of multidimensional deviations of functional pathways and automatic differentiation of abnormal dynamic types. This effectively solves the technical defects of existing technologies that cannot distinguish functional pathway abnormalities of different causes and cannot judge the nature of deviations based on the individual's own homeostasis. It provides a stable and reliable quantitative analysis basis for individualized dynamic monitoring of gut microbiota function.

[0050] Second, this invention performs cross-pathway aggregation analysis on deviation analysis results of individual functional pathways, calculates the proportion of abnormal pathways, and generates an overall functional status rating of the individual's gut microbiota. At the same time, it identifies the core stable pathway set within the baseline period and triggers graded early warnings based on their deviations. This can transform fine-grained pathway analysis results into intuitive overall health status assessment conclusions, lock in the core supporting modules of individual microbiota functional homeostasis to achieve early warning of systemic imbalances, and complete the confidence labeling of deviation events through reference data from healthy populations, greatly improving the practicality of the analysis results and the efficiency of clinical decision support.

[0051] Other advantages, objectives and features of the invention will be set forth in part in the description which follows, and in part will be apparent to those skilled in the art from the following examination or study, or may be learned from the practice of the invention. Attached Figure Description

[0052] To more clearly illustrate the technical solutions in the embodiments of the present invention or the prior art, the accompanying drawings used in the description of the embodiments or the prior art will be briefly introduced below. Obviously, the drawings described below are merely some embodiments of the present invention. For those skilled in the art, other drawings can be obtained based on these drawings without any creative effort.

[0053] Figure 1 This is a flowchart illustrating the overall process of the dynamic analysis method for multi-time-series gut microbiota gene detection data of the present invention.

[0054] Figure 2 This is a schematic diagram of the functional pathway temporal three-component decomposition and individualized temporal baseline model construction of the present invention;

[0055] Figure 3 This is a schematic diagram of the new time point deviation calculation and anomaly type classification process of the present invention. Detailed Implementation

[0056] To further illustrate the technical means and effects of the present invention in achieving its intended purpose, the following detailed description of the specific implementation methods, structures, features, and effects of the present invention, in conjunction with the accompanying drawings and preferred embodiments, is provided below.

[0057] Example

[0058] This embodiment specifically discloses a dynamic analysis method for multi-time-series gut microbiota gene detection data, applied to the longitudinal dynamic monitoring and functional abnormality identification of individual gut microbiota metagenomic sequencing data. Based on metagenomic sequencing data from fecal samples at multiple time points from the same individual, this embodiment achieves individualized dynamic assessment of gut microbiota functional status through the entire process of constructing a functional pathway abundance matrix, identifying individualized healthy baseline data windows, completing time-series three-component decomposition and baseline model establishment, quantifying deviations from new samples, and automatically classifying abnormality types. This provides a practical and reproducible quantitative analysis solution for gut microbiota gene detection services. Figure 1 As shown, the overall process of the analysis method in this embodiment includes five core steps, namely, abundance matrix construction, baseline window identification, individualized baseline model establishment, new sample deviation calculation, deviation event classification and result output.

[0059] Specifically, the analysis process in this embodiment begins with the acquisition and preprocessing of sequencing data. First, fecal samples from the same individual at multiple consecutive or non-consecutive time points are collected. Each sample undergoes the same collection, preservation, and nucleic acid extraction process to ensure consistent processing conditions across different time points and eliminate interference from batch processing differences on subsequent analysis results. Metagenomic sequencing is performed on the fecal samples at each time point, obtaining the corresponding raw sequencing data. A second-generation high-throughput sequencing platform is used, and the sequencing data volume meets the minimum coverage requirements for metagenomic functional annotation.

[0060] For example, the raw sequencing data at each time point undergoes standardized preprocessing. The preprocessing workflow includes sequencing adapter sequence removal, low-quality read filtering, and host genome sequence alignment and removal to obtain high-quality, clean sequencing data for each time point. The criteria for low-quality read filtering are that reads with an average quality value below a set threshold are removed. Host genome sequence alignment uses homology alignment tools to align to a human reference genome; all matching reads are removed to avoid interference from host-derived sequences on the microbial community data.

[0061] Metagenomic functional annotation was performed on clean sequencing data at each time point after preprocessing. A universal microbial functional database was used for the annotation, and the process included gene prediction, construction of a non-redundant gene set, and gene function alignment annotation. The result was the abundance values ​​of multiple functional pathways at each time point. The annotation system for functional pathways adopted a recognized metabolic pathway annotation system to ensure the universality and comparability of the annotation results.

[0062] After completing functional annotation and abundance extraction at all time points, a functional pathway abundance matrix is ​​constructed. Each row in the matrix corresponds to a single functional pathway, and each column corresponds to a time point in the sample collection. The value at each position in the matrix represents the abundance value of the corresponding functional pathway at that time point. The functional pathway annotation system remains completely consistent across all time points, ensuring that abundance values ​​in the same row correspond to the same functional pathway, and abundance values ​​in the same column correspond to all functional pathways at the same time point, thus achieving numerical comparability between different time points and different functional pathways.

[0063] After constructing the functional pathway abundance matrix, the baseline data window identification step is performed. The baseline data window provides the data foundation for the subsequent establishment of individualized time-series baseline models. Its core is a continuous set of time points in which the abundance of functional pathways in the microbial community remains stable when the individual is in a healthy state.

[0064] Specifically, the first step is to mark the time points of an individual's health status. Based on the individual's clinical health records and self-health status feedback, the health status at each time point is determined. Time points in which the individual has no clear disease onset, no history of antibiotic use or other drugs that affect the gut microbiota, and no drastic changes in diet and lifestyle are selected as the health status time points.

[0065] From all health status time points, consecutive time points that meet the preset stability criteria are identified as the baseline data window. The core indicator for the preset stability criteria is that the coefficient of variation of abundance of a preset core functional pathway set is less than a preset threshold. The preset core functional pathway set is a set of core functional pathways related to human basal metabolism and intestinal barrier homeostasis. The functional pathways within this set generally remain stable in healthy individuals and are not easily disturbed by temporary external factors. The coefficient of variation of abundance is the ratio of the standard deviation to the mean of the abundance value of a single functional pathway within the consecutive time points to be evaluated, used to quantify the degree of fluctuation in the abundance of functional pathways.

[0066] For example, for all consecutive combinations of health status time points, the abundance coefficient of variation (COP) of each functional pathway within the core functional pathway set is calculated sequentially. When, within a consecutive time point, the abundance COPs of functional pathways exceeding a set proportion in the core functional pathway set are all less than a preset threshold, that consecutive time point is identified as a candidate baseline period data window. Among multiple candidate baseline period data windows, the window with the longest time span and the largest number of time points is preferentially selected as the final baseline period data window to ensure that subsequent time series decomposition and baseline model establishment have sufficient time series data support.

[0067] After determining the baseline data window, time series analysis is performed on the time series of each functional pathway within the baseline data window, and an individualized time series baseline model is established based on the decomposition results. This step is the core of the entire analysis method, as it separates different dynamic components from the functional pathway abundance time series through time series decomposition, providing an individualized reference benchmark for subsequent deviation assessment. Figure 2 As shown, this step completes the three-component decomposition of a single functional pathway time series and constructs the baseline reference system corresponding to each of the three components.

[0068] Specifically, the time-series decomposition employs a seasonal trend decomposition method based on local weighted regression. For each functional pathway within the baseline data window, the time-series decomposition operation is performed individually. For functional pathway p, its abundance value at time point t is decomposed into the sum of three independent components, as shown in the following formula:

[0069]

[0070] In the formula, The abundance value of functional pathway p at time point t. Let p be the trend component value of the functional pathway at time point t. Let be the periodic component value of functional pathway p at time point t. This represents the residual component value of functional pathway p at time point t.

[0071] It is understandable that the trend component reflects the long-term evolution direction of functional pathway abundance within the baseline data window, unaffected by short-term periodic fluctuations and random disturbances; the periodic component reflects the fixed-period repetitive fluctuation pattern of functional pathway abundance within the baseline period, corresponding to the regular changes in gut microbiota function brought about by individual periodic behaviors such as diet and lifestyle; the residual component reflects the residual irregular fluctuations of functional pathway abundance after removing the trend and periodic components, corresponding to the random changes in abundance caused by temporary and irregular factors.

[0072] For example, the execution flow of time series decomposition is as follows: First, the trend component sequence is extracted from the time series by smoothing the fit through local weighted regression; the trend component sequence is subtracted from the original time series to obtain the detrended sequence; the detrended sequence is then subjected to periodic smoothing to extract the periodic component sequence; finally, the periodic component sequence is subtracted from the detrended sequence to obtain the remainder component sequence. During the decomposition process, the period length is determined according to the time interval of sample collection to ensure that the period length matches the actual sampling frequency and improve the accuracy of the decomposition results.

[0073] After completing the time series decomposition, individualized time series baseline models are established for each functional pathway based on the three component sequences obtained from the decomposition. The individualized time series baseline model consists of three parts: the baseline statistical characteristics of the trend component, the baseline envelope of the periodic component, and the baseline statistical distribution of the residual component. These three parts correspond to the individualized reference benchmarks of the three components, respectively.

[0074] First, baseline statistical features of trend components are constructed. The starting and ending values ​​of the trend component sequence for each functional pathway p within the baseline data window are extracted, and the trend slope is calculated. The trend slope represents the change in the trend component value per unit time, and is used as the baseline feature of the trend component of functional pathway p. Specifically, the trend slope is calculated by dividing the difference between the ending and starting values ​​of the trend component sequence by the time span of the baseline data window to obtain the change in the trend component per unit time. A positive trend slope indicates a long-term upward trend in the abundance of the functional pathway within the baseline period; a negative trend slope indicates a long-term downward trend; and a trend slope close to 0 indicates no significant long-term trend change in the abundance of the functional pathway within the baseline period. This quantitative indicator of trend slope accurately depicts the long-term evolution of functional pathways under the steady state of the baseline period, providing a benchmark for subsequent trend deviation assessment.

[0075] Next, the baseline envelope of the periodic components is constructed. The peak and trough values ​​of each period in the periodic component sequence of each functional pathway p within the baseline data window are extracted. The peak values ​​of each period are connected sequentially to form the upper envelope, and the trough values ​​are connected sequentially to form the lower envelope. Specifically, the periodic component sequence is first divided into periods, with the length of each period consistent with the period length set during the time-series decomposition. Within each complete period, the maximum value of the periodic component is extracted as the peak point of that period, and the minimum value is extracted as the trough point. The peak points of all periods are linearly connected in chronological order to form the upper envelope of the periodic components. Similarly, the trough points of all periods are linearly connected in chronological order to form the lower envelope of the periodic components. The interval between the upper and lower envelopes represents the normal fluctuation range of the periodic components of the functional pathway. When the periodic component value at a new time point is within this interval, it indicates that its periodic fluctuation conforms to the normal rhythm of the baseline period; when the periodic component value exceeds this interval, it indicates that its periodic fluctuation is abnormal.

[0076] Finally, the baseline statistical distribution of the residual components is constructed. The standard deviation of the residual component sequences for each functional pathway p within the baseline data window is calculated. Based on standard deviation Establish reference ranges for the residual components of functional pathway p. The reference ranges for the residual components are as follows: k is a preset multiple. Specifically, the mean and standard deviation of the residual component sequences within the baseline period are first calculated. The mean of the residual component sequences theoretically approaches 0, so the standard deviation is used as the core statistic to construct the reference range. The preset multiple k can be set according to the sensitivity requirements of the detection. The larger the k value, the wider the reference range and the higher the detection specificity; the smaller the k value, the narrower the reference range and the higher the detection sensitivity. The reference range of the residual components is the normal range of random fluctuations of this functional pathway. When the residual component value at a new time point is within this range, it means that its random fluctuation is in line with the normal level of the baseline period; when the residual component value exceeds this range, it means that an acute disturbance exceeding the normal random fluctuation range has occurred.

[0077] When new time-point metagenomic sequencing data for an individual arrives, a dynamic analysis and deviation calculation step is performed for the new sample. This step, based on the established individualized time-series baseline model, processes the sequencing data at the new time point, quantifies the deviation of the three components of each functional pathway, and provides a quantitative basis for subsequent abnormality classification.

[0078] Specifically, the metagenomic sequencing data at the new time points were first processed using the same preprocessing and functional annotation procedures as the baseline samples. Abundance values ​​for each functional pathway at the new time points were extracted to ensure complete comparability with the baseline data. The abundance values ​​of the functional pathways at the new time points were then appended to the end of the corresponding time series data to form an updated time series. This updated time series includes all baseline time points and the newly added time points.

[0079] For each updated time series of functional pathways, the seasonal trend decomposition method based on local weighted regression, which is completely consistent with the baseline period, is executed again to obtain the trend component value, periodic component value, and residual component value corresponding to the new time point. The parameter settings for the time series decomposition are completely consistent with the baseline period decomposition process to ensure that the three components obtained by decomposition have the same physical meaning and comparability with the baseline period components.

[0080] Based on the established individualized time-series baseline model, the deviation of the trend component, the deviation of the periodic component, and the deviation of the residual component at each functional path at the new time point are calculated respectively.

[0081] First, the trend component deviation is calculated. The trend component deviation is the absolute value of the difference between the updated trend slope of functional pathway p after the addition of the new time point and the trend slope of functional pathway p within the baseline data window. Specifically, based on the updated time series, the trend component series including the new time point is extracted, and the updated trend slope is calculated. The calculation method for the updated trend slope is exactly the same as that for the baseline trend slope. The difference between the updated trend slope and the baseline trend slope is calculated, and the absolute value of the difference is taken as the trend component deviation. The larger the value of the trend component deviation, the more significant the change in the long-term evolution trend of the functional pathway after the addition of the new time point compared to the baseline steady state.

[0082] Next, the deviation of the periodic component is calculated. The deviation of the periodic component is the penetration depth of the periodic component value at the new time point relative to the baseline envelope of the periodic component. Specifically, when the periodic component value at the new time point is within the interval between the upper and lower envelopes of the baseline period, the deviation of the periodic component is 0, representing no deviation. When the periodic component value at the new time point is greater than the value at the corresponding time position of the upper envelope, the deviation of the periodic component is the difference between the periodic component value at the new time point and the corresponding value of the upper envelope. When the periodic component value at the new time point is less than the value at the corresponding time position of the lower envelope, the deviation of the periodic component is the difference between the corresponding value of the lower envelope and the periodic component value at the new time point. The larger the value of the deviation of the periodic component, the more severe the deviation of the periodic component value at the new time point from the normal fluctuation boundary of the baseline period, and the higher the degree of cyclical rhythm disorder.

[0083] Finally, the deviation of the residual components is calculated. The deviation of the residual components is the absolute value of the Z-score of the residual component values ​​at the new time point. Specifically, the Z-score is calculated by subtracting the mean of the residual component series at the baseline period from the residual component value at the new time point, and then dividing by the standard deviation of the residual component series at the baseline period. The absolute value of the calculated Z-score is taken as the deviation of the residual component. The larger the value of the deviation of the residual component, the greater the degree to which the residual component value at the new time point exceeds the normal random fluctuation range of the baseline period, and the greater the amplitude of the acute disturbance.

[0084] After calculating the deviation of the three components at the new time point, the deviation event classification and result output steps are performed. This step, based on the combination pattern of the three component deviations, automatically classifies functional pathway deviation events and outputs the corresponding analysis results. Simultaneously, it performs cross-pathway aggregation analysis and population reference validation, providing complete analytical conclusions for gut microbiota gene detection services. Figure 3 As shown, this step is based on the comparison results of the deviation of the three components with the corresponding thresholds to complete the automatic determination and classification of the anomaly type.

[0085] Specifically, corresponding judgment thresholds are pre-set for the deviation of the trend component, the deviation of the periodic component, and the deviation of the residual component, including a preset trend deviation threshold, a preset periodic deviation threshold, and a preset disturbance deviation threshold. The three thresholds correspond to the abnormal judgment thresholds for the deviation of the three components. When the deviation value exceeds the corresponding threshold, the component is judged to have an abnormal deviation.

[0086] Based on the combination pattern of the trend component deviation, periodic component deviation, and residual component deviation of each functional pathway, the deviation events of each functional pathway are classified into one of a number of preset abnormality types. The preset abnormality types include trend deviation type, rhythm disorder type deviation type, acute disturbance type deviation type, compound deviation type, and normal fluctuation type.

[0087] When the deviation of the trend component of a functional pathway exceeds a preset trend deviation threshold, while the deviations of the periodic component and the residual component do not exceed their respective preset thresholds, the deviation event of that functional pathway is classified as a trend deviation type. A trend deviation type indicates that the long-term evolution trend of the functional pathway has changed significantly compared to the baseline steady state, while periodic fluctuations and random disturbances remain normal. This type of abnormality typically corresponds to a persistent shift in gut microbiota function caused by long-term changes in diet, lifestyle, or chronic health status.

[0088] When the deviation of the periodic component of a functional pathway exceeds a preset periodic deviation threshold, while the deviation of the trend component does not exceed a preset trend deviation threshold, the deviation event of that functional pathway is classified as a periodic rhythm disorder type. A periodic rhythm disorder type indicates that the periodic fluctuation pattern of the functional pathway has significantly deviated from the baseline steady state, while the long-term evolution trend remains normal. This type of abnormality typically corresponds to the disruption of the periodic regularity of gut microbiota function caused by individual circadian rhythm disturbances and dietary instability.

[0089] When only one component of a functional pathway deviates beyond a preset perturbation deviation threshold, while the deviations of the trend component and the periodic component do not exceed their respective preset thresholds, the deviation event of that functional pathway is classified as an acute perturbation type. An acute perturbation type indicates that the functional pathway exhibits only a transient abnormality exceeding the normal random fluctuation range, while the long-term trend and periodic rhythm remain normal. This type of abnormality typically corresponds to acute fluctuations in gut microbiota function caused by transient factors such as temporary dietary stimuli or short-term medication.

[0090] When at least two of the trend component deviation, periodic component deviation, and residual component deviation of a functional pathway simultaneously exceed their respective preset thresholds, the deviation event of that functional pathway is classified as a composite deviation type. A composite deviation type indicates that the functional pathway exhibits abnormal deviations in multiple dimensions simultaneously, resulting in a holistic change in the dynamic pattern of gut microbiota function. This type of abnormality typically corresponds to relatively severe gut microbiota dysbiosis or changes in health status.

[0091] When the deviation of the trend component, periodic component, and residual component of a functional pathway does not exceed their respective preset thresholds, the deviation event of that functional pathway is classified as a normal fluctuation type. A normal fluctuation type indicates that none of the three dimensions of the functional pathway show abnormal deviations, and the functional state of the microbial community remains within the steady-state range of the baseline period.

[0092] After classifying deviation events for a single functional pathway, the deviation analysis results are output on a per-functional-pathway basis. The analysis results include the deviation values ​​of the three components of each functional pathway, the deviation event classification results, and an explanation of the degree of abnormality.

[0093] After outputting deviation analysis results at the functional pathway level, cross-pathway aggregation analysis is performed on the deviation analysis results of all functional pathways. First, the percentage of abnormal pathways is calculated, which is the proportion of functional pathways with deviation events to the total number of functional pathways. Based on the percentage of abnormal pathways, an overall functional status rating of the individual's gut microbiota is generated. Specifically, multiple intervals for the percentage of abnormal pathways corresponding to different status ratings are pre-defined. Based on the interval in which the calculated percentage of abnormal pathways falls, the corresponding overall functional status rating is matched. The status rating includes four levels: normal, mildly abnormal, moderately abnormal, and severely abnormal, providing users with an intuitive overall assessment of the gut microbiota's functional status.

[0094] Cross-pathway aggregation analysis also includes the monitoring and early warning of core stable pathway sets. First, core stable pathway sets are identified from the baseline data window. These sets consist of multiple functional pathways whose baseline statistical characteristics of the trend component meet a preset stability condition, and whose baseline statistical distribution of the residual components meets a preset low volatility condition. The preset stability condition is that the absolute value of the baseline trend slope of the functional pathway is lower than a preset stability threshold, indicating that the functional pathway has no significant long-term trend change during the baseline period. The preset low volatility condition is that the standard deviation of the baseline residual component sequence of the functional pathway is lower than a preset volatility threshold, indicating that the random fluctuation amplitude of the functional pathway during the baseline period is extremely low. The core stable pathway set is the core supporting module for the functional homeostasis of an individual gut microbiota, and its stability directly reflects the basic health status of the gut microbiota. When the number of functional pathways in the core stable pathway set exhibiting deviation events exceeds a preset threshold, a corresponding level of early warning is triggered. The warning level is positively correlated with the number of core stable pathways exhibiting deviations, achieving early warning of gut microbiota imbalance.

[0095] The method in this embodiment also includes population-referenced validation and confidence level labeling for deviation events. When a functional pathway is determined to have a deviation event, the abundance distribution data of that functional pathway in a population healthy population database is queried. The abundance value of the functional pathway at the new time point is compared with the abundance distribution data, and a population-referenced validation result is output. The population healthy population database is a large-sample database of gut microbiota functional pathway abundance in healthy individuals. The database contains the abundance distribution statistics of each functional pathway in the healthy population, including the mean, standard deviation, and normal reference range. When the abundance value of the functional pathway at the new time point exceeds the normal reference range for healthy individuals, the population-referenced validation result indicates an anomaly; when the abundance value of the functional pathway at the new time point is within the normal reference range for healthy individuals, the population-referenced validation result indicates normality. When both the deviation event determined by the individualized time-series baseline model and the population-referenced validation result indicate an anomaly, the confidence level of the deviation event is marked as high confidence; when only the individualized time-series baseline model determines a deviation event, but the population-referenced validation result indicates normality, the confidence level of the deviation event is marked as individual-specific deviation, providing a confidence level reference for the interpretation of the test report.

[0096] In summary, this embodiment details the complete implementation process of a dynamic analysis method for multi-time-series gut microbiota gene detection data. It achieves standardized integration of multi-time-point data by constructing a functional pathway abundance matrix, establishes a reference benchmark consistent with individual homeostatic characteristics by identifying individualized baseline data windows, isolates different dynamic components of functional pathway abundance through time-series three-component decomposition and establishes corresponding baseline models, accurately quantifies the degree of abnormality through multi-dimensional deviation calculation, automatically distinguishes abnormality types through combination pattern classification, assesses overall functional status and provides early warning through cross-pathway aggregation analysis, and marks the confidence level of deviation events through population-based reference validation. The method of this embodiment can be directly applied to the longitudinal data analysis stage of gut microbiota gene detection services, shifting the reference system for dynamic assessment of microbiota function from population statistics to individual historical homeostasis. This effectively distinguishes abnormal deviations caused by different dynamic factors, providing a feasible and reproducible quantitative analysis solution for personalized gut microbiota health monitoring. All processes in this embodiment can be implemented through computer programs, possessing good executability and reproducibility, and can be directly integrated into gut microbiota gene detection data analysis platforms to meet the application needs of commercial detection services.

[0097] The above description is merely a preferred embodiment of the present invention and is not intended to limit the present invention in any way. Although the present invention has been disclosed above with reference to preferred embodiments, it is not intended to limit the present invention. Any person skilled in the art can make some modifications or alterations to the above-disclosed technical content to create equivalent embodiments without departing from the scope of the present invention. Any simple modifications, equivalent changes and alterations made to the above embodiments based on the technical essence of the present invention without departing from the scope of the present invention shall still fall within the scope of the present invention.

Claims

1. A dynamic analysis method for multi-time-series gut microbiota gene detection data, characterized in that, The method includes the following steps: Step S1: Obtain metagenomic sequencing data of fecal samples from multiple time points of the same individual, perform functional annotation on the metagenomic sequencing data at each time point, extract the abundance values ​​of multiple functional pathways at each time point, and construct a functional pathway abundance matrix. The rows of the functional pathway abundance matrix correspond to functional pathways, and the columns correspond to time points. Step S2: From all time points of an individual, identify the continuous time points where the individual is in a healthy state and the abundance value of the functional pathways meets the preset stability criteria as the baseline period data window. The preset stability criteria include that the abundance variation coefficient of the preset core functional pathway set among multiple functional pathways is less than a preset threshold. Step S3: Perform time series decomposition on the time series of each functional pathway in the baseline data window, decomposing the time series of each functional pathway into trend component series, periodic component series, and residual component series. Based on the decomposed trend component series, periodic component series, and residual component series, establish an individualized time series baseline model for each functional pathway. The individualized time series baseline model includes the statistical characteristics of the trend component baseline, the envelope of the periodic component baseline, and the statistical distribution of the residual component baseline. Step S4: When the metagenomic sequencing data of an individual at a new time point arrives, the functional pathway abundance value at the new time point is added to the time series of the corresponding functional pathway. The time series after the addition is decomposed again. Based on the individualized time series baseline model, the deviation of the trend component, the deviation of the periodic component, and the deviation of the residual component at the new time point on each functional pathway are calculated respectively. Step S5: Based on the combination pattern of the trend component deviation, periodic component deviation, and residual component deviation of each functional pathway, classify the deviation events of each functional pathway into one of the preset multiple abnormal types, and output the deviation analysis results in units of functional pathways.

2. The method for dynamic analysis of multi-time-series gut microbiota gene detection data according to claim 1, characterized in that, In step S3, the time series of each functional pathway is decomposed using a seasonal trend decomposition method based on local weighted regression, which determines the abundance value of functional pathway p at time point t. Decomposed into: In the formula, The abundance value of functional pathway p at time point t. Let p be the trend component value of the functional pathway at time point t. Let be the periodic component value of functional pathway p at time point t. The residual component value of functional pathway p at time point t; The trend component value The periodic component values ​​reflect the long-term evolution direction of functional pathway abundance within the baseline data window. The periodic fluctuation pattern reflecting the abundance of functional pathways, the residual component values It reflects the residual irregular fluctuations after removing trend and periodic components.

3. The method for dynamic analysis of multi-time series gut microbiota gene detection data according to claim 1, characterized in that, The trend component baseline statistical features in step S3 include: extracting the trend component sequence of each functional pathway p within the baseline data window. The starting and ending values ​​are used to calculate the trend slope, which is the change in the trend component value per unit time. The trend slope is used as the baseline feature of the trend component of the functional pathway p.

4. The method for dynamic analysis of multi-time series gut microbiota gene detection data according to claim 1, characterized in that, The periodic component baseline envelope in step S3 is established in the following manner: Extract the periodic component sequence of each functional pathway p within the baseline data window. The peaks and troughs of each period are connected sequentially to form the upper envelope, and the troughs of each period are connected sequentially to form the lower envelope.

5. The method for dynamic analysis of multi-time series gut microbiota gene detection data according to claim 1, characterized in that, The baseline statistical distribution of the remaining components in step S3 is established in the following way: Calculate the remainder component sequence for each functional pathway p within the baseline data window. Standard deviation Based on standard deviation Establish reference ranges for the residual components of functional pathway p. The reference ranges for the residual components are as follows: k is a preset multiple.

6. The method for dynamic analysis of multi-time series gut microbiota gene detection data according to claim 1, characterized in that, The trend component deviation in step S4 is the absolute value of the difference between the update trend slope of functional path p after the addition of the new time point and the trend slope of functional path p within the baseline data window. The periodic component deviation is the penetration depth of the periodic component value at a new time point relative to the periodic component baseline envelope. The deviation of the remainder component is the absolute value of the Z-score of the remainder component at the new time point.

7. The method for dynamic analysis of multi-time series gut microbiota gene detection data according to claim 1, characterized in that, The preset multiple abnormal types mentioned in step S5 include: Trend deviation type, that is, the deviation of the trend component exceeds the preset trend deviation threshold, while the deviation of the periodic component and the deviation of the residual component do not exceed their respective preset thresholds; The periodic imbalance type of deviation means that the deviation of the periodic component exceeds the preset periodic deviation threshold while the deviation of the trend component does not exceed the preset trend deviation threshold. Acute disturbance type deviation, i.e., only the deviation of the remaining component exceeds the preset disturbance deviation threshold; Composite deviation type, that is, at least two of the deviations of the trend component deviation, periodic component deviation and residual component deviation exceed their respective preset thresholds at the same time. Normal fluctuation type, that is, the deviation of trend component, periodic component and residual component does not exceed their respective preset thresholds.

8. The method for dynamic analysis of multi-time series gut microbiota gene detection data according to claim 1, characterized in that, After outputting the deviation analysis results in terms of functional pathways in step S5, the method further includes: aggregating the deviation analysis results of all functional pathways across pathways, calculating the percentage of abnormal pathways, where the percentage of abnormal pathways is the proportion of the number of functional pathways with deviation events to the total number of functional pathways, and generating an overall functional status rating of the individual gut microbiota based on the percentage of abnormal pathways.

9. The method for dynamic analysis of multi-time series gut microbiota gene detection data according to claim 8, characterized in that, The cross-pathway aggregation also includes: The core stable pathway set is identified from the baseline period data window. The core stable pathway set consists of multiple functional pathways whose trend component baseline statistical characteristics satisfy a preset stability condition and whose residual component baseline statistical distribution satisfies a preset low volatility condition. When the number of functional pathways experiencing deviation events in the core stable pathway set exceeds a preset threshold, a warning of the corresponding level is triggered.

10. The method for dynamic analysis of multi-time-series gut microbiota gene detection data according to claim 1, characterized in that, The method further includes: When a functional pathway is identified as having a deviation event, the abundance distribution data of the functional pathway in the population healthy population database is queried, the abundance value of the functional pathway at the new time point is compared with the functional pathway abundance distribution data, and the population reference validation result is output. When both the deviation event identified by the individualized time-series baseline model and the population reference validation results indicate anomalies, the confidence level of the deviation event is marked as high confidence.

Citation Information

Patent Citations

  • Detection system and method for intestinal flora and dynamic database

    CN105046094A

  • A dynamic observation device and detection method based on time-series metagenomic data

    CN112151117B