Bioinformatics-based microbial community dynamic monitoring method
By constructing multidimensional feature maps and dynamic change maps through high-frequency data collection, the problems of time delay and data processing complexity in the dynamic monitoring of microbial communities in existing technologies are solved. This achieves highly timely and automated monitoring, which can accurately identify changes in microbial community function and is suitable for ecological and environmental monitoring and disease prevention.
Patent Information
- Application Number
- CN202510734952.4
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-06-04
- Publication Date
- 2026-02-10
- Estimated Expiration
- 2045-06-04
AI Technical Summary
Existing methods for monitoring the dynamics of microbial communities suffer from large time delays, cumbersome data processing procedures, lack of real-time performance and versatility, making it difficult to achieve high-frequency, continuous, and highly sensitive monitoring.
By collecting trace sample data at high frequency, a microbial data stream is constructed, metagenomic marker amplification is performed, a multidimensional feature mapping map is built, stable core microbial community modules and dynamic fluctuation regions are identified, and a set of microbial community temporal behavior vectors is generated using sliding window analysis and distributed clustering methods, combined with a multi-scale trend matching algorithm, to construct a dynamic change map of the microbial community.
It enables highly timely and automated monitoring of microbial communities, accurately identifies key dynamic signal segments and functional drift nodes, and improves the accuracy and real-time performance of monitoring. It is applicable to fields such as ecological monitoring, environmental protection, and disease prevention.
Smart Images

Figure CN120636541B_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of microbial communities, specifically to a method for dynamic monitoring of microbial communities based on bioinformatics. Background Technology
[0002] Microbial communities play a crucial role in various fields such as environmental ecology, human health, and agricultural production. Their community structure and dynamic succession processes are often closely related to host conditions or environmental changes. Accurate monitoring of microbial community changes over time not only helps reveal the functional mechanisms of micro-ecosystems but also provides important evidence for disease early warning, environmental pollution source tracing, and agricultural soil improvement. Existing methods for monitoring microbial communities mainly rely on high-throughput sequencing technology combined with traditional bioinformatics workflows to analyze the community structure of samples. Although these methods have achieved some success in analyzing the species composition of static samples, they still have many limitations in monitoring dynamic changes of communities at high frequency, continuously, and with high sensitivity. Most existing technologies involve batch sequencing of samples collected at multiple time points and inferring dynamic changes of the community indirectly based on comparative analysis of the results at each time point. This approach not only suffers from large time delays and cumbersome data processing workflows but also heavily relies on parameter presets and artificial species annotation, lacking real-time performance and universality. Therefore, it is essential to design a bioinformatics-based method for monitoring the dynamics of microbial communities with high timeliness and automated processing capabilities. Summary of the Invention
[0003] (a) Technical problems to be solved
[0004] To address the shortcomings of existing technologies, this invention provides a bioinformatics-based method for dynamic monitoring of microbial communities, which has the advantages of high timeliness and automated processing capabilities, thus solving the problems mentioned in the background technology.
[0005] (II) Technical Solution
[0006] To achieve the aforementioned goals of high timeliness and automated processing capabilities, this invention provides the following technical solution: a bioinformatics-based method for monitoring the dynamics of microbial communities, comprising the following steps:
[0007] We acquired high-frequency data based on trace samples, constructed a microbial data stream for time-series analysis, and performed rapid metagenomic marker amplification on each sampling unit in the data stream to obtain a preliminary feature matrix.
[0008] Based on the co-occurrence frequency of microbial functional genes in the preliminary feature matrix, a multidimensional feature map is constructed, and the feature map is divided into graph structures according to the feature synergy degree to identify stable core microbial community modules and dynamic fluctuation regions.
[0009] Based on the dynamic fluctuation region, real-time pattern recognition is performed on the trend of bacterial community abundance changes. Through continuous window sliding analysis, the growth and deceleration rates, abundance gradients and related factor changes of bacterial communities are extracted, thereby identifying key dynamic signal segments.
[0010] For key dynamic signal segments, a distributed clustering method based on variation information entropy regulation is adopted to reconstruct the evolution trajectory of the bacterial community, generate a set of temporal behavior vectors of the bacterial community, and adaptively remove noise samples and redundant distribution points based on fluctuation characteristics.
[0011] Based on the set of microbial community temporal behavior vectors, a multi-scale trend matching algorithm is used to dynamically align with a pre-constructed reference model to identify the differences in the evolutionary patterns of microbial communities at different time scales and to label the functional drift nodes of the microbial community.
[0012] Based on the functional drift nodes and behavioral vector sets, and combined with the coupling characteristics between microbial community structure changes and functional annotations, a dynamic change map of the microbial community is constructed.
[0013] Preferably, step S1 further includes periodically extracting trace samples from the target microbial community or sample, performing continuous or timed sampling, generating data units for each sampling, constructing a time-series data stream, selecting specific metagenomic markers for amplification, obtaining community information using efficient amplification methods, and integrating the sampling results into a feature matrix through polymerase chain reaction and sequencing analysis.
[0014] Preferably, step S2 further includes obtaining characteristic data of the microbial community through metagenomic marker amplification and integrating it into a preliminary feature matrix, calculating the co-occurrence frequency among functional genes, identifying gene modules with strong synergistic relationships by calculating the synergy among genes, and analyzing the stability and volatility of these modules in time-series data.
[0015] Preferably, the formula for calculating the co-occurrence frequency among functional genes is:
[0016]
[0017] In the formula, N is the total number of samples, and 1 is an indicator function that takes the value of 1 if and only if functional genes A and B are both present in the i-th sample, and 0 otherwise. and , respectively, represent the abundance of functional genes A and B in the i-th sample.
[0018] Preferably, step S3 further includes identifying dynamic fluctuation regions in the microbial community through graph structure partitioning and synergy analysis; extracting time series data of abundance in the dynamic fluctuation regions; capturing the abundance change trend in real time through a sliding window analysis method; identifying the acceleration and deceleration intervals by analyzing the abundance changes in each window; and defining key dynamic signal segments.
[0019] Preferably, step S4 further includes identifying key dynamic signal segments of the bacterial community in the time series through real-time pattern recognition and sliding window analysis methods, using variation information entropy to measure abundance fluctuations, thereby regulating the sensitivity of the clustering process, analyzing the bacterial community time series data through distributed clustering methods, reconstructing the bacterial community evolution trajectory, converting the evolution trajectory of each dynamic signal segment into a time series behavior vector, and identifying noise samples and redundant distribution points through anomaly detection algorithms.
[0020] Preferably, step S5 further includes constructing a set of temporal behavioral vectors of the microbial community to capture key dynamic features of the microbial community evolution, arranging these vectors in chronological order to form a microbial community evolution trajectory matrix, and dynamically aligning it with a reference model through time regularization. Through multi-scale trend feature analysis, the evolutionary deviation between the microbial community and the reference model is evaluated, high-sensitivity intervals of evolutionary differences are identified, and changes in the abundance of functional genes are further analyzed to determine functional drift nodes.
[0021] Preferably, step S6 further includes identifying functional drift nodes by analyzing the set of temporal behavioral vectors of the microbial community at multiple scales, and revealing the intrinsic connection between functional and structural changes in the microbial community by combining the relationships between abundance changes, functional gene expression, species diversity, and functional diversity through coupled feature analysis.
[0022] Preferably, the graph structure constructed by the co-occurrence frequency of microbial functional genes is analyzed using graph convolutional networks to identify the complex relationships between microbial community modules. Through multi-layer feature extraction of graph convolutional networks, the changing trends of microbial community modules in time series are analyzed in depth, and modules related to functional drift are automatically identified, thereby realizing the joint analysis of functional modules and community structure.
[0023] Preferably, in the clustering process of the microbial community temporal behavior vector, an adaptive adjustment mechanism is adopted to adaptively select the clustering method and the number of clusters according to the characteristics of each dynamic signal segment, optimize the clustering results, and automatically adjust the algorithm parameters by comparing the stability and signal-to-noise ratio of each clustering result to optimize the final microbial community evolution trajectory reconstruction result.
[0024] (III) Beneficial Effects
[0025] Compared with existing technologies, this invention provides a bioinformatics-based method for monitoring the dynamics of microbial communities, which has the following advantages:
[0026] This invention, through high-frequency data acquisition and real-time analysis, enables precise monitoring of microbial community changes across different time scales, identifying key dynamic signal segments and functional drift nodes. The method utilizes metagenomic marker amplification technology to obtain a preliminary feature matrix, constructs a multidimensional feature map based on co-occurrence frequencies, and then identifies stable core communities and dynamically fluctuating regions. Within these regions, advanced techniques such as sliding window analysis, information entropy-controlled distributed clustering, and multi-scale trend matching algorithms, combined with the coupling characteristics of community structure and function, are employed to accurately reconstruct the evolutionary trajectory of the community and generate dynamic change maps. This not only improves the timeliness and accuracy of microbial community monitoring but also identifies and tracks key functional change nodes, possessing broad application value, particularly in ecological monitoring, environmental protection, and disease prevention. Attached Figure Description
[0027] Figure 1 This is a schematic diagram of the method of the present invention. Detailed Implementation
[0028] The technical solutions of the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of the present invention, and not all embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of the present invention.
[0029] Example 1
[0030] Please see Figure 1 As shown in the embodiment of the present invention, the method for monitoring the dynamics of microbial communities based on bioinformatics includes the following steps:
[0031] S1: Acquire high-frequency data based on trace samples, construct a microbial data stream for time-series analysis, and perform rapid metagenomic marker amplification on each sampling unit in the data stream to obtain a preliminary feature matrix.
[0032] A trace sample is extracted from the target microbial community or sample. Sampling is performed continuously or at timed intervals to ensure the temporal sequence of the data. Samples are collected periodically from the environment using sensors, automated sampling devices, or manual methods. The sampling time interval and frequency need to be set according to the experimental objectives and the dynamic change cycle of the microbial community to ensure sufficient data points for time-series analysis. The results of each sampling form a data unit, and all sampled data units are arranged in chronological order to form a time-series data stream. Each unit in the data stream contains multiple indicators of the microbial community, such as genomic characteristics, abundance information, and species composition. To ensure data consistency, the data from each sampling needs to be standardized, including handling environmental noise, sample dilution, and inter-sample differences, to ensure high comparability. Specific metagenomic markers, such as 16S rRNA genes, 18S rRNA genes, and ITS regions, are selected as markers for the microbial community. These markers have good conservation and species specificity, making them suitable for rapid identification and classification. Metagenomic markers are amplified in each sampling unit using polymerase chain reaction or other amplification techniques. This process typically employs efficient and rapid amplification methods to ensure experimental speed and data accuracy. The amplified products are analyzed using sequencing or other molecular biology techniques to obtain characteristic information of the microbial community at each sampling point. The amplification results from each sampling are integrated into a feature matrix. Rows in the feature matrix represent different sampling units, and columns represent different microbial markers. Necessary preprocessing is performed on the data matrix, such as removing low-abundance species, noise removal, and missing value imputation. Preliminary cluster analysis and principal component analysis are performed on the feature matrix to extract the dynamic changes of the microbial community at different time points. The preliminary feature matrix provides the foundation for subsequent data analysis. Time-series analysis methods, such as dynamic time warping, Fourier transform, and regression models, are used to analyze the patterns of microbial community changes over time, revealing the fluctuations, stability, and correlations with environmental factors in community composition. Based on metagenomic marker information, functional changes in the microbial community, such as metabolic pathways and drug resistance, are predicted. This process can be achieved through database comparisons or gene function annotation.
[0033] S2: Based on the co-occurrence frequency of microbial functional genes in the preliminary feature matrix, a multidimensional feature map is constructed, and the feature map is divided into graph structures according to the feature synergy degree to identify stable core microbial community modules and dynamic fluctuation regions.
[0034] Microbial community feature data obtained through metagenomic marker amplification are integrated into a preliminary feature matrix. Each row in the matrix represents a sampling unit, and each column represents the abundance or expression of a microbial functional gene. This feature matrix should contain sufficient information to reflect the genomic and functional characteristics of each sample. Based on the microbial functional gene data in the preliminary feature matrix, the co-occurrence frequency between functional genes is calculated. Co-occurrence frequency refers to the frequency of occurrence of two genes or genomic modules in the same sampling unit. Calculation methods include correlation coefficient, cosine similarity, or Jaccard similarity. The co-occurrence frequency matrix is transformed into a network graph, where nodes represent microbial functional genes, and edge weights represent the co-occurrence frequency between functional genes. Genes with high co-occurrence frequencies form strong connections, while genes with low co-occurrence frequencies have weaker connections. Based on the co-occurrence matrix, a multidimensional feature mapping graph is constructed. Let the preliminary feature matrix be...
[0035] Where N is the number of samples, G is the number of functional genes, and each element Given the abundance value of the j-th functional gene in the i-th sample, construct a co-occurrence matrix of functional genes based on M. The formula is:
[0036]
[0037] In the formula, The co-occurrence frequency of functional genes a and b. The larger the value, the more likely functional genes a and b tend to co-occur;
[0038] Define the multidimensional feature map as follows In the formula, V is the set of functional gene nodes, E is the set of edges connecting the functional sets, and W is the weight function of the edges. ;
[0039] Algorithm for constructing a multidimensional feature map: Input the initial feature matrix M, and set the abundance threshold. Calculate C according to the above formula, set a co-occurrence frequency threshold, treat each functional gene as a node, connect functional gene pairs whose co-occurrence frequency exceeds the threshold, and assign edge weights to obtain a multidimensional feature mapping graph. ;
[0040] The nodes in the mapping graph represent different microbial functional genes, and the edge weights between nodes are defined by co-occurrence frequencies. Node characteristics include gene abundance and expression levels, forming a feature map in a high-dimensional space. To facilitate subsequent analysis, dimensionality reduction methods such as principal component analysis, t-SNE, or UMAP are used to map the data in the high-dimensional feature space to a two- or three-dimensional space, helping to visualize the distribution of functional gene characteristics in the microbial community. Based on co-occurrence frequency data, the synergy between functional genes is calculated. Synergy is an indicator of the cooperative relationship between genes in the community, reflecting the degree of their joint role. The synergy between genes is assessed by calculating the similarity between nodes. A synergy matrix is formed using the synergy calculation results. The elements in the matrix represent the strength of the synergistic relationship between two functional genes. Core microbial community modules typically exhibit strong synergy and abundance across multiple samplings, demonstrating high stability and high co-occurrence frequencies. By analyzing the changes in synergy of gene modules under different time periods or environmental conditions, functional gene modules exhibiting fluctuating or dynamic changes under specific conditions are identified. Dynamic fluctuation regions are typically characterized by frequent co-occurrence changes and low stability.
[0041] Quantitative indicators for distinguishing stable core microbial communities from dynamically fluctuating regions include:
[0042] 1. Mean co-occurrence frequency: The average co-occurrence frequency among all functional genes within a module. If the mean remains at a high level at multiple sampling time points, it indicates that there is a strong synergistic relationship among genes within the module.
[0043] 2. Coordination volatility: The degree of variation in the coordination degree within a module over different time periods. By calculating the magnitude of change in the module's coordination degree over each time period, a lower volatility indicates a more stable module; a higher volatility indicates that the module structure is in a dynamic state of change.
[0044] 3. Frequency of occurrence: The proportion of times a module appears across all sampling time points. A high and stable frequency of occurrence indicates that the module has temporal consistency and environmental adaptability, and belongs to a stable core module; modules that appear or disappear frequently tend to be classified as dynamic fluctuation areas.
[0045] 4. Stability of co-occurrence relationships: The degree of consistency of the co-occurrence relationships of functional genes within a module over time. If the co-occurrence intensity of most gene pairs in a module changes little over different time periods, it is determined to be a stable core microbial community module; if the changes are frequent and large, it is determined to be a dynamic fluctuation region.
[0046] When a module has a high mean synergy, low volatility, high frequency of occurrence, and stable co-occurrence relationship, it is identified as a stable core microbial community module.
[0047] When the module has a low mean synergy, high volatility, low frequency of occurrence, and unstable co-occurrence relationship, it is identified as a dynamic fluctuation region.
[0048] This study utilizes Graph Convolutional Networks (GCNs) to analyze the graph structure constructed from the co-occurrence frequencies of microbial functional genes. Microbial functional genes are treated as nodes in a graph, with edges between nodes representing gene co-occurrence relationships. Through multi-layer feature extraction using GCNs, the nodes and their adjacency relationships in the graph structure are iteratively learned, thereby capturing the complex synergistic effects and network structure among functional genes. In multi-level convolutional operations, GCNs can progressively extract the temporal features of gene modules, analyzing the evolutionary trends and dynamic changes of each module at different time points. By learning higher-order relationships between nodes, GCNs can automatically identify key modules related to functional drift, which may reflect significant changes in microbial community structure or function. By jointly analyzing changes in functional modules and microbial community structure, potential ecological turning points and key functional drift events can be identified, providing accurate identification and early warning capabilities for the dynamic monitoring of microbial communities.
[0049] By calculating the co-occurrence frequency of microbial functional genes and constructing a multidimensional feature mapping map, combined with feature synergy analysis, stable core microbial community modules and dynamically changing fluctuating regions within a microbial community can be effectively identified. This not only helps reveal the functional characteristics and community structure of microbial communities but also provides more refined analytical tools for microbial ecology research, supporting applications such as community function prediction, environmental monitoring, and microbiome optimization.
[0050] S3: Based on the dynamic fluctuation region, perform real-time pattern recognition on the trend of bacterial community abundance changes. Through continuous window sliding analysis, extract the growth and deceleration rate, abundance gradient and related factor change characteristics of the bacterial community, and then identify key dynamic signal segments.
[0051] When analyzing dynamic fluctuation regions, the following method is used to extract bacterial community abundance change features and identify key dynamic signal segments. The formula for bacterial community growth / deceleration rate is:
[0052]
[0053] In the formula, For time bacterial abundance at any given time For time intervals;
[0054] Within a continuously sliding window W, the abundance gradient formula is:
[0055]
[0056] In the formula, are the bacterial abundance at the beginning and end of the window, respectively, and n is the number of time points in the window span;
[0057] Key dynamic signal segments are identified using the following thresholds and rules:
[0058] If the absolute value of the bacterial community growth / deceleration rate is greater than a set threshold, it is identified as a potential dynamic change point.
[0059] If the absolute value within a continuous sliding window exceeds a threshold, it is determined to be a potential dynamic segment.
[0060] If a potential dynamic segment exists continuously for more than the minimum duration, it is determined to be a critical dynamic signal segment;
[0061] By using graph structure partitioning and synergy analysis, dynamic fluctuation regions within the microbial community were identified. These fluctuation regions represent areas where microbial abundance changes significantly over different time periods, potentially caused by environmental factors, ecological disturbances, or community interactions. Within these dynamic fluctuation regions, time-series abundance data was extracted, including abundance changes for individual functional genes or species, to further analyze abundance trends. The abundance data from these regions was then formatted into a time-series format, generating microbial abundance information for each time point. Each time point corresponds to specific characteristics of the microbial community, such as species abundance and gene expression levels. A sliding window analysis method was employed to analyze abundance trends in real time. A sliding window defines a fixed-size window within the time series, which slides progressively across the data sequence. Whenever the window slides to a new position, the abundance change characteristics within that window are extracted. The window size needs to be determined based on the rate of data change and experimental design, typically analyzed at different time scales to capture dynamic fluctuations at different frequencies. The step size of the sliding window also needs to be set according to the characteristics of the data. Too small a step size may lead to overly granular data, while too large a step size may miss important dynamic changes. For the abundance data within each sliding window, the rate of increase or decrease of bacterial community abundance is calculated. This is achieved by calculating the slope of abundance change, or by using a more complex dynamic model to quantify the rate of change. The growth rate interval refers to the part of the abundance that increases rapidly within a certain period, while the deceleration interval refers to the region where the growth rate of abundance slows down or decreases. The abundance gradient is an important indicator for measuring the change of bacterial community abundance in a time series. By calculating the magnitude of abundance change within each sliding window, the abundance gradient of the bacterial community at different time periods is obtained. For example, the absolute value of abundance change, the relative rate of change, the peak value, etc. The rate of change of abundance is obtained by calculating the difference between the maximum and minimum values within each window, or by using the local derivative. In continuous window analysis, we can define key dynamic signal segments as critical periods of bacterial community abundance change. These signal segments are usually related to ecological changes, external disturbances, or biological processes. Statistical analysis is used to identify those time periods with significant abundance changes as dynamic signal segments. Machine learning algorithms can be used to further refine and identify time periods with special dynamic patterns, such as sudden and sharp increases or decreases in abundance.
[0062] By analyzing the trend of microbial community abundance changes through continuous window sliding analysis, it is possible to extract the characteristics of microbial community growth and deceleration rates, abundance gradients, and changes in related factors, and further identify key dynamic signal segments. This method can monitor and analyze the dynamic changes of microbial communities in real time over different time periods, helping to identify key fluctuation periods and sensitive stages of microbial communities and their correlation with environmental factors, thus providing important evidence for microbial community management, ecological intervention, and disease prevention and control.
[0063] S4: For key dynamic signal segments, a distributed clustering method based on variation information entropy regulation is adopted to reconstruct the evolution trajectory of the bacterial community, generate a set of temporal behavior vectors of the bacterial community, and adaptively remove noise samples and redundant distribution points based on fluctuation characteristics.
[0064] Variational information entropy is used to quantify the uncertainty of bacterial community abundance changes over time. The abundance change sequence of the bacterial community is defined as follows within a key dynamic signal segment: The formula for the information entropy of mutation is:
[0065]
[0066] In the formula, Let K be the probability of the occurrence of the k-th discrete abundance interval, where K is the total number of intervals after the abundance change is discretized.
[0067] The process of adaptively removing noisy samples and redundant distribution points based on fluctuation characteristics is as follows:
[0068] For each temporal behavior vector (i.e., the microbial community characteristics of each sampling point), a local sliding window is constructed, and the local variation information entropy of the samples within the window is calculated;
[0069] If the local entropy of a sampling point is higher than the global entropy mean, it is judged as an abnormal fluctuation sample and marked as noise.
[0070] If the distance between adjacent samples is less than the set micro-difference and the change in local entropy is less than the micro-amplitude, then it is determined to be a redundant point;
[0071] The distributed clustering process is as follows:
[0072] The set of microbial community temporal behavior vectors is divided into several subsets according to time intervals or feature space density. The partitioning principle ensures high correlation within subsets and few cross-subset boundaries. Each subset is assigned to an independent computing node, and a local clustering algorithm is executed independently on each node. An adaptive variation information entropy control mechanism is applied within each subset to dynamically adjust local clustering parameters. The clustering results of each node are collected, and the overlap of clusters on cross-subset boundaries is detected. Highly overlapping local clusters are merged to generate the final global microbial community evolution trajectory cluster.
[0073] Based on real-time pattern recognition, a sliding window analysis method was used to identify key dynamic signal segments in the microbial community over time. These signal segments reflect dramatic changes in microbial abundance and may be associated with changes in the ecological environment and external disturbances. Key signal segments with biological significance were selected based on the rate and magnitude of abundance changes and their correlation with environmental factors. Signal segments exhibiting significant fluctuations or long-term stable dynamic patterns were chosen for further analysis. Information entropy, a measure of system uncertainty, reflects the degree of data disorder. In microbial community time-series data analysis, information entropy can be used to measure the variability of abundance changes. Within key dynamic signal segments, the variational information entropy was calculated based on the time-series changes in microbial abundance. A higher variational information entropy indicates more dramatic changes in microbial abundance, higher uncertainty, and potentially more ecological disturbances or community structure adjustments. Time-series-based dynamic entropy calculation methods, such as Shannon entropy and Renyi entropy, were used to measure the volatility of microbial abundance within each time period. The sensitivity of the clustering process was adjusted by regulating the variational information entropy. In regions with high entropy, clustering methods may be more refined to capture more subtle changes, while in regions with low entropy, coarser clustering simplifies the model and avoids overfitting noisy data. Based on a strategy of regulating variational information entropy, a distributed clustering method is used to analyze the time-series data of bacterial communities and reconstruct their evolutionary trajectory. The advantage of distributed clustering is that it effectively handles large-scale data while avoiding memory overflow and computational bottlenecks. The abundance information of the bacterial community within each time period is treated as a data point, and the clustering algorithm groups it into several categories according to the temporal variation characteristics of the data. Each category represents the evolutionary trajectory of the bacterial community within a specific time period. Based on the clustering results, the evolutionary trajectory of the bacterial community within each dynamic signal segment is transformed into a behavioral vector. Each vector contains the trend of bacterial abundance change, the rate of increase or decrease, and the correlation characteristics with environmental factors within a specific time period. The temporal behavior vectors consist of the following features: the rate of abundance change, such as the slope of acceleration and deceleration; the magnitude of abundance change, such as peak values, trough values, and abundance gradients; the degree of influence of environmental factors, such as the correlation between temperature, pH, and humidity; and the cooperative relationships between communities, such as the co-occurrence of core bacterial communities. All identified temporal behavior vectors are aggregated into a set to form a global temporal behavior feature map, providing a foundation for subsequent analysis and pattern recognition. By analyzing the clustering results, samples exhibiting anomalous fluctuations or inconsistent with the mainstream evolutionary trend of the community are identified. These samples are usually noisy data, possibly originating from measurement errors, environmental interference, etc. Distance-based anomaly detection algorithms are used to identify noisy samples, which will appear as outliers in the temporal behavior vector set. During the clustering process, some redundant distribution points may be generated. These points are close to other clusters but fail to form an effective community pattern. By calculating the similarity between vectors, redundant distribution points are identified and removed, reducing data redundancy.
[0074] In the clustering process of the microbial community's temporal behavior vectors, an adaptive adjustment mechanism is employed to optimize the clustering results. Based on the characteristics of each dynamic signal segment, the most suitable clustering method and number of clusters are automatically selected. Preliminary clustering is performed on the temporal behavior vectors of each dynamic signal segment. Then, the effectiveness of each clustering scheme is evaluated by comparing the stability and signal-to-noise ratio of the clustering results under different clustering methods and numbers of clusters. The clustering scheme with higher stability and better signal-to-noise ratio is selected as the final result. Based on this, the algorithm automatically adjusts clustering parameters, such as the number of clusters and the distance metric, to further optimize the clustering process, ensuring a more accurate reconstruction of the final microbial community evolutionary trajectory. Through this adaptive adjustment mechanism, the clustering results can better reflect the evolutionary characteristics of the microbial community in different dynamic signal segments, providing efficient evolutionary path analysis.
[0075] By employing a distributed clustering method based on variational information entropy regulation, we can accurately reconstruct the evolutionary trajectory of microbial communities and generate a set of temporal behavioral vectors. This method not only identifies dynamic evolutionary patterns of microbial communities from large-scale data but also improves the accuracy and reliability of data analysis by adaptively eliminating noisy samples and redundant points. This process provides crucial technical support for a deeper understanding of the temporal behavior, ecological changes, and relationships with environmental factors of microbial communities.
[0076] S5: Based on the set of microbial community temporal behavior vectors, a multi-scale trend matching algorithm is used to dynamically align with a pre-built reference model to identify the differences in the evolutionary patterns of microbial communities at different time scales and to label the functional drift nodes of the microbial community.
[0077] When performing multi-scale trend matching analysis based on the time-series behavioral vector set of microbial communities, the first step is to extract multi-scale features from the behavioral vector sequence. Specifically, multiple different time window scales are set, and a sliding window calculation is applied to the continuous behavioral vector sequence at each scale to extract local change features. Each local change feature is used to form a trend feature set by calculating the average change direction and magnitude of the behavioral vectors within the window, thereby capturing the change patterns of microbial community evolution at different time scales.
[0078] After extracting multi-scale trend features, a dynamic time warping algorithm is used to dynamically align the microbial community behavior vector sequence with a pre-constructed reference model. During the alignment process, a smoothed time penalty weight is set according to the time step difference, the matching distance between the two sets of trend features at each time point is calculated, and the optimal matching path is searched through dynamic programming to obtain the overall alignment relationship. The weighted dynamic time warping method can effectively avoid excessive distortion caused by time scale differences, thereby improving the accuracy and stability of evolutionary pattern matching.
[0079] After dynamic alignment is completed, the differences in the evolutionary patterns of microbial communities at different time scales are further identified based on the matching errors in the alignment path. Specifically, the global deviation is measured by calculating the overall alignment error, while the local evolutionary deviation rate is evaluated within a local window. When the local deviation rate exceeds a preset threshold and is accompanied by significant changes in the functional characteristics of the microbial community, evolutionary drift is determined to exist in that time period. Based on the magnitude, duration, and degree of impact of the drift on microbial community function, the corresponding time points or intervals are marked as functional drift nodes, and relevant characteristic change information is recorded.
[0080] Indicators of differences in evolutionary patterns include, but are not limited to:
[0081] Normalized alignment distance: measures the overall deviation between the overall evolutionary trajectory and the reference model;
[0082] Local evolution deviation rate: quantifies the dramatic changes in local evolution trends in the short term;
[0083] Magnitude of change in functional characteristics: Assess the magnitude of shift at the functional level based on changes in the abundance or diversity of functional genes;
[0084] The criteria for judgment include: the normalized alignment distance exceeds the reference threshold; the local evolutionary deviation rate exceeds the mean plus twice the standard deviation in multiple consecutive windows; and the change rate of functional gene abundance or diversity index exceeds 30% in a short period of time.
[0085] The set of temporal behavioral vectors of the microbial community represents the key dynamic features of microbial community evolution at each time period, including abundance increase / decrease rates, changes in abundance gradients, behavioral patterns of key species or genes, and correlations with environmental factors. These vectors are arranged in chronological order to form a complete microbial community evolutionary trajectory matrix. A reference model for comparison is pre-constructed, consisting of standard behavioral vectors from healthy micro-ecosystems, known typical microbial community evolution curves, and typical behavioral temporal templates trained based on historical data. Evolutionary patterns exhibit different expression characteristics at different time scales.
[0086] Microscale: Short-term fluctuations in the microbial community at the hourly / daily level
[0087] Mesoscale: Weekly / monthly adjustments in gut microbiota structure
[0088] Macroscale: Long-term ecological evolution at the quarterly / annual level
[0089] The behavioral vector sequence is segmented using sliding windows of varying lengths. Trend features are extracted within each window, preserving trend information at each scale to form a multi-scale trend feature tensor. The trend features at each time scale are dynamically time-warped and aligned with the reference model, allowing for nonlinear registration to address the inconsistency in the temporal rate of behavioral evolution. This yields an alignment error matrix, used to assess the deviation of the current microbial community from the reference model. The alignment distance is the shortest path distance after DTW alignment, reflecting the degree of deviation from the reference model in evolutionary trends. The multi-scale difference index integrates the alignment errors at various time scales, measuring the overall shift between the evolutionary trend and the reference evolutionary model. Time periods where the alignment error is significantly higher than the average at a given time scale are marked as high-sensitivity intervals for evolutionary differences. If high sensitivity occurs simultaneously at multiple scales, this period may represent a critical ecological turning point. Within these high-sensitivity intervals, the direction and magnitude of changes in the abundance of microbial functional genes are further analyzed. When significant expression shifts occur in certain core functional genes, these are identified as functional drift nodes. Drift determination metrics include: Functional shift angle: the change in the angle between the feature vector of a functional class gene and its historical mean in the vector space, used to identify the direction of the change trend. Drift intensity score: a comprehensive evaluation based on differential expression levels, drift duration, and the weights of the functional pathways involved. A threshold is set; exceeding this threshold indicates a drift node. Drift nodes are labeled in the original time-series behavioral vector, supporting visualization of the microbial community evolution trajectory and drift events along dimensions such as time axis, functional category, and species population.
[0090] By comparing microbial community behavior vectors with a reference model through multi-scale trend analysis and dynamic alignment, the evolutionary deviations of microbial communities at different time scales are accurately identified, and functional drift nodes are accurately labeled based on functional hierarchy analysis. This mechanism is applicable to complex and dynamic scenarios such as health status monitoring, early disease screening, and ecological disturbance tracking, and has strong real-time performance and biological interpretability.
[0091] S6: Based on the functional drift nodes and behavioral vector set, and combined with the coupling characteristics between microbial community structure changes and functional annotations, construct a dynamic change map of the microbial community.
[0092] In the dynamic monitoring of microbial communities, functional drift nodes refer to key points in time where the functional characteristics of the microbial community undergo significant changes beyond the normal fluctuation range, accompanied by marked shifts in community structure. Functional drift nodes typically reflect important events in the microbial community's response to environmental changes, disruption of internal homeostasis, or functional transfer.
[0093] The criteria for identifying functionally drifting nodes include:
[0094] In the set of behavioral vectors, the deviation rate of the local evolutionary trend exceeds the mean plus twice the standard deviation;
[0095] Meanwhile, in the corresponding functional gene set, the rate of change in functional abundance exceeded 30%;
[0096] The community functional diversity index showed changes exceeding a preset threshold.
[0097] In multi-timescale analyses, the consistency of the trend of change is met, that is, similar drift patterns are detected at at least two different scales.
[0098] The process of labeling functional drift nodes at different time scales using a multi-scale time window sliding algorithm is as follows:
[0099] Set multiple sliding time windows and extract the behavior vector sequences within different durations;
[0100] At each scale, local deviation detection and functional change analysis are performed independently, and possible drift candidate nodes are initially labeled.
[0101] The consistency of nodes at different time scales is evaluated by using a scale fusion algorithm, and nodes with high consistency are given higher confidence weights.
[0102] Finally, based on the confidence level, functional drift nodes that are officially labeled are selected and scale source information is added to form multi-scale labeling results.
[0103] Quantitative indicators of the degree of functional change include:
[0104] Functional abundance change rate:
[0105]
[0106] In the formula, This represents the percentage change in the mean abundance of the same functional genes before and after the drift node. and These are the mean functional abundance values at two time points before and after the drift node;
[0107] Magnitude of change in functional diversity: This is quantified by comparing the changes in functional diversity indices. For example, a change in the absolute value of the difference in the Shannon index that exceeds a set threshold is considered a significant change.
[0108] Coupling degree change: measures the consistency between functional changes and structural changes, by calculating the cosine similarity change between the functional gene change vector and the species abundance change vector.
[0109] The process of constructing a dynamically changing map is as follows:
[0110] Construct a time-weighted multi-attribute graph G=(V,E,A,T), where: V is the set of nodes, each node v∈V represents the functional state of a microorganism at a specific time point; E is the set of edges, representing the evolutionary relationship between adjacent nodes, and the weight of the edges represents the degree of evolutionary deviation; A is the set of node attributes, including species structure feature vectors, functional gene abundance vectors, diversity indices, etc.; T is the set of time labels, marking the actual sampling time corresponding to each node;
[0111] Each node v contains: Microbial community structure vector, Function annotation vectors Functional diversity index;
[0112] node and The edge between weight Defined as the weighted difference between overall structural and functional changes:
[0113]
[0114] The process of constructing a dynamic change map is as follows:
[0115] Based on the set of temporal behavioral vectors of the microbial community, structural and functional features are extracted at each sampling time point to construct nodes;
[0116] In chronological order, for nodes at adjacent time points, calculate the distance of changes in community structure and functional abundance, generate weighted edges, and assign evolutionary bias weights.
[0117] For each node, identify and label the functional drift nodes based on the rate of functional change, the rate of structural change, and the fluctuation of diversity.
[0118] Based on the locations where the evolutionary path changes drastically, the process is dynamically divided into sub-stages, forming dynamic sub-graphs of continuous or abrupt changes, which facilitates subsequent stage-by-stage analysis.
[0119] Output a complete map of the dynamic changes of the microbial community G=(V,E,A,T).
[0120] The process of employing a multi-level coupling analysis strategy is as follows:
[0121] For each pair of consecutive nodes, calculate the structural change and the functional change respectively;
[0122] The synchronicity index of the two changes is calculated and defined as the degree of consistency of the change curves;
[0123] Based on the species-function mapping relationship, a species-function association matrix is constructed, and local coupling degree is calculated by combining species change and function change using matrix mapping.
[0124] When labeling nodes for functional drift, both structural drift and functional drift are considered simultaneously. A joint drift threshold is set, meaning that a node is officially labeled as a functional drift node only when both structure and function have changed.
[0125] The evolutionary trajectory and functional changes of the microbial community along a timeline are plotted as a continuous time series map, with time on the x-axis and microbial abundance, functional gene abundance, and other characteristics on the y-axis. Functional drift nodes are marked on the map, showing the structural and functional changes of the microbial community at each node. Dendrograms, pie charts, or heatmaps are used to illustrate the changes in microbial community structure and function. For example, heatmaps can be used to show the abundance changes of various species and their associated functional gene changes at different time points. Maps are also created showing the changes in functional modules during microbial community evolution, such as changes in metabolic pathways, material transformation capacity, and resistance characteristics. Based on the coupling characteristics of functional drift nodes, structural changes, and functional modules, an evolutionary path map of the microbial community along the time dimension is plotted. Each node in the map represents a time point or drift node, and the lines connecting the nodes represent the trend of microbial community evolution. Functional drift nodes and structural changes are displayed synchronously, using color, size, or other graphical elements to visually reflect key changes.
[0126] By constructing a dynamic map of microbial communities, the complex relationship between temporal behavior, structural changes, and functional evolution of microbial communities can be visualized and systematized. This map not only provides a new perspective for understanding the dynamic evolution of microbial communities but also offers strong support for research in disease surveillance, environmental intervention, and microbial therapy.
[0127] It should be noted that, in this document, relational terms such as "first" and "second" are used only to distinguish one entity or operation from another, and do not necessarily require or imply any such actual relationship or order between these entities or operations. Furthermore, the terms "comprising," "including," or any other variations thereof are intended to cover non-exclusive inclusion, such that a process, method, article, or apparatus that comprises a list of elements includes not only those elements but also other elements not expressly listed, or elements inherent to such process, method, article, or apparatus.
[0128] Although embodiments of the invention have been shown and described, it will be understood by those skilled in the art that various changes, modifications, substitutions and alterations can be made to these embodiments without departing from the principles and spirit of the invention, the scope of which is defined by the appended claims and their equivalents.
Claims
1. A method for monitoring the dynamics of microbial communities based on bioinformatics, characterized in that, Includes the following steps: S1: Acquire high-frequency data based on trace samples, construct a microbial data stream for time-series analysis, and perform rapid metagenomic marker amplification on each sampling unit in the data stream to obtain a preliminary feature matrix; S2: Based on the co-occurrence frequency of microbial functional genes in the preliminary feature matrix, a multidimensional feature map is constructed, and the feature map is divided into graph structures according to the feature synergy degree to identify stable core microbial community modules and dynamic fluctuation regions. S3: Based on the dynamic fluctuation region, perform real-time pattern recognition on the trend of bacterial community abundance changes. Extract the growth and deceleration rate, abundance gradient and related factor change characteristics of bacterial community through continuous window sliding analysis, and then identify key dynamic signal segments. S4: For key dynamic signal segments, a distributed clustering method based on variation information entropy regulation is adopted to reconstruct the evolution trajectory of the bacterial community, generate a set of temporal behavior vectors of the bacterial community, and adaptively remove noise samples and redundant distribution points based on fluctuation characteristics. Variation information entropy is used to quantify the uncertainty of bacterial community abundance changes over time; the abundance change sequence of the bacterial community is defined as follows within the key dynamic signal segment: The formula for the information entropy of mutation is: ; In the formula, Let K be the probability of the occurrence of the k-th discrete abundance interval, where K is the total number of intervals after the abundance change is discretized. The process of adaptively removing noisy samples and redundant distribution points based on fluctuation characteristics is as follows: For each temporal behavior vector, a local sliding window is constructed, and the local variation information entropy of the samples within the window is calculated; If the local entropy of a sampling point is higher than the global entropy mean, it is judged as an abnormal fluctuation sample and marked as noise. If the distance between adjacent samples is less than the set micro-difference and the change in local entropy is less than the micro-amplitude, then it is determined to be a redundant point; The distributed clustering process is as follows: The set of microbial community temporal behavior vectors is divided into several subsets according to time intervals or feature space density. The partitioning principle ensures high correlation within subsets and few cross-subset boundaries. Each subset is assigned to an independent computing node, and a local clustering algorithm is executed independently on each node. An adaptive variation information entropy control mechanism is applied within each subset to dynamically adjust local clustering parameters. The clustering results of each node are collected, and the overlap of clusters on cross-subset boundaries is detected. Highly overlapping local clusters are merged to generate the final global microbial community evolution trajectory cluster. S5: Based on the set of microbial community temporal behavior vectors, a multi-scale trend matching algorithm is used to dynamically align with a pre-built reference model to identify the differences in the evolutionary patterns of microbial communities at different time scales and to label the functional drift nodes of the microbial community. When performing multi-scale trend matching analysis based on the time-series behavioral vector set of the microbial community, multi-scale feature extraction is performed on the behavioral vector sequence. Multiple different time window scales are set, and sliding window calculation is applied to the continuous behavioral vector sequence at each scale to extract local change features. Each local change feature is calculated by the average change direction and change amplitude of the behavioral vector within the window to form a set of trend features, thereby capturing the change pattern of microbial community evolution at different time scales. After extracting multi-scale trend features, a dynamic time warping algorithm is used to dynamically align the microbial community behavior vector sequence with the pre-constructed reference model. During the alignment process, a smooth time penalty weight is set according to the time step difference, the matching distance of the two sets of trend features at each time point is calculated, and the optimal matching path is searched through dynamic programming to obtain the overall alignment relationship. S6: Based on the functional drift nodes and behavioral vector set, and combined with the coupling characteristics between microbial community structure changes and functional annotations, construct a dynamic change map of the microbial community.
2. The method for monitoring the dynamics of microbial communities based on bioinformatics according to claim 1, characterized in that, S1 further includes periodically extracting trace samples from the target microbial community or sample, performing continuous or timed sampling, generating data units for each sampling, constructing a time-series data stream, selecting specific metagenomic markers for amplification, obtaining community information using efficient amplification methods, and integrating the sampling results into a feature matrix through polymerase chain reaction and sequencing analysis.
3. The method for monitoring the dynamics of microbial communities based on bioinformatics according to claim 1, characterized in that, S2 further includes obtaining characteristic data of microbial communities through metagenomic marker amplification and integrating them into a preliminary feature matrix, calculating the co-occurrence frequency among functional genes, identifying gene modules with strong synergistic relationships by calculating the synergy between genes, and analyzing the stability and volatility of these modules in time-series data.
4. The bioinformatics-based method for monitoring the dynamics of microbial communities according to claim 3, characterized in that, The formula for calculating the co-occurrence frequency among functional genes is: ; In the formula, N is the total number of samples, and 1 is an indicator function that takes the value of 1 if and only if functional genes A and B are both present in the i-th sample, and 0 otherwise. and , respectively, represent the abundance of functional genes A and B in the i-th sample.
5. The method for monitoring the dynamics of microbial communities based on bioinformatics according to claim 1, characterized in that, S3 further includes identifying dynamic fluctuation regions in the microbial community through graph structure partitioning and synergy analysis; extracting time series data of abundance in the dynamic fluctuation regions; capturing the trend of abundance change in real time through the sliding window analysis method; identifying the acceleration and deceleration intervals by analyzing the abundance change in each window; and defining key dynamic signal segments.
6. The method for monitoring the dynamics of microbial communities based on bioinformatics according to claim 1, characterized in that, S4 further includes identifying key dynamic signal segments of the bacterial community in the time series through real-time pattern recognition and sliding window analysis methods, using variation information entropy to measure abundance fluctuations, thereby regulating the sensitivity of the clustering process, analyzing the bacterial community time series data through distributed clustering methods, reconstructing the bacterial community evolution trajectory, converting the evolution trajectory of each dynamic signal segment into a time series behavior vector, and identifying noise samples and redundant distribution points through anomaly detection algorithms.
7. The method for monitoring the dynamics of microbial communities based on bioinformatics according to claim 1, characterized in that, S5 further includes constructing a set of temporal behavioral vectors of the microbial community to capture key dynamic features of the microbial community evolution, arranging these vectors in chronological order to form a microbial community evolution trajectory matrix, and dynamically aligning it with a reference model through time regularization. Through multi-scale trend feature analysis, the evolutionary deviation between the microbial community and the reference model is evaluated, high-sensitivity intervals of evolutionary differences are identified, and further analysis of changes in the abundance of functional genes is conducted to determine functional drift nodes.
8. The method for monitoring the dynamics of microbial communities based on bioinformatics according to claim 1, characterized in that, S6 further includes identifying functional drift nodes by analyzing the set of temporal behavioral vectors of the microbial community at multiple scales, and revealing the intrinsic connection between functional and structural changes in the microbial community by combining the relationships between abundance changes, functional gene expression, species diversity and functional diversity through coupled feature analysis.
9. The method for monitoring the dynamics of microbial communities based on bioinformatics according to claim 1, characterized in that, This study utilizes graph convolutional networks to analyze the graph structure constructed from the co-occurrence frequency of microbial functional genes, identifies the complex relationships between microbial community modules, and conducts in-depth analysis of the temporal variation trends of microbial community modules through multi-layer feature extraction of graph convolutional networks. It also automatically identifies key modules related to functional drift, achieving joint analysis of functional modules and community structure.
10. The method for monitoring the dynamics of microbial communities based on bioinformatics according to claim 1, characterized in that, In the clustering process of microbial community temporal behavior vectors, an adaptive adjustment mechanism is adopted. The clustering method and number of clusters are adaptively selected according to the characteristics of each dynamic signal segment to optimize the clustering results. The adaptive adjustment mechanism automatically adjusts the algorithm parameters by comparing the stability and signal-to-noise ratio of each clustering result, thereby optimizing the final microbial community evolution trajectory reconstruction result.
Citation Information
Patent Citations
Method and system for identification and classification of operational taxonomic units in a metagenomic sample
CN110021344A
Microbial community clustering analysis method
CN116992314A