A longitudinal microbiome data multivariate count data processing method and system

By employing a two-stage segmented processing method, combined with Gaussian variational approximation and moment estimation, the problem of efficient analysis of multivariate longitudinal microbiome count data was solved, resulting in a significant improvement in computational speed and estimation efficiency.

CN122224288APending Publication Date: 2026-06-16SUZHOU UNIV
View PDF 0 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
SUZHOU UNIV
Filing Date
2026-03-27
Publication Date
2026-06-16

AI Technical Summary

Technical Problem

Existing technologies suffer from high computational costs and low efficiency when processing multivariate longitudinal microbiome data, especially when dealing with high-dimensional integrals in multivariate Poisson mixed-effects models and lacking effective tools, making it difficult to perform efficient data estimation.

Method used

A two-stage segmented processing method is adopted. First, the parameters of the univariate Posson mixed effect model are optimized by Gaussian variational approximation. Then, the moment estimation method of inter-response dependence is used to estimate the cross-response correlation, decompose the high-dimensional integral problem, and improve the computational efficiency.

Benefits of technology

While ensuring estimation accuracy, the analysis speed of microbiome big data was improved, achieving a two-order-of-magnitude increase in computing speed and improving estimation efficiency.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN122224288A_ABST
    Figure CN122224288A_ABST
Patent Text Reader

Abstract

The application discloses a longitudinal microbiome data multivariate count data processing method and system, relates to the technical field of biological information processing, and comprises the following steps: acquiring a microbiome multivariate count data set; wherein the microbiome multivariate count data set comprises a plurality of types of microbiome count data collected in a plurality of time periods for a plurality of microbiome samples and corresponding fixed effect covariates; inputting the microbiome multivariate count data set into a pre-established multivariate Birkhoff mixed effect model for two-stage segmentation processing, and outputting microbiome data estimated values, wherein the microbiome data estimated values comprise first-stage estimated values and second-stage estimated values; the multivariate Birkhoff mixed effect model is used for determining the dependency relationship between the plurality of types of microbiome count data collected in the plurality of time periods for the plurality of microbiome samples and the corresponding fixed effect covariates, and capturing the intra-group correlation and cross-response correlation in the data through a random effect structure; and the estimation efficiency of microbiome big data is improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of bioinformatics processing technology, specifically a method and system for processing multivariate counting data of longitudinal microbiome data. Background Technology

[0002] Microbiome studies often generate complex multivariate longitudinal count data, such as abundance data of multiple microorganisms (e.g., Clostridium difficile, Enterococcus) observed at multiple time points. Current techniques for analyzing such data face challenges due to the complex dependencies inherent in multivariate longitudinal count data. Therefore, multivariate Poisson mixed-effects models (mPMMs) are generally used to model these dependencies. However, existing mPMMs involve high-dimensional integrals over random effects or latent variables, which are difficult to handle and lack closed-form solutions. Numerical methods, such as the Gauss-Hermite quadrature and Laplace approximation, still face problems such as high computational costs when dealing with the complex integrals in mPMMs. Furthermore, there is a lack of readily available tools to implement these methods for mPMMs. Although sampling-based MCMC methods have been extended to multivariate settings using the R package MCMCglmm, they suffer from computational inefficiency, resulting in low estimation efficiency when estimating microbiome data. Summary of the Invention

[0003] To address the shortcomings mentioned in the background section, the present invention aims to provide a method and system for processing multivariate counting data of longitudinal microbiome data.

[0004] Firstly, the objective of this invention can be achieved through the following technical solution: a method for processing multivariate count data of longitudinal microbiome data, the method comprising the following steps: Obtain a multivariate microbial count dataset; wherein, the multivariate microbial count dataset includes multiple microbial count data collected from multiple microbial samples over multiple time periods and corresponding fixed-effect covariates; The microbial multivariate count dataset is input into a pre-established multivariate Boson mixed-effects model for two-stage segmentation processing, and the output is the microbial data estimate, wherein the microbial data estimate includes the first stage estimate and the second stage estimate. The multivariate Boson mixed-effects model is used to determine the dependencies between various microbial count data collected from multiple microbial samples over multiple time periods and the corresponding fixed-effects covariates, and to capture within-group correlations and cross-response correlations in the data through a random-effects structure.

[0005] In conjunction with the first aspect, in some implementations of the first aspect, the method further includes: within the two-stage segmented processing, the first stage uses Gaussian variational approximation to perform parameter optimization estimation on the univariate Boson mixed-effects model combined with microbial count data to obtain the first-stage estimated value.

[0006] In conjunction with the first aspect, in some implementations of the first aspect, the method further includes: within the two-stage segmented processing, the second stage employs the moment estimation method for inter-response dependencies, and combines the first-stage estimation value to estimate the cross-response correlation between multiple microbial count data collected from multiple microbial samples over multiple time periods, thereby obtaining the second-stage estimation value.

[0007] In conjunction with the first aspect, in some implementations of the first aspect, the method further includes: the parameters in the multivariate Posson mixed-effects model include fixed-effects coefficient parameters and random-effects covariance parameters, wherein the random-effects covariance parameters include random-effects variance and correlation coefficient parameters.

[0008] In conjunction with the first aspect, in some implementations of the first aspect, the method further includes: the Gaussian variational approximation process, comprising the following steps: A variational distribution is constructed for each microbial sample; The variational lower bound of the univariate Posson mixture effect model is generated based on the variational distribution and the univariate Posson mixture effect model. The fixed effects coefficient and random effects variance parameters of the univariate Posson mixed-effects model are optimized by maximizing the variational lower bound of the model until the convergence condition is met. The estimated values ​​of the fixed effects coefficient and random effects variance parameters are then obtained as the first-stage estimates. The variational distribution is a Gaussian distribution.

[0009] In conjunction with the first aspect, in some implementations of the first aspect, the method further includes: a process for estimating the moment of the inter-response dependency, comprising the following steps: Based on various microbial count data collected over multiple time periods, the cross-product sample moments and the sum-product sample moments between different microbial count data are calculated. Based on the estimated fixed-effects coefficient parameters obtained from the first-stage estimation and the known covariate data, the model prediction moments between different microbial count data are calculated. Based on the cross-product sample moments, the sum-product sample moments, the model prediction moments, and the estimated values ​​of the random effects variance parameters from the first stage estimation, the estimated values ​​of the random effects correlation coefficient parameters are calculated using the moment estimation formula and used as the estimated values ​​for the second stage.

[0010] In conjunction with the first aspect, in some implementations of the first aspect, the method further includes: the calculation expression for the Gaussian variational approximation is as follows: In the formula, Let represent the marginal likelihood function of the univariate Poisson mixed-effects model for the k-th microorganism. Represents the probability density function. Let represent the probability density function of the introduced Gaussian variational distribution. This represents the lower bound function of the Gaussian variational approximation for the k-th microorganism. This represents the k-th microbial count data measured at time point j for the i-th sample. Let represent the fixed-effects covariate vector of the i-th sample at time point j concerning the k-th microorganism. This represents the vector of fixed-effect coefficients for the k-th microorganism. This represents the variance of the random effects of the k-th microorganism. This represents the random effect of the k-th microorganism in the i-th sample. This represents the mean of the variational distribution introduced for the k-th microorganism in the i-th sample. The model parameters represent the variational distribution variance introduced for the k-th microorganism in the i-th sample. Variational parameters First-stage parameter estimates It is obtained by maximizing the variational lower bound.

[0011] In conjunction with the first aspect, in some implementations of the first aspect, the method further includes: the calculation process of the moment estimation of the inter-response dependencies, as follows: In the formula, and Let the cross product sample moments and the sum product sample moments be represented, respectively, between the k-th and l-th microbial count data. and This represents the model prediction moments between the k-th and l-th microbial count data. This represents the estimated correlation coefficient between the k-th and l-th microorganisms.

[0012] Secondly, in order to achieve the above objectives, the present invention discloses a multivariate counting data processing system for longitudinal microbiome data, comprising: The data acquisition module is used to acquire a multivariate microbial count dataset; wherein, the multivariate microbial count dataset includes multiple microbial count data collected from multiple microbial samples over multiple time periods and corresponding fixed-effect covariates; The parameter estimation module is used to input the microbial multivariate count dataset into a pre-established multivariate Boson mixed effect model for two-stage segmentation processing and output microbial data estimates, wherein the microbial data estimates include first-stage estimates and second-stage estimates. The multivariate Boson mixed-effects model is used to determine the dependencies between various microbial count data collected from multiple microbial samples over multiple time periods and the corresponding fixed-effects covariates, and to capture within-group correlations and cross-response correlations in the data through a random-effects structure.

[0013] In another aspect of the present invention, in order to achieve the above-mentioned objective, a terminal device is disclosed, including a memory, a processor, and a computer program stored in the memory and capable of running on the processor. The memory stores the computer program capable of running on the processor, and when the processor loads and executes the computer program, it employs a multivariate counting data processing method for longitudinal microbiome data as described above.

[0014] The beneficial effects of this invention are: This invention decomposes the high-dimensional integral problem into two stages: low-dimensional optimization and moment estimation. While ensuring estimation accuracy, it improves the computation speed by two orders of magnitude, enabling efficient analysis of large-scale microbiome data and enhancing the estimation efficiency of large-scale microbiome data. Attached Figure Description

[0015] To more clearly illustrate the technical solutions in the embodiments of the present invention or the prior art, the drawings used in the description of the embodiments or the prior art will be briefly introduced below. Obviously, for those skilled in the art, other drawings can be obtained based on these drawings without creative effort. Figure 1 This is a schematic diagram of the method flow of the present invention; Figure 2 This is a schematic diagram comparing the running time of the present invention with that of the prior art; Figure 3 This is a schematic diagram comparing the estimation accuracy of the first microbial immobility effect coefficient parameter between the present invention and existing technologies; Figure 4 This is a schematic diagram comparing the estimation accuracy of the second microbial immobility effect coefficient parameter between the present invention and the prior art; Figure 5 This is a schematic diagram comparing the estimation accuracy of random effects variance and correlation coefficient parameters of the present invention and existing technologies; Figure 6 This is a schematic diagram of the system structure of the present invention. Detailed Implementation

[0016] The technical solutions of the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of the present invention, and not all embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of the present invention.

[0017] Example 1: like Figure 1 As shown, a method for processing multivariate count data of longitudinal microbiome data includes the following steps: S101: Obtain a multivariate microbial count dataset; wherein, the multivariate microbial count dataset includes multiple microbial count data collected from multiple microbial samples over multiple time periods and corresponding fixed-effect covariates; Multivariate longitudinal count data, which involves repeated measurements of multiple outcomes over time, are prevalent in fields such as medical research, economics, and social sciences. While the statistical literature offers various methodologies for analyzing such data, modeling their complex dependency structures remains a key challenge. Poisson mixture models have proven effective for univariate longitudinal count data, handling intra-subject correlations well while maintaining interpretability. However, univariate PMMs for each outcome inherently fail to capture inter-response dependencies.

[0018] S102: Input the microbial multivariate count dataset into a pre-established multivariate Boson mixed-effects model for two-stage segmentation processing, and output the microbial data estimate, wherein the microbial data estimate includes the first-stage estimate and the second-stage estimate. The multivariate Boson mixed-effects model is used to determine the dependencies between various microbial count data collected from multiple microbial samples over multiple time periods and the corresponding fixed-effects covariates, and to capture within-group correlations and cross-response correlations in the data through a random-effects structure.

[0019] Within the two-stage segmented processing, the first stage uses Gaussian variational approximation to optimize and estimate the parameters of the univariate Boson mixed-effects model in combination with microbial count data, thus obtaining the first-stage estimate.

[0020] Within the two-stage segmented processing, the second stage employs the moment estimation method for inter-response dependencies. It combines the estimates from the first stage to estimate the cross-response correlation between various microbial count data collected from multiple microbial samples over multiple time periods, thus obtaining the second-stage estimates.

[0021] The parameters in the multivariate Posson mixed-effects model include fixed-effects coefficient parameters and random-effects covariance parameters, wherein the random-effects covariance parameters include random-effects variance and correlation coefficient parameters.

[0022] The Gaussian variational approximation process includes the following steps: A variational distribution is constructed for each microbial sample; The variational lower bound of the univariate Posson mixture effect model is generated based on the variational distribution and the univariate Posson mixture effect model. The fixed effects coefficient and random effects variance parameters of the univariate Posson mixed-effects model are optimized by maximizing the variational lower bound of the model until the convergence condition is met. The estimated values ​​of the fixed effects coefficient and random effects variance parameters are then obtained as the first-stage estimates. The variational distribution is a Gaussian distribution.

[0023] The computational expression for the Gaussian variational approximation is as follows: In the formula, Let represent the marginal likelihood function of the univariate Poisson mixed-effects model for the k-th microorganism. Represents the probability density function. Let represent the probability density function of the introduced Gaussian variational distribution. This represents the lower bound function of the Gaussian variational approximation for the k-th microorganism. This represents the k-th microbial count data measured at time point j for the i-th sample. Let represent the fixed-effects covariate vector of the i-th sample at time point j concerning the k-th microorganism. This represents the vector of fixed-effect coefficients for the k-th microorganism. This represents the variance of the random effects of the k-th microorganism. This represents the random effect of the k-th microorganism in the i-th sample. This represents the mean of the variational distribution introduced for the k-th microorganism in the i-th sample. The model parameters represent the variational distribution variance introduced for the k-th microorganism in the i-th sample. Variational parameters First-stage parameter estimates It is obtained by maximizing the variational lower bound.

[0024] The process of moment estimation of inter-response dependencies includes the following steps: Based on various microbial count data collected over multiple time periods, the cross-product sample moments and the sum-product sample moments between different microbial count data are calculated. Based on the estimated fixed-effects coefficient parameters obtained from the first-stage estimation and the known covariate data, the model prediction moments between different microbial count data are calculated. Based on the cross-product sample moments, the sum-product sample moments, the model prediction moments, and the estimated values ​​of the random effects variance parameters from the first stage estimation, the estimated values ​​of the random effects correlation coefficient parameters are calculated using the moment estimation formula and used as the estimated values ​​for the second stage.

[0025] The calculation process for the moment estimation of the dependence between responses is as follows: In the formula, and Let the cross product sample moments and the sum product sample moments be represented, respectively, between the k-th and l-th microbial count data. and This represents the model prediction moments between the k-th and l-th microbial count data. This represents the estimated correlation coefficient between the k-th and l-th microorganisms.

[0026] Specifically, the present invention will be further illustrated below through embodiments: Multivariate Poisson mixed-effects models extend univariate models by simultaneously analyzing multiple counts. This model explicitly quantifies the associations between outcomes and assesses covariate effects across the entire response vector. However, a key challenge lies in marginal likelihood, which involves high-dimensional integrals over random effects or latent variables, which are cumbersome and lack closed-form solutions. Some numerical methods, such as the Gauss-Hermite quadrature and the Laplace approximation, still face problems such as high computational cost when dealing with complex integrals in mPMMs. Furthermore, there is a lack of readily available tools to implement these methods for mPMMs. Although sampling-based MCMC methods have been extended to multivariate settings via the R package MCMCglmm, they suffer from computational inefficiency.

[0027] Variational approximation offers a computationally efficient approach by transforming the integral problem into an optimization problem. This method approximates the tractable likelihood by minimizing the Kullback-Leibler divergence between a tractable variational distribution and the true posterior of the random effects, providing a computational advantage in high-dimensional settings.

[0028] This invention proposes a frequency-based method for fitting mPMM, integrating Gaussian variational approximation and moment-based estimation. This method offers higher computational efficiency than MCMC while maintaining consistency of estimators. It decomposes the estimation problem by employing variational approximation to estimate fixed-effects coefficients and random-effects variances, and by using moments to estimate correlation coefficients between random effects to capture inter-response dependencies, thus resolving the estimation problem. Crucially, this method is not a simple extension of Gaussian variational approximation to mPMM; it allows for separate univariate estimations while preserving inter-response dependencies. It also differs from the classical variational Bayesian inference framework, which does not guarantee consistency of mPMM parameter estimators.

[0029] The method of this invention was applied to a diet and microbiome dataset. This dataset contained daily 24-hour dietary records and fecal shotgun metagenomic data from 34 healthy human subjects over 17 days. *Clostridium difficile* and *Enterococcus* have been shown to have a synergistic effect in enhancing pathogenicity; however, how they are influenced by specific variables remains unclear. This invention used *Clostridium difficile* and *Enterococcus* as response variables and butyric acid, caprylic acid, dodecanoic acid, added vitamin E, niacin, and eicosapentaenoic acid as covariates. Significant variables (pMCMC < 0.05) were selected using bidirectional stepwise regression using the R package MCMCglmm, and are summarized in Table 1. The invention then fitted the following mPMM, incorporating random effects on the subjects: As shown in Table 2, the two methods yielded comparable parameter estimates. GVA-M (the method of this invention) ) and MCMC ( 63) The estimates of the correlations with random effects were all positive, consistent with known microbial synergies (Enterococcus enhances Clostridium difficile pathogenicity). However, GVA-M achieved a 31× speedup over MCMC, reducing computation time from 14.86 seconds to 0.48 seconds, demonstrating its high efficiency in microbiome data analysis.

[0030] Table 1 Table 2 This embodiment uses a binary case (K=2) as an example to conduct a simulation experiment to verify the effectiveness of the method proposed in this invention. It should be noted that for the high-dimensional case of K>2, those skilled in the art can directly extend the same algorithm flow, which will not be elaborated here.

[0031] Consider a bivariate Poisson mixture model, where the random effects follow a bivariate normal distribution, and the fixed effects coefficients and covariate parameters are configured according to typical settings. Under different individual sample sizes (from small to large samples), the method of this invention is compared and evaluated with existing MFVB methods (an approximate inference method based on mean-field variational Bayes) and MCMC methods.

[0032] The experiment was based on 100 independent, repeated simulations. Evaluation metrics included computation time, relative bias of parameter estimation, and mean square error. Figure 2 As shown, the method of this invention significantly outperforms the comparative methods in terms of computational efficiency, reducing the average running time by approximately 89% compared to the MFVB method and by over 98% compared to the MCMC method. Furthermore, this computational advantage remains stable as the sample size increases. Regarding parameter estimation accuracy, as... Figures 3 to 5 As shown, the method of the present invention can obtain estimation results comparable to those of the comparative methods.

[0033] The above experimental data fully verify the effectiveness and superiority of the method proposed in this invention.

[0034] Example 2: To achieve the above objective, such as Figure 6 As shown, based on Embodiment 1, this invention discloses a multivariate counting data processing system for longitudinal microbiome data, comprising: The data acquisition module 11 is used to acquire a microbial multivariate count dataset; wherein, the microbial multivariate count dataset includes multiple microbial count data collected from multiple microbial samples over multiple time periods and corresponding fixed-effect covariates; The parameter estimation module 12 is used to input the microbial multivariate count dataset into a pre-established multivariate Boson mixed effect model for two-stage segmentation processing and output microbial data estimates, wherein the microbial data estimates include first-stage estimates and second-stage estimates. The multivariate Boson mixed-effects model is used to determine the dependencies between various microbial count data collected from multiple microbial samples over multiple time periods and the corresponding fixed-effects covariates, and to capture within-group correlations and cross-response correlations in the data through a random-effects structure.

[0035] Based on the same inventive concept, this invention also provides a computer device, comprising: one or more processors, and a memory for storing one or more computer programs; the programs include program instructions, and the processor executes the program instructions stored in the memory. The processor may be a Central Processing Unit (CPU), or other general-purpose processors, digital signal processors (DSPs), application-specific integrated circuits (ASICs), field-programmable gate arrays (FPGAs), or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components, etc. It is the computing and control core of the terminal, used to implement one or more instructions, specifically for loading and executing one or more instructions stored in a computer storage medium to implement the above-described method.

[0036] It should be further explained that, based on the same inventive concept, the present invention also provides a computer storage medium storing a computer program, which, when executed by a processor, performs the above-described method. This storage medium can be any combination of one or more computer-readable media. The computer-readable medium can be a computer-readable signal medium or a computer-readable storage medium. The computer-readable storage medium can be, for example, but not limited to, an electrical, magnetic, optical, electromagnetic, infrared, or semiconductor system, apparatus, or device, or any combination thereof. More specific examples of computer-readable storage media (a non-exhaustive list) include: an electrical connection having one or more wires, a portable computer disk, a hard disk, random access memory (RAM), read-only memory (ROM), erasable programmable read-only memory (EPROM or flash memory), optical fiber, portable compact disk read-only memory (CD-ROM), optical storage device, magnetic storage device, or any suitable combination thereof. In the present invention, the computer-readable storage medium can be any tangible medium containing or storing a program that can be used by or in conjunction with an instruction execution system, apparatus, or device.

[0037] In the description of this specification, references to terms such as "an embodiment," "example," "specific example," etc., indicate that a specific feature, structure, material, or characteristic described in connection with that embodiment or example is included in at least one embodiment or example of this disclosure. In this specification, the illustrative expressions of the above terms do not necessarily refer to the same embodiment or example. Furthermore, the specific features, structures, materials, or characteristics described may be combined in any suitable manner in one or more embodiments or examples.

[0038] The foregoing has shown and described the basic principles, main features, and advantages of this disclosure. Those skilled in the art should understand that this disclosure is not limited to the above embodiments. The embodiments and descriptions in the specification are merely illustrative of the principles of this disclosure. Various changes and modifications can be made to this disclosure without departing from its spirit and scope, and all such changes and modifications fall within the scope of this disclosure as claimed.

Claims

1. A method for processing multivariate count data of longitudinal microbiome data, characterized in that, The method includes the following steps: Obtain a multivariate microbial count dataset; wherein, the multivariate microbial count dataset includes multiple microbial count data collected from multiple microbial samples over multiple time periods and corresponding fixed-effect covariates; The microbial multivariate count dataset is input into a pre-established multivariate Boson mixed-effects model for two-stage segmentation processing, and the output is the microbial data estimate, wherein the microbial data estimate includes the first stage estimate and the second stage estimate. The multivariate Boson mixed-effects model is used to determine the dependencies between various microbial count data collected from multiple microbial samples over multiple time periods and the corresponding fixed-effects covariates, and to capture within-group correlations and cross-response correlations in the data through a random-effects structure.

2. The method for processing multivariate counting data of longitudinal microbiome data according to claim 1, characterized in that, Within the two-stage segmented processing, the first stage uses Gaussian variational approximation to optimize and estimate the parameters of the univariate Boson mixed-effects model in combination with microbial count data, thereby obtaining the first-stage estimated value.

3. The method for processing multivariate counting data of longitudinal microbiome data according to claim 2, characterized in that, Within the two-stage segmented processing, the second stage employs the moment estimation method for inter-response dependencies, combining the first-stage estimates to estimate the cross-response correlation between various microbial count data collected from multiple microbial samples over multiple time periods, thus obtaining the second-stage estimates.

4. The method for processing multivariate counting data of longitudinal microbiome data according to claim 1, characterized in that, The parameters in the multivariate Boson mixed-effects model include fixed-effects coefficient parameters and random-effects covariance parameters, wherein the random-effects covariance parameters include random-effects variance and correlation coefficient parameters.

5. The method for processing multivariate counting data of longitudinal microbiome data according to claim 3, characterized in that, The Gaussian variational approximation process includes the following steps: A variational distribution is constructed for each microbial sample; The variational lower bound of the univariate Posson mixture effect model is generated based on the variational distribution and the univariate Posson mixture effect model. The fixed effects coefficient and random effects variance parameters of the univariate Posson mixed-effects model are optimized by maximizing the variational lower bound of the model until the convergence condition is met. The estimated values ​​of the fixed effects coefficient and random effects variance parameters are then obtained as the first-stage estimates. The variational distribution is a Gaussian distribution.

6. The method for processing multivariate counting data of longitudinal microbiome data according to claim 5, characterized in that, The process of the moment estimation method for inter-response dependencies includes the following steps: Based on various microbial count data collected over multiple time periods, the cross-product sample moments and the sum-product sample moments between different microbial count data are calculated. Based on the estimated fixed-effects coefficient parameters obtained from the first-stage estimation and the known covariate data, the model prediction moments between different microbial count data are calculated. Based on the cross-product sample moments, the sum-product sample moments, the model prediction moments, and the estimated values ​​of the random effects variance parameters from the first stage estimation, the estimated values ​​of the random effects correlation coefficient parameters are calculated using the moment estimation formula and used as the estimated values ​​for the second stage.

7. The method for processing multivariate counting data of longitudinal microbiome data according to claim 6, characterized in that, The calculation expression for the Gaussian variational approximation is as follows: In the formula, Let represent the marginal likelihood function of the univariate Poisson mixed-effects model for the k-th microorganism. Represents the probability density function. Let represent the probability density function of the introduced Gaussian variational distribution. This represents the lower bound function of the Gaussian variational approximation for the k-th microorganism. This represents the k-th microbial count data measured at time point j for the i-th sample. Let represent the fixed-effects covariate vector of the i-th sample at time point j concerning the k-th microorganism. This represents the vector of fixed-effect coefficients for the k-th microorganism. This represents the variance of the random effects of the k-th microorganism. This represents the random effect of the k-th microorganism in the i-th sample. This represents the mean of the variational distribution introduced for the k-th microorganism in the i-th sample. The model parameters represent the variational distribution variance introduced for the k-th microorganism in the i-th sample. Variational parameters First-stage parameter estimates It is obtained by maximizing the variational lower bound.

8. The method for processing multivariate counting data of longitudinal microbiome data according to claim 7, characterized in that, The calculation process for the moment estimation of the inter-response dependency is as follows: In the formula, and Let the cross product sample moments and the sum product sample moments be represented, respectively, between the k-th and l-th microbial count data. and This represents the model prediction moments between the k-th and l-th microbial count data. This represents the estimated correlation coefficient between the k-th and l-th microorganisms.

9. A multivariate counting data processing system for longitudinal microbiome data, employing the multivariate counting data processing method for longitudinal microbiome data as described in any one of claims 1 to 8, characterized in that, include: The data acquisition module is used to acquire a multivariate microbial count dataset; wherein, the multivariate microbial count dataset includes multiple microbial count data collected from multiple microbial samples over multiple time periods and corresponding fixed-effect covariates; The parameter estimation module is used to input the microbial multivariate count dataset into a pre-established multivariate Boson mixed effect model for two-stage segmentation processing and output microbial data estimates, wherein the microbial data estimates include first-stage estimates and second-stage estimates. The multivariate Boson mixed-effects model is used to determine the dependencies between various microbial count data collected from multiple microbial samples over multiple time periods and the corresponding fixed-effects covariates, and to capture within-group correlations and cross-response correlations in the data through a random-effects structure.

10. A terminal device, comprising a memory, a processor, and a computer program stored in the memory and capable of running on the processor, characterized in that, The memory stores a computer program that can run on a processor. When the processor loads and executes the computer program, it employs a multivariate counting data processing method for longitudinal microbiome data according to any one of claims 1 to 8.