Longitudinal data high-dimensional parameter estimation method based on Bayesian empirical likelihood and variational inference
By using the Bayesian empirical likelihood and variational inference methods, an approximate posterior distribution is constructed, which solves the problem of difficult calculation of the posterior distribution in high-dimensional medical data and achieves fast and accurate disease risk diagnosis.
Patent Information
- Application Number
- CN202510910233.3
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-07-02
- Publication Date
- 2025-10-17
- Estimated Expiration
- Not applicable · inactive patent
AI Technical Summary
Traditional Bayesian likelihood is difficult to calculate the posterior distribution in high-dimensional medical data analysis, has low accuracy, poor real-time performance, and cannot quickly and accurately diagnose patients' disease risks.
A method based on Bayesian empirical likelihood and variational inference is used to construct a theoretical posterior distribution by standardizing the data. The variational inference method is used to generate an approximate posterior distribution, estimate high-dimensional parameters, and generate a high-dimensional parameter estimation table to assist in disease risk diagnosis.
It achieves fast and accurate calculation of the posterior distribution of Bayesian likelihood estimation, improving the accuracy and efficiency of disease risk diagnosis.
Smart Images

Figure CN120809151A_ABST
Abstract
Description
TECHNICAL FIELD
[0001] The application belongs to the field of intelligent auxiliary diagnosis, and particularly relates to a longitudinal data high-dimensional parameter estimation method based on Bayesian empirical likelihood and variational inference. BACKGROUND
[0002] With the development of information medical technology, a Bayesian likelihood auxiliary medical diagnosis technology appears. Traditional Bayesian likelihood strictly follows the hierarchical modeling principle in high-dimensional medical data analysis: firstly, a joint posterior distribution containing hyperparameters is constructed; then, a Markov chain Monte Carlo algorithm is iteratively sampled to generate a dependent sample chain to approximate the posterior distribution. However, with the increase of data dimension, the posterior distribution is difficult to calculate, the accuracy is low, and the real-time performance is poor. SUMMARY
[0003] Therefore, it is necessary to provide a longitudinal data high-dimensional parameter estimation method based on Bayesian empirical likelihood and variational inference, which can quickly calculate the posterior distribution of Bayesian likelihood estimation and has the effect of quickly and accurately diagnosing the disease risk of a patient.
[0004] In a first aspect, the application provides a longitudinal data high-dimensional parameter estimation method based on Bayesian empirical likelihood and variational inference, comprising:
[0005] obtaining longitudinal data and performing standardization processing on the longitudinal data to obtain standardized data; the longitudinal data is data used for recording physiological states of each individual at different time points in history obtained from a medical database;
[0006] based on the standardized data, obtaining a theoretical posterior distribution of the longitudinal data through Bayesian empirical likelihood estimation;
[0007] based on the theoretical posterior distribution, constructing an approximate posterior distribution similar to the theoretical posterior distribution through a variational inference method;
[0008] based on the approximate posterior distribution, estimating high-dimensional parameters of the longitudinal data to obtain a high-dimensional parameter estimation table corresponding to the longitudinal data; the high-dimensional parameter estimation table is used for auxiliary diagnosis of disease risks of the patient corresponding to the longitudinal data.
[0009] Further, obtaining longitudinal data and performing standardization processing on the longitudinal data to obtain standardized data comprises:
[0010] extracting longitudinal data from a medical database and sorting the longitudinal data according to time stamps to obtain an original longitudinal data set;
[0011] identifying missing values in the original longitudinal data set and performing missing value filling to obtain a complete longitudinal data set;
[0012] time aligning the complete longitudinal data set to obtain a time-aligned longitudinal data set;
[0013] based on the time-aligned longitudinal data set, standardized data is calculated by the following formula:
[0014]
[0015] wherein, is standardized data, representing standardized data of physiological indicator j of patient i at time point t, is an original measurement value of physiological indicator j of patient i at time point t, is a global mean value of physiological indicator j, is a global standard deviation of physiological indicator j, is an intermediate value after global standardization, is an individual historical mean value of physiological indicator j of patient i, is an individual historical standard deviation of physiological indicator j of patient i.
[0016] Further, based on the approximate posterior distribution, high-dimensional parameters of the longitudinal data are estimated to obtain a high-dimensional parameter estimation table, and the method further comprises:
[0017] obtaining original data of a new patient, and standardizing the original data of the new patient to obtain new standardized data;
[0018] sampling a parameter posterior mean, a parameter posterior covariance and a covariance parameter from the approximate posterior distribution, and constructing them as effective parameters;
[0019] based on the new standardized data and the effective parameters, predicting a disease risk to obtain a disease probability matrix;
[0020] based on the disease probability matrix and the high-dimensional parameter table estimation table, generating a risk assessment report.
[0021] Further, based on the standardized data, a theoretical posterior distribution of the longitudinal data is obtained by Bayesian empirical likelihood estimation, comprising:
[0022] based on the standardized data, a generalized estimation equation is defined by the following formula:
[0023]
[0024] wherein, U i (θ) is an estimation function of the i th individual, D i is a derivative matrix, V i is a covariance matrix, μ i (θ) is a marginal mean vector, Y i is an observation variable of a response variable of the i th individual, obtained from the standardized data;
[0025] constructing a decorrelation function based on generalized estimating equations and working correlation matrix;
[0026] constructing an empirical likelihood function based on the decorrelation function;
[0027] setting a prior distribution based on a preset medical prior knowledge base to obtain a parameter prior;
[0028] constructing a posterior distribution based on the empirical likelihood function, the decorrelation function and the parameter prior through the following formula to obtain a theoretical posterior distribution:
[0029] p (θ, γ, R | data) ∝ L e (θ) · p (θ | γ) p (γ) · p (R)
[0030] wherein, p (θ, γ, R | data) is the theoretical posterior distribution, L e (θ) is the empirical likelihood function, p (θ | γ) p (γ) is the parameter prior, and p (R) is the correlation prior.
[0031] Further, based on the decorrelation function, an empirical likelihood function is constructed, including:
[0032] solving the following formula by Lagrange multiplier method to obtain the empirical weight:
[0033]
[0034] wherein, λ (θ) satisfies the constraint formula:
[0035]
[0036] wherein, p i (θ) is the empirical weight of the i th observation, λ (θ) is the Lagrange multiplier vector, n is the sample number, U i * (θ) is the decorrelation function;
[0037] obtaining the empirical likelihood function through the following formula according to the empirical weight:
[0038]
[0039] wherein, L e (θ) is the empirical likelihood function, and p i (θ) is the empirical weight of the i th observation.
[0040] Further, based on the theoretical posterior distribution, an approximate posterior distribution similar to the theoretical posterior distribution is constructed by variational inference method, including:
[0041] Based on the theoretical posterior distribution, a variational distribution family is selected, and the variational parameters of the variational distribution family are initialized to obtain initial parameters;
[0042] Based on the theoretical posterior distribution and the variational distribution family, an evidence lower bound is constructed through the following formula to obtain a computable ELBO expression:
[0043]
[0044] Wherein, L(λ) is the objective function of the evidence lower bound, λ is the Lagrange multiplier vector, K is the number of Monte Carlo sampling, n is the sample number, θ (k) is the parameter vector of the kth sampling, U i * (θ) is a decorrelation function, γ is a prior distribution hyperparameter, R is a working correlation matrix, p(θ|γ), p(γ) is a parameter prior, q(θ,γ,R) is a variational distribution;
[0045] Based on the initial parameters and the computable ELBO expression, the variational parameters are optimized through the following formula to obtain optimized parameters:
[0046]
[0047] Wherein, λ (t+1) is the optimized parameter, λ (t) is the parameter before optimization, t is the iteration number, p t is an adaptive step size, F -1 is the inverse of the information matrix, is the gradient of ELBO with respect to λ;
[0048] Based on the optimized parameters, a complete posterior distribution is constructed to obtain an approximate posterior distribution.
[0049] Further, based on the optimized parameters, a complete posterior distribution is constructed to obtain an approximate posterior distribution, including:
[0050] The parameter posterior mean vector, the parameter posterior covariance matrix, the variable selection probability vector, the inverse Wishart degree of freedom and the inverse Wishart scale matrix are obtained from the optimized parameters;
[0051] Based on the parameter posterior mean vector, the parameter posterior covariance matrix, the variable selection probability vector, the inverse Wishart degree of freedom and the inverse Wishart scale matrix, the approximate posterior distribution is constructed through the following formula:
[0052]
[0053] Wherein, q * (θ,γ,R) is the approximate posterior distribution, is the parameter posterior mean vector, is a parameter posterior covariance matrix, is a variable selection probability vector, v * is an inverse Wishart degree of freedom, Ψ * is an inverse Wishart scale matrix, γ is a hyperparameter of the prior distribution, and R is a working correlation matrix.
[0054] Further, based on the approximate posterior distribution, the high-dimensional parameters of the longitudinal data are estimated, and a high-dimensional parameter estimation table is obtained, including:
[0055] Key statistics in the approximate posterior distribution are extracted; the key statistics include: posterior mean, marginal variance, and selection probability;
[0056] Based on the key statistics, the point estimate vector and the standard error vector are calculated by the following formula:
[0057]
[0058] wherein, is one of the elements of the point estimate vector, is the posterior mean, SE j is one of the elements of the standard error vector, is the marginal variance;
[0059] Based on the key statistics, a confidence interval is constructed using the normal approximation method;
[0060] The key statistics, the point estimate vector, the standard error vector, and the confidence interval are integrated to obtain the high-dimensional parameter estimation table.
[0061] In a second aspect, the present application also provides a longitudinal data high-dimensional parameter estimation device based on Bayesian empirical likelihood and variational inference, comprising:
[0062] A standardization module is configured to obtain longitudinal data and perform standardization processing on the longitudinal data to obtain standardized data; the longitudinal data is data used to record physiological states of each individual at different time points in history obtained from a medical database;
[0063] A Bayesian module is configured to obtain a theoretical posterior distribution of the longitudinal data by Bayesian empirical likelihood estimation based on the standardized data;
[0064] A variational inference module is configured to construct an approximate posterior distribution similar to the theoretical posterior distribution by a variational inference method based on the theoretical posterior distribution;
[0065] An estimation module is configured to estimate high-dimensional parameters of the longitudinal data based on the approximate posterior distribution, and obtain a high-dimensional parameter estimation table corresponding to the longitudinal data; the high-dimensional parameter estimation table is used to assist in diagnosing the disease risk of the patient corresponding to the longitudinal data.
[0066] In a third aspect, the present application also provides a computer device comprising a memory and a processor, wherein the memory stores a computer program, and the processor implements the steps of the method provided in the first aspect of the present application when executing the computer program.
[0067] The longitudinal data high-dimensional parameter estimation method based on Bayesian empirical likelihood and variational inference provided by the present application comprises the following steps: obtaining longitudinal data, and performing standardization processing on the longitudinal data to obtain standardized data; the longitudinal data is data of each individual at different time points in history for recording physiological state obtained from a medical database; based on the standardized data, a theoretical posterior distribution of the longitudinal data is obtained through Bayesian empirical likelihood estimation; based on the theoretical posterior distribution, an approximate posterior distribution similar to the theoretical posterior distribution is constructed through a variational inference method; based on the approximate posterior distribution, high-dimensional parameters of the longitudinal data are estimated to obtain a high-dimensional parameter estimation table corresponding to the longitudinal data; and the high-dimensional parameter estimation table is used as a technical means for assisting in diagnosing disease risk of the patient corresponding to the longitudinal data. The posterior distribution of the Bayesian likelihood estimation can be quickly calculated, and the effect of quickly and accurately diagnosing the disease risk of the patient is achieved. BRIEF DESCRIPTION OF DRAWINGS
[0068] In order to more clearly illustrate the technical solutions in the embodiments of the present application or the related art, the drawings needed to be used in the embodiment or related art description will be briefly introduced. Obviously, the drawings in the following description are only some embodiments of the present application, and other drawings can be obtained by those skilled in the art without creative labor.
[0069] Figure 1 A longitudinal data high-dimensional parameter estimation method based on Bayesian empirical likelihood and variational inference provided by the present application is shown in the flowchart.
[0070] Figure 2 A longitudinal data high-dimensional parameter estimation device structure provided by the present application is shown in the structural diagram. DETAILED DESCRIPTION
[0071] In order to make the purpose, technical scheme and advantages of the present application more clear, the present application will be further described in detail below in combination with the drawings and embodiments. It should be understood that the specific embodiments described herein are only used to explain the present application, and are not used to limit the present application.
[0072] In one embodiment, as Figure 1As shown, a longitudinal data high-dimensional parameter estimation method based on Bayesian empirical likelihood and variational inference is provided. In this embodiment, the method is applied to a terminal. It should be understood that the method can also be applied to a server, and can also be applied to a system including a terminal and a server, and can be realized through the interaction of the terminal and the server. The method includes the following steps:
[0073] In step 101, longitudinal data is obtained, and the longitudinal data is standardized to obtain standardized data. The longitudinal data is data obtained from a medical database for recording physiological states of each individual at different time points in history.
[0074] The standardized data is physiological index data that eliminates dimensional differences and individual baseline fluctuations. The terminal performs data cleaning on the obtained longitudinal data, which can include timestamp sorting, missing value filling, and time alignment. The data after time alignment is double-standardized to eliminate dimensional differences between the longitudinal data, so that the longitudinal data is comparable, and standardized data is obtained.
[0075] In step 102, based on the standardized data, the theoretical posterior distribution of the longitudinal data is obtained through Bayesian empirical likelihood function estimation.
[0076] Specifically, the theoretical posterior distribution is the true probability distribution of the parameter θ given the data. Bayesian empirical likelihood is a hybrid estimation method combining empirical likelihood and Bayesian framework. Based on the standardized data, a generalized estimating equation (GEE) is constructed, the generalized estimating equation is decorrelated, a correlation likelihood function is constructed, a parameter prior distribution is introduced, and a theoretical posterior distribution is constructed. The parameter prior distribution can include conjugate distribution, non-informative prior, spike-and-plate prior, horseshoe prior, etc.
[0077] In step 103, based on the theoretical posterior distribution, an approximate posterior distribution similar to the theoretical posterior distribution is constructed by a variational inference method.
[0078] Specifically, the approximate posterior distribution is a distribution that is easier to handle than the theoretical posterior distribution. Variational inference is a method for approximating complex posteriors with simple distributions. The terminal selects a variational distribution family and constructs an evidence lower bound (ELBO), and then optimizes the parameters of the variational distribution based on the evidence lower bound to output an approximate posterior distribution similar to the theoretical posterior distribution.
[0079] In step 104, based on the approximate posterior distribution, the high-dimensional parameters of the longitudinal data are estimated to obtain a high-dimensional parameter estimation table corresponding to the longitudinal data. The high-dimensional parameter estimation table is used to assist in diagnosing the disease risk of the patient corresponding to the longitudinal data.
[0080] Specifically, a high-dimensional parameter estimate table is a structured output of key statistics. High-dimensional parameters are parameters whose dimensions far exceed the number of samples. The terminal extracts statistics from the approximate posterior distribution, constructs credible intervals, and then integrates them into a high-dimensional parameter estimate table.
[0081] The high-dimensional parameter estimation method for longitudinal data based on Bayesian empirical likelihood and variational inference provided in this embodiment obtains longitudinal data and performs standardization on the longitudinal data to obtain standardized data; the longitudinal data is data obtained from a medical database for recording the physiological state of each individual at different time points in history; based on the standardized data, the theoretical posterior distribution of the longitudinal data is obtained by estimating the Bayesian empirical likelihood function; based on the theoretical posterior distribution, an approximate posterior distribution similar to the theoretical posterior distribution is constructed by the variational inference method; based on the approximate posterior distribution, the high-dimensional parameters of the longitudinal data are estimated to obtain a high-dimensional parameter estimation table corresponding to the longitudinal data; the high-dimensional parameter estimation table is used to assist in diagnosing the disease risk of the patient corresponding to the longitudinal data. Through the above technical means, the posterior distribution of the Bayesian likelihood estimate can be quickly calculated, achieving the effect of quickly and accurately diagnosing the patient's disease risk.
[0082] In one embodiment, obtaining longitudinal data and performing normalization processing on the longitudinal data to obtain normalized data includes:
[0083] Step 201 : extract longitudinal data from a medical database and sort them by timestamp to obtain an original longitudinal data set.
[0084] Specifically, the original longitudinal dataset is a data table sorted in ascending timestamp order, where each row may contain a patient ID, a timestamp, and the measured value of a specific physiological indicator. For example, the terminal retrieves all historical records of the target patient from the medical database. Fields may include patient identifier, measurement time, and physiological indicator value. The terminal then groups each patient, sorting the data within the group based on the timestamp field.
[0085] Step 202 : Identify missing values in the original longitudinal dataset and perform missing value filling to obtain a complete longitudinal dataset.
[0086] Specifically, a complete longitudinal data set is a longitudinal data table with no missing values, and its completeness criterion is that all cells contain valid values and the values are within the reasonable range for the human body. The terminal scans each column of data, marking empty values or values outside the clinically reasonable range. Optionally, for missing values of the same indicator for the same patient, the arithmetic mean of the measurements at the most recent time point is taken. Optionally, if other indicators at the same time point are complete, the missing values are calculated using a preset clinical relationship model, and the original values of the modified values outside the clinically reasonable range are marked for manual review.
[0087] Step 203, time alignment is performed on the complete longitudinal data set to obtain a time-aligned longitudinal data set.
[0088] The time-aligned longitudinal data set is a data table in which all patient data are mapped to a uniform time grid. The numerical alignment can be performed by dividing the time point sequence at fixed intervals, with the earliest and latest measurement times as boundaries.
[0089] Step 204, based on the time-aligned longitudinal data set, the standardized data is calculated by the following formula:
[0090]
[0091] wherein, is the standardized data, representing the standardized data of the physiological indicator j of the patient i at the time point t, is the original measurement value of the physiological indicator j of the patient i at the time point t, is the global mean value of the physiological indicator j, is the global standard deviation of the physiological indicator j, is the intermediate value after global standardization, is the individual historical mean value of the physiological indicator j of the patient i, is the individual historical standard deviation of the physiological indicator j of the patient i.
[0092] Specifically, the standardized data is a collection of final standardized values. Global standardization is performed: the mean value and the standard deviation of the indicator j of all patients at all time points are calculated, and then the global standardization is performed to eliminate the dimensional differences of the data of all patients, so that the longitudinal data are comparable across individuals; and individual standardization is performed: the mean value and the standard deviation of the indicator j of the patient i in the historical window are calculated, and then the standardized data representing the standard deviation multiple of the current value deviating from the individual historical baseline is obtained based on the intermediate value standardization.
[0093] The embodiment accurately quantifies patient-specific abnormalities through double standardization, provides good data for subsequent steps, and improves the accuracy of parameter estimation.
[0094] In one embodiment, based on the approximate posterior distribution, the high-dimensional parameters of the longitudinal data are estimated to obtain a high-dimensional parameter estimation table, and the method further comprises:
[0095] Step 301, obtaining the original data of the new patient, and standardizing the original data of the new patient to obtain new standardized data.
[0096] The new standardized data is a collection of standardized values of physiological indicators of the new patient, and is isomorphic with the standardized data. The terminal reads the original physiological indicator data of the new patient from the medical device and / or system interface, and performs the same steps as the aforementioned longitudinal data processing on the original physiological indicator data to obtain the new standardized data.
[0097] Step 302, sampling parameter posterior mean, parameter posterior covariance and covariance parameter from the approximate posterior distribution, and constructing as effective parameters.
[0098] Specifically, the effective parameters are a parameter set that can include the parameter posterior mean, the parameter posterior covariance and the covariance parameter. For example, the terminal randomly extracts 500 groups of parameter posterior mean, parameter posterior covariance and covariance parameter from the approximate posterior distribution, and integrates the sampled parameters by type.
[0099] Step 303, predicting the disease risk based on the new standardized data and the effective parameters, and obtaining a disease probability matrix.
[0100] The disease probability matrix can include a time point, a disease type and a probability value. Based on the new standardized data and the effective parameters collected from the similar posterior distribution, the value of Bayesian prediction is calculated to obtain the disease probability matrix.
[0101] Step 304, generating a risk assessment report based on the disease probability matrix and the high-dimensional parameter estimation table.
[0102] Specifically, the risk assessment report is a clinical decision support document for assisting disease diagnosis. The terminal extracts point estimates from the high-dimensional parameter estimation table to determine the influence direction of the index, extracts selection probability to identify key warning indicators, and then generates a risk assessment report.
[0103] The embodiment predicts the disease risk by associating the high-dimensional parameter estimation value with the new patient data, so that the parameter estimation has actual use value, and provides a fast and accurate method for diagnosing the disease risk of a new patient.
[0104] In one embodiment, based on the standardized data, the theoretical posterior distribution of the longitudinal data is obtained by Bayesian empirical likelihood function estimation, including:
[0105] Step 401, based on the standardized data, the generalized estimation equation is defined by the following formula:
[0106]
[0107] Wherein, U i (θ) is the estimation function of the i th individual, D i is the derivative matrix, V i is the covariance matrix, μ i (θ) is the marginal mean vector, Y i is the observation variable of the response variable of the i th individual, obtained from the standardized data.
[0108] Specifically, the generalized estimating equation is a mathematical definition of an individual estimating function, which is used to establish the relationship between the parameter theta and the standardized data. The standardized data is input into the equation to construct the equation.
[0109] Step 402, based on the generalized estimating equation and the work-related matrix, a decorrelation function is constructed.
[0110] Specifically, the decorrelation function is an orthogonalized estimating function that can eliminate the bias caused by improper setting of the work-related matrix. Based on the correlation structure parameters of the work-related matrix, a sensitivity matrix is calculated, and then a projection matrix is generated through a formula to output the decorrelation function. For example, if the preset work-related matrix is an independent structure, the correlation structure parameter is 0, but the real data has correlation.
[0111] Step 403, based on the decorrelation function, an empirical likelihood function is constructed.
[0112] Among them, the empirical likelihood function is a probability model based on data rather than a preset distribution. The terminal solves the empirical weight and then generates the likelihood function, which can avoid the error of normal assumption without assuming the form of data distribution.
[0113] Step 404, based on the preset medical prior knowledge base, the prior distribution is set to obtain the parameter prior.
[0114] Specifically, the parameter prior can include: variable selection prior based on the medical knowledge base; parameter conditional prior based on the conjugate prior; correlation structure prior based on the time series characteristics. For example: the conjugate prior acts on the parameter conditional prior, which can choose normal distribution, Gamma distribution and Beta distribution, which can make the posterior distribution of the same family as the prior distribution, and then simplify the posterior calculation; non-informative prior can act on parameter conditional prior or correlation prior, which can choose uniform distribution and Jeffreys prior; spike and slab prior can act on variable selection prior; horseshoe prior can act on parameter shrinkage prior; correlation structure prior can act on correlation prior.
[0115] Step 405, based on the empirical likelihood function, the decorrelation function and the parameter prior, the posterior distribution is constructed by the following formula to obtain the theoretical posterior distribution:
[0116] p(θ,γ,R|data)∝L e (θ)·p(θ|γ)p(γ)·p(R)
[0117] Among them, p(θ,γ,R|data) is the theoretical posterior distribution, L e (θ) is the empirical likelihood function, p(θ|γ)p(γ) is the parameter prior, and p(R) is the correlation prior.
[0118] Specifically, the theoretical posterior distribution is a parameter joint probability model. The terminal integrates the likelihood and the prior, wherein the empirical likelihood function carries data information, the parameter prior introduces medical field knowledge, and the relevant prior constraint time correlation structure.
[0119] The embodiment introduces a decorrelation function to eliminate the estimation deviation caused by the setting error of the working correlation matrix, avoids distribution hypothesis error and overfitting through empirical likelihood and prior constraint, improves the accuracy of parameter estimation, and further improves the accuracy of subsequent disease risk prediction.
[0120] In one embodiment, based on the decorrelation function, an empirical likelihood function is constructed, including:
[0121] Step 501: The following formula is solved by the Lagrange multiplier method to obtain the empirical weight:
[0122]
[0123] Wherein, λ(θ) satisfies the constraint formula:
[0124]
[0125] Wherein, p i (θ) is the empirical weight of the i-th observation, λ(θ) is the Lagrange multiplier vector, n is the sample number, U i * (θ) is the decorrelation function.
[0126] Specifically, the empirical weight is the probability weight of the i-th observation data point under the parameter θ, and the Lagrange multiplier vector is a vector parameter used for constraint optimization. The goal of the empirical weight is to maximize the empirical likelihood while satisfying the constraint condition, that is, the sum of the weighted decorrelation functions of the observation data is zero and the probability sum is 1. The terminal establishes a Lagrange function containing two constraints, takes the partial derivative of the probability and sets it to zero to obtain the empirical weight expression, uses the Newton iteration method to solve the nonlinear equation, and obtains the empirical weight.
[0127] Step 502: According to the empirical weight, the empirical likelihood function is obtained by the following formula:
[0128]
[0129] Wherein, L e (θ) is the empirical likelihood function, p i (θ) is the empirical weight of the i-th observation.
[0130] Specifically, the empirical likelihood function is based on a probability model of data, the empirical weight is regarded as a discrete probability mass, the empirical likelihood is defined as the product of the weights, the calculation is carried out using the logarithmic form, when the constraint is completely satisfied, the likelihood value is maximum, and the size of the likelihood value reflects the fitting degree of the parameter theta and the data.
[0131] The embodiment solves the distribution assumption problem of the traditional parameter likelihood in high-dimensional medical data by constructing a probability model without distribution assumption through empirical likelihood, maintains statistical efficiency, and improves the accuracy of parameter estimation.
[0132] In one of the embodiments, based on the theoretical posterior distribution, an approximate posterior distribution similar to the theoretical posterior distribution is constructed through a variational inference method, including:
[0133] Step 601, based on the theoretical posterior distribution, a variational distribution family is selected, and the variational parameters of the variational distribution family are initialized to obtain initial parameters.
[0134] Specifically, the initial parameters are initial guess values of the variational parameters, the variational distribution family is a set of parameterized probability distributions used to approximate the complex posterior distribution, which can adopt the product form of normal, Bernoulli and inverse Wishart, and the variational parameters are a set of parameters that control the shape of the variational distribution. According to the structural characteristics of the theoretical posterior distribution, the terminal selects a decomposable variational distribution family and performs parameter initialization. Exemplarily, the normal distribution is used to fit the symmetric distribution characteristics of the continuous parameter theta, the Bernoulli distribution is used to model the binary variable selection, and the inverse Wishart distribution is used to describe the correlation working matrix.
[0135] Step 602, based on the theoretical posterior distribution and the variational distribution family, the lower bound of the evidence is constructed through the following formula to obtain a computable ELBO expression:
[0136]
[0137] Wherein, L(λ) is the objective function of the lower bound of the evidence, λ is the Lagrange multiplier vector, K is the number of Monte Carlo sampling, n is the sample number, theta (k) is the parameter vector of the kth sampling, U i * (θ) is a decorrelation function, gamma is a prior distribution hyperparameter, R is a working correlation matrix, p(θ|gamma), p(gamma) is a parameter prior, q(θ, gamma, R) is a variational distribution.
[0138] Specifically, the evidence lower bound (ELBO) is an optimization objective function of the variational inference, maximizing the ELBO is equivalent to minimizing the KL divergence, and the ELBO expression can be calculated by a numerical solution form of Monte Carlo sampling. For example, the first term in the ELBO expression is the decomposition value of the expected likelihood, and the second and third terms together form the KL regularization term. By sampling K sets of parameters from the current variational distribution, approximating the expected term, and then decomposing the KL regularization term, the problem of the non-differentiable empirical likelihood function can be solved.
[0139] Step 603, based on the initial parameters and the computable ELBO expression, the variational parameters are optimized by the following formula to obtain the optimized parameters:
[0140]
[0141] where λ (t+1) is the optimized parameter, λ (t) is the parameter before optimization, t is the iteration number, p t is the adaptive step size, F -1 is the inverse of the information matrix, and the gradient of ELBO with respect to λ.
[0142] Specifically, the adaptive step size is a dynamically adjusted learning rate, and the optimized parameter is the variational parameter after iteration convergence. The gradient of ELBO with respect to the variational parameter is calculated, the value of the Fisher information matrix is estimated, and the preconditioned gradient is calculated. The Riemannian optimization strategy is used for optimization. When the relative change threshold is less than 10 -5 , it is judged to be converged, or when the maximum preset iteration number is reached, it is judged to be converged.
[0143] Step 604, based on the optimized parameters, the complete posterior distribution is constructed to obtain the approximate posterior distribution.
[0144] where the approximate posterior distribution is the optimized probability model. The terminal extracts parameters from the final iteration result, constructs a distribution using the extracted parameters, and then verifies the distribution by calculating the ELBO value and estimating the KL divergence to obtain the approximate posterior distribution.
[0145] The embodiment generates an approximate posterior distribution by introducing the method of variational inference, avoids the calculation of the theoretical posterior distribution which is difficult to calculate, and improves the real-time performance and processing speed of parameter estimation.
[0146] In one embodiment, based on the optimized parameters, the complete posterior distribution is constructed to obtain the approximate posterior distribution, including:
[0147] Step 701, obtaining the parameter posterior mean vector, the parameter posterior covariance matrix, the variable selection probability vector, the inverse Wishart degree of freedom and the inverse Wishart scale matrix from the optimized parameters.
[0148] Specifically, the parameter posterior mean vector is the posterior expectation vector of the continuous parameter θ, the parameter posterior covariance matrix is the posterior covariance matrix of θ, the variable selection probability vector refers to the marginal probability that the index j is selected into the model, and the inverse Wishart degree of freedom and the inverse Wishart scale matrix are the inverse Wishart distribution degree of freedom and the scale matrix of the inverse Wishart distribution of the working correlation matrix R. The variational parameters are directly read from the optimized parameters, and it is ensured that the parameter posterior covariance matrix and the inverse Wishart scale matrix are positive definite matrices, and the variable selection probability vector ranges between 0 and 1.
[0149] Step 702, based on the parameter posterior mean vector, the parameter posterior covariance matrix, the variable selection probability vector, the inverse Wishart degree of freedom and the inverse Wishart scale matrix, the approximate posterior distribution is constructed by the following formula:
[0150]
[0151] Wherein, q * (θ,γ,R) is the approximate posterior distribution, is the parameter posterior mean vector, is the parameter posterior covariance matrix, is the variable selection probability vector, v * is the inverse Wishart degree of freedom, Ψ * is the inverse Wishart scale matrix, γ is the prior distribution hyperparameter, and R is the working correlation matrix.
[0152] Specifically, the multivariate normal distribution, the Bernoulli component and the inverse Wishart component are constructed, and the joint probability density is integrated as q * (θ,γ,R)=p(θ)×∏ j p(γ j )×p(r).
[0153] The embodiment constructs the approximate posterior distribution by the normal distribution, the Bernoulli distribution and the inverse Wishart distribution, and improves the accuracy and real-time performance of parameter estimation.
[0154] In one of the embodiments, based on the approximate posterior distribution, the high-dimensional parameters of the longitudinal data are estimated to obtain a high-dimensional parameter estimation table, including:
[0155] Step 801, extracting key statistics in the approximate posterior distribution; the key statistics include: posterior mean, marginal variance and selection probability.
[0156] In particular, the posterior mean is the expected value of the parameter θ j , the marginal variance is the variance of the parameter θ j , and the selection probability is the variable γ j =1. Each element of the parameter posterior mean vector is directly read as the posterior mean, the diagonal element of the covariance matrix is extracted as the marginal variance, and the element value of the variable selection probability vector is directly read as the selection probability.
[0157] Step 802, based on the key statistics, the point estimate vector and the standard error vector are calculated by the following formula:
[0158]
[0159] wherein, is one of the elements of the point estimate vector, is the posterior mean, SE j is one of the elements of the standard error vector, is the marginal variance.
[0160] In particular, the point estimate vector is the best single-value estimate of the parameter, and the standard error vector is the standard error of the point estimate. The posterior mean is directly used as the point estimate to generate the point estimate vector, and the standard deviation is calculated to generate the standard error vector.
[0161] Step 803, based on the key statistics, the confidence interval is constructed using the normal approximation method.
[0162] wherein, the confidence interval is the interval estimate of the parameter θ j , and the posterior distribution of θ j is assumed to be approximately normal, and the 95% confidence interval is calculated. For example, when the confidence level is adjusted to 99%: When the confidence level is adjusted to 90%:
[0163] Step 804, the key statistics, the point estimate vector, the standard error vector, and the confidence interval are integrated to obtain the high-dimensional parameter estimation table.
[0164] The posterior mean, the marginal variance, the point estimate, the standard error, the confidence interval, and the selection probability are integrated into a structured table, and the significance is marked based on the value of the selection probability to obtain the high-dimensional parameter estimation table.
[0165] This embodiment provides support for subsequent high-dimensional parameter auxiliary diagnosis and rapid decision-making through the trinity of point estimate, interval, and probability.
[0166] It should be understood that although the steps in the flowcharts involved in the embodiments described above are shown in sequence according to the arrows, the steps are not necessarily executed in the order indicated by the arrows. Unless otherwise specified herein, the execution of the steps is not strictly limited in sequence, and the steps can be executed in other orders. Moreover, at least some of the steps in the flowcharts involved in the embodiments described above can include multiple steps or multiple stages, which are not necessarily executed at the same time, but can be executed at different times, and the execution order of the steps or stages is not necessarily sequential, but can be alternately or alternately executed with at least part of other steps or steps or stages in other steps.
[0167] Based on the same inventive concept, the embodiments of the present application also provide a Bayesian empirical likelihood and variational inference based longitudinal data high-dimensional parameter estimation device for implementing the above-mentioned Bayesian empirical likelihood and variational inference based longitudinal data high-dimensional parameter estimation method. The implementation scheme for solving the problem provided by the device is similar to the implementation scheme described in the above method, so the specific limitations in one or more Bayesian empirical likelihood and variational inference based longitudinal data high-dimensional parameter estimation device embodiments provided below can be referred to the limitations of the Bayesian empirical likelihood and variational inference based longitudinal data high-dimensional parameter estimation method described above, which will not be repeated here.
[0168] In one exemplary embodiment, as shown in Figure 2 a Bayesian empirical likelihood and variational inference based longitudinal data high-dimensional parameter estimation device 900 is provided, comprising:
[0169] a standardization module 901 configured to obtain longitudinal data and perform standardization processing on the longitudinal data to obtain standardized data; the longitudinal data is data for recording physiological states of each individual at different time points in history obtained from a medical database;
[0170] a Bayesian module 902 configured to obtain a theoretical posterior distribution of the longitudinal data by Bayesian empirical likelihood estimation based on the standardized data;
[0171] a variational inference module 903 configured to construct an approximate posterior distribution similar to the theoretical posterior distribution by a variational inference method based on the theoretical posterior distribution;
[0172] an estimation module 904 configured to estimate high-dimensional parameters of the longitudinal data based on the approximate posterior distribution to obtain a high-dimensional parameter estimation table corresponding to the longitudinal data; the high-dimensional parameter estimation table is used to assist in diagnosing the disease risk of the patient corresponding to the longitudinal data.
[0173] Further, the standardization module 901 is further configured to:
[0174] Extract longitudinal data from the medical database and sort by timestamp to obtain an original longitudinal data set;
[0175] Identify missing values in the original longitudinal data set and fill in missing values to obtain a complete longitudinal data set;
[0176] Time align the complete longitudinal data set to obtain a time-aligned longitudinal data set;
[0177] Based on the time-aligned longitudinal data set, the standardized data is calculated by the following formula:
[0178]
[0179] wherein, is the standardized data, representing the standardized data of physiological indicator j of patient i at time point t, is the original measurement value of physiological indicator j of patient i at time point t, is the global mean of physiological indicator j, is the global standard deviation of physiological indicator j, is the intermediate value after global standardization, is the individual historical mean of physiological indicator j of patient i, is the individual historical standard deviation of physiological indicator j of patient i.
[0180] Further, the device further comprises a diagnosis module for:
[0181] Obtaining original data of a new patient and standardizing the original data of the new patient to obtain new standardized data;
[0182] Sampling the parameter posterior mean, the parameter posterior covariance and the covariance parameter from the approximate posterior distribution, and constructing them as effective parameters;
[0183] Based on the new standardized data and the effective parameters, predicting the disease risk to obtain a disease probability matrix;
[0184] Based on the disease probability matrix and the high-dimensional parameter table estimation table, generating a risk assessment report.
[0185] Further, the Bayesian module 902 is also used for:
[0186] Based on the standardized data, the generalized estimating equation is defined by the following formula:
[0187]
[0188] wherein, U i (θ) is the estimation function of the i-th individual, D i is the derivative matrix, V iCovariance matrix, μ i (θ) is a marginal mean vector, Y i is an observed variable of the response variable of the i th individual, obtained from the standardized data;
[0189] Based on the generalized estimating equation and the working correlation matrix, a decorrelation function is constructed;
[0190] Based on the decorrelation function, an empirical likelihood function is constructed;
[0191] Based on the preset medical prior knowledge base, a prior distribution is set to obtain a parameter prior;
[0192] Based on the empirical likelihood function, the decorrelation function and the parameter prior, a posterior distribution is constructed by the following formula to obtain a theoretical posterior distribution:
[0193] p(θ,γ,R|data)∝L e (θ)·p(θ|γ)p(γ)·p(R)
[0194] Wherein, p(θ,γ,R|data) is a theoretical posterior distribution, L e (θ) is an empirical likelihood function, p(θ|γ)p(γ) is a parameter prior, and p(R) is a correlation prior.
[0195] Further, the Bayesian module 902 is also used for:
[0196] The following formula is solved by the Lagrange multiplier method to obtain an empirical weight:
[0197]
[0198] Wherein, λ(θ) satisfies the constraint formula:
[0199]
[0200] Wherein, p i (θ) is the empirical weight of the i th observation, λ(θ) is the Lagrange multiplier vector, n is the sample number, and U i * (θ) is a decorrelation function;
[0201] According to the empirical weight, the empirical likelihood function is obtained by the following formula:
[0202]
[0203] Wherein, L e (θ) is an empirical likelihood function, and p i (θ) is the empirical weight of the i th observation.
[0204] Further, the variational inference module 903 is also used for:
[0205] based on the theoretical posterior distribution, selecting a variational distribution family, and initializing variational parameters of the variational distribution family to obtain initial parameters;
[0206] based on the theoretical posterior distribution and the variational distribution family, constructing an evidence lower bound through the following formula to obtain a computable ELBO expression:
[0207]
[0208] wherein L(λ) is a target function of the evidence lower bound, λ is a Lagrange multiplier vector, K is a number of Monte Carlo sampling, n is a sample number, θ (k) is a parameter vector of the kth sampling, U i * (θ) is a decorrelation function, γ is a prior distribution hyperparameter, R is a working correlation matrix, p(θ|γ), p(γ) is a parameter prior, q(θ,γ,R) is a variational distribution;
[0209] based on the initial parameters and the computable ELBO expression, optimizing the variational parameters through the following formula to obtain optimized parameters:
[0210]
[0211] wherein λ (t+1) is the optimized parameters, λ (t) is the parameters before optimization, t is an iteration number, p t is an adaptive step size, F -1 is an inverse of an information matrix, is a gradient of ELBO with respect to λ;
[0212] based on the optimized parameters, constructing a complete posterior distribution to obtain an approximate posterior distribution.
[0213] Further, the variational inference module 903 is further configured to:
[0214] obtaining a parameter posterior mean vector, a parameter posterior covariance matrix, a variable selection probability vector, an inverse Wishart degree of freedom, and an inverse Wishart scale matrix from the optimized parameters;
[0215] based on the parameter posterior mean vector, the parameter posterior covariance matrix, the variable selection probability vector, the inverse Wishart degree of freedom, and the inverse Wishart scale matrix, constructing the approximate posterior distribution through the following formula:
[0216]
[0217] wherein q * (θ,γ,R) is the approximate posterior distribution, is a parameter posterior mean vector, is a parameter posterior covariance matrix, is a variable selection probability vector, v * is an inverse Wishart degree of freedom, Ψ * is an inverse Wishart scale matrix, γ is a prior distribution hyperparameter, and R is a working correlation matrix.
[0218] Further, the estimation module 904 is further configured to:
[0219] extract key statistics in the approximate posterior distribution; the key statistics include: a posterior mean, a marginal variance, and a selection probability;
[0220] based on the key statistics, calculate a point estimate vector and a standard error vector by the following formula:
[0221]
[0222] wherein, is one of the elements of the point estimate vector, is the posterior mean, SE j is one of the elements of the standard error vector, is the marginal variance;
[0223] based on the key statistics, construct a confidence interval using a normal approximation method;
[0224] integrate the key statistics, the point estimate vector, the standard error vector, and the confidence interval to obtain a high-dimensional parameter estimation table.
[0225] In one embodiment, a computer device is provided, comprising a memory and a processor, the memory stores a computer program, and the processor implements the steps of the longitudinal data high-dimensional parameter estimation method based on Bayesian empirical likelihood and variational inference as described above when executing the computer program.
[0226] In one embodiment, a computer readable storage medium is provided, which stores a computer program, and the computer program is executed by a processor to implement the steps in each of the method embodiments described above.
[0227] For the device embodiment, since it basically corresponds to the method embodiment, the relevant part can be seen from the part of the method embodiment. The device embodiment described above is only schematic, wherein the components shown as separate components can or can not be physically separate, and the components shown as a unit can or can not be a physical unit, i.e., can be located in one place, or can be distributed to multiple network units. Part or all of the modules can be selected according to actual needs to achieve the purpose of the present disclosure. Those skilled in the art can understand and implement it without creative labor.
[0228] The above-described embodiments only express several implementation manners of the present application, which are described in detail, but should not be understood as a limitation on the patent scope of the application. It should be pointed out that, for those skilled in the art, without departing from the concept of the present application, several modifications and improvements can be made, which are all within the protection scope of the present application.
Claims
1. A high-dimensional parameter estimation method for longitudinal data based on Bayesian empirical likelihood and variational inference, characterized by: The method comprises: Acquiring longitudinal data and performing standardization processing on the longitudinal data to obtain standardized data; the longitudinal data is data obtained from a medical database for recording the physiological status of each individual at different time points in history; Based on the standardized data, the theoretical posterior distribution of the longitudinal data is obtained by Bayesian empirical likelihood function estimation; Based on the theoretical posterior distribution, constructing an approximate posterior distribution similar to the theoretical posterior distribution through a variational inference method; Based on the approximate posterior distribution, the high-dimensional parameters of the longitudinal data are estimated to obtain a high-dimensional parameter estimation table corresponding to the longitudinal data; the high-dimensional parameter estimation table is used to assist in diagnosing the patient's disease risk corresponding to the longitudinal data.
2. The method for estimating high-dimensional parameters of longitudinal data based on Bayesian empirical likelihood and variational inference according to claim 1, characterized in that: The acquiring of longitudinal data and performing standardization processing on the longitudinal data to obtain standardized data includes: Extracting longitudinal data from the medical database and sorting them by timestamp to obtain an original longitudinal dataset; Identifying missing values in the original longitudinal dataset and performing missing value filling to obtain a complete longitudinal dataset; performing time alignment on the complete longitudinal dataset to obtain a time-aligned longitudinal dataset; Based on the time-aligned longitudinal dataset, the standardized data is calculated using the following formula: in, is the standardized data, representing the standardized data of physiological index j of patient i at time point t, is the original measurement value of the physiological index j of patient i at time point t, is the global mean of physiological index j, is the global standard deviation of physiological index j, is the median value after global normalization, is the individual historical mean of physiological index j of patient i, is the individual historical standard deviation of physiological indicator j of patient i.
3. The method for high-dimensional parameter estimation of longitudinal data based on Bayesian empirical likelihood and variational inference according to claim 1, characterized in that: After estimating the high-dimensional parameters of the longitudinal data based on the approximate posterior distribution and obtaining a high-dimensional parameter estimation table, the method further includes: Acquiring new patient original data, and standardizing the new patient original data to obtain new standardized data; Sampling parameter posterior means, parameter posterior covariances, and covariance parameters from the approximate posterior distribution and constructing them as effective parameters; Predicting the disease risk based on the new standardized data and the effective parameters to obtain a disease probability matrix; A risk assessment report is generated based on the disease probability matrix and the high-dimensional parameter table estimation table.
4. The method for estimating high-dimensional parameters of longitudinal data based on Bayesian empirical likelihood and variational inference according to claim 1, characterized in that: The method of obtaining the theoretical posterior distribution of the longitudinal data based on the standardized data by estimating the Bayesian empirical likelihood function includes: Based on the standardized data, the generalized estimating equation is defined by the following formula: Among them, U i (θ) is the estimation function of the i-th individual, D i is the derivative matrix, V i is the covariance matrix, μ i (θ) is the marginal mean vector, Y i is the observed variable of the response variable of the i-th individual, obtained from the standardized data; constructing a decorrelation function based on the generalized estimating equation and the working correlation matrix; constructing an empirical likelihood function based on the decorrelation function; A prior distribution is set based on a preset medical prior knowledge base to obtain parameter priors; Based on the empirical likelihood function, the decorrelation function and the parameter prior, the posterior distribution is constructed by the following formula to obtain the theoretical posterior distribution: p(θ,γ,R|data)∝L e (θ)·p(θ|γ)p(γ)·p(R) Among them, p(θ,γ,R|data) is the theoretical posterior distribution, L e (θ) is the empirical likelihood function, p(θ|γ)p(γ) is the parameter prior, and p(R) is the correlation prior.
5. The method for estimating high-dimensional parameters of longitudinal data based on Bayesian empirical likelihood and variational inference according to claim 4, characterized in that: The constructing of an empirical likelihood function based on the decorrelation function includes: The Lagrange multiplier method is used to solve the following formula to obtain the empirical weight: Among them, λ(θ) satisfies the constraint formula: Among them, p i (θ) is the empirical weight of the i-th observation, λ(θ) is the Lagrange multiplier vector, n is the number of samples, U i * (θ) is the decorrelation function; According to the empirical weights, the empirical likelihood function is obtained by the following formula: Among them, L e (θ) is the empirical likelihood function, p i (θ) is the empirical weight of the i-th observation.
6. The method for estimating high-dimensional parameters of longitudinal data based on Bayesian empirical likelihood and variational inference according to claim 1, characterized in that: The method of constructing an approximate posterior distribution similar to the theoretical posterior distribution by a variational inference method based on the theoretical posterior distribution includes: Based on the theoretical posterior distribution, a variational distribution family is selected, and variational parameters of the variational distribution family are initialized to obtain initial parameters; Based on the theoretical posterior distribution and the variational distribution family, the lower bound of evidence is constructed by the following formula to obtain a computable ELBO expression: Where L(λ) is the objective function of the evidence lower bound, λ is the Lagrange multiplier vector, K is the Monte Carlo sampling number, n is the number of samples, θ (k) is the parameter vector of the kth sampling, U i * (θ) is the decorrelation function, γ is the hyperparameter of the prior distribution, R is the working correlation matrix, p(θ|γ) and p(γ) are parameter priors, and q(θ,γ,R) is the variational distribution; Based on the initial parameters and the computable ELBO expression, the variational parameters are optimized using the following formula to obtain the optimized parameters: Among them, λ (t+1) is the optimization parameter, λ (t) is the parameter before optimization, t is the number of iterations, p t is the adaptive step size, F -1 is the inverse of the information matrix, is the gradient of ELBO with respect to λ; Based on the optimized parameters, a complete posterior distribution is constructed to obtain the approximate posterior distribution.
7. The method for estimating high-dimensional parameters of longitudinal data based on Bayesian empirical likelihood and variational inference according to claim 6, characterized in that: The constructing a complete posterior distribution based on the optimized parameters to obtain the approximate posterior distribution includes: Obtain parameter posterior mean vector, parameter posterior covariance matrix, variable selection probability vector, inverse Wishart degrees of freedom and inverse Wishart scaling matrix from the optimized parameters; Based on the parameter posterior mean vector, the parameter posterior covariance matrix, the variable selection probability vector, the inverse Wishart degrees of freedom and the inverse Wishart scaling matrix, the approximate posterior distribution is constructed by the following formula: Among them, q * (θ,γ,R) is the approximate posterior distribution, is the parameter posterior mean vector, is the parameter posterior covariance matrix, Select the probability vector for the variable, v * is the inverse Wishart degree of freedom, Ψ * is the inverse Wishart scaling matrix, γ is the prior distribution hyperparameter, and R is the working correlation matrix.
8. The method for estimating high-dimensional parameters of longitudinal data based on Bayesian empirical likelihood and variational inference according to any one of claims 1 to 7, characterized in that: The method of estimating high-dimensional parameters of the longitudinal data based on the approximate posterior distribution to obtain a high-dimensional parameter estimation table includes: Extracting key statistics from the approximate posterior distribution; the key statistics include: posterior mean, marginal variance and selection probability; Based on the key statistics, the point estimate vector and standard error vector are calculated using the following formula: in, is one of the elements of the point estimate vector, is the posterior mean, SE j is one of the elements of the standard error vector, is the marginal variance; Based on the key statistics, credible intervals were constructed using the normal approximation method; The key statistics, the point estimate vector, the standard error vector, and the credible interval are integrated to obtain the high-dimensional parameter estimation table.
9. A device for estimating high-dimensional parameters of longitudinal data based on Bayesian empirical likelihood and variational inference, characterized in that: The device comprises: a standardization module, configured to obtain longitudinal data and perform standardization processing on the longitudinal data to obtain standardized data; the longitudinal data is data obtained from a medical database to record the physiological status of each individual at different time points in history; A Bayesian module, configured to obtain a theoretical posterior distribution of the longitudinal data based on the standardized data by estimating a Bayesian empirical likelihood function; A variational inference module, configured to construct an approximate posterior distribution similar to the theoretical posterior distribution through a variational inference method based on the theoretical posterior distribution; An estimation module is used to estimate the high-dimensional parameters of the longitudinal data based on the approximate posterior distribution, and obtain a high-dimensional parameter estimation table corresponding to the longitudinal data; the high-dimensional parameter estimation table is used to assist in diagnosing the patient's disease risk corresponding to the longitudinal data.
10. A computer device comprising a memory and a processor, wherein the memory stores a computer program, wherein: When the processor executes the computer program, the steps of the method according to any one of claims 1 to 8 are implemented.
Citation Information
Cited By
Bayesian disease risk high-precision mapping method and system
CN122158146A