A biomarker aggregation method based on expectation-maximization iteration
By aggregating biomarkers from multiple data centers based on the desired maximization iteration, the problem of batch effects affecting parameter estimation accuracy is solved, and efficient biomarker polymerization and accurate parameter estimation are achieved.
Patent Information
- Application Number
- CN202510054052.5
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-01-14
- Publication Date
- 2025-05-16
- Estimated Expiration
- 2045-01-14
AI Technical Summary
How to effectively aggregate biomarkers from multiple data centers to improve the accuracy of subsequent biomarker parameter estimation, which is affected by batch effects.
Using a method based on the expected maximization iteration, a sample likelihood function is constructed and an expected maximization iteration algorithm is performed to obtain target estimation parameters and characterize the effect of the biomarkers on disease outcomes by obtaining local measurements of biomarkers from multiple data sources and the reference measurements performed for sampling assays.
Effectively reduces the batch effect, improves the accuracy of estimating biomarker parameters, and realizes effective polymerization of biomarkers from multiple data centers.
Smart Images

Figure CN119474650B_ABST
Abstract
Description
Technical Field
[0001] The present application relates to the technical field of medical data processing, and in particular to a biomarker aggregation method based on expectation maximization iteration. Background Art
[0002] Biomarkers are biological molecules found in blood, other body fluids or tissues, and they play a vital role in medical research and clinical practice. They can serve as indicators that can objectively measure and evaluate normal biological processes, pathological processes or responses to drug interventions.
[0003] In medical research and clinical trials, pooling biomarker data from multiple sources to explore the relationship between biomarkers and clinical outcomes is a commonly used data analysis method that significantly improves the ability to detect meaningful associations between biomarker data and clinical outcomes by increasing sample size and improving statistical power.
[0004] However, biomarkers from different data sources have batch effects due to different corresponding research conditions, which affects the accuracy of the parameter estimation corresponding to the final biomarker and disease outcome. Therefore, how to effectively aggregate biomarkers from multiple data centers to improve the accuracy of subsequent biomarker parameter estimation has become an urgent problem to be solved. Summary of the invention
[0005] The main purpose of this application is to provide a biomarker aggregation method based on expectation maximization iteration to solve the technical problem of how to effectively aggregate biomarkers from multiple data centers to improve the accuracy of subsequent biomarker parameter estimation.
[0006] To achieve the above objectives, the present application proposes a biomarker aggregation method based on expectation maximization iteration, which comprises:
[0007] Obtaining local measurements corresponding to biomarkers from multiple data sources;
[0008] Sampling and measuring the biomarkers to obtain reference measurement values;
[0009] An expectation-maximization iterative algorithm is executed according to the reference measurement value and the local measurement value to obtain a target estimation parameter; the target estimation parameter characterizes the effect of the biomarker on the disease outcome.
[0010] In one embodiment, the step of performing an expectation-maximization iterative algorithm according to the reference measurement value and the local measurement value to obtain a target estimation parameter comprises:
[0011] Obtain the study covariates corresponding to the biomarkers;
[0012] constructing a sample likelihood function corresponding to the biomarker according to the local measurement value, the reference measurement value and the research covariate;
[0013] Based on the sample likelihood function, the biological effect of the biomarker is iterated by expectation maximization to obtain target estimation parameters.
[0014] In one embodiment, the step of constructing a sample likelihood function corresponding to the biomarker according to the local measurement value, the reference measurement value and the research covariate comprises:
[0015] constructing an initial likelihood function corresponding to the biomarker based on the local measurement value, the study covariate and the reference measurement value;
[0016] A sample likelihood function corresponding to the biomarker is constructed according to a preset independence hypothesis and the initial likelihood function.
[0017] In one embodiment, the step of constructing the sample likelihood function corresponding to the biomarker according to the preset independence hypothesis and the initial likelihood function includes:
[0018] Get the target regression model corresponding to the current parameter estimate;
[0019] A sample likelihood function corresponding to the biomarker is constructed according to the preset independence hypothesis, the initial likelihood function and the target regression model.
[0020] In one embodiment, the step of performing expectation maximization iteration on the biological effect of the biomarker based on the sample likelihood function to obtain a target estimation parameter comprises:
[0021] Initialize parameters based on the reference measurement values to obtain initial estimated parameters;
[0022] Performing expectation maximization iterative update on the initial estimation parameters according to the current estimated outcome type and the sample likelihood function;
[0023] When it is detected that the iteration converges or reaches the preset maximum number of iterations, the current estimated parameters are used as the target estimated parameters.
[0024] In one embodiment, the step of iteratively updating the initial estimated parameters by expectation maximization according to the current estimated outcome type and the sample likelihood function comprises:
[0025] If the current estimated outcome type is a continuous outcome, determining a corresponding estimated statistic according to the conditional distribution of the initial estimated parameters;
[0026] The initial estimated parameters are iteratively updated by expectation maximization through an iterative reweighted least squares algorithm, a first solution, the sample likelihood function and the estimated statistic.
[0027] In one embodiment, the step of iteratively updating the initial estimated parameters by expectation maximization according to the current estimated outcome type and the sample likelihood function further includes:
[0028] If the current estimated outcome type is a binary outcome, a preset Laplace operation is performed on the initial estimated parameters to obtain a target normal posterior distribution;
[0029] Obtaining the importance weight corresponding to the target normal posterior distribution;
[0030] Obtaining a joint posterior expectation based on the importance weights;
[0031] The initial estimated parameters are iteratively updated by expectation maximization through the Newton-Raphson algorithm, the second solution, the sample likelihood function and the joint posterior expectation.
[0032] In addition, to achieve the above-mentioned purpose, the present application also proposes a biomarker aggregation device based on expectation maximization iteration, and the biomarker aggregation device based on expectation maximization iteration includes:
[0033] A data acquisition module, used to acquire local measurement values corresponding to biomarkers from multiple data sources;
[0034] A reference measurement module, used to perform sampling measurement on the biomarker to obtain a reference measurement value;
[0035] A data estimation module is used to perform an expectation-maximization iterative algorithm based on the reference measurement value and the local measurement value to obtain a target estimation parameter; the target estimation parameter characterizes the effect of the biomarker on the disease outcome.
[0036] In addition, to achieve the above-mentioned purpose, the present application also proposes a biomarker aggregation device based on expectation-maximization iteration, the device comprising: a memory, a processor, and a biomarker aggregation program based on expectation-maximization iteration stored in the memory and executable on the processor, the biomarker aggregation program based on expectation-maximization iteration being configured to implement the steps of the biomarker aggregation method based on expectation-maximization iteration as described above.
[0037] In addition, to achieve the above-mentioned purpose, the present application also proposes a storage medium, which is a storage medium on which a biomarker aggregation program based on expectation maximization iteration is stored. When the biomarker aggregation program based on expectation maximization iteration is executed by a processor, the steps of the biomarker aggregation method based on expectation maximization iteration as described above are implemented.
[0038] The present application provides a biomarker aggregation method based on expectation maximization iteration, which includes obtaining local measurement values corresponding to biomarkers from multiple data sources; sampling and measuring the biomarkers to obtain reference measurement values; executing the expectation maximization iteration algorithm based on the reference measurement values and the local measurement values to obtain target estimation parameters; the target estimation parameters characterize the effect influence corresponding to the biomarker. The present application first measures some biomarkers, obtains reference measurement values that are not affected by batch effects, and then iterates the reference measurement values and local measurement values through the expectation maximization algorithm, and finally obtains effective target estimation parameters between biomarkers and disease outcomes, thereby achieving effective aggregation of biomarkers from multiple data sources. BRIEF DESCRIPTION OF THE DRAWINGS
[0039] The accompanying drawings, which are incorporated in and constitute a part of this specification, illustrate embodiments consistent with the present application and, together with the description, serve to explain the principles of the present application.
[0040] In order to more clearly illustrate the embodiments of the present application or the technical solutions in the prior art, the drawings required for use in the embodiments or the description of the prior art will be briefly introduced below. Obviously, for ordinary technicians in this field, other drawings can be obtained based on these drawings without paying any creative labor.
[0041] Figure 1 This is a first flow chart of the first embodiment of the biomarker aggregation method based on expectation maximization iteration of the present application;
[0042] Figure 2 This is a second flow chart of the first embodiment of the biomarker aggregation method based on expectation maximization iteration of the present application;
[0043] Figure 3 This is a first flow chart of the second embodiment of the biomarker aggregation method based on expectation maximization iteration of the present application;
[0044] Figure 4 This is a second flow chart of the second embodiment of the biomarker aggregation method based on expectation maximization iteration of the present application;
[0045] Figure 5 This is a third flow chart of the second embodiment of the biomarker aggregation method based on expectation maximization iteration of the present application;
[0046] Figure 6 This is a schematic diagram of the module structure of a biomarker aggregation device based on expectation maximization iteration according to an embodiment of the present application;
[0047] Figure 7Schematic diagram of the device structure of the hardware operating environment involved in the biomarker aggregation method based on expectation-maximization iteration in the embodiment of the present application.
[0048] The purpose, features and advantages of this application will be further described in conjunction with the embodiments and with reference to the accompanying drawings. DETAILED DESCRIPTION
[0049] It should be understood that the specific embodiments described herein are only used to explain the technical solutions of the present application and are not used to limit the present application.
[0050] In order to better understand the technical solution of the present application, a detailed description will be given below in conjunction with the accompanying drawings and specific implementation methods.
[0051] The main solution of this application is: to obtain local measurement values corresponding to biomarkers from multiple data sources; to sample and measure biomarkers to obtain reference measurement values; to perform an expectation-maximization iterative algorithm based on the reference measurement values and local measurement values to obtain target estimation parameters; the target estimation parameters characterize the effect influence corresponding to the biomarker.
[0052] Currently, pooling biomarker data from multiple research centers can effectively improve the statistical power and precision of quantifying biomarker-disease associations. However, due to the variability of biomarkers between different research centers, this batch effect will affect the accuracy of parameter estimates between subsequent biomarkers and disease outcomes.
[0053] To address this problem, the present application can re-measure a sample subset of biomarkers from multiple data sources in a local reference laboratory to generate reference measurements that are assumed to be unaffected by batch effects, and effectively combine these reference measurements with the richer but potentially biased local measurements of each data source, thereby improving the efficiency of subsequent biomarker parameter estimation and achieving effective aggregation of biomarkers from multiple data sources.
[0054] Based on this, this application further developed an expectation maximization-based biomarker pooling (EMBP) method to collect biomarker data from multiple research sources. This method innovatively incorporates unmeasured reference measurements as latent variables into a statistical model and uses the expectation maximization (EM) algorithm to effectively estimate parameters, addressing the biomarker data integration problem from a new perspective.
[0055] Specifically, the present application first obtains reference measurement values corresponding to some biomarkers that are assumed to be unaffected by batch effects by performing local sampling measurements on biomarkers from multiple data centers; then, based on the reference measurement values, the unmeasured biomarkers are incorporated into a statistical model as latent variables, and an expectation-maximization iterative algorithm is executed using the reference measurement values and local measurement values to iterate the biological effects of the biomarkers, ultimately obtaining effective target estimation parameters between biomarkers from different data sources and disease outcomes, thereby achieving effective aggregation of biomarkers from multiple data centers.
[0056] It should be noted that the execution subject of this embodiment can be a biomarker aggregation system based on expectation maximization iteration, or a computing service device with data processing, network communication and program running functions, such as a tablet computer, a personal computer, a mobile phone, etc., or a biomarker aggregation device based on expectation maximization iteration that can achieve the above functions, etc. This embodiment does not specifically limit this. The following takes the biomarker aggregation based on expectation maximization iteration (referred to as the aggregation device) as the execution subject as an example to illustrate this embodiment and the following embodiments.
[0057] Based on this, the present application embodiment provides a multi-center data aggregation method, referring to Figure 1 , Figure 1 This is a first flow chart of the first embodiment of the multi-center data aggregation method of the present application.
[0058] In this embodiment, the biomarker aggregation method based on expectation-maximization iteration includes steps S10 to S30:
[0059] Step S10, obtaining local measurement values corresponding to biomarkers from multiple data sources;
[0060] It can be understood that the above-mentioned biomarkers from multiple data sources indicate that the biomarkers may be research indicators from different clinical research centers or different laboratories, and the above-mentioned local measurement values may be measurement values obtained after specific determination of the biomarkers in the research center or laboratory corresponding to the data source.
[0061] It should be understood that the purpose of aggregating the above-mentioned biomarkers from different data centers is to estimate the degree of influence between biomarkers and disease outcomes. Therefore, if data sets from multiple data centers are reasonably used in the process of impact analysis, it can be regarded as completing the effective aggregation of biomarkers from multiple data centers.
[0062] Step S20, sampling and measuring the biomarkers to obtain reference measurement values;
[0063] It is easy to understand that there are batch effects in biomarkers from different data centers. Therefore, in order to ensure the accuracy of parameter estimation of biomarkers after aggregation, this embodiment can reduce this difference by re-measuring and analyzing the biomarkers.
[0064] However, it is costly to reanalyze all biological samples from multiple research centers in a single reference laboratory, and not all biological samples can be remotely operated. Therefore, limited by the impact of cost and distance, this embodiment can extract part of the data from the biomarkers as a sample subset for re-measurement, that is, sample and measure the biomarkers, obtain the reference measurement values corresponding to the sample subset biomarkers, and assume that the biomarkers that have not been re-measured and analyzed (which can be referred to as the remaining biomarkers later) and the reference measurement values corresponding to the current reference laboratory (that is, the data center or laboratory where the current aggregation device is located) are regarded as estimable hidden variables, and based on this assumption, a calibration model for a specific study is established for each data center.
[0065] It is understandable that this embodiment assumes that represents the research indicators from S different laboratories, namely the above biomarkers, then for the research Individuals in , the disease outcome can be expressed as , the biomarker measurement from the local reference laboratory (i.e., the reference measurement mentioned above) can be expressed as , the biomarker measurements from different study-specific data centers (i.e., the local measurements mentioned above) can be expressed as . Combining the above analysis, we can see that for the research All biomarkers in the However, only some biomarkers have corresponding reference measurement values after the above sampling determination. In addition, it is assumed in this embodiment that for the study Individuals in , Representatives can obtain reference measurements, otherwise . And the reference measurement value The distribution of is consistent across studies, whereas local measurements The distribution of was heterogeneous in different studies.
[0066] Step S30, executing an expectation-maximization iterative algorithm according to the reference measurement value and the local measurement value to obtain a target estimation parameter; the target estimation parameter characterizes the effect of the biomarker on the disease outcome.
[0067] It is understandable that this embodiment can calibrate and summarize the local measurement values corresponding to the remaining biomarkers based on the reference measurement values of the biomarkers that have been re-measured and analyzed in the current reference laboratory to estimate the reference measurement values corresponding to the remaining biomarkers. Therefore, in this embodiment, in order to effectively aggregate and define biomarker-disease associations from multiple clinical studies, an expectation maximization-based biomarker pool (EMBP) method is proposed to collect biomarker data from multiple research sources.
[0068] In one possible implementation, refer to Figure 2 , Figure 2 This is a second flow chart of the first embodiment of the multi-center data aggregation method of the present application. In this embodiment, step S30 may include steps A1 to A3:
[0069] Step A1, obtaining the research covariates corresponding to the biomarkers;
[0070] It should be understood that the above-mentioned research covariates may be other factors that affect the disease outcome of the individual i in the study, such as the individual's age, gender, race, etc. A vector representing the study covariates.
[0071] Step A2, constructing a sample likelihood function corresponding to the biomarker according to the local measurement value, the reference measurement value and the research covariate;
[0072] It is easy to understand that this embodiment can innovatively incorporate the reference measurement values of the remaining biomarkers as latent variables into a statistical model, namely, the above-mentioned sample likelihood function, based on the reference measurement values obtained through sampling and the research covariates, and use the expectation maximization (EM) algorithm to effectively estimate the biological effects between biomarkers and disease outcomes on the basis of the sample likelihood function, and obtain the above-mentioned target estimation parameters, thereby achieving effective aggregation of biomarkers from multiple data centers.
[0073] In a feasible implementation manner, in this embodiment, step A2 includes steps A21-A22:
[0074] Step A21, constructing an initial likelihood function corresponding to the biomarker based on the local measurement value, the research covariate and the reference measurement value;
[0075] It can be understood that in order to perform calibration analysis on the local measurement value based on the reference measurement value, this embodiment can preliminarily establish a two-level research-biological specimen model to describe the relationship between the reference measurement, the local measurement and the disease outcome, that is, the above-mentioned initial likelihood function, and the initial likelihood function containing all biomarker corresponding samples is expressed as follows:
[0076]
[0077]
[0078]
[0079]
[0080]
[0081] (1)
[0082] In the formula, For research The total number of samples in For all disease outcomes corresponding to the biomarkers, are all local measurements corresponding to the biomarkers, are all the study covariates corresponding to the biomarkers, is the reference measurement value corresponding to the biomarker determined by sampling, Potential reference measurements for the remaining biomarkers that were not sampled.
[0083] For convenience, can represent the set of all estimated parameters, .
[0084] However, from formula (1), we can see that the above initial likelihood function contains not only the potential reference measurement value and the parameters to be estimated , and does not include any specific calculation formula. Therefore, this embodiment can segment the initial likelihood function and further refine formula (1) to obtain a sample likelihood function that can execute the expectation maximization algorithm.
[0085] Step A22: constructing a sample likelihood function corresponding to the biomarker according to a preset independence hypothesis and the initial likelihood function.
[0086] It is understandable that the above-mentioned assumption of independence can be used to assume that the disease outcome and local measurements Measured with a given reference is conditionally independent, which means:
[0087]
[0088] At the same time, for convenience, this embodiment can also assume that the covariate There is no effect of study-specific measurement bias, which means:
[0089]
[0090] Therefore, by combining formulas (2) and (3), formula (1) can be segmented to obtain the sample likelihood function, which can be expressed as follows:
[0091]
[0092] At this time, based on formula (4), it can be observed that the sample likelihood function It consists of three components, namely biomarker-disease association , Reference-Local Measurement Association and reference prior .
[0093] It is understandable that in order to facilitate the subsequent execution of the expectation maximization algorithm, this embodiment needs to associate the biomarker-disease , Reference-Local Measurement Association and reference prior All are converted into specific calculation methods, wherein the present embodiment may use a regression model to specifically describe the association between the biomarker and the disease in the target association function. Therefore, the present embodiment needs to determine the target regression model corresponding to the current parameter estimation.
[0094] In a feasible manner, in this embodiment, step A22 includes steps A221-A222:
[0095] Step A221, obtaining the target regression model corresponding to the current parameter estimation;
[0096] Step A222: constructing a sample likelihood function corresponding to the biomarker according to a preset independence hypothesis, the initial likelihood function and the target regression model.
[0097] It should be noted that, in this embodiment, the disease outcome corresponding to the biomarker can be a continuous outcome, a binary outcome, or a survival outcome. Therefore, this embodiment can determine the current corresponding target outcome type based on the purpose of parameter estimation of biomarkers currently aggregating multiple data sources, and then determine the corresponding target regression model based on the target outcome type.
[0098] It is easy to understand that under this framework, two different parameter estimation models can be developed in this embodiment: the EMBP continuous model for continuous outcomes and the EMBP binary model for binary outcomes. Both models can be equipped with bootstrapping techniques to estimate the confidence interval of the biological effect of the biomarker to improve the reliability of subsequent parameter estimation.
[0099] Specifically, in a first feasible implementation, this embodiment may use a linear regression model with a random intercept term to describe the correlation between biomarkers, diseases and continuous outcomes. In this case, the target regression model may be expressed as follows:
[0100]
[0101]
[0102] in, is the study-specific intercept, ,and It is the most critical parameter.
[0103] It should be understood that formula (5) mainly requires estimating , which can be the logarithm of the odds ratio (OR) describing the relationship between the biomarker and the disease.
[0104] In a second possible implementation, this embodiment can use a logistic regression model with a random intercept term to describe the association between the biomarker and the disease, as shown below:
[0105]
[0106] in, is the study-specific intercept, It is the inverse of the logit function (logistic regression model).
[0107] In addition to the target regression model, the model used to describe the reference-local measurement association in this embodiment is called a calibration model. and local measurements There is a linear correlation between them, and both follow a normal distribution. Then the calibration model can be expressed as:
[0108]
[0109] Furthermore, referring to the prior It can be expressed as:
[0110]
[0111] In the above formula, from the above analysis, we can know that , , , , , , and can be regarded as The elements contained in are parameters that characterize the degree of influence of biomarkers on disease outcomes.
[0112] Therefore, in summary, this embodiment can refine the above initial likelihood function based on formula (5) / (6), formula (7) and formula (8) to obtain a sample likelihood function that can be used for subsequent expectation maximization calculation: .
[0113] Step A3, performing expectation maximization iteration on the biological effect of the biomarker based on the sample likelihood function to obtain target estimation parameters.
[0114] It is easy to understand that since the sample likelihood function involves unknown values , so the maximum likelihood estimation introduces a challenging integral calculation. Therefore, this embodiment can use the expectation maximization algorithm to estimate all unknown parameters based on the sample likelihood function .
[0115] In this embodiment, first, local sampling and measurement of biomarkers from multiple data centers are performed to obtain reference measurement values corresponding to some biomarkers, which are assumed to be unaffected by batch effects; then, based on the reference measurement values, a sample likelihood function is constructed using the unmeasured biomarkers as latent variables, and finally, the sample likelihood function is used to execute an expectation-maximization iterative algorithm to iterate the biological effects of the biomarkers, ultimately obtaining effective target estimation parameters between biomarkers from different data sources and disease outcomes, thereby achieving effective aggregation of biomarkers from multiple data centers.
[0116] This embodiment provides a biomarker aggregation method based on expectation maximization iteration, which includes: obtaining local measurement values corresponding to biomarkers from multiple data sources; sampling and measuring the biomarkers to obtain reference measurement values; obtaining research covariates corresponding to the biomarkers; constructing an initial likelihood function corresponding to the biomarkers based on the local measurement values, the research covariates and the reference measurement values; constructing a sample likelihood function corresponding to the biomarkers according to the preset independent hypothesis and the initial likelihood function; performing expectation maximization iteration on the biological effect of the biomarkers based on the sample likelihood function to obtain target estimation parameters; the target estimation parameters characterize the effect of the biomarkers on the disease outcome. This embodiment performs local sampling and measurement on biomarkers from multiple data centers to obtain reference measurement values corresponding to some biomarkers that are assumed to be unaffected by batch effects; then constructing a sample likelihood function based on the reference measurement values using the unmeasured biomarkers as potential variables, and finally executing an expectation maximization iteration algorithm using the sample likelihood function to iterate the biological effect of the biomarkers in a loop, and finally obtaining effective target estimation parameters between biomarkers from different data sources and disease outcomes, thereby achieving effective aggregation of biomarkers from multiple data centers.
[0117] Based on the first embodiment of the present application, in the second embodiment of the present application, the same or similar contents as those in the above-mentioned embodiment 1 can be referred to the above introduction and will not be repeated later.
[0118] Based on the first embodiment, please refer to Figure 3 , Figure 3 This is a first flow chart of the second embodiment of the biomarker aggregation method based on expectation maximization iteration of the present application. In this embodiment, step A3 may include steps B1 to B3:
[0119] Step B1, initializing parameters based on the reference measurement values to obtain initial estimated parameters;
[0120] Step B2, performing expectation maximization iterative update on the initial estimation parameters according to the current estimated outcome type and the sample likelihood function;
[0121] Step B3: When it is detected that the iteration converges or reaches a preset maximum number of iterations, the current estimated parameters are used as target estimated parameters.
[0122] It is easy to understand that this embodiment can perform effective parameter estimation on the biological effects of biomarkers by using the EM algorithm. First, this embodiment can assign a reasonable value to each parameter based on the reference measurement value, and then iteratively perform the E step (Expectation step) and the M step (Maximization step) to estimate the parameters.
[0123] From the above analysis, it can be seen that the current estimated outcome type in this embodiment can be divided into a continuous outcome or a binary outcome. Therefore, if the current estimated outcome type is a continuous outcome, the method of initializing the parameters in this embodiment can be: the ordinary least square (OLS) result corresponding to the reference measurement value of the biomarker is the above-mentioned parameter to be estimated Set the initial value.
[0124] If the currently estimated outcome type is a binary outcome, the method of initializing the parameters may be: setting initial values for the parameters to be estimated based on the maximum likelihood estimation results corresponding to the reference measurement values of the biomarkers.
[0125] Then, this embodiment can repeat the E step and the M step to perform the expectation maximization iterative update until iterative convergence is detected or the preset maximum number of iterations is reached, then the expectation maximization iterative loop can be exited, and the latest current estimated parameters obtained are used as the final target estimated parameters.
[0126] Therefore, in a feasible implementation mode, referring to Figure 4 , Figure 4 This is a second flow chart of the second embodiment of the biomarker aggregation method based on expectation maximization iteration of the present application. In this embodiment, step B2 includes steps C1 to C2:
[0127] Step C1, if the current estimated outcome type is a continuous outcome, determining a corresponding estimated statistic according to the conditional distribution of the initial estimated parameters;
[0128] Step C2, performing expectation maximization iterative update on the initial estimated parameters by using an iterative reweighted least squares algorithm, the first solution, the sample likelihood function and the estimated statistic.
[0129] It should be noted that if the current estimated outcome type is a continuous outcome, In the E step, this embodiment uses the given Determine the corresponding estimated statistics. In the first step, it corresponds to the initial statistics. The conditional distribution of updates sufficient statistics: of , and can be expressed as follows:
[0130]
[0131] In the M step, this embodiment can perform numerical iterative update by giving the maximization equation (11) corresponding to the estimated statistic:
[0132] ; (11)
[0133] Where n is the total number of samples corresponding to the biomarker, For research The total number of samples corresponding to the biomarkers in It is expressed as:
[0134] ; (12)
[0135] at this time, Some parameters in can be directly solved as follows:
[0136] (13)
[0137] It is easy to understand that formula (13) can be the first solution formula mentioned above, and the remaining parameters related to the biomarker-disease association model are difficult to derive in closed form. Therefore, this embodiment can use the Iterative Reweighted Least Squares (IRLS) method to solve The other part of the parameters in , this process is achieved by iteratively applying the following formula (14):
[0138] (14)
[0139] Here, r represents the iteration step, and the initial value of each parameter in formula (14) can be set to the result of the previous m steps.
[0140] In another possible implementation, referring to Figure 5 , Figure 5 This is a third flow chart of the second embodiment of the biomarker aggregation method based on expectation maximization iteration of the present application. In this embodiment, step B2 also includes steps D1 to D4:
[0141] Step D1, if the current estimated outcome type is a binary outcome, a preset Laplace operation is performed on the initial estimated parameters to obtain a target normal posterior distribution;
[0142] It should be noted that if the current estimated outcome type is a binary outcome, then for the In the E step, we first need to obtain the posterior distribution expressed by the following formula (15):
[0143] ; (15)
[0144] in, is the log-likelihood function.
[0145] It should be understood that formula (15) is a non-normalized version of the posterior distribution and cannot be in an exact form. Therefore, this embodiment chooses to use Laplace approximation (LA) to obtain a normal approximation as an alternative to obtain the above target normal posterior distribution, which is expressed as:
[0146] ; (16)
[0147] in, is the mode of the posterior distribution, which can be obtained using the Newton-Raphson algorithm, and for The inverse of the second derivative of the posterior density function.
[0148] Step D2, obtaining the importance weight corresponding to the target normal posterior distribution;
[0149] Step D3, obtaining a joint posterior expectation based on the importance weights;
[0150] Step D4, performing expectation maximization iterative update on the initial estimated parameters by using the Newton-Raphson algorithm, the second solution, the sample likelihood function and the joint posterior expectation.
[0151] It should be noted that after obtaining an approximate representation of the posterior distribution, this embodiment also needs to calculate the posterior expectation of the joint log-likelihood. In many cases, LA can provide an accurate estimate of this posterior expectation. However, when there is a strong biomarker-disease correlation, or an unbalanced outcome distribution, the skewness of the posterior distribution increases significantly, making the estimate from LA inaccurate. Therefore, in this embodiment, importance sampling can be used to further refine the estimate of the posterior expectation.
[0152] Specifically, this embodiment can use the above target normal posterior distribution as the sample The expected distribution of the number of times, obtain the sample , and calculate their importance sampling weights, that is, obtain the above importance weights, and express them as follows:
[0153] ; (17)
[0154] in,
[0155]
[0156] Combining equations (17) and (18), the posterior expectation of the joint log-likelihood, i.e., the joint posterior expectation, can be expressed as follows:
[0157]
[0158] In addition, in order to balance the computational complexity and the estimation accuracy, this embodiment may also use a smaller N for iteration in the early stage of the EM algorithm, and gradually increase the size of N as the training proceeds.
[0159] In the M step, this embodiment can perform numerical iteration update by maximizing equation (19). At this time, this embodiment can also assume that the sample first-order moment estimate and the second moment estimate for:
[0160] ; (20)
[0161] At the same time, it can be directly obtained in closed form The estimated values of some parameters in are expressed as follows:
[0162] (twenty one)
[0163] It is easy to understand that formula (21) can be the second solution above, The remaining parameters in can not be obtained in a closed form because they involve logical transformation. At this time, the present embodiment can use the Newton-Raphson algorithm to iteratively solve them.
[0164] First, this embodiment can be used To express The remaining parameters in , and the reference measurements and covariates are combined into a vector , then we get the following expression:
[0165] ;(twenty two)
[0166] ;(twenty three)
[0167] Then use the value obtained from the previous M step to perform Initialization.
[0168] Afterwards, the present embodiment can obtain the following formulas (24) and (25):
[0169] (twenty four)
[0170] In addition, some parameters in formula (24) can be expressed as:
[0171] (25)
[0172] At this time, we can use formulas (24) and (25) to 's update until its value no longer changes, completing its iterative estimation.
[0173] In this embodiment, the pseudo code of the detailed steps of EMBP for continuous outcomes can be expressed as follows:
[0174] {
[0175] #Method 1 Expectation-maximization of a biomarker pool with continuous outcomes (EMBPc)
[0176] enter: , , , , maximum number of iterations , relative tolerance , smoothing constant ;
[0177] Output: Target estimated parameters ;
[0178] Results of ordinary least squares using all samples with available reference measurements , and set ;
[0179] While and do;
[0180] (Step E), given , calculate sufficient statistics using formulas (9) and (10) ,and ;
[0181] (M steps), given , through formula (13) , , , , ;
[0182] (M steps), through the IRLS algorithm, combined with formula (14) iterative update , , , iteratively until convergence;
[0183] ;
[0184] End while;
[0185] Return ;
[0186] }
[0187] The pseudo code of the detailed steps of EMBP for binary outcomes can be expressed as follows:
[0188] {
[0189] #Method 2 Expectation-Maximization of Biomarker Pools with Binary Outcomes (EMBPb)
[0190] enter: , , , , maximum number of iterations , relative tolerance , smoothing constant ;
[0191] Output: Target estimated parameters ;
[0192] Results of Ordinary Least Squares and Maximum Likelihood Estimation of Samples , and set ;
[0193] While and do;
[0194] (Step E), given , use the Laplace approximation to calculate the parameters of the normal approximation , obtain the target normal posterior distribution ;
[0195] (Step E), sampling by and through
[0196] Formula (17) calculates the normalized importance sampling weight ;
[0197] (M steps), update by formula (21) , , , , ;
[0198] (M steps), using the Newton-Raphson algorithm, iteratively update the parameters through formula (24) , , , until convergence;
[0199] ;
[0200] End while;
[0201] Return ;
[0202] }
[0203] It is easy to understand that the above are only two feasible implementations corresponding to step B2 provided in this embodiment, and this embodiment does not specifically limit the specific implementation of step B2. Therefore, this embodiment can provide a specific implementation method of biomarker aggregation based on expectation maximization iteration related to two disease outcomes, so as to effectively estimate the biomarkers based on the target regression model in the sample likelihood function.
[0204] This embodiment discloses that the parameters are initialized based on the reference measurement value to obtain the initial estimated parameters; if the current estimated outcome type is a continuous outcome, the corresponding estimated statistic is determined according to the conditional distribution of the initial estimated parameters; the initial estimated parameters are iteratively updated by expectation maximization through iterative reweighted least squares algorithm, the first solution, the sample likelihood function and the estimated statistic. If the current estimated outcome type is a binary outcome, the initial estimated parameters are subjected to a preset Laplace operation to obtain the target normal posterior distribution; the importance weight corresponding to the target normal posterior distribution is obtained; the joint posterior expectation is obtained based on the importance weight; the initial estimated parameters are iteratively updated by expectation maximization through the Newton-Raphson algorithm, the second solution, the sample likelihood function and the joint posterior expectation. When iterative convergence is detected or the preset maximum number of iterations is reached, the current estimated parameters are used as the target estimated parameters. This embodiment can provide a specific implementation method for the aggregation of biomarkers related to two disease outcomes based on expectation maximization iterations, thereby effectively estimating the biomarkers based on the target regression model in the sample likelihood function.
[0205] It should be noted that the above examples are only used to understand the present application and do not constitute a limitation on the biomarker aggregation method based on expectation maximization iteration of the present application. More simple transformations based on this technical concept are all within the scope of protection of the present application.
[0206] The present application also provides a biomarker aggregation device based on expectation maximization iteration, please refer to Figure 6 , Figure 6 Schematic diagram of the module structure of the biomarker aggregation device based on expectation maximization iteration in the embodiment of the present application. In this embodiment, the biomarker aggregation device based on expectation maximization iteration includes:
[0207] A data acquisition module 601 is used to acquire local measurement values corresponding to biomarkers from multiple data sources;
[0208] A reference measurement module 602 is used to perform sampling measurement on the biomarker to obtain a reference measurement value;
[0209] The data estimation module 603 is used to perform an expectation-maximization iterative algorithm based on the reference measurement value and the local measurement value to obtain a target estimation parameter; the target estimation parameter represents the effect of the biomarker on the disease outcome.
[0210] In a feasible implementation, in this embodiment, the data estimation module 603 is also used to obtain research covariates corresponding to the biomarkers;
[0211] The data estimation module 603 is further used to construct a sample likelihood function corresponding to the biomarker according to the local measurement value, the reference measurement value and the research covariate;
[0212] The data estimation module 603 is further used to perform expectation maximization iteration on the biological effect of the biomarker based on the sample likelihood function to obtain target estimation parameters.
[0213] In a feasible implementation, in this embodiment, the data estimation module 603 is further used to construct an initial likelihood function corresponding to the biomarker based on the local measurement value, the research covariate and the reference measurement value;
[0214] The data estimation module 603 is further used to construct a sample likelihood function corresponding to the biomarker according to a preset independent hypothesis and the initial likelihood function.
[0215] In a feasible implementation, in this embodiment, the data estimation module 603 is also used to obtain a target regression model corresponding to the current parameter estimation;
[0216] The data estimation module 603 is further used to construct a sample likelihood function corresponding to the biomarker according to a preset independence hypothesis, the initial likelihood function and the target regression model.
[0217] In a feasible implementation manner, in this embodiment, the data estimation module 603 is further used to perform parameter initialization based on the reference measurement value to obtain initial estimation parameters;
[0218] The data estimation module 603 is further used to perform expectation maximization iterative update on the initial estimation parameters according to the current estimated outcome type and the sample likelihood function;
[0219] The data estimation module 603 is further configured to use the current estimation parameters as target estimation parameters when it is detected that the iteration converges or a preset maximum number of iterations is reached.
[0220] In a feasible implementation, in this embodiment, the data estimation module 603 is further used to determine a corresponding estimation statistic according to the conditional distribution of the initial estimation parameter if the current estimated outcome type is a continuous outcome;
[0221] The data estimation module 603 is further used to perform expectation maximization iterative update on the initial estimation parameters through an iterative reweighted least squares algorithm, the first solution, the sample likelihood function and the estimation statistic.
[0222] In a feasible implementation, in this embodiment, the data estimation module 603 is further used to perform a preset Laplace operation on the initial estimation parameters to obtain a target normal posterior distribution if the current estimated outcome type is a binary outcome;
[0223] The data estimation module 603 is also used to obtain the importance weight corresponding to the target normal posterior distribution;
[0224] The data estimation module 603 is further used to obtain a joint posterior expectation based on the importance weight;
[0225] The data estimation module 603 is further used to perform expectation maximization iterative update on the initial estimation parameters through the Newton-Raphson algorithm, the second solution, the sample likelihood function and the joint posterior expectation.
[0226] The biomarker aggregation device based on expectation maximization iteration provided by the present application adopts the biomarker aggregation method based on expectation maximization iteration in the above-mentioned embodiment, which can solve the technical problem of how to effectively realize biomarker aggregation. Compared with the prior art, the beneficial effects of the biomarker aggregation device based on expectation maximization iteration provided by the present application are the same as the beneficial effects of the biomarker aggregation method based on expectation maximization iteration provided by the above-mentioned embodiment, and the other technical features in the biomarker aggregation device based on expectation maximization iteration are the same as the features disclosed in the above-mentioned embodiment method, which will not be repeated here.
[0227] The present application provides a biomarker aggregation device based on expectation-maximization iteration, and the biomarker aggregation device based on expectation-maximization iteration includes: at least one processor; and a memory connected to the at least one processor in communication; wherein the memory stores instructions that can be executed by the at least one processor, and the instructions are executed by the at least one processor so that the at least one processor can execute the biomarker aggregation method based on expectation-maximization iteration in the above-mentioned embodiment 1.
[0228] Reference below Figure 7, which shows a schematic diagram of the structure of a biomarker aggregation device based on expectation maximization iteration suitable for implementing the embodiment of the present application. The biomarker aggregation device based on expectation maximization iteration in the embodiment of the present application may include but is not limited to mobile terminals such as mobile phones, laptops, digital broadcast receivers, PDAs (Personal Digital Assistants), PADs (Portable Application Descriptions), PMPs (Portable Media Players), vehicle-mounted terminals (such as vehicle-mounted navigation terminals), etc., and fixed terminals such as digital TVs, desktop computers, etc. Figure 7 The illustrated biomarker aggregation device based on expectation-maximization iteration is merely an example and should not bring any limitation to the functions and scope of use of the embodiments of the present application.
[0229] like Figure 7 As shown, the biomarker aggregation device based on expectation maximization iteration may include a processing device 1001 (e.g., a central processing unit, a graphics processing unit, etc.), which can perform various appropriate actions and processes according to a program stored in a read-only memory (ROM: Read Only Memory) 1002 or a program loaded from a storage device 1003 to a random access memory (RAM: Random Access Memory) 1004. Various programs and data required for the operation of the biomarker aggregation device based on expectation maximization iteration are also stored in RAM1004. The processing device 1001, ROM1002, and RAM1004 are connected to each other via a bus 1005. An input / output (I / O) interface 1006 is also connected to the bus. Typically, the following systems may be connected to the I / O interface 1006: input devices 1007 including, for example, a touch screen, a touchpad, a keyboard, a mouse, an image sensor, a microphone, an accelerometer, a gyroscope, etc.; output devices 1008 including, for example, a liquid crystal display (LCD), a speaker, a vibrator, etc.; storage devices 1003 including, for example, a magnetic tape, a hard disk, etc.; and communication devices 1009. The communication device 1009 may allow the expectation maximization iteration-based biomarker aggregation device to communicate wirelessly or wired with other devices to exchange data. Although the figure shows an expectation maximization iteration-based biomarker aggregation device with various systems, it should be understood that it is not required to implement or have all the systems shown. More or fewer systems may be implemented or have alternatively.
[0230] In particular, according to the embodiments disclosed in the present application, the process described above with reference to the flowchart can be implemented as a computer software program. For example, the embodiments disclosed in the present application include a biomarker aggregation program product based on expectation maximization iteration, which includes a biomarker aggregation program based on expectation maximization iteration carried on a computer-readable medium, and the biomarker aggregation program based on expectation maximization iteration contains a program code for executing the method shown in the flowchart. In such an embodiment, the biomarker aggregation program based on expectation maximization iteration can be downloaded and installed from the network through a communication device, or installed from a storage device 1003, or installed from ROM1002. When the biomarker aggregation program based on expectation maximization iteration is executed by the processing device 1001, the above-mentioned functions defined in the method of the embodiment disclosed in the present application are executed.
[0231] The biomarker aggregation device based on expectation maximization iteration provided by the present application adopts the biomarker aggregation method based on expectation maximization iteration in the above-mentioned embodiment, which can solve the technical problem of how to effectively realize biomarker aggregation. Compared with the prior art, the beneficial effects of the biomarker aggregation device based on expectation maximization iteration provided by the present application are the same as the beneficial effects of the biomarker aggregation method based on expectation maximization iteration provided by the above-mentioned embodiment, and the other technical features in the biomarker aggregation device based on expectation maximization iteration are the same as the features disclosed in the method of the previous embodiment, which will not be repeated here.
[0232] It should be understood that the various parts disclosed in this application can be implemented by hardware, software, firmware or a combination thereof. In the description of the above embodiments, specific features, structures, materials or characteristics can be combined in any one or more embodiments or examples in a suitable manner.
[0233] The above are only specific implementations of the present application, but the protection scope of the present application is not limited thereto. Any technician familiar with the technical field can easily think of changes or substitutions within the technical scope disclosed in the present application, which should be included in the protection scope of the present application. Therefore, the protection scope of the present application should be based on the protection scope of the claims.
[0234] The present application provides a storage medium having computer-readable program instructions (i.e., an expectation-maximization-iteration-based biomarker aggregation program) stored thereon, and the computer-readable program instructions are used to execute the expectation-maximization-iteration-based biomarker aggregation method in the above-mentioned embodiment.
[0235] The storage medium provided in the present application may be, for example, a USB flash drive, but is not limited to electrical, magnetic, optical, electromagnetic, infrared, or semiconductor systems, systems or devices, or any combination of the above. More specific examples of storage media may include, but are not limited to: an electrical connection with one or more wires, a portable computer disk, a hard disk, a random access memory (RAM: RandomAccess Memory), a read-only memory (ROM: Read Only Memory), an erasable programmable read-only memory (EPROM: Erasable Programmable Read Only Memory or flash memory), an optical fiber, a portable compact disk read-only memory (CD-ROM: CD-Read Only Memory), an optical storage device, a magnetic storage device, or any suitable combination of the above. In this embodiment, the storage medium may be any tangible medium containing or storing a program, which may be used by or in combination with an instruction execution system, system or device. The program code contained on the storage medium may be transmitted using any appropriate medium, including but not limited to: wires, optical cables, RF (Radio Frequency: Radio Frequency), etc., or any suitable combination of the above.
[0236] The above storage medium may be included in the expectation-maximization-based iterative biomarker aggregation device; or may exist independently without being assembled into the expectation-maximization-based iterative biomarker aggregation device.
[0237] The storage medium carries one or more programs. When the one or more programs are executed by the expectation-maximization iteration-based biomarker aggregation device, the expectation-maximization iteration-based biomarker aggregation device performs: expectation-maximization iteration-based biomarker aggregation.
[0238] The expected maximization iteration-based biomarker aggregation program code for performing the operation of the present application can be written in one or more programming languages or a combination thereof, including object-oriented programming languages such as Java, Smalltalk, C++, and conventional procedural programming languages such as "C" language or similar programming languages. The program code can be executed entirely on the user's computer, partially on the user's computer, as a separate software package, partially on the user's computer and partially on a remote computer, or entirely on a remote computer or server. In the case of a remote computer, the remote computer can be connected to the user's computer through any type of network, including a local area network (LAN) or a wide area network (WAN), or can be connected to an external computer (for example, using an Internet service provider to connect through the Internet).
[0239] The flowcharts and block diagrams in the accompanying drawings illustrate the possible implementation architecture, functions and operations of the biomarker aggregation program products based on the expected maximization iteration according to the systems, methods and various embodiments of the present application. In this regard, each box in the flowchart or block diagram can represent a module, a program segment, or a part of a code, which contains one or more executable instructions for implementing the specified logical function. It should also be noted that in some alternative implementations, the functions marked in the box can also occur in a different order from the order marked in the accompanying drawings. For example, two boxes represented in succession can actually be executed substantially in parallel, and they can sometimes be executed in the opposite order, depending on the functions involved. It should also be noted that each box in the block diagram and / or flowchart, and the combination of boxes in the block diagram and / or flowchart can be implemented with a dedicated hardware-based system that performs a specified function or operation, or can be implemented with a combination of dedicated hardware and computer instructions.
[0240] The modules involved in the embodiments of the present application may be implemented by software or hardware, wherein the name of the module does not limit the unit itself in some cases.
[0241] The readable storage medium provided in the present application is a storage medium, which stores computer-readable program instructions for executing the above-mentioned biomarker aggregation method based on expectation maximization iteration (i.e., biomarker aggregation program based on expectation maximization iteration), which can solve the technical problem of how to effectively achieve biomarker aggregation. Compared with the prior art, the beneficial effects of the storage medium provided in the present application are the same as the beneficial effects of the biomarker aggregation method based on expectation maximization iteration provided in the above-mentioned embodiment, and will not be repeated here.
[0242] The above are only some embodiments of the present application, and are not intended to limit the patent scope of the present application. All equivalent structural changes made using the contents of the present application specification and drawings under the technical concept of the present application, or direct / indirect application in other related technical fields are included in the patent protection scope of the present application.
Claims
1. A biomarker aggregation method based on expectation maximization iteration, characterized in that: The method comprises: Obtaining local measurements corresponding to biomarkers from multiple data sources; Sampling and measuring the biomarkers to obtain reference measurement values; Executing an expectation-maximization iterative algorithm according to the reference measurement value and the local measurement value to obtain a target estimation parameter; the target estimation parameter represents the effect of the biomarker on the disease outcome; The step of performing an expectation maximization iterative algorithm according to the reference measurement value and the local measurement value to obtain a target estimation parameter comprises: Obtain the study covariates corresponding to the biomarkers; constructing a sample likelihood function corresponding to the biomarker according to the local measurement value, the reference measurement value and the research covariate; Performing expectation maximization iteration on the biological effect of the biomarker based on the sample likelihood function to obtain a target estimation parameter; The step of performing expectation maximization iteration on the biological effect of the biomarker based on the sample likelihood function to obtain target estimation parameters comprises: Initialize parameters based on the reference measurement values to obtain initial estimated parameters; Performing expectation maximization iterative update on the initial estimation parameters according to the current estimated outcome type and the sample likelihood function; When it is detected that the iteration converges or reaches the preset maximum number of iterations, the current estimated parameters are used as the target estimated parameters; The step of performing expectation maximization iterative updating on the initial estimation parameters according to the current estimated outcome type and the sample likelihood function further includes: If the current estimated outcome type is a binary outcome, a preset Laplace operation is performed on the initial estimated parameters to obtain a target normal posterior distribution; Obtaining the importance weight corresponding to the target normal posterior distribution; Obtaining a joint posterior expectation based on the importance weights; The initial estimated parameters are iteratively updated by expectation maximization through the Newton-Raphson algorithm, the second solution, the sample likelihood function and the joint posterior expectation.
2. The biomarker aggregation method based on expectation maximization iteration according to claim 1, characterized in that: The step of constructing the sample likelihood function corresponding to the biomarker according to the local measurement value, the reference measurement value and the research covariate comprises: constructing an initial likelihood function corresponding to the biomarker based on the local measurement value, the study covariate and the reference measurement value; A sample likelihood function corresponding to the biomarker is constructed according to a preset independence hypothesis and the initial likelihood function.
3. The biomarker aggregation method based on expectation maximization iteration according to claim 2, characterized in that: The step of constructing the sample likelihood function corresponding to the biomarker according to the preset independent hypothesis and the initial likelihood function comprises: Get the target regression model corresponding to the current parameter estimate; A sample likelihood function corresponding to the biomarker is constructed according to the preset independence hypothesis, the initial likelihood function and the target regression model.
4. The biomarker aggregation method based on expectation maximization iteration according to claim 3, characterized in that: The step of iteratively updating the initial estimated parameters by performing expectation maximization according to the current estimated outcome type and the sample likelihood function comprises: If the current estimated outcome type is a continuous outcome, determining a corresponding estimated statistic according to the conditional distribution of the initial estimated parameters; The initial estimated parameters are iteratively updated by expectation maximization through an iterative reweighted least squares algorithm, a first solution, the sample likelihood function and the estimated statistic.
5. A biomarker aggregation device based on expectation maximization iteration, characterized in that: The biomarker aggregation device based on expectation maximization iteration includes: A data acquisition module, used to acquire local measurement values corresponding to biomarkers from multiple data sources; A reference measurement module, used to perform sampling measurement on the biomarker to obtain a reference measurement value; A data estimation module, configured to perform an expectation-maximization iterative algorithm based on the reference measurement value and the local measurement value to obtain a target estimation parameter; the target estimation parameter represents the effect of the biomarker on the disease outcome; The data estimation module is further used to obtain a research covariate corresponding to the biomarker; construct a sample likelihood function corresponding to the biomarker according to the local measurement value, the reference measurement value and the research covariate; and perform expectation maximization iteration on the biological effect of the biomarker based on the sample likelihood function to obtain a target estimation parameter; The data estimation module is further used to initialize parameters based on the reference measurement values to obtain initial estimation parameters; perform expectation maximization iterative update on the initial estimation parameters according to the current estimation outcome type and the sample likelihood function; and use the current estimation parameters as target estimation parameters when it is detected that the iteration converges or reaches a preset maximum number of iterations; The data estimation module is also used to perform a preset Laplace operation on the initial estimation parameters to obtain a target normal posterior distribution if the current estimated outcome type is a binary outcome; obtain the importance weight corresponding to the target normal posterior distribution; obtain the joint posterior expectation based on the importance weight; and perform expectation maximization iterative update on the initial estimation parameters through the Newton-Raphson algorithm, the second solution, the sample likelihood function and the joint posterior expectation.
6. A biomarker aggregation device based on expectation maximization iteration, characterized in that: The device includes: a memory, a processor, and an expectation-maximization iteration-based biomarker aggregation program stored in the memory and executable on the processor, wherein the expectation-maximization iteration-based biomarker aggregation program is configured to implement the steps of the expectation-maximization iteration-based biomarker aggregation method as described in any one of claims 1 to 4.
7. A storage medium, characterized in that: The storage medium stores an expectation-maximization iteration-based biomarker aggregation program, which, when executed by a processor, implements the steps of the expectation-maximization iteration-based biomarker aggregation method as described in any one of claims 1 to 4.
Citation Information
Patent Citations
Method for generating diagnostic model, and diagnostic method and device using same
CN116052777A
Model parameter estimation method based on FCM and EM
CN117493924A
SuStaIn-Advance-based disease staging and subtype identification method
CN118116582A