A multi-center data aggregation method

Through the multi-center data aggregation method, the Bayesian algorithm is used to aggregate biomarker data, which solves the problem of low aggregation accuracy of data from different sources, and improves the accuracy of parameter estimation between biomarkers and disease outcomes.

CN119479830BActive Publication Date: 2025-05-16PEKING UNIV
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202510060768.6
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-01-15
Publication Date
2025-05-16
Estimated Expiration
2045-01-15

AI Technical Summary

Technical Problem

The prior art is difficult to effectively aggregate biomarker data from different sources, resulting in the impact of parameter estimation accuracy between biomarkers and disease outcomes.

Method used

A multi-center data aggregation method is proposed to obtain the reference measurement value by obtaining the local measurement value of biomarkers from multiple data centers, performing sampling measurements, and performing Bayesian posterior estimation based on the reference measurement value and local measurement values ​​is used to obtain the target estimation parameters.

Benefits of technology

Improve the accuracy of biomarker parameter estimation and realize effective aggregation of biomarker data from multiple data centers.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119479830B_ABST
    Figure CN119479830B_ABST
Patent Text Reader

Abstract

The present application discloses a multi-center data aggregation method, which relates to the field of medical data processing technology. The method includes: obtaining local measurement values ​​corresponding to biomarkers from multiple data centers; sampling and measuring the biomarkers to obtain some reference measurement values; using the Bayesian algorithm to estimate the posterior distribution of the biological effects of the biomarkers based on the local measurement values, reference measurement values ​​and research covariates to obtain target estimation parameters. The present application first measures some biomarkers, obtains reference measurement values ​​that are not affected by batch effects, and then performs Bayesian posterior estimation of the biological effects of the biomarkers based on the reference measurement values, local measurement values ​​and research covariates, and finally obtains effective target estimation parameters between the biomarkers and disease outcomes, thereby achieving effective aggregation of biomarkers from multiple data sources.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present application relates to the technical field of medical data processing, and in particular to a multi-center data aggregation method. Background Art

[0002] Biomarkers, as indicators that can objectively measure and evaluate normal biological processes, pathological processes, or responses to drug interventions, play an important role in medical research and clinical practice. Their utility as surrogate endpoints has been increasingly recognized, especially when validated through rigorous studies. Biomarkers can conduct clinical trials in smaller cohorts in a shorter period of time, effectively improving trial efficiency. Therefore, determining the correlation between biomarkers and clinical outcomes is particularly important in drug development and regulatory decision-making. In order to improve the statistical power and precision of the correlation, combining biomarker data from multiple data centers for analysis has become a popular method for quantifying biomarker-disease associations.

[0003] However, this approach is complicated by inter-study variability, primarily due to batch effects in biomarkers from different data centers due to differences in analytical methods. For example, measurements of circulating 25-hydroxyvitamin D (25(OH)D) can show up to 40% variability in analyses performed at different data centers. This variability can affect the accuracy of the parameter estimates that ultimately map the biomarker to the disease outcome.

[0004] Therefore, how to effectively aggregate biomarker data from different sources to improve the accuracy of subsequent biomarker parameter estimation has become a pressing issue to be addressed. Summary of the invention

[0005] The main purpose of this application is to provide a multi-center data aggregation method, which aims to solve the technical problem of how to effectively aggregate biomarker data from different sources to improve the accuracy of subsequent biomarker parameter estimation.

[0006] To achieve the above objectives, this application proposes a multi-center data aggregation method, which includes:

[0007] Obtain local measurements corresponding to biomarkers from multiple data centers;

[0008] Sampling and measuring the biomarkers to obtain reference measurement values;

[0009] A Bayesian posterior estimation of the biological effect of the biomarker is performed based on the reference measurement value and the local measurement value by a Bayesian algorithm to obtain a target estimation parameter; the target estimation parameter characterizes the effect of the biomarker on the disease outcome.

[0010] In one embodiment, the step of performing a Bayesian posterior estimation of the biological effect of the biomarker based on the reference measurement value and the local measurement value by a Bayesian algorithm to obtain a target estimation parameter comprises:

[0011] Obtaining research covariates corresponding to the biomarkers;

[0012] constructing local estimation information corresponding to the biomarker according to the study covariate, the reference measurement value and the local measurement value;

[0013] A Bayesian posterior estimation is performed on the biomarker based on the local estimation information to obtain a target estimation parameter.

[0014] In one embodiment, the step of performing Bayesian posterior estimation on the biomarker based on the local estimation information to obtain target estimation parameters comprises:

[0015] Performing Bayesian sampling on the biomarker using a preset sampler and the local estimation information to obtain a biomarker posterior sample;

[0016] Statistical estimation is performed on the biomarker posterior samples to obtain target estimation parameters.

[0017] In one embodiment, the step of constructing the local estimation information corresponding to the biomarker according to the study covariate, the reference measurement value and the local measurement value comprises:

[0018] constructing a sample likelihood function corresponding to the biomarker according to the research covariate, the reference measurement value and the local measurement value;

[0019] Constructing a joint prior distribution corresponding to the local measurement values;

[0020] Local estimation information corresponding to the biomarker is constructed based on the sample likelihood function and the joint prior distribution.

[0021] In one embodiment, the step of constructing a sample likelihood function corresponding to the biomarker according to the study covariate, the reference measurement value and the local measurement value comprises:

[0022] constructing an initial likelihood function corresponding to the biomarker based on the local measurement value, the study covariate and the reference measurement value;

[0023] A sample likelihood function corresponding to the biomarker is constructed according to a preset independence hypothesis and the initial likelihood function.

[0024] In one embodiment, the step of constructing the sample likelihood function corresponding to the biomarker according to the preset independence hypothesis and the initial likelihood function includes:

[0025] Get the target regression model corresponding to the current parameter estimate;

[0026] A sample likelihood function corresponding to the biomarker is constructed according to the preset independence hypothesis, the initial likelihood function and the target regression model.

[0027] In one embodiment, the step of obtaining a target regression model corresponding to the current parameter estimate includes:

[0028] Get the target outcome type corresponding to the current parameter estimate;

[0029] The target regression model is determined according to the target outcome type.

[0030] In addition, to achieve the above purpose, the present application also proposes a multi-center data aggregation device, which includes:

[0031] A data acquisition module, used to obtain local measurements and study covariates corresponding to biomarkers from multiple data centers;

[0032] A reference measurement module, used to perform sampling measurement on the biomarker to obtain a reference measurement value;

[0033] A data estimation module is used to estimate the posterior distribution of the local measurement value based on the reference measurement value and the research covariate through a Bayesian algorithm to obtain a target estimation parameter; the target estimation parameter characterizes the effect of the biomarker on the disease outcome.

[0034] In addition, to achieve the above-mentioned objectives, the present application also proposes a multi-center data aggregation device, which includes: a memory, a processor, and a multi-center data aggregation program stored in the memory and executable on the processor, and the multi-center data aggregation program is configured to implement the steps of the multi-center data aggregation method as described above.

[0035] In addition, to achieve the above-mentioned purpose, the present application also proposes a storage medium, which is a storage medium on which a multi-center data aggregation program is stored. When the multi-center data aggregation program is executed by a processor, the steps of the multi-center data aggregation method as described above are implemented.

[0036] The present application provides a multi-center data aggregation method, which includes: obtaining local measurement values ​​corresponding to biomarkers from multiple data centers; sampling and measuring the biomarkers to obtain some reference measurement values; using the Bayesian algorithm to estimate the posterior distribution of the biological effects of the biomarkers based on the local measurement values, reference measurement values ​​and research covariates to obtain target estimation parameters. The present application first measures some biomarkers, obtains reference measurement values ​​that are not affected by batch effects, and then regards the reference measurement values ​​of the unmeasured biomarkers as estimable latent variables based on the reference measurement values, and performs Bayesian posterior estimation of the biological effects of the biomarkers based on the reference measurement values, local measurement values ​​and research covariates, and finally obtains effective target estimation parameters between the biomarkers and disease outcomes, thereby achieving effective aggregation of biomarkers from multiple data sources. BRIEF DESCRIPTION OF THE DRAWINGS

[0037] The accompanying drawings, which are incorporated in and constitute a part of this specification, illustrate embodiments consistent with the present application and, together with the description, serve to explain the principles of the present application.

[0038] In order to more clearly illustrate the embodiments of the present application or the technical solutions in the prior art, the drawings required for use in the embodiments or the description of the prior art will be briefly introduced below. Obviously, for ordinary technicians in this field, other drawings can be obtained based on these drawings without paying any creative labor.

[0039] Figure 1 This is a first flow chart of the first embodiment of the multi-center data aggregation method of the present application;

[0040] Figure 2 This is a second flow chart of the first embodiment of the multi-center data aggregation method of the present application;

[0041] Figure 3 This is a first flow chart of the second embodiment of the multi-center data aggregation method of the present application;

[0042] Figure 4 This is a second flow chart of the second embodiment of the multi-center data aggregation method of the present application;

[0043] Figure 5 This is a third flow chart of the second embodiment of the multi-center data aggregation method of the present application;

[0044] Figure 6 This is a schematic diagram of the module structure of a multi-center data aggregation device according to an embodiment of the present application;

[0045] Figure 7 Schematic diagram of the device structure of the hardware operating environment involved in the multi-center data aggregation method in the embodiment of the present application.

[0046] The purpose, features and advantages of this application will be further described in conjunction with the embodiments and with reference to the accompanying drawings. DETAILED DESCRIPTION

[0047] It should be understood that the specific embodiments described herein are only used to explain the technical solutions of the present application and are not used to limit the present application.

[0048] In order to better understand the technical solution of the present application, a detailed description will be given below in conjunction with the accompanying drawings and specific implementation methods.

[0049] The main solution of this application is: to obtain local measurement values ​​corresponding to biomarkers from multiple data centers; to perform sampling measurements on biomarkers to obtain reference measurement values; to perform Bayesian posterior estimation of the biological effects of biomarkers based on reference measurement values ​​and local measurement values ​​through a Bayesian algorithm to obtain target estimation parameters; the target estimation parameters characterize the effect of biomarkers on disease outcomes.

[0050] Currently, pooling biomarkers from multiple research centers can effectively improve the statistical power and precision of quantifying biomarker-disease associations. However, due to the variability of biomarkers between different research centers, this batch effect will affect the accuracy of parameter estimates between subsequent biomarkers and disease outcomes.

[0051] To solve this problem, the present application can calibrate the reference analysis before pooling biomarkers to standardize the biomarkers of the local research center. Based on this, the present application further developed a new Bayesian Biomarker Pooling (BBP) method to collect biomarkers from multiple research sources or multiple data centers. This method regards the reference measurement values ​​of local biological specimens that have not been reanalyzed as estimable latent variables, thereby performing a Bayesian posterior distribution of the biomarker-disease association, and finally determining the effective estimation parameters corresponding to the biological effects of the biomarkers and disease outcomes based on the results of the posterior distribution analysis.

[0052] Specifically, the present application first obtains reference measurement values ​​corresponding to some biomarkers that are assumed to be unaffected by batch effects by performing local sampling measurements on biomarkers from multiple data centers; then, based on the determined reference measurement values, the reference measurement values ​​of the unmeasured biomarkers are regarded as estimable latent variables, thereby constructing a probability graph relationship between the reference measurement values, local measurement values, and biological effects of the biomarkers; finally, based on the probability graph relationship, a Bayesian posterior estimation of the biological effects of the biomarkers is performed to obtain effective target estimation parameters between the biomarkers from different data centers and disease outcomes, thereby achieving effective aggregation of biomarkers from multiple data centers.

[0053] It should be noted that the execution subject of this embodiment can be a multi-center data aggregation system, or a computing service device with data processing, network communication and program running functions, such as a tablet computer, a personal computer, a mobile phone, etc., or a multi-center data aggregation device capable of realizing the above functions, etc., and this embodiment does not specifically limit this. The following takes the multi-center data aggregation device (referred to as the aggregation device) as the execution subject as an example to illustrate this embodiment and the following embodiments.

[0054] Based on this, the present application embodiment provides a multi-center data aggregation method, referring to Figure 1 , Figure 1 This is a first flow chart of the first embodiment of the multi-center data aggregation method of the present application.

[0055] In this embodiment, the multi-center data aggregation method includes steps S10 to S30:

[0056] Step S10, obtaining local measurement values ​​corresponding to biomarkers from multiple data centers;

[0057] It can be understood that the above-mentioned biomarkers from multiple data centers indicate that the biomarkers can be research indicators from different data sources, different clinical research centers or different laboratories, and the above-mentioned local measurement values ​​can be measurement values ​​obtained after a specific determination of the biomarkers in the laboratory where the data originated.

[0058] It should be understood that the purpose of aggregating the above-mentioned biomarkers from different data centers is to estimate the degree of influence between biomarkers and disease outcomes. Therefore, if data sets from multiple data centers are reasonably used in the process of impact analysis, it can be regarded as completing the effective aggregation of biomarkers from multiple data centers.

[0059] Step S20, sampling and measuring the biomarkers to obtain reference measurement values;

[0060] It is easy to understand that there are batch effects in biomarkers from different data centers. Therefore, in order to ensure the accuracy of parameter estimation of biomarkers after aggregation, this embodiment can reduce this difference by re-measuring and analyzing the biomarkers.

[0061] However, it is costly to reanalyze all biological samples from multiple research centers in a single reference laboratory, and not all biological samples can be remotely operated. Therefore, limited by the impact of cost and distance, this embodiment can extract part of the data from the biomarkers as a sample subset for re-measurement, that is, sample and measure the biomarkers, obtain the reference measurement values ​​corresponding to the sample subset biomarkers, and assume that the biomarkers that have not been re-measured and analyzed, which can be referred to as the remaining biomarkers here, are regarded as estimable latent variables corresponding to the reference measurement values ​​of the current reference laboratory (that is, the data center or laboratory where the current aggregation device is located), and establish a calibration model for a specific study for each data center based on this assumption.

[0062] It is understandable that this embodiment assumes that represents the research indicators from S different laboratories, namely the above biomarkers, then for the research Individuals in , the disease outcome can be expressed as , the biomarker measurement from the local reference laboratory (i.e., the reference measurement mentioned above) can be expressed as , the biomarker measurements from different study-specific data centers (i.e., the local measurements mentioned above) can be expressed as . Combining the above analysis, we can see that for the research All biomarkers in the However, only some biomarkers have corresponding reference measurement values ​​after the above sampling determination. In addition, it is assumed in this embodiment that for the study Individuals in , Representatives can obtain reference measurements, otherwise . And the reference measurement value The distribution of is consistent across studies, whereas local measurements The distribution of was heterogeneous in different studies.

[0063] Step S30, performing Bayesian posterior estimation on the biomarker based on the reference measurement value and the local measurement value by using a Bayesian algorithm to obtain a target estimation parameter; the target estimation parameter characterizes the effect of the biomarker on the disease outcome.

[0064] It can be understood that the present embodiment can regard the reference measurement values ​​of the biomarkers that have not been re-measured as estimable latent variables. Specifically, the present embodiment can calibrate and summarize the reference measurement values ​​corresponding to the remaining biomarkers based on the reference measurement values ​​of the biomarkers that have been re-measured and analyzed. Commonly used calibration methods may include a two-stage method and a summary method. However, such methods are usually based on two basic assumptions: 1) the noise in the calibration model is moderate, or it is assumed that the association between the disease outcome and the reference measurement of the biomarker is not too strong; 2) the disease is rare.

[0065] However, these assumptions may not be generally valid in the context of clinical development projects. Typically, the association between biomarkers and disease manifestations is obvious, and the patient cohorts used to explore these biomarker-disease associations are usually extracted from clinical studies. That is, the applicability of traditional calibration methods is general, and this embodiment needs to propose a calibration analysis method that is adapted to the robustness of biomarker-disease associations commonly encountered in clinical studies.

[0066] Therefore, in this example, in order to pool and qualify biomarker-disease associations from multiple clinical studies, a novel Bayesian biomarker pooling (BBP) method was introduced, which aggregated biomarkers from multiple data centers using a Bayesian algorithm.

[0067] In one possible implementation, refer to Figure 2 , Figure 2 This is a second flow chart of the first embodiment of the multi-center data aggregation method of the present application. In this embodiment, step S30 may include steps A1 to A3:

[0068] Step A1, obtaining the research covariate corresponding to the biomarker;

[0069] It should be understood that the above-mentioned research covariates may be other factors that affect the disease outcome of the individual i in the study, such as the individual's age, gender, race, etc. A vector representing the study covariates.

[0070] Step A2, constructing local estimation information corresponding to the biomarker according to the research covariate, the reference measurement value and the local measurement value;

[0071] Step A3: Perform Bayesian posterior estimation on the biomarker based on the local estimation information to obtain target estimation parameters.

[0072] It is easy to understand that in this embodiment, the reference measurement of the biological specimen that has not been reanalyzed is used as an observable latent variable, and the posterior distribution corresponding to the remaining biomarkers and the corresponding assumed reference measurement values ​​is constructed based on the already determined reference measurement values ​​and the research covariates, that is, the above-mentioned local estimation information. Based on the local estimation information, a Bayesian posterior estimation of the biological effect of the biomarker, that is, the effect between the biomarker and the disease outcome is performed, and the above-mentioned target estimation parameters can be obtained, thereby realizing the effective aggregation of biomarkers from multiple data centers.

[0073] In a feasible implementation manner, in this embodiment, step A3 may include steps A31-A32:

[0074] Step A31, performing Bayesian sampling on the biomarker by using a preset sampler and the local estimation information to obtain a biomarker posterior sample;

[0075] Step A32, performing statistical estimation on the biomarker posterior samples to obtain target estimation parameters.

[0076] It can be understood that the present application can use the sampler corresponding to the Markov Chain Monte Carlo (MCMC) method as the above-mentioned preset sampler, such as NUTS (No-U-Turn Sampler) to perform Bayesian sampling on the local estimation information to obtain the corresponding biomarker posterior samples.

[0077] Specifically, this embodiment can calculate the sample mean of the biomarker posterior samples based on the MCMC sampler as the point estimate of the target estimation parameter, and calculate the 95% highest posterior density interval corresponding to the biomarker posterior samples as the interval estimate of the target estimation parameter.

[0078] In this embodiment, a sampling measurement is performed on some biomarkers among the biomarkers collected from multiple data centers to obtain reference measurement values ​​corresponding to the part of the biomarkers, which are assumed to be unaffected by batch effects; then, based on the reference measurement values, a Bayesian posterior estimation is performed on the potential reference measurement values ​​of the remaining biomarkers that have not been re-measured and analyzed, and the potential reference measurement values ​​of the remaining biomarkers are converted into estimable quantities to obtain the corresponding joint posterior distribution, i.e., local estimation information; finally, a statistical estimation of the biological effects of the biomarkers is performed based on the local estimation information, and ultimately an effective Bayesian posterior estimation parameter between the biomarkers and the disease outcomes is obtained, thereby achieving effective aggregation of biomarkers from multiple data sources.

[0079] The present embodiment provides a multi-center data aggregation method, which includes: obtaining local measurement values ​​corresponding to biomarkers from multiple data centers; sampling and measuring the biomarkers to obtain reference measurement values; obtaining research covariates corresponding to the biomarkers; constructing local estimation information corresponding to the biomarkers based on the research covariates, reference measurement values ​​and local measurement values; performing Bayesian sampling on the biomarkers through a preset sampler and local estimation information to obtain biomarker posterior samples; performing statistical estimation on the biomarker posterior samples to obtain target estimation parameters; the target estimation parameters characterize the effect of the biomarker on the disease outcome. This embodiment performs sampling measurement on some biomarkers collected from multiple data centers to obtain reference measurement values ​​corresponding to these biomarkers, which are assumed to be unaffected by batch effects; then, based on the reference measurement values, Bayesian posterior estimation is performed on the potential reference measurement values ​​of the remaining biomarkers that have not been re-measured and analyzed, and the potential reference measurement values ​​of the remaining biomarkers are converted into estimable quantities to obtain the corresponding joint posterior distribution, i.e., local estimation information; finally, based on the local estimation information, statistical estimation of the biological effects of the biomarkers is performed, and finally effective Bayesian posterior estimation parameters between the biomarkers and the disease outcomes are obtained, thereby achieving effective aggregation of biomarkers from multiple data sources.

[0080] Based on the first embodiment of the present application, in the second embodiment of the present application, the same or similar contents as those in the above-mentioned embodiment 1 can be referred to the above introduction and will not be repeated later.

[0081] Based on the first embodiment, please refer to Figure 3 , Figure 3 This is a first flow chart of the second embodiment of the multi-center data aggregation method of the present application. In this embodiment, step A2 includes steps B1 to B3:

[0082] Step B1, constructing a sample likelihood function corresponding to the biomarker according to the research covariate, the reference measurement value and the local measurement value;

[0083] In one possible implementation, refer to Figure 4 , Figure 4 This is a second flow chart of the second embodiment of the multi-center data aggregation method of the present application. In this embodiment, step B1 may include steps C1-C2:

[0084] Step C1, constructing an initial likelihood function corresponding to the biomarker based on the local measurement value, the research covariate and the reference measurement value;

[0085] It can be understood that in order to perform calibration analysis on the local measurement values ​​of the remaining biomarkers based on the determined reference measurement values ​​and to estimate the reference measurement values ​​of the remaining biomarkers, this embodiment can preliminarily establish a two-level research-biological specimen model to describe the relationship between the reference measurement, the local measurement and the disease outcome, that is, the above-mentioned initial likelihood function, and the initial likelihood function containing the samples corresponding to all biomarkers is expressed as follows:

[0086]

[0087]

[0088]

[0089]

[0090]

[0091] (1)

[0092] In the formula, for convenience, represents the set of all estimated parameters, For research The total number of samples in For all disease outcomes corresponding to the biomarkers, are all local measurements corresponding to the biomarkers, are all the study covariates corresponding to the biomarkers, is the reference measurement value corresponding to the biomarker determined by sampling, Potential reference measurements for the remaining biomarkers that were not sampled.

[0093] However, from formula (1), we can see that the above initial likelihood function contains not only the potential reference measurement value and the parameters that need to be estimated , and does not include any specific calculation formula. Therefore, this embodiment can segment the initial likelihood function and further refine formula (1) to obtain a sample likelihood function that can be used for Bayesian calculation.

[0094] Step C2: constructing a sample likelihood function corresponding to the biomarker according to a preset independence hypothesis and the initial likelihood function.

[0095] It is understandable that the above-mentioned assumption of independence can be used to assume that the disease outcome and local measurements Measured with a given reference is conditionally independent, which means:

[0096]

[0097] At the same time, for convenience, this embodiment can also assume that the covariate There is no effect of study-specific measurement bias, which means:

[0098]

[0099] Therefore, by combining formulas (2) and (3), formula (1) can be segmented to obtain the sample likelihood function, which can be expressed as follows:

[0100]

[0101] At this point, it can be observed that the sample likelihood function It consists of three components, namely biomarker-disease association , Reference-Local Measurement Association and reference prior .

[0102] In one possible implementation, refer to Figure 5 , Figure 5 This is a third flow chart of the second embodiment of the multi-center data aggregation method of the present application. In this embodiment, step C2 may include steps C21-C22:

[0103] Step C21, obtaining the target regression model corresponding to the current parameter estimation;

[0104] It is understandable that in order to facilitate the subsequent Bayesian calculation, this embodiment needs to associate the biomarker-disease , Reference-Local Measurement Association and reference prior All are converted into specific calculation methods, wherein this embodiment can use a regression model to specifically describe the association between biomarkers and diseases in the target association function. Therefore, this embodiment needs to obtain the target regression model corresponding to the current parameter estimation.

[0105] In a feasible implementation manner, in this embodiment, step C21 may include steps C211 to C212:

[0106] Step C211, obtaining the target outcome type corresponding to the current parameter estimation;

[0107] Step C212, determining a target regression model according to the target outcome type.

[0108] It should be noted that the disease outcome corresponding to the biomarker can be a continuous outcome, a binary outcome, or a survival outcome. Therefore, this embodiment can determine the current corresponding target outcome type based on the purpose of aggregating biomarkers from multiple data centers for parameter estimation, and then determine the corresponding target regression model based on the target outcome type.

[0109] In a first feasible implementation, if the target outcome type is a binary outcome, this embodiment can use a logistic regression model with a random intercept term to describe the association between the biomarker and the disease. In this case, the target regression model can be expressed as follows:

[0110]

[0111] in, is the study-specific intercept, It is the inverse of the logit function (logistic regression model).

[0112] It should be understood that formula (5) mainly requires estimating , which can be the logarithm of the odds ratio (OR) describing the relationship between the biomarker and the disease.

[0113] In a second feasible implementation, if the target outcome type is a continuous outcome, this embodiment may use a linear regression model to describe the association between the biomarker and the disease. In this case, the target regression model may be expressed as follows:

[0114]

[0115] In a third feasible implementation, if the target outcome type is a survival outcome, this embodiment may use a Weibull regression model to describe the association between the biomarker and the disease. In this case, the target regression model may be expressed as follows:

[0116]

[0117] Step C22, constructing a sample likelihood function corresponding to the biomarker according to a preset independence hypothesis, the initial likelihood function and the target regression model.

[0118] It is understood that after the biomarker-disease association is determined by the target regression model, the model used to describe the reference-local measurement association can be called a calibration model. At this time, the present embodiment can assume that the reference measurement value and local measurements There is a linear correlation between them, and both follow a normal distribution, so the calibration model can be expressed as:

[0119]

[0120] Furthermore, the above reference prior It can be expressed as:

[0121]

[0122] In the above formula, from the above analysis, we can know that , , , , , , , and can be regarded as The elements contained in are parameters that characterize the degree of influence of biomarkers on disease outcomes.

[0123] Therefore, in summary, this embodiment can refine the above initial likelihood function based on formula (5) / (6) / (7), formula (8) and formula (9) to obtain a sample likelihood function that can be used for subsequent Bayesian calculations: .

[0124] Step B2, constructing a joint prior distribution corresponding to the local measurement value;

[0125] It is important to understand that since the sample likelihood function involves unknown values Therefore, the use of the above maximum likelihood estimation introduces a challenging integral calculation. In this embodiment, after treating the reference measurement values ​​corresponding to the remaining biomarkers that have not been re-measured and analyzed as latent variables, it is necessary to convert the biological effect of the biomarker into an estimable quantity by constructing an appropriate prior distribution, that is, the above joint prior distribution.

[0126] At the same time, this embodiment assumes that and The joint prior distribution of is factorized and independent of the covariates , then the above joint prior distribution can be expressed as:

[0127]

[0128] In the formula, , , , , , , for The specific parameters contained in are all parameters that characterize the degree of influence of biomarkers on disease outcomes.

[0129] It will be appreciated that since this example assumes that the reference measurements are consistent across studies, and should be identically distributed. Therefore , ,and It is not independent.

[0130] At this time, based on the assumption of normal distribution, this embodiment can express the prior distribution of these three variables as follows:

[0131]

[0132] in, and It can be set based on reference measurements from real-world scenarios or, without loss of generality, can be determined using conjugate priors such as normal and inverse gamma distributions.

[0133] At the same time, this embodiment can also place a non-informative prior, and express it as follows:

[0134]

[0135] And for and , this embodiment can use information prior or weak information prior to ensure the stability of statistical performance. For example, assuming and It obeys a standard normal distribution, which is expressed as follows:

[0136]

[0137] in, , is the number of covariates.

[0138] And for the parameters specific to study S , , and In a feasible implementation, this embodiment may assume a “random effect” model so that any vector between them can be sampled from a common distribution.

[0139] In another possible implementation, since in the Bayesian framework, it is not necessary to assume that the experiment is sampled from a super population. Based on the qualitative assumption of "exchangeability", this embodiment can determine that there are no differences between different studies and continue to simply assume that the parameters of each study are independent samples from a super population distribution controlled by some unknown super parameters. For example, , and It obeys the normal distribution and is expressed as:

[0140]

[0141]

[0142]

[0143] And for , which can be made to obey the inverse gamma distribution and expressed as:

[0144]

[0145] In the formula , ; , ; , ; and They are respectively ; b; ; The corresponding hyperparameters can also be considered as unknown parameters part of express It obeys the inverse gamma distribution and can be expressed as follows:

[0146]

[0147] Similarly, for and The prior distribution of , this embodiment assumes that an appropriate weak information prior is assigned to these hyperparameters. , and The prior distribution of can be expressed using a standard normal distribution, as follows:

[0148]

[0149] For the above , , , and , then it can be represented by a semi-Cauchy distribution:

[0150]

[0151] In the formula, That is the corresponding , , , .

[0152] Therefore, by combining the above formulas (10) to (20), we can obtain the expression of the joint prior distribution.

[0153] Step B3: constructing local estimation information corresponding to the biomarker based on the sample likelihood function and the joint prior distribution.

[0154] It is easy to understand that in this embodiment, the posterior distribution of the assumed reference measurement values ​​corresponding to the remaining biomarkers can be constructed based on the sample likelihood function and the joint prior distribution to obtain local estimation information.

[0155] Specifically, according to the Bayesian formula, this embodiment can obtain unknown target estimation parameters and assumed reference measurement values ​​based on the target regression model and the joint prior distribution The unnormalized joint posterior distribution of , that is, the local estimation information, is expressed as follows:

[0156]

[0157] In this embodiment, a two-level study-biospecimen model is established to describe the relationship between reference measurements, local measurements and disease outcomes, and the above-mentioned sample likelihood function is obtained. The reference measurement values ​​corresponding to the remaining biomarkers that have not been re-measured and analyzed are then treated as latent variables, and converted into estimable variables by constructing an appropriate prior distribution, namely the above-mentioned joint prior distribution. Finally, based on the sample likelihood function and the joint prior distribution, a non-normalized joint posterior distribution between the assumed reference measurement values ​​of the remaining biomarkers and the target estimated parameters is constructed, namely the local estimation information, thereby effectively estimating the biological effects of the biomarkers and providing a robust analysis framework for integrating biomarkers from multiple data centers.

[0158] This embodiment discloses constructing an initial likelihood function corresponding to a biomarker based on local measurements, research covariates, and reference measurements; obtaining the target outcome type corresponding to the current parameter estimate; determining the target regression model according to the target outcome type; constructing a sample likelihood function corresponding to the biomarker according to the preset independent assumption, the initial likelihood function, and the target regression model. Constructing a joint prior distribution corresponding to the local measurements; constructing local estimation information corresponding to the biomarker based on the sample likelihood function and the joint prior distribution. This embodiment describes the relationship between the reference measurement, the local measurement, and the disease outcome by establishing a sample likelihood function, and then treating the reference measurement values ​​corresponding to the remaining biomarkers that have not been re-measured and analyzed as latent variables, converting them into estimable variables by constructing an appropriate joint prior distribution, and finally constructing a non-normalized joint posterior distribution between the assumed reference measurement values ​​of the remaining biomarkers and the target estimation parameters based on the sample likelihood function and the joint prior distribution, obtaining local estimation information, thereby effectively estimating the biological effects of the biomarkers, and providing a robust analysis framework for integrating biomarkers in multiple data centers.

[0159] It should be noted that the above examples are only used to understand the present application and do not constitute a limitation on the multi-center data aggregation method of the present application. More simple transformations based on this technical concept are all within the scope of protection of the present application.

[0160] This application also provides a multi-center data aggregation device, please refer to Figure 6 , Figure 6 This is a schematic diagram of the module structure of a multi-center data aggregation device in an embodiment of the present application. In this embodiment, the multi-center data aggregation device includes:

[0161] A data acquisition module 601 is used to acquire local measurement values ​​corresponding to biomarkers from multiple data centers;

[0162] A reference measurement module 602 is used to perform sampling measurement on the biomarker to obtain a reference measurement value;

[0163] The data estimation module 603 is used to perform Bayesian posterior estimation on the biological effect of the biomarker based on the reference measurement value and the local measurement value through a Bayesian algorithm to obtain a target estimation parameter; the target estimation parameter characterizes the effect of the biomarker on the disease outcome.

[0164] As an implementable method, in this embodiment, the data estimation module 603 is also used to obtain the research covariate corresponding to the biomarker;

[0165] The data estimation module 603 is further used to construct local estimation information corresponding to the biomarker based on the research covariate, the reference measurement value and the local measurement value;

[0166] The data estimation module 603 is further configured to perform Bayesian posterior estimation on the biomarker based on the local estimation information to obtain target estimation parameters.

[0167] As an implementable method, in this embodiment, the data estimation module 603 is further used to perform Bayesian sampling on the biomarker through a preset sampler and the local estimation information to obtain a biomarker posterior sample;

[0168] The data estimation module 603 is further used to perform statistical estimation on the biomarker posterior samples to obtain target estimation parameters.

[0169] As an implementable method, in this embodiment, the data estimation module 603 is further used to construct a sample likelihood function corresponding to the biomarker according to the research covariate, the reference measurement value and the local measurement value;

[0170] The data estimation module 603 is further used to construct a joint prior distribution corresponding to the local measurement value;

[0171] The data estimation module 603 is further configured to construct local estimation information corresponding to the biomarker based on the sample likelihood function and the joint prior distribution.

[0172] As an implementable method, in this embodiment, the data estimation module 603 is further used to construct an initial likelihood function corresponding to the biomarker based on the local measurement value, the research covariate and the reference measurement value;

[0173] The data estimation module 603 is further used to construct a sample likelihood function corresponding to the biomarker according to a preset independent hypothesis and the initial likelihood function.

[0174] As an implementable method, in this embodiment, the data estimation module 603 is also used to obtain a target regression model corresponding to the current parameter estimation;

[0175] The data estimation module 603 is further used to construct a sample likelihood function corresponding to the biomarker according to a preset independence hypothesis, the initial likelihood function and the target regression model.

[0176] As an implementable method, in this embodiment, the data estimation module 603 is also used to obtain the target outcome type corresponding to the current parameter estimation;

[0177] The data estimation module 603 is also used to determine a target regression model according to the target outcome type.

[0178] The multi-center data aggregation device provided by the present application adopts the multi-center data aggregation method in the above embodiment, which can solve the technical problem of multi-center data aggregation. Compared with the prior art, the beneficial effects of the multi-center data aggregation device provided by the present application are the same as the beneficial effects of the multi-center data aggregation method provided by the above embodiment, and other technical features in the multi-center data aggregation device are the same as the features disclosed in the above embodiment method, which will not be repeated here.

[0179] The present application provides a multi-center data aggregation device, which includes: at least one processor; and a memory communicatively connected to the at least one processor; wherein the memory stores instructions executable by the at least one processor, and the instructions are executed by the at least one processor so that the at least one processor can execute the multi-center data aggregation method in the above-mentioned embodiment one.

[0180] Reference below Figure 7 , which shows a schematic diagram of the structure of a multi-center data aggregation device suitable for implementing the embodiment of the present application. The multi-center data aggregation device in the embodiment of the present application may include but is not limited to mobile terminals such as mobile phones, laptops, digital broadcast receivers, PDAs (Personal Digital Assistants), PADs (Portable Application Descriptions), PMPs (Portable Media Players), vehicle-mounted terminals (such as vehicle-mounted navigation terminals), etc., and fixed terminals such as digital TVs, desktop computers, etc. Figure 7 The multi-center data aggregation device shown is only an example and should not bring any limitation to the functions and scope of use of the embodiments of the present application.

[0181] like Figure 7As shown, the multi-center data aggregation device may include a processing device 1001 (such as a central processing unit, a graphics processor, etc.), which can perform various appropriate actions and processes according to a program stored in a read-only memory (ROM: Read Only Memory) 1002 or a program loaded from a storage device 1003 to a random access memory (RAM: Random Access Memory) 1004. In RAM1004, various programs and data required for the operation of the multi-center data aggregation device are also stored. The processing device 1001, ROM1002 and RAM1004 are connected to each other through a bus 1005. An input / output (I / O) interface 1006 is also connected to the bus. Generally, the following systems can be connected to the I / O interface 1006: an input device 1007 including, for example, a touch screen, a touch pad, a keyboard, a mouse, an image sensor, a microphone, an accelerometer, a gyroscope, etc.; an output device 1008 including, for example, a liquid crystal display (LCD: Liquid Crystal Display), a speaker, a vibrator, etc.; a storage device 1003 including, for example, a magnetic tape, a hard disk, etc.; and a communication device 1009. The communication device 1009 can allow the multi-center data aggregation device to communicate with other devices wirelessly or wired to exchange data. Although the multi-center data aggregation device with various systems is shown in the figure, it should be understood that it is not required to implement or have all the systems shown. More or fewer systems can be implemented or have alternatively.

[0182] In particular, according to the embodiments disclosed in the present application, the process described above with reference to the flowchart can be implemented as a computer software program. For example, the embodiments disclosed in the present application include a multi-center data aggregation program product, which includes a multi-center data aggregation program carried on a computer-readable medium, and the multi-center data aggregation program contains program code for executing the method shown in the flowchart. In such an embodiment, the multi-center data aggregation program can be downloaded and installed from the network through a communication device, or installed from a storage device 1003, or installed from ROM1002. When the multi-center data aggregation program is executed by the processing device 1001, the above-mentioned functions defined in the method of the embodiment disclosed in the present application are executed.

[0183] The multi-center data aggregation device provided by the present application adopts the multi-center data aggregation method in the above embodiment, which can solve the technical problem of how to effectively aggregate biomarkers from different sources. Compared with the prior art, the beneficial effects of the multi-center data aggregation device provided by the present application are the same as the beneficial effects of the multi-center data aggregation method provided by the above embodiment, and the other technical features in the multi-center data aggregation device are the same as the features disclosed in the method of the previous embodiment, which will not be repeated here.

[0184] It should be understood that the various parts disclosed in this application can be implemented by hardware, software, firmware or a combination thereof. In the description of the above embodiments, specific features, structures, materials or characteristics can be combined in any one or more embodiments or examples in a suitable manner.

[0185] The above are only specific implementations of the present application, but the protection scope of the present application is not limited thereto. Any technician familiar with the technical field can easily think of changes or substitutions within the technical scope disclosed in the present application, which should be included in the protection scope of the present application. Therefore, the protection scope of the present application should be based on the protection scope of the claims.

[0186] The present application provides a storage medium having computer-readable program instructions (ie, a multi-center data aggregation program) stored thereon, wherein the computer-readable program instructions are used to execute the multi-center data aggregation method in the above-mentioned embodiment.

[0187] The storage medium provided in the present application may be, for example, a USB flash drive, but is not limited to electrical, magnetic, optical, electromagnetic, infrared, or semiconductor systems, systems or devices, or any combination of the above. More specific examples of storage media may include, but are not limited to: an electrical connection with one or more wires, a portable computer disk, a hard disk, a random access memory (RAM: RandomAccess Memory), a read-only memory (ROM: Read Only Memory), an erasable programmable read-only memory (EPROM: Erasable Programmable Read Only Memory or flash memory), an optical fiber, a portable compact disk read-only memory (CD-ROM: CD-Read Only Memory), an optical storage device, a magnetic storage device, or any suitable combination of the above. In this embodiment, the storage medium may be any tangible medium containing or storing a program, which may be used by or in combination with an instruction execution system, system or device. The program code contained on the storage medium may be transmitted using any appropriate medium, including but not limited to: wires, optical cables, RF (Radio Frequency: Radio Frequency), etc., or any suitable combination of the above.

[0188] The above storage medium may be included in the multi-center data aggregation device; or may exist independently without being assembled into the multi-center data aggregation device.

[0189] The storage medium carries one or more programs. When the one or more programs are executed by the multi-center data aggregation device, the multi-center data aggregation device can effectively aggregate biomarkers from different sources.

[0190] The multi-center data aggregation program code for performing the operations of the present application can be written in one or more programming languages ​​or a combination thereof, including object-oriented programming languages ​​such as Java, Smalltalk, C++, and conventional procedural programming languages ​​such as "C" language or similar programming languages. The program code can be executed entirely on the user's computer, partially on the user's computer, as a separate software package, partially on the user's computer and partially on a remote computer, or entirely on a remote computer or server. In the case of a remote computer, the remote computer can be connected to the user's computer through any type of network, including a local area network (LAN) or a wide area network (WAN), or can be connected to an external computer (for example, using an Internet service provider to connect through the Internet).

[0191] The flow chart and block diagram in the accompanying drawings illustrate the possible architecture, function and operation of the system, method and multi-center data aggregation program product according to various embodiments of the present application. In this regard, each box in the flow chart or block diagram can represent a module, a program segment or a part of a code, and the module, the program segment or a part of the code contains one or more executable instructions for realizing the specified logical function. It should also be noted that in some alternative implementations, the function marked in the box can also occur in a different order from that marked in the accompanying drawings. For example, two boxes represented in succession can actually be executed substantially in parallel, and they can sometimes be executed in the opposite order, depending on the functions involved. It should also be noted that each box in the block diagram and / or flow chart, and the combination of the boxes in the block diagram and / or flow chart can be implemented with a dedicated hardware-based system that performs a specified function or operation, or can be implemented with a combination of dedicated hardware and computer instructions.

[0192] The modules involved in the embodiments of the present application may be implemented by software or hardware, wherein the name of the module does not limit the unit itself in some cases.

[0193] The readable storage medium provided in this application is a storage medium, which stores computer-readable program instructions (i.e., a multi-center data aggregation program) for executing the above-mentioned multi-center data aggregation method, and can solve the technical problem of how to effectively aggregate biomarkers from different sources. Compared with the prior art, the beneficial effects of the storage medium provided in this application are the same as the beneficial effects of the multi-center data aggregation method provided in the above-mentioned embodiment, and will not be repeated here.

[0194] The above are only some embodiments of the present application, and are not intended to limit the patent scope of the present application. All equivalent structural changes made using the contents of the present application specification and drawings under the technical concept of the present application, or direct / indirect application in other related technical fields are included in the patent protection scope of the present application.

Claims

1. A multi-center data aggregation method, characterized in that: The multi-center data aggregation method comprises: Obtain local measurements corresponding to biomarkers from multiple data centers; Sampling and measuring the biomarkers to obtain reference measurement values; Obtaining research covariates corresponding to the biomarkers; constructing local estimation information corresponding to the biomarker according to the study covariate, the reference measurement value and the local measurement value; A Bayesian posterior estimation is performed on the biomarker based on the local estimation information to obtain a target estimation parameter; the target estimation parameter characterizes the effect of the biomarker on the disease outcome.

2. The multi-center data aggregation method according to claim 1, characterized in that: The step of performing Bayesian posterior estimation on the biomarker based on the local estimation information to obtain target estimation parameters comprises: Performing Bayesian sampling on the biomarker using a preset sampler and the local estimation information to obtain a biomarker posterior sample; Statistical estimation is performed on the biomarker posterior samples to obtain target estimation parameters.

3. The multi-center data aggregation method according to claim 2, characterized in that: The step of constructing the local estimation information corresponding to the biomarker according to the research covariate, the reference measurement value and the local measurement value comprises: constructing a sample likelihood function corresponding to the biomarker according to the research covariate, the reference measurement value and the local measurement value; Constructing a joint prior distribution corresponding to the local measurement values; Local estimation information corresponding to the biomarker is constructed based on the sample likelihood function and the joint prior distribution.

4. The multi-center data aggregation method according to claim 3, characterized in that: The step of constructing the sample likelihood function corresponding to the biomarker according to the research covariate, the reference measurement value and the local measurement value comprises: constructing an initial likelihood function corresponding to the biomarker based on the local measurement value, the study covariate and the reference measurement value; A sample likelihood function corresponding to the biomarker is constructed according to a preset independence hypothesis and the initial likelihood function.

5. The multi-center data aggregation method according to claim 4, characterized in that: The step of constructing the sample likelihood function corresponding to the biomarker according to the preset independent hypothesis and the initial likelihood function comprises: Get the target regression model corresponding to the current parameter estimate; A sample likelihood function corresponding to the biomarker is constructed according to the preset independence hypothesis, the initial likelihood function and the target regression model.

6. The multi-center data aggregation method according to claim 1, characterized in that: The step of obtaining the target regression model corresponding to the current parameter estimation includes: Get the target outcome type corresponding to the current parameter estimate; The target regression model is determined according to the target outcome type.

7. A multi-center data aggregation device, characterized in that: The multi-center data aggregation device comprises: A data acquisition module, used to obtain local measurements and study covariates corresponding to biomarkers from multiple data centers; A reference measurement module, used to perform sampling measurement on the biomarker to obtain a reference measurement value; A data estimation module is used to obtain the research covariate corresponding to the biomarker; construct the local estimation information corresponding to the biomarker according to the research covariate, the reference measurement value and the local measurement value; perform Bayesian posterior estimation on the biomarker based on the local estimation information to obtain the target estimation parameter; the target estimation parameter characterizes the effect of the biomarker on the disease outcome.

8. A multi-center data aggregation device, characterized in that: The device comprises: a memory, a processor, and a multi-center data aggregation program stored in the memory and executable on the processor, wherein the multi-center data aggregation program is configured to implement the steps of the multi-center data aggregation method according to any one of claims 1 to 6.

9. A storage medium, characterized in that: The storage medium stores a multi-center data aggregation program, which, when executed by a processor, implements the steps of the multi-center data aggregation method according to any one of claims 1 to 6.

Citation Information

Patent Citations

  • Schizophrenia classification method and system based on multi-center model

    CN113197578A

  • Data quality evaluation method, system and equipment for multi-center research

    CN115050479A