Multi-party mixed data provenance method and system based on latent group instrumental variables
By employing a multi-party hybrid data tracing method based on latent group instrumental variables, the problem of unreliable causal analysis results in medical record data was solved, enabling accurate tracing of medical record data and precise treatment plan recommendations.
Patent Information
- Application Number
- CN202210836782.7
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2022-07-15
- Publication Date
- 2026-01-06
- Estimated Expiration
- 2042-07-15
AI Technical Summary
The lack of unified data collection standards in medical records leads to heterogeneity in data sources from different medical institutions, resulting in unreliable causal analysis results. How to trace the types of diagnosis and treatment methods or medical institutions in medical record data has become an urgent problem to be solved.
A multi-party hybrid data source tracing method based on latent group instrumental variables is adopted. By using representation learning and expectation-maximization algorithms, medical record data are divided into multiple sample groups with different potential covariate-intervention relationships. The latent group instrumental variables are recovered and embedded into downstream prediction or recommendation tasks to identify differences in diagnosis and treatment methods among different medical institutions.
It enables accurate tracing of medical record data, identifies differences in diagnostic and treatment methods among different medical institutions, and provides precise treatment plan recommendations through joint learning, thereby improving the credibility of causal analysis and the accuracy of treatment plans.
Smart Images

Figure CN115188484B_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of causal inference, and more particularly to a method and system for tracing the source of medical records based on potential group instrumental variables. Background Technology
[0002] Causal inference is a powerful explanatory tool that plays a crucial role in decision-making across various fields, such as precision medicine, policy decisions, precision recommendations, and improvements to teaching strategies. While randomized controlled trials are considered the gold standard for identifying treatment / intervention effects or potential outcome functions, their time-consuming, costly nature, and ethical implications often limit causal analysis to observational data. However, the lack of standardized data collection methods means that data from multiple sources with varying causal relationships are frequently mixed together, posing additional challenges to causal analysis.
[0003] Considering the prevalent confounding biases and unobserved biases in causal data, instrumental variables are the most classic and reliable method for removing unobserved confounding effects from observed datasets. However, instrumental variable methods often rely on instrumental variables selected by human experts. Sometimes, due to limitations in human prior knowledge, these instrumental variables may not strictly fulfill their intended function in the algorithm, leading to unreliable estimates. One solution is to integrate instrumental variables from a large pool of available candidates using testing or screening methods. However, another problem arises: finding a suitable pool of instrumental variable candidates is not easy.
[0004] Medical record datasets contain patient information such as physical characteristics (weight, height, age, gender, occupation, and relevant examination results), verbal descriptions (patient descriptions of their symptoms and medical history), doctor-suggested treatment plans, and treatment outcomes obtained through follow-up visits. This data is often collected and compiled by different medical institutions. These institutions, due to differences in treatment philosophies, technical expertise, and medical equipment, may employ different treatment allocation mechanisms for the same type of disease; that is, treatment allocation mechanisms may exhibit heterogeneity across different institutions. Heterogeneous treatment allocation mechanisms represent different causal relationships between interventional and confounding variables from different data sources, with each mechanism corresponding to a specific treatment method. However, due to the lack of standardized medical record data collection practices, the medical institutions from which different medical records originate and their corresponding treatment methods are often missing from the records, resulting in a dataset containing multiple causal relationships. Therefore, tracing the origins of medical record samples within a dataset to identify the corresponding treatment methods or medical institutions is a pressing technical challenge. Summary of the Invention
[0005] The purpose of this invention is to overcome the current lack of standardized medical record data collection in the medical field, which leads to historically collected medical record data originating from different medical institutions and exhibiting differences in diagnostic and treatment methods, thus introducing additional biases into the estimation of the true potential outcomes. This invention proposes a multi-party mixed data tracing method and system based on latent group instrumental variables. It can divide data into multiple potential covariate-intervention relationship sample groups based on a heterogeneous treatment plan allocation mechanism, and use representation learning and expectation-maximization algorithms to recover the indicators of these subgroups from the mixed overall data. These subgroup indicators are then embedded as instrumental variables in downstream prediction or recommendation tasks for more accurate prediction of potential outcome functions.
[0006] The specific technical solution adopted in this invention is as follows:
[0007] In a first aspect, the present invention provides a multi-party mixed data source tracing method based on latent group instrumental variables, which includes the following steps:
[0008] S1. Obtain a medical record dataset from multiple medical institutions for source tracing and identification. Each medical record contains medical information, a treatment plan provided by the medical institution, and the treatment results after treatment according to the treatment plan. The medical information includes the patient's physical signs and oral information.
[0009] S2. Within the range of cluster hyperparameter values, select a candidate value for the number of clusters, use the treatment plan given by the medical institution as the intervention variable, use the corresponding disease information as the confusion variable, and map the disease information observed in the medical record dataset to a representation space through representation learning.
[0010] S3. Fix the expectation and covariance matrix in the representation space obtained in S2, and use the expectation maximization algorithm to identify the heterogeneous treatment plan allocation mechanism corresponding to the number of candidate values of the cluster. The heterogeneous treatment plan allocation mechanism represents the different causal relationships between intervention variables and confounding variables from different data sources. Each heterogeneous treatment plan allocation mechanism corresponds to a diagnosis and treatment method.
[0011] S4. Traverse all candidate values for the number of clusters within the range of the cluster hyperparameter. Perform S2 and S3 for each candidate value for the number of clusters to obtain the heterogeneous treatment plan allocation mechanism corresponding to each candidate value for the number of clusters. For each candidate value for the number of clusters, divide the samples in the medical record dataset into a sample subgroup with the same number of samples as the candidate value for the number of clusters. Then, based on the correlation independence index, select an optimal number of clusters and a latent group instrumental variable corresponding to each sample under the optimal number of clusters from all candidate values for the number of clusters. Use the latent group instrumental variable as an indicator variable for different sources in the multi-source mixed medical record data. Cluster all medical record data in the multi-source mixed medical record data to form multiple groups. Medical record data in each group that have the same diagnosis and treatment methods belong to the same heterogeneous treatment plan allocation mechanism, thereby tracing the differences in diagnosis and treatment methods of different medical institutions in the medical record dataset.
[0012] As a preferred embodiment of the first aspect above, in S1, each medical record includes patient physical signs such as weight, height, age, gender, occupation, and relevant examination results; the oral information comes from the patient's description of their symptoms and past medical history; and the treatment results come from follow-up results.
[0013] As a preferred embodiment of the first aspect above, S2 specifically includes the following sub-steps:
[0014] S201. For each candidate value K of the number of clusters within the range of cluster hyperparameter values, a representation learning algorithm is used to treat all observed patient signs and verbal information as confounding variables, projecting them onto independent representation spaces of each dimension through a mapping function, and learning non-independent data and complex multivariate interaction terms together as a noise term:
[0015]
[0016] Where X represents the patient's condition information used as a confounding variable in the medical record data, and T represents the treatment plan used as an intervention variable in the medical record data, ∈ TZError terms that represent unobserved patient signs and verbal information or errors caused by measurement errors; This represents the heterogeneous treatment allocation mechanism corresponding to K potential group instrumental variables Z∈{1,...,K}. Its inputs are the confusion variable X, K is the number of candidate clusters currently selected; z is the instantiation of the latent group instrumental variable Z; R is the final learned representation space, R j This represents the j-th component of the data representation space R, j∈{1,...,m} R}, m R For the total representation dimension, α zj These are the linear fitting coefficients corresponding to the characterization, β z It is a noise term learned jointly by non-independent data and complex multivariate interaction terms, 1 [Z=z] It is a conditional function, that is, the actual treatment plan allocation mechanism Z = z corresponding to sample data X and T is 1 when Z = z and 0 otherwise;
[0017] S202. Based on the representation space R finally learned in S201, calculate the expectation and covariance of the representation space:
[0018]
[0019] Where r i σ is the representation vector of the i-th sample, σ(R,R) is the covariance matrix, and n is the total number of samples in the medical record dataset.
[0020] S203. The formulas for calculating the likelihood function and log-likelihood function of complete data (where z comes from latent instrumental variable modeling) are defined as follows:
[0021]
[0022]
[0023] Where: t is the instantiation of the intervention variable T, r is the instantiation of the representation space R, t i ,r i ,z i Let t, r, and z be the values corresponding to the i-th sample, respectively. It is the joint probability distribution of t, r, z given the distribution parameter θ, π k It is t i ,r i Source: group z i The probability of =k Given the distribution parameter {μ k ,Σ k}Down t i ,r i The joint probability distribution, μk ,Σ k These are the mean and variance, respectively. It is a conditional function, i.e., z i =1 when k is equal to 0 otherwise, k∈{1,...,K}.
[0024] As a preferred embodiment of the first aspect above, S3 specifically includes the following sub-steps:
[0025] S301. Initialize heterogeneous data distribution with random numbers. Where K is the number of clusters selected in S2;
[0026] S302, Representation space information obtained using S202 Reinitialize the heterogeneous data distribution θ to θ (0) ={π (0) ,μ (0) ,Σ (0)}:
[0027]
[0028] in, These are the mean of T, the variance of T, and the covariance matrix of T and R, respectively, for randomly initialized T. yes transpose;
[0029] S303. Begin executing the desired step in the s-th iteration, i.e., estimate θ based on the given observed data {T,R} and the current heterogeneous data distribution. (s) The expected value of the log-likelihood function for the complete data is:
[0030]
[0031] Among them, expectations It is the i-th sample in the k-th group with respect to θ (s) Conditional probability distribution:
[0032]
[0033] in, Let be the probability that sample t, r originates from group z = i, and the sum of the conditional probabilities of the K groups is 1. Given distribution parameters The joint probability distribution of T and R;
[0034] S304. Continue with the maximization step in the s-th iteration, that is, estimate θ based on the given observed data {T,R} and the current heterogeneous data distribution. (s) Maximize the expected value of the log-likelihood function Q(θ,θ) of the complete data.(s) And update the heterogeneous data distribution estimate to θ. (s+1) :
[0035] θ (s+1) =argmax θ Q(θ,θ (s) )
[0036] Where θ (s+1) The parameters in the solution are obtained as follows:
[0037]
[0038]
[0039]
[0040] in This indicates concatenating T and R along the feature dimension. It is a matrix, M 2 =MM T ;
[0041] S305. In the expectation-maximization algorithm, the expectation step S304 and the maximization step S305 are iteratively executed until a distributed convergent solution corresponding to the current K value is obtained. From θ * Characterize the different causal relationships and their corresponding distributions among intervention variables and confounding variables from different data sources.
[0042] As a preferred embodiment of the first aspect above, step S4 specifically includes the following sub-steps:
[0043] S401. Traverse all candidate values for the number of clusters within the range of the cluster hyperparameter K, and execute S2 and S3 for each candidate K value to obtain a distributed convergent solution θ for the complete data. * ={π * ,μ * ,Σ * Based on the convergent solution of the distribution corresponding to each K value, reconstruct the latent group instrumental variable corresponding to each medical record data sample in the medical record dataset:
[0044]
[0045] Where the subscript i represents the parameter corresponding to the i-th sample, i = 1, 2, ..., n;
[0046] S402. For all cluster hyperparameter K values, use the correlation independence index MMD as the screening index, and select the cluster hyperparameter K that minimizes MMD as the optimal number of clusters.
[0047]
[0048] K * =argmin K MMD K (Z,R),K={1,2,…,10}
[0049] in, K represents the mean of the representation R corresponding to all samples in the k-th sample subgroup. * The optimal number of clusters;
[0050] S403, Select the optimal number of clusters K * The latent group instrumental variable z for each sample below i Z is the optimal potential group instrumental variable. * At the same time, with z i As different source indicator variables for the mixed data from multiple sources, the medical record data from different sources in the medical record dataset are clustered into multiple groups. The medical record data in each group have the same source indicator variable, which means that the medical record data in the same group used the same diagnosis and treatment methods, that is, they belong to the same heterogeneous treatment plan allocation mechanism, thereby realizing the traceability of the diagnosis and treatment method category in each medical record data.
[0051] As a preferred embodiment of the first aspect above, the representation learning algorithm employs variational autoencoder, principal component analysis, correlation minimization representation learning, or representation based on prior knowledge.
[0052] Secondly, the present invention provides a multi-party hybrid data tracing system based on latent group instrumental variables, comprising:
[0053] The dataset acquisition module is used to acquire medical record datasets from multiple medical institutions for subsequent source tracing and identification. Each medical record contains medical information, a treatment plan provided by the medical institution, and the treatment results after treatment according to the treatment plan. The medical information includes the patient's physical signs and oral information.
[0054] The representation module is used to select a candidate value for the number of clusters within the range of cluster hyperparameter values, use the treatment plan given by the medical institution as the intervention variable, use the corresponding disease information as the confusion variable, and map the disease information observed in the medical record dataset to a representation space through representation learning.
[0055] The expectation-maximization algorithm module is used to fix the expectation and covariance matrix in the representation space obtained in the representation module, and use the expectation-maximization algorithm to identify the heterogeneous treatment plan allocation mechanism corresponding to the candidate value of the number of clusters. The heterogeneous treatment plan allocation mechanism represents the different causal relationships between intervention variables and confounding variables from different data sources. Each heterogeneous treatment plan allocation mechanism corresponds to a diagnosis and treatment method.
[0056] The grouping and tracing module is used to traverse all candidate values for the number of clusters within the range of cluster hyperparameter values. For each candidate value, the representation module and the expectation-maximization algorithm module are executed to obtain the heterogeneous treatment plan allocation mechanism corresponding to each candidate value. For each candidate value, the samples in the medical record dataset are divided into sample subgroups with the same number of candidate values. Then, based on the correlation independence index, an optimal number of clusters and the latent group instrumental variable corresponding to each sample under the optimal number of clusters are selected from all candidate values. The latent group instrumental variable is used as an indicator variable for different sources in the multi-source mixed medical record data. All medical record data in the multi-source mixed medical record data are clustered to form multiple groups. Medical record data in each group that have the same diagnosis and treatment means belong to the same heterogeneous treatment plan allocation mechanism, thereby tracing the differences in diagnosis and treatment methods among different medical institutions in the medical record dataset.
[0057] Thirdly, the present invention provides a precision treatment plan recommendation system, which includes:
[0058] The joint learning module is used to obtain the latent group instrumental variables corresponding to each sample under the optimal number of clusters obtained by any of the multi-party mixed data source tracing and identification methods described in the first aspect above. The obtained latent group instrumental variables are embedded into the instrumental variable regression method and combined with multi-party knowledge for joint learning to obtain the counterfactual prediction function.
[0059] The treatment plan recommendation module is used to input the disease information of the target case as a confounding variable into the counterfactual prediction function to obtain the predicted treatment outcome value that the target case can achieve under each treatment method, which is used as a reference when selecting treatment methods.
[0060] As a preferred option in the third aspect mentioned above, each learning sample needs to be input with the medical record data corresponding to the patient's condition information, the treatment plan given by the medical institution, the optimal potential group instrumental variable, and the treatment result after treatment according to the treatment plan, so as to learn a counterfactual prediction function that can predict the treatment results under different treatment methods based on the patient's condition information.
[0061] As a preferred embodiment of the third aspect above, the instrumental variable regression method includes two-stage least squares regression, polynomial-based two-stage least squares regression, kernel-based two-stage least squares regression, deep learning algorithm-based least squares regression, and two-stage least squares regression based on adversarial moment conditions.
[0062] Traditional instrumental variable methods often rely on instrumental variables selected by human experts. However, these variables may not always fulfill their intended function within the algorithm due to limitations in prior human knowledge, leading to unreliable estimates. In contrast to existing technologies, this invention addresses the issue of identifying the source of treatment methods in large datasets shared by multiple medical institutions in a healthcare setting. To estimate the latent outcome function unaffected by confounding variables, it proposes a method for recovering latent group instrumental variables and inferring latent outcome functions based on heterogeneous treatment / intervention data. This method does not rely on human expert knowledge for instrumental variable specification or candidate set generation. This method identifies differences in treatment methods across different medical institutions, enabling clustering of medical record data according to these methods and facilitating source identification. Furthermore, based on the identified differences in treatment methods across different medical institutions, this invention can further integrate multi-party knowledge for joint learning, facilitating the recommendation of optimal treatment plans and supporting precise treatment for each patient. Attached Figure Description
[0063] Figure 1 This is a flowchart of a multi-source mixed data tracing method based on potential group instrumental variables.
[0064] Figure 2 This is a schematic diagram of a multi-party hybrid data tracing system based on potential group instrumental variables.
[0065] Figure 3 A schematic diagram of the modules of a precision treatment recommendation system.
[0066] Figure 4 This is a schematic diagram of heterogeneous treatment / intervention data in the embodiments.
[0067] Figure 5 This is a visualization of the accuracy of multi-source data tracing and identification in the embodiment and its corresponding results. Detailed Implementation
[0068] The present invention will be further described and illustrated below with reference to the accompanying drawings and specific embodiments.
[0069] like Figure 1As shown, in a preferred embodiment of the present invention, a multi-party hybrid data tracing method based on latent group instrumental variables is provided. This method targets a large dataset shared by multiple medical institutions in a healthcare setting. It traces and identifies the source of each medical record, revealing differences in treatment methods among different institutions. By combining multi-party knowledge through joint learning, it aims to assist in providing precise treatment for each patient. The multi-party hybrid data tracing method specifically includes the following steps:
[0070] S1. Obtain a medical record dataset from multiple medical institutions for source tracing and identification. Each medical record contains medical information, a treatment plan provided by the medical institution, and the treatment results after treatment according to the treatment plan. The medical information includes the patient's physical signs and oral information.
[0071] In this embodiment, each medical record includes patient vital signs such as weight, height, age, gender, occupation, and relevant examination results. The oral information comes from the patient's description of their symptoms and past medical history, and the treatment results come from follow-up results.
[0072] S2. Within the range of cluster hyperparameter values, select a candidate value for the number of clusters. Use the treatment plan provided by the medical institution as the intervention variable (also known as the treatment / intervention variable) and the corresponding condition information as the confusion variable. Through representation learning, map the condition information observed in the medical record dataset to a representation space.
[0073] In this embodiment, step S2 specifically includes the following sub-steps:
[0074] S201. For each candidate value K of the number of clusters within the range of cluster hyperparameter values, a representation learning algorithm is used to treat all observed patient signs and verbal information as confounding variables, projecting them onto independent representation spaces of each dimension through a mapping function, and learning non-independent data and complex multivariate interaction terms together as a noise term:
[0075]
[0076] Where X represents the patient's condition information used as a confounding variable in the medical record data, and T represents the treatment plan used as an intervention variable in the medical record data, ∈ TZ Error terms that represent unobserved patient signs and verbal information or errors caused by measurement errors; This represents the heterogeneous treatment allocation mechanism corresponding to K potential group instrumental variables Z∈{1,...,K}. Its inputs are the confusion variable X, K is the number of candidate clusters currently selected; z is an instantiation of the potential group instrument variable Z; h zj (X) is the polynomial representation learning function of variable j, ξ zj,1x represents the linear fitting coefficients of the corresponding polynomial; j This refers to the j-th dimension in X, where the superscript indicates the power; the specific maximum power needs to be optimized based on the fitting results. R is the final learned representation space. j This represents the j-th component of the data representation space R, j∈{1,...,m} R}, m R For the total representation dimension, α zj These are the linear fitting coefficients corresponding to the characterization; β z It is a noise term learned jointly by non-independent data and complex multivariate interaction terms, 1 [Z=z] It is a conditional function, that is, the actual treatment plan allocation mechanism Z = z corresponding to sample data X and T is 1 when Z = z and 0 otherwise.
[0077] In this embodiment, the above representation learning algorithm can be implemented in advance using a variational autoencoder, principal component analysis, correlation minimization representation learning, or representation based on prior knowledge.
[0078] S202. Based on the representation space R finally learned in S201, calculate the expectation and covariance of the representation space:
[0079]
[0080] Where r i σ is the representation vector of the i-th sample, σ(R,R) is the covariance matrix, and n is the total number of samples in the medical record dataset.
[0081] S203. The formulas for calculating the likelihood function and log-likelihood function of complete data (where z comes from latent instrumental variable modeling) are defined as follows:
[0082]
[0083]
[0084] Where: t is the instantiation of the intervention variable T, r is the instantiation of the representation space R, t i ,r i ,z i Let t, r, and z be the values corresponding to the i-th sample, respectively. It is the joint probability distribution of t, r, z given the distribution parameter θ, π k It is t i ,r i Source: group z i The probability of =k Given the distribution parameter {μ k ,Σ k}Down t i ,r iThe joint probability distribution, μ k ,Σ k These are the mean and variance, respectively. It is a conditional function, i.e., z i =1 when k is equal to 0 otherwise, k∈{1,...,K}.
[0085] It should be noted that in this invention, medical record data with source medical institution labels are considered complete data, while medical record data without source labels are considered incomplete data. The aforementioned medical record dataset is incomplete data and requires subsequent modeling to obtain its complete form. In S203, only the likelihood function and log-likelihood function calculation formulas for complete data are defined. However, since the medical record dataset is incomplete at this point, the likelihood function and log-likelihood function for complete data cannot be directly obtained and need to be solved subsequently.
[0086] S3. Fix the expectation and covariance matrix in the representation space obtained in S2, and use the expectation maximization algorithm to identify the heterogeneous treatment plan allocation mechanism corresponding to the number of candidate values of the cluster. The heterogeneous treatment plan allocation mechanism represents the different causal relationships between intervention variables and confounding variables from different data sources. Each heterogeneous treatment plan allocation mechanism corresponds to a diagnosis and treatment method.
[0087] It's important to note that the diagnostic and treatment methods used by medical institutions are closely related to their treatment philosophies, areas of expertise, and medical equipment. Different institutions using different methods may offer different treatment plans for the same condition. For example, for a given condition, institution A typically uses blood tests first and then determines the treatment plan based on the results; institution B typically uses a combination of observation, auscultation, inquiry, and palpation; and institution C typically uses ultrasound or radiology equipment for examination and then determines the treatment plan based on the results. These three diagnostic and treatment methods correspond to three heterogeneous treatment plan allocation mechanisms. That is, when faced with a treatment plan (confounding variable), a medical institution will provide a treatment plan based on its own diagnostic and treatment methods (intervention variable), and different causal relationships exist between the intervention variable and the confounding variable across different medical institutions.
[0088] In this embodiment, step S3 specifically includes the following sub-steps:
[0089] S301. Initialize heterogeneous data distribution with random numbers. Where K is the number of clusters selected in S2, which is a candidate value.
[0090] S302, Representation space information obtained using S202 Reinitialize the heterogeneous data distribution θ to θ (0) ={π (0) ,μ (0) ,Σ(0)}:
[0091]
[0092] in, These are the mean of T, the variance of T, and the covariance matrix of T and R, respectively, for randomly initialized T. yes The transpose of .
[0093] S303. Begin executing the desired step in the s-th iteration, i.e., estimate θ based on the given observed data {T,R} and the current heterogeneous data distribution. (s) The expected value of the log-likelihood function for the complete data is:
[0094]
[0095] Among them, expectations It is the i-th sample in the k-th group with respect to θ (s) Conditional probability distribution:
[0096]
[0097] in, Let be the probability that sample t, r originates from group z = i (where i = 1, 2, ..., K), and the sum of the conditional probabilities of the K groups is 1. Given distribution parameters The joint probability distribution of T and R.
[0098] S304. Continue with the maximization step in the s-th iteration, that is, estimate θ based on the given observed data {T,R} and the current heterogeneous data distribution. (s) Maximize the expected value of the log-likelihood function Q(θ,θ) of the complete data. (s) And update the heterogeneous data distribution estimate to θ. (s+1) :
[0099] θ (s+1) =argmax θ Q(θ,θ (s) )
[0100] Where θ (s+1) The parameters in the solution are obtained as follows:
[0101]
[0102]
[0103]
[0104] in This indicates concatenating T and R along the feature dimension. It is a matrix, M 2 =MM T .
[0105] S305. In the expectation-maximization algorithm, the expectation step S304 and the maximization step S305 are iteratively executed until a distributed convergent solution corresponding to the current K value is obtained. From θ * Characterize the different causal relationships and their corresponding distributions among intervention variables and confounding variables from different data sources.
[0106] S4. Traverse all candidate values for the number of clusters within the range of the cluster hyperparameter. Perform S2 and S3 for each candidate value for the number of clusters to obtain the heterogeneous treatment plan allocation mechanism corresponding to each candidate value for the number of clusters. For each candidate value for the number of clusters, divide the samples in the medical record dataset into a sample subgroup with the same number of samples as the candidate value for the number of clusters. Then, based on the correlation independence index, select an optimal number of clusters and a latent group instrumental variable corresponding to each sample under the optimal number of clusters from all candidate values for the number of clusters. Use the latent group instrumental variable as an indicator variable for different sources in the multi-source mixed medical record data. Cluster all medical record data in the multi-source mixed medical record data to form multiple groups. Medical record data in each group that have the same diagnosis and treatment methods belong to the same heterogeneous treatment plan allocation mechanism, thereby tracing the differences in diagnosis and treatment methods of different medical institutions in the medical record dataset.
[0107] In this embodiment, step S4 specifically includes the following sub-steps:
[0108] S401. Traverse all candidate values for the number of clusters within the range of the cluster hyperparameter K, and execute S2 and S3 for each candidate K value to obtain a distributed convergent solution θ for the complete data. * ={π * ,μ * ,Σ * Based on the convergent solution of the distribution corresponding to each K value, reconstruct the latent group instrumental variable corresponding to each medical record data sample in the medical record dataset:
[0109]
[0110] Where the subscript i represents the parameter corresponding to the i-th sample, i = 1, 2, ..., n;
[0111] S402. For all cluster hyperparameter K values, use the correlation independence index MMD as the screening index, and select the cluster hyperparameter K that minimizes MMD as the optimal number of clusters.
[0112]
[0113] K * =argmin K MMD K (Z,R),K={1,2,…,10}
[0114] in, K represents the mean of the representation R corresponding to all samples in the k-th sample subgroup. * The optimal number of clusters;
[0115] S403, Select the optimal number of clusters K * The latent group instrumental variable z for each sample below i Z is the optimal potential group instrumental variable. * At the same time, with z i As different source indicator variables for the mixed data from multiple sources, the medical record data from different sources in the medical record dataset are clustered into multiple groups. The medical record data in each group have the same source indicator variable, which means that the medical record data in the same group used the same diagnosis and treatment methods, that is, they belong to the same heterogeneous treatment plan allocation mechanism, thereby realizing the traceability of the diagnosis and treatment method category in each medical record data.
[0116] It should be noted that the source tracing methods S1 to S4 described above can be used to identify heterogeneous treatment allocation mechanisms (corresponding to the medical institutions' diagnostic and treatment methods) in a medical record dataset without source labels by reconstructing latent group instrumental variables. This involves identifying different causal relationships between interventional and confounding variables in the medical record data, thereby classifying the medical record data according to the category of diagnostic and treatment methods. Medical record data of the same category can be considered as a sample subgroup, where samples can be considered to have the same diagnostic and treatment methods, meaning their heterogeneous treatment allocation mechanisms are consistent. Therefore, medical record data can determine its own category of diagnostic and treatment methods, achieving source tracing.
[0117] Furthermore, if the heterogeneous treatment plan allocation mechanisms of different medical institutions within the medical record dataset are different, then the classification categories can be directly mapped to medical institutions, meaning that the medical record data samples in each sample subgroup all originate from the same medical institution. If some medical record data samples in a sample subgroup have medical institution source labels (complete data), while other medical record data samples do not have medical institution source labels (incomplete data), then the medical institution source labels of the complete data can be used to supplement the medical institution source labels of the incomplete data within a sample subgroup, thus achieving traceability of the institutional origin of the medical record data.
[0118] Similarly, based on the same inventive concept, such as Figure 2As shown, another preferred embodiment of the present invention also provides a multi-party mixed data tracing system based on latent group instrumental variables, corresponding to the multi-party mixed data tracing method based on latent group instrumental variables provided in the above embodiments, which includes:
[0119] The dataset acquisition module is used to acquire medical record datasets from multiple medical institutions for subsequent source tracing and identification. Each medical record contains medical information, a treatment plan provided by the medical institution, and the treatment results after treatment according to the treatment plan. The medical information includes the patient's physical signs and oral information.
[0120] The representation module is used to select a candidate value for the number of clusters within the range of cluster hyperparameter values, use the treatment plan given by the medical institution as the intervention variable, use the corresponding disease information as the confusion variable, and map the disease information observed in the medical record dataset to a representation space through representation learning.
[0121] The expectation-maximization algorithm module is used to fix the expectation and covariance matrix in the representation space obtained in the representation module, and use the expectation-maximization algorithm to identify the heterogeneous treatment plan allocation mechanism corresponding to the candidate value of the number of clusters. The heterogeneous treatment plan allocation mechanism represents the different causal relationships between intervention variables and confounding variables from different data sources. Each heterogeneous treatment plan allocation mechanism corresponds to a diagnosis and treatment method.
[0122] The grouping and tracing module is used to traverse all candidate values for the number of clusters within the range of cluster hyperparameter values. For each candidate value, the representation module and the expectation-maximization algorithm module are executed to obtain the heterogeneous treatment plan allocation mechanism corresponding to each candidate value. For each candidate value, the samples in the medical record dataset are divided into sample subgroups with the same number of candidate values. Then, based on the correlation independence index, an optimal number of clusters and the latent group instrumental variable corresponding to each sample under the optimal number of clusters are selected from all candidate values. The latent group instrumental variable is used as an indicator variable for different sources in the multi-source mixed medical record data. All medical record data in the multi-source mixed medical record data are clustered to form multiple groups. Medical record data in each group that have the same diagnosis and treatment means belong to the same heterogeneous treatment plan allocation mechanism, thereby tracing the differences in diagnosis and treatment methods among different medical institutions in the medical record dataset.
[0123] Since the principle of the above-mentioned multi-party mixed data tracing method based on latent group instrumental variables is similar to that of the multi-party mixed data tracing system based on latent group instrumental variables in the above-mentioned embodiment of the present invention, the specific implementation forms of each module of the system in this embodiment that are not fully described can also be referred to the specific implementation forms of the method shown in S1 to S4 above, and the repeated parts will not be described again.
[0124] In another embodiment of the present invention, based on the multi-party hybrid data tracing and identification method shown in S1 to S4 above, a more precise treatment plan recommendation can be achieved through step S5. The specific steps are as follows:
[0125] First, obtain the latent group instrumental variables corresponding to each sample under the optimal number of clusters obtained by the multi-party hybrid data tracing and identification method described in S1 to S4 of the aforementioned embodiments. Embed the obtained latent group instrumental variables into the instrumental variable regression method and combine them with multi-party knowledge for joint learning to obtain the counterfactual prediction function.
[0126] Then, the disease information of the target case is input into the counterfactual prediction function as a confounding variable to obtain the predicted treatment outcome value that the target case can achieve under each treatment method, which is used as a reference when selecting treatment methods.
[0127] Additionally, in another embodiment of the present invention, such as Figure 3 As shown, based on the same inventive concept as the aforementioned precision treatment plan recommendation, a precision treatment plan recommendation system is also provided, which includes:
[0128] The joint learning module is used to obtain the latent group instrumental variables corresponding to each sample under the optimal number of clusters obtained by the multi-party hybrid data source tracing and identification method described in the previous embodiment. The obtained latent group instrumental variables are embedded into the instrumental variable regression method and combined with multi-party knowledge for joint learning to obtain the counterfactual prediction function.
[0129] The treatment plan recommendation module is used to input the disease information of the target case as a confounding variable into the counterfactual prediction function to obtain the predicted treatment outcome value that the target case can achieve under each treatment method, which is used as a reference when selecting treatment methods.
[0130] In the aforementioned precision treatment recommendation method and system, when combining multi-party knowledge for joint learning, each learning sample needs to be input with the patient's condition information corresponding to the medical record data, the treatment plan given by the medical institution, the optimal potential group instrumental variable, and the treatment result after treatment according to the treatment plan, so as to learn a counterfactual prediction function that can predict the treatment results under different diagnostic and treatment methods based on the patient's condition information.
[0131] The instrumental variable regression methods used in the aforementioned precision treatment recommendation methods and systems include two-stage least squares regression, polynomial-based two-stage least squares regression, kernel-based two-stage least squares regression, deep learning algorithm-based least squares regression, and two-stage least squares regression based on adversarial moment conditions.
[0132] It should be noted that in the above-described precision treatment plan recommendation method and system of the present invention, the device only provides the predicted treatment outcome value that the target case can achieve under each treatment method, but the specific treatment method can be selected by the patient or doctor. This precision treatment plan recommendation method and system can be applied to the field of auxiliary medicine, as well as non-medical fields such as scientific research.
[0133] It should also be noted that in the systems of the above embodiments, each module is executed sequentially as a program module, thus essentially performing a data processing flow. Those skilled in the art will understand that, for ease of description and brevity, the specific working process of the systems described above can be referred to the corresponding processes in the foregoing method embodiments, and will not be repeated here. In the embodiments provided in this application, the division of steps or modules in the methods and systems is merely a logical functional division; in actual implementation, there may be other division methods. For example, multiple modules or steps may be combined or integrated together, and a module or step may also be split.
[0134] The present invention will now demonstrate the application effect of the multi-party mixed data tracing method and the precision treatment plan recommendation method based on potential group instrumental variables in the above embodiments on a specific dataset through a specific example, so as to facilitate understanding of the essence of the present invention.
[0135] Example
[0136] This embodiment uses the publicly available Infant Health and Development Program (IHDP) dataset and the PM2.5 concentration impact on cardiovascular mortality (PM-CMR) dataset. Figure 4 As shown, taking PM2.5 concentration in the PM-CMR dataset as an example, it demonstrates the existence of a latent cohort instrumental variable in heterogeneous treatment / intervention data, and the latent cohort instrumental variable in IHDP is similar.
[0137] The Infant Health and Development Program (IHDP) dataset contains 747 twin samples, each including 6 continuous variables and 19 discrete variables related to the infant's and mother's qualities before treatment / intervention. It aims to investigate the impact of an early special home visit teacher on the infant's future intellectual development. This example divides the dataset into training, validation, and test sets based on a 63% / 27% / 10% ratio. Similar to previous work, this example makes hypotheses about the potential outcome function and then generates the corresponding data for this example. Figure 4 A semi-synthetic dataset.
[0138] The PM2.5 concentration impact on cardiovascular mortality (PM-CMR) dataset contains data from 2132 cities, each including six continuous variables related to cardiovascular disease before treatment / intervention. It aims to study the impact of PM2.5 concentration on cardiovascular mortality. This example uses a 63% / 27% / 10% ratio to partition the dataset into training, validation, and test sets. Similar to previous work, this example makes assumptions about its potential outcome function and then generates the corresponding dataset for this example. Figure 4 A semi-synthetic dataset.
[0139] To objectively evaluate the performance of this algorithm, this example performs 10 random data shuffling and model retraining operations on both datasets above, and uses LatGIV... EM The mean and standard deviation (mean(std)) of the local latent outcome function fitting MSE error were calculated over 10 experiments using a regression method with nine different downstream instrumental variables embedded.
[0140] For two datasets, the accuracy of multi-source data tracing and identification and its corresponding visualization are as follows: Figure 5 As shown in the figure, LatGIV EM This is the source tracing method proposed in the foregoing embodiments of the present invention, LatGIV. KM This indicates that the method used in this invention directly uses K-Means clustering to obtain clusters as potential instrumental variables. It can be seen that the method used in this invention achieves an accuracy of about 80%, while directly using the K-Means algorithm for clustering can only achieve an accuracy of less than 60%, and for 5 data sources, it cannot even guarantee an accuracy of 30%.
[0141] Furthermore, based on the multi-party hybrid data tracing method provided in this embodiment of the invention, the recommendation accuracy of the precision treatment plan recommendation method was further tested. It obtains a counterfactual prediction function by combining the optimal group instrumental variables obtained from the tracing with multi-party knowledge through joint learning. This counterfactual prediction function is then used to predict the treatment outcome for different cases under each treatment method. Based on the predicted values, the optimal treatment plan is selected, providing a precision treatment plan recommendation for each patient. Specific results are shown in Table 1:
[0142] Table 1 LatGIV EM The mean (std) error of precision treatment recommendations provided on the IHDP and PM-CMR datasets.
[0143]
[0144] In the table, NoneIV means no instrumental variables are used, while the potential instrumental variables for UAS are from the following references [1], for WAS they are from the following references [2], for ModeIV they are from the following references [3], for AutoIV they are from the following references [4], and for LatGIV they are from the following references [4]. KM The potential instrumental variables come from the K-Means algorithm and LatGIV. EM The potential instrumental variables are derived from the precision treatment recommendation method proposed in the foregoing embodiments of this invention, where TrueIV refers to prior knowledge of the group instrumental variables.
[0145] In addition, along the horizontal axis of the table, Poly2SLS is the most classic two-stage instrumental variable regression method for predicting potential outcomes (i.e., the outcome of precision treatment recommendations), KernelIV's potential outcome estimation method comes from the following reference [5], DeepIV's potential outcome estimation method comes from the following reference [6], and DeepGMM's potential outcome estimation method comes from the following reference [7].
[0146] The specific references listed above are as follows:
[0147] [1].Neil M Davies,Stephanie von Hinke Kessler Scholder,HelmutFarbmacher,Stephen Burgess,Frank Windmeijer,and George Davey Smith.2015.Themany weak instruments problem and Mendelian randomization.Statistics inmedicine 34,3(2015),454–468.
[0148] [2].Stephen Burgess,Frank Dudbridge,and Simon GThompson.2016.Combining information on multiple instrumental variables inMendelian randomization: comparison of allele score and summarized datamethods.Statistics in medicine 35,11(2016),1880–1906.
[0149] [3].Jason S Hartford, Victor Veitch, Dhanya Sridhar, and Kevin Leyton-Brown. 2021. Valid causal inference with (some) invalid instruments. InInternational Conference on Machine Learning. PMLR, 4096–4106.
[0150] [4]. Junkun Yuan, Anpeng Wu, Kun Kuang, Bo Li, Runze Wu, Fei Wu, and LanfenLin. 2022. Auto IV: Counterfactual Prediction via Automatic InstrumentalVariable Decomposition. ACM Transactions on Knowledge Discovery from Data (TKDD) 16, 4 (2022), 1–20.
[0151] [5].Rahul Singh,Maneesh Sahani,and Arthur Gretton.2019.Kernelinstrumental variable regression.In NeurIPS 2019.4593–4605.
[0152] [6].Jason Hartford,Greg Lewis,Kevin Leyton-Brown,and MattTaddy.2017.DeepIV: A flexible approach for counterfactual prediction.In ICML2017.
[0153] [7].Andrew Bennett,Nathan Kallus,and Tobias Schnabel.2019.Deepgeneralized method of moments for instrumental variable analysis.In NeurIPS2019.
[0154] Therefore, it can be seen that the present invention has better recommendation accuracy compared with the estimation methods in the prior art.
[0155] The embodiments described above are merely two preferred embodiments of the present invention, and are not intended to limit the invention. Those skilled in the art can make various changes and modifications without departing from the spirit and scope of the invention. Therefore, all technical solutions obtained by equivalent substitution or equivalent transformation fall within the protection scope of the present invention.
Claims
1. A multi-party mixed data provenance method based on latent group instrumental variables, characterized in that, The method comprises the following steps: S1, obtaining medical record data sets from multiple medical institutions for traceability identification, wherein each medical record data comprises condition information, a treatment plan given by a medical institution, and a treatment result after treatment according to the treatment plan, and the condition information comprises patient signs and oral information; S2, selecting a cluster number candidate value in a cluster super parameter value range, taking the treatment plan given by the medical institution as an intervention variable, and taking the corresponding condition information as a confounding variable, and mapping the observed condition information in the medical record data set to a representation space through representation learning; S3, fixing the expectation and covariance matrix in the representation space obtained in S2, and identifying the heterogeneous treatment plan allocation mechanism corresponding to the cluster number candidate value by using an expectation maximization algorithm, wherein the heterogeneous treatment plan allocation mechanism represents different causal relationships between the intervention variables and the confounding variables from different data sources, and each heterogeneous treatment plan allocation mechanism corresponds to a diagnosis and treatment method; S4, traversing all cluster number candidate values in the cluster super parameter value range, and performing S2 and S3 for each cluster number candidate value to obtain a heterogeneous treatment plan allocation mechanism corresponding to each cluster number candidate value; for each cluster number candidate value, the samples in the medical record data set are divided into sample subgroups with the same number as the cluster number candidate value, and then a best cluster number and a latent group instrumental variable corresponding to each sample under the best cluster number are selected from all cluster number candidate values based on a correlation independence index; the latent group instrumental variable is used as a different source indication variable in the mixed medical record data of multiple parties, and all medical record data in the mixed medical record data of multiple parties are clustered and divided into multiple groups, and the medical record data in each group have the same diagnosis and treatment method, i.e., belong to the same heterogeneous treatment plan allocation mechanism, so that the differences in diagnosis and treatment methods of different medical institutions in the medical record data set are traced. S2 specifically comprises the following sub-steps: S201, for each cluster number candidate value K in the cluster super parameter value range, all observed patient signs and oral information are taken as confounding variables, are mapped to a representation space of each dimension by a mapping function, and a noise item is learned together with non-independent data and a multivariate complex interaction term: where X is the condition information in the medical record data as the confounder, T is the treatment regimen in the medical record data as the intervention variable, ∈ TZ represents the unobserved patient signs and oral information or error terms caused by measurement errors; represents the heterogeneous treatment assignment mechanism corresponding to the potential K potential latent group instrumental variable Z ∈ {1,..., K} The input is the confounder X, K is the selected class cluster number to be selected value; z is the instantiation of the latent group instrumental variable Z. R is the final learned representation space, R j represents the jth component of the data representation space R, j ∈ {1,..., m R}, m R is the total representation dimension, α zj is the linear fitting coefficient corresponding to the representation, β z is the noise term learned jointly by the dependent data and the multivariate complex interaction term, 1 [Z=z] is the conditional function, i.e., 1 when the corresponding true treatment scheme assignment mechanism Z = z between the sample data X and T is 1, otherwise 0; S202, based on the representation space R finally learned in S201, the expectation and covariance of the representation space are calculated: where r i is the representation vector of the i-th sample, σ(R, R) is the covariance matrix, and n is the total number of samples in the medical record dataset. S203, the calculation formulas of the likelihood function and the log-likelihood function of the complete data are defined as follows: Where: t is the instantiation of the intervention variable T, r is the instantiation of the representation space R, t i ,r i ,z i Let t, r, and z be the values corresponding to the i-th sample, respectively. It is the joint probability distribution of t, r, z given the distribution parameter θ, π k It is t i ,r i Source: group z i The probability of =k Given the distribution parameter {μ k ,Σ k }Down t i ,r i The joint probability distribution, μ k ,Σ k These are the mean and variance, respectively. It is a conditional function, i.e., z i =1 when k is equal to 0 otherwise, k∈{1,...,K}; S3 specifically comprises the following sub-steps: S301、use random number to initialize heterogeneous data distribution Wherein K is the selected class cluster number in the S2 selected value S302, using the representation space information obtained in S202 reinitializing the heterogeneous data distribution θ to θ (0) = {π (0) , μ (0) , Σ (0)} wherein are the mean of T, the variance of T, and the covariance matrix of T and R, respectively, initialized randomly, is the transpose of . S303、Start to execute the expected step in the s-th iteration, that is, estimate θ according to the given observation data {T, R} and the current heterogeneous data distribution estimation (s) The expected value of the log-likelihood function of the complete data is calculated as follows: where the expectation is the conditional probability distribution of the ith sample on the kth group with respect to θ (s) : wherein, is the probability that sample t, r originated from group z = i, and the sum of the conditional probabilities over K groups is 1, is the joint probability distribution of T and R given the distribution parameters T and R; S304, continue to execute the maximization step in the s-th iteration, i.e., estimate θ according to the given observation data {T, R} and the current estimate of the heterogeneous data distribution (s) , maximize the log-likelihood function of the complete data expectation Q(θ, θ (s) ) and update the estimate of the heterogeneous data distribution to θ (s+1) : θ (s+1) = argmax θ Q(θ,θ (s) ) where θ (s+1) The parameters in are solved as follows: wherein represents concatenating T and R in the feature dimension direction, is a matrix, M 2 = MM T ; S305. In the expectation-maximization algorithm, the expectation step S304 and the maximization step S305 are iteratively executed until a distributed convergent solution corresponding to the current K value is obtained. From θ * Characterize the different causal relationships and their corresponding distributions among intervention variables and confounding variables from different data sources.
2. The latent group instrument variable based multi-party mixed data provenance method of claim 1, wherein, In S1, the patient signs in each medical record data include weight, height, age, gender, work nature and related examination results, the oral information is derived from the patient's description of his / her own condition and past medical history, and the treatment result is derived from a follow-up result.
3. The latent group instrument variable based multi-party mixed data provenance method of claim 1, wherein, S4 specifically comprises the following sub-steps: S401. Traverse all candidate values for the number of clusters within the range of the cluster hyperparameter K, and execute S2 and S3 for each candidate K value to obtain a distributed convergent solution θ for the complete data. * ={π * ,μ * ,Σ * Based on the convergent solution of the distribution corresponding to each K value, reconstruct the latent group instrumental variable corresponding to each medical record data sample in the medical record dataset: Wherein, subscript i represents a parameter corresponding to the ith sample, i = 1, 2, …, n; S402, for all cluster parameter K values, use the correlation independent index MMD (maximum mean difference) as the screening index, select the cluster parameter K that makes MMD minimum as the best cluster number; K * = argmin K MMD K (Z, R), K = {1, 2, …, 10} wherein, represents the mean value of the representation R corresponding to all samples in the kth sample subgroup, K * is the optimal number of clusters; S403、selecting the optimal cluster number K * The potential group tool variable z corresponding to each sample i The optimal potential group tool variable Z * At the same time, z i As a different source indicating variable of multi-party mixed data, all medical record data of different sources in the medical record data set are clustered and divided into multiple groups, the medical record data in each group has the same source indicating variable, represents that the medical record data in the same group adopts the same diagnosis and treatment means, that is, belongs to the same kind of heterogeneous treatment scheme allocation mechanism, thereby realizing the traceability of the diagnosis and treatment means category in each medical record data.
4. The latent group instrument variable based multi-party mixed data provenance method of claim 1, wherein, The representation learning algorithm adopts a variational autoencoder, principal component analysis, correlation minimization representation learning, or prior knowledge-based representation.
5. A multi-party mixed data provenance system based on latent group instrumental variables, the system comprising: Comprise: A data set acquisition module for acquiring medical record data sets from multiple medical institutions for subsequent traceability identification, wherein each medical record data includes condition information, a treatment plan given by a medical institution, and a treatment result after treatment according to the treatment plan, and the condition information includes patient signs and spoken information; A representation module for selecting a cluster quantity candidate value within a cluster parameter value range, using the treatment plan given by the medical institution as the intervention variable and the corresponding condition information as the confounding variable, and mapping the observed condition information in the medical record data set to a representation space through representation learning; The specific steps are: S201, for each cluster quantity candidate value K in the cluster parameter value range, all observed patient signs and spoken information are taken as confounding variables, and are projected onto a representation space of independent dimensions through a mapping function by a representation learning algorithm, and a noise term is learned from non-independent data and multivariate complex interaction terms: where X is the condition information in the medical record data as the confounder, T is the treatment regimen in the medical record data as the intervention variable, ∈ TZ represents the unobserved patient signs and oral information or error terms caused by measurement errors; represents the heterogeneous treatment assignment mechanism corresponding to the potential K potential latent group instrumental variables Z ∈ {1,..., K} The input is the confounder X, and K is the selected number of class clusters to be selected. Z is an instantiation of the latent group instrumental variable Z; R is the final learned representation space, R j represents the jth component of the data representation space R, j ∈ {1,..., m R}, m R is the total representation dimension, a zj is the linear fitting coefficient corresponding to the representation, b z is the noise term learned jointly by the dependent data and the multivariate complex interaction term, 1 [Z=z] is the conditional function, i.e., 1 when the corresponding true treatment scheme assignment mechanism Z = z between the sample data X and T is 1, otherwise 0; S202, based on the representation space R finally learned in S201, the expectation and covariance of the representation space are calculated: where r i is the representation vector of the i-th sample, σ(R, R) is the covariance matrix, and n is the total number of samples in the medical record dataset. S203, the calculation formulas of the likelihood function and the log-likelihood function of the complete data are defined as follows: where: t is an instantiation of the intervention variable T, r is an instantiation of the representation space R, t i ,r i ,z i are t, r and z, respectively, corresponding to the i-th sample, is the joint probability distribution of t, r, z given the distribution parameters θ, π k is the probability that t i ,r i comes from the group z i =k, is the joint probability distribution of t, r given the distribution parameters {μ k ,Σ k}, μ i ,Σ i are the mean and variance, respectively, k k i is the conditional function, i.e., 1 if z TZ =k and 0 otherwise, k ∈ {1,..., K}. An expectation maximization algorithm module for fixing the expectation and covariance matrix in the representation space obtained by the representation module, identifying the heterogeneous treatment plan allocation mechanism corresponding to the cluster quantity candidate value using an expectation maximization algorithm, the heterogeneous treatment plan allocation mechanism representing different causal relationships between intervention variables and confounding variables from different data sources, each heterogeneous treatment plan allocation mechanism corresponding to a diagnosis and treatment method; The specific steps are: S201, for each cluster quantity candidate value K in the cluster parameter value range, all observed patient signs and spoken information are taken as confounding variables, and are projected onto a representation space of independent dimensions through a mapping function by a representation learning algorithm, and a noise term is learned from non-independent data and multivariate complex interaction terms: where X is the confounder information in the medical record data, T is the treatment regimen in the medical record data, ∈ TZ represents the unobserved patient signs and oral information or error terms caused by measurement errors; represents the heterogeneous treatment assignment mechanism corresponding to the potential K latent cluster instrumental variable Z ∈ {1,..., K} whose input is the confounder X, K is the number of selected class clusters to be selected value; z is the instantiation of the latent cluster instrumental variable Z; R is the final learned representation space, R j represents the jth component of the data representation space R, j ∈ {1,..., m R}, m R is the total representation dimension, α zj is the linear fitting coefficient corresponding to the representation, β z is the noise term learned by the non-independent data and the multivariate complex interaction term, 1 [Z=z] is the conditional function, that is, 1 when the corresponding true treatment assignment mechanism Z = z between the sample data X and T, otherwise 0; S202, based on the representation space R finally learned in S201, the expectation and covariance of the representation space are calculated: where r i is the representation vector of the i-th sample, σ(R, R) is the covariance matrix, and n is the total number of samples in the medical record dataset. S203, the calculation formulas of the likelihood function and the log-likelihood function of the complete data are defined as follows: where: t is an instantiation of the intervention variable T, r is an instantiation of the representation space R, t i ,r i ,z i are t, r and z, respectively, corresponding to the i-th sample, is the joint probability distribution of t, r, z given the distribution parameters θ, π k is the probability that t i ,r i originated from the group z i =k, is the joint probability distribution of t, r given the distribution parameters {μ k ,Σ k}, μ i ,r i ,Σ k ,Σ k are the mean and variance, respectively, is the conditional function, i.e., 1 if z i =k and 0 otherwise, k ∈ {1,..., K}. The packet tracing module is configured to traverse all cluster quantity candidate values in a cluster hyperparameter value range, execute the representation module and the expectation maximization algorithm module for each cluster quantity candidate value respectively, and obtain a heterogeneous treatment scheme allocation mechanism corresponding to each cluster quantity candidate value; for each cluster quantity candidate value, the sample in the medical record data set is divided into sample subgroups with the same number as the cluster quantity candidate value, and then a best cluster quantity and a latent group instrumental variable corresponding to each sample under the best cluster quantity are selected from all cluster quantity candidate values based on a correlation independence index; the latent group instrumental variable is used as a different source indicating variable in the multi-party mixed medical record data, all medical record data in the multi-party mixed medical record data is clustered and divided to form a plurality of groups, the medical record data in each group has the same diagnosis and treatment means, i.e., belongs to the same heterogeneous treatment scheme allocation mechanism, and thus the difference of diagnosis and treatment means of different medical institutions in the medical record data set is traced.
6. A precision therapy recommendation system characterized by, The joint learning module is configured to obtain the latent group instrumental variable corresponding to each sample under the best cluster quantity obtained by the multi-party mixed data tracing and identification method according to any one of claims 1 to 4, embed the obtained latent group instrumental variable into a tool variable regression method, and perform joint learning combined with multi-party knowledge to obtain an counterfactual prediction function; The treatment scheme recommendation module is configured to input the disease information of a target case as a confounding variable into the counterfactual prediction function to obtain a treatment result prediction value of the target case under each diagnosis and treatment means, which is used as a reference when selecting diagnosis and treatment means. When the joint learning combined with multi-party knowledge is performed, the disease information corresponding to the medical record data, the treatment scheme given by the medical institution, the best latent group instrumental variable, and the treatment result after treatment according to the treatment scheme are input into each learning sample, so that the counterfactual prediction function capable of predicting the treatment result under different diagnosis and treatment means based on the disease information is learned.
7. The precision therapy recommendation system of claim 6, wherein, The tool variable regression method includes a two-stage least square regression method, a two-stage least square regression method based on a polynomial, a two-stage least square regression method based on a kernel method, a least square regression method based on a deep learning algorithm, and a two-stage least square regression method based on a counterfactual condition.
8. The precision therapy recommendation system of claim 6, wherein,
Citation Information
Patent Citations
Method for disease diagnosis and treatment scheme based on generalized neural network clustering
CN104915560A
Similarity analysis method and system for patients suffering from cardio-cerebral vascular diseases
CN106778042A