Medical data enhancement method and device based on causal inference, storage medium and equipment
Generating high-quality medical data through causal inference and generative adversarial networks, the problem of scarcity of data in the diagnosis of rare diseases is solved, and the performance and robustness of the AI model are improved.
Patent Information
- Application Number
- CN202510448893.4
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-04-10
- Publication Date
- 2025-08-29
AI Technical Summary
The existing medical data augmentation technology has insufficient data learning ability and diagnostic accuracy in the diagnosis of rare diseases. Traditional methods lack medical regularity and the generated data quality is low.
By constructing a causal graph, causal modeling and effect estimation are carried out, and combined with the potential outcome framework and generative adversarial network, high-quality virtual cases and enhanced data are generated to ensure that the data complies with medical laws.
It improves the performance and generalization capabilities of the AI model, enhances the quality and robustness of the data, and optimizes the applicability of tasks in different scenarios.
Smart Images

Figure CN120565099A_ABST
Abstract
Description
Technical Field
[0001] The present application relates to the field of computer technology, and in particular to a medical data enhancement method, apparatus, storage medium, and device based on causal inference. Background Art
[0002] In the field of medical artificial intelligence, especially in scenarios such as rare disease modeling and clinical decision support, the performance of deep learning models is highly dependent on large-scale, high-quality medical data. For example, in the diagnosis of rare diseases, the small number of patients and extremely limited clinical samples lead to a scarcity of data for training AI models, which seriously affects the model's learning ability and diagnostic accuracy.
[0003] Among related technologies, medical data augmentation primarily relies on traditional methods such as image transformation, interpolation, and data synthesis. These methods increase data volume through simple geometric transformations (such as rotation, scaling, and cropping), interpolation algorithms (such as linear interpolation and spline interpolation), or rule-based data synthesis (such as random noise injection). Traditional image transformation and interpolation methods lack rationality and interpretability in the medical field and are prone to introducing noise or fabricated data that does not conform to medical laws, resulting in low-quality generated data and an inability to effectively improve the performance of AI models. Summary of the Invention
[0004] The embodiments of the present application provide a method, apparatus, storage medium, and device for medical data enhancement based on causal inference. To provide a basic understanding of some aspects of the disclosed embodiments, a brief summary is provided below. This summary is not intended to be a comprehensive review, identify key / important elements, or delineate the scope of protection for these embodiments. Its sole purpose is to present some concepts in a simplified form, serving as a prelude to the detailed description that follows.
[0005] In a first aspect, an embodiment of the present application provides a medical data enhancement method based on causal inference, the method comprising:
[0006] Acquire and preprocess the original medical data of existing actual cases to obtain the medical data to be enhanced;
[0007] Construct a causal graph based on the medical data to be augmented; the causal graph is used to represent the causal relationship between diseases, symptoms, and treatment plans;
[0008] Estimating the causal effect of the augmented medical data to obtain a weighted data set;
[0009] Based on the causal graph and the weighted data set, a potential outcome framework is used to perform counterfactual inference and obtain virtual cases.
[0010] Generate enhanced medical datasets based on virtual cases and preset generative adversarial networks.
[0011] Optionally, construct a causal graph based on the medical data to be enhanced, including:
[0012] Use structural equation modeling to model the causal relationship of augmented medical data and construct an original causal graph;
[0013] Access existing literature and expert knowledge;
[0014] According to existing literature and expert knowledge, the edges and directions of the original causal graph are optimized to obtain the optimized causal graph.
[0015] Optionally, causal effect estimation is performed on the enhanced medical data to obtain a weighted dataset, including:
[0016] The instrumental variable method is used to identify and estimate the impact of confounding factors on augmented medical data, and to obtain adjusted causal effect estimates;
[0017] Using the causal effect estimation results, propensity score matching is performed on the augmented medical data to simulate the clinical trial environment and obtain a matched data set;
[0018] Inverse probability weighting is applied to the matched dataset to eliminate sample selection bias and obtain a weighted dataset.
[0019] Optionally, based on the causal graph and the weighted dataset, a potential outcome framework is used to perform counterfactual inference to obtain hypothetical cases, including:
[0020] Based on the causal diagram and weighted data sets, construct the potential outcome data for each patient under different treatment options;
[0021] Using synthetic control methods, virtual cases are constructed by combining data on each patient's potential outcomes under different treatment options.
[0022] Optionally, generate an enhanced medical dataset based on the virtual case and the preset generative adversarial network, including:
[0023] Initialize the preset generative adversarial network, which includes a generator and a discriminator;
[0024] Using virtual cases to train the generator so that it generates simulated data samples based on causal relationships;
[0025] Using the original medical data and the simulated data samples to train a discriminator so that the discriminator can distinguish between the original medical data and the simulated data samples;
[0026] Through the Wasserstein GAN training method and preset medical prior knowledge, the generator and discriminator are alternately optimized to output a fake dataset close to the real data distribution;
[0027] The forged dataset is merged with the original medical data to obtain the enhanced medical dataset.
[0028] Optionally, original medical data from existing actual cases is obtained and preprocessed to obtain the medical data to be enhanced, including:
[0029] Obtain original medical data of existing actual cases;
[0030] Fill missing values, correct outliers, and standardize features in the original medical data to obtain cleaned and standardized data files;
[0031] The data in the cleaned standardized data file is used as the medical data to be enhanced.
[0032] Optionally, the method further includes:
[0033] Create a scenario model for the disease scenario to which the original medical data belongs;
[0034] Input the enhanced medical dataset into the scene model and output the model's loss value;
[0035] When the loss value reaches the minimum, a pre-trained scene model is generated.
[0036] In a second aspect, an embodiment of the present application provides a medical data enhancement device based on causal inference, the device comprising:
[0037] The original medical data acquisition module is used to acquire and pre-process the original medical data of existing actual cases to obtain the medical data to be enhanced;
[0038] A causal graph construction module is used to construct a causal graph based on the medical data to be enhanced; wherein the causal graph is used to represent the causal relationship between diseases, symptoms and treatment plans;
[0039] A causal effect estimation module is used to estimate the causal effect of the augmented medical data and obtain a weighted data set;
[0040] The counterfactual inference module is used to perform counterfactual inference based on the causal graph and the weighted data set using the potential outcome framework to obtain virtual cases;
[0041] The medical dataset generation module is used to generate an enhanced medical dataset based on virtual cases and a preset generative adversarial network.
[0042] In a third aspect, an embodiment of the present application provides a computer storage medium, which stores a plurality of instructions suitable for being loaded by a processor and executing the above-mentioned method steps.
[0043] In a fourth aspect, an embodiment of the present application provides a device, which may include: a processor and a memory; wherein the memory stores a computer program, and the computer program is suitable for being loaded by the processor and executing the above-mentioned method steps.
[0044] The technical solutions provided by the embodiments of the present application may have the following beneficial effects:
[0045] In an embodiment of the present application, on the one hand, a causal graph is constructed using the medical data to be enhanced. The causal graph is used to characterize the causal relationship between diseases, symptoms, and treatment plans. By modeling the causal relationship, the introduction of data that does not conform to medical laws can be avoided, so that the generated data is of high quality and the performance of the AI model can be effectively improved. On the other hand, by using a potential outcome framework for counterfactual inference, the potential outcome data of each patient under different treatment plans can be simulated, so that the results tend to be realistic. Training the model based on the results can improve the generalization ability of the AI model. On the other hand, by combining a preset generative adversarial network, high-quality medical data synthesis can be achieved, data distribution drift can be avoided, high-quality enhanced data can be generated, and tasks in different scenarios can be optimized, thereby improving the robustness and clinical applicability of the AI model.
[0046] It should be understood that the foregoing general description and the following detailed description are exemplary and explanatory only and are not restrictive of the present application. BRIEF DESCRIPTION OF THE DRAWINGS
[0047] The accompanying drawings, which are incorporated in and constitute a part of this specification, illustrate embodiments consistent with the present application and, together with the description, serve to explain the principles of the present application.
[0048] Figure 1 This is a flow chart of a medical data enhancement method based on causal inference provided in an embodiment of the present application;
[0049] Figure 2 This is a flow chart of a scenario model training method provided in an embodiment of the present application;
[0050] Figure 3 This is a schematic diagram of the structure of a medical data enhancement device based on causal inference provided by this application;
[0051] Figure 4 It is a structural diagram of a device provided in an embodiment of the present application. DETAILED DESCRIPTION
[0052] The following description and the drawings sufficiently illustrate specific embodiments of the application to enable those skilled in the art to practice them.
[0053] It should be clear that the embodiments described are only part of the embodiments of this application, not all of the embodiments. Based on the embodiments in this application, all other embodiments obtained by ordinary technicians in this field without making creative work are within the scope of protection of this application.
[0054] When the following description refers to the accompanying drawings, unless otherwise indicated, the same numbers in different drawings represent the same or similar elements. The embodiments described in the following exemplary embodiments do not represent all embodiments consistent with the present application. Instead, they are merely examples of devices and methods consistent with certain aspects of the present application, as detailed in the appended claims.
[0055] In the description of this application, it should be understood that the terms "first", "second", etc. are used for descriptive purposes only and should not be understood as indicating or implying relative importance. For those of ordinary skill in the art, the specific meanings of the above terms in this application can be understood according to specific circumstances. In addition, in the description of this application, unless otherwise specified, "multiple" refers to two or more. "And / or" describes the association relationship of associated objects, indicating that three relationships may exist. For example, A and / or B can represent: A exists alone, A and B exist at the same time, and B exists alone. The character " / " generally indicates that the previous and subsequent associated objects are in an "or" relationship.
[0056] The present application provides a medical data enhancement method, device, storage medium and equipment based on causal inference to solve the problems existing in the above-mentioned related technical problems. In an embodiment of the present application, on the one hand, a causal graph is constructed through the medical data to be enhanced. The causal graph is used to characterize the causal relationship between diseases, symptoms and treatment plans. Through causal relationship modeling, the introduction of data that does not conform to medical laws can be avoided, so that the generated data is of high quality and the performance of the AI model can be effectively improved. On the other hand, the potential result framework is used for counterfactual inference to simulate the potential result data of each patient under different treatment plans, so that the results tend to be realistic. The model is trained based on the result, which can improve the generalization ability of the AI model. On the other hand, by combining the preset generative adversarial network, high-quality medical data synthesis is achieved, data distribution drift is avoided, high-quality enhanced data can be generated, and tasks in different scenarios can be optimized, thereby improving the robustness and clinical applicability of the AI model. The following exemplary embodiments are used for detailed explanation.
[0057] The following will be combined with the Figure 1 -Attached Figure 2This paper details the causal inference-based medical data enhancement method provided in the embodiments of this application. This method can be implemented using a computer program and run on a von Neumann-based causal inference-based medical data enhancement device. This computer program can be integrated into an application or run as a standalone tool application.
[0058] See Figure 1 , provides a flow chart of a medical data enhancement method based on causal inference for an embodiment of the present application. Figure 1 As shown, the method of the embodiment of the present application may include the following steps:
[0059] S101, acquiring and preprocessing original medical data of existing actual cases to obtain medical data to be enhanced;
[0060] Among them, raw medical data refers to unprocessed patient data obtained from hospital information systems (HIS), electronic medical records (EMR), pictured accredited clinical imaging systems (PACS) or other medical data sources. These data may include the patient's personal information, medical history, symptoms, test results, treatment plans, imaging data, etc. For example, raw medical data includes the patient's age, gender, disease diagnosis, laboratory test results (such as blood routine, biochemical indicators), medical images (such as X-rays, CT, MRI), etc. Preprocessing refers to a series of data cleaning and conversion operations on raw medical data to make it meet the requirements of subsequent data enhancement and model training. Medical data to be enhanced refers to medical data that has been preprocessed. These data already have a certain quality and consistency, but there may still be problems such as insufficient data volume and uneven data distribution, and need to be further expanded and optimized through data enhancement technology.
[0061] In some embodiments of the present application, the specific process of obtaining and preprocessing the original medical data of existing actual cases to obtain the medical data to be enhanced includes: obtaining the original medical data of existing actual cases; filling missing values, correcting outliers, and standardizing features in the original medical data to obtain a cleaned standardized data file; and using the data in the cleaned standardized data file as the medical data to be enhanced.
[0062] In a possible implementation, raw medical data is obtained from a hospital database or API interface, the obtained raw medical data is cleaned and converted, and the cleaned and converted data is saved as medical data to be enhanced.
[0063] S102, constructing a causal graph based on the medical data to be enhanced; wherein the causal graph is used to represent the causal relationship between the disease, symptoms, and treatment options;
[0064] A causal graph is a directed acyclic graph (DAG) that represents the causal relationships between variables. In medicine, a causal graph can represent the causal logic between variables such as diseases, symptoms, and treatment plans.
[0065] In some embodiments of the present application, the specific process of constructing a causal graph based on the medical data to be enhanced includes: using the structural equation modeling method to model the causal relationship of the medical data to be enhanced to construct an original causal graph; obtaining existing literature and expert knowledge; and optimizing the edges and directions of the original causal graph based on the existing literature and expert knowledge to obtain an optimized causal graph.
[0066] Among them, Structural Equation Modeling (SEM) is a multivariate statistical analysis technique used to analyze the causal relationship between variables. It combines factor analysis and path analysis, can handle multiple dependent variables and multiple independent variables at the same time, and can take into account measurement errors. For example, in the medical field, SEM can be used to analyze complex causal relationships between variables such as diseases, symptoms, and treatment plans. The original causal graph is a causal relationship graph constructed based on data and preliminary analysis, which represents the preliminary causal relationship between variables. For example: a simple causal graph may include three nodes: disease (D), symptoms (S), and treatment plan (T), as well as directed edges between them.
[0067] In an embodiment of the present application, a causal graph is constructed using the medical data to be enhanced. The causal graph is used to characterize the causal relationship between diseases, symptoms, and treatment plans. By modeling the causal relationship, the introduction of data that does not conform to medical laws can be avoided, so that the generated data is of high quality and can effectively improve the performance of the AI model.
[0068] S103, estimating the causal effect of the medical data to be enhanced to obtain a weighted data set;
[0069] Causal effect estimation uses statistical methods to estimate the causal impact of one variable (e.g., treatment regimen) on another variable (e.g., disease outcome). Common methods include instrumental variable (IV) methods, propensity score matching (PSM), and inverse probability weighting (IPW).
[0070] In some embodiments of the present application, the specific process of estimating the causal effect of the medical data to be enhanced and obtaining a weighted data set includes: using the instrumental variable method to identify and estimate the impact of confounding factors on the medical data to be enhanced, and obtaining an adjusted causal effect estimation result; using the causal effect estimation result to perform propensity score matching on the medical data to be enhanced to simulate the clinical trial environment, and obtain a matched data set; applying inverse probability weighting to the matched data set to eliminate sample selection bias, and obtain a weighted data set.
[0071] S104, based on the causal graph and the weighted data set, a potential outcome framework is used to perform counterfactual inference and obtain a virtual case;
[0072] In some embodiments of the present application, based on the causal graph and the weighted data set, a potential outcome framework is used to perform counterfactual inference, and the specific process of obtaining a virtual case includes: constructing the potential outcome data of each patient under different treatment plans based on the causal graph and the weighted data set; using the synthetic control method to combine the potential outcome data of each patient under different treatment plans to construct a virtual case.
[0073] Among them, the potential outcome framework is a causal inference method used to estimate the potential outcomes of an individual under different treatment conditions. Y(1) represents the outcome when the individual receives the treatment, and Y(0) represents the outcome when the individual does not receive the treatment. Counterfactual inference is to estimate the outcomes of individuals under treatment conditions that they have not actually experienced through the potential outcome framework. Counterfactual inference can help understand the potential impact of different treatment options on individuals. The synthetic control method is a method for constructing virtual cases, which simulates the outcomes of a virtual individual under different treatment conditions by combining data from multiple similar individuals.
[0074] In some embodiments of the present application, a causal graph is constructed based on medical knowledge and data to represent the causal relationship between variables such as diseases, symptoms, and treatment plans. For example, diseases cause symptoms, and treatment plans affect diseases and symptoms. The original data set is weighted using the inverse probability weighting (IPW) method to reduce sample selection bias. The weight of each sample is the inverse of its probability of being selected. For each patient, based on the causal graph and the weighted data set, its potential outcomes under different treatment plans are estimated. For example, for a patient who actually received treatment, the possible outcome Y(0) when he did not receive treatment is estimated; for a patient who did not receive treatment, the possible outcome Y(1) when he received treatment is estimated. Using the synthetic control method, the potential outcome data of each patient under different treatment plans are combined to construct a virtual case.
[0075] The specific steps are as follows:
[0076] Select similar patients: For each target patient, a group of patients with similar characteristics are selected from the weighted dataset as a control group.
[0077] Constructing a virtual case: By combining the data from the control group with data from other patients using methods such as weighted averaging or interpolation, we construct a virtual case of the target patient before treatment. For example, if the target patient actually received treatment, we can construct a virtual case of their pre-treatment condition, including information such as hypothetical symptoms and disease progression.
[0078] In an embodiment of the present application, the potential outcome framework is used for counterfactual inference, which can simulate the potential outcome data of each patient under different treatment plans, so that the results tend to be realistic. Training the model based on the results can improve the generalization ability of the AI model.
[0079] S105: Generate an enhanced medical data set based on the virtual case and the preset generative adversarial network.
[0080] In some embodiments of the present application, the specific process of generating an enhanced medical data set based on virtual cases and a preset generative adversarial network includes: initializing the preset generative adversarial network, which includes a generator and a discriminator; using virtual cases to train the generator so that the generator generates simulated data samples based on causal relationships; using original medical data and simulated data samples to train the discriminator so that the discriminator distinguishes between original medical data and simulated data samples; alternately optimizing the generator and the discriminator through the Wasserstein GAN training method and preset medical prior knowledge to output a forged data set that is close to the real data distribution; merging the forged data set with the original medical data to obtain an enhanced medical data set.
[0081] The pre-existing generative adversarial network (GAN) is a deep learning model consisting of a generator and a discriminator. The goal of the generator is to generate fake data that is as close to real data as possible, while the goal of the discriminator is to distinguish between real and fake data. The Wasserstein GAN (WGAN) training method is an improved GAN training method that uses the Wasserstein distance (also known as the Earth Mover distance) to optimize the generator and discriminator, making the generated data closer to the real data distribution.
[0082] For example, a neural network architecture for a generator and a discriminator is constructed. The generator is responsible for generating simulated data samples based on causal relationships, while the discriminator is responsible for distinguishing between real data and simulated data. The generator is trained using virtual cases previously generated through counterfactual inference and synthetic control methods. These virtual cases provide additional training samples, helping the generator learn how to generate data that conforms to medical laws. The discriminator is trained by mixing original medical data with simulated data samples generated by the generator. The discriminator's goal is to accurately distinguish between real and fake data. Using the Wasserstein GAN (WGAN) training method, the generator and discriminator are optimized by minimizing the Wasserstein distance between generated and real data. At the same time, medical prior knowledge is incorporated to ensure that the generated data conforms to medical laws. For example, if a disease is known to be associated with specific symptoms, the data generated by the generator should reflect this law. After multiple iterations of training, the generator is able to generate high-quality fake data samples that are close in distribution to real data. The generated fake data set is merged with the original medical data to form an enhanced medical data set. This merged data set, which contains both real and fake data, is used to train more robust medical diagnosis models.
[0083] Furthermore, after obtaining the enhanced medical dataset, the enhanced medical dataset can be used for model training. The training process includes: creating a scenario model under the disease scenario to which the original medical data belongs; inputting the enhanced medical dataset into the scenario model and outputting the loss value of the model; when the loss value reaches the minimum, generating a pre-trained scenario model.
[0084] In other embodiments, when the loss value has not reached the minimum, the step of inputting the enhanced medical dataset into the scene model is continued until the loss value of the model reaches the minimum.
[0085] The scene model may be created by using a decision tree neural network or other neural networks.
[0086] Furthermore, after obtaining the pre-trained scenario model, high-quality enhanced data can be generated to optimize tasks such as medical image analysis, epidemic prediction, and rare disease modeling, thereby improving the robustness and clinical applicability of the AI model.
[0087] In one application scenario, pre-trained scenario models can be used to address the scarcity of rare disease data and improve the AI model's learning ability with few samples. Causal and counterfactual inferences can be used to generate virtual cases, improving the accuracy of rare disease diagnoses.
[0088] In another application scenario, causal inference is combined with GANs to improve the quality and diversity of epidemic datasets and optimize disease transmission models. Through counterfactual inference, the impact of different interventions (such as vaccines and quarantine measures) on the spread of the epidemic can be simulated, improving the scientific nature of public health decision-making.
[0089] In another application scenario, causal GANs are used to generate medical imaging data to improve the robustness of medical imaging AI models. Combined with causal inference, the synthesis of lesion areas is optimized, improving the AI model's ability to recognize medical images.
[0090] In another application scenario, causal inference can be used to estimate individual treatment effects, improving personalized medical decision-making capabilities. Combined with reinforcement learning, treatment plans can be optimized to improve patient prognosis.
[0091] In the examples of this application, causal structure modeling, counterfactual inference, synthetic control methods, causal GANs, and other technologies are used to generate high-quality synthetic data in the absence of medical data, thereby improving the learning ability of AI models. This method is applicable to rare disease modeling, epidemic prediction, medical image enhancement, personalized treatment optimization, and other fields, providing an innovative data enhancement strategy for intelligent healthcare.
[0092] In an embodiment of the present application, on the one hand, a causal graph is constructed using the medical data to be enhanced. The causal graph is used to characterize the causal relationship between diseases, symptoms, and treatment plans. By modeling the causal relationship, the introduction of data that does not conform to medical laws can be avoided, so that the generated data is of high quality and the performance of the AI model can be effectively improved. On the other hand, by using a potential outcome framework for counterfactual inference, the potential outcome data of each patient under different treatment plans can be simulated, so that the results tend to be realistic. Training the model based on the results can improve the generalization ability of the AI model. On the other hand, by combining a preset generative adversarial network, high-quality medical data synthesis can be achieved, data distribution drift can be avoided, high-quality enhanced data can be generated, and tasks in different scenarios can be optimized, thereby improving the robustness and clinical applicability of the AI model.
[0093] See Figure 2 , provides a flow chart of a scene model training method for an embodiment of the present application.
[0094] like Figure 2 As shown, the method of the embodiment of the present application may include the following steps:
[0095] S201, creating a scenario model for the disease scenario to which the original medical data belongs;
[0096] S202, inputting the enhanced medical data set into the scenario model and outputting the loss value of the model;
[0097] S203: When the loss value reaches the minimum, generate a pre-trained scene model.
[0098] In an embodiment of the present application, on the one hand, a causal graph is constructed using the medical data to be enhanced. The causal graph is used to characterize the causal relationship between diseases, symptoms, and treatment plans. By modeling the causal relationship, the introduction of data that does not conform to medical laws can be avoided, so that the generated data is of high quality and the performance of the AI model can be effectively improved. On the other hand, by using a potential outcome framework for counterfactual inference, the potential outcome data of each patient under different treatment plans can be simulated, so that the results tend to be realistic. Training the model based on the results can improve the generalization ability of the AI model. On the other hand, by combining a preset generative adversarial network, high-quality medical data synthesis can be achieved, data distribution drift can be avoided, high-quality enhanced data can be generated, and tasks in different scenarios can be optimized, thereby improving the robustness and clinical applicability of the AI model.
[0099] The following are device embodiments of the present application, which can be used to implement the method embodiments of the present application. For details not disclosed in the device embodiments of the present application, please refer to the method embodiments of the present application.
[0100] See Figure 3 , which shows a schematic diagram of the structure of a medical data enhancement device based on causal inference, provided by an exemplary embodiment of the present application. This medical data enhancement device based on causal inference can be implemented as all or part of a device through software, hardware, or a combination of both. The device 1 includes a raw medical data acquisition module 10, a causal graph construction module 20, a causal effect estimation module 30, a counterfactual inference module 40, and a medical dataset generation module 50.
[0101] The original medical data acquisition module 10 is used to acquire and pre-process the original medical data of existing actual cases to obtain the medical data to be enhanced;
[0102] A causal graph construction module 20 is used to construct a causal graph based on the medical data to be enhanced; wherein the causal graph is used to represent the causal relationship between diseases, symptoms and treatment plans;
[0103] A causal effect estimation module 30 is used to perform causal effect estimation on the medical data to be enhanced to obtain a weighted data set;
[0104] A counterfactual inference module 40 is configured to perform counterfactual inference based on the causal graph and the weighted data set using a potential outcome framework to obtain a virtual case;
[0105] The medical data set generation module 50 is used to generate an enhanced medical data set based on the virtual case and the preset generative adversarial network.
[0106] It should be noted that the causal inference-based medical data enhancement device provided in the above embodiment only uses the division of the above-mentioned functional modules as an example when executing the causal inference-based medical data enhancement method. In actual applications, the above-mentioned functions can be assigned to different functional modules as needed, that is, the internal structure of the device can be divided into different functional modules to complete all or part of the functions described above. In addition, the causal inference-based medical data enhancement device provided in the above embodiment and the causal inference-based medical data enhancement method embodiment are based on the same concept. The implementation process is detailed in the method embodiment and will not be repeated here.
[0107] The serial numbers of the above embodiments of the present application are for description only and do not represent the advantages or disadvantages of the embodiments.
[0108] In an embodiment of the present application, on the one hand, a causal graph is constructed using the medical data to be enhanced. The causal graph is used to characterize the causal relationship between diseases, symptoms, and treatment plans. By modeling the causal relationship, the introduction of data that does not conform to medical laws can be avoided, so that the generated data is of high quality and the performance of the AI model can be effectively improved. On the other hand, by using a potential outcome framework for counterfactual inference, the potential outcome data of each patient under different treatment plans can be simulated, so that the results tend to be realistic. Training the model based on the results can improve the generalization ability of the AI model. On the other hand, by combining a preset generative adversarial network, high-quality medical data synthesis can be achieved, data distribution drift can be avoided, high-quality enhanced data can be generated, and tasks in different scenarios can be optimized, thereby improving the robustness and clinical applicability of the AI model.
[0109] The present application also provides a computer-readable medium having program instructions stored thereon, which, when executed by a processor, implement the medical data enhancement method based on causal inference provided by the above-mentioned various method embodiments.
[0110] The present application also provides a computer program product comprising instructions, which, when executed on a computer, enables the computer to execute the medical data enhancement method based on causal inference of each of the above method embodiments.
[0111] See Figure 4 , which is a schematic diagram of the structure of a device provided in the embodiment of the present application. Figure 4 As shown, the device 1000 may include: at least one processor 1001 , at least one network interface 1004 , a user interface 1003 , a memory 1005 , and at least one communication bus 1002 .
[0112] The communication bus 1002 is used to implement the connection and communication between these components.
[0113] The user interface 1003 may include a display screen (Display) and a camera (Camera). Optionally, the user interface 1003 may also include a standard wired interface and a wireless interface.
[0114] The network interface 1004 may optionally include a standard wired interface or a wireless interface (such as a WI-FI interface).
[0115] The processor 1001 may include one or more processing cores. The processor 1001 utilizes various interfaces and circuits to connect the various components within the entire device 1000. It executes instructions, programs, code sets, or instruction sets stored in the memory 1005, and accesses data stored in the memory 1005 to perform various functions and process data within the device 1000. Optionally, the processor 1001 may be implemented in the form of at least one hardware component: a digital signal processing (DSP), a field-programmable gate array (FPGA), or a programmable logic array (PLA). The processor 1001 may integrate one or a combination of a central processing unit (CPU), a graphics processing unit (GPU), and a modem. The CPU primarily handles operating devices, user interfaces, and applications; the GPU is responsible for rendering and drawing content displayed on the display; and the modem handles wireless communications. It is understood that the modem may not be integrated into the processor 1001 and may be implemented as a separate chip.
[0116] Among them, the memory 1005 may include a random access memory (RAM) or a read-only memory (Read-Only Memory). Optionally, the memory 1005 includes a non-transitory computer-readable storage medium. The memory 1005 can be used to store instructions, programs, codes, code sets or instruction sets. The memory 1005 may include a program storage area and a data storage area, wherein the program storage area may store instructions for implementing the operating device, instructions for at least one function (such as a touch function, a sound playback function, an image playback function, etc.), instructions for implementing the above-mentioned various method embodiments, etc.; the data storage area may store data involved in the above-mentioned various method embodiments, etc. The memory 1005 may also be optionally at least one storage device located away from the aforementioned processor 1001. As Figure 4 As shown, the memory 1005 as a computer storage medium may include an operating device, a network communication module, a user interface module, and a medical data enhancement application based on causal inference.
[0117] exist Figure 4 In the device 1000 shown, the user interface 1003 is mainly used to provide an input interface for the user and obtain user input data; and the processor 1001 can be used to call the medical data enhancement application based on causal inference stored in the memory 1005 and specifically perform the following operations:
[0118] Acquire and preprocess the original medical data of existing actual cases to obtain the medical data to be enhanced;
[0119] Construct a causal graph based on the medical data to be augmented; the causal graph is used to represent the causal relationship between diseases, symptoms, and treatment plans;
[0120] Estimating the causal effect of the augmented medical data to obtain a weighted data set;
[0121] Based on the causal graph and the weighted data set, a potential outcome framework is used to perform counterfactual inference and obtain virtual cases.
[0122] Generate enhanced medical datasets based on virtual cases and preset generative adversarial networks.
[0123] In one embodiment, when constructing a causal graph based on the medical data to be enhanced, the processor 1001 specifically performs the following operations:
[0124] Use structural equation modeling to model the causal relationship of augmented medical data and construct an original causal graph;
[0125] Access existing literature and expert knowledge;
[0126] According to existing literature and expert knowledge, the edges and directions of the original causal graph are optimized to obtain the optimized causal graph.
[0127] In one embodiment, when the processor 1001 performs causal effect estimation on the medical data to be enhanced to obtain a weighted data set, the processor 1001 specifically performs the following operations:
[0128] The instrumental variable method is used to identify and estimate the impact of confounding factors on augmented medical data, and to obtain adjusted causal effect estimates;
[0129] Using the causal effect estimation results, propensity score matching is performed on the augmented medical data to simulate the clinical trial environment and obtain a matched data set;
[0130] Inverse probability weighting is applied to the matched dataset to eliminate sample selection bias and obtain a weighted dataset.
[0131] In one embodiment, when processor 1001 performs counterfactual inference based on the causal graph and the weighted data set using the potential outcome framework to obtain a virtual case, the following operations are specifically performed:
[0132] Based on the causal diagram and weighted data sets, construct the potential outcome data for each patient under different treatment options;
[0133] Using synthetic control methods, virtual cases are constructed by combining data on each patient's potential outcomes under different treatment options.
[0134] In one embodiment, when the processor 1001 generates an enhanced medical dataset based on a virtual case and a preset generative adversarial network, the processor 1001 specifically performs the following operations:
[0135] Initialize the preset generative adversarial network, which includes a generator and a discriminator;
[0136] Using virtual cases to train the generator so that it generates simulated data samples based on causal relationships;
[0137] Using the original medical data and the simulated data samples to train a discriminator so that the discriminator can distinguish between the original medical data and the simulated data samples;
[0138] Through the Wasserstein GAN training method and preset medical prior knowledge, the generator and discriminator are alternately optimized to output a fake dataset close to the real data distribution;
[0139] The forged dataset is merged with the original medical data to obtain the enhanced medical dataset.
[0140] In one embodiment, when the processor 1001 acquires and pre-processes the original medical data of an existing actual case to obtain the medical data to be enhanced, the processor 1001 specifically performs the following operations:
[0141] Obtain original medical data of existing actual cases;
[0142] Fill missing values, correct outliers, and standardize features in the original medical data to obtain cleaned and standardized data files;
[0143] The data in the cleaned standardized data file is used as the medical data to be enhanced.
[0144] In one embodiment, the processor 1001 further performs the following operations:
[0145] Create a scenario model for the disease scenario to which the original medical data belongs;
[0146] Input the enhanced medical dataset into the scene model and output the model's loss value;
[0147] When the loss value reaches the minimum, a pre-trained scene model is generated.
[0148] In an embodiment of the present application, on the one hand, a causal graph is constructed using the medical data to be enhanced. The causal graph is used to characterize the causal relationship between diseases, symptoms, and treatment plans. By modeling the causal relationship, the introduction of data that does not conform to medical laws can be avoided, so that the generated data is of high quality and the performance of the AI model can be effectively improved. On the other hand, by using a potential outcome framework for counterfactual inference, the potential outcome data of each patient under different treatment plans can be simulated, so that the results tend to be realistic. Training the model based on the results can improve the generalization ability of the AI model. On the other hand, by combining a preset generative adversarial network, high-quality medical data synthesis can be achieved, data distribution drift can be avoided, high-quality enhanced data can be generated, and tasks in different scenarios can be optimized, thereby improving the robustness and clinical applicability of the AI model.
[0149] Those skilled in the art will appreciate that all or part of the processes in the above-described method embodiments can be implemented by instructing related hardware through a computer program. The program for medical data enhancement based on causal inference can be stored in a computer-readable storage medium. When executed, the program can include the processes in the above-described method embodiments. The storage medium for the program for medical data enhancement based on causal inference can be a magnetic disk, an optical disk, a read-only memory, or a random access memory.
[0150] The above disclosure is only a preferred embodiment of the present application, and certainly cannot be used to limit the scope of rights of the present application. Therefore, equivalent changes made according to the claims of the present application are still within the scope covered by the present application.
Claims
1. A medical data enhancement method based on causal inference, characterized in that: The method comprises: Acquire and preprocess the original medical data of existing actual cases to obtain the medical data to be enhanced; Constructing a causal graph based on the medical data to be enhanced; wherein the causal graph is used to represent the causal relationship between diseases, symptoms, and treatment plans; performing causal effect estimation on the medical data to be enhanced to obtain a weighted data set; Based on the causal graph and the weighted data set, counterfactual inference is performed using a potential outcome framework to obtain a virtual case; An enhanced medical data set is generated based on the virtual case and the preset generative adversarial network.
2. The method according to claim 1, characterized in that The step of constructing a causal graph based on the medical data to be enhanced includes: Using a structural equation modeling method to perform causal relationship modeling on the medical data to be enhanced, and constructing an original causal graph; Access existing literature and expert knowledge; Based on the existing literature and expert knowledge, the edges and directions of the original causal graph are optimized to obtain an optimized causal graph.
3. The method according to claim 1, characterized in that The estimating the causal effect of the medical data to be enhanced to obtain a weighted data set includes: Using the instrumental variable method, identify and estimate the impact of confounding factors on the medical data to be enhanced, and obtain adjusted causal effect estimation results; Using the causal effect estimation result, performing propensity score matching on the medical data to be enhanced to simulate a clinical trial environment to obtain a matched data set; Inverse probability weighting is applied to the matched data set to eliminate sample selection bias, thereby obtaining a weighted data set.
4. The method according to claim 1, wherein Based on the causal graph and the weighted data set, counterfactual inference is performed using a potential outcome framework to obtain a virtual case, including: constructing potential outcome data for each patient under different treatment plans based on the causal graph and the weighted data set; Using synthetic control methods, virtual cases are constructed by combining data on each patient's potential outcomes under different treatment options.
5. The method according to claim 1, characterized in that Generating an enhanced medical data set based on the virtual case and the preset generative adversarial network includes: Initialize the preset generative adversarial network, which includes a generator and a discriminator; Using the virtual case to train the generator so that the generator generates simulated data samples according to the causal relationship; training the discriminator using the original medical data and the simulated data samples so that the discriminator distinguishes between the original medical data and the simulated data samples; By using the Wasserstein GAN training method and preset medical prior knowledge, the generator and the discriminator are alternately optimized to output a forged dataset that is close to the real data distribution; The forged data set is merged with the original medical data to obtain an enhanced medical data set.
6. The method according to claim 1, characterized in that The acquisition and preprocessing of original medical data of existing actual cases to obtain the medical data to be enhanced includes: Obtain original medical data of existing actual cases; Filling missing values, correcting outliers, and standardizing features in the original medical data to obtain a cleaned standardized data file; The data in the cleaned standardized data file is used as the medical data to be enhanced.
7. The method according to claim 1, characterized in that The method further comprises: Create a scenario model for the disease scenario to which the original medical data belongs; Inputting the enhanced medical data set into the scenario model and outputting the loss value of the model; When the loss value reaches the minimum, a pre-trained scene model is generated.
8. A medical data enhancement device based on causal inference, characterized in that: The device comprises: The original medical data acquisition module is used to acquire and pre-process the original medical data of existing actual cases to obtain the medical data to be enhanced; A causal graph construction module, configured to construct a causal graph based on the medical data to be enhanced; wherein the causal graph is used to represent the causal relationship between diseases, symptoms, and treatment plans; a causal effect estimation module, configured to perform causal effect estimation on the medical data to be enhanced to obtain a weighted data set; a counterfactual inference module, configured to perform counterfactual inference based on the causal graph and the weighted data set using a potential outcome framework to obtain a virtual case; The medical data set generation module is used to generate an enhanced medical data set based on the virtual case and a preset generative adversarial network.
9. A computer storage medium, characterized in that The computer storage medium stores a plurality of instructions, and the instructions are suitable for being loaded by a processor and executing the method according to any one of claims 1 to 7.
10. A device, characterized in that include: A processor and a memory; wherein the memory stores a computer program, and the computer program is suitable for being loaded by the processor and executing the method according to any one of claims 1 to 7.
Citation Information
Patent Citations
General multi-disease prediction system generated based on causal verification data
CN114664452A
Disease prediction method and system based on causal inference and dynamic integration of multiple tags
CN116364274A
Integrated adaptive neural fuzzy system applied to diabetes analysis
CN116629311A
Method and device for training debiased multi-modal large language model for medical care
CN119361165A
Cited By
Medical synthetic data analysis method and device based on causal reasoning, and medium
CN121075691A