Coronary heart disease interpretability prediction method based on LLM data enhancement and causal Bayesian inference
Through LLM-based data enhancement and causal Bayesian inference methods, the dataset imbalance and interpretability problems in CHD prediction are solved, high-quality CHD prediction and causal relationship revelation are achieved, the missed diagnosis rate is reduced and explainable diagnostic support is provided.
Patent Information
- Application Number
- CN202510750368.8
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-06-06
- Publication Date
- 2025-09-16
AI Technical Summary
Existing technologies have problems of data set imbalance and poor interpretability in the prediction of coronary heart disease. Traditional models are insensitive to minority samples and lack the ability to reveal causal relationships, resulting in a high rate of missed diagnosis and difficulty in providing explainable diagnostic support.
LLM-based data augmentation and causal Bayesian inference methods were used. Minority samples were generated through MCMC sampling, combined with LLM for data augmentation. The Bayesian Information Criterion and LLM prior knowledge were used to discover causal structures, and a highly interpretable coronary heart disease prediction model was constructed.
It achieves high-quality data enhancement for coronary heart disease and accurate prediction of causal relationships, reduces the missed diagnosis rate, provides explainable diagnostic logic support, and improves the stability and credibility of the model.
Smart Images

Figure CN120656717A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to an interpretable prediction method for coronary heart disease based on LLM data enhancement and causal Bayesian inference, belonging to the intersection of artificial intelligence and medical health, and specifically to a data enhancement and interpretable prediction method for early prediction of coronary heart disease, belonging to the field of computer-implemented disease prediction technology and medical diagnosis auxiliary systems. Background Art
[0002] Coronary Heart Disease (CHD) is one of the cardiovascular diseases with the highest morbidity and mortality rates worldwide, and its early warning is crucial for clinical intervention. In recent years, with the rapid accumulation of large-scale clinical data such as electronic health records (EHR) and electrocardiograms (ECG), automated CHD detection methods based on machine learning and deep learning have been widely studied. For example, models such as logistic regression, random forest, and convolutional neural networks all achieve CHD risk prediction by extracting and classifying features of indicators such as blood sugar, blood pressure, cholesterol, or ECG signals in EHR. However, existing technologies still have the following shortcomings:
[0003] The problem of dataset imbalance is serious. This issue exists not only in CHD research but also in the medical field generally. CHD cases often make up a minority of datasets, which causes the model to learn more features of negative samples during training, leading to model bias. Numerous review articles have shown that traditional machine learning methods (such as random forests and support vector machines) can achieve accuracy rates exceeding 80% on standard electrocardiogram (ECG) and clinical indicator datasets, but are insensitive to minority samples (CHD-positive) and have a high rate of missed diagnoses. SMOTE (Synthetic Minority Oversampling Technique) proposed by Chawla et al. (2002) is a classic method for dealing with sample imbalance in the medical field. It significantly improves the detection performance of minority classes by generating synthetic samples in the feature space; Apostolopoulos et al. (2020) further verified the effectiveness of SMOTE on coronary artery disease datasets, but also pointed out that redundant and overfitted samples are easily generated in high-dimensional feature scenarios; recently, Kosolwattana et al. (2023) proposed the self-inspected adaptive SMOTE (SASMOTE) for datasets with ultra-high imbalance ratios, which can effectively improve the F1-score, but cannot simultaneously meet the authenticity and diversity of synthetic samples.
[0004] Existing models have poor interpretability. Most machine learning, deep learning, or integrated models treat the prediction task as a "black box." For example, the hybrid random forest and deep learning framework proposed by Vishnupriya et al. (2025) achieved an overall accuracy of 92% on a public CHD dataset. However, these models only predict the target variable based on correlations among feature variables, but fail to reveal causal relationships between clinical feature variables. This makes it difficult to output interpretable diagnostic logic, explain the pathological mechanisms of the disease, and provide mechanistic support for clinical decision-making.
[0005] There are currently a few studies exploring the causal relationship between variables in the prediction process, such as the Peter-Clark (PC) algorithm and the Greedy Equivalence Search (GES) algorithm. However, these studies primarily rely on data-driven search and ignore the integration of prior knowledge, not to mention the recently rapidly developing Large Language Model (LLM). Relying solely on data-driven search can lead to unstable model performance. For example, the CSD (Computational Casual Structure Discovery) method applied to HER by Xinpeng Shen et al. (2021) can reveal potential causal relationships, but its performance is unstable in scenarios with sparse clinical data and high noise. Alternatively, they rely on manually defined priors, but the huge workload makes it difficult to carry out on a large scale. Summary of the Invention
[0006] The technical problem solved by the present invention is: to overcome the deficiencies of the existing technology and propose an interpretable prediction method for coronary heart disease based on LLM data enhancement and causal Bayesian inference.
[0007] The technical solution of the present invention is:
[0008] An interpretable prediction method for coronary heart disease based on LLM data enhancement and causal Bayesian inference, the method includes the following steps:
[0009] The first step is to obtain the electrocardiogram dataset ECG dataset Including the majority of and diseased minorities Due to the characteristics of medical data, usually That is, the number of sick people is much smaller than the number of non-sick people. Data preprocessing is performed, which includes missing value processing, outlier processing and data normalization processing in sequence, and the processed data set G is obtained. The data set G includes the majority class without and diseased minorities
[0010] In the second step, according to the diseased minority class in the dataset G obtained in the first step, Randomly initialize an index subset s of size k (0) ={i1,i2,...,i k}. Make a distribution proposal, construct a neighbor subset s′, and (t-1) Randomly select a position p from the set, and randomly select a new index from the remaining unselected indexes to replace the position. Take the new subset S′ as the candidate state, and calculate the score difference between it and the current state ΔS=S(S′)-S(S (t-1) ), the S(·) operator is used to calculate the comprehensive score of coverage, diversity and quantile coverage. According to the acceptance and rejection rules, the index in the subset is adjusted, and the above proposal-acceptance process is repeated for T steps to obtain a Markov chain on the "subset space" Finally, the subset with the highest score is used as the few-sample seed set T is the number of rounds of distribution proposals; t = 1, 2, 3, ..., T;
[0011] The third step is to use the small sample seed set obtained in the second step Prompt is designed based on the data information. LLM is used for data enhancement using the Few-shot method to obtain a sample set with positive results for coronary heart disease.
[0012] The fourth step is to calculate the diseased minority class in the data set G obtained in the first step. The mean μ and covariance matrix ∑ of Each generated sample Calculate Mahalanobis distance Set the threshold τ and keep The sample, that is Combined non-infectious majority Minority with illness and Get a balanced dataset
[0013] Step 5: Balance the data set obtained in step 4 The Bayesian Information Criterion, Bayesian Dirichlet Equivalence and LLM structural prior knowledge scores are weighted, and the hill climbing algorithm is used to find the optimal network structure based on the weighted results. And use Bayesian Estimator to calculate the optimal network structure Each node X k Conditional distribution probability (CPD), according to the optimal network structure obtained and CPD to predict coronary heart disease;
[0014] Step 6: Optimal network structure Each node X k The do(·) operator is used to simulate external interventions, and the difference between the expected values of the two extreme intervention conditions is used to represent the average treatment effect, which quantifies the effect of the characteristic X. k The magnitude of the causal effect on the risk of coronary heart disease.
[0015] Step 7: Causal Bayesian network based on the learning in step 5 Perform individual causal analysis on each patient and calculate the individual characteristic influence score P(Y=1|X k =x)-P(Y=1), this score is used to measure the feature X k The contribution of the value of to the individual's risk of coronary heart disease is completed, and the interpretability analysis of coronary heart disease based on LLM data enhancement and causal Bayesian inference is completed.
[0016] In the first step, the method for handling missing values is: for any record, if two or more variables are missing, delete the record. Otherwise, fill the missing value with the median. The method for handling outliers is: calculate the first quartile Q1 and the third quartile Q3 in each variable, thereby obtaining the interquartile range IQR, IQR = Q3-Q1. Set the multiple iqr_multiplier, the upper bound of the outlier boundary upperbound = Q1 + iqr_multiplier × IQR, the lower bound of the outlier boundary lowerbound = Q1-iqr_multiplier × IQR, replace all values above the upper bound with the upper bound value, and replace all values below the lower bound with the lower bound value. The data normalization processing method is: for each variable, use the Min-Max method to linearly scale the original data x′ to the [0,1] interval based on the minimum and maximum values of the column, and obtain x norm ,
[0017] In the second step, the coverage is calculated as follows: let the entire data set have n samples, the index set of the selected k samples is S, and the corresponding feature vector is For any sample x in the data i (i=1,…,n), calculate its minimum Euclidean distance d to the selected sample set i =min j∈S ||x i -x j ||2, and The calculation of diversity is: only meaningful when k>1, calculate each pair of distances between the selected samples and then take the average, that is, The calculation of quantile coverage is as follows: for each numerical feature X, calculate the specified quantile set {QX (f)|f∈{0.1,0.25,0.5,0.75,0.9}}, let the minimum and maximum values of the selected samples on this feature be and Calculate the number of quantiles covered, Then calculate the ratio r of the value range of the selected sample on this feature to the value range of the original data X , The proportion of the number of quantiles that will be covered Because there are 5 quantiles in total) and the range coverage ratio r X Average to get the percentile coverage score of the feature Finally, according to the weight ω of each feature X (Default is 1), weighted average of all feature scores is obtained S(·)=ω1×coverage+ω2×diversity+ω3×quantile coverage. The acceptance-rejection rule is:
[0018]
[0019] In the third step, the design of prompt refers to the use of role-playing, few-shot and clear instructions. In the data augmentation scenario in the medical field, the role-playing can be designed as follows: "You are a medical data augmentation assistant. I need you to generate medical data of coronary heart disease patients that conform to a specific distribution." The few-shot method is to use the data obtained in the second step to Add to the prompt. Explicit instructions fully describe the task that the LLM needs to complete, such as "I will provide several real coronary heart disease sample data as a reference. Please generate new samples based on these samples, ensuring that the generated data is medically reasonable and has a feature distribution similar to the original data. Please generate the following columns (and only these columns)."
[0020] In the fourth step, the diseased minority class in the dataset G The mean Covariance matrix Mahalanobis distance
[0021] In the fifth step, the Bayesian Information Criterion (BIC) calculation formula is as follows:
[0022]
[0023] Where m is the total number of nodes in the Bayesian network, that is, the number of variables, X k represents the kth node, Pa(X k ) represents node X k In structure The parent node set in Indicates that the data set is balanced Only X k The part of observations related to its superset, is the parameter estimate obtained by maximum likelihood estimation (MLE) after the structure is fixed, λ(X k ,Pa(X k )) is with X k and its superset together define the number of independent parameters. For example, if X k There are r states, Pa(X k ) has q combinations, then the number of parameters is (r-1)×q, In general, the first term lnp(·) is the log-likelihood, which indicates the goodness of fit to the data; the second term is the "complexity penalty" that encourages sparse structures.
[0024] The calculation formula of Bayesian Dirichlet equivalent uniformity (BDeu) is as follows:
[0025]
[0026] where q k Represents the parent set Pa(X k ), the total number of possible value combinations, r k Represents variable X k The number of states itself (the number of categories after discretization), j = 1, ..., q k , enumerate the j-th parent node configuration; Indicates the total number of samples in the data when the parent configuration is the jth type; N jr Indicates that in the data, the parent configuration is the jth type and X k Take the count of the rth state; represents the equivalent “virtual” count assigned to each (j,r) configuration in the Dirichlet prior (assuming uniform distribution); represents the total prior count under the j-th parent configuration; Γ(·) represents the Gamma function, which is used in the closed-form expression of the Dirichlet–polynomial marginal likelihood; ESS (Equivalent Sample Size) is a hyperparameter representing the equivalent sample size in the prior, which defaults to 1 in the pgmpy library.
[0027] In the LLM structure prior knowledge scoring, the LLM API is called to obtain a JSON file containing the prior knowledge edge relationship. The specific scoring calculation formula is as follows: S prior (X k ,Pa(X k ))
[0028]
[0029] Where must_edges = (A, B), ... represents the edge set that "must appear" in the LLM prior knowledge; forbidden_edges = (C, D), ... represents the edge set that "cannot appear" in the LLM prior knowledge; 1 (·) Represents the indicator function, which takes 1 when the condition is met, otherwise takes 0; ω r Indicates that when a necessary edge A→B actually appears in Pa(X k ) in the middle, add ω r If it does not appear, then reduce ω r points; similarly, ω p Indicates that when a forbidden edge C→D appears in Pa(X k ) in the middle, additional reduction p point.
[0030] Then weighted where ω bdeu ,ω bic and ω prior are the weights of the three scores, satisfying ω bdeu +ω bic +ω prior =1. hybrid As the goal of hill climbing search, by adding / deleting / flipping an edge locally, the scores are continuously compared and finally converge to the structure with the highest score. In the best network We use Bayesian Estimator to learn the conditional distribution probability of all nodes, train a Bayesian network model, and then use VariableElimination as the inference engine. For each sample to be predicted, we extract all observed features X except the target variable Y from the network nodes. i =x i As evidence e, then for each i Conditioning is performed on the factors of X, keeping only i =x i The corresponding probability value, the remaining entries are equivalent to multiplying by zero, and the branch is removed. Next, all the remaining implicit variables Z that are not query variables and evidence variables are factored out in turn: select an implicit variable Z from the current factor set Φ, and find all factors containing Z Multiply them together to get a new factor φ′, and then add φ1,…,φ kRemove it from Φ, add φ′ and continue to eliminate the next hidden variable, gradually eliminating all irrelevant variables from the joint distribution, leaving only those factors related to Y and evidence. Finally, we get two unnormalized factor values about Y, corresponding to φ(Y=0) and φ(Y=1), and then normalize Take P(Y=1|e) as the posterior probability, and then binarize the probability based on the threshold (the default is 0.5): when P≥threshold, the predicted label is 1, otherwise it is predicted to be 0.
[0031] In the sixth step, the two extreme intervention conditions refer to X k The highest and lowest discrete values of To represent the highest discrete value, use represents the lowest discrete value, and the probability of observing Y = 1 under these interventions is and The average treatment effect where E[·] represents the expected value operator;
[0032] Specifically, represents a clinical dataset, where x i ∈R d represents the feature vector of the i-th patient, y i ∈{0,1} indicates whether the patient has coronary heart disease (1 means sick, 0 means not sick). Due to the serious imbalance in sample distribution in the coronary heart disease dataset, that is, To solve this problem, we first use the LLM data augmentation method (M-LLM) based on Markov chain Monte Carlo sampling (MCMC) to generate high-quality minority class samples. Sampling to obtain a small number of sample seed sets To approximate the real data distribution, and use LLM to generate candidate synthetic sample sets To ensure the authenticity and consistency of synthetic samples, we apply statistical constraint filters Filter effective samples to obtain enhanced sets So the final dataset is At the same time, in order to enhance the interpretability and robustness of the model, the prompt words are designed to obtain the prior information provided by LLM, guiding the clinical variable set X = {X1, X2, ..., X d Causal structure discovery on} is achieved by using a weighted hybrid scoring method of Bayesian Information Criterion, Bayesian Dirichlet Equivalence and LLM structural prior knowledge scoring to find a directed acyclic graph using a hill climbing algorithm. Based on the causal diagram Perform Bayesian causal inference to predict coronary heart disease and build the final coronary heart disease prediction model The model integrates LLM for data enhancement and relies on the prior knowledge it provides for causal inference, solving the problem of data set imbalance while achieving accurate and explainable predictions.
[0033] Beneficial effects
[0034] The method of the present invention is mainly based on the two modules of M-LLM and CPF to design an innovative method for sample processing and causal structure discovery of small sample imbalanced datasets. The M-LLM module samples the minority class samples in the imbalanced dataset based on MCMC, selects the seed samples with overall representativeness, uses LLM to perform data enhancement through the Few-shot method, and sets filters. The CPF module ensures the authenticity and distribution consistency of generated data. It incorporates domain prior knowledge provided by LLM, combining the advantages of data-driven and LLM prior knowledge to perform Bayesian network structure discovery. It proposes a hybrid objective function and hybrid local scoring, and uses BIC, BDeu, and LLM-prior scoring to support a hill climbing algorithm to search for the optimal causal structure. Ultimately, predictions are achieved based on this network.
[0035] The CHD detection method based on M-LLM data enhancement and causal Bayesian inference in the present invention has the following significant advantages over the existing technology:
[0036] High-quality, distribution-adaptive data augmentation: MCMC sampling is used to automatically select minority class seeds, combined with LLM for few-shot data generation, and then filtered with Mahalanobis distance statistical constraints to ensure that the synthesized samples meet medical rationality in high-dimensional space while maintaining consistency with the original data distribution, significantly reducing the risk of overfitting.
[0037] Hybrid structural learning driven by domain priors and data: The prior structural information provided by LLM is introduced into the hybrid scoring Bayesian network. The BIC / BDeu scores and prior knowledge are balanced through adjustable weights, which offsets the limitations of data-driven search alone. The causal directed acyclic graph (DAG) is generated to meet the constraints of the same domain, improving the accuracy and credibility of causal structure discovery.
[0038] Interpretable causal inference and quantification of intervention effects: Based on the discovered causal graph, ATE is calculated using the Bayesian causal inference framework, which can not only explain the causal impact of various clinical characteristics on CHD risk, but also provide probabilistic guidance for potential intervention measures (such as adjusting a certain indicator), which is significantly better than the traditional black-box classification model.
[0039] Minority class seed sampling replacement: Importance sampling or Gaussian mixture model (GMM) can be used instead of MCMC to select representative samples of minority classes;
[0040] Diversified LLM replacement and hint design: In the data augmentation module, you can choose different LLMs such as GPT-4, Llama-2, and Claude, and use zero- / one- / multi-shot hint strategies to optimize the quality of synthetic samples;
[0041] Diversified filters: In the statistical constraint phase, in addition to the Mahalanobis distance, anomaly detectors based on the Wasserstein distance or isolation forest can also be introduced as filters;
[0042] Alternative causal structure learning algorithms: Hybrid scoring BN can be replaced by a PC-algorithm with prior constraints or a width-limited Greedy Equivalence Search (GES) algorithm, both of which incorporate LLM priors in the objective function.
[0043] Replacement of intervention effect estimation methods: ATE calculation can adopt Pearl's G-computation, inverse probability weighting (IPW) and other technologies to adapt to different clinical scenarios. BRIEF DESCRIPTION OF THE DRAWINGS
[0044] Figure 1 This is the T-SNE display diagram;
[0045] Figure 2 To select the distribution diagram of each characteristic of the sample;
[0046] Figure 3 A causal relationship diagram for a record with a target variable of 1;
[0047] Figure 4 A causal relationship diagram for a "target variable is 0" record. DETAILED DESCRIPTION
[0048] The present invention will be further described below with reference to the accompanying drawings and embodiments.
[0049] Example
[0050] An interpretable prediction method for coronary heart disease based on LLM data enhancement and causal Bayesian inference, the method includes the following steps:
[0051] The first step is to obtain the electrocardiogram dataset ECG dataset Including the majority of and diseased minorities Due to the characteristics of medical data, usually That is, the number of sick people is much smaller than the number of non-sick people. Data preprocessing is performed, which includes missing value processing, outlier processing and data normalization processing in sequence, and the processed data set G is obtained. The data set G includes the majority class without and diseased minorities
[0052] In the second step, according to the diseased minority class in the dataset G obtained in the first step, Randomly initialize an index subset S of size k (0) ={i1,i2,...,i k}. Make a distribution proposal, construct a neighbor subset S′, and from the current subset S (t-1) Randomly select a position p from the set, and randomly select a new index from the remaining unselected indexes to replace the position. Take the new subset S′ as the candidate state, and calculate the score difference between it and the current state ΔS=S(S′)-S(S (t-1) ), the S(·) operator is used to calculate the comprehensive score of coverage, diversity and quantile coverage. According to the acceptance and rejection rules, the index in the subset is adjusted, and the above proposal-acceptance process is repeated for T steps to obtain a Markov chain on the "subset space" Finally, the subset with the highest score is used as the few-sample seed set like Figure 1 and Figure 2 As shown, T is the number of rounds of distribution proposals; t = 1, 2, 3, ..., T;
[0053] The third step is to use the small sample seed set obtained in the second step Prompt is designed based on the data information. LLM is used for data enhancement using the Few-shot method to obtain a sample set with positive results for coronary heart disease.
[0054] Although we have adopted a few-shot method and detailed prompt words, and with the continuous advancement of technology, the semantic understanding ability of LLM is getting better and better, strictly speaking, we still cannot completely avoid the occurrence of LLM hallucination. In a research paper "Survey of Hallucination in Natural Language Generation" from the Artificial Intelligence Research Center, hallucination is defined as "the generated content does not match the provided source content or is meaningless." Therefore, in order to ensure the authenticity and distribution consistency of the synthesized samples, we introduced Mahalanobis Distance (MD) as a filter For the dataset generated by LLM Filter. First calculate the diseased minority class The mean μ and covariance matrix ∑ of :
[0055]
[0056] Among them, μ is the arithmetic mean of all minority class samples in each dimension, which is used to characterize the central position of the sample in each dimension; Σ is used to measure the linear relationship between the dimensions of the sample and the degree of dispersion of its individual variables. As the denominator, it ensures an unbiased estimate of the population covariance (i.e., Bessel correction). Each generated sample Calculate Mahalanobis distance:
[0057]
[0058] By setting the threshold τ, the samples, thereby obtaining a data enhancement set that conforms to the original data distribution If you do not set the threshold manually, you can also use the "95% quantile of the chi-square distribution when the degrees of freedom are equal to the number of selected features and the significance level is 0.05" as the threshold.
[0059] After data augmentation, we get a balanced distribution dataset A common approach is to use traditional machine learning for prediction. However, considering the "black box" effect of machine learning and deep learning that is difficult to explain, we proposed a "causal perception prediction framework with LLM prior guidance" to discover the causal relationship diagram between clinical multivariate variables by utilizing the LLM's vast knowledge base and Bayesian network. Based on this structural diagram, interpretable Bayesian causal inference is performed, which is divided into two parts: (a) structural learning guided by LLM prior; (b) average treatment effect.
[0060] (a) LLM prior-guided structure learning
[0061] A Bayesian network is a directed graph that plots the random variables involved in a research system based on whether they are conditionally independent. It is mainly used to describe the conditional dependencies between random variables, with circles representing random variables and arrows representing conditional dependencies. Existing Bayesian network structure learning mostly relies on data-driven (such as PC and GES algorithms), which is easily affected by sample size and noise, and difficult to incorporate domain expert knowledge. In order to combine the characteristics of data-driven and LLM priors, we used a hybrid scoring BN for structure learning, that is, weighted the Bayesian information criterion, Bayes Dirichlet equivalent uniformity, and LLM structure prior knowledge scores to obtain a hybrid score. The specific calculation formula is as follows:
[0062] ① The Bayesian Information Criterion (BIC) calculation formula is as follows:
[0063]
[0064] Where m is the total number of nodes in the Bayesian network, that is, the number of variables, X k represents the kth node, Pa(X k ) represents node X k In structure The parent node set in Indicates that the data set is balanced Only X k The part of observations related to its superset, is the parameter estimate obtained by maximum likelihood estimation (MLE) after the structure is fixed, λ(X k ,Pa(X k )) is with X k and its superset together define the number of independent parameters. For example, if X k There are r states, Pa(X k ) has q combinations, then the number of parameters is (r-1)×q, In general, the first term lnp(·) is the log-likelihood, which indicates the goodness of fit to the data; the second term is the "complexity penalty" that encourages sparse structures.
[0065] ② The Bayesian Dirichlet equivalent uniform calculation formula is as follows:
[0066]
[0067] where q k Represents the parent set Pa(X k ), the total number of possible value combinations, r k Represents variable X k The number of states itself (the number of categories after discretization), j = 1, ..., q k , enumerate the j-th parent node configuration; Indicates the total number of samples in the data when the parent configuration is the jth type; N jr Indicates that in the data, the parent configuration is the jth type and X k Take the count of the rth state; represents the equivalent “virtual” count assigned to each (j,r) configuration in the Dirichlet prior (assuming uniform distribution); represents the total prior count under the j-th parent configuration; Γ(·) represents the Gamma function, which is used for the closed-form expression of the Dirichlet–multinomial marginal likelihood; ESS is a hyperparameter representing the equivalent sample size in the prior, which defaults to 1 in the pgmpy library.
[0068] ③ In the LLM structure prior knowledge scoring, call the LLM API to obtain a JSON file containing the prior knowledge edge relationship. The specific scoring calculation formula is as follows:
[0069]
[0070] Where must_edges = (A, B), ... represents the edge set that "must appear" in the LLM prior knowledge; forbidden_edges = (C, D), ... represents the edge set that "cannot appear" in the LLM prior knowledge; 1 (·) Represents the indicator function, which takes 1 when the condition is met, otherwise takes 0; ω r Indicates that when a necessary edge A→B actually appears in Pa(X k ) in the middle, add ω r If it does not appear, then reduce ω r points; similarly, ω p Indicates that when a forbidden edge C→D appears in Pa(X k ) in the middle, additional reduction p point.
[0071] ④Then weighted where ω bdeu ,ω bic and ω prior are the weights of the three scores, satisfying ω bdeu +ω bic +ω prior =1, with S hybrid As the goal of hill climbing search, by adding / deleting / flipping an edge locally, the scores are continuously compared and finally converge to the structure with the highest score. In the best network We use the Bayesian Estimator to learn the CPT of all nodes, train a Bayesian network model, and then use the VariableElimination provided by the pgmpy library as the inference engine. For each sample to be predicted, we extract all observed features X except the target variable Y from the network nodes. i =x i As evidence e, then for each i Conditioning is performed on the factors of X, keeping only i =x i The corresponding probability value, the remaining entries are equivalent to multiplying by zero to remove the branch. Next, all the remaining non-query variables and non-evidence variables are factored out in turn: select a hidden variable Z from the current factor set Φ, and find all factors containing Z Multiply them together to get a new factor φ′, and then add φ1,…,φ k Remove it from Φ, add φ′ and continue to eliminate the next hidden variable, gradually eliminating all irrelevant variables from the joint distribution, leaving only those factors related to Y and evidence. Finally, we get two unnormalized factor values about Y, corresponding to φ(Y=0) and φ(Y=1), and then normalize Take P(Y=1|e) as the posterior probability, and then binarize the probability based on the threshold (taken as 0.45 in this experiment. In the medical field, the recall rate is increased as much as possible to avoid missed diagnoses): when P ≥ the threshold, the predicted label is 1, otherwise it is predicted as 0.
[0072] (b) Average treatment effect
[0073] Based on the discovered causal structure, this study introduces the do(·) operator proposed by Pearl to simulate external intervention and to optimize the network structure. Each node X k The do(·) operator is used to simulate external interventions, and the average treatment effect (ATE) is expressed as the difference in expected values between two extreme intervention conditions, which quantifies the effect of the characteristic X. k The magnitude of the causal effect on the risk of coronary heart disease. Specifically, we let and Represents X k The probability of observing Y = 1 under these interventions is expressed as and Based on these intervention probabilities, X k The ATE calculation formula is as follows:
[0074]
[0075] where E[·] represents the expected value operator. This measures the effect of an external k How much does the probability of the coronary heart disease prediction result change when it is set to its highest and lowest possible values? Finally, the prediction is made. Two examples are shown below (such as Figure 3 and Figure 4 shown):
[0076] Causal analysis of case 27 revealed that the primary driver of disease was the QRS axis in the electrocardiogram (ECG). A QRS axis value of 0 (using the data discretized into five bins, i.e., quantiles within the overall data distribution, when constructing the Bayesian network) increased the patient's probability of a CHD diagnosis by approximately 28% compared to baseline. Next, the QRS duration, at a value of 4, also contributed a positive 12.3% to the risk, while the P–R interval (at a value of 3) contributed approximately 7.5%. Furthermore, ECG timing indicators such as the QT interval, heart rate, P wave duration, and T wave duration all contributed slightly to the risk (between 1% and 5%). As a posterior effect of this diagnosis, RV5 amplitude (at a value of 2) slightly reduced the risk (approximately –1.1%), while SV1 amplitude had a negligible effect. This shows that the QRS axis, QRS duration, and P–R interval are the most critical predictors in this case, and their changes have the most significant impact on the diagnosis of coronary artery disease. Similarly, we can perform a similar analysis on another case. We can see in this figure that the QRS axis, QRS duration, and other ECG duration indicators have minimal impact on the risk of disease, which is consistent with the fact that the target variable itself is 0 (i.e., no coronary artery disease).
[0077]
[0078]
[0079] Table 1 Prompt design
[0080] In summary, the above are only preferred embodiments of the present invention and are not intended to limit the scope of protection of the present invention. Any modifications, equivalent replacements, improvements, etc. made within the spirit and principles of the present invention should be included in the scope of protection of the present invention.
Claims
1. An interpretable prediction method for coronary heart disease based on LLM data enhancement and causal Bayesian inference, characterized by The steps of the method include: The first step is to obtain the electrocardiogram dataset ECG dataset Including the majority of and diseased minorities For the electrocardiogram dataset Perform data preprocessing to obtain the preprocessed data set G, which includes the majority class and diseased minorities The second step is to randomly initialize the diseased minority class in the dataset G obtained in the first step. Get an index subset S of size k (0) ={i1,i2,...,i k }, for the index subset S (0) ={i1,i2,...,i k } Make a distribution proposal and get a Markov chain on the "subset space" Get the subset with the highest score as the few-sample seed set T is the number of rounds of distribution proposals; t = 1, 2, 3, ..., T; The third step is to use the Few-shot method to use LLM to the few sample seed sets obtained in the second step. Perform data enhancement to obtain a sample set with positive results for coronary heart disease Step 4: Calculate the sample set obtained in step 3 Each sample Mahalanobis distance Set the threshold τ and keep All samples of form a sample set Combined non-infectious majority Minority with illness and Get a balanced dataset Step 5: Balance the data set obtained in step 4 The Bayesian Information Criterion, Bayesian Dirichlet Equivalence and LLM structural prior knowledge scores are weighted, and the hill climbing algorithm is used to find the optimal network structure based on the weighted results. Using Bayesian Estimator to calculate the optimal network structure Each node X k The conditional distribution probability of the optimal network structure is obtained and conditional distribution probability to predict coronary heart disease; Step 6: Based on the optimal network structure in step 5 Perform individual causal analysis on each patient and calculate the individual node influence score P(Y=1|X k =x)-P(Y=1), the individual node influence score is used to measure the influence of node X k The contribution of the value of to the individual's risk of coronary heart disease is completed, and the interpretability analysis of coronary heart disease based on LLM data enhancement and causal Bayesian inference is completed.
2. The interpretable prediction method for coronary heart disease based on LLM data enhancement and causal Bayesian inference according to claim 1 is characterized by: In the first step, preprocessing includes missing value processing, outlier processing and data normalization processing.
3. The interpretable prediction method for coronary heart disease based on LLM data enhancement and causal Bayesian inference according to claim 2 is characterized by: The method of handling missing values is as follows: for any record, if more than two variables are missing, the record is deleted; otherwise, the missing value is filled with the median; The method for handling outliers is as follows: calculate the first quartile Q1 and the third quartile Q3 in each variable to obtain the interquartile range IQR, IQR = Q3-Q1; set the multiple iqr_multiplier, the upper bound of the outlier boundary upperbound = Q1+iqr_multiplier×IQR, the lower bound of the outlier boundary lowerbound = Q1-iqr_multiplier×IQR, replace all values above the upper bound with the upper bound value, and replace all values below the lower bound with the lower bound value; The data normalization method is: for each variable, use the Min-Max method to linearly scale the original data x′ to the interval [0,1] based on the minimum and maximum values of the column, and obtain x norm , 4. The interpretable prediction method for coronary heart disease based on LLM data enhancement and causal Bayesian inference according to claim 3 is characterized by: In the second step, the method for making distribution proposals is: From the current subset S (t-1) Randomly select a position p from the list, randomly select a new index from the remaining unselected indexes to replace the position, and obtain a new subset S′ as the candidate state. Calculate the score difference ΔS=S(S′)-S(S (t-1) ), the S(·) operator is used to calculate the comprehensive score of coverage, diversity and quantile coverage, and adjust the index in the subset according to the acceptance and rejection rules.
5. The interpretable prediction method for coronary heart disease based on LLM data enhancement and causal Bayesian inference according to claim 4 is characterized by: The calculation method of coverage is: let the entire data set have n samples, the index set of the selected k samples is S, and the corresponding feature vector is For any sample x in the data i (i=1,…,n), calculate x i The minimum Euclidean distance d to the selected sample set i =min j∈S ||x i -x j ||2, The calculation method of diversity is: only meaningful when k>1, calculate each pair of distances between the selected samples and then take the average, that is, The calculation method of quantile coverage is: for each numerical feature X, calculate the specified quantile set {Q X (f)|f∈{0.1,0.25,0.5,0.75,0.9}}, let the minimum and maximum values of the selected samples on this feature be and Calculate the number of quantiles covered, Then calculate the ratio r of the value range of the selected sample on this feature to the value range of the original data X , The proportion of the number of quantiles that will be covered and range coverage ratio r X Average to get the percentile coverage score of the feature Finally, according to the weight ω of each feature X , taking a weighted average of the scores of all features, S(·) = ω1 × coverage + ω2 × diversity + ω3 × quantile coverage; The accept-reject rules are as follows:
6. The interpretable prediction method for coronary heart disease based on LLM data enhancement and causal Bayesian inference according to claim 4, characterized in that: In the third step, LLM is used to analyze the small sample seed set The input prompt for data enhancement is based on the few sample seed sets Design based on data information.
7. The interpretable prediction method for coronary heart disease based on LLM data enhancement and causal Bayesian inference according to claim 6, characterized in that: The input prompt is designed using role-playing, few-shot, and explicit instruction methods. In the data augmentation scenario in the medical field, the role-play is: "You are a medical data augmentation assistant. I need you to generate medical data of coronary heart disease patients that conforms to the set distribution." Few-sho is Add to the prompt word; Explicit instructions are complete descriptions of the tasks that the LLM needs to complete.
8. The interpretable prediction method for coronary heart disease based on LLM data enhancement and causal Bayesian inference according to claim 6, characterized in that: In the fourth step, Mahalanobis distance Among them, the minority of patients The mean Covariance matrix 9. The interpretable prediction method for coronary heart disease based on LLM data enhancement and causal Bayesian inference according to claim 6, characterized in that: In the fifth step, the Bayesian Information Criterion calculation formula is as follows: Among them, m is the total number of nodes in the Bayesian network, that is, the number of variables, X k represents the kth node, Pa(X k ) represents node X k In structure The parent node set in Indicates that the data set is balanced Only X k The part of observations related to its superset, is the parameter estimate obtained by maximum likelihood estimation after the structure is fixed, λ(X k ,Pa(X k )) is with X k The number of independent parameters defined by its superset, if X k There are r states, Pa(X k ) has q combinations, then the number of parameters is (r-1)×q, represents the total number of observed samples; the first term lnp(·) is the log-likelihood, which indicates the data fit; the second term is the "complexity penalty", which encourages sparse structure; The Bayesian Dirichlet equivalent uniform calculation formula is as follows: Among them, q k Represents the parent set Pa(X k ), the total number of possible value combinations, r k Represents variable X k The number of states itself, j = 1,…,q k , enumerate the j-th parent node configuration; Indicates the total number of samples in the data when the parent configuration is the jth type; N jr Indicates that in the data, the parent configuration is the jth type and X k Take the count of the rth state; represents the equivalent "virtual" count assigned to each (j,r) configuration in the Dirichlet prior; represents the total prior count under the j-th parent configuration; Γ(·) represents the Gamma function, which is used for the closed-form expression of the Dirichlet–polynomial marginal likelihood; ESS is a hyperparameter representing the equivalent sample size in the prior, which defaults to 1 in the pgmpy library; In the LLM structure prior knowledge scoring, the LLM API is called to obtain a JSON file containing the prior knowledge edge relationship. The specific scoring calculation formula is as follows: Among them, must_edges = (A, B), ... represents the edge set that "must appear" in the LLM prior knowledge; forbidden_edges = (C, D), ... represents the edge set that "cannot appear" in the LLM prior knowledge; 1 (·) Represents the indicator function, which takes 1 when the condition is met, otherwise takes 0; ω k Indicates that when a necessary edge A→B actually appears in Pa(X k ) in the middle, add ω r If it does not appear, then reduce ω r points; similarly, ω p Indicates that when a forbidden edge C→D appears in Pa(X k ) in the middle, additional reduction p point; Then weighted Among them, ω bdeu ,ω bic and ω prior are the weights of the three scores, satisfying ω bdeu +ω bic +ω prior =1, with S hybrid As the goal of hill climbing search, by adding / deleting / flipping an edge locally, the scores are continuously compared and finally converge to the structure with the highest score.
10. The interpretable prediction method for coronary heart disease based on LLM data enhancement and causal Bayesian inference according to claim 9, characterized in that: In the best network We use Bayesian Estimator to learn the conditional probability distribution of all nodes, train a Bayesian network model, and then use VariableElimination provided by the pgmpy library as the inference engine. For each sample to be predicted, we extract all observed features X except the target variable Y from the network nodes. i =x i As evidence e, then for each i Conditioning is performed on the factors of X, keeping only i =x i The corresponding probability value, the remaining entries are equivalent to multiplying by zero to remove the branch; Next, perform factor elimination on all other implicit variables Z that are not query variables and evidence variables: select an implicit variable Z from the current factor set Φ and find all factors containing Z. Multiply them together to get a new factor φ′, and then add φ1,…,φ k Remove it from Φ, add φ′ and continue to eliminate the next latent variable, gradually eliminating all irrelevant variables from the joint distribution, leaving only those factors that are related to Y and the evidence; Finally, we get two unnormalized factor values about Y, corresponding to φ(Y=0) and φ(Y=1), and then normalize them. y∈{0,1}, take P(Y=1|e) as the posterior probability, and then binarize the probability according to the threshold: when P≥threshold, the predicted label is 1, otherwise it is predicted to be 0.