Drug Recommendation System and Method Based on Causal Inference

Through a deep learning model based on causal inference, processing patient characteristics and medical examination data, the problem of causal relationship in the drug recommendation system is solved, and more accurate and reliable drug recommendation is achieved.

CN119811576BActive Publication Date: 2025-07-01NORTHEASTERN UNIV CHINA
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202510292960.8
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-03-13
Publication Date
2025-07-01
Estimated Expiration
2045-03-13

AI Technical Summary

Technical Problem

The existing drug recommendation system fails to fully consider the causal relationship between the patient and the drug, which makes the recommendation system susceptible to data deviations and noise interference, affecting the accuracy and applicability of the recommendation.

Method used

A deep learning model based on causal inference is adopted, including a patient status encoder, a medical examination autoencoder and a recommendation module, to estimate the true causal relationship between the patient-drug through causal inference, and a multi-layer perceptron and variational autoencoder is used to process patient characteristics and medical examination data, combining comparative learning and recommendation loss to optimize drug recommendations.

Benefits of technology

By eliminating the effects of unobserved covariates, the accuracy and robustness of drug recommendations are improved, and the predictive ability of the model under complex clinical data is improved.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119811576B_ABST
    Figure CN119811576B_ABST
Patent Text Reader

Abstract

The present invention provides a drug recommendation system and method based on causal inference, which relates to the technical field of data mining. This system uses a deep learning model, specifically including a patient status encoder, a medical examination autoencoder, and a recommendation module; by inputting patient features X , project it into a low-dimensional embedding space using a neural network encoder X . Use an encoder-decoder structure to reconstruct medical examination data R , and obtain an alternative low-dimensional latent variable R of Z R , and calculate the reconstruction loss; use contrastive learning to optimize the low-dimensional latent variable Z R , and calculate the contrastive learning loss; according to the patient features X and the latent variable of the medical examination Z R predict suitable drugs, and calculate the recommendation loss; jointly optimize the reconstruction loss, the contrastive learning loss, and the recommendation loss to optimize the parameters of the drug recommendation system, and establish the causal relationship between the patient and the drug P ( Y | do ( X )) to improve the recommendation ability of the drug recommendation system. Finally, use the drug recommendation system to recommend drugs by combining patient features and medical examination data.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of data mining, and in particular to a drug recommendation system and method based on causal inference. Background Art

[0002] Drug recommendation aims to provide the most suitable drug advice for doctors and patients according to the health status of patients, and avoid side effects and adverse reactions caused by improper drug use. Existing drug recommendation systems are mainly based on electronic medical record data, which are usually organized at the patient level, visit level, and event level. By analyzing the patient's current visit information or multiple visit records, the system can construct an expression of the patient's health status and provide personalized drug recommendations accordingly.

[0003] As a method for studying causal relationships, causal inference focuses on the causal mechanism between variables and attempts to reveal the true impact of a certain intervention on a specific outcome. By introducing causal inference, the problem of data bias can be effectively alleviated, such as distinguishing the spurious associations caused by confounding factors between the irrelevant manifestations of patients and the recommended drugs. In addition, causal inference can also help the system make reasonable inferences in the case of missing data, improving the robustness and credibility of the recommendation.

[0004] However, most existing drug recommendation algorithms adopt data-driven supervised learning methods. Specifically, these algorithms use patient data as the supervision signal for model training. However, the problems of bias and noise in the data are likely to interfere with model training and affect its performance in the deployment stage. For example, due to unobserved covariates, certain drugs may be effective in a specific patient group, but may not be applicable in other groups. This data bias will weaken the generalization ability of the model, resulting in a decrease in the accuracy and applicability of the recommendation. In addition, the medical records of patients may be missing or contain irrelevant diagnostic information, which further affects the reliability of drug recommendation. Summary of the Invention

[0005] Aiming at the deficiencies of the prior art, the present invention provides a drug recommendation system and method based on causal inference, which solves the problem that the existing drug recommendation methods do not fully consider the causal relationship between patients and drugs, resulting in the recommendation system being easily interfered by data bias and noise. The present invention estimates the true causal relationship between patients and drugs through the front-door criterion in causal inference, and implements a deep learning model to improve the accuracy and reliability of the recommendation.

[0006] A drug recommendation system based on causal inference uses a deep learning model, specifically including a patient status encoder, a medical examination autoencoder, and a recommendation module;

[0007] The patient status encoder takes the patient feature X as input and uses a multi-layer perceptron to map the high-dimensional patient feature X to a low-dimensional latent space, generating a latent variable Z of the patient status X , denoted as: Z X = f enc (X, θ X );

[0008] where f enc is the neural network model of the patient status encoder, and θ X are the model parameters;

[0009] The medical examination autoencoder processes the medical examination data R using an encoder-decoder structure, extracts key features and compresses them into a low-dimensional space, and constructs an alternative low-dimensional latent variable Z R ;

[0010] where the encoder-decoder structure adopts a variational autoencoder (VAE) structure, specifically including two modules: an encoder and a decoder, as follows:

[0011] Z R = f vae-enc (X, φ enc )

[0012]

[0013] where f vae-enc and f vae-dec are the encoder and decoder, and φ enc and φ dec are the encoder and decoder parameters, is the reconstructed medical examination data;

[0014] In the variational autoencoder (VAE), the prior distribution of the alternative low-dimensional latent variable Z R is set to a standard multivariate normal distribution:

[0015] P(Z R ) = N(0, I)

[0016] where N(0, I) represents a multivariate normal distribution with a mean of 0 and a covariance matrix of the identity matrix I;

[0017] The encoder takes the medical examination data as input and generates the posterior distribution q(Z R |R) of the alternative low-dimensional latent variable Z R as the output of the encoder:

[0018] q(Z R |R) = N(μ, diag(σ 2 ))

[0019] Among them, μ and σ are the mean and standard deviation of the surrogate low-dimensional latent variable Z R respectively, and diag represents a diagonal matrix;

[0020] The decoder takes the low-dimensional latent variable Z R as input, maps the encoded latent variable back to the original input space, and reconstructs the original data to generate a reconstruction vector for the medical examination data R, that is, the reconstructed medical examination data

[0021] The training objective of the variational autoencoder VAE is to minimize the total loss function L VAE , and the total loss function L VAE includes the reconstruction loss and the KL divergence loss:

[0022] The reconstruction loss is the negative expectation of the conditional log-likelihood:

[0023]

[0024] Among them, denotes the expected value, and logP(R|Z R ) is the log-likelihood term, representing the probability of the observed data R under the condition of the low-dimensional latent variable Z R ;

[0025] If the conditional distribution P(R|Z R ) is assumed to be a normal distribution, it is expanded into the form of mean squared error:

[0026]

[0027] Among them, R i and are the original and reconstructed medical examination data of the i-th sample;

[0028] The KL divergence loss is specifically:

[0029]

[0030] Among them, D KL represents the Kullback-Leibler divergence; d is the dimension of the latent variable, and are the mean and standard deviation of the i-th latent variable respectively, and q(Z R |R) is the posterior distribution of the low-dimensional latent variable Z R given the medical examination R.

[0031] By balancing the reconstruction loss and the KL divergence loss through the hyperparameter β, the total loss function L VAE is:

[0032] LVAE = L reconstruct + β·L KL

[0033] Use contrastive learning to further optimize the surrogate low-dimensional latent variable Z R , capturing the causal association between medical examination data R and patient characteristics X, that is, optimizing the contrastive probability The ratio of the posterior probability P(X|Z R ) and the prior probability P(X);

[0034] In the said contrastive learning: The positive samples are data pairs (X, R) from the same patient, R is the medical examination data of the target patient, and X is the corresponding patient characteristics; The negative samples are randomly selected patient characteristics X' and the medical examination data R of the target patient combined as (X', R); Contrastive learning calculates the contrastive probability by pulling the distance of the positive sample pairs closer and pushing the distance of the negative sample pairs farther Specifically, the InfoNCE loss function is used as the contrastive learning loss, and its definition is as follows:

[0035]

[0036] where τ is the temperature coefficient, i, j represent the positive and negative sample indices, X i , X j are the positive and negative samples of the indices, Z Ri is the surrogate latent variable of the medical examination R of patient i, N is the number of negative samples, and sim(·, ·) is the similarity function; The similarity function sim(·, ·) represents the similarity between the surrogate variable and the patient characteristics, and is calculated using the cosine similarity:

[0037]

[0038] The said recommendation module is a collaborative filtering model, that is, a factorization machine FM; Using the patient state latent variable Z X and the low-dimensional latent variable Z R as inputs, providing drug recommendations for each patient, and the output of the recommendation module is the conditional probability P(Y|R, X), representing the recommendation probability of each drug y ∈ Y; Specifically as follows:

[0039] Y = f recommend (Z X , Z R ; φ Y )

[0040] where f recommend is the recommendation module; φ Y is the parameter of the recommendation module, and the output is a multinomial probability distribution, representing the recommendation probability P(Y|R, X) of each drug;

[0041] Use the recommendation loss Lrecommendation Measure the deviation between the predicted drug recommendation probability of the model and the actual drug label, which is defined as the multi-class cross-entropy loss:

[0042]

[0043] where N represents the number of patients; K is the number of drugs; i is the patient index; k is the drug category index, and y i,k represents the true label of patient i for drug k; P(y i,k |R i ,X i ) represents the predicted probability of patient i for drug k;

[0044] Combine the reconstruction loss, contrastive learning loss, and recommendation loss to construct a joint optimization objective L total :

[0045] L total = λ1·L reconstruct + λ2·L contrast + λ3·L recommendation

[0046] In the formula, λ1, λ2, and λ3 measure the importance weights of the reconstruction loss, contrastive learning loss, and recommendation loss respectively;

[0047] In the causal graph corresponding to the drug recommendation system, the variable X represents patient characteristics, the variable R represents medical examination data, Y represents the set of drugs suitable for the patient, and C represents unobserved covariates; the drug recommendation system aims to mine the causal relationship of X→Y and identify the causal effect P(Y|do(X)); there is a causal path: X→R→Y, where the set of drugs Y suitable for the patient is conditional on patient characteristics X and medical examination results R;

[0048] Based on the front-door adjustment in causal inference, the following causal effect is obtained:

[0049]

[0050] where P(Y|do(X)) represents the causal effect of X on Y; do(X) represents an intervention operation on X; Z R is a low-dimensional latent variable substituting for R, and P(Z R ) and P(X) are the probability distributions of the low-dimensional latent variable Z R and X; P(X|Z R ) represents the probability distribution of X given Z R , and P(Y|X,R) represents the probability distribution of Y given X and R.

[0051] Train the drug recommendation system to optimize the three probability terms P(Z R ), and P(Y|X,R) to obtain the causal effect of X on Y and eliminate the confounding effect caused by the unobserved covariate C; specifically, approximate estimation of the causal effect is performed by optimizing the above reconstruction loss, contrast learning loss, and recommendation loss.

[0052] A drug recommendation method based on causal inference, implemented based on the aforementioned drug recommendation system based on causal inference, specifically including the following steps:

[0053] Step 1: Input patient features X, and project X into the patient state latent variable Z using a neural network encoder X ;

[0054] Step 2: Use an encoder-decoder structure to reconstruct the medical examination data R to obtain an alternative low-dimensional latent variable Z of R R , and calculate the reconstruction loss;

[0055] Step 3: Calculate the contrast learning loss, and optimize the low-dimensional latent variable Z through the contrast learning loss R ;

[0056] Step 4: According to the patient state latent variable Z X and the latent variable Z of the medical examination R predict suitable drugs and calculate the recommendation loss;

[0057] Step 5: Jointly optimize the reconstruction loss, contrast learning loss, and recommendation loss to optimize the parameters of the drug recommendation system, and improve the recommendation ability of the drug recommendation system by optimizing the causal relationship P(Y|do(X)) between the patient and the drug;

[0058] Step 6: After the drug recommendation system is optimized, ignore the contrast learning, and use the drug recommendation system to perform drug recommendation by combining the patient features and medical examination data.

[0059] The beneficial effects of adopting the above technical solutions are as follows:

[0060] The present invention provides a drug recommendation system and method based on causal inference. The present invention utilizes the front-door criterion in causal inference to eliminate the bias and noise effects caused by unobserved covariates C. In model training, first, the patient features X are mapped to a low-dimensional embedding space through a neural network encoder; then, an encoder-decoder structure and a reconstruction loss are used to learn the low-dimensional latent variables as the compressed representation of R; through the contrastive learning loss, it is ensured that the generated low-dimensional latent variables can capture the causal correlation between R and X. According to the patient features and medical examinations, the appropriate drug Y is predicted and the recommendation loss is calculated. Finally, the drug recommendation model integrating causal relationships is obtained by jointly optimizing the reconstruction loss, the contrastive loss, and the recommendation loss. Compared with existing methods, the present invention can embed the potential causal mechanism in the data into the drug recommendation model, eliminate the influence caused by unobserved covariates, reduce the bias and noise in the data, and improve the recommendation ability of the model. This method not only enhances the robustness of the model when facing complex clinical data but also improves the prediction accuracy in practical applications. BRIEF DESCRIPTION OF THE DRAWINGS

[0061] Figure 1 It is a schematic diagram of the causal relationship of the present invention;

[0062] Figure 2 It is a schematic diagram of the drug recommendation system of the present invention;

[0063] Figure 3 It is the overall flowchart of the drug recommendation method of the present invention. DETAILED DESCRIPTION OF THE EMBODIMENTS

[0064] The following combines the drawings and embodiments to further describe in detail the specific embodiments of the present invention. The following embodiments are used to illustrate the present invention but are not used to limit the scope of the present invention.

[0065] A drug recommendation system based on causal inference uses a deep learning model. An embodiment of the present invention discloses a drug recommendation system as Figure 2 shown, specifically including a patient status encoder, a medical examination autoencoder, and a recommendation module; the training process in this embodiment is as Figure 3 shown.

[0066] The patient status encoder inputs patient features X, including the demographic information (such as age, gender) of the patient, medical history data, and other health features. A multi-layer perceptron is used to map the high-dimensional patient features X to a low-dimensional latent space to generate the latent variable Z of the patient status X , and the encoding process is expressed as:

[0067] Z X = f enc (X, θ X )

[0068] where f encThe neural network model for the patient status encoder, θ X is the model parameter;

[0069] The medical examination autoencoder uses an encoder-decoder structure to process the medical examination data R, extracts key features and compresses them into a low-dimensional space representation, and constructs an alternative low-dimensional latent variable Z that can effectively reflect the causal information related to the task R to achieve effective compression and characterization of the medical examination data;

[0070] The encoder-decoder structure adopts a variational autoencoder (VAE) structure, which specifically includes two modules: an encoder and a decoder, as follows:

[0071] Z R = f vae-enc (X, φ enc )

[0072]

[0073] where f vae-enc and f vae-dec are the encoder and decoder, and φ enc and φ dec are the parameters of the encoder and decoder, is the reconstructed medical examination data;

[0074] In the variational autoencoder (VAE), the prior distribution of the alternative low-dimensional latent variable Z R is set to a standard multivariate normal distribution:

[0075] P(Z R ) = N(0, I)

[0076] where N(0, I) represents a multivariate normal distribution with a mean of 0 and a covariance matrix of the identity matrix I;

[0077] The encoder inputs the medical examination data and generates the posterior distribution q(Z R |R) of the alternative low-dimensional latent variable Z R as the output of the encoder:

[0078] q(Z R |R) = N(μ, diag(σ 2 ))

[0079] where μ and σ are the mean and standard deviation of the alternative low-dimensional latent variable Z R respectively, and diag represents a diagonal matrix;

[0080] The decoder uses the low-dimensional latent variable Z RAs the input, it maps the encoded latent variables back to the original input space, reconstructs the original data, and generates a reconstructed vector for the medical examination data R, that is, the reconstructed medical examination data

[0081] The training objective of the variational autoencoder VAE is to minimize the total loss function L VAE , the total loss function L VAE includes a reconstruction loss and a KL divergence loss:

[0082] The reconstruction loss measures the consistency between the generated by the decoder and the medical examination data R, which is defined as the negative expectation of the conditional log-likelihood:

[0083]

[0084] If the conditional distribution P(R|Z R ) is assumed to be a normal distribution, it is expanded into the mean squared error form:

[0085]

[0086] where, R i and are the original and reconstructed medical examination data of the i-th sample.

[0087] The KL divergence loss measures the difference between the posterior distribution q(Z R |R) of the latent variable and the prior distribution P(Z R ), specifically:

[0088]

[0089] where, D KL represents the Kullback-Leibler divergence; d is the dimension of the latent variable, and are the mean and standard deviation of the i-th latent variable respectively.

[0090] By balancing the reconstruction loss and the KL divergence loss through the hyperparameter β, the total loss function L VAE has:

[0091] L VAE = L reconstruct + β·L KL

[0092] Through the above design, the present invention uses the variational autoencoder (VAE) structure to effectively compress and characterize the medical examination data, and converts it into low-dimensional latent variables Z R . The latent variables can not only compactly express the data features, but also contain causal information related to the task, which helps to improve the recommendation accuracy.

[0093] Further optimize the surrogate low-dimensional latent variable Z using contrastive learning R , capture the causal association between medical examination data R and patient characteristics X, that is, the derived contrast probability term By leveraging the similarities and differences between data samples, contrastive learning can enhance the representational ability of the surrogate variable for causal information.

[0094] In the contrastive learning: The positive samples are data pairs (X, R) from the same patient, where R is the medical examination data of the target patient and X is the corresponding patient characteristics. Since these two originate from the same patient and there is a direct causal association, they are defined as positive sample pairs. The goal is to enable the encoder to capture this causal association through contrastive learning and manifest it in the surrogate variable Z R . The negative samples are randomly selected combinations of patient characteristics and the medical examination data of the target patient, i.e., (X′, R); due to the randomly selected patient characteristics having no significant association with the positive sample data, they are defined as negative sample pairs. Contrastive learning calculates the contrast probability by pulling the distance of the positive sample pairs closer and pushing the distance of the negative samples farther Specifically, the InfoNCE loss function is adopted as the contrastive learning loss, and its definition is as follows:

[0095]

[0096] where τ is the temperature coefficient that controls the sensitivity of the positive and negative sample distributions in contrastive learning, usually a hyperparameter less than 1 (e.g., 0.1). N is the number of negative samples; i, j represent the positive and negative sample indices used to construct the negative sample set. sim(·, ·) is the similarity function.

[0097] The similarity function sim(·, ·) represents the similarity between the surrogate variable and patient characteristics and is calculated using the cosine similarity:

[0098]

[0099] The recommendation module is a collaborative filtering model, namely the factorization machine FM. It aims to generate personalized drug recommendation results based on patient characteristics and medical examination data. To achieve drug recommendation, the patient status latent variable Z X and the latent variable Z R are used as inputs to capture their potential feature interactions, and then precise drug recommendations are provided for each patient to best match their clinical needs; the output of the recommendation module is the conditional probability P(Y|R, X), representing the recommendation probability of each drug y ∈ Y; specifically as follows:

[0100] Y = f recommend (Z X , Z R ; φY )

[0101] where f recommend is the recommendation module; φ Y is the parameter of the recommendation module, and the output is a multi-class probability distribution, representing the recommendation probability P(Y|R,X) of each drug;

[0102] The recommendation loss L recommendation is used to measure the deviation between the drug recommendation probability predicted by the model and the actual drug label, and is defined as the multi-class cross-entropy loss:

[0103]

[0104] where N represents the number of patients; K is the number of drugs; i is the patient index; k is the drug category index, and y i,k represents the true label (usually 0 or 1) of patient i on drug category k; P(y i,k |R i ,X i ) represents the predicted probability of patient i on drug k;

[0105] The optimization objective of this loss function is to maximize the predicted probability of the model for the actual drugs used by the patients, while minimizing the probability of misrecommending the unused drugs.

[0106] Combining the reconstruction loss, the contrastive learning loss, and the recommendation loss, a joint optimization objective L total is constructed as follows:

[0107] L total = λ1·L reconstruct + λ2·L contrast + λ3·L recommendation

[0108] In the formula, λ1, λ2, and λ3 measure the importance weights of the reconstruction loss, the contrastive learning loss, and the recommendation loss respectively; by tuning these weights through experiments, the balance between different tasks can be achieved, ensuring that the model performance reaches the optimal level on each subtask. Through the optimization of the total loss L total , the accuracy of drug recommendation is improved, the influence of irrelevant noise and bias is eliminated, and more reliable personalized treatment suggestions are provided for patients.

[0109] The causal graph corresponding to the drug recommendation system is as shown in Figure 1As shown, the variable X represents patient characteristics, the variable R represents medical examination data (including laboratory tests, imaging examinations, etc.), Y represents the set of drugs suitable for the patient, and C represents unobserved covariates; the drug recommendation system aims to mine the causal relationship of X→Y and identify the causal effect P(Y|do(X)); due to the existence of unobserved covariates C, a spurious association between X and Y is generated through the path X←C→Y, which damages the recommendation performance; in clinical practice, the results of medical examinations can objectively reflect the patient's health status and provide important support for medical diagnosis. Therefore, there is a causal path: X→R→Y, where the set of drugs Y suitable for the patient is conditional on the patient characteristics X and the medical examination results R, and R serves as a mediating variable, providing strong evidence to support the treatment drug Y.

[0110] Based on the front-door criterion in causal inference theory, eliminate the bias and noise effects brought by unobserved covariates C; in the causal inference framework, the path X→R→Y is the front-door path, and the influence brought by covariates C is effectively eliminated through the front-door adjustment technique to accurately estimate the causal relationship of X→Y.

[0111] The following causal effect is obtained based on the front-door adjustment in causal inference:

[0112]

[0113] Among them, P(Y|do(X)) represents the causal effect of X on Y; do(X) represents an intervention operation on X; Z R is a low-dimensional latent variable substituting for R, used to approximately calculate the causal effect; P(Z R ) and P(X) are the probability distributions of the latent variable Z R and X; P(X|Z R ) represents the probability distribution of X given Z R , and P(Y|X,R) represents the probability distribution of Y given X and R.

[0114] Train the drug recommendation system to optimize the three probability terms P(Z R ), and P(Y|X,R), obtain the causal effect of X on Y, and eliminate the confounding effect generated by unobserved covariates C; specifically, approximately estimate the causal effect by optimizing the reconstruction loss, contrastive learning loss, and recommendation loss.

[0115] A drug recommendation method based on causal inference is implemented based on the aforementioned drug recommendation system based on causal inference, and specifically includes the following steps:

[0116] Step 1: Input patient characteristics X (such as age, gender, medical history, etc.), and project X into the patient state latent variable Z using a neural network encoder (such as a multi-layer perceptron) X ;

[0117] Step 2: Use an encoder-decoder structure (such as a variational autoencoder) to reconstruct the medical examination data R, and obtain the alternative low-dimensional latent variable Z of R R , and calculate the reconstruction loss;

[0118] Step 3: Calculate the contrastive learning loss, and optimize the low-dimensional latent variable Z through the contrastive learning loss (such as InfoNCE) R , ensuring that the generated alternative variables can learn the causal relationship of X→R;

[0119] Step 4: According to the patient status latent variable Z X and the latent variable Z of the medical examination R predict the appropriate drugs, and calculate the recommendation loss;

[0120] Step 5: Jointly optimize the reconstruction loss, contrastive learning loss, and recommendation loss to optimize the parameters of the drug recommendation system. By optimizing the causal relationship P(Y|do(X)) between the patient and the drug, improve the recommendation ability of the drug recommendation system.

[0121] Step 6: After the drug recommendation system is optimized, ignore the contrastive learning, and use the drug recommendation system to recommend drugs by combining the patient characteristics and medical examination data.

[0122] The drug recommendation system proposed in the embodiments of the present invention has been verified and applied on the open-source electronic health record dataset MIMIC-III. To ensure the effectiveness and consistency of the data, the dataset was preprocessed (filtering patients with only single admissions), and finally 5,449 patients and their 14,141 admission records were selected. In this embodiment, the Diagnosis and Procedures data in the MIMIC-III dataset were used to extract the disease characteristics and medical examination information of the patients as input features respectively, and a drug recommendation model was constructed and optimized. The drug recommendation effect was comprehensively evaluated through indicators such as Jaccard similarity score (Jaccard), F1 score (F1 Score), and the area under the Precision-Recall curve (PRAUC). The specific comparison is shown in Table 1 below.

[0123]

[0124] As can be seen from the results in the table, the embodiments of the present invention have significantly improved various indicators compared with other methods, especially achieving the best performance in Jaccard, F1 Score, and PRAUC. At the same time, on the basis of the factorization machine, the reconstruction loss and the contrastive learning loss are respectively added, and the performance is improved, indicating the effectiveness of each module in the invention. In summary, the drug recommendation system based on causal inference of the present invention can better integrate patient characteristics and medical examinations, recommend drug regimens more accurately, and has broad practical application value and promotion prospects.

[0125] The above description is only the preferred embodiments of the present disclosure and the description of the applied technical principles. Those skilled in the art should understand that the scope of the invention involved in the embodiments of the present disclosure is not limited to the technical solutions formed by the specific combination of the above technical features, and should also cover other technical solutions formed by any combination of the above technical features or their equivalent features without departing from the above inventive concept. For example, the technical solutions formed by mutually replacing the above features with the technical features (but not limited to) disclosed in the embodiments of the present disclosure that have similar functions.

Claims

1. A drug recommendation system based on causal inference, characterized in that: Use deep learning models, including patient state encoder, medical examination autoencoder, and recommendation module; The patient state encoder inputs the patient feature X, uses a multi-layer perceptron to map the high-dimensional patient feature X to a low-dimensional latent space, and generates a latent variable Z of the patient state. X , expressed as: Z X =f enc (X,θ X ) Among them, f enc is the neural network model of the patient state encoder, θ X are model parameters; The medical examination autoencoder uses an encoder-decoder structure to process the medical examination data R, extract key features and compress them into a low-dimensional space, and construct an alternative low-dimensional latent variable Z R ; The encoder-decoder structure adopts a variational autoencoder VAE structure, which specifically includes two modules, an encoder and a decoder, as shown below: Z R =f vae-enc (X,φ enc ) Among them, f vae-enc and f vae-dec is the encoder and decoder, φ enc and φ dec are encoder and decoder parameters, Medical examination data for reconstruction; In the variational autoencoder VAE, replace the low-dimensional latent variable Z R The prior distribution of is set to the standard multivariate normal distribution: P(Z R )=N(0,I) Where N(0,I) represents a multivariate normal distribution with a mean of 0 and a covariance matrix of the identity matrix I; The recommendation module is a collaborative filtering model, namely, a factorization machine FM; the patient status latent variable Z X and low-dimensional latent variables Z R As input, a drug recommendation is provided for each patient. The output of the recommendation module is the conditional probability P(Y|R,X), which represents the recommendation probability of each drug y∈Y; specifically, as follows: Y=f recommend (Z X ,Z R ;φ Y ) where f recommend It is a recommended module; Y It is the parameter of the recommendation module, and the output is a multi-classification probability distribution, which indicates the recommendation probability of each drug P(Y|R,X); The encoder inputs medical examination data to generate a replacement low-dimensional latent variable Z R The posterior distribution q(Z R |R) as the output of the encoder: q(Z R |R)=N(μ,diag(σ 2 )) Among them, μ and σ are alternative low-dimensional latent variables Z R The mean and standard deviation of , diag represents a diagonal matrix; The decoder uses a low-dimensional latent variable Z R As input, the encoded latent variables are mapped back to the original input space, and the original data is reconstructed to generate the reconstruction vector of the medical examination data R, that is, the reconstructed medical examination data The training goal of the variational autoencoder VAE is to minimize the total loss function L VAE , the total loss function L VAE Including reconstruction loss and KL divergence loss: The reconstruction loss is the negative expectation of the conditional log-likelihood: in, represents the expected value, logP(R|Z R ) is the log-likelihood term, which indicates that in the low-dimensional latent variable Z R The probability of observing data R under the condition of ; If the conditional distribution P(R|Z R ) is assumed to be normally distributed and expanded into the mean square error form: Among them, R i and is the original and reconstructed medical examination data of the i-th sample; The KL divergence loss is specifically: Among them, D KL represents the Kullback-Leibler divergence; d is the dimension of the latent variable, and are the mean and standard deviation of the ith latent variable, q(Z R |R) is the low-dimensional latent variable Z for a given medical examination R R The posterior distribution of By balancing the reconstruction loss and KL divergence loss through the hyperparameter β, the total loss function L VAE have: L VAE =L reconstruct +β·L KL Use contrastive learning to further optimize the alternative low-dimensional latent variable Z R , capturing the causal relationship between medical examination data R and patient characteristics X, that is, optimizing the comparison probability The posterior probability P(X|Z R ) and the prior probability P(X); In the contrastive learning, the positive sample is a data pair (X, R) from the same patient, where R is the medical examination data of the target patient and X is the corresponding patient feature; the negative sample is a combination of a randomly selected patient feature X′ and the medical examination data R of the target patient, i.e. (X', R); contrastive learning calculates the contrast probability by pulling the distance of the positive sample pair closer and pushing the distance of the negative sample farther. Specifically, the InfoNCE loss function is used as the contrastive learning loss, which is defined as follows: Among them, τ is the temperature coefficient, i,j represents the positive and negative sample indexes, and X i ,X j are the positive and negative samples of the index, is the surrogate latent variable of the medical examination R of patient i, N is the number of negative samples, and sim(·,·) is the similarity function; the similarity function sim(·,·) represents the similarity between the surrogate variable and the patient's characteristics, and is calculated using cosine similarity: In the recommendation module, the recommendation loss L is used recommendation Measures the deviation between the drug recommendation probability predicted by the model and the actual drug label, which is defined as the multi-classification cross entropy loss: Where N is the number of patients; K is the number of drugs; i is the patient index; k is the drug category index, y is i,k represents the true label of patient i on drug k; P(y i,k |R i ,X i ) represents the predicted probability of patient i on drug k; Combining reconstruction loss, contrastive learning loss and recommendation loss, construct a joint optimization target L total : L total =λ1·L reconstruct +λ2·L contrast +λ3·L recommendation Where λ1, λ2, and λ3 respectively measure the importance weights of reconstruction loss, contrastive learning loss, and recommendation loss; In the causal graph corresponding to the drug recommendation system, variable X represents patient characteristics, variable R represents medical examination data, Y represents the set of drugs suitable for the patient, and C represents unobserved covariates; the drug recommendation system aims to explore the causal relationship from X to Y and identify the causal effect P(Y|do(X)); there is a causal path: X→R→Y, in which the set of drugs suitable for the patient Y is conditional on the patient characteristics X and the medical examination results R; Based on the front-door adjustment in causal inference, the following causal effect is obtained: Among them, P(Y|do(X)) represents the causal effect of X on Y; do(X) represents the intervention operation on X; Z R is a low-dimensional latent variable to replace R, P(Z R ) and P(X) are low-dimensional latent variables Z R and the probability distribution of X; P(X|Z R ) indicates that given Z R P(Y|X,R) represents the probability distribution of Y given X and R. Train the drug recommendation system to optimize the three probability terms P(Z R ), and P(Y|X,R) to obtain the causal effect of X on Y and eliminate the confounding effect of the unobserved covariate C. Specifically, the causal effect is approximately estimated by optimizing the above reconstruction loss, contrastive learning loss, and recommendation loss.

2. A drug recommendation method based on causal inference, implemented by a drug recommendation system based on causal inference as claimed in claim 1, characterized in that: The specific steps include: Step 1: Input patient feature X and use the neural network encoder to project X into the patient state latent variable Z X ; Step 2: Use the encoder-decoder structure to reconstruct the medical examination data R and obtain the alternative low-dimensional latent variable Z of R R , and calculate the reconstruction loss; Step 3: Calculate the contrastive learning loss and optimize the low-dimensional latent variable Z through contrastive learning loss R ; Step 4: According to the patient status latent variable Z X and the latent variable Z of medical examination R Predict appropriate medications and calculate recommendation losses; Step 5: Jointly optimize the reconstruction loss, contrastive learning loss, and recommendation loss to optimize the parameters of the drug recommendation system, and improve the recommendation ability of the drug recommendation system by optimizing the causal relationship P(Y|do(X)) between patients and drugs; Step 6: After the drug recommendation system is optimized, ignore contrastive learning and use the drug recommendation system to make drug recommendations by combining patient characteristics and medical examination data.

Citation Information

Patent Citations

  • Out-of-domain intention detection method based on comparative learning and variational auto-encoder

    CN117744665A

  • Drug recommendation method based on causal comparison and message passing

    CN118969175A