Immunosuppressant medication recommendation system based on deep reinforcement learning

Through the deep reinforcement learning medication recommendation system, the problem of individual differences in immunosuppressant medication regimens is solved, personalized medication strategy optimization is achieved, the accuracy and safety of medication regimens are improved, the risks are reduced, and it is suitable for patients at different treatment stages.

CN120376033BActive Publication Date: 2025-09-05SICHUAN ACADEMY OF MEDICAL SCI SICHUAN PROVINCIAL PEOPLES HOSPITAL
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202510864034.3
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-06-26
Publication Date
2025-09-05
Estimated Expiration
2045-06-26

AI Technical Summary

Technical Problem

In the existing technology, the dosage regimen of immunosuppressants is difficult to accurately adjust according to the individual differences of patients, resulting in substandard or excessive blood drug concentrations, increasing the risk of medication for patients and the difficulty of formulating dosage regimens.

Method used

A medication recommendation system based on deep reinforcement learning is adopted. By building a medication recommendation model, combining patient characteristics and blood drug concentration prediction model, using reward mechanism to optimize medication regimen, simulating clinical expert decision-making, and generating personalized medication strategies.

Benefits of technology

It improves the accuracy and safety of the dosing regimen, reduces the risks caused by fluctuations in blood drug concentrations, reduces the workload of medical staff, enhances the stability and adaptability of the dosing regimen, and is suitable for individualized guidance at different treatment stages.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120376033B_ABST
    Figure CN120376033B_ABST
Patent Text Reader

Abstract

The present application discloses an immunosuppressant dosing recommendation system based on deep reinforcement learning, which relates to the field of medical drug management technology, and includes a model construction module, an information acquisition module and a scheme recommendation module; the model construction module includes a strategy submodule, an environment submodule, a reward submodule and a training submodule; the strategy submodule constructs a dosing recommendation model, and the dosing recommendation model is used to output a dosing recommendation scheme after receiving patient characteristics; the environment submodule predicts the patient's blood drug concentration after taking the drug based on the patient characteristics and the dosing recommendation scheme; the reward submodule evaluates the patient's blood drug concentration after taking the drug to obtain a reward or punishment value; the training submodule trains the dosing recommendation model based on the reward or punishment value, and sends the dosing recommendation model to the scheme recommendation module when the training termination condition is met; the scheme recommendation module inputs the target characteristics of the target patient sent by the information acquisition module into the dosing recommendation model to obtain a target dosing scheme, so as to reduce the difficulty of decision-making on the immunosuppressant dosing scheme.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present application relates to the technical field of medical drug management, and specifically to an immunosuppressant drug administration recommendation system based on deep reinforcement learning. Background Art

[0002] Immunosuppressants refer to drugs that have an inhibitory effect on the body's immune response. For example, they are widely used to treat immune rejection reactions after solid organ transplantation and hematopoietic stem cell transplantation. The dosing regimen of immunosuppressants is a key factor affecting the patient's treatment effect. When the patient's blood drug concentration after taking the drug is lower than the effective concentration, the patient's risk of rejection reaction increases; when the patient's blood drug concentration after taking the drug is not less than the toxic concentration, the patient's risk of adverse events such as infection increases. The narrow therapeutic window of immunosuppressants makes it difficult to decide on the patient's immunosuppressant dosing regimen.

[0003] In current clinical practice, dosage is primarily determined by patient weight. However, due to individual differences in age, transplant type, time since surgery, and concomitant medications, a minority of patients can achieve target blood drug concentrations when a dosing regimen based solely on weight is formulated. Most patients will still require secondary or even multiple adjustments based on post-dose blood drug concentration test results. This increases the risk of medication use and the difficulty of formulating a dosing regimen with each adjustment.

[0004] Therefore, there is an urgent need for an immunosuppressant dosing recommendation system to solve the technical problem of the difficulty in making dosing regimen decisions. Summary of the Invention

[0005] The present invention provides an immunosuppressant dosing recommendation system based on deep reinforcement learning, which optimizes the dosing recommendation model through reinforcement learning training to assist medical staff in formulating immunosuppressant dosing plans.

[0006] The present invention seeks protection for an immunosuppressant dosing recommendation system based on deep reinforcement learning, comprising a model building module, an information acquisition module and a regimen recommendation module; the model building module comprises a strategy submodule, an environment submodule, a reward submodule and a training submodule; the strategy submodule constructs a dosing recommendation model, which is used to output a dosing recommendation regimen after receiving patient characteristics; the environment submodule predicts the patient's blood drug concentration after taking the drug based on the patient characteristics and the dosing recommendation regimen; the reward submodule evaluates the patient's blood drug concentration after taking the drug to obtain a reward or punishment value; the training submodule trains the dosing recommendation model based on the reward or punishment value, and sends the dosing recommendation model to the regimen recommendation module when the training termination condition is met; the regimen recommendation module inputs the target characteristics of the target patient sent by the information acquisition module into the dosing recommendation model to obtain a target dosing regimen.

[0007] In one embodiment of the present application, the recommended dosing regimen includes the recommended dosage of each drug in all administration methods, and the blood drug concentration of the patient after taking the drug includes the blood drug concentration change trend within a preset time period.

[0008] In one embodiment of the present application, the recommended dosing regimen also includes the recommended frequency of each drug in all dosing methods.

[0009] In one embodiment of the present application, the reward submodule includes a concentration constraint submodule and a reward and punishment calculation submodule; the concentration constraint submodule includes a standard concentration acquisition unit; the standard concentration acquisition unit matches the corresponding target blood drug concentration range according to the patient characteristics; the reward and punishment calculation submodule calculates the concentration component according to the proportion of time that is consistent with the target blood drug concentration range in the blood drug concentration change trend within a preset time period, and generates corresponding reward and punishment values ​​according to the concentration component.

[0010] In one embodiment of the present application, the standard concentration acquisition unit also includes a rule matching subunit, a reference acquisition subunit and a standard generation subunit; the rule matching subunit generates a rule concentration range based on the concentration matching results of the patient characteristics in the medication rule set; the reference acquisition subunit queries the database based on the patient characteristics to obtain a reference patient, and obtains a reference concentration range based on the actual blood drug concentration of the reference patient in the corresponding treatment stage; the reference patient includes a historical patient whose patient portrait is the same as the patient portrait of the corresponding patient and who has taken normal medication in the corresponding treatment stage; the standard generation subunit obtains the target blood drug concentration range based on the rule concentration range and all reference concentration ranges.

[0011] In one embodiment of the present application, the standard generation sub-unit also takes the union of the rule concentration range and all reference concentration ranges as the target blood drug concentration range, divides the target blood drug concentration range into several target sub-ranges according to the overlap between the rule concentration range and all reference concentration ranges, and obtains the range weight according to the number of sets of the rule concentration range and all reference concentration ranges that include each target sub-range. The target sub-ranges with more sets have corresponding range weights with larger values; the reward and punishment calculation module calculates the concentration component according to the proportion of time that the blood drug concentration change trend conforms to each target sub-range and the corresponding range weight.

[0012] In one embodiment of the present application, the reference acquisition subunit calculates the feature similarity between the patient portrait of a historical patient who took the drug normally in each corresponding treatment stage and the patient portrait of the corresponding patient, and uses the historical patient whose feature similarity is not less than the similarity threshold as the reference patient; the calculation method of the concentration component includes:

[0013] ;

[0014] ;

[0015] Where S represents the concentration component, φ and f are mapping functions, M represents the number of reference subranges, ω j represents the range weight of the j-th target subrange, p j represents the proportion of time when the trend of blood drug concentration changes conforms to the jth target sub-range, ω0 represents the out-of-range weight, whose value is less than 0, p0 represents the proportion of time when the trend of blood drug concentration changes does not conform to the target blood drug concentration range, N represents the number of sets of the jth target sub-range, s k Represents the feature similarity of the kth concentration range containing the jth target subrange.

[0016] In one embodiment of the present application, the reward submodule also includes a safety constraint submodule, which obtains a safety component based on the medication safety matching results of the medication recommendation scheme and the medication rule set, and the reward and punishment calculation submodule generates corresponding reward and punishment values ​​based on the concentration component and the safety component.

[0017] This application has the following beneficial effects:

[0018] 1. The medication recommendation model, acting as an intelligent agent, learns the optimal strategy for recommending immunosuppressant medication regimens to patients within the environment. Through reinforcement learning and a trial-and-error approach, it improves the medication recommendation model's decision-making capabilities, simulating the decision-making methods of clinical experts. The reward function, serving as the basis for training and updating the medication recommendation model, enables it to distinguish between optimal and actual medication regimens based on the medication recommendations generated by clinical experts. Specifically, the medication recommendation regimen that gives the maximum reward value is considered the optimal medication regimen for the patient. This gradually optimizes the target recommendation regimen, reducing the difficulty for clinicians in developing immunosuppressant medication regimens for their patients.

[0019] 2. The peaks and troughs in the complete blood drug concentration trend after medication administration obtained by the blood drug concentration prediction model reflect the individual differences in drug absorption, distribution, and metabolism among different patient groups. Compared with a single blood drug concentration value given at a fixed preset time point, it more realistically reflects the dynamic process of the drug in the body, reduces the probability of situations where the blood drug concentration temporarily reaches the target but there is actually insufficient dosage or toxicity risk, and improves the safety, stability, and accuracy of drug administration.

[0020] 3. The strategy submodule not only gives the recommended dose of immunosuppressants but also generates the recommended frequency, thus realizing a dose and frequency combination scheme based on medication cycles. This technical solution significantly reduces the frequency of manual operations and workload of medical staff. In particular, for formulating medication plans for patients with stable conditions or in the late stages of recovery, the cyclical plan is conducive to improving clinical work efficiency and operational convenience.

[0021] 4. Introducing evaluation indicators of medication safety into the reward submodule enables the strategy submodule to not only focus on achieving blood drug concentration targets when generating dosing recommendations, but also proactively identify dosing recommendations that may pose unsafe medication risks, guiding the strategy submodule to recommend medication plans with better safety, thereby enhancing the clinical usability and patient safety of the recommendation results.

[0022] 5. The reward submodule designed based on time series can more realistically reflect the metabolic dynamics of drugs in patients and the triggering mechanism of toxic and side effects. It comprehensively evaluates the length of time that patients' blood drug concentrations stay in each risk interval, thereby realizing a fine-grained analysis of patients' drug exposure process. This method reduces the scoring misjudgment caused by occasional abnormalities such as noise interference and patient status fluctuations, and misjudges accidental blood drug concentration fluctuations as rejection reactions or poisoning reactions, thereby enhancing the stability of the medication recommendation model.

[0023] 6. The concentration acquisition unit matches a more accurate target blood drug concentration range based on the individual differences of the patients. On the one hand, it not only improves the system's evaluation of the safety and rationality of the medication regimen, but also reduces the difficulty of collecting the medication rule set. In particular, when immunosuppressants are used in combination, the coverage of the medication rule set becomes more difficult. Introducing the individual difference matching of historical patients expands the use scenarios of the medication recommendation model to combination medication, reducing the limitations brought by the evaluation of a single drug.

[0024] 7. The reward and punishment calculation module constructs personalized blood drug concentration risk stratification intervals through an interval overlapping voting mechanism. The design method of this reward function takes into account the differences in control targets at different stages of the immune transplantation treatment process. Under the premise of ensuring the safety of patient medication, this design method enables the blood drug concentration evaluation to be dynamically adjusted with the treatment stage. The reference concentration range reflects the acceptable range of blood drug concentration fluctuations for patients in the current treatment stage. In particular, when a larger number of reference concentration ranges overlap in intervals outside the regular concentration range, the risk of rejection and / or poisoning reactions for the patient is lower when the blood drug concentration is within this concentration range, thereby enhancing the adaptability and individualized guidance capabilities of the model in actual clinical applications. For example, in the early stage after transplantation surgery, the patient's body is at high risk of rejection of the transplanted organ, and the control of blood drug concentrations needs to be more stringent. The patient's medication compliance is required to be high. The difference between the intersection and union of all reference ranges is small, and the target concentration range should be tight to ensure the immunosuppressive effect; in the later stage of recovery, the body gradually establishes adaptability to the transplanted organ, and even short-term fluctuations in blood drug concentrations are not likely to cause rejection. At this stage, more attention should be paid to the risk of long-term cumulative toxicity. The difference between the intersection and union of all reference ranges increases, and the target blood drug concentration range is appropriately relaxed to balance efficacy and toxicity.

[0025] 8. The reward and punishment calculation module uses reference patients who are highly similar to the patient's patient portrait to have a greater influence on the final risk interval division result, while reference patients with lower similarity have a smaller influence weight, thereby reducing the interference of non-representative historical data on the risk judgment result. BRIEF DESCRIPTION OF THE DRAWINGS

[0026] The accompanying drawings, which are incorporated in and constitute a part of this specification, illustrate embodiments consistent with the present disclosure and, together with the description, serve to explain the principles of the present disclosure.

[0027] Figure 1 This is a schematic diagram of the data transmission process of the immunosuppressant dosage recommendation system based on deep reinforcement learning involved in an embodiment of the present application;

[0028] Figure 2 A schematic diagram of the structure of the immunosuppressant drug administration recommendation system involved in an embodiment of the present application;

[0029] Figure 3 Schematic diagram of data transmission in a concentration constraint module according to an embodiment of the present application;

[0030] Figure 4 Schematic diagram of another structure of the immunosuppressant drug administration recommendation system involved in an embodiment of the present application;

[0031] Figure 5 This is a schematic diagram of the range weights involved in the embodiments of this application;

[0032] Figure 6 A schematic diagram of the structure of an electronic device involved in an embodiment of the present application;

[0033] Symbols in the figure: x1, x2, x3, x4, x5 - example values ​​of blood drug concentration, L4 - regular concentration range, L1 - first reference concentration range, L2 - second reference concentration range, L3 - third reference concentration range. DETAILED DESCRIPTION

[0034] The present invention provides an immunosuppressant dosing recommendation system based on deep reinforcement learning. To make the above-mentioned objects, features, and advantages of this application more readily apparent, the technical solutions in the embodiments of this application will be clearly and completely described below in conjunction with the accompanying drawings of the embodiments of this application. It should be understood that the described embodiments are only some of the embodiments of this application, and not all of them. The components of the embodiments of this application generally described and illustrated in the drawings herein may be arranged and designed in a variety of different configurations. Therefore, reference to the terms "one embodiment," "some embodiments," "implementation," "embodiment," "illustrative embodiment," "example," "specific example," or "some examples," etc., in the following detailed description of the embodiments of this application provided in the drawings, is not intended to limit the scope of the claimed application, but merely indicates that the specific features, structures, or characteristics described in conjunction with such embodiment or example are included in at least one embodiment or example of the present invention. Furthermore, the specific features, structures, or characteristics described may be combined in any suitable manner in any one or more embodiments or examples. All other embodiments derived by those skilled in the art based on the embodiments of this application without inventive effort are within the scope of protection of this application.

[0035] It should be noted that similar reference numerals and letters represent similar items in the following figures. Therefore, once an item is defined in one figure, it does not need to be further defined or explained in subsequent figures. At the same time, in the description of this application, relational terms such as "first" and "second" are only used to distinguish one entity or operation from another entity or operation, and do not necessarily require or imply any actual relationship or order between these entities or operations.

[0036] The present invention claims protection for an immunosuppressant drug administration recommendation system based on deep reinforcement learning, Figure 1 and attached Figure 2 As shown, it includes a model building module, an information acquisition module and a solution recommendation module.

[0037] It should be noted that the model construction module includes a strategy submodule, an environment submodule, a reward submodule and a training submodule. The strategy submodule is used to construct a medication recommendation model. The input of the medication recommendation model includes patient characteristics, and the output includes a medication recommendation scheme. The environment submodule is used to construct and train a blood drug concentration prediction model. The blood drug concentration prediction model is used to predict the blood drug concentration of the patient after taking the medication according to the corresponding medication scheme. The reward submodule is used to evaluate the blood drug concentration of the patient after taking the medication, and generate an evaluation result based on a preset reward function, that is, to obtain the corresponding reward and punishment value. The training submodule obtains a training data set uploaded by the user for training the medication recommendation model. At the same time, the training submodule also uses reinforcement learning to update the medication recommendation model of the strategy submodule based on the reward value and the value function. When the preset training termination condition is met, the trained medication recommendation model is sent to the scheme recommendation module.

[0038] It should be noted that the information acquisition module is used to obtain characteristics of the target patient, obtain target characteristics, and send the target characteristics to the regimen recommendation module. The regimen recommendation module inputs the received target characteristics into a trained medication recommendation model to output a medication regimen for the target patient, i.e., the target medication regimen.

[0039] It should be noted that the strategy submodule can be pre-configured with a machine learning model, such as a neural network model, to generate the medication recommendation model based on user input requirements. For example, user-specified features can be used as input features of the medication recommendation model, and the output and objective function of the medication recommendation model can be generated based on the user-specified prediction task. The strategy submodule can also directly access the medication recommendation model constructed by the user.

[0040] It should be noted that the environmental submodule is used to simulate the patient's blood drug concentration after medication. The blood drug concentration prediction model can be constructed using a neural network model. The input of this blood drug concentration prediction model is the patient's characteristics and the immunosuppressant dosage regimen, and the output is the patient's blood drug concentration after using the dosage regimen. The blood drug concentration prediction model can be trained based on the historical patient characteristics, actual dosage regimen, and blood drug concentration after medication in historical medical records as a prediction training dataset.

[0041] It should be noted that the reward submodule is used to evaluate patients' blood drug concentrations. Reference materials such as literature, clinical diagnosis and treatment guidelines, drug instructions, and medication standards can be systematically searched. The search results are evaluated using the Cochrane Systematic Review Tool and the Guidelines Research and Evaluation System II (AGREE II) to ensure the quality and applicability of the reference materials and exclude reference materials of low quality or that do not meet the screening criteria. Patient medication rules are extracted from the screened reference materials and combined with clinical expert opinions to obtain a preliminary medication rule set. Each medication rule in the preliminary medication rule set is evaluated using the Delphi expert consultation method. Specifically, a questionnaire is designed based on all medication evaluation rules and distributed to an expert panel. After multiple rounds of questionnaire completion, collection, collation, and analysis, the expert panel reaches a consensus. The final rule results are then collated and refined to obtain the medication rule set. The reward submodule matches the medication rule set according to the patient characteristics to obtain the target blood drug concentration range of the corresponding patient; compares the patient's output in the environment submodule with the target blood drug concentration range, and if the patient's output in the environment submodule falls within the target blood drug concentration range, the reward submodule gives a reward; if the patient's output in the environment submodule does not fall within the target blood drug concentration range, the reward submodule gives a penalty.

[0042] It should be noted that the training submodule obtains a recommended training dataset input by the user, and the recommended dataset is used to train the medication recommendation model. The training submodule can train the medication recommendation model based on a value function such as a Q-Learning algorithm, a Sarsa algorithm, or a policy gradient algorithm.

[0043] It should be noted that the training submodule first trains the medication recommendation model based on the imitation learning method. Imitation learning can imitate and learn the medication regimen of clinical experts to reduce the learning difficulty and learning efficiency of the medication recommendation model in the initial learning stage.

[0044] It should be noted that the patient characteristics include basic characteristics and diagnostic characteristics, but are not limited to these. The basic characteristics include gender, age, past medical history, family medical history, etc., but are not limited to these. The diagnostic characteristics include the name of the disease, treatment plan, and pre-drug examination results, etc., but are not limited to these. The treatment method includes the time series of all treatment plans used after the patient is admitted to the hospital. For example, the treatment plan for a liver transplant patient includes drug A on the first day of admission, drug A and drug B on the second day, drug B on the third day, liver transplant surgery on the fourth day, drug B on the fifth day, and drug C on the sixth day, which can be recorded as (A, AB, B, liver transplant surgery, B, C). The pre-drug examination results include blood drug concentrations before medication, but are not limited to these.

[0045] It should be noted that if the currently formulated dosing regimen is for the patient's first use of an immunosuppressant, the corresponding pre-drug blood drug concentration for the patient can be defaulted to 0.

[0046] In a first embodiment, the strategy submodule is used to generate a recommended dosage for each pre-set medication based on patient characteristics. The environment submodule's input includes the patient characteristics and the recommended dosing regimen, and its output includes the patient's blood drug concentration at a pre-set time point after using the recommended dosing regimen. For example, the blood drug concentration at one hour after dosing.

[0047] In this embodiment, the reward submodule includes a concentration constraint module and a reward and punishment calculation module. The concentration constraint module includes a standard concentration acquisition unit. The standard concentration acquisition unit has a pre-built medication rule set built in, matches the medication rule set according to the patient's characteristics, and generates a corresponding target blood drug concentration range based on all successfully matched medication rule sets. The reward and punishment calculation module is used to generate corresponding reward and punishment values ​​based on the comparison results of the patient's blood drug concentration after medication and the target blood drug concentration range. For example, if the patient's blood drug concentration meets the target blood drug concentration range, the probability of the patient having rejection and toxic reactions is considered to be low risk, and the reward and punishment calculation module gives a reward value of 1. If the patient's blood drug concentration does not meet the target blood drug concentration range, the probability of the patient having rejection and toxic reactions is considered to be high risk, and the reward and punishment calculation module gives a penalty value of -1.

[0048] It should be noted that the number of preset drugs can be one or more. The preset drugs can be preset based on the immunosuppressive drugs of the medical institution. This application does not further limit the preset drugs and quantities.

[0049] It should be noted that the preset time point can be obtained by presetting, or can be used as an input feature of the blood drug concentration prediction model to obtain a user-specified time point, or other feasible implementation methods. This application does not further limit the setting method and specific value of the preset time point.

[0050] In this embodiment, the reward submodule also includes a safety constraint submodule, but it may not be limited to this. The safety constraint submodule obtains a safety component based on the matching result between the recommended medication regimen and the medication regimen rule. The safety component is used to evaluate the safety level of the medication recommendation regimen. The safety level assessment criteria include whether the recommended medication regimen of the immunosuppressant interacts with the patient's historical medication or other medications, whether there are drug interactions when different immunosuppressant drugs are used at the same time, etc., but it may not be limited to this. If the recommended medication regimen will produce adverse interactions with other drugs currently being used by the patient, the reward and punishment calculation submodule will give a penalty value. The specific value of the penalty value can be obtained based on the degree of adverse effects of the interaction between drugs. For example, if the interaction between drugs is to reduce the efficacy of immunosuppressants or other drugs, a penalty value of -0.1 is given; if the interaction between drugs is to cause a risk of death, a penalty value of -1 is given.

[0051] In this embodiment, the reward and punishment calculation module can obtain the result of linear weighted calculation of the concentration component and the safety component according to importance.

[0052] In this embodiment, the dosing recommendation model includes several dosage recommendation sub-models, each of which outputs the mean and variance of the immunosuppressant dosage distribution for the patient's corresponding preset medication. The strategy sub-module considers the mean of the immunosuppressant dosage distribution as the best estimate of the dosage. After constructing a normal distribution of the patient population profile based on the mean and variance output by the dosing recommendation model, the recommended dosage for the patient corresponding to the preset medication is sampled from the normal distribution based on the probability density.

[0053] It should be noted that, taking cyclosporine and tacrolimus as an example, the medication recommendation model includes a first dose recommendation sub-model and a second dose recommendation sub-model, but it is certainly not limited to this. The first dose recommendation sub-model is used to generate the mean and variance of the patient's use of cyclosporine based on the patient's characteristics. The second dose recommendation sub-model is used to generate the recommended dose of tacrolimus for the patient based on the patient's characteristics. The strategy sub-module constructs a normal distribution of the patient's use of cyclosporine and a normal distribution of the patient's use of tacrolimus, respectively, and samples the patient's recommended dose of cyclosporine and the recommended dose of tacrolimus from the corresponding normal distributions. Different patients have individual differences, which leads to different therapeutic effects of the same dose on different patients. The probability distribution of the patient's medication dose is predicted by the medication recommendation model to simulate the dosage preferences of the corresponding patient group when using each preset drug. The action prediction in the continuous space is provided based on probability density sampling to achieve adaptation to individual differences in patients.

[0054] In the second embodiment, the difference from the first embodiment is that the output of the environment submodule includes the changing trend of the patient's blood drug concentration within a preset time period after using the recommended dosing regimen. For example, the changing trend of blood drug concentration within 5 hours after taking the medicine. After the standard concentration acquisition unit of the reward submodule obtains the corresponding target blood drug concentration range based on the patient's characteristics, the reward and punishment calculation module generates corresponding reward and punishment values ​​based on whether there is a moment in the blood drug concentration changing trend that does not conform to the target blood drug concentration range. If the patient's blood drug concentration changing trends are all in line with the target blood drug concentration range, it means that the probability of the patient having an abnormal medication event is small, and the reward and punishment calculation module gives a reward value of 1. If there is a moment in the blood drug concentration changing trend that does not conform to the target blood drug concentration range, it means that the patient has a high probability of an abnormal medication event, and the reward and punishment calculation module gives a penalty value of -1.

[0055] It should be noted that the preset time period can be obtained by presetting, or can be used as an input feature of the blood drug concentration prediction model to obtain a user-specified time period, or other feasible implementation methods. This application does not further limit the setting method and specific duration of the preset time period.

[0056] It should be noted that the blood drug concentration prediction model is constructed based on the long short-term memory neural network model. The training data set can be obtained from historical medical records, and the blood drug concentration values ​​of historical patients at different time points are used as the time series characteristics of blood drug concentration.

[0057] In the third embodiment, the difference from the second embodiment is that the strategy submodule is also used to generate the recommended frequency of each preset drug for the patient based on the patient's characteristics, and the strategy submodule generates the recommended dosing plan based on the recommended dosage and recommended frequency of each preset drug.

[0058] In this embodiment, the medication recommendation model further includes a first frequency recommendation sub-model and a second frequency recommendation sub-model, although this is not limited thereto. The first frequency recommendation sub-model is used to generate a recommended frequency of cyclosporine use for the patient based on patient characteristics. The second frequency recommendation sub-model is used to generate a recommended frequency of tacrolimus use for the patient based on patient characteristics.

[0059] In this embodiment, the reward and punishment calculation module obtains the concentration component according to the time proportion that meets the target blood drug concentration range in the corresponding blood drug concentration change trend output by the environment submodule.

[0060] It should be noted that, the smaller the proportion of time within the preset time period that is within the target blood drug concentration range, the smaller the value of the corresponding concentration component; the greater the proportion of time within the preset time period that is within the target blood drug concentration range, the greater the value of the corresponding concentration component.

[0061] In this embodiment, the value range of the concentration component is preset to [-1, 1]. If the value of the concentration component is greater than 0, it means that the reward and punishment calculation module gives a corresponding reward value on the concentration component. For example, when the value of the concentration component is 1, it means that the patient's blood drug concentration change trend in the preset time period is consistent with the target blood drug concentration range. If the value of the concentration component is less than 0, it means that the reward and punishment calculation module gives a corresponding penalty value on the concentration component. For example, when the value of the concentration component is -1, it means that the patient's blood drug concentration change trend in the preset time period is not consistent with the target blood drug concentration range. The calculation method of the concentration component S1 can be:

[0062] ;

[0063] ;

[0064] Wherein, T represents the total number of moments in the preset time period, i represents the sequence number of the moment, and c i represents the blood drug concentration value at the i-th moment in the output of the environmental submodule, C represents the target blood drug concentration range, I is an indicative function, and A represents the input condition of the indicative function I.

[0065] It should be noted that the above-described concentration component calculation method uses a linear mapping approach to convert the proportion of time within the preset time period that meets the target blood drug concentration range into a corresponding reward or penalty value. Nonlinear methods such as tangent function mapping or z-score mapping can also be used to calculate reward or penalty values. This embodiment does not further limit the specific mapping method.

[0066] In the fourth embodiment, refer to the attached Figure 3 and attached Figure 4 As shown, the difference from the third embodiment is that the standard concentration acquisition unit also includes a rule matching subunit, a reference acquisition subunit and a standard generation subunit, of course, it is not limited to this. The rule matching subunit is used to generate a rule concentration range based on the concentration matching results of the patient characteristics in the medication rule set. The reference acquisition subunit queries the database based on the patient characteristics to obtain a reference patient, and obtains a reference concentration range based on the actual blood drug concentration of the reference patient in the corresponding treatment stage. The reference patient includes a patient profile that is the same as the patient profile of the corresponding patient and a historical patient who has taken the drug normally in the corresponding treatment stage. The standard generation subunit obtains the target blood drug concentration range based on the rule concentration range and all reference concentration ranges.

[0067] It should be noted that the medication rule set includes recommended blood drug concentration ranges for patients with different characteristics, such as the recommended blood drug concentration for adult patients within one week after surgery and the recommended blood drug concentration for patients with abnormal liver function. The rule matching subunit obtains corresponding concentration matching results based on patient characteristics and generates corresponding rule concentration ranges based on the concentration matching results.

[0068] In this embodiment, the reference acquisition subunit selects historical patients who did not experience immune reactions or toxic reactions during the treatment phase with the corresponding patient from historical medical records to obtain candidate patients, i.e., the candidate patients are the corresponding historical patients who took the medication normally. The reference acquisition subunit calculates the feature similarity between the patient profile of each candidate patient and the patient profile of the corresponding patient, and selects the candidate patients whose feature similarity is not less than the similarity threshold as the reference patients.

[0069] It should be noted that the patient profile is constructed by collecting data on the patient's basic information, disease status, medical behavior, and treatment process to construct a personalized information model of the patient. The neural network model can be used to extract the feature vector of the patient profile of the candidate patient and the feature vector of the patient profile of the corresponding patient. The feature similarity is calculated based on cosine similarity. The reference acquisition subunit is used to obtain the corresponding reference concentration range based on the maximum and minimum blood drug concentrations of each reference patient at the corresponding treatment stage.

[0070] It should be noted that the similarity threshold can be obtained by presetting. The target blood drug concentration range can be obtained according to the intersection of the rule concentration range and all the reference concentration ranges.

[0071] In this embodiment, the target blood drug concentration range of the standard generation subunit is regarded as a reference range according to the rule concentration range and all the reference concentration ranges, which is recorded as L, that is, L={L1, L2, ..., L i ,…,L n}, where L i represents the i-th reference range, and n represents the number of reference ranges. The union of all reference ranges is used as the target blood drug concentration range. The standard generation subunit also includes dividing the target blood drug concentration range into a number of target sub-ranges based on the overlap of all reference ranges, and obtaining the range weight of the corresponding target sub-range based on the number of sets of all reference ranges that include each target sub-range. The target sub-range with more sets has a corresponding larger range weight.

[0072] It should be noted that the value of n is 4, as an example. Figure 5As shown, the values ​​of the blood drug concentration example values ​​x1, x2, x3, x4 and x5 decrease in sequence. When the value of n is 4, the reference range set L={L1, L2, L3, L4} is described, and the regular concentration range is recorded as L4, that is, L4=[x2, x4]. The first reference concentration range is recorded as L1, that is, L1=[x1, x3], the second reference concentration range is recorded as L2, that is, L2=[x2, x3], and the third reference concentration range is recorded as L3, that is, L3=[x2, x5]. It can be seen that the target blood drug concentration range is [x1, x5]. According to the overlap of all reference ranges, the target blood drug concentration range can be divided into 4 target sub-ranges, namely [x1, x2], [x2, x3], [x3, x4] and [x4, x5]. The target sub-range [x1, x2] corresponds to being only included in the first reference concentration range L1, and the target sub-range [x4, x5] corresponds to being only included in the third reference concentration range L3. Then the number of sets corresponding to the target sub-ranges [x1, x2] and [x4, x5] is 1, and the corresponding range weight has the smallest value. The target sub-range [x3, x4] corresponds to being included in both the regular concentration range L4 and the third reference concentration range L3. Then the number of sets corresponding to the target sub-range [x3, x4] is 3. Similarly, the number of sets corresponding to the target sub-range [x2, x3] is 4, and the corresponding range weight has the largest value. The range weight of the corresponding target sub-range can be calculated through normalization. That is, the range weights of the target sub-ranges [x1, x2] and [x4, x5] are 0.125, the range weight of the target sub-range [x3, x4] is 0.25, and the range weight of the target sub-range [x2, x3] is 0.5.

[0073] In this embodiment, the reward and punishment calculation module calculates the concentration component S in the following manner:

[0074] ;

[0075] Where φ is a mapping function, which can be converted to the preset value range of reward and punishment values ​​by linear mapping, tangent function mapping, and z-score mapping; M represents the number of reference sub-ranges; ω j represents the range weight of the j-th target sub-range; p j It represents the proportion of time when the trend of blood drug concentration changes conforms to the jth target sub-range; ω0 represents the out-of-range weight, and p0 represents the proportion of time when the trend of blood drug concentration changes does not conform to the target blood drug concentration range.

[0076] It should be noted that ω0 is used to set a penalty value for the time that is not within the target blood drug concentration range, that is, its value is less than any ω jWhen the value of ω0 is 0, the dosing recommendation model only focuses on the proportion of time that the target blood concentration is reached. When the value of ω0 is negative, the dosing recommendation model penalizes the proportion of time that is not within the target blood concentration range, thereby reflecting the risk of the dosing recommendation plan, improving the differentiation between the risk of using medication and the risk of not using medication, and guiding the strategic learning direction of the dosing recommendation model.

[0077] In this embodiment, the range weight ω j The calculation method is:

[0078] ;

[0079] Among them, f is the mapping function, which can be converted to the preset range of range weight by linear mapping, tangent function mapping and z-score mapping; N represents the number of sets of the j-th target sub-range, s k Represents the feature similarity of the kth concentration range containing the jth target subrange.

[0080] It should be noted that in the example above where n is 4, the description continues with f as the normalization operation. The range weight of the target subrange [x1, x2] is the result of the normalization operation based on the feature similarity of the reference patients corresponding to the first reference concentration range. The range weights of other target subranges are calculated similarly and are not further described here.

[0081] Refer to the attached Figure 6 As shown, an embodiment of the present application provides an electronic device, including: a processor and a memory, the processor and the memory are interconnected and communicate with each other through a communication bus and / or other forms of connection mechanisms (not shown), the memory stores a computer program executable by the processor, and when the computing device is running, the processor executes the computer program to execute the system in any optional implementation mode of the above embodiment.

[0082] An embodiment of the present application provides a storage medium, wherein when the computer program is executed by a processor, the system of any optional implementation of the above embodiment is executed. The storage medium can be implemented by any type of volatile or non-volatile storage device or a combination thereof, such as static random access memory (SRAM), electrically erasable programmable read-only memory (EEPROM), erasable programmable read-only memory (EPROM), programmable read-only memory (PROM), read-only memory (ROM), magnetic storage, flash memory, magnetic disk or optical disk.

[0083] In the embodiments provided in this application, it should be understood that the disclosed system can be implemented in other ways. The system embodiments described above are merely schematic. For example, the division of the modules is only a logical function division, and can be implemented in another way. For example, multiple modules or components can be combined or integrated into another system, or some features can be ignored or not executed. Another point is that the mutual coupling or direct coupling or communication connection shown or discussed can be through some communication interface, indirect coupling or communication connection of the system or unit, which can be electrical, mechanical or other forms.

[0084] In addition, the units described as separate components may or may not be physically separate, and the components shown as units may or may not be physical units, that is, they may be located in one place or distributed across multiple network units. Some or all of the units may be selected according to actual needs to achieve the purpose of the solution of this embodiment.

[0085] Furthermore, the functional modules in the various embodiments of the present application can be integrated together to form an independent part, or each module can exist independently, or two or more modules can be integrated to form an independent part.

[0086] Flowcharts are used herein to illustrate the steps of the methods of the embodiments of the present disclosure. It should be understood that the preceding or following steps do not necessarily need to be performed in exact order. Instead, the various steps may be evaluated in reverse order or simultaneously. Furthermore, other operations may be added to these processes.

[0087] Unless otherwise defined, all terms used herein have the same meaning as commonly understood by those skilled in the art to which the present disclosure belongs. It should also be understood that terms such as those defined in common dictionaries should be interpreted as having a meaning consistent with their meaning in the context of the relevant art and should not be interpreted in an idealized or highly formal sense unless expressly defined as such herein.

[0088] The above is a detailed introduction to the immunosuppressant dosing recommendation system based on deep reinforcement learning. This article uses specific examples to illustrate the principles and implementation methods of this application. The description of the above embodiments is only an embodiment of this application. It is only used to help understand the immunosuppressant dosing recommendation system based on deep reinforcement learning of this application and is not used to limit the scope of protection of this application. At the same time, for those skilled in the art, this application can have various changes and variations. Any modifications and equivalent substitutions made within the spirit and principles of this application should be included in the scope of protection of this application.

Claims

1. An immunosuppressant drug administration recommendation system based on deep reinforcement learning, characterized by: It includes a model construction module, an information acquisition module, and a regimen recommendation module; the model construction module includes a strategy submodule, an environment submodule, a reward submodule, and a training submodule; the strategy submodule constructs a medication recommendation model, which is used to output a medication recommendation regimen after receiving patient characteristics; The environmental submodule predicts the patient's blood drug concentration after medication based on the patient's characteristics and the recommended medication regimen; the reward submodule evaluates the patient's blood drug concentration after medication and obtains the reward and punishment value; The training submodule trains the medication recommendation model based on the reward and punishment values, and sends the medication recommendation model to the regimen recommendation module when the training termination condition is met; the regimen recommendation module inputs the target features of the target patient sent by the information acquisition module into the medication recommendation model to obtain the target medication regimen; The recommended dosing regimen includes the recommended dosage of each drug in all administration methods, and the patient's blood drug concentration after taking the drug, including the trend of blood drug concentration changes within the preset time period; The reward submodule includes a concentration constraint module and a reward and penalty calculation module; the concentration constraint module includes a standard concentration acquisition unit; the standard concentration acquisition unit matches the corresponding target blood drug concentration range according to the patient's characteristics; the reward and penalty calculation module calculates the concentration component based on the proportion of time that the blood drug concentration change trend within the preset time period is within the target blood drug concentration range, and generates corresponding reward and penalty values ​​based on the concentration component; The standard concentration acquisition unit further includes a rule matching subunit, a reference acquisition subunit and a standard generation subunit; The rule matching subunit generates a rule concentration range based on the concentration matching results of the patient characteristics in the medication rule set; the reference acquisition subunit queries the database based on the patient characteristics to obtain a reference patient, and obtains a reference concentration range based on the actual blood drug concentration of the reference patient at the corresponding treatment stage; the reference patient includes a patient with a patient profile that is the same as the patient profile of the corresponding patient and a history of normal medication use at the corresponding treatment stage; the standard generation subunit obtains a target blood drug concentration range based on the rule concentration range and all reference concentration ranges; The standard generation sub-unit also takes the union of the rule concentration range and all reference concentration ranges as the target blood drug concentration range, divides the target blood drug concentration range into several target sub-ranges according to the overlap between the rule concentration range and all reference concentration ranges, and obtains the range weight according to the number of sets of the rule concentration range and all reference concentration ranges that contain each target sub-range. The target sub-ranges with more sets have corresponding range weights with larger values; the reward and punishment calculation module calculates the concentration component according to the proportion of time that the blood drug concentration change trend conforms to each target sub-range and the corresponding range weight.

2. The immunosuppressant dosage recommendation system based on deep reinforcement learning according to claim 1, characterized in that: The recommended dosing schedule also includes the recommended frequency of each drug for all routes of administration.

3. The immunosuppressant dosage recommendation system based on deep reinforcement learning according to claim 1 or 2, characterized in that: The reference acquisition subunit calculates the feature similarity between the patient portraits of historical patients who normally take medication in each corresponding treatment stage and the patient portraits of the corresponding patients, and uses the historical patients whose feature similarity is not less than the similarity threshold as reference patients; The calculation methods of concentration components include: ; ; Where S represents the concentration component, φ and f are mapping functions, M represents the number of reference subranges, ω j represents the range weight of the j-th target subrange, p j represents the proportion of time when the trend of blood drug concentration changes conforms to the jth target sub-range, ω0 represents the out-of-range weight, whose value is less than 0, p0 represents the proportion of time when the trend of blood drug concentration changes does not conform to the target blood drug concentration range, N represents the number of sets of the jth target sub-range, s k Represents the feature similarity of the kth concentration range containing the jth target subrange.

4. The immunosuppressant dosage recommendation system based on deep reinforcement learning according to claim 1, characterized in that: The reward submodule also includes a safety constraint submodule, which obtains a safety component based on the medication safety matching results of the medication recommendation scheme and the medication rule set, and a reward and punishment calculation submodule generates corresponding reward and punishment values ​​based on the concentration component and the safety component.

Citation Information

Patent Citations

  • Recommendation system for constructing T2DM patient drug regimen based on deep learning and reinforcement learning

    CN117894424A

  • Training method, optimization method and device of anti-epileptic drug administration strategy optimization model

    CN118588226A