Immunosuppressant administration recommendation system based on deep reinforcement learning

Through the deep reinforcement learning drug delivery recommendation system, the dosing plan of immunosuppressants is optimized, which solves the problems of difficulty in making dosing plans and inaccurate control of blood drug concentration in the prior art, and realizes personalized, safe and stable drug delivery plan recommendations.

CN120376033AActive Publication Date: 2025-07-25SICHUAN ACADEMY OF MEDICAL SCI SICHUAN PROVINCIAL PEOPLES HOSPITAL

Patent Information

Application Number
CN202510864034.3
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-06-26
Publication Date
2025-07-25
Estimated Expiration
2045-06-26

AI Technical Summary

Technical Problem

In the prior art, it is difficult to make the dosing regimen of immunosuppressants, and it is difficult to accurately control the blood drug concentration of patients, resulting in an increase in rejection and the risk of poisoning.

Method used

The drug delivery recommendation system based on deep reinforcement learning is adopted, and the drug delivery plan is optimized through the strategy submodule, the environmental submodule, the reward submodule and the training submodule, the decision-making of clinical experts is simulated, and a personalized drug delivery recommendation plan is generated. The drug delivery strategy is optimized by combining blood drug concentration prediction and safety evaluation.

Benefits of technology

It reduces the difficulty of clinical physicians' administration decision-making, improves the safety and accuracy of the dosing regimen, reduces the risks caused by fluctuations in blood drug concentration, and improves the stability and individualized guidance capabilities of the dosing regimen.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120376033A_ABST
    Figure CN120376033A_ABST
Patent Text Reader

Abstract

The invention discloses an immunosuppressor administration recommendation system based on deep reinforcement learning, and relates to the technical field of medical medication management, and the system comprises a model construction module, an information acquisition module and a scheme recommendation module. The model construction module comprises a strategy sub-module, an environment sub-module, a reward sub-module and a training sub-module; the strategy sub-module constructs a drug administration recommendation model, and the drug administration recommendation model is used for receiving patient characteristics and then outputting a drug administration recommendation scheme; the environment sub-module predicts the blood concentration of the patient after medication according to the patient characteristics and the medication recommendation scheme; the reward sub-module evaluates the blood concentration of the patient after medication to obtain a reward and punishment value; the training sub-module trains a drug administration recommendation model based on the reward and punishment values, and sends the drug administration recommendation model to the scheme recommendation module when a training termination condition is met; and the scheme recommendation module inputs the target characteristics of the target patient sent by the information acquisition module into a drug administration recommendation model to obtain a target drug administration scheme so as to reduce the decision-making difficulty of the immunosuppressor drug administration scheme.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the technical field of medical drug management, and particularly to an immunosuppressant administration recommendation system based on deep reinforcement learning. Background Art

[0002] Immunosuppressants refer to drugs that have an inhibitory effect on the body's immune response, such as being widely used to treat immune rejection reactions after solid organ transplantation and hematopoietic stem cell transplantation. And the administration regimen of immunosuppressants is a key factor affecting the treatment effect of patients. When the blood drug concentration of a patient after taking the drug is lower than the effective concentration, the risk of the patient having a rejection reaction increases; when the blood drug concentration of a patient after taking the drug is not lower than the toxic concentration, the risk of the patient having adverse events such as infection increases. The narrow therapeutic window of immunosuppressants makes it difficult to make decisions on the administration regimen of immunosuppressants for patients.

[0003] In existing clinical practice, the administration dose is mainly determined according to the patient's body weight. However, due to individual differences in different patients in terms of age, type of transplantation, time after surgery, and concomitant medications, only a few patients' blood drug concentrations can reach the target concentration range when formulating the administration regimen according to body weight, and most patients still need to be adjusted twice or even multiple times according to the blood drug concentration test results after taking the drug. At this time, the risk of patient medication and the difficulty of formulating the administration regimen will increase with the increase in the number of adjustments.

[0004] Therefore, there is an urgent need for an immunosuppressant administration recommendation system to solve the technical problem of difficult decision-making on the administration regimen. Summary of the Invention

[0005] The present invention provides an immunosuppressant administration recommendation system based on deep reinforcement learning, which optimizes the administration recommendation model through the training method of reinforcement learning to assist medical staff in formulating the administration regimen of immunosuppressants.

[0006] The present invention claims an immunosuppressant administration recommendation system based on deep reinforcement learning, including a model construction module, an information acquisition module, and a scheme recommendation module; the model construction module includes a policy sub-module, an environment sub-module, a reward sub-module, and a training sub-module; the policy sub-module constructs an administration recommendation model, and the administration recommendation model is used to output an administration recommendation scheme after receiving patient characteristics; the environment sub-module predicts the blood drug concentration of a patient after taking the drug according to the patient characteristics and the administration recommendation scheme; the reward sub-module evaluates the blood drug concentration of a patient after taking the drug to obtain a reward and punishment value; the training sub-module trains the administration recommendation model based on the reward and punishment value, and sends the administration recommendation model to the scheme recommendation module when the training termination condition is met; the scheme recommendation module inputs the target characteristics of the target patient sent by the information acquisition module into the administration recommendation model to obtain the target administration scheme.

[0007] In an embodiment of the present application, the drug administration recommendation plan includes the recommended doses of each drug in all administration methods, and the blood drug concentration after the patient takes the drug includes the blood drug concentration change trend within a preset time period.

[0008] In an embodiment of the present application, the drug administration recommendation plan further includes the recommended frequencies of each drug in all administration methods.

[0009] In an embodiment of the present application, the reward sub-module includes a concentration constraint grandchild module and a reward and punishment calculation grandchild module; the concentration constraint grandchild module includes a standard concentration acquisition unit; the standard concentration acquisition unit matches the corresponding target blood drug concentration range according to the patient characteristics; the reward and punishment calculation grandchild module calculates the concentration component according to the proportion of the time that meets the target blood drug concentration range in the blood drug concentration change trend within a preset time period, and generates the corresponding reward and punishment value according to the concentration component.

[0010] In an embodiment of the present application, the standard concentration acquisition unit further includes a rule matching subunit, a reference acquisition subunit, and a standard generation subunit; the rule matching subunit generates a rule concentration range according to the concentration matching result of the patient characteristics in the drug administration rule set; the reference acquisition subunit queries the database according to the patient characteristics to obtain a reference patient, and obtains the reference concentration range according to the actual blood drug concentration of the reference patient in the corresponding treatment stage; the reference patient includes a historical patient with the same patient portrait as the corresponding patient and taking medicine normally in the corresponding treatment stage; the standard generation subunit obtains the target blood drug concentration range according to the rule concentration range and all reference concentration ranges.

[0011] In an embodiment of the present application, the standard generation subunit also takes the union of the rule concentration range and all reference concentration ranges as the target blood drug concentration range, divides the target blood drug concentration range into several target sub-ranges according to the overlapping situation of the rule concentration range and all reference concentration ranges, and obtains the range weight according to the number of sets containing each target sub-range in the rule concentration range and all reference concentration ranges; the target sub-range with a larger number of sets corresponds to a larger range weight; the reward and punishment calculation grandchild module calculates the concentration component according to the proportion of the time that the blood drug concentration change trend meets each target sub-range and the corresponding range weight respectively.

[0012] In an embodiment of the present application, the reference acquisition subunit calculates the feature similarity between the patient portrait of each historical patient taking medicine normally in the corresponding treatment stage and the patient portrait of the corresponding patient respectively, and takes the historical patient with the feature similarity not less than the similarity threshold as the reference patient; the calculation method of the concentration component includes: ; ; where S represents the concentration component, φ and f are mapping functions, M represents the number of reference sub-ranges, and ω jRepresents the range weight of the j-th target sub-range, p j Represents the proportion of time when the blood drug concentration change trend conforms to the j-th target sub-range. ω0 represents the out-of-range weight, whose value is less than 0. p0 represents the proportion of time when the blood drug concentration change trend does not conform to the target blood drug concentration range. N represents the number of sets of the j-th target sub-range, s k Represents the feature similarity of the k-th concentration range including the j-th target sub-range.

[0013] In an embodiment of the present application, the reward sub-module further includes a safety constraint grandchild module. The safety constraint grandchild module obtains a safety component according to the medication safety matching result between the drug administration recommendation plan and the drug administration rule set. The reward and punishment calculation grandchild module generates a corresponding reward and punishment value according to the concentration component and the safety component.

[0014] The present application has the following beneficial effects: 1. The drug administration recommendation model as an agent learns the optimal strategy for recommending immunosuppressant medication plans for patients in the environment. By reinforcement learning, it adopts a trial-and-error method to improve the decision-making ability of the drug administration recommendation model and simulate the decision-making method of clinical experts. The reward function serves as the basis for training and updating the drug administration recommendation model, enabling the drug administration recommendation model to distinguish the difference between the best drug administration plan and the actual drug administration plan on the basis of learning how clinical experts generate drug administration plans. That is, the drug administration recommendation plan corresponding to the maximum reward value given by the reward function is the best drug administration plan for this patient, thereby gradually optimizing the target recommendation plan and reducing the decision-making difficulty for clinicians to formulate immunosuppressant medication plans for patients.

[0015] 2. The peak and trough corresponding time points in the complete blood drug concentration change trend obtained by the blood drug concentration prediction model after the patient takes the medicine reflect the individual differences in drug absorption, distribution, and metabolism among different patient groups. Compared with the single blood drug concentration value given at a fixed preset time point, it more truly reflects the dynamic process of the drug in the body, reduces the probability of situations such as the blood drug concentration briefly reaching the standard but actually having insufficient dosage or toxicity risks, and improves the safety, stability, and accuracy of drug administration.

[0016] 3. The policy sub-module generates a recommended frequency while giving the recommended dose of the immunosuppressant, realizing a dose and frequency combination plan in units of the medication cycle. This technical solution significantly reduces the manual operation frequency and workload of medical staff. Especially for formulating medication plans for patients with stable conditions and in the late recovery stage, the periodic plan is beneficial to improving clinical work efficiency and operation convenience.

[0017] 4. Introduce evaluation indicators of drug use safety in the reward sub-module, so that when the strategy sub-module generates a drug administration recommendation plan, it not only focuses on achieving the target blood drug concentration, but also can actively identify drug administration recommendation plans with potential unsafe drug use risks, guiding the strategy sub-module to prefer to recommend drug use plans with better safety, thereby enhancing the clinical usability and patient safety of the recommendation results.

[0018] 5. Designing the reward sub-module based on time series can more realistically reflect the mechanism of drug metabolism dynamics and side effects in patients, comprehensively evaluate the residence time of the patient's blood drug concentration in each risk interval, so as to realize the fine-grained analysis of the patient's drug exposure process. This method reduces the misjudgment of scores caused by accidental abnormalities such as noise interference and patient state fluctuations, misjudging the accidental blood drug concentration fluctuations as rejection reactions or poisoning reaction states, and enhancing the stability of the drug administration recommendation model.

[0019] 6. The concentration acquisition unit matches a more accurate target blood drug concentration range based on the individual differences of patients. On the one hand, it not only improves the safety and rationality evaluation of the drug use plan by this system, but on the other hand, it also reduces the difficulty of collecting the drug administration rule set. Especially when immunosuppressants are used in combination, the coverage difficulty of the drug administration rule set increases. Introducing the individual differences of historical patients to match expands the application scenario of the drug administration recommendation model to combination drug use, reducing the limitations brought by single drug evaluation.

[0020] 7. The reward and punishment calculation sub-module constructs a personalized blood drug concentration risk stratification interval through the interval overlap voting mechanism. This design method of the reward function considers the differences in control objectives at different stages during immunosuppressive therapy. On the premise of ensuring patient drug use safety, this design method enables the blood drug concentration evaluation to be dynamically adjusted with the treatment stage. The reference concentration range reflects the acceptable range of blood drug concentration fluctuations of the patient at the current treatment stage. Especially when more reference concentration ranges overlap in the interval outside the regular concentration range, for the corresponding patient, the risk of rejection reaction and / or poisoning reaction when the blood drug concentration is in this concentration range is lower, thereby enhancing the adaptability and individualized guidance ability of the model in actual clinical applications. For example, in the early stage after transplantation surgery, the patient's body has a high risk of rejection reaction to the transplanted organ, the control of blood drug concentration needs to be more strict, the requirement for the patient's medication compliance is high, the difference between the intersection and union of all reference ranges is small, and the target concentration interval should be relatively tight to ensure the immunosuppressive effect; while in the later stage of recovery, the body gradually establishes adaptability to the transplanted organ, and even if the blood drug concentration fluctuates briefly, it is not easy to cause rejection reaction. At this stage, more attention should be paid to the risk of long-term cumulative toxicity, and the difference between the intersection and union of all reference ranges increases, and the target blood drug concentration interval is appropriately relaxed to balance the curative effect and toxicity.

[0021] 8. Reward and Punishment Calculation Sub-module The reference patients highly similar to the patient portrait of the patient have a greater influence on the final risk interval division result, while the reference patients with lower similarity have a smaller influence weight, thereby reducing the interference of non-representative historical data on the risk judgment result. Brief Description of the Drawings

[0022] The accompanying drawings herein are incorporated into the specification and form a part of this specification, showing embodiments consistent with the present disclosure and used together with the specification to explain the principles of the present disclosure.

[0023] Figure 1 It is a schematic diagram of the data transmission process of the immunosuppressant administration recommendation system based on deep reinforcement learning according to an embodiment of the present application; Figure 2 It is a schematic structural diagram of an immunosuppressant administration recommendation system according to an embodiment of the present application; Figure 3 It is a schematic diagram of data transmission in the concentration constraint sub-module according to an embodiment of the present application; Figure 4 It is another schematic structural diagram of an immunosuppressant administration recommendation system according to an embodiment of the present application; Figure 5 It is a schematic diagram of range weights according to an embodiment of the present application; Figure 6 It is a schematic structural diagram of an electronic device according to an embodiment of the present application; Reference signs in the figures: x1, x2, x3, x4, x5 - example values of blood drug concentration, L4 - regular concentration range, L1 - first reference concentration range, L2 - second reference concentration range, L3 - third reference concentration range. Detailed Embodiments

[0024] The present invention provides an immunosuppressant administration recommendation system based on deep reinforcement learning. To make the above objects, features, and advantages of the present application more obvious and understandable, the technical solutions in the embodiments of the present application will be clearly and completely described below with reference to the accompanying drawings in the embodiments of the present application. Obviously, the described embodiments are only a part of the embodiments of the present application, rather than all the embodiments. Usually, the components of the embodiments of the present application described and shown in the accompanying drawings here can be arranged and designed in various different configurations. Therefore, with reference to terms such as "one embodiment", "some embodiments", "embodiment", "embodiments", "schematic embodiments", "example", "specific example", or "some examples", etc., the following detailed description of the embodiments of the present application provided in the accompanying drawings is not intended to limit the scope of the present application to be protected, but only represents that the specific features, structures, or characteristics described in connection with the embodiment or example are included in at least one embodiment or example of the present invention. Moreover, the specific features, structures, or characteristics described can be combined in a suitable manner in any one or more embodiments or examples. Based on the embodiments of the present application, all other embodiments obtained by those skilled in the art without creative efforts belong to the scope of protection of the present application.

[0025] It should be noted that: similar reference numerals and letters denote similar items in the following drawings. Therefore, once an item is defined in one drawing, it does not need to be further defined and explained in subsequent drawings. At the same time, in the description of the present application, relational terms such as "first", "second", etc. are only used to distinguish one entity or operation from another entity or operation, and do not necessarily require or imply any such actual relationship or order between these entities or operations.

[0026] The present invention claims protection for an immunosuppressant administration recommendation system based on deep reinforcement learning. Referring to the attached Figure 1 and the attached Figure 2 as shown, it includes a model construction module, an information acquisition module, and a solution recommendation module.

[0027] It should be noted that the model construction module includes a policy sub-module, an environment sub-module, a reward sub-module, and a training sub-module. The policy sub-module is used to construct a drug administration recommendation model. The input of the drug administration recommendation model includes patient characteristics, and the output includes a drug administration recommendation plan. The environment sub-module is used to construct and train a blood drug concentration prediction model. The blood drug concentration prediction model is used to predict the blood drug concentration of a patient after taking medicine according to the corresponding drug administration plan. The reward sub-module is used to evaluate the blood drug concentration of a patient after taking medicine, and generate an evaluation result based on a preset reward function, that is, obtain a corresponding reward and punishment value. The training sub-module obtains a training data set uploaded by a user for training the drug administration recommendation model. At the same time, the training sub-module also uses the method of reinforcement learning to update the drug administration recommendation model of the policy sub-module based on the reward value and the value function. When the preset training termination condition is satisfied, the trained drug administration recommendation model is sent to the plan recommendation module.

[0028] It should be noted that the information acquisition module is used to acquire the characteristics of a target patient, obtain target characteristics, and send the target characteristics to the plan recommendation module. The plan recommendation module inputs the received target characteristics into the trained drug administration recommendation model to output a drug administration plan for the target patient, that is, the target drug administration plan.

[0029] It should be noted that the policy sub-module can be pre-set with machine learning models such as neural network models, and generate the drug administration recommendation model according to the requirements input by the user. For example, the features specified by the user are used as the input features of the drug administration recommendation model, and the output and objective function of the drug administration recommendation model are generated according to the prediction task specified by the user. The policy sub-module can also directly obtain the drug administration recommendation model constructed by the user.

[0030] It should be noted that the environment sub-module is used to simulate the blood drug concentration of a patient after taking medicine. The blood drug concentration prediction model can be constructed through a neural network model. The input of this blood drug concentration prediction model is the patient characteristics and the drug administration plan of the immunosuppressant, and the output is the blood drug concentration of the patient after using this drug administration plan. The characteristics, actual drug administration plan, and blood drug concentration after taking medicine of historical patients in the historical medical records can be used as a prediction training data set, and the blood drug concentration prediction model is trained according to this prediction training data set.

[0031] It should be noted that the reward sub-module is used to evaluate the blood drug concentration of the patient. Reference materials such as literature, clinical practice guidelines, drug instructions, and medication standards can be systematically retrieved. The Cochrane systematic review tool and the Appraisal of Guidelines for Research & Evaluation System II (AGREE II) are used to evaluate the retrieval results to ensure the quality and applicability of the reference materials, and to exclude reference materials with low quality or those that do not meet the screening criteria. The medication rules of the patient are extracted from the screened reference materials, and combined with the opinions of clinical experts to obtain a preliminary set of medication rules. The Delphi expert consultation method is used to evaluate each medication rule in the preliminary set of medication rules, that is, a questionnaire is designed according to all medication evaluation rules, and the questionnaire is distributed to the expert group. After multiple rounds of questionnaire filling, collection, collation, and analysis, until the opinions of the expert group tend to be consistent, the final rule results are sorted out and improved to obtain a set of dosing rules. The reward sub-module matches the set of dosing rules according to the patient characteristics to obtain the target blood drug concentration range for the corresponding patient; compares the output of the patient in the environment sub-module with the target blood drug concentration range. If the output of the patient in the environment sub-module belongs to the target blood drug concentration range, the reward sub-module gives a reward. If the output of the patient in the environment sub-module does not belong to the target blood drug concentration range, the reward sub-module gives a penalty.

[0032] It should be noted that the training sub-module obtains the recommended training data set input by the user, and this recommended data set is used to train the dosing recommendation model. The training sub-module can train the dosing recommendation model according to value functions such as the Q-Learning algorithm, the Sarsa algorithm, or the policy gradient algorithm.

[0033] It should be noted that the training sub-module first trains the dosing recommendation model based on the method of imitation learning. Imitation learning can imitate and learn the dosing schemes of clinical experts to reduce the learning difficulty and learning efficiency of the dosing recommendation model in the initial learning stage.

[0034] It should be noted that the patient characteristics include basic characteristics and diagnostic characteristics, and of course it can be not limited to this. The basic characteristics include gender, age, past medical history, family medical history, etc., and of course it can be not limited to this. The diagnostic characteristics include the name of the disease suffered, the treatment plan, and the examination and test results before medication, etc., and of course it can be not limited to this. The treatment method includes the time series of the use of all treatment plans after the patient is admitted to the hospital. For example, the treatment plan for a liver transplant patient includes using drug A on the first day after admission, using drug A and drug B on the second day, using drug B on the third day, having a liver transplant surgery on the fourth day, using drug B on the fifth day, and using drug C on the sixth day, which can be correspondingly recorded as (A, AB, B, liver transplant surgery, B, C). The examination and test results before medication include the blood drug concentration before medication, and of course it can be not limited to this.

[0035] It should be noted that if the currently formulated drug administration plan is the first time for the patient to use an immunosuppressant, the pre-drug blood drug concentration corresponding to the patient can be defaulted to 0.

[0036] In the first embodiment, the policy sub-module is used to generate the recommended dose of each preset drug for the patient according to the patient characteristics. The input of the environment sub-module includes the patient characteristics and the drug administration recommendation plan, and the output includes the blood drug concentration of the patient at a preset time point after using the drug administration recommendation plan. For example, the blood drug concentration 1 hour after taking the drug.

[0037] In this embodiment, the reward sub-module includes a concentration constraint grandchild module and a reward and punishment calculation grandchild module. The concentration constraint module includes a standard concentration acquisition unit. The standard concentration acquisition unit has a pre-constructed drug administration rule set built in, matches the drug administration rule set according to the patient characteristics, and generates a corresponding target blood drug concentration range according to all successfully matched drug administration rule sets. The reward and punishment calculation grandchild module is used to generate a corresponding reward and punishment value according to the comparison result between the blood drug concentration of the patient after taking the drug and the target blood drug concentration range. For example, if the blood drug concentration of the patient meets the target blood drug concentration range, the probability of the patient having rejection reaction and toxicity reaction is regarded as a low risk, and the reward and punishment calculation grandchild module gives a reward value of 1. If the blood drug concentration of the patient does not meet the target blood drug concentration range, the probability of the patient having rejection reaction and toxicity reaction is regarded as a high risk, and the reward and punishment calculation grandchild module gives a punishment value of -1.

[0038] It should be noted that the number of the preset drugs can be one or multiple. The preset drugs can be preset according to the immunosuppressants in the medical institution. The present application does not further limit the preset drugs and the quantity.

[0039] It should be noted that the preset time point can be obtained by a pre-set method, or can be used as an input feature of the blood drug concentration prediction model to obtain the time point specified by the user, or other feasible implementation manners. The present application does not further limit the setting method and the specific value of the preset time point.

[0040] In this embodiment, the reward sub-module further includes a safety constraint sub-sub-module, although it is not limited thereto. The safety constraint sub-sub-module obtains a safety component according to the matching result between the drug administration recommendation plan and the drug administration plan rule. The safety component is used to evaluate the safety level of the drug use recommendation plan. The evaluation criteria for the safety level include whether the drug use recommendation plan of the immunosuppressant interacts with the patient's historical medications or other medications, whether there are drug interactions when different immunosuppressant drugs are used simultaneously, etc., although it is not limited thereto. If the drug administration recommendation plan will have an adverse interaction with other drugs that the patient is currently using, the reward and punishment calculation sub-sub-module gives a punishment value. The specific value of the punishment value can be obtained according to the adverse impact degree of the drug interaction. For example, if the drug interaction is to reduce the efficacy of the immunosuppressant or other drugs, a punishment value of -0.1 is given; if the drug interaction poses a lethal risk, a punishment value of -1 is given.

[0041] In this embodiment, the reward and punishment calculation sub-sub-module can be obtained based on the linear weighted calculation result of the concentration component and the safety component according to the importance.

[0042] In this embodiment, the drug administration recommendation model includes a number of dose recommendation sub-models. The output of each dose recommendation sub-model includes the mean and variance of the immunosuppressant dose distribution of the patient for the corresponding preset drug. The strategy sub-module regards the mean of the immunosuppressant dose distribution as the best estimate of the dose. After constructing the normal distribution of the patient population portrait according to the mean and variance output by the drug administration recommendation model, the recommended dose of the patient for the corresponding preset drug is sampled from the normal distribution according to the probability density.

[0043] It should be noted that taking cyclosporine and tacrolimus as examples, the drug administration recommendation model includes a first dose recommendation sub-model and a second dose recommendation sub-model, although it is not limited thereto. The first dose recommendation sub-model is used to generate the mean and variance of the patient's use of cyclosporine according to the patient characteristics. The second dose recommendation sub-model is used to generate the recommended dose of the patient's use of tacrolimus according to the patient characteristics. The strategy sub-module constructs the dose normal distribution of the patient's use of cyclosporine and the dose normal distribution of the patient's use of tacrolimus respectively, and samples the recommended dose of the patient's use of cyclosporine and the recommended dose of the patient's use of tacrolimus from the corresponding dose normal distributions respectively. Due to individual differences among different patients, the same dose may have different treatment effects on different patients. By predicting the probability distribution of the patient's drug dosage through the drug administration recommendation model, the dosage preferences of the corresponding patient population for each preset drug are simulated, and action prediction in a continuous space is provided through probability density sampling, realizing adaptation to the individual differences of patients.

[0044] In the second embodiment, the difference from the first embodiment is that the output of the environment sub-module includes the change trend of the patient's blood drug concentration within a preset time period after using the drug administration recommendation plan. For example, the change trend of the blood drug concentration within 5 hours after taking the drug. After the standard concentration acquisition unit of the reward sub-module obtains the corresponding target blood drug concentration range according to the patient characteristics, the reward and punishment calculation sub-module generates the corresponding reward and punishment value according to whether there is a moment when the blood drug concentration change trend does not conform to the target blood drug concentration range. If the blood drug concentration change trend of the patient conforms to the target blood drug concentration range, it indicates that the probability of the patient having an abnormal drug use event is small, and the reward and punishment calculation sub-module gives a reward value of 1. If there is a moment when the blood drug concentration change trend does not conform to the target blood drug concentration range, it indicates that the probability of the patient having an abnormal drug use event is large, and the reward and punishment calculation sub-module gives a punishment value of -1.

[0045] It should be noted that the preset time period can be obtained by pre-setting, or can be used as an input feature of the blood drug concentration prediction model to obtain the time period specified by the user, or other feasible implementation methods. The present application does not further limit the setting method and specific duration of the preset time period.

[0046] It should be noted that the blood drug concentration prediction model is constructed based on a long short-term memory neural network model, and the training data set can be obtained from historical medical records, and the blood drug concentration values of historical patients at different time points are used as the time series features of the blood drug concentration.

[0047] In the third embodiment, the difference from the second embodiment is that the policy sub-module is further configured to generate the recommended frequency of each preset drug for the patient according to the patient characteristics, and the policy sub-module generates the drug administration recommendation plan according to the recommended dose and recommended frequency of each preset drug.

[0048] In this embodiment, the drug administration recommendation model further includes a first frequency recommendation sub-model and a second frequency recommendation sub-model, and of course, it may not be limited to this. The first frequency recommendation sub-model is used to generate the recommended frequency of cyclosporine for the patient according to the patient characteristics. The second frequency recommendation sub-model is used to generate the recommended frequency of tacrolimus for the patient according to the patient characteristics.

[0049] In this embodiment, the reward and punishment calculation sub-module obtains the concentration component according to the proportion of the time within the corresponding blood drug concentration change trend output by the environment sub-module that conforms to the target blood drug concentration range.

[0050] It should be noted that the smaller the proportion of the time within the preset time period that conforms to the target blood drug concentration range, the smaller the value of the corresponding concentration component; the larger the proportion of the time within the preset time period that conforms to the target blood drug concentration range, the larger the value of the corresponding concentration component.

[0051] In this embodiment, the value range of the concentration component is preset to [-1, 1]. When the value of the concentration component is greater than 0, it indicates that the reward and punishment calculation sub-module gives a corresponding reward value on the concentration component. For example, when the value of the concentration component is 1, it indicates that the blood drug concentration change trend of the patient within the preset time period conforms to the target blood drug concentration range. When the value of the concentration component is less than 0, it indicates that the reward and punishment calculation sub-module gives a corresponding punishment value on the concentration component. For example, when the value of the concentration component is -1, it indicates that the blood drug concentration change trend of the patient within the preset time period does not conform to the target blood drug concentration range. The calculation method of the concentration component S1 can be: ; ; where T represents the total number of time points in the preset time period, i represents the serial number of the time point, c i represents the blood drug concentration value at the i-th time point in the output of the environment sub-module, C represents the target blood drug concentration range, I is an indicator function, and A represents the input condition of the indicator function I.

[0052] It should be noted that in the above calculation method of the concentration component, the proportion of the time within the preset time period that conforms to the target blood drug concentration range is converted into the corresponding reward and punishment value by using a linear mapping method. Nonlinear methods such as tangent function mapping or z-score mapping can also be used to calculate the reward and punishment value. This embodiment does not further limit the specific mapping method.

[0053] In the fourth embodiment, referring to Attachments Figure 3 and Attachments Figure 4 shown, the difference from the third embodiment is that the standard concentration acquisition unit further includes a rule matching sub-unit, a reference acquisition sub-unit, and a standard generation sub-unit, and of course it may not be limited to this. The rule matching sub-unit is used to generate a rule concentration range according to the concentration matching result of the patient characteristics in the drug administration rule set. The reference acquisition sub-unit queries the database according to the patient characteristics to obtain a reference patient, and obtains a reference concentration range according to the actual blood drug concentration of the reference patient in the corresponding treatment stage. The reference patient includes historical patients with the same patient portrait as the patient portrait of the corresponding patient and taking medicine normally in the corresponding treatment stage. The standard generation sub-unit obtains the target blood drug concentration range according to the rule concentration range and all reference concentration ranges.

[0054] It should be noted that the dosing rule set includes the recommended blood drug concentration ranges corresponding to patients with different characteristics. For example, the recommended blood drug concentration for adult patients within one week after surgery, the recommended blood drug concentration for patients with abnormal liver function, etc. The rule matching subunit matches the corresponding concentration matching result according to the patient characteristics, and generates the corresponding rule concentration range according to the concentration matching result.

[0055] In this embodiment, the reference obtaining subunit screens out historical patients who did not have immune reactions and toxicity reactions during the treatment stage corresponding to the corresponding patients from the historical medical records to obtain candidate patients, that is, the candidate patients are historical patients who used drugs normally corresponding to them. The reference obtaining subunit calculates the feature similarity between the patient portraits of each candidate patient and the patient portrait of the corresponding patient respectively, and takes the candidate patients with the feature similarity not less than the similarity threshold as the reference patients.

[0056] It should be noted that the patient portrait is constructed by collecting data on aspects such as the patient's basic information, disease status, medical treatment behavior, and treatment process to build a personalized information model of the patient. The feature vectors of the patient portraits of the candidate patients and the feature vectors of the patient portraits of the corresponding patients can be extracted respectively through a neural network model. The feature similarity is calculated based on the cosine similarity. The reference obtaining subunit is used to obtain the corresponding reference concentration range according to the maximum and minimum blood drug concentrations of each reference patient during the corresponding treatment stage.

[0057] It should be noted that the similarity threshold can be obtained by a pre-set method. The target blood drug concentration range can be obtained according to the intersection of the rule concentration range and all the reference concentration ranges.

[0058] In this embodiment, the standard generating subunit regards the target blood drug concentration range according to the rule concentration range and all the reference concentration ranges as the reference range, denoted as L, that is, L = {L1, L2,..., L i ,..., L n}, where L i represents the i-th reference range, and n represents the number of reference ranges. The union of all reference ranges is used as the target blood drug concentration range. The standard generating subunit also includes dividing the target blood drug concentration range into several target sub-ranges according to the overlapping situation of all reference ranges, and obtaining the range weight corresponding to each target sub-range according to the number of sets containing each target sub-range in all reference ranges. The target sub-range with a larger number of sets corresponds to a larger range weight value.

[0059] It should be noted that taking n = 4 as an example, referring to the appendix Figure 5As shown, where the example values of the blood drug concentration x1, x2, x3, x4, and x5 decrease in sequence. When the value of n is 4, it indicates that the reference range set L = {L1, L2, L3, L4}. The rule concentration range is denoted as L4, that is, L4 = [x2, x4]. The first reference concentration range is denoted as L1, that is, L1 = [x1, x3]. The second reference concentration range is denoted as L2, that is, L2 = [x2, x3]. The third reference concentration range is denoted as L3, that is, L3 = [x2, x5]. Thus, it can be seen that the target blood drug concentration range is [x1, x5]. According to the overlapping situation of all reference ranges, the target blood drug concentration range can be divided into 4 target sub-ranges, which are [x1, x2], [x2, x3], [x3, x4], and [x4, x5] respectively. The target sub-range [x1, x2] is only included in the first reference concentration range L1, and the target sub-range [x4, x5] is only included in the third reference concentration range L3. Then the number of sets corresponding to the target sub-ranges [x1, x2] and [x4, x5] is 1, and the value of the corresponding range weight is the smallest. The target sub-range [x3, x4] is simultaneously included in the rule concentration range L4 and the third reference concentration range L3. Then the number of sets corresponding to the target sub-range [x3, x4] is 3. Similarly, it can be known that the number of sets corresponding to the target sub-range [x2, x3] is 4, and the value of the corresponding range weight is the largest. The range weight of the corresponding target sub-range can be calculated through a normalization operation. That is, the range weights of the target sub-ranges [x1, x2] and [x4, x5] are 0.125, the range weight of the target sub-range [x3, x4] is 0.25, and the range weight of the target sub-range [x2, x3] is 0.5.

[0060] In this embodiment, when the reward and punishment calculation sub-module calculates the concentration component S, the calculation method is as follows: ; where φ is a mapping function, and methods such as linear mapping, tangent function mapping, and z-score mapping can be selected to convert it to the preset value range of the reward and punishment value; M represents the number of reference sub-ranges; ω j represents the range weight of the jth target sub-range; p j represents the proportion of time when the change trend of the blood drug concentration conforms to the jth target sub-range; ω0 represents the out-of-range weight, and p0 represents the proportion of time when the change trend of the blood drug concentration does not conform to the target blood drug concentration range.

[0061] It should be noted that ω0 is used to set a penalty value for the time that does not belong to the target blood drug concentration range, that is, its value is less than any ω jWhen the value of ω0 is 0, it indicates that the drug administration recommendation model only focuses on the proportion of time within the standard range. When the value of ω0 is negative, it means that the drug administration recommendation model penalizes the proportion of time outside the target blood drug concentration range, thereby reflecting the risk of the drug administration recommendation plan, improving the discrimination between the drug administration with risk and without risk, and guiding the strategy learning direction of the drug administration recommendation model.

[0062] In this embodiment, the range weight ω j is calculated as follows: ; where f is a mapping function, and methods such as linear mapping, tangent function mapping, and z-score mapping can be selected to convert it to the preset range of the range weight; N represents the number of sets of the j-th target sub-range, and s k represents the feature similarity of the k-th concentration range including the j-th target sub-range.

[0063] It should be noted that in the above example where the value of n is 4, the normalization operation is continued with f as an example. The range weight of the target sub-range [x1, x2] is the result of the normalization operation of the feature similarity of the reference patient corresponding to the first reference concentration range. The calculation methods of the range weights of other target sub-ranges are similar and will not be described in detail here.

[0064] Referring to the appendix Figure 6 shown, the embodiment of the present application provides an electronic device, including: a processor and a memory. The processor and the memory are interconnected and communicate with each other through a communication bus and / or other forms of connection mechanisms (not marked). The memory stores a computer program executable by the processor. When the computing device runs, the processor executes the computer program to execute the system in any optional implementation manner of the above embodiment.

[0065] An embodiment of the present application provides a storage medium. When the computer program is executed by a processor, it executes the system in any optional implementation manner of the above embodiments. Among them, the storage medium can be implemented by any type of volatile or non-volatile storage device or a combination thereof, such as static random access memory (SRAM for short), electrically erasable programmable read-only memory (EEPROM for short), erasable programmable read-only memory (EPROM for short), programmable read-only memory (PROM for short), read-only memory (ROM for short), magnetic memory, flash memory, magnetic disk or optical disc.

[0066] In the embodiments provided by the present application, it should be understood that the disclosed system can be implemented in other ways. The system embodiments described above are merely illustrative. For example, the division of the modules is only a logical function division. In actual implementation, there may be other division methods. For another example, multiple modules or components can be combined or integrated into another system, or some features can be ignored or not executed. Another point is that the displayed or discussed couplings or direct couplings or communication connections to each other can be through some communication interfaces. The indirect couplings or communication connections of the systems or units can be in electrical, mechanical or other forms.

[0067] In addition, the units described as separate components may or may not be physically separated. The components displayed as units may or may not be physical units, that is, they may be located in one place, or they may be distributed to multiple network units. Some or all of the units can be selected according to actual needs to achieve the purpose of the solution of this embodiment.

[0068] Furthermore, the functional modules in each embodiment of the present application can be integrated together to form an independent part, or each module can exist alone, or two or more modules can be integrated to form an independent part.

[0069] Flowcharts are used herein to illustrate the steps of the methods through the embodiments of the present disclosure. It should be understood that the previous or subsequent steps do not necessarily need to be carried out precisely in sequence. On the contrary, they can be carried out in reverse order or evaluated simultaneously. At the same time, other operations can also be added to these processes.

[0070] Unless otherwise defined, all terms used herein have the same meaning as commonly understood by one of ordinary skill in the art to which this disclosure belongs. It should also be understood that terms such as those defined in commonly used dictionaries should be interpreted as having a meaning that is consistent with their meaning in the context of the relevant art and should not be interpreted in an idealized or overly formal sense unless expressly so defined herein.

[0071] The above has introduced in detail the provided immunosuppressant administration recommendation system based on deep reinforcement learning. Specific examples are used herein to elaborate on the principles and implementation manners of this application. The description of the above embodiments is only for the embodiments of this application and is only used to help understand the immunosuppressant administration recommendation system based on deep reinforcement learning of this application, and is not used to limit the protection scope of this application; at the same time, for those skilled in the art, various changes and modifications can be made to this application. Any modification or equivalent replacement made within the spirit and principle of this application shall be included within the protection scope of this application.

Claims

1. An immunosuppressant administration recommendation system based on deep reinforcement learning, characterized in that, It includes a model construction module, an information acquisition module, and a solution recommendation module; the model construction module includes a strategy sub-module, an environment sub-module, a reward sub-module, and a training sub-module; the strategy sub-module constructs a drug administration recommendation model, and the drug administration recommendation model is used to output a drug administration recommendation solution after receiving patient characteristics; The environment sub-module predicts the blood drug concentration after the patient takes the medicine according to the patient characteristics and the drug administration recommendation solution; the reward sub-module evaluates the blood drug concentration after the patient takes the medicine to obtain a reward and punishment value; The training sub-module trains the drug administration recommendation model based on the reward and punishment value, and sends the drug administration recommendation model to the solution recommendation module when the training termination condition is met; the solution recommendation module inputs the target characteristics of the target patient sent by the information acquisition module into the drug administration recommendation model to obtain a target drug administration solution.

2. The immunosuppressant administration recommendation system based on deep reinforcement learning according to claim 1, characterized in that The drug administration recommendation solution includes the recommended doses of each drug in all administration methods, and the blood drug concentration after the patient takes the medicine includes the blood drug concentration change trend within a preset time period.

3. The immunosuppressant administration recommendation system based on deep reinforcement learning according to claim 2, wherein The drug administration recommendation solution also includes the recommended frequencies of each drug in all administration methods.

4. The immunosuppressant administration recommendation system based on deep reinforcement learning according to claim 2 or 3, characterized in that, The reward sub-module includes a concentration constraint grandchild module and a reward and punishment calculation grandchild module; the concentration constraint grandchild module includes a standard concentration acquisition unit; the standard concentration acquisition unit matches the corresponding target blood drug concentration range according to the patient characteristics; the reward and punishment calculation grandchild module calculates the concentration component according to the proportion of the time within the target blood drug concentration range in the blood drug concentration change trend within a preset time period, and generates the corresponding reward and punishment value according to the concentration component.

5. The immunosuppressive agent administration recommendation system based on deep reinforcement learning according to claim 4, wherein The standard concentration acquisition unit also includes a rule matching sub-unit, a reference acquisition sub-unit, and a standard generation sub-unit; the rule matching sub-unit generates a rule concentration range according to the concentration matching result of the patient characteristics in the drug administration rule set; the reference acquisition sub-unit queries the database according to the patient characteristics to obtain a reference patient, and obtains a reference concentration range according to the actual blood drug concentration of the reference patient in the corresponding treatment stage; the reference patient includes a historical patient with the same patient portrait as the corresponding patient and taking medicine normally in the corresponding treatment stage; the standard generation sub-unit obtains the target blood drug concentration range according to the rule concentration range and all reference concentration ranges.

6. The immunosuppressant administration recommendation system based on deep reinforcement learning according to claim 5, wherein The standard generation sub-unit also takes the union of the rule concentration range and all reference concentration ranges as the target blood drug concentration range, divides the target blood drug concentration range into several target sub-ranges according to the overlapping situation between the rule concentration range and all reference concentration ranges, and obtains the range weight according to the number of sets containing each target sub-range in the rule concentration range and all reference concentration ranges; the larger the number of sets, the larger the range weight corresponding to the target sub-range; the reward and punishment calculation grandchild module calculates the concentration component according to the proportion of the time when the blood drug concentration change trend conforms to each target sub-range and the corresponding range weight respectively.

7. The immunosuppressant administration recommendation system based on deep reinforcement learning according to claim 6, characterized in that, The reference acquisition sub-unit calculates the feature similarity between the patient portrait of each historical patient taking medicine normally in the corresponding treatment stage and the patient portrait of the corresponding patient respectively, and takes the historical patient with the feature similarity not less than the similarity threshold as the reference patient; The calculation method of the concentration component includes: ; ; Where S represents the concentration component, φ and f are mapping functions, M represents the number of reference sub-ranges, and ω j represents the range weight of the j-th target sub-range, and p j represents the proportion of time during which the trend of the blood drug concentration conforms to the j-th target sub-range. ω0 represents the out-of-range weight, whose value is less than 0, and p0 represents the proportion of time during which the trend of the blood drug concentration does not conform to the target blood drug concentration range. N represents the number of sets of the j-th target sub-range, and s k represents the feature similarity of the k-th concentration range that includes the j-th target sub-range.

8. The immunosuppressant administration recommendation system based on deep reinforcement learning according to claim 1, wherein The reward sub-module further includes a safety constraint grandchild module. The safety constraint grandchild module obtains a safety component based on the medication safety matching result between the medication recommendation plan and the medication rule set. The reward and punishment calculation grandchild module generates a corresponding reward and punishment value based on the concentration component and the safety component.

Citation Information

Patent Citations

  • Method and device for determining medication scheme of patient

    CN113255735A

  • Recommendation system for constructing T2DM patient drug regimen based on deep learning and reinforcement learning

    CN117894424A

  • Training method, optimization method and device of anti-epileptic drug administration strategy optimization model

    CN118588226A

  • Method and device for determining drug regimen of patient

    WO2022227198A1

Cited By

  • Personalized drug recommendation method based on multi-target deep reinforcement learning

    CN121768572A