Dynamic medication dosage optimization method fused with reinforcement learning

By dynamically optimizing drug dosage through the fusion of reinforcement learning algorithms, the efficacy problem caused by dosage mismatch is solved, achieving precise drug use and reducing side effects.

CN120913746APending Publication Date: 2025-11-07YUEYANG MATERNAL & CHILD HEALTH HOSPITAL
View PDF 0 Cites 2 Cited by

Patent Information

Application Number
CN202511126210.X
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-08-12
Publication Date
2025-11-07

AI Technical Summary

Technical Problem

Current treatment regimens do not take into account individual patient differences in dosage, leading to poor efficacy or excessive side effects.

Method used

By employing a fusion reinforcement learning approach, and configuring reinforcement learning algorithms corresponding to the patient's symptoms, control parameters are determined based on the patient's physiological indicators and basic information, and medication dosage is dynamically optimized.

Benefits of technology

This ensures that the medication dosage is matched with the patient's actual condition, guarantees that the efficacy is consistent with the patient's physical condition, and avoids poor efficacy or excessive side effects.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120913746A_ABST
    Figure CN120913746A_ABST
Patent Text Reader

Abstract

The invention provides a dynamic medication dosage optimization method fused with reinforcement learning. The method comprises the following steps: acquiring a first physiological index parameter of a target patient; obtaining a first diagnosis report of the target patient, wherein the first diagnosis report comprises basic information of the target patient and first disease information of the target patient; the first disease information comprises a first disease type and first disease description information; determining a first reinforcement learning algorithm corresponding to the first disease type; determining a first control parameter of the first reinforcement learning algorithm according to the first physiological index parameter and the basic information; and performing operation on the first disease description information through the first reinforcement learning algorithm and the first control parameter to obtain a first medication dosage parameter. Based on the application, the drug effect can be consistent with the physical condition of a patient, and poor drug effect and excessive side effects caused by too strong drug effect are avoided.
Need to check novelty before this filing date? Find Prior Art

Description

TECHNICAL FIELD

[0001] The present application relates to the medical technical field, and in particular to a dynamic medication dosage optimization method fusing reinforcement learning. BACKGROUND

[0002] In the current treatment scheme, the medication dosage is mostly based on the experience of medication dosage, and in specific practice, the individual differences of patients are not considered, and the medication dosage remains unchanged in the process. Since the medication dosage remains unchanged, especially after taking medicine for a period of time, the body tends to recover, and the use of the same medication dosage will cause some side effects. SUMMARY

[0003] The embodiment of the present application provides a dynamic medication dosage optimization method fusing reinforcement learning, a reinforcement learning algorithm corresponding to a patient's illness is configured, corresponding control parameters are configured based on the physical condition and basic information of the patient, so that the reinforcement learning algorithm is consistent with the actual situation of the patient, and then the algorithm is operated based on the configured algorithm to obtain an accurate medication dosage, so that the medication dosage corresponds to the actual situation of the patient, helps to ensure the drug efficacy, makes the drug efficacy consistent with the physical condition of the patient, avoids poor drug efficacy, and avoids excessive drug efficacy to produce too many side effects.

[0004] The embodiment of the present application provides a dynamic medication dosage optimization method fusing reinforcement learning, the method comprises:

[0005] Obtain a first physiological index parameter of a target patient;

[0006] Obtain a first diagnosis report of the target patient, the first diagnosis report comprising: basic information of the target patient and first illness information of the target patient; the first illness information comprising a first illness type and first illness description information;

[0007] Determine a first reinforcement learning algorithm corresponding to the first illness type;

[0008] Determine a first control parameter of the first reinforcement learning algorithm according to the first physiological index parameter and the basic information;

[0009] Operate the first illness description information through the first reinforcement learning algorithm and the first control parameter to obtain a first medication dosage parameter.

[0010] In the embodiment of the present application, first, the first physiological index parameter of the target patient is acquired, and then the first diagnosis report of the target patient is acquired, the first diagnosis report comprising: basic information of the target patient and first disease information of the target patient, the first disease information comprising a first disease type and first disease description information, then, a first reinforcement learning algorithm corresponding to the first disease type is determined, and a first control parameter of the first reinforcement learning algorithm is determined according to the first physiological index parameter and the basic information, finally, the first disease description information is operated through the first reinforcement learning algorithm and the first control parameter to obtain a first drug dosage parameter, on the one hand, the first physiological index parameter represents the physical condition of the target patient, the first diagnosis report represents the basic information of the patient and the related disease (the first disease type) and the disease related performance (the first disease description information), and the reinforcement learning algorithm corresponding to the disease is configured, on the other hand, the corresponding control parameter is configured based on the physical condition and the basic information of the patient, so that the reinforcement learning algorithm is consistent with the actual situation of the patient, and then the disease related performance is operated based on the configured algorithm to obtain an accurate drug dosage, so as to ensure that the drug dosage corresponds to the actual situation of the patient, which is helpful to ensure the drug efficacy and make the drug efficacy consistent with the physical condition of the patient, thereby avoiding poor drug efficacy and avoiding excessive drug efficacy to produce excessive side effects. BRIEF DESCRIPTION OF DRAWINGS

[0011] In order to more clearly illustrate the technical solutions in the embodiments of the present application or the prior art, the following will briefly introduce the drawings needed to be used in the embodiments or the prior art description. Obviously, the drawings in the following description are only some embodiments of the present application, and other drawings can also be obtained by those skilled in the art without creative labor.

[0012] Figure 1 is a structural schematic diagram of a dynamic drug dosage optimization system fusing reinforcement learning provided by the embodiment of the present application;

[0013] Figure 2 is a flowchart of a dynamic drug dosage optimization method fusing reinforcement learning provided by the embodiment of the present application;

[0014] Figure 3 is a structural schematic diagram of a dynamic drug dosage optimization device fusing reinforcement learning provided by the embodiment of the present application;

[0015] Figure 4 is a structural schematic diagram of an electronic device provided by the embodiment of the present application. DETAILED DESCRIPTION

[0016] In the following, the technical solutions in the embodiments of the present application will be described clearly and completely in conjunction with the drawings in the embodiments of the present application, so that those skilled in the art can better understand the present application. Obviously, the described embodiments are only some of the embodiments of the present application, but not all the embodiments. Based on the embodiments in the present application, all other embodiments obtained by those skilled in the art without creative work fall within the scope of the present application.

[0017] The terms "first", "second", and the like in the description and claims of the present application and the above drawings are used to distinguish different objects, and are not used to describe a specific order. In addition, the terms "include" and "have" and any variations thereof are intended to cover non-exclusive inclusion. For example, a process, method, system, product or device including a series of steps or units is not limited to the listed steps or units, but can optionally include other steps or units not listed, or can optionally include other steps or units inherent to the process, method, product or device.

[0018] It should be understood that the term "and / or" herein is only used to describe the association relationship of the associated objects, which means that there can be three relationships, for example, A and / or B can mean that A exists alone, A and B exist together, and B exists alone. In addition, the character " / " in this paper represents that the front and rear associated objects are a "or" relationship. "Multiple" in the embodiments of the present application means two or more.

[0019] The "at least one" or similar expressions in the embodiments of the present application means any combination of these items, including any combination of single item or multiple items, means one or more, and multiple means two or more. For example, at least one of a, b or c can represent the following seven cases: a, b, c, a and b, a and c, b and c, a, b and c. Wherein, each of a, b and c can be an element or a set containing one or more elements.

[0020] The "connection" appearing in the embodiments of the present application means direct connection or indirect connection and various connection modes to realize communication between devices, which is not limited in the embodiments of the present application.

[0021] In this paper, "embodiment" means that the specific features, structures or characteristics described in conjunction with the embodiment can be included in at least one embodiment of the present application. The phrase appears at various places in the specification does not necessarily refer to the same embodiment, nor is it an independent or alternative embodiment to other embodiments. Those skilled in the art explicitly and implicitly understand that the embodiments described herein can be combined with other embodiments.

[0022] In the embodiment of the present application, the reinforcement learning algorithm can include at least one of the following: a neural network algorithm, a value-based method (such as Q-learning, DQN), a policy-based method (such as policy gradient), an Actor-Critic method (such as A2C, PPO), and a hybrid method (such as DDPG, TRPO), and of course, other reinforcement learning algorithms can also be included.

[0023] In the embodiment of the present application, the electronic device can include various computer devices, such as servers, medical devices, smart phones, vehicle-mounted devices, wearable devices, smart watches, smart glasses, wireless Bluetooth earphones, computing devices, and other various forms of user equipment (UE), mobile stations (MS), and the like, without limitation.

[0024] The embodiment of the present application will be described in detail below.

[0025] Please refer to Figure 1 , Figure 1 is a structural schematic diagram of a dynamic medication dosage optimization system fusing reinforcement learning provided by the embodiment of the present application. The dynamic medication dosage optimization system fusing reinforcement learning includes an electronic device and a physiological state sensor.

[0026] The physiological state sensor can include at least one of the following: a temperature sensor, an in-vivo chip, a micro robot, an electrode, and the like, without limitation.

[0027] The physiological state sensor can collect physiological index parameters at certain time intervals, which can be pre-set or system default. The physiological index parameters can include at least one of the following: temperature, brain waves, blood pressure, blood sugar, blood lipids, platelets, adrenaline content, and the number of B cells, without limitation.

[0028] Based on the electronic device in the dynamic medication dosage optimization system fusing reinforcement learning, the following functions can be performed:

[0029] Obtain the first physiological index parameter of the target patient;

[0030] Obtain the first diagnosis report of the target patient, which includes the basic information of the target patient and the first disease information of the target patient; the first disease information includes the first disease type and the first disease description information;

[0031] Determine the first reinforcement learning algorithm corresponding to the first disease type;

[0032] Determine the first control parameter of the first reinforcement learning algorithm according to the first physiological index parameter and the basic information;

[0033] The first disease description information is calculated through the first reinforcement learning algorithm and the first control parameter to obtain a first drug dosage parameter.

[0034] In a specific implementation, first, a first physiological index parameter of a target patient is acquired, and then a first diagnosis report of the target patient is acquired, the first diagnosis report including basic information of the target patient and first disease information of the target patient, the first disease information including a first disease type and first disease description information. Next, a first reinforcement learning algorithm corresponding to the first disease type is determined, and a first control parameter of the first reinforcement learning algorithm is determined according to the first physiological index parameter and the basic information. Finally, the first disease description information is calculated through the first reinforcement learning algorithm and the first control parameter to obtain a first drug dosage parameter. On one hand, the first physiological index parameter represents the physical condition of the target patient, the first diagnosis report represents the basic information of the patient and the related disease (the first disease type) and the disease-related performance (the first disease description information), and the reinforcement learning algorithm corresponding to the disease is configured. On the other hand, the corresponding control parameter is configured based on the physical condition of the patient and the basic information, so that the reinforcement learning algorithm is consistent with the actual situation of the patient. Then, the disease-related performance is calculated based on the configured algorithm to obtain an accurate drug dosage, so as to ensure that the drug dosage corresponds to the actual situation of the patient, which helps to ensure the drug efficacy and make the drug efficacy consistent with the physical condition of the patient, thereby avoiding poor drug efficacy and excessive drug efficacy to produce excessive side effects.

[0035] Please refer to Figure 2 , Figure 2 is a flowchart of a dynamic drug dosage optimization method fusing reinforcement learning provided by an embodiment of the present application, as shown in the figure, the dynamic drug dosage optimization method fusing reinforcement learning includes:

[0036] S201, a first physiological index parameter of a target patient is acquired.

[0037] The first physiological index parameter is used to represent the physiological condition of the target patient, and the first physiological index parameter can include at least one of the following: temperature, brain wave, blood pressure, blood sugar, blood lipid, electrocardiogram, blood platelet, adrenaline content, and bag cell number, etc., which are not limited herein.

[0038] In a specific implementation, the first physiological index parameter can be a single index parameter or a composite index parameter.

[0039] Preferably, the step of acquiring the first physiological index parameter of the target patient can include the following steps:

[0040] The first physiological index parameter of the target patient is acquired at each time of drug use of the target patient.

[0041] In a specific implementation, at each time of medication of the target patient, the first physiological index parameter of the target patient can be acquired first, i.e., the current physical condition of the target patient can be understood through the first physiological index parameter.

[0042] Preferably, the method further comprises:

[0043] acquiring a first time of the last medication of the target patient;

[0044] determining a time interval between the current time and the first time;

[0045] when the time interval is greater than a first preset duration, performing the step of acquiring the first physiological index parameter of the target patient.

[0046] Preferably, the first preset duration can be pre-set or system default.

[0047] In a specific implementation, the first time of the last medication of the target patient can be acquired, and then the current time can be acquired, and the time interval between the current time and the first time can be determined. When the time interval is greater than the first preset duration, it can be understood that the efficacy of the last medication has been fully exerted, and then the step of acquiring the first physiological index parameter of the target patient can be performed again to complete the determination of the dose of the medication again, so as to ensure accurate medication and drug efficacy.

[0048] Preferably, the method further comprises the following steps:

[0049] counting the number of medications of the target patient;

[0050] determining the first preset duration according to the number of medications.

[0051] In a specific implementation, the number of medications of the target patient can be counted, and a mapping relationship between a preset number of medications and a preset duration can be pre-stored, and the first preset duration corresponding to the counted number of medications can be determined based on the mapping relationship, so that the time interval of medication corresponds to the number of medications. Since the recovery degree of the target patient is different for different number of medications, the drug efficacy also changes to some extent, and then the time interval of medication corresponds to the number of medications, so that the drug efficacy can be fully exerted, and accurate medication can be ensured to ensure drug efficacy.

[0052] Preferably, the method further comprises the following steps:

[0053] acquiring a set of physiological index parameters before the current time, the set of physiological index parameters comprising a plurality of historical physiological index parameters, each historical physiological index parameter corresponding to an acquisition time;

[0054] According to the plurality of historical physiological index parameters and the collection time points thereof, a first physiological index fitting straight line is obtained; the horizontal axis of the first physiological index fitting straight line is time, and the vertical axis is a physiological index parameter;

[0055] A first slope of the first fitting straight line is obtained.

[0056] The first preset time length is determined according to the first slope.

[0057] In a specific implementation, a set of physiological index parameters before the current time can be obtained, the set of physiological index parameters including a plurality of historical physiological index parameters, each historical physiological index parameter corresponding to a collection time point. Then, according to the plurality of historical physiological index parameters and the collection time points thereof, a first physiological index fitting straight line is obtained. Specifically, the plurality of historical physiological index parameters and the collection time points thereof can be regarded as a plurality of coordinate points. Then, based on the plurality of coordinate points, the first physiological index fitting straight line is obtained. The horizontal axis of the first physiological index fitting straight line is time, and the vertical axis is a physiological index parameter. Then, a first slope of the first fitting straight line is obtained. The first slope reflects the body change of the target patient. Then, a first preset time length is determined according to the first slope. That is, a mapping relationship between a preset slope and a preset time length can be pre-stored. Based on the mapping relationship, the first preset time length corresponding to the first slope is determined. In this way, the medication interval corresponds to the body change of the target patient, which helps to improve the medication effect.

[0058] S202, a first diagnosis report of the target patient is obtained, the first diagnosis report including: basic information of the target patient and first disease information of the target patient; the first disease information including a first disease type and first disease description information.

[0059] The first diagnosis report can be a current diagnosis report. For example, the target patient can also be diagnosed before each medication to obtain a current diagnosis report. Alternatively, the first diagnosis report can be a diagnosis report of the most recent time.

[0060] The first diagnosis report includes: basic information of the target patient and first disease information of the target patient. The first disease information includes a first disease type and first disease description information. The first disease type is used to describe the specific disease type of the patient. The first disease description information is used to describe the external manifestation and / or internal manifestation of the disease. The first disease description information is used to describe the disease condition change corresponding to the first disease type of the target patient.

[0061] The basic information of the target patient can include at least one of the following: height, weight, gender, and the like, which are not limited herein.

[0062] S203, a first reinforcement learning algorithm corresponding to the first disease type is determined.

[0063] The mapping relationship between the preset stored disease type and the reinforcement learning algorithm can be pre-stored, that is, the first reinforcement learning algorithm corresponding to the first disease type can be determined based on the mapping relationship. In a specific implementation, a corresponding reinforcement learning algorithm can be configured for different disease types.

[0064] In S204, a first control parameter of the first reinforcement learning algorithm is determined according to the first physiological index parameter and the basic information.

[0065] In a specific implementation, the first control parameter is used to control the execution effect of the first reinforcement learning algorithm, and the execution effect can include at least one of the following: execution speed, execution accuracy, and the like.

[0066] Specifically, the first control parameter of the first reinforcement learning algorithm can be determined according to the first physiological index parameter and the basic information, so that the execution effect of the first reinforcement learning algorithm corresponds to the actual situation of the patient.

[0067] Preferably, the step of determining the first control parameter of the first reinforcement learning algorithm according to the first physiological index parameter and the basic information can be performed in the following manner:

[0068] A reference control parameter of the first reinforcement learning algorithm is determined according to the basic information.

[0069] A first physical condition evaluation value corresponding to the first physiological index parameter is determined.

[0070] A first adjustment parameter is determined according to the first physical condition evaluation value.

[0071] The reference control parameter is adjusted according to the first adjustment parameter to obtain the first control parameter.

[0072] In a specific implementation, the reference control parameter is used to control the execution effect of the first reinforcement learning algorithm, and the execution effect can include at least one of the following: execution speed, execution accuracy, and the like.

[0073] The mapping relationship between the preset basic information and the control parameter of the first reinforcement learning algorithm can be pre-stored, and the first control parameter corresponding to the basic information of the target patient can be determined based on the mapping relationship.

[0074] Correspondingly, the mapping relationship between the preset physiological index parameter and the body condition evaluation value can also be stored in advance, and the first body condition evaluation value corresponding to the first physiological index parameter can be determined based on the mapping relationship. When the first physiological index parameter is a single index parameter, the first body condition evaluation value corresponding to the first physiological index parameter can be directly determined based on the mapping relationship. When the first physiological index parameter is a plurality of index parameters, the body condition evaluation value corresponding to each physiological index parameter in the first physiological index parameter can be determined based on the mapping relationship, a plurality of body condition evaluation values are obtained, and the first body condition evaluation value is obtained by performing weighted operation based on the plurality of body condition evaluation values.

[0075] Further, the mapping relationship between the preset body condition evaluation value and the adjustment parameter is stored in advance, the first adjustment parameter corresponding to the first body condition evaluation value is determined based on the mapping relationship, and finally, part or all of the control parameters in the reference control parameter are adjusted according to the first adjustment parameter to obtain the first control parameter, that is, the first control parameter=(1+first adjustment parameter)*reference control parameter. For example, the value range of the first adjustment parameter is-0.1-0.1. Further, the model effect can be dynamically optimized based on the body condition of the target patient, so that the model later configures the drug dose to be consistent with the body condition, thereby ensuring that the drug dose corresponds to the actual situation of the patient, which helps to ensure the drug effect, so that the drug effect is consistent with the body condition of the patient, thereby avoiding poor drug effect and avoiding excessive drug effect to produce excessive side effects.

[0076] S205, the first drug dosage parameter is obtained by operating the first disease description information through the first reinforcement learning algorithm and the first control parameter.

[0077] The first drug dosage parameter can include a drug dosage parameter for one drug, which can be understood as a main drug component for the first symptom type or a component combination of the drug for the first symptom type.

[0078] The first drug dosage parameter can include a drug name, a drug dosage, a drug taking method, and the like.

[0079] In specific implementations, the first disease description information can be feature extracted to obtain first feature information, the first reinforcement learning algorithm is configured by using the first control parameter to obtain a configured model, the first feature information is input into the model to obtain the first drug dosage parameter.

[0080] The first feature information can include at least one of a feature value, a feature vector, a keyword, and the like, without limitation.

[0081] In the embodiment of the present application, first, a first physiological index parameter of a target patient is acquired, and then a first diagnosis report of the target patient is acquired, the first diagnosis report including basic information of the target patient and first disease information of the target patient, the first disease information including a first disease type and first disease description information; then, a first reinforcement learning algorithm corresponding to the first disease type is determined, and a first control parameter of the first reinforcement learning algorithm is determined according to the first physiological index parameter and the basic information; finally, the first disease description information is operated through the first reinforcement learning algorithm and the first control parameter to obtain a first drug dosage parameter. On the one hand, the first physiological index parameter represents the physical condition of the target patient, the first diagnosis report represents the basic information of the patient and the related disease (the first disease type) and the disease-related performance (the first disease description information), and the reinforcement learning algorithm corresponding to the disease is configured. On the other hand, the corresponding control parameter is configured based on the physical condition and the basic information of the patient, so that the reinforcement learning algorithm is consistent with the actual situation of the patient. Then, the disease-related performance is operated based on the configured algorithm to obtain an accurate drug dosage, so as to ensure that the drug dosage corresponds to the actual situation of the patient, which helps to ensure the drug efficacy and make the drug efficacy consistent with the physical condition of the patient, thereby avoiding poor drug efficacy and avoiding excessive drug efficacy to produce excessive side effects.

[0082] Preferably, the method further comprises the following steps:

[0083] According to the first drug dosage parameter, a corresponding drug is configured;

[0084] After the target patient takes the drug, a second time is recorded;

[0085] Within a first time period after the second time, physiological index parameters of the target patient are acquired every preset time interval to obtain a plurality of physiological index parameters, each physiological index parameter corresponding to a collection time;

[0086] According to the set of physiological index parameters and the collection time corresponding to each physiological index parameter, a second physiological index fitting straight line is obtained, the horizontal axis of the second physiological index fitting straight line being time and the vertical axis being a physiological index parameter;

[0087] A slope of the second physiological index fitting straight line is acquired to obtain a second slope;

[0088] A first deviation between the second slope and a preset threshold is determined;

[0089] According to the first deviation, the first control parameter is optimized to obtain a second control parameter.

[0090] The preset time interval can be pre-set or system default. The preset threshold can be pre-set or system default. The preset threshold can be understood as a standard drug effect. The preset threshold can be an empirical value or a statistical value. The preset threshold can be related to the number of medication times. Different medication times correspond to different preset thresholds.

[0091] In a specific implementation, after step S205, the corresponding drug can be configured according to the first medication dosage parameter, and after the target patient takes the drug, the second time can be recorded. The second time can be understood as the time of taking the drug. In the first time period after the second time, the first time period can be understood as the time period from the current drug taking to the next drug taking. The physiological index parameter of the target patient is obtained every preset time interval, and a plurality of physiological index parameters are obtained. Each physiological index parameter corresponds to a collection time. Then, the physiological index parameter set and the collection time corresponding to each physiological index parameter are fitted to obtain a second physiological index fitting straight line. That is, each physiological index parameter and the collection time corresponding thereto can be regarded as a coordinate point. Therefore, a plurality of coordinate points can be obtained from a plurality of physiological index parameters. The second physiological index fitting straight line can be obtained by fitting based on the plurality of coordinate points. The horizontal axis of the second physiological index fitting straight line is time and the vertical axis is the physiological index parameter. Then, the slope of the second physiological index fitting straight line is obtained to obtain a second slope.

[0092] Next, the first deviation between the second slope and the preset threshold can be determined. The first deviation = (second slope - preset threshold) / preset threshold. The first deviation reflects the deviation degree between the actual drug effect and the standard drug effect. Then, the first control parameter is optimized according to the first deviation to obtain a second control parameter, so that the next medication is closer to the actual situation of the target patient.

[0093] Next, when the next medication is performed, the second diagnosis report of the target patient can be obtained. The second diagnosis report can be the current diagnosis report. Specifically, the target patient can also be diagnosed before each medication to obtain the current diagnosis report. The second diagnosis report can include second symptom description information corresponding to the first symptom type. The second symptom description information is used to describe the external manifestation of the symptom and / or the internal manifestation. The second symptom description information is used to describe the change of the second symptom type of the target patient.

[0094] The second symptom description information is operated by the first reinforcement learning algorithm and the second control parameter to obtain a second medication dosage parameter. In this way, the drug effect dynamic optimization model parameter of the target patient can be obtained, so that the medication dosage is dynamically optimized, which helps to improve the recovery effect of the target patient.

[0095] Preferably, the step of optimizing the first control parameter according to the first difference to obtain the second control parameter can be implemented in the following manner:

[0096] determining a first optimization parameter corresponding to the first deviation degree;

[0097] optimizing the first control parameter according to the first optimization parameter to obtain the second control parameter.

[0098] In a specific implementation, a mapping relationship between a preset deviation degree and an optimization parameter can be pre-stored, the first optimization parameter corresponding to the first deviation degree is determined based on the mapping relationship, and part or all of the first control parameters are optimized according to the first optimization parameter to obtain the second control parameter. For example, the value range of the first optimization parameter is -0.02-0.02, and the specific optimization process can also be: second control parameter=(1+first optimization parameter)*first control parameter. The dynamic optimization model parameter of the target patient can be based on the target patient, so that the medication dose is dynamically optimized, which helps to improve the recovery effect of the target patient.

[0099] Please refer to Figure 3 , Figure 3 is a structure diagram of a dynamic medication dose optimization device 300 fusing reinforcement learning provided by an embodiment of the present application; the dynamic medication dose optimization device 300 fusing reinforcement learning comprises an acquisition unit 301, a determination unit 302, and an operation unit 303, wherein:

[0100] The acquisition unit 301 is configured to acquire a first physiological index parameter of a target patient, and acquire a first diagnosis report of the target patient, wherein the first diagnosis report comprises basic information of the target patient and first disease information of the target patient, and the first disease information comprises a first disease type and first disease description information.

[0101] The determination unit 302 is configured to determine a first reinforcement learning algorithm corresponding to the first disease type, and determine a first control parameter of the first reinforcement learning algorithm according to the first physiological index parameter and the basic information.

[0102] The operation unit 303 is configured to perform operation on the first disease description information by using the first reinforcement learning algorithm and the first control parameter to obtain a first medication dose parameter.

[0103] The dynamic medication dose optimization device 300 fusing reinforcement learning in the embodiment of the application first acquires a first physiological index parameter of a target patient, then acquires a first diagnosis report of the target patient, the first diagnosis report including basic information of the target patient and first illness information of the target patient, the first illness information including a first illness type and first illness description information, then determines a first reinforcement learning algorithm corresponding to the first illness type, and determines a first control parameter of the first reinforcement learning algorithm according to the first physiological index parameter and the basic information, and finally performs operation on the first illness description information by the first reinforcement learning algorithm and the first control parameter to obtain a first medication dose parameter. On one hand, the first physiological index parameter represents the physical condition of the target patient, the first diagnosis report represents the basic information of the patient and the related illness (the first illness type) and the illness-related performance (the first illness description information), and the reinforcement learning algorithm corresponding to the illness is configured. On the other hand, the corresponding control parameter is configured based on the physical condition and the basic information of the patient, so that the reinforcement learning algorithm is consistent with the actual situation of the patient, and the illness-related performance is operated based on the configured algorithm to obtain an accurate medication dose, so as to ensure that the medication dose corresponds to the actual situation of the patient, which helps to ensure the drug efficacy and make the drug efficacy consistent with the physical condition of the patient, thereby avoiding poor drug efficacy and avoiding excessive drug efficacy to produce excessive side effects.

[0104] Preferably, the acquisition unit 301 implements the following manner when performing the acquisition of the first physiological index parameter of the target patient:

[0105] At each time of medication of the target patient, the first physiological index parameter of the target patient is acquired.

[0106] Preferably, the dynamic medication dose optimization device 300 fusing reinforcement learning is further specifically used for:

[0107] The first time of the most recent medication of the target patient is acquired.

[0108] The time interval between the current time and the first time is determined.

[0109] When the time interval is greater than a first preset time length, the step of acquiring the first physiological index parameter of the target patient is performed.

[0110] Preferably, the dynamic medication dose optimization device 300 fusing reinforcement learning is further specifically used for:

[0111] The number of medications of the target patient is counted.

[0112] The first preset time length is determined according to the number of medications.

[0113] Preferably, the dynamic medication dosage optimization device 300 that fuses reinforcement learning is also specifically used for:

[0114] Obtaining a physiological indicator parameter set before the current time, the physiological indicator parameter set including a plurality of historical physiological indicator parameters, each historical physiological indicator parameter corresponding to an acquisition time;

[0115] Fitting according to the plurality of historical physiological indicator parameters and the acquisition times thereof to obtain a first physiological indicator fitting straight line, the horizontal axis of the first physiological indicator fitting straight line being time and the vertical axis being a physiological indicator parameter;

[0116] Obtaining a first slope of the first fitting straight line;

[0117] Determining the first preset time length according to the first slope.

[0118] Preferably, when the determining unit 302 executes the determining of the first control parameter of the first reinforcement learning algorithm according to the first physiological indicator parameter and the basic information, the determining unit 302 can be implemented in the following manner:

[0119] Determining a reference control parameter of the first reinforcement learning algorithm according to the basic information;

[0120] Determining a first body condition evaluation value corresponding to the first physiological indicator parameter;

[0121] Determining a first adjustment parameter according to the first body condition evaluation value;

[0122] Adjusting the reference control parameter according to the first adjustment parameter to obtain the first control parameter.

[0123] Preferably, the dynamic medication dosage optimization device 300 that fuses reinforcement learning is also specifically used for:

[0124] Configuring a corresponding medication according to the first medication dosage parameter;

[0125] Recording a second time after the target patient takes the medication;

[0126] Obtaining a physiological indicator parameter of the target patient every preset time interval within a first time period after the second time to obtain a plurality of physiological indicator parameters, each physiological indicator parameter corresponding to an acquisition time;

[0127] Fitting according to the physiological indicator parameter set and the acquisition time corresponding to each physiological indicator parameter to obtain a second physiological indicator fitting straight line, the horizontal axis of the second physiological indicator fitting straight line being time and the vertical axis being a physiological indicator parameter;

[0128] obtaining a slope of the fitting straight line of the second physiological index, to obtain a second slope;

[0129] determining a first deviation degree between the second slope and a preset threshold value;

[0130] optimizing the first control parameter according to the first deviation degree, to obtain a second control parameter.

[0131] Preferably, the dynamic medication dose optimization device 30 of the fusion reinforcement learning can be implemented in the following manner when performing the optimization of the first control parameter according to the first difference value to obtain the second control parameter:

[0132] determining a first optimization parameter corresponding to the first deviation degree;

[0133] optimizing the first control parameter according to the first optimization parameter, to obtain the second control parameter.

[0134] Please refer to Figure 4 , Figure 4 is a structural schematic diagram of an electronic device provided by an embodiment of the application, which includes a processor, a memory, a communication interface and one or more programs, the one or more programs are stored in the memory and configured to be executed by the processor, and the program includes instructions for executing the following steps:

[0135] obtaining a first physiological index parameter of a target patient;

[0136] obtaining a first diagnosis report of the target patient, the first diagnosis report including basic information of the target patient and first disease information of the target patient, and the first disease information including a first disease type and first disease description information;

[0137] determining a first reinforcement learning algorithm corresponding to the first disease type;

[0138] determining a first control parameter of the first reinforcement learning algorithm according to the first physiological index parameter and the basic information;

[0139] operating the first disease description information through the first reinforcement learning algorithm and the first control parameter to obtain a first medication dose parameter.

[0140] The one or more programs are stored in the memory and configured to be executed by the processor, and the one or more programs include instructions for executing any step of the above method embodiments.

[0141] The processor can be a general processor, a central processing unit (CPU), an application-specific integrated circuit (ASIC), a digital signal processor (DSP), a field programmable gate array (FPGA) or other programmable logic device, a transistor logic device, a hardware component, or any combination thereof. It can implement or execute various exemplary logical blocks, units, and circuits described in combination with the disclosure of the present application. The processor can also be a combination of computing functions, such as a combination of one or more microprocessors, a combination of a DSP and a microprocessor, and the like. The communication unit can be a communication interface, a transceiver, a transceiver circuit, and the like, and the storage unit can be a memory.

[0142] The memory can be a volatile memory or a non-volatile memory, or can include both volatile and non-volatile memories. The non-volatile memory can be a read-only memory (ROM), a programmable ROM (PROM), an erasable PROM (EPROM), an electrically EPROM (EEPROM), or a flash memory. The volatile memory can be a random access memory (RAM) used as an external cache. By way of example, but not limitation, many forms of random access memory (RAM) are available, such as static random access memory (SRAM), dynamic random access memory (DRAM), synchronous DRAM (SDRAM), double data rate SDRAM (DDR SDRAM), enhanced SDRAM (ESDRAM), synchlink DRAM (SLDRAM), and direct rambus RAM (DR RAM).

[0143] It can be understood that the electronic device can further include more or less structural elements than those in the above structural block diagram, for example, a communication module (a Wi-Fi module), a physical key, a Bluetooth module, a sensor, a display module, a speaker, a power module, etc., which are not limited herein. It can be understood that the electronic device can be equipped with, for example, a system architecture as described above. Figure 1 The system architecture is described above.

[0144] The embodiment of the present application further provides a computer storage medium, wherein the computer storage medium stores a computer program for electronic data exchange, and the computer program causes a computer to execute part or all steps of any method described in the above method embodiments.

[0145] The embodiment of the present application further provides a computer program product, which includes a non-transitory computer-readable storage medium storing a computer program, and the computer program is operable to cause a computer to execute part or all steps of any method described in the above method embodiments. The computer program product can be a software installation package.

[0146] It should be noted that, for the above method embodiments, in order to simply describe, they are all described as a series of action combinations, but those skilled in the art should know that the present application is not limited to the action sequence described, because according to the present application, some steps can be performed in other sequences or simultaneously. Secondly, those skilled in the art should know that the embodiments described in the specification all belong to preferred embodiments, and the actions and modules involved are not necessarily necessary for the present application.

[0147] In the above embodiments, the description of each embodiment has its own focus, and the parts not described in detail in a certain embodiment can be referred to the related description of other embodiments.

[0148] In several embodiments provided by the present application, it should be understood that the disclosed device can be implemented in other ways. For example, the device embodiments described above are only schematic. The division of the above units is only a logical function division. There can be another division manner for actual implementation, for example, a plurality of units or components can be combined or integrated into another system, or some features can be ignored or not executed. In addition, the displayed or discussed mutual coupling or direct coupling or communication connection can be indirect coupling or communication connection through some interfaces, devices or units, and can be electrical or other forms.

[0149] The units described as separate components above can or can not be physically separate, and the components shown as units can or can not be physical units, i.e., can be located in one place, or can be distributed to multiple network units. Some or all of the units can be selected according to actual needs to achieve the purpose of the embodiment.

[0150] In addition, each functional unit in each embodiment of the present application can be integrated in one processing unit, or each unit can exist physically, or two or more units can be integrated in one unit. The integrated unit can be realized in the form of hardware or in the form of a software functional unit.

[0151] The integrated unit, if realized in the form of a software functional unit and sold or used as an independent product, can be stored in a computer readable storage medium. Based on this understanding, the technical solutions of the present application essentially or the part that contributes to the prior art or the whole or part of the technical solutions can be embodied in the form of a software product. The computer software product is stored in a storage medium and includes a plurality of instructions for causing a computer device (which can be a personal computer, a server or a network device, etc.) to execute all or part of the steps of the above-mentioned method of each embodiment of the present application. The aforementioned storage medium includes: a U disk, a read-only memory (ROM), a random access memory (RAM), a mobile hard disk, a magnetic disk or an optical disk, and various media that can store program codes.

[0152] Those skilled in the art can understand that all or part of the steps of the various methods of the above embodiments can be completed by programs instructing related hardware, and the programs can be stored in a computer readable storage medium, which can include: a flash disk, a read-only memory (ROM), a random access memory (RAM), a magnetic disk or an optical disk, etc.

[0153] The embodiments of the present application are described in detail above, and the principles and implementation manners of the present application are described by applying specific examples. The above description of the embodiments is only used to help understand the method of the present application and its core idea; at the same time, for those skilled in the art, according to the idea of the present application, the specific implementation manner and application range will be changed; in view of the above, the content of the specification should not be understood as a limitation of the present application.

Claims

1. A dynamic medication dosage optimization method fusing reinforcement learning, characterized in that, The method comprises: acquiring a first physiological indicator parameter of a target patient; acquiring a first diagnosis report of the target patient, the first diagnosis report comprising: basic information of the target patient and first disease information of the target patient; the first disease information comprising a first disease type and first disease description information; determining a first reinforcement learning algorithm corresponding to the first disease type; determining a first control parameter of the first reinforcement learning algorithm according to the first physiological indicator parameter and the basic information; operating the first disease description information through the first reinforcement learning algorithm and the first control parameter to obtain a first medication dosage parameter.

2. The method of claim 1, wherein, The acquiring of the first physiological indicator parameter of the target patient comprises: acquiring the first physiological indicator parameter of the target patient at each medication time of the target patient.

3. The method of claim 1 or 2, wherein, The method further comprises: acquiring a first time of the most recent medication of the target patient; determining a time interval between a current time and the first time; when the time interval is greater than a first preset time length, executing the step of acquiring the first physiological indicator parameter of the target patient.

4. The method of claim 3, wherein, The method further comprises: counting a medication frequency of the target patient; determining the first preset time length according to the medication frequency.

5. The method of claim 3, wherein, The method further comprises: acquiring a set of physiological indicator parameters before the current time, the set of physiological indicator parameters comprising a plurality of historical physiological indicator parameters, each historical physiological indicator parameter corresponding to a collection time; fitting according to the plurality of historical physiological indicator parameters and the collection times thereof to obtain a first physiological indicator fitting straight line; the horizontal axis of the first physiological indicator fitting straight line is time, and the vertical axis is a physiological indicator parameter; acquiring a first slope of the first fitting straight line; determining the first preset time length according to the first slope.

6. The method of claim 1 or 2, wherein, The determining of the first control parameter of the first reinforcement learning algorithm according to the first physiological indicator parameter and the basic information comprises: determining a reference control parameter of the first reinforcement learning algorithm according to the basic information; determining a first body condition evaluation value corresponding to the first physiological indicator parameter; determining a first adjustment parameter according to the first body condition evaluation value; adjusting the reference control parameter according to the first adjustment parameter to obtain the first control parameter.

7. The method of claim 1 or 2, wherein, The method further comprises: configuring a corresponding drug according to the first medication dosage parameter; recording a second time after the target patient takes the drug; acquiring physiological indicator parameters of the target patient every preset time interval within a first time period after the second time to obtain a plurality of physiological indicator parameters, each physiological indicator parameter corresponding to a collection time; fitting according to the set of physiological indicator parameters and the collection time corresponding to each physiological indicator parameter to obtain a second physiological indicator fitting straight line, the horizontal axis of the second physiological indicator fitting straight line being time and the vertical axis being a physiological indicator parameter; acquiring a slope of the second physiological indicator fitting straight line to obtain a second slope; determining a first deviation between the second slope and a preset threshold; optimizing the first control parameter according to the first deviation to obtain a second control parameter.

8. The method of claim 7, wherein, The optimizing the first control parameter according to the first difference value comprises: determining a first optimization parameter corresponding to the first deviation degree; optimizing the first control parameter according to the first optimization parameter to obtain the second control parameter.

Citation Information

Cited By

  • Critical patient analgesic drug dosage dynamic regulation and control system based on reinforcement learning

    CN121528415A

  • Individualized formula dynamic adjustment and optimization method based on reinforcement learning

    CN122224427A