Reinforcement learning based method for individualized dosage dynamic optimization of low molecular heparin for cancer patients
By using improved CWGAN-GP to generate time-series data and the EA-DRL model, the problems of data imbalance and insufficient model generalization in the administration of low molecular weight heparin to cancer patients were solved, enabling individualized dosage optimization, reducing the risk of thrombosis and bleeding, and improving the safety and efficacy of treatment.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- SICHUAN CANCER HOSPITAL
- Filing Date
- 2026-02-02
- Publication Date
- 2026-04-21
AI Technical Summary
Existing technologies suffer from data imbalance in the administration of low molecular weight heparin to cancer patients, failing to generate time-series data that aligns with clinical realities. Traditional GAN generation models are unstable, have simplistic reward function designs, and cannot simultaneously balance thrombosis and bleeding risks. Furthermore, traditional reinforcement learning models lack generalization ability in dynamic scenarios, leading to inappropriate administration methods and insufficient safety and efficacy.
A modified Conditional Wasserstein Generative Adversarial Network (CWGAN-GP) is used to generate temporal synthetic data. A reinforcement learning environment is constructed by combining conditional constraints and gradient penalty mechanisms. A comprehensive reward function that integrates state recognition rewards and environment adaptive rewards is defined. An Environment Adaptive Deep Reinforcement Learning (EA-DRL) model is used to integrate a domain adversarial generalization neural network to optimize the dosage of low molecular weight heparin.
The generation of high-quality time-series data improved the stability and accuracy of model training, enabled personalized and precise drug administration, reduced the incidence of adverse events such as thrombosis and bleeding, and improved treatment efficacy and clinical drug safety.
Smart Images

Figure CN121617545B_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the fields of medical health and artificial intelligence technology, and in particular to a method for dynamic optimization of individualized dosage of low molecular weight heparin for cancer patients based on reinforcement learning. Background Technology
[0002] Cancer patients are 4-7 times more likely to develop venous thrombosis (VTE) than non-cancer patients, and VTE has become the second leading cause of death among cancer patients. Low molecular weight heparin (LMWH) is the first-line anticoagulant for thrombosis prevention in cancer patients and is widely used in clinical practice. However, cancer patients have complex coagulation systems, and more than 90% have abnormal coagulation parameters. The risk of thrombosis recurrence and massive bleeding is high during anticoagulation therapy, with incidence rates of 20.7% and 12.4%, respectively.
[0003] Currently, the dosage of anticoagulant wound healing (LMWH) is mainly determined based on body weight, fixed daily dose, or clinical experience, which is highly subjective and uncertain, resulting in an inappropriate medication rate as high as 80.04%. Furthermore, different patients respond significantly differently to the same dose. Existing research on machine learning for predicting anticoagulant dosage mainly focuses on citrate and warfarin, with very few studies specifically targeting LMWH. The few existing studies also have significant limitations: they only focus on the "effective dose" for thrombosis prevention, neglecting the trade-off of the "safe dose" for bleeding; they do not include key indicators such as anti-X factor activity; the sample size is small and external clinical validation has not been conducted, leading to insufficient model generalization ability; they only target a single type of LMWH, failing to achieve personalized dosing; and they lack model visualization, limiting clinical application.
[0004] Further analysis reveals significant shortcomings in existing technologies at the data processing level: In data related to LMWH treatment in cancer patients, the proportion of samples showing effectiveness with low doses, high bleeding risk, and specific tumor stages is extremely low, highlighting a significant data imbalance. Traditional data augmentation methods struggle to generate time-series data that accurately reflects clinical practice, and commonly used generative adversarial networks (GANs) suffer from pattern collapse (generated data has a unidirectional distribution, failing to cover all features of real data) and training instability (loss function fluctuates wildly, making convergence difficult), failing to meet the high-quality data requirements of dosage optimization models. At the model construction level, traditional reinforcement learning models struggle to adapt to the dynamic scenario of "individualized dosage optimization for cancer patients," failing to accommodate fluctuations in patient physiological indicators and changes in treatment stages, exhibiting insufficient generalization ability and poor decision robustness. Furthermore, existing reward function designs are simplistic, failing to simultaneously balance thrombosis and bleeding risks, and thus unable to accurately guide the model to learn the optimal dosage decision strategy. Additionally, the model is prone to getting trapped in local optima during training, affecting the accuracy of the final dosage decision.
[0005] Therefore, there is an urgent need for a method that can solve the problem of data imbalance, adapt to the dynamic clinical environment, balance the risks of thrombosis and bleeding, and combine individual dynamic indicators of patients to achieve precise optimization of LMWH dosage, so as to overcome the shortcomings of existing administration methods and improve the safety and effectiveness of anticoagulation therapy. Summary of the Invention
[0006] The core objective of this invention is to address the critical pain point of risk warning in medication for cancer-chronic disease comorbidities by providing a reinforcement learning-based method for dynamic optimization of individualized low molecular weight heparin (LMWH) dosage for cancer patients. This method combines real-time physiological, pathological, and laboratory indicators of patients to dynamically optimize the LMWH dosage, achieving individualized and precise drug administration, reducing the incidence of adverse events such as thrombosis and bleeding, and improving treatment efficacy and clinical drug safety.
[0007] To achieve the above objectives, the following technical solution is employed: a reinforcement learning-based method for dynamic optimization of individualized low-molecular-weight heparin dosage for cancer patients, comprising the following steps:
[0008] S1. Collect multidimensional clinical data of cancer patients and preprocess them; use an improved conditional Wasserstein generative adversarial network to enhance the preprocessed data and generate time-series synthetic data to supplement the scarce samples of low-dose effectiveness, high bleeding risk and special tumor stages.
[0009] S2. Based on the multi-dimensional clinical data and the temporal synthetic data, a reinforcement learning environment is constructed, including: defining a state space, with the selected individual dynamic indicators of patients as state variables; defining an action space, with the dosage of low molecular weight heparin (LMWH) as action variables, including dose-level actions and dose-adjustment direction actions; and defining a reward function, using a comprehensive reward function that integrates state recognition rewards and environment adaptive rewards.
[0010] S3. Construct an environment-adaptive deep reinforcement learning (EA-DRL) model as a dose optimization model; train and optimize the dose optimization model in the reinforcement learning environment until it converges and obtains the optimal dose decision strategy.
[0011] S4. Input the real-time dynamic clinical data of the patient to be decided into the trained dose optimization model, and output an individualized low molecular weight heparin (LMWH) dosage recommendation. This dosage recommendation is the optimal decision after comprehensively balancing thrombosis prevention and bleeding risk under the current condition.
[0012] Furthermore, in step S1, the collected multidimensional clinical data specifically includes: basic information, tumor-related indicators, physiological indicators, laboratory test indicators, medication history and past medical history collected based on the REDCap platform; wherein, the laboratory test indicators include complete blood count, coagulation indicators containing anti-X factor activity, liver and kidney function indicators and thromboelastography parameters.
[0013] Furthermore, in step S1, the improved Conditional Wasserstein Generative Adversarial Network achieves data augmentation through a dual mechanism of conditional constraints and gradient penalties. The conditional constraint mechanism involves using core clinical indicators representing patient characteristics as conditional variables, simultaneously inputting them into the generator and discriminator, to guide the generator to produce time-series synthetic data that is associated with the conditional variables and conforms to clinical reality. The gradient penalty mechanism involves introducing a gradient-based penalty term into the discriminator's loss function to force the discriminator to satisfy the Lipschitz continuity condition.
[0014] Furthermore, the discriminator's loss function for:
[0015] ;
[0016] Loss function of generator for:
[0017] ;
[0018] in, For conditional variables, it refers to the key feature vectors of cancer patients input into the generator and discriminator; It is a real data sample taken from a real clinical dataset of cancer patients; The probability distribution followed by real-world clinical datasets of cancer patients; For generator G under condition Below, based on random noise The generated synthetic data sample; For the discriminator to judge real samples Under given conditions The following discrimination results; To obtain from interpolation distribution Interpolated samples from the middle sample; It is the discriminator's interpolation of samples Under given conditions The following discrimination results; The discriminator analyzes the generated samples. Given the same conditions The following evaluation output; For discriminator output Relative to its input interpolation samples The gradient vector; δ is the gradient penalty coefficient; It represents the mathematical expectation.
[0019] Furthermore, in step S2, the state space is defined as follows: by using a weighted normalized mutual information feature selection method on the multi-dimensional clinical data, the values of M key features related to the individualized dosage optimization target at time t are selected and constructed as state variables. The key characteristics include: age, weight, tumor stage, anti-X factor activity, coagulation function indicators, liver and kidney function indicators, type of medication and treatment stage.
[0020] Furthermore, in step S2, the action space is defined as follows: the low molecular weight heparin dosing decision at time t is defined as an action variable. The action variable Actions based on discrete specific dose levels With discrete dose adjustment direction Together they constitute; wherein, the specific dose level action Within a preset clinically commonly used dose range, the dose adjustment direction is divided according to a preset gradient. This includes "increase", "maintain", or "decrease".
[0021] Furthermore, in step S2, the integrated reward function that combines state recognition reward and environment adaptation reward is used to simultaneously evaluate the effectiveness of dosage decisions in balancing thrombosis and bleeding risks, adapting to individual differences, and optimizing treatment efficiency. The state recognition reward is calculated based on the thrombosis-bleeding risk distance matrix obtained through support vector data description, and is used to reward or penalize the difference between the currently selected dose and the theoretically optimal dose. The environment adaptation reward is used to reward or penalize whether the current dose adjustment direction conforms to the expected adjustment direction determined by the patient's current state. The integrated reward function also includes auxiliary reward terms positively correlated with shortening hospital stays and reducing treatment costs, and risk penalty terms negatively correlated with the occurrence of thrombosis or bleeding events.
[0022] Furthermore, in step S3, the dose optimization model is based on a duel dual deep Q network architecture and incorporates a domain adversarial generalized neural network. The dose optimization model evaluates the long-term value of selecting a specific dosing dose in a given patient state by combining a state value function, an advantage function, and an environment-dependent value function. The environment-dependent value function is used to quantify the impact of individual patient environmental differences on dose decisions.
[0023] Furthermore, the training objective of the dose optimization model is to minimize the total training loss function, which includes a reinforcement learning loss based on temporal difference error and a domain adversarial loss to force the feature extraction layer to generate domain-invariant features.
[0024] Furthermore, in step S3, when training the dose optimization model, an Ornstein-Uhlenbeck process is used to generate time-dependent exploratory noise, and this noise is superimposed on the action output by the model to increase the exploratory nature of the strategy and avoid the training process from getting stuck in local optima.
[0025] Compared with the prior art, the present invention achieves the following beneficial effects:
[0026] 1. This invention proposes an improved Conditional Wasserstein Generative Adversarial Network (CWGAN-GP), which integrates Conditional GAN (CGAN) and Gradient Penalized Wasserstein GAN (WGAN-GP). This addresses the pattern collapse and training instability issues inherent in traditional GANs when generating small-sample, imbalanced temporal data. It generates temporal synthetic data that conforms to the original data distribution in the scenario of optimizing LMWH dose for cancer patients, and focuses on supplementing scarce samples such as those effective for low doses, high bleeding risk, and special tumor stages. This alleviates the negative impact of data imbalance on subsequent reinforcement learning model training and provides high-quality data support for model training.
[0027] 2. This invention improves the quality and training stability of synthetic data through a dual mechanism of "conditional constraints + gradient penalty". It introduces the conditional constraint mechanism of CGAN to enable the generator to generate time-series data that conforms to the specific characteristics of patients, avoiding the generation of irrelevant or contradictory samples. It uses the gradient penalty mechanism of WGAN-GP to replace the traditional WGAN weight pruning strategy, solving the problem of model capacity reduction caused by weight pruning and ensuring that the generated data is highly consistent with the distribution of real data.
[0028] 3. This invention establishes a comprehensive reward function that integrates State Recognition Reward (SIR) and Environment Adaptive Reward (EAR), while simultaneously balancing the risks of thrombosis and bleeding, accurately guiding the model to learn the optimal dosage decision strategy, and avoiding decision bias caused by a single reward function.
[0029] 4. This invention proposes an Environment Adaptive Deep Reinforcement Learning (EA-DRL) model, which is based on a duel dual deep Q network and integrates a domain adversarial generalization neural network. This model addresses the problems of insufficient generalization ability and poor decision robustness of traditional reinforcement learning models in dynamic clinical scenarios. It enables the model to adaptively adapt to individual differences among patients and environmental changes at different treatment stages, and ultimately outputs accurate and stable dosage decisions.
[0030] 5. This invention introduces the Ornstein-Uhlenbeck process to generate noise, which is then superimposed on the output action of the actor network to avoid the model training getting stuck in local optima, thereby further improving the accuracy and reliability of dosage decisions.
[0031] In summary, this invention aims to dynamically optimize the dosage of LMWH by combining real-time physiological, pathological, and laboratory indicators of patients, thereby achieving individualized and precise drug administration, reducing the incidence of adverse events such as thrombosis and bleeding, and improving treatment efficacy and clinical drug safety.
[0032] It should be understood that the description in the Summary of the Invention is not intended to limit the key or essential features of the embodiments of the present invention, nor is it intended to restrict the scope of the invention. Other features of the invention will become readily apparent from the following description. Attached Figure Description
[0033] The above and other features, advantages, and aspects of the various embodiments of the present invention will become more apparent from the accompanying drawings and the following detailed description. The drawings are provided for a better understanding of the invention and are not intended to limit the invention. In the drawings, the same or similar reference numerals denote the same or similar elements, wherein:
[0034] Figure 1 This is a flowchart illustrating the method for dynamic optimization of individualized low molecular weight heparin dosage for cancer patients based on reinforcement learning, as provided in an embodiment of the present invention.
[0035] Figure 2 This is a schematic diagram of the workflow of the individualized dose dynamic optimization system provided in this embodiment of the invention;
[0036] Figure 3 This is a schematic diagram of the Environment Adaptive Deep Reinforcement Learning (EA-DRL) model architecture according to an embodiment of the present invention. Detailed Implementation
[0037] To make the objectives, technical solutions, and advantages of the embodiments of the present invention clearer, the technical solutions of the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of the present invention, and not all embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of the present invention.
[0038] Furthermore, the term "and / or" in this article is merely a description of the relationship between related objects, indicating that three relationships can exist. For example, A and / or B can represent: A existing alone, A and B existing simultaneously, or B existing alone. Additionally, the character " / " in this article generally indicates that the preceding and following related objects have an "or" relationship.
[0039] Figure 1 This is a flowchart illustrating the method for dynamic optimization of individualized low molecular weight heparin dosage for cancer patients based on reinforcement learning, as provided in an embodiment of the present invention. Figure 2This is a schematic diagram of the workflow of the individualized dose dynamic optimization system provided in this embodiment of the invention. Figure 2 It demonstrates the complete workflow from data to decision, covering Figure 1 The system-level interactions of each step are shown. A reinforcement learning-based method for dynamic optimization of individualized low-molecular-weight heparin dosage for cancer patients includes the following steps:
[0040] S1. Collect multidimensional clinical data of cancer patients and preprocess them; use an improved conditional Wasserstein generative adversarial network to enhance the preprocessed data and generate time-series synthetic data to supplement the scarce samples of low-dose effectiveness, high bleeding risk and special tumor stages.
[0041] Step S1 is used to perform data acquisition, preprocessing, and data augmentation.
[0042] S1.1 Raw Data Acquisition:
[0043] A prospective cohort was built based on the REDCap platform to collect data on cancer patients' basic information, tumor-related indicators, physiological indicators, laboratory test indicators, medication history, and past medical history, forming the original dataset.
[0044] (1) Basic information: For example, patient's unique code, gender, age (years), height (cm), weight (kg), body mass index (BMI), ethnicity, date of admission, data collection time, etc.
[0045] (2) Tumor-related indicators: For example, tumor characteristics: primary tumor location (e.g., lung cancer, pancreatic cancer, breast cancer, etc.), pathological type (e.g., adenocarcinoma, squamous cell carcinoma, etc.), histological grade. Disease status: tumor stage determined according to the Recognition of Contributing Efficacy in Solid Tumors (RECIST) or the corresponding hematologic malignancy criteria (e.g., stage I, II, III, IV); whether it is active (yes / no, referring to the presence of radiological progression or new lesions); whether it invades organs or blood vessels (yes / no, such as tumor invasion of large blood vessels). Treatment information: current anti-tumor treatment regimen (chemotherapy, targeted therapy, immunotherapy, etc.), number of treatment cycles.
[0046] (3) Physiological indicators: such as body temperature (°C), heart rate (beats / min), respiratory rate (breaths / min), blood pressure (systolic / diastolic, mmHg), blood oxygen saturation (%), level of consciousness, pain score (such as NRS score), etc.
[0047] (4) Laboratory test indicators: Complete blood count: white blood cell count (WBC), absolute neutrophil count (ANC), hemoglobin (Hb), platelet count (PLT), etc. Coagulation indicators: prothrombin time (PT), activated partial thromboplastin time (APTT), international normalized ratio (INR), anti-X factor activity (Anti-Xa, unit: IU / mL), D-dimer, fibrinogen (FIB), etc. Liver and kidney function: alanine aminotransferase (ALT), aspartate aminotransferase (AST), total bilirubin (TBil), serum creatinine (Cr), estimated glomerular filtration rate (eGFR), etc. Thromboelastography parameters: reaction time (R value), clotting time (K value), angle (α angle), maximum amplitude (MA value), etc.
[0048] (5) Medication history: For example, the type of low molecular weight heparin currently and previously used (such as enoxaparin, nadroparin, etc.), dosage (mg / kg or IU), route of administration, start and end time; other anticoagulant / antiplatelet drugs used at the same time (such as warfarin, rivaroxaban, aspirin, etc.); recent surgical history (especially tumor-related surgery) and time.
[0049] (6) Past medical history and risk factors: For example, a clear history of venous thromboembolism (VTE) (type, time), arterial thrombosis; history of bleeding diseases (such as gastrointestinal bleeding, intracranial hemorrhage); underlying diseases such as hypertension, diabetes, chronic liver disease, and chronic kidney disease; smoking history and drinking history; family history of hereditary thrombophilia.
[0050] The REDCap platform enables mandatory data entry, logical validation, and real-time quality control of data fields, forming a structured raw dataset that provides a reliable source for subsequent analysis. The data items listed above are exemplary parameters collected in this embodiment of the invention to construct a comprehensive state space. In practical applications, the types and ranges of collected parameters can be adaptively increased or decreased according to specific clinical scenarios, data accessibility, and model optimization needs, and should not be construed as the sole limitation on the data composition required by this invention. It should be understood that this system has obtained explicit informed consent and necessary authorization from relevant users before collecting, processing, and using multi-dimensional clinical data, and all data processing activities comply with relevant laws and regulations on personal information protection.
[0051] S1.2, Preprocessing of the original dataset:
[0052] Outliers and missing values were removed, numerical data were standardized and normalized, and categorical data underwent encoding transformation. A weighted normalized mutual information feature selection method was used to screen features relevant to the dose optimization objective. The comprehensive evaluation of these features... With weighting coefficients Calculate using the following formula:
[0053] ;
[0054] Features to be evaluated The overall score is used to measure the importance of the feature to the dose optimization target; the higher the score, the more the feature should be retained. The i-th feature to be evaluated, such as the cancer patient's age, weight, anti-X factor activity, tumor stage, and other specific indicators. Features to be evaluated The correlation with the dose optimization objective quantifies the degree of influence of this feature on individualized dose decision-making for LMWH; a higher value indicates a stronger correlation. Features to be evaluated The weighting coefficients are used to balance the relationship between feature correlation and feature redundancy. Features to be evaluated Redundancy between the selected feature set and the selected feature set; a higher value indicates that the feature has more duplicate information with the selected features. Selected feature set The number of features contained therein. Selected feature set A certain feature Information entropy measures the uncertainty of a feature. Features to be evaluated Information entropy measures the uncertainty of a feature. Dosage optimization target With the features to be evaluated The mutual information quantifies the dependency between the two. The goal of dose optimization is to achieve precise individualized dose optimization of LMWH (while balancing thrombosis prevention and bleeding risk). Dosage optimization target Information entropy measures the uncertainty of the target variable.
[0055] Among them, feature correlation This is measured by calculating the normalized mutual information between this feature and the dose optimization target (usually represented by binary or multi-class labels such as whether an adverse event occurred or whether the target anti-Xa activity was achieved), i.e. Feature redundancy By calculating this feature and the set of selected features It is measured by the average mutual information of all features, that is: .
[0056] This method eliminates feature redundancy, enhances feature discrimination, and provides high-quality input for model training.
[0057] S1.3, Data Augmentation:
[0058] Step S1.3 proposes an improved Conditional Wasserstein Generative Adversarial Network (CWGAN-GP), which integrates Conditional GAN (CGAN) and Gradient Penalized Wasserstein GAN (WGAN-GP). This addresses the problems of pattern collapse (generated data distribution is singular and cannot cover all features of real data) and training instability (loss function fluctuates wildly and is difficult to converge) in the generation of small-sample, imbalanced temporal data by traditional GANs. It generates temporal synthetic data that conforms to the original data distribution in the scenario of optimizing LMWH dose for cancer patients, and focuses on supplementing scarce samples such as low-dose effectiveness, high bleeding risk, and special tumor stages, thereby mitigating the negative impact of data imbalance on the subsequent training of reinforcement learning models.
[0059] Specifically, this improved CWGAN-GP model achieves enhanced synthetic data quality and training stability through a dual mechanism of "conditional constraints + gradient penalty":
[0060] On the one hand, the conditional constraint mechanism of CGAN is introduced, using key characteristics of cancer patients (such as tumor stage, anti-X factor activity level, liver and kidney function classification, etc.) as conditional variables. Input to generator With discriminator In this process, the generator is able to specifically generate time-series data that conforms to the characteristics of specific patients. At this point, the generator... The mapping function is ( It is random noise. (For generator parameters), discriminator The evaluation function is ( For input data, (For discriminator parameters), conditional variables guide the generated data to maintain consistency in feature correlation with real patient data, avoiding the generation of irrelevant or contradictory samples (such as "advanced tumor + normal antibodies"). "Factor activity" and similar combined data that do not conform to clinical practice.
[0061] On the other hand, the gradient penalty mechanism of WGAN-GP is used to replace the traditional WGAN weight pruning strategy, solving the problem of model capacity reduction caused by weight pruning. By introducing a gradient penalty term into the discriminator loss function, the discriminator is forced to satisfy the 1-Lipschitz continuity condition, making the model training process more stable and ensuring that the distribution of generated data is closer to the distribution of real data. The discriminator loss function and generator loss function of this model are defined as follows:
[0062] Wherein, the loss function of the discriminator for:
[0063] ;
[0064] Wherein, the generator's loss function for:
[0065] ;
[0066] The loss function value of the discriminator is used as a parameter to optimize the discriminator. It measures the discriminator's ability to distinguish between real data and generated data. The smaller the value, the stronger the discriminator's ability to distinguish between real data and generated data. The loss function value of the generator is used as a parameter to optimize the generator and measures how realistic the data generated by the generator is. The smaller the value, the closer the data synthesized by the generator is to the real data distribution. : Mathematical expectation, which represents the average of the expression within parentheses over its random variable distribution. Real data samples taken from real patient clinical datasets, such as real clinical indicator data of cancer patients (e.g., age, anti-X factor activity, tumor stage, etc.). : The probability distribution followed by real clinical datasets of cancer patients. A conditional discriminator network takes data samples and a condition variable y as input and outputs a scalar to evaluate the performance of the input samples under given conditions. The probability or confidence level of belonging to the true data. The discriminator analyzes the input samples. (under the conditions) The closer the output value is to 1, the more the discriminator considers the sample to be real data. From interpolation distribution Interpolated samples from the middle sample. : The probability distribution of the interpolated data, which is obtained by connecting a line to the real samples With generated samples The result is obtained by uniform sampling along the straight line, i.e. ,in It follows a uniform distribution on the interval [0, 1]. Conditional generator networks This indicates that the input to the conditional generator network is random noise. and condition variables The output is a generated synthetic data sample, such as simulated clinical indicator data of cancer patients. The random noise vector in the input generator, usually sampled from a simple prior distribution (such as the standard normal distribution), is the source of randomness in generating synthetic data and is used to drive the generator to produce diverse synthetic data. Conditional variables refer to the key feature vectors of cancer patients (e.g., a combination of tumor stage, anti-X factor activity level, liver and kidney function grade, etc.) input into the generator and discriminator, which are used to guide the direction of data generation to be consistent with the specific patient characteristics. The gradient penalty coefficient is a hyperparameter greater than 0. It can be set to 10 based on the characteristics of clinical data. It is used to adjust the weight of the gradient penalty term in the total loss, balance the discriminator's discriminative ability and training stability, and ensure that the discriminator satisfies the 1-Lipschitz continuity condition. Discriminator output Relative to its input interpolation samples The gradient vector reflects the sensitivity of the discriminator output to changes in the input samples. : The L2 norm (i.e., Euclidean length) of the gradient vector of the discriminator, used to measure the magnitude of the gradient. : Gradient penalty term. This term is used to force the discriminator to satisfy the 1-Lipschitz continuity condition, that is, the L2 norm of its gradient should be close to 1 everywhere in the input space. This is a key mechanism for stable training of the WGAN-GP framework.
[0067] Discriminator Network Its function is to distinguish whether the input sample is real patient data or data synthesized by the generator. Generator network The function is in the condition Under guidance, based on random noise Generate synthetic samples that conform to the distribution of real data (such as supplementing scarce data on patients with low-dose efficacy and high bleeding risk).
[0068] In a preferred embodiment of the present invention, both the generator G and the discriminator D employ a neural network structure including a Long Short-Term Memory (LSTM) network layer to capture the temporal dependencies of patient clinical indicators. Conditional variables With random noise vector The input layers of the generator G are concatenated to form a single input to G; condition variables... Similarly, with input data (or generate data) The concatenation is performed at the input layer of the discriminator D. The model is trained using the Adam optimizer with a learning rate of 0.0001 to 0.0005, a batch size of 32 to 128, and a gradient penalty coefficient δ of 10.
[0069] S2. Based on the multi-dimensional clinical data and the temporal synthetic data, a reinforcement learning environment is constructed, including: defining a state space, with the selected individual dynamic indicators of patients as state variables; defining an action space, with the dosage of low molecular weight heparin (LMWH) as action variables, including dose-level actions and dose-adjustment direction actions; and defining a reward function, using a comprehensive reward function that integrates state recognition rewards and environment adaptive rewards.
[0070] Step S2 is used to build the reinforcement learning environment.
[0071] S2.1, Define the state space:
[0072] Specifically, the state space is defined as follows: by using a weighted normalized mutual information feature selection method on multi-dimensional clinical data, the values of M key features related to the individualized dosage optimization target at time t are selected and constructed as state variables. The specific process is as follows:
[0073] Using pretreated individual dynamic indicators of patients as state variables It covers factors such as age, weight, tumor stage, anti-X factor activity, coagulation function indicators, liver and kidney function, type of medication and treatment stage, and is expressed as follows:
[0074] ;
[0075] The state variables at time t, which comprehensively represent the patient's key physiological and pathological states at that time, serve as the input basis for the reinforcement learning model to make dosage decisions. The M key features selected after weighted normalization mutual information feature screening, such as age, weight, tumor stage, anti-X factor activity, prothrombin time, liver and kidney function indicators, type of medication and treatment stage, reflect the patient's physiological and pathological state in real time. : Time markers represent a specific moment in the patient's treatment process (such as the 3rd day of treatment, after the 5th test, etc.), reflecting the dynamic changes in the patient's condition; The number of key features, determined by the feature selection process, reflects the total number of core indicators related to LMWH dose optimization.
[0076] S2.2, Define the action space:
[0077] Specifically, the action space is defined as follows: the low molecular weight heparin dosing decision at time t is defined as the action variable. The action variable Actions based on discrete specific dose levels With discrete dose adjustment direction Together they constitute; among them, specific dosage levels of action. Within a preset clinically commonly used dose range, the dose adjustment direction is divided according to a preset gradient. This includes options such as "increase," "maintain," or "decrease." The specific process is as follows:
[0078] LMWH dosage was used as an action variable. Covering the range of commonly used clinical dosages (such as...) ), according to the preset gradient (e.g. The dose levels are divided, and the action space expression is as follows:
[0079] ;
[0080] : Action variable at time t, i.e., the LMWH dosing decision output by the reinforcement learning model at that time. The specific dose level action at time t is based on the clinically commonly used dose range (e.g., ...). ) according to the preset gradient (e.g. ) can be divided into specific dosage values, such as "0.5mg / kg" and "0.8mg / kg". The dose adjustment direction at time t includes three types: "increase," "maintain," and "decrease," used to indicate the adjustment trend of the dose relative to the previous time (e.g., from...). Adjust to hour, (For "increase").
[0081] S2.3, Define the reward function:
[0082] A comprehensive reward function integrating State Identification Reward (SIR) and Environment Adaptive Reward (EAR) is constructed to simultaneously evaluate the effectiveness of dosage decisions in balancing thrombotic and bleeding risks, adapting to individual differences, and optimizing treatment efficiency. The State Identification Reward (SIR) is calculated based on the thrombotic-bleeding risk distance matrix obtained through support vector data description and is used to reward or penalize the difference between the currently selected dose and the theoretically optimal dose. The Environment Adaptive Reward (EAR) rewards or penalizes whether the current dose adjustment direction aligns with the expected adjustment direction determined by the patient's current state. The comprehensive reward function also includes auxiliary reward terms positively correlated with shortened hospital stays and reduced treatment costs, and risk penalty terms negatively correlated with the occurrence of thrombotic or bleeding events. The specific formula for the comprehensive reward function is as follows:
[0083] Reward (Base length of stay - Actual length of stay) (Baseline treatment cost - actual treatment cost) Punishment for thrombosis Punishment for bleeding;
[0084] Among them, Reward: the comprehensive reward function value, is used to evaluate the merits of the reinforcement learning model in selecting the dosage action at time t. The higher the value, the more the dosage decision is in line with the individualized needs of the patient (balancing the risk of thrombosis and bleeding, optimizing treatment efficiency and cost).
[0085] Status recognition reward Based on the design of the distance between multiple risk modes, the formula is as follows:
[0086] ;
[0087] : Thrombosis-bleeding risk interval matrix, each element Representative dose The corresponding risk interval values are calculated using the support vector data description method. The matrix elements are the interval values between the thrombosis risk and the bleeding risk corresponding to different doses. The larger the interval, the better the risk balance effect of the dose. The risk interval corresponding to the optimal dose, i.e. Zhongyu The corresponding matrix elements are the optimal doses. The corresponding risk interval value. The specific dose level action selected by the model at time t (e.g.) ). The optimal dose for the current patient condition, i.e., the LMWH dosing dose that achieves the best balance between the risk of thrombosis and the risk of bleeding (optimized based on clinical guidelines and real patient data). Exponential functions are used to amplify differences in rewards or penalties, enhancing the model's ability to distinguish between optimal and non-optimal doses. Thrombosis-bleeding risk gap matrix The maximum value of all elements is used to standardize the penalty for non-optimal doses. It is a matrix Action with the currently selected dose The corresponding element value, i.e. .
[0088] In a preferred embodiment of the present invention, the thrombosis-bleeding risk distance matrix The construction method includes: based on historical patient data, for each dose *a*, constructing risk feature sets for thrombotic and bleeding events respectively; using support vector data description, calculating the minimum enclosing hypersphere of the two risk feature sets in the feature space; the matrix element D(a) is the distance between the centers of the two hyperspheres at that dose. A larger distance indicates a stronger ability of the dose to distinguish between the two risks in the feature space, i.e., a better risk balance potential.
[0089] Environmental Adaptive Rewards The guided model adapts to individual patient differences, and the formula is:
[0090] ;
[0091] The dose adjustment direction (increase, maintain, decrease) selected by the model at time t. Dosage adjustments should be tailored to the individual patient's characteristics and based on the patient's current condition (e.g., low anti-factor X activity). The adjustment direction is "increase", indicating abnormal liver and kidney function. The adjustment direction is set to "reduction". The weighting coefficient of auxiliary reward items related to the length of hospital stay is set according to the priority of "shortening the length of hospital stay" in clinical practice (e.g., 0.1-0.3). The weighting coefficient of auxiliary reward items related to treatment costs is set according to the priority of medical resource optimization (e.g., 0.05-0.2). The weighting coefficient for the thrombosis penalty item is set according to the clinical priority of "reducing the risk of thrombosis" (e.g., 0.8-1.2). The weighting coefficient for the bleeding penalty item is set according to the clinical priority of "reducing the risk of bleeding" (e.g., 0.8-1.2).
[0092] Among them, the direction of dose adjustment that conforms to the individual characteristics of the patient. This is determined based on pre-defined clinical rules and logic. For example: if a patient's current anti-X factor activity is below the lower limit of the target range, then... "Increase" indicates an increase; if liver and kidney function indicators (such as eGFR) decrease significantly, then... The target is "reduction"; if all indicators are within the target range, then... For "maintain". These rules are predefined during system initialization.
[0093] S3. Construct an environment-adaptive deep reinforcement learning (EA-DRL) model as a dose optimization model; train and optimize the dose optimization model in the reinforcement learning environment until it converges and obtains the optimal dose decision strategy.
[0094] Step S3 is used to train and optimize the reinforcement learning model.
[0095] The Environment Adaptive Deep Reinforcement Learning (EA-DRL) model proposed in this invention constructs an LMWH dose optimization model. Based on a duel-style dual deep Q-network framework, it integrates a domain adversarial generalization neural network. The core of this model is to solve the problems of insufficient generalization ability and poor decision robustness of traditional reinforcement learning models in the dynamic scenario of "individualized dose optimization for cancer patients" due to fluctuations in patients' physiological indicators and changes in treatment stages. The model achieves adaptive adaptation to individual differences among patients and environmental changes at different treatment stages, and ultimately outputs accurate and stable dose decisions.
[0096] like Figure 3 As shown, in a preferred embodiment of the present invention, the specific architecture and training mechanism of the Environment Adaptive Deep Reinforcement Learning (EA-DRL) model are as follows:
[0097] 1. Model Architecture:
[0098] The EA-DRL model uses a Dueling Double DQN as its backbone and integrates a domain discriminator to form a domain adversarial generalization mechanism. The entire model contains a shared feature extraction layer and three functional branches:
[0099] Shared feature extraction layer: Consists of several fully connected layers, with parameters as follows Its input is a state variable. The output is a high-dimensional feature representation. .
[0100] State value function : A parameter The defined fully connected network branch takes shared features as input. Output a scalar , indicating state Its inherent value.
[0101] Advantage function A fully connected network branch defined by parameter β, with shared features as input. Output a vector with the same dimension as the action space. Each element corresponds to a specific action. The advantage value.
[0102] Environment-dependent value function : A parameter The fully connected network branch is defined, and its input also consists of shared features. Output a scalar It is used to model state-added values that are independent of the baseline state value, resulting from patient-specific environments (such as different tumor types and differences in hospital treatment standards).
[0103] Domain discriminator: A separate neural network whose input is the feature representation output by the shared feature extraction layer. Its task is to determine which "domain" a feature originates from (e.g., from the dataset of hospital A or the dataset of hospital B). Through adversarial training, the feature extraction layer is forced to generate domain-invariant features, thereby improving the model's generalization ability in new domains (new hospitals, new patient groups).
[0104] The dose optimization model evaluates the long-term value of selecting a specific dosing dose in a given patient state by combining the state value function, the advantage function, and the environment-dependent value function. The environment-dependent value function quantifies the impact of individual patient environmental differences on dosing decisions. The overall Q-function formula is as follows:
[0105] ;
[0106] The model's total Q-function is used to evaluate the state at time t. Select action The expected cumulative reward comprehensively reflects the long-term value of the dosage decision. Patient state variables at time t include key features selected (such as age, anti-X factor activity, tumor stage, etc.). Dose-action variables at any given time, including specific dose levels and adjust direction . : Parameters of the model feature extraction layer, used to extract deep semantic information (such as the correlation between features) of patient state features. The fully connected layer parameters of the state value function affect the assessment of the baseline value of the patient's current state. The fully connected layer parameters of the dominance function affect the judgment of the relative superiority or inferiority of different doses. : Parameters of the environment-dependent value function, used to adapt to individual patient differences (such as different tumor types and underlying diseases). State-value function, outputs the patient's state at time t. The fundamental value (unrelated to specific actions) reflects the overall treatment prospects in this state. The dominance function outputs the state at time t. Select action The advantage value relative to other actions; the higher the value, the better the dosage action is compared to other options. The mean of the advantage function is the average of the advantage values over all possible actions. It is used to center the advantage function to highlight the differences between actions. : Environment-dependent value function, used to quantify the impact of individual patient environment (such as specific tumor type, abnormal liver and kidney function) on dosage decision, and enhance the model's adaptability to individualized scenarios.
[0107] 2. Domain-based adversarial generalization mechanism:
[0108] Domain definition: Training data is divided into different "domains" based on their source and meta-features that may introduce distributional differences. For example, data collected from different medical centers may be considered as different domains, or data from patients with different types of primary tumors (such as lung cancer and colorectal cancer) may be considered as different domains.
[0109] Network Structure: Following the shared feature extraction layer, a neighborhood discriminator branch is introduced. This neighborhood discriminator is an independent fully connected neural network whose input is the feature representation output from the shared feature extraction layer. Its task is to predict the "domain" label to which the feature sample belongs as accurately as possible.
[0110] Adversarial training objective: During training, the classification loss of the neighborhood discriminator is backpropagated to the shared feature extraction layer through a gradient reversal layer. This forces the shared feature extraction layer to not only learn features useful for dosage decisions during optimization, but also to as much as possible obfuscate the neighborhood discriminator, i.e., learn to extract feature representations that are insensitive to the "neighborhood" and have generalizability. The loss of the neighborhood discriminator... The standard cross-entropy loss function is used.
[0111] Overall training objective: The total training loss of the model. It consists of two parts: first, the temporal difference error loss L(θ) for reinforcement learning (used to optimize the dosage strategy); and second, the domain adversarial loss. (Used to optimize feature generalization). The two are combined through a weighted sum: , where λ is a hyperparameter that weighs the two tasks.
[0112] The training objective of the temporal difference error loss L(θ) in reinforcement learning is to minimize a loss function constructed based on the temporal difference error. This loss function is defined by calculating the difference between the Q-value estimated at the current time step of the Q-network output and the target Q-value calculated based on the next time step state, the target Q-network, and the immediate reward. Specifically, the loss function for model training is:
[0113] ;
[0114] Temporal difference error loss in reinforcement learning is used to estimate the loss of a Q-network and measures the deviation between the estimated Q-value and the target Q-value. A smaller value indicates more accurate model predictions. Its parameters are... . : The total set of parameters for the model, including the weights and biases of all network layers. E: Mathematical expectation, used to calculate the average value of the loss function on the training samples. Instant reward at any moment, i.e., the action performed. The reward value obtained afterwards (calculated by the comprehensive reward function). Target Q network: used to generate stable target Q values and avoid fluctuations during training. : Estimation Q-network, used to estimate Q-values and update parameters in real time. : The set of parameters for the target Q-network. : Estimate the set of parameters for the Q-network. Patient status at any given time (performing actions) The next state after that. : The specific dosage level and adjustment direction at any given time. The action that maximizes the estimated Q-network output value (specific dose level and adjustment direction) is the optimal candidate action in the current state.
[0115] Loss of the neighborhood discriminator The cross-entropy loss function, a standard feature in classification tasks, is used to measure the discrepancy between the domain discriminator's predictions and the true domain labels. Suppose the training data is divided into N distinct domains (e.g., datasets from different hospitals or different tumor types). For the feature representations output by the shared feature extraction layer... Its actual domain labels use one-hot encoded vectors. It means that among them and Specifically, if The true source is the k-th domain, then ,the remaining (j≠k) are all 0. The neighborhood discriminator is based on the input. The predicted neighborhood probability distribution is an N-dimensional vector. ,in This indicates the predicted probability that the sample belongs to the j-th domain, and Then the domain combat loss. Defined as:
[0116] ;
[0117] in, It represents the natural logarithm.
[0118] Loss of the neighborhood discriminator Its parameters are those of the domain discriminator itself (denoted as φ), and the loss function is used to train the domain discriminator to accurately distinguish the domain from which features originate. Simultaneously, through a gradient reversal layer, The gradient is used to update the parameters ω of the shared feature extraction layer, thereby driving the layer to learn to generate domain-insensitive feature representations. This means achieving domain-specific adversarial generalization.
[0119] The weight coefficient λ of the domain adversarial loss is used to balance the reinforcement learning task and the domain generalization task. In the early stage of training, a smaller λ (e.g., 0.1) can be set to prioritize learning the dosage strategy; in the later stage of training, λ is gradually increased (e.g., up to 1.0) to enhance the model's ability to extract domain-invariant features. The specific scheduling strategy for λ can be determined through validation set performance.
[0120] Environment-dependent value function It is a fully connected network branch defined by parameter γ. It shares the same input features as the state value function V and the advantage function J. . The output is a scalar that adaptively adjusts the baseline state value to reflect the inherent value shifts caused by patient-specific circumstances, such as specific complications or rare tumor subtypes. This value is added directly to the baseline Q value calculated from V and J, as shown in the total Q function formula, thus enabling the final value assessment to dynamically adapt to individualized environmental differences.
[0121] Training methods: parameters Parameters of shared feature extraction layer and value function branch And the parameters φ of the domain discriminator, together during model training, are minimized by minimizing the total loss. Perform end-to-end optimization. Domain adversarial training promotes... These become domain-invariant features, while the K branch is responsible for resolving environment-sensitive added values from these general features.
[0122] Furthermore, when training the dose optimization model, an Ornstein-Uhlenbeck process is used to generate time-dependent exploratory noise, which is then superimposed on the actions output by the model to increase the exploratory nature of the strategy and avoid the training process from getting stuck in local optima.
[0123] Noise is generated by introducing an Ornstein-Uhlenbeck process. During the training phase, this time-related noise is superimposed on the action selected by the argmax operation of the Q function to increase the exploratory nature of the policy and avoid getting trapped in local optima during training. The formula is as follows:
[0124] ;
[0125] : The noise value at any given moment is used to increase the randomness of action selection. : Noise value at time t. Mean reversion rate, controlling noise towards the mean. The convergence speed (e.g., set to 0.15 to ensure the noise does not deviate excessively from a reasonable range). : The mean of the noise (e.g., set to 0, so that the noise fluctuates around zero). : Noise variance (e.g., set to 0.2 to control the dispersion of noise and balance exploration and utilization). Random noise that follows a standard normal distribution (mean 0, variance 1) provides a source of randomness for noise sequences.
[0126] During training, the model selects the drug dosage action based on the patient's current state, and the environment provides corresponding reward values. The parameters are updated using gradient descent to continuously optimize the action selection strategy until the model converges and the optimal dosage decision strategy is obtained.
[0127] S4. Input the real-time dynamic clinical data of the patient to be decided into the trained dose optimization model, and output an individualized low molecular weight heparin (LMWH) dosage recommendation. This dosage recommendation is the optimal decision after comprehensively balancing thrombosis prevention and bleeding risk under the current condition.
[0128] In step S4, dynamic data during the patient's treatment process is collected, input into the trained model, and an individualized LMWH (Less-than-Hyperhydrone) dosage recommendation is output. This recommended dosage is the optimal decision obtained after simultaneously optimizing the goals of "preventing thrombosis" (effectiveness) and "avoiding bleeding" (safety) during model training, thereby achieving dynamic and precise drug delivery.
[0129] It should be noted that the various embodiments in this specification are described in a progressive manner, with each embodiment focusing on the differences from other embodiments. Similar or identical parts between embodiments can be referred to interchangeably. For the apparatus disclosed in the embodiments, since it corresponds to the method disclosed in the embodiments, the description is relatively simple, and relevant parts can be referred to the method section.
[0130] It should also be noted that, in the embodiments of this application, relational terms such as "first" and "second" are used only to distinguish one entity or operation from another entity or operation, and do not necessarily require or imply any such actual relationship or order between these entities or operations. Furthermore, the terms "comprising," "including," or any other variations thereof are intended to cover non-exclusive inclusion, such that a process, method, article, or apparatus that comprises a list of elements includes not only those elements but also other elements not expressly listed, or elements inherent to such a process, method, article, or apparatus. Without further limitations, an element defined by the phrase "comprising one..." does not exclude the presence of other identical elements in the process, method, article, or apparatus that includes said element.
[0131] The above description of the disclosed embodiments enables those skilled in the art to make or use this application. Various modifications to these embodiments will be readily apparent to those skilled in the art, and the general principles defined in the embodiments of this application may be implemented in other embodiments without departing from the spirit or scope of this application. Therefore, this application is not to be limited to the embodiments shown in this application, but is to be accorded the widest scope consistent with the principles and novel features disclosed in the embodiments of this application.
Claims
1. A method for dynamic optimization of individualized low molecular weight heparin dosage for cancer patients based on reinforcement learning, characterized in that, Includes the following steps: S1. Collect and preprocess multidimensional clinical data from cancer patients; The preprocessed data was augmented using an improved conditional Wasserstein generative adversarial network to generate time-series synthetic data to supplement scarce samples with low-dose efficacy, high bleeding risk, and specific tumor stages. In step S1, the improved conditional Wasserstein generative adversarial network achieves data augmentation through a dual mechanism of conditional constraints and gradient penalties. The condition constraint mechanism is as follows: core clinical indicators that characterize patient features are used as condition variables and simultaneously input into the generator and discriminator to guide the generator to produce time-series synthetic data that is associated with the condition variables and conforms to clinical reality. The gradient penalty mechanism is as follows: a gradient-based penalty term is introduced into the loss function of the discriminator to force the discriminator to satisfy the Lipschitz continuity condition. S2. Based on the multi-dimensional clinical data and the temporal synthetic data, a reinforcement learning environment is constructed, including: defining a state space, with the selected individual dynamic indicators of patients as state variables; defining an action space, with the dosage of low molecular weight heparin (LMWH) as action variables, including dose-level actions and dose-adjustment direction actions; and defining a reward function, using a comprehensive reward function that integrates state recognition rewards and environment adaptive rewards. S3. Construct an environment-adaptive deep reinforcement learning (EA-DRL) model as a dose optimization model; train and optimize the dose optimization model in the reinforcement learning environment until it converges and obtains the optimal dose decision strategy. In step S3, the dose optimization model is based on a duel dual deep Q network architecture and incorporates a domain adversarial generalization neural network. The dosage optimization model evaluates the long-term value of selecting a specific dosing dose in a given patient state by combining the state value function, the advantage function, and the environment-dependent value function. The environment-dependent value function is used to quantify the impact of individual patient environmental differences on dosage decisions. S4. Input the real-time dynamic clinical data of the patient to be decided into the trained dose optimization model, and output an individualized low molecular weight heparin (LMWH) dosage recommendation. This dosage recommendation is the optimal decision after comprehensively balancing thrombosis prevention and bleeding risk under the current condition.
2. The method for dynamic optimization of individualized low molecular weight heparin dosage for cancer patients based on reinforcement learning according to claim 1, characterized in that, In step S1, the collected multidimensional clinical data specifically includes: basic information, tumor-related indicators, physiological indicators, laboratory test indicators, medication history and past medical history collected based on the REDCap platform; wherein, the laboratory test indicators include complete blood count, coagulation indicators containing anti-X factor activity, liver and kidney function indicators and thromboelastography parameters.
3. The method for dynamic optimization of individualized low molecular weight heparin dosage for cancer patients based on reinforcement learning according to claim 2, characterized in that, loss function of discriminator for: ; Loss function of generator for: ; in, For conditional variables, it refers to the key feature vectors of cancer patients input into the generator and discriminator; It is a real data sample taken from a real clinical dataset of cancer patients; The probability distribution followed by real-world clinical datasets of cancer patients; For generator G under condition Below, based on random noise The generated synthetic data sample; For the discriminator to judge real samples Under given conditions The following judgment results; To obtain from interpolation distribution Interpolated samples from the middle sample; It is the discriminator's interpolation of samples Under given conditions The following judgment results; The discriminator analyzes the generated samples. Given the same conditions The following evaluation output; For discriminator output Relative to its input interpolation samples The gradient vector; δ is the gradient penalty coefficient; It represents the mathematical expectation.
4. The method for dynamic optimization of individualized low molecular weight heparin dosage for cancer patients based on reinforcement learning according to claim 1, characterized in that, In step S2, the state space is defined as follows: By employing a weighted normalized mutual information feature selection method on the multidimensional clinical data, M key feature values at time t related to the individualized dosage optimization target are selected and constructed as state variables. ; The key characteristics include: age, weight, tumor stage, anti-X factor activity, coagulation function indicators, liver and kidney function indicators, type of medication and treatment stage.
5. The method for dynamic optimization of individualized low molecular weight heparin dosage for cancer patients based on reinforcement learning according to claim 1 or 4, characterized in that, In step S2, the action space is defined as follows: The decision to administer low molecular weight heparin at time t is defined as an action variable. The action variable Actions based on discrete specific dose levels With discrete dose adjustment direction Together they constitute; wherein, the specific dose level action Within a preset clinically commonly used dose range, the dose adjustment direction is divided according to a preset gradient. This includes "increase", "maintain", or "decrease".
6. The method for dynamic optimization of individualized low molecular weight heparin dosage for cancer patients based on reinforcement learning according to claim 5, characterized in that, In step S2, the integrated reward function that combines the state recognition reward and the environment adaptation reward is used to simultaneously evaluate the effectiveness of dose decision in balancing thrombosis and bleeding risks, adapting to individual differences, and optimizing treatment efficiency. The state recognition reward is calculated based on the thrombosis-bleeding risk distance matrix obtained by the support vector data description method, and is used to reward or punish the difference between the currently selected dose and the theoretical optimal dose. The environmental adaptive reward is used to reward or penalize whether the current dose adjustment direction conforms to the expected adjustment direction determined by the patient's current state; The comprehensive reward function also includes auxiliary reward items that are positively correlated with shortening hospital stays and reducing treatment costs, as well as risk penalty items that are negatively correlated with the occurrence of thrombotic or bleeding events.
7. The method for dynamic optimization of individualized low molecular weight heparin dosage for cancer patients based on reinforcement learning according to claim 1, characterized in that, The training objective of the dose optimization model is to minimize the total training loss function, which includes a reinforcement learning loss based on temporal difference error and a domain adversarial loss used to force the feature extraction layer to generate domain-invariant features.
8. The method for dynamic optimization of individualized low molecular weight heparin dosage for cancer patients based on reinforcement learning according to claim 1, characterized in that, In step S3, when training the dose optimization model, an Ornstein-Uhlenbeck process is used to generate time-dependent exploratory noise, and this noise is superimposed on the action output by the model to increase the exploratory nature of the strategy and avoid the training process from getting stuck in local optima.
Citation Information
Patent Citations
Diagnosis and treatment large model decision optimization method based on interaction feedback
CN118748075A
Adaptive radiotherapy dose optimization method and device, and storage medium
CN119896823A