Method and device for generating personalized levodopa prescription recommendation for Parkinson's disease
By employing a two-stage reinforcement learning approach, combined with virtual patient modeling and state transition models, a personalized levodopa prescription strategy is generated. This addresses the shortcomings of existing technologies in providing personalized and dynamic recommendations for Parkinson's disease medication, thereby improving the accuracy and adaptability of medication use.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- INST OF AUTOMATION CHINESE ACAD OF SCI
- Filing Date
- 2025-12-15
- Publication Date
- 2026-04-28
AI Technical Summary
Existing methods for recommending Parkinson's disease medications lack personalization and dynamism, fail to effectively address individual patient differences, and lack continuous temporal modeling and multi-step sequential optimization of the drug efficacy process, resulting in inaccurate and dynamistic recommendations.
A two-stage reinforcement learning approach is adopted. First, offline pre-training is performed through virtual patient modeling and state transition modeling. Then, online optimization is performed in model-free reinforcement learning to generate personalized levodopa prescription strategies, including dosage and decision interval.
It enables personalized and dynamic levodopa medication guidance, improving the accuracy and adaptability of prescriptions and meeting the long-term medication needs of Parkinson's disease patients.
Smart Images

Figure CN121938552A_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of artificial intelligence technology, and in particular to a method and apparatus for generating personalized levodopa prescription recommendations for Parkinson's disease. Background Technology
[0002] Currently, clinical treatment for Parkinson's disease primarily focuses on relieving motor symptoms, and levodopa is widely used at different stages of the disease due to its significant efficacy. However, long-term use often leads to complications such as motor fluctuations and dyskinesia, and the efficacy varies significantly among different patients.
[0003] Therefore, achieving personalized and dynamic levodopa medication is a significant challenge in the precision treatment of Parkinson's disease. On the one hand, existing research largely relies on drug combinations or clinical experience to provide patients with fixed prescriptions, failing to address the challenges of individual patient differences. On the other hand, some studies have attempted to use knowledge graphs and large models for medication recommendations. While these can cover multiple diseases, their reasoning logic is static, their interpretability is poor, and they struggle to model the continuous evolution of patient states. This makes it impossible to achieve personalized dynamic adjustments to medication through multi-step sequential decision-making, and compared to reinforcement learning, they lack the interactive ability to continuously optimize prescriptions.
[0004] This shows that the recommended methods for Parkinson's disease medication in related technologies have a technical problem of poor targeting. Summary of the Invention
[0005] This invention provides a method and apparatus for generating personalized levodopa prescription recommendations for Parkinson's disease, which addresses the shortcomings of existing Parkinson's disease medication recommendation methods in terms of their poor targeting, and enables the provision of more scientific, personalized and dynamic levodopa medication guidance for Parkinson's disease patients.
[0006] This invention provides a method for generating personalized levodopa prescription recommendations for Parkinson's disease, comprising the following steps: Acquiring patient data, wherein the patient data includes the patient's static characteristics and historical medication data, the static characteristics including age, gender, and disease duration, and the historical medication data including levodopa dosage records; Modeling a virtual patient based on the patient data to obtain a virtual patient model, wherein the virtual patient model is used to simulate the pharmacokinetic behavior of the levodopa preparation in the patient and generate a motor symptom score; Inputting a preset medication behavior, the motor symptom score, and the patient data into a state transition model to obtain a prediction result of symptom change after a target time output by the state transition model; Performing multi-stage reinforcement learning through the state transition model based on the patient data, the medication behavior, and the prediction result of symptom change after the target time to obtain a personalized levodopa prescription strategy of the reinforcement learning model, wherein the personalized levodopa prescription strategy is used to output a levodopa prescription recommendation based on the input patient data, the levodopa prescription recommendation including the dosage and decision interval of the levodopa preparation.
[0007] According to the present invention, a method for generating personalized levodopa prescription recommendations for Parkinson's disease includes, in which, based on the patient data, virtual patient modeling is performed to obtain a virtual patient model. This includes: acquiring pharmacokinetic parameters of different types of levodopa preparations, wherein the pharmacokinetic parameters include time to peak concentration, peak concentration, and half-life; performing levodopa pharmacokinetic modeling based on the historical medication data and the pharmacokinetic parameters to obtain the total equivalent concentration of levodopa in the blood; performing blood-central nervous system transport modeling based on the total equivalent concentration of levodopa and the static characteristics to obtain the levodopa concentration in the central nervous system; and performing modeling of the patient's motor function response based on the levodopa concentration in the central nervous system to obtain a motor symptom score, wherein the motor symptom score includes bradykinesia and dyskinesia scores.
[0008] According to the present invention, a method for generating personalized levodopa prescription recommendations for Parkinson's disease includes a state transition model comprising a multilayer perceptron and a gated recurrent unit. The method involves inputting a preset medication behavior, the motor symptom score, and the patient data into the state transition model to obtain a predicted result of symptom change after a target time, as output by the state transition model. This includes: performing a nonlinear mapping on the preset medication behavior, the motor symptom score, and the patient data using the multilayer perceptron to obtain static features; performing time-dependent modeling on the preset medication behavior, the motor symptom score, and the patient data using the gated recurrent unit to obtain time-series features; concatenating and fusing the time-series features and the static features, and performing feature interaction and mapping using the multilayer perceptron to obtain a predicted result of symptom change after the target time.
[0009] According to the present invention, a method for generating personalized levodopa prescription recommendations for Parkinson's disease is provided. The action space of the reinforcement learning model includes: the dosage and decision interval of different types of levodopa preparations; the state space of the reinforcement learning model includes: the patient data, the medication behavior, and the predicted results of symptom changes after the target time; the reward function of the reinforcement learning model includes: a penalty for irregular medication, a reward for improved motor function, and a reward for maintaining motor function.
[0010] According to the method for generating personalized levodopa prescription recommendations for Parkinson's disease provided by the present invention, the irregular medication penalty term is represented by the following formula: The reward for improved motor function is expressed by the following formula: The motor function maintenance reward is represented by the following formula: in, This represents the penalty value for irregular medication use. Indicates the penalty coefficient. Indicates time in hours. Indicates the current medication time. Indicates the date of the last medication administration; This indicates the reward value for improved motor function. The reward coefficient representing improvement in bradykinesia. The bradykinesia score indicates the current duration of medication administration. The bradykinesia score indicates the time since the last medication administration. The reward coefficient representing the improvement in dyskinesia. Dyskinesia score indicating the current medication duration. Dyskinesia score indicating the time since the last medication administration; This indicates the reward value for maintaining motor function. This represents the penalty coefficient for maintaining bradykinesia. This represents the penalty coefficient for maintaining dyskinesia.
[0011] According to the present invention, a method for generating personalized levodopa prescription recommendations for Parkinson's disease includes the following steps: obtaining a reinforcement learning model by performing multi-stage reinforcement learning based on the state transition model using the patient data, the medication behavior, and the predicted symptom changes after the target time; generating an initial dosing strategy by performing offline pre-training based on the state transition model; and fine-tuning the initial dosing strategy by deploying it in the environment of the virtual patient model using model-free reinforcement learning to obtain a personalized levodopa prescription strategy.
[0012] This invention also provides a device for generating personalized levodopa prescription recommendations for Parkinson's disease, comprising the following modules: an acquisition module for acquiring patient data, wherein the patient data includes the patient's static characteristics and historical medication data, the static characteristics including age, gender, and disease duration, and the historical medication data including levodopa medication records; a modeling module for performing virtual patient modeling based on the patient data to obtain a virtual patient model, wherein the virtual patient model is used to simulate the pharmacokinetic behavior of the levodopa preparation in the patient's body and generate a motor symptom score; and a prediction module for predicting the pre-set medication behavior... The system inputs the motor symptom score and the patient data into a state transition model to obtain a predicted result of symptom change after a target time, output by the state transition model. A strategy module is used to perform multi-stage reinforcement learning based on the patient data, medication behavior, and the predicted result of symptom change after the target time through the state transition model, to obtain a personalized levodopa prescription strategy for the reinforcement learning model. This personalized levodopa prescription strategy is used to output a levodopa prescription recommendation based on the input patient data, and the levodopa prescription recommendation includes the dosage and decision interval of the levodopa preparation.
[0013] The present invention also provides an electronic device, including a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein the processor executes the computer program to implement the method for generating personalized levodopa prescription recommendations for Parkinson's disease as described above.
[0014] The present invention also provides a non-transitory computer-readable storage medium having a computer program stored thereon, which, when executed by a processor, implements the method for generating personalized levodopa prescription recommendations for Parkinson's disease as described above.
[0015] The present invention also provides a computer program product, comprising a computer program that, when executed by a processor, implements the method for generating personalized levodopa prescription recommendations for Parkinson's disease as described above.
[0016] The present invention provides a method and apparatus for generating personalized levodopa prescription recommendations for Parkinson's disease. First, by acquiring the patient's static characteristics (such as age, gender, and disease duration) and historical medication data, a comprehensive individualized basis is provided for personalized recommendations, ensuring that prescription suggestions are based on the patient's specific situation. Next, based on a virtual patient model, the patient's static characteristics (such as age, gender, and disease duration), historical medication data, and changes in motor function scores are acquired, providing a comprehensive individualized basis for personalized recommendations and ensuring that prescription suggestions are based on the patient's specific situation. Then, by integrating preset medication behaviors, motor symptom scores, and patient data through a state transition model, the change in symptoms after a target time is predicted, enabling a prospective assessment of medication effectiveness and helping to identify the optimal intervention time. Finally, multi-stage reinforcement learning is used to optimize the prescription strategy based on the above prediction results, generating personalized prescription recommendations that can automatically output levodopa dosage and decision intervals based on input patient data, thereby significantly improving the accuracy, adaptability, and therapeutic effect of the prescription. Attached Figure Description
[0017] To more clearly illustrate the technical solutions in this invention or the prior art, the drawings used in the description of the embodiments or the prior art will be briefly introduced one by one below. Obviously, the drawings described below are some embodiments of this invention. For those skilled in the art, other drawings can be obtained from these drawings without creative effort.
[0018] Figure 1 This is a flowchart illustrating the method for generating personalized levodopa prescription recommendations for Parkinson's disease provided by the present invention.
[0019] Figure 2 This is a pharmacokinetic simulation diagram of three levodopa preparations provided by the present invention.
[0020] Figure 3 This is a schematic diagram of the GRU–MLP hybrid state transition model for predicting changes in patient clinical status provided by the present invention.
[0021] Figure 4 This is a network framework diagram of the reinforcement learning model provided by the present invention.
[0022] Figure 5 This is a flowchart of the control scheme for generating personalized levodopa prescription recommendations for Parkinson's disease provided by the present invention.
[0023] Figure 6This is a comparison chart of the pharmacokinetic simulations of levodopa IR, CR and Rytary in the blood and central nervous system provided by this invention, as well as the prediction and observation results of virtual patient motor symptoms based on the state transition model.
[0024] Figure 7 The results show the comparison between levodopa concentration simulation and motor symptom scores under the reinforcement learning strategy provided by this invention.
[0025] Figure 8 This is a schematic diagram of the module of the device for generating personalized levodopa prescription recommendations for Parkinson's disease provided by the present invention.
[0026] Figure 9 This is a schematic diagram of the physical structure of the electronic device provided by the present invention. Detailed Implementation
[0027] To make the objectives, technical solutions, and advantages of this invention clearer, the technical solutions of this invention will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some, not all, of the embodiments of this invention. All other embodiments obtained by those skilled in the art based on the embodiments of this invention without creative effort are within the scope of protection of this invention.
[0028] This invention proposes a personalized levodopa prescription recommendation method for Parkinson's disease based on two-stage reinforcement learning: First, a model-based reinforcement learning method is proposed, utilizing simulated medication behavior and symptom response data from virtual patients to construct a state transition model for the first stage of prescription strategy learning. Second, a model-free reinforcement learning method is proposed, involving online interactive optimization with virtual patients to achieve the second stage of prescription strategy learning. This method ensures the safety and rationality of drug dosage and decision intervals during the interaction process while improving strategy optimization efficiency, providing Parkinson's disease patients with more scientific and dynamic medication guidance.
[0029] The following outlines typical solutions and characteristics of other existing technologies, which stand in stark contrast to the reinforcement learning prescription recommendation model of this application.
[0030] In the development of drugs for Parkinson's disease, some patents have attempted to enhance efficacy, slow disease progression, or reduce side effects through drug combinations. For example, an oral drug composition combining tapentamol and levodopa is used to treat Parkinson's disease and related disorders, aiming to reduce unresponsive periods, alleviate symptom fluctuations, and improve motor and non-motor symptoms in late-stage patients. A combination of levodopa and icariin has been validated in animal experiments to delay or treat dyskinesia caused by long-term levodopa use, leveraging the synergistic protective effects of traditional Chinese medicine active ingredients and small-molecule drugs. Pharmacologically, these approaches achieve enhanced efficacy or reduced side effects through compound drugs. Their advantage lies in the relatively direct pathway and potential for commercialization. However, their drawback is that they focus on static combinations of fixed components and dosages, making it difficult to achieve dynamic management of patients during actual medication. Therefore, although combination therapies have shown some efficacy in experiments and early clinical trials, they still cannot meet the long-term, dynamic, and personalized medication needs of Parkinson's disease patients.
[0031] In the field of Parkinson's disease management, some studies attempt to construct relationship networks between disease, symptoms, and medications based on knowledge graphs, or rely on medical databases for rule matching to generate recommended treatment plans for patients. For example, related technologies construct and utilize medical knowledge graphs to model the association between patient pathology data and drug information, thereby generating highly accurate personalized medication recommendations, reducing misuse of medications, and improving the intelligence level of medical decision-making. Related technologies utilize feature word matching models based on pathological diagnostic data and patient personal records to intelligently generate prescription recommendations, simplifying the prescription generation process, improving efficiency, and reducing repeated modifications due to physician errors. The advantage of such methods lies in their reliance on structured medical knowledge systems, possessing strong interpretability, and the ability to quickly cover known medication rules and adverse reaction information. However, because these methods are mainly based on static rules or graph paths for reasoning, they lack the ability to dynamically model the disease evolution process and individual drug efficacy responses, leading to a "one-size-fits-all" problem in clinical practice and difficulty in adapting to individual patient differences.
[0032] Some studies have attempted to use deep learning models to model historical medication data and clinical assessment information of patients with various diseases, aiming to automate and intelligently recommend prescriptions. For example, a prescription recommendation method and system for chronic obstructive pulmonary disease (COPD) based on an end-to-end deep learning model outputs a probabilistic prescription through steps such as symptom information vectorization, syndrome type prediction, and feature fusion. Another prescription recommendation method and system based on neural networks integrates patient complaints, examination results, and basic statistical information, combined with prescription information from similar patients, to generate more reasonable and accurate final prescriptions for doctors. A medication recommendation method and device based on supervised learning decomposes disease and symptom descriptions into medication questions and trains the model's medical reasoning ability, thereby analyzing medication requests and outputting more accurate drug recommendations. These models typically employ supervised learning strategies, relying on existing samples to train fixed mapping relationships. They cannot dynamically update strategies during interactions with new patients and lack the ability to model the medication-feedback-re-decision process, thus failing to achieve multi-step sequential decision-making and personalized adjustments. Meanwhile, in real clinical scenarios, patients' symptoms and responses are highly heterogeneous, and relying solely on a single information input makes it difficult to capture long-term efficacy fluctuations, which is not conducive to achieving the goal of continuous optimization of prescription recommendations.
[0033] Furthermore, related technologies provide a prescription recommendation model for Parkinson's disease patients, which is essentially a deep learning-based prescription recommendation method, and its shortcomings are consistent with the problems mentioned above. Related technologies have explored reinforcement learning-based prescription recommendation approaches in other disease scenarios, but they still have two shortcomings compared to this invention: First, such methods usually optimize strategies directly in clinical interactions, while this invention uses a "two-stage reinforcement learning" design to perform offline modeling and optimization based on data before the model has achieved stability, and only enters the interaction stage with patients after the model is reliable, thereby improving safety; Second, existing methods mostly focus on single-time decisions, while this invention can achieve sequential optimization of prescriptions on a 24-hour or even multi-day scale, that is, systematically model the "medication-feedback-re-decision" cycle, thereby supporting multi-step sequential decision-making and individualized dynamic adjustments.
[0034] Based on a review of existing technologies, current personalized medication recommendations for Parkinson's disease mainly suffer from the following problems: First, they primarily focus on drug combinations or universal prescription recommendations, relying mainly on fixed formulas, rule matching, or static data modeling, lacking sufficient consideration of individual patient differences; second, existing methods mostly remain at the single-step decision-making level, lacking a closed-loop mechanism of "medication-feedback-re-decision," and cannot achieve dynamic prescription adjustments based on long-term feedback; third, they do not consider the interaction safety of reinforcement learning models before they have achieved stability. These shortcomings result in insufficient personalization and dynamism in the recommendation results, making it difficult to meet the actual needs of Parkinson's disease patients for long-term medication. Therefore, this invention proposes a prescription recommendation method based on two-stage reinforcement learning, which, while ensuring the reliability of the prescription during the interaction process, achieves personalized levodopa medication guidance and online optimization, thereby filling a key gap in existing technologies.
[0035] In summary, while numerous patents and published literature exist regarding Parkinson's disease drug combinations, knowledge graph reasoning, deep learning modeling, and reinforcement learning prescription recommendation, few methods can simultaneously achieve precise adaptation to individual patient differences, continuous temporal modeling of the pharmacodynamic process, multi-step sequential optimization of prescriptions, and safety assurance during clinical interaction stages. Therefore, there remains significant room for improvement in personalized dynamic medication management for Parkinson's disease. The "two-stage reinforcement learning" prescription recommendation method proposed in this invention first establishes a state transition model through model-based reinforcement learning and performs offline pre-training to ensure the stability and safety of the strategy before interaction. Then, model-free reinforcement learning is used to achieve sequential optimization of the prescription under controlled conditions, thereby improving convergence efficiency while ensuring safety, and providing Parkinson's disease patients with more scientific, personalized, and dynamic levodopa medication guidance.
[0036] Figure 1 This is a flowchart illustrating the method for generating personalized levodopa prescription recommendations for Parkinson's disease provided by the present invention. Figure 1 As shown, the method includes the following steps.
[0037] Step 101: Obtain patient data, which includes the patient's static characteristics and historical medication data.
[0038] Static characteristics include age, sex, and disease duration, while historical medication data includes levodopa medication records.
[0039] In this embodiment of the invention, static features cover basic information such as age, gender, and disease course. Static features are used to reflect individual differences in patients and the stage of disease progression. Historical medication data includes medication records of levodopa preparations (such as L-dopa IR, L-dopa CR, and Rytary), specifically involving administration time, dosage type, and dosage history.
[0040] Step 102: Based on the patient data, perform virtual patient modeling to obtain the virtual patient model.
[0041] Among them, the virtual patient model is used to simulate the pharmacokinetic behavior of levodopa in patients and generate motor symptom scores.
[0042] In this embodiment of the invention, virtual patient modeling is performed on patient data based on physiological research patterns and pharmacological statistics to construct a virtual patient model capable of highly simulating an individual's dynamic response to levodopa preparations. The virtual patient model transforms the patient's static characteristics and medication records into quantifiable symptom outputs by simulating three physiological processes.
[0043] The metabolic process of levodopa in vivo was simulated. Due to the differences in release characteristics and onset time among commonly used levodopa formulations in clinical practice, the virtual patient model used a standardized levodopa equivalent dose for calculation. This method can accurately characterize the entire process of different formulations from oral ingestion and gastrointestinal absorption to entry into the bloodstream, and reflect the dynamic characteristics of their concentration in the blood over time. For example, immediate-release formulations have a rapid onset but short duration of action, while sustained-release formulations have the opposite effect.
[0044] The virtual patient model simulates the transport of drugs from the bloodstream to the central nervous system. This process is not instantaneous but involves a significant time lag. The model introduces a special mathematical function to simulate this "acceleration followed by deceleration" transport characteristic and takes into account individual patient differences, allowing virtual patients of different ages, genders, and medication histories to exhibit varying drug delivery efficiencies, thus more realistically reflecting the complexity of the organism.
[0045] The output of the virtual patient model is reflected in the quantitative scores of motor symptoms, mainly including bradykinesia and dyskinesia scores. The virtual patient model calculates changes in symptom scores in a non-linear manner based on the drug concentration reaching the central nervous system. When the drug concentration is below the effective threshold, bradykinesia symptoms worsen; when the concentration is within the ideal therapeutic window, symptoms are significantly improved; and if the concentration exceeds the safe upper limit, the risk of dyskinesia increases significantly. Each virtual patient's symptom evolution is based on their unique physiological parameters and medication history, ensuring the individualization and realism of the simulation.
[0046] Step 103: Input the preset medication behavior, motor symptom scores and patient data into the state transition model to obtain the predicted symptom change after the target time output by the state transition model.
[0047] The inputs to the state transition model include: the patient's static characteristics, medication behavior, and motor symptom scores; the patient's static characteristics include: age, sex, and disease duration; medication behavior includes the specific time of administration and the dosage of each levodopa preparation (such as L-dopa IR, CR, Rytary) on the day, and these records are further converted into idealized drug concentrations in the blood within a specific time window in the past; motor symptom scores mainly include the continuous changes in bradykinesia and dyskinesia scores within a set time range.
[0048] To overcome the limitations of traditional single-step state prediction in capturing the continuous evolution of Parkinson's disease symptoms, the state transition model employs an iterative prediction strategy. Unlike simple mappings that only predict the state at the next moment, this model can continuously simulate the symptom evolution trajectory across multiple future time steps based on the current state and medication behavior. This approach better aligns with the inherent laws governing the natural development of a patient's physiological state over time, effectively reducing the risk of error accumulation in long-sequence predictions.
[0049] The state transition model is implemented using an advanced GRU-MLP hybrid neural network architecture. This architecture consists of two main components: a gated recurrent unit module responsible for modeling the temporal input data, capturing the dynamic patterns and dependencies of motor symptom scores and drug concentrations over time; and a multilayer perceptron module performing nonlinear mapping and high-level feature extraction on all input features. Subsequently, the temporal features extracted by the GRU are concatenated and fused with the static features extracted by the MLP to ultimately generate accurate predictions of symptom changes at future target time points (e.g., two minutes later), namely, changes in bradykinesia scores and dyskinesia scores.
[0050] Step 104: Using the state transition model, multi-stage reinforcement learning is performed based on patient data, medication behavior, and the predicted changes in symptoms after the target time to obtain a personalized levodopa prescription strategy from the reinforcement learning model.
[0051] The personalized levodopa prescription strategy is used to output levodopa prescription recommendations based on the input patient data. The levodopa prescription recommendations include the dosage and decision interval of the levodopa preparation.
[0052] In this embodiment of the invention, a reinforcement learning model is constructed with the goal of generating personalized prescription strategies. The action space of the model is carefully designed and includes two types of decision variables: one is the dosage of three types of levodopa preparations (L-dopa IR, L-dopa CR, and Rytary), each of which provides multiple discrete dosage levels to choose from; the other is the time interval for the next dosing decision (i.e., the decision interval), which can be selected from multiple time points ranging from a few minutes to several hours.
[0053] The state vector of the reinforcement learning model integrates the patient's static features (such as age, gender, and disease duration), dynamic symptom features (such as bradykinesia scores, dyskinesia scores, and their trends over a past period), and daily medication history. Based on this, a comprehensive reward function is designed to guide the policy optimization direction. This function mainly considers three objectives: penalizing irregular medication behavior to encourage patients to adhere to a reasonable medication cycle; providing positive rewards for improvements in motor function symptoms; and rewarding the maintenance of good symptom levels. Through the weighted summation of multi-objective rewards, the model is guided to learn a prescription strategy that is both effective in relieving symptoms and safe and reliable.
[0054] The training process of the reinforcement learning model employs a unique two-stage strategy to balance learning efficiency and clinical safety. In the first stage, the model-based pre-training stage, the reinforcement learning agent does not directly interact with real virtual patients but learns offline in a simulated environment constructed by the aforementioned state transition model. The state transition model can predict symptom changes in the near future based on the input patient state and medication behavior, thus providing the agent with an efficient and low-risk learning environment. In this stage, the agent learns a relatively stable and safe initial prescription strategy through extensive interaction with this simulated environment. Subsequently, in the second stage, the model-free fine-tuning stage, the agent deploys the pre-trained strategy to a high-fidelity virtual patient model for online interaction. In this stage, the strategy is finely adjusted based on feedback obtained from interacting with a more realistic environment, thereby correcting potential modeling biases in the first stage's simulated environment and further improving the adaptability and effectiveness of the strategy in real dynamic situations.
[0055] Through the two-stage learning process described above, the reinforcement learning model ultimately arrives at a personalized levodopa prescription strategy. The essence of this personalized levodopa prescription strategy is an intelligent decision function that automatically outputs the optimal levodopa prescription recommendation based on the patient's state data at any given time. This recommendation includes not only the specific dosage of various levodopa preparations to be administered immediately, but also the suggested next decision interval.
[0056] Through the embodiments of this invention, firstly, by acquiring the patient's static characteristics (such as age, gender, and disease course) and historical medication data, a comprehensive individualized basis is provided for personalized recommendations, ensuring that prescription suggestions are based on the patient's specific situation. Next, based on this data, a virtual patient model is created to simulate the pharmacokinetic behavior of levodopa in the patient's body and generate a motor symptom score, thereby transforming real patient characteristics into a quantifiable dynamic model, providing reliable input for subsequent predictions. Then, by integrating preset medication behavior, motor symptom scores, and patient data through a state transition model, the amount of symptom change after a target time is predicted, achieving a prospective assessment of medication effectiveness and helping to identify the optimal intervention time. Finally, multi-stage reinforcement learning is used to optimize the prescription strategy based on the above prediction results, generating personalized prescription recommendations that can automatically output levodopa dosage and decision intervals based on input patient data, thereby significantly improving the accuracy, adaptability, and therapeutic effect of the prescription.
[0057] According to the present invention, a method for generating personalized levodopa prescription recommendations for Parkinson's disease is provided, which involves virtual patient modeling based on patient data to obtain a virtual patient model, including: Obtain pharmacokinetic parameters for different types of levodopa formulations, including time to peak concentration, peak concentration, and half-life. Based on historical medication data and pharmacokinetic parameters, a pharmacokinetic model of levodopa was performed to obtain the total equivalent concentration of levodopa in the blood. Based on the total equivalent concentration and static characteristics of levodopa, the blood-nerve central transport process of the drug was modeled to obtain the concentration of levodopa in the central nervous system. Based on the concentration of levodopa in the central nervous system, the patient's motor function response was modeled to obtain a motor symptom score, which included bradykinesia and dyskinesia scores.
[0058] In this embodiment of the invention, to assess patients' symptom responses to various levodopa formulations, a levodopa pharmacokinetic conversion model was constructed to uniformly convert the actual dosage of levodopa into a levodopa equivalent dose (LED). The LED is a standardized indicator widely used in clinical practice to assess the overall dopaminergic effect of Parkinson's disease drugs, and it can quantify the patient's pharmacodynamic response at any time point without considering differences in specific formulations.
[0059] Given the significant differences in release mechanisms, absorption rates, and final dilution processes among different formulations, this invention uses the following two formulas to calculate LED performance: in, Indicates the current time. This indicates the date of the patient's last medication use. This indicates the time required for the levodopa drug equivalent to reach its peak concentration. The peak conversion equivalent of a unit of levodopa is determined by the type of levodopa drug. This indicates the previous dose of levodopa.
[0060] The levodopa drugs used in this invention include L-dopa IR, L-dopa CR, and Rytary, and the parameters corresponding to their 100mg concentrations are shown in Table 1. After a patient ingests multiple levodopa drugs, the levodopa concentration in the blood meets the following requirements: in, , and The conversion factor for the three classes of levodopa drugs. , and LEDs are for three classes of levodopa drugs.
[0061] Table 1. Pharmacokinetic parameters of three classes of levodopa drugs
[0062] In this embodiment of the invention, a blood-central transport model was constructed to describe the dynamic process of levodopa administration, from absorption into the bloodstream to its final delivery to the target region of the central nervous system. Given the prevalent lag effect in drug transport within the body, this invention assumes that the transport process exhibits a nonlinear characteristic of "initially accelerating, then decelerating." To characterize this time-dependent transport delay effect, the blood-central transport model introduces a chi-square distribution kernel function to weight the levodopa concentration in the blood, as shown in the following formula: in, This indicates the real-time concentration of levodopa in the patient's central nervous system (CNS). This indicates the corresponding concentration in the blood. Represents the gamma function, where the parameter Characterize the time lag in drug delivery, which is affected by individual differences (such as age, gender, and medication history). This refers to the noise introduced during dopaminergic transport.
[0063] refer to Figure 2 , Figure 2 This is a pharmacokinetic simulation diagram of three levodopa preparations provided by the present invention.
[0064] Based on the aforementioned blood-central transport model, an 80-year-old female virtual subject with a 5-year history of Parkinson's disease was constructed, and she was set to take 100 mg doses of L-dopa IR, L-dopa CR, and Rytary at specific time intervals. The absorption process of these three types of drugs, and the dynamic concentration changes of the drugs in the blood and central nervous system (CNS) were studied. Figure 2 As shown.
[0065] To more comprehensively evaluate patients' motor function, the Bradykinesia Score and Dyskinesia Score are introduced to measure the symptom response characteristics of different patients under different CNS dopa equivalent concentrations. The lower the score, the better the patient's condition in that aspect, which can be summarized by the following formula: in, This represents the score for slow motion at time t. This represents the motion anomaly score at time t. and The change in score due to introduced noise and Related to the concentration of dopa equivalents in the central nervous system, satisfying: in, This indicates the concentration of levodopa equivalents in the central nervous system; =500ng / mL and =2000ng / mL represents the threshold concentration of L-dopa in the CNS: the former is the minimum effective concentration, and the latter is the upper limit of the risk of increased peak dose dyskinesia when the value is exceeded.
[0066] This invention incorporates parameters such as patient age, gender, and medication history to construct a levodopa pharmacokinetic model encompassing four processes: oral intake, gastrointestinal absorption, blood circulation, and effect manifestation. This model simulates the dynamic changes in drug concentration across different dosage forms (IR, CR, rytary) and introduces bradykinesia and dyskinesia scores as indicators for evaluating motor function. This virtual patient model not only serves as the primary source of data but also provides a near-realistic interactive environment, supporting the training and optimization of subsequent prescribing strategies.
[0067] Through the embodiments of the present invention, by integrating the pharmacokinetic parameters of multiple types of levodopa preparations, a dynamic model from blood drug concentration to central nervous system transport is established, and individual characteristics and motor symptom responses are correlated, thereby achieving accurate quantitative simulation of patient pharmacodynamic response and providing a high-fidelity, computable virtual experimental environment for personalized prescription recommendations.
[0068] According to the present invention, a method for generating personalized levodopa prescription recommendations for Parkinson's disease is provided, wherein the state transition model includes: a multilayer perceptron and a gated recurrent unit; By inputting pre-defined medication behaviors, motor symptom scores, and patient data into a state transition model, the predicted changes in symptoms after a target time, as output by the state transition model, are obtained, including: Static features are obtained by nonlinearly mapping preset medication behavior, motor symptom scores, and patient data using a multilayer perceptron. By using a gated loop unit, the preset medication behavior, motor symptom scores, and patient data are modeled in a time-dependent manner to obtain time-series features; The temporal and static features are concatenated and fused, and the features are interacted and mapped through a multilayer perceptron to obtain the predicted results of the symptom changes after the target time.
[0069] In clinical settings, policy learning often cannot directly obtain a complete model of a patient, but can only make inferences based on limited rehabilitation data. Therefore, in the early stages of policy training, a state transition model needs to be constructed using limited patient data, and this model interacts with reinforcement learning algorithms in a model-driven simulation environment. Once the model accuracy reaches the expected level, the optimized policy is then applied to patient interaction to reduce risk and improve policy convergence efficiency. Based on this idea, this invention first constructs an iterative state transition model to predict the dynamic evolution of symptoms in Parkinson's disease patients under drug intervention.
[0070] Traditional step-by-step state transition methods focus on predicting the process of state S1 transitioning to state S2 after executing action A1, i.e. However, the symptom evolution in PD patients is persistent and temporally correlated, making single-step prediction insufficient to fully characterize its dynamic features. Therefore, this invention employs an iterative prediction strategy, continuously simulating the state evolution process across multiple time steps in a single computation based on the model. This allows for a closer reflection of the patient's condition over time and effectively reduces the cumulative error in long-term predictions.
[0071] Based on the above analysis, the information available for the state transition model includes: the patient's static characteristics, such as age, gender and disease course; medication records (behavior), including the time and dosage of administration on the day; (iii) history of symptom changes, i.e. changes in bradykinesia and dyskinesia scores within a set time window.
[0072] Specifically, patient static characteristics and medication records remain unchanged between two adjacent decision points; the prediction target is the change in two symptom scores after the decision. To improve prediction accuracy, the medication record can first be converted into an idealized blood L-dopa concentration over the past 30 minutes; simultaneously, the difference between bradykinesia and dyskinesia scores is used as one of the inputs to the state transition model. The model output is defined as the change in the two scores two minutes later, denoted as... and .
[0073] in, This indicates the change in bradykinesia score and dyskinesia score after two minutes. Indicates the patient's static characteristics, Indicates from the current moment Rewinding to the ideal levodopa concentration in the patient's blood 30 minutes prior, Indicates a slow historical movement score. This indicates the historical dyskinesia score. This indicates the change in the historical movement slowness score. This indicates the change in historical dyskinesia scores.
[0074] Based on the above inputs and outputs, this invention constructs a network architecture consisting of one classifier and two feature extractors. First, a multilayer perceptron performs a nonlinear mapping on all input features to extract high-level representations. Second, a gated recurrent unit models the time-series input containing five variables, capturing its temporal dependencies and dynamic change patterns. Finally, the temporal features extracted by the gated recurrent unit are concatenated and fused with static features, and the multilayer perceptron further completes feature interaction and mapping, outputting the symptom change within the next two minutes. The prediction results.
[0075] refer to Figure 3 , Figure 3 This is a schematic diagram of the GRU–MLP hybrid state transition model for predicting changes in patient clinical status provided by the present invention.
[0076] In this invention, to achieve the initial learning of prescription strategies through model-based reinforcement learning in the first stage, a state transition model is constructed using simulated medication behavior and symptom response data from virtual patients to simulate the dynamic evolution of patient symptoms. Addressing the shortcomings of traditional "point-to-point" prediction methods, this invention proposes a hybrid network structure based on GRU–MLP to iteratively predict patient symptom changes at the minute-level, utilizing the continuity and temporal sequence of drug efficacy and symptoms to predict the subject's state. This method achieves an MSE of 1.22 × 10⁻⁶ on the virtual dataset. -4This significantly improves the ability to dynamically characterize the progression of Parkinson's disease.
[0077] According to the present invention, a method for generating personalized levodopa prescription recommendations for Parkinson's disease is provided, wherein the action space of the reinforcement learning model includes: the dosage of different types of levodopa preparations and the decision interval; The state space of a reinforcement learning model includes: patient data, medication behavior, and the predicted changes in symptoms after the target time. The reward function of a reinforcement learning model includes: a penalty for irregular medication use, a reward for improved motor function, and a reward for maintaining motor function.
[0078] This invention constructs a reinforcement learning model to generate personalized, dynamically adaptive levodopa dosing strategies for Parkinson's disease patients. This aims to balance bradykinesia and dyskinesia symptom control throughout the day, reducing the risks associated with overdose or underdose. To this end, the model's action space is designed to include the current dosage of the three levodopa formulations and the time interval for the next decision. The selectable values are shown in the table below: Table 2 Action Space of Reinforcement Learning Models
[0079] Subsequently, this invention models the patient's immediate state as a multi-dimensional vector, encompassing demographic information (age, gender, disease duration), recent physiological and symptom characteristics (bradykinesia and dyskinesia scores in the past 30 minutes), and symptom variability (…). The system can measure the patient's physiological and symptom status at a single point in time, and calculate the theoretical dopa concentration in the blood over the past 30 minutes based on the medication used that day, thus providing key information for strategic decision-making.
[0080] Based on the above, this invention designs the rewards and punishments of the reinforcement learning model from three aspects: regular medication use, improvement of motor function scores, and maintenance of motor function scores.
[0081] According to the method for generating personalized levodopa prescription recommendations for Parkinson's disease provided by the present invention, the penalty for irregular medication use is represented by the following formula: The reward for improved motor function is expressed by the following formula: The reward for maintaining motor function is expressed by the following formula: in, This represents the penalty value for irregular medication use. Indicates the penalty coefficient. Indicates time in hours. Indicates the current medication time. Indicates the date of the last medication administration; This indicates the reward value for improved motor function. The reward coefficient representing improvement in bradykinesia. The bradykinesia score indicates the current duration of medication administration. The bradykinesia score indicates the time since the last medication administration. The reward coefficient representing the improvement in dyskinesia. Dyskinesia score indicating the current medication duration. Dyskinesia score indicating the time since the last medication administration; This indicates the reward value for maintaining motor function. This represents the penalty coefficient for maintaining bradykinesia. This represents the penalty coefficient for maintaining dyskinesia.
[0082] In this embodiment of the invention, the penalty for irregular medication use is as follows: Based on clinical experience and needs, doctors often require patients to take medication regularly every 2-4 hours. Therefore, if the model's medication use time exceeds this range, a penalty will be applied based on the linear correlation between the model and the time exceeded. Motor function score improvement reward: If the simulation device observes an improvement in bradykinesia and dyskinesia scores over a period of time, a reward is given; otherwise, a penalty is imposed. Motor function score maintenance reward: If the simulation device observes that the bradykinesia score and dyskinesia score have remained at a good level over a period of time, a reward is given; otherwise, a penalty is given. The total immediate reward at the current time point is obtained by weighted summation of the three rewards. To reflect the impact of future rewards on current decisions and balance the differences in returns across different decision intervals, a time-dependent discounted return is introduced for cumulative calculation: Where γ represents the basic discount factor, This represents the time interval between two consecutive decisions. This is the time normalization constant. Regarding the model, this invention employs a two-stage Proximal Policy Optimization (PPO) framework.
[0083] refer to Figure 4 , Figure 4 This invention provides a reinforcement learning model network framework.
[0084] like Figure 4 As shown, PPO generates a dosing strategy through iterative policy updates and sets up a policy-value network that uses a shared backbone structure (multilayer perceptron) to extract state features, predict L-dopa IR, L-dopa CR, Rytary dose and decision interval, and includes a value output head for state assessment.
[0085] In order to achieve personalized prescription recommendations for levodopa through the embodiments of the present invention, the present invention uses the patient's static characteristics (age, gender, disease course), dynamic symptom characteristics (e.g., changes in bradykinesia and dyskinesia scores within 30 minutes, i.e., the predicted results of symptom changes after the target time), and daily medication as the state vector input of the reinforcement learning model, and designs a multi-objective reward function covering three aspects: regular medication, improvement of motor function, and maintenance of motor function, to guide the optimization of the strategy.
[0086] According to the present invention, a method for generating personalized levodopa prescription recommendations for Parkinson's disease involves multi-stage reinforcement learning using a state transition model based on patient data, medication behavior, and predictions of symptom changes after a target time. The resulting reinforcement learning model includes: An initial drug delivery strategy is generated through offline pre-training based on a state transition model. The initial dosing strategy was deployed in a virtual patient model environment and fine-tuned using model-free reinforcement learning to obtain a personalized levodopa prescription strategy.
[0087] In this embodiment of the invention, the training process is divided into two stages: Phase 1 (Model-based pre-training): The reinforcement learning agent first interacts with a symptom state transition model built from clinical data, which can predict the dynamic changes in patient symptoms, thus enabling the agent to learn an initial dosing strategy based solely on clinical data.
[0088] Phase Two (Model-Free Fine-Tuning): After the pre-training achieves a certain performance, the strategy is deployed in a high-fidelity virtual patient environment. The strategy is further modified and optimized by directly interacting with the virtual patient to improve its adaptability and effectiveness in real dynamic situations.
[0089] This invention proposes a personalized levodopa prescription recommendation framework for Parkinson's disease based on two-stage reinforcement learning. The first stage utilizes an offline pre-training environment constructed from simulated clinical data to obtain an initial prescription recommendation strategy. The second stage involves fine-tuning the strategy using model-free reinforcement learning in virtual patients to further correct modeling biases and improve strategy performance. This design balances safety and convergence speed, avoiding early high-risk exploration in real patients.
[0090] refer to Figure 5 , Figure 5 This is a flowchart of the control scheme for generating personalized levodopa prescription recommendations for Parkinson's disease provided by this invention. The virtual patient environment construction process is achieved sequentially through three steps: levodopa pharmacokinetic modeling, blood-nerve central transport process modeling, and motor function response modeling. The prescription recommendation model training process begins in the virtual patient environment, then constructs a patient state transition model based on data, and employs a two-stage training strategy: first, model-based reinforcement learning pre-training, and then fine-tuning of the prescription recommendation model based on model-free reinforcement learning.
[0091] Through the aforementioned technical approach, this invention achieves joint optimization of the dosage and next decision interval for three types of levodopa preparations, constructing a cross-time-step "decision-dosing-feedback-re-decision" optimization mechanism. This mechanism enables prescription strategies to be continuously updated and adaptively adjusted on daily or even multi-day scales, achieving multi-step sequential decision-making and continuous feedback regulation, thereby providing dynamic prescription recommendations and offering a sustainable and practical intelligent decision-making framework for long-term medication management of Parkinson's disease.
[0092] To address the shortcomings of existing personalized levodopa prescription recommendation methods, such as the inability to fully consider individual differences among Parkinson's patients and the inability to continuously optimize and dynamically adjust prescriptions on a 24-hour or multi-day scale, this invention proposes a personalized levodopa prescription recommendation method for Parkinson's disease based on two-stage reinforcement learning, and achieves the following technical effects and advantages.
[0093] In the pharmacokinetic analysis, the absorption and metabolism of three classes of levodopa formulations—IR, CR, and rytary—were considered simultaneously. Through equivalent LED conversion and blood-central nervous system transport modeling, a unified description of multiple formulations was achieved. Simulation results showed that the peak concentration, half-life, and duration of action of the three classes of formulations were reasonably reflected. The correlation between drug concentration curves and motor symptom scores was consistent with clinical observations, providing a solid foundation for joint optimization across formulations.
[0094] The proposed GRU–MLP hybrid state transition model fully integrates temporal dependency modeling and nonlinear feature interaction, enabling a more accurate characterization of symptom evolution in Parkinson's patients under different drug administration conditions. Training results on 14,700 state transition samples show that the model reaches early cessation at 70 epochs, exhibiting [the following characteristics]. , , The predicted curves are highly consistent with the actual symptom changes, demonstrating that the model has a strong ability to capture long-term trends and fit short-term fluctuations.
[0095] refer to Figure 6 , Figure 6 This is a comparison chart of the pharmacokinetic simulations of levodopa IR, CR and Rytary in the blood and central nervous system provided by this invention, as well as the prediction and observation results of virtual patient motor symptoms based on the state transition model.
[0096] The reinforcement learning component employs a two-stage framework of "model-driven pre-training and model-free fine-tuning." First, an initial policy is learned using a data-driven state transition model, increasing the cumulative reward from 67.08 to 105.46. Subsequently, further interaction in a virtual patient environment leads to a stable increase in the reward value to 111.70 without significant regression. This divide-and-conquer training process effectively avoids the potential risks of early direct interaction with patients, ensuring the stability and safety of policy convergence.
[0097] This invention introduces a sequential decision-making mechanism into prescription recommendation strategies, constructing a time-step sequential optimization framework of "decision-medication-feedback-re-decision." Unlike traditional methods that rely on single-step static recommendations, this mechanism evaluates the subsequent impact of each decision action through long-term reward estimation using reinforcement learning, thereby learning the optimal long-term prescription strategy through multiple rounds of interaction. In strategy application testing, an 80-year-old female virtual patient showed a rapid increase in blood and CNS levodopa concentrations under the reinforcement learning strategy, which remained within a narrow range of fluctuation. Correspondingly, bradykinesia and dyskinesia scores remained at low levels, demonstrating that this invention not only improves motor function but also effectively reduces the risk of peak dyskinesia. The results indicate that the sequential optimization mechanism proposed in this invention can achieve more scientific, forward-looking, and globally optimal medication prescription recommendations while considering both long-term efficacy and short-term stability.
[0098] refer to Figure 7 , Figure 7 The results show the comparison between levodopa concentration simulation and motor symptom scores under the reinforcement learning strategy provided by this invention.
[0099] The state representation vector integrates patient static characteristics (age, gender, disease duration), dynamic symptom characteristics (changes in bradykinesia and dyskinesia scores over 30 minutes), and drug concentration information, thus ensuring the policy's adaptability to heterogeneous patient populations. Furthermore, a time-normalized discount factor is used to handle unequal interval decisions, enabling the model to demonstrate robust policy optimization capabilities across different medication cycles.
[0100] The two-stage reinforcement learning framework proposed in this invention is not only applicable to the three commonly used levodopa formulations, but can also be extended to other disease drugs or combination drug scenarios. Its three-layer architecture of virtual patient-state transition model-reinforcement learning provides an interface and path for the subsequent introduction of multimodal monitoring data (such as wearable devices), adaptive safety boundary constraints, and personalized long-term follow-up management, and has high potential for engineering and clinical application.
[0101] The following describes the apparatus for generating personalized levodopa prescription recommendations for Parkinson's disease provided by the present invention. The apparatus for generating personalized levodopa prescription recommendations for Parkinson's disease described below can be referred to in correspondence with the method for generating personalized levodopa prescription recommendations for Parkinson's disease described above.
[0102] refer to Figure 8 , Figure 8 This is a schematic diagram of the module of the device for generating personalized levodopa prescription recommendations for Parkinson's disease provided by the present invention.
[0103] The acquisition module 801 is used to acquire patient data, which includes the patient's static characteristics and historical medication data. The static characteristics include age, gender and disease course, and the historical medication data includes medication records of levodopa preparations. Modeling module 802 is used to model virtual patients based on patient data to obtain a virtual patient model. The virtual patient model is used to simulate the pharmacokinetic behavior of levodopa in the patient and generate motor symptom scores. The prediction module 803 is used to input preset medication behavior, motor symptom scores and patient data into the state transition model to obtain the prediction results of the symptom change after the target time output by the state transition model. The strategy module 804 is used to perform multi-stage reinforcement learning based on patient data, medication behavior, and the predicted changes in symptoms after the target time through the state transition model, to obtain a personalized levodopa prescription strategy for the reinforcement learning model. The personalized levodopa prescription strategy is used to output a levodopa prescription recommendation based on the input patient data. The levodopa prescription recommendation includes the dosage of the levodopa preparation and the decision interval.
[0104] Specifically, the device for generating personalized levodopa prescriptions for Parkinson's disease provided by the present invention can realize all the method steps implemented in the above-described method embodiment for generating personalized levodopa prescriptions for Parkinson's disease, and can achieve the same technical effect. Here, the parts that are the same as those in the method embodiment and the beneficial effects will not be described in detail.
[0105] Figure 9 This is a schematic diagram of the physical structure of the electronic device provided by the present invention, such as... Figure 9As shown, the electronic device may include: a processor 910, a communications interface 920, a memory 930, and a communications bus 940, wherein the processor 910, the communications interface 920, and the memory 930 communicate with each other through the communications bus 940. The processor 910 can call logic instructions in the memory 930 to execute a method for generating personalized levodopa prescription recommendations for Parkinson's disease. This method includes: acquiring patient data, including static characteristics and historical medication data; static characteristics including age, gender, and disease duration; and historical medication data including levodopa dosage records; modeling a virtual patient based on the patient data to obtain a virtual patient model, which simulates the pharmacokinetic behavior of levodopa in the patient and generates a motor symptom score; inputting preset medication behavior, motor symptom scores, and patient data into a state transition model to obtain a prediction result of symptom change after a target time, output by the state transition model; and performing multi-stage reinforcement learning through the state transition model based on patient data, medication behavior, and the prediction result of symptom change after the target time to obtain a personalized levodopa prescription strategy from the reinforcement learning model. This personalized levodopa prescription strategy is used to output a levodopa prescription recommendation based on the input patient data, and the levodopa prescription recommendation includes the dosage and decision interval of the levodopa preparation.
[0106] Furthermore, the logical instructions in the aforementioned memory 930 can be implemented as software functional units and, when sold or used as independent products, can be stored in a computer-readable storage medium. Based on this understanding, the technical solution of the present invention, in essence, or the part that contributes to the prior art, or a part of the technical solution, can be embodied in the form of a software product. This computer software product is stored in a storage medium and includes several instructions to cause a computer device (which may be a personal computer, server, or network device, etc.) to execute all or part of the steps of the methods of the various embodiments of the present invention. The aforementioned storage medium includes various media capable of storing program code, such as USB flash drives, portable hard drives, read-only memory (ROM), random access memory (RAM), magnetic disks, or optical disks.
[0107] On the other hand, the present invention also provides a computer program product, which includes a computer program that can be stored on a non-transitory computer-readable storage medium. When the computer program is executed by a processor, the computer is able to execute the method for generating personalized levodopa prescription recommendations for Parkinson's disease provided by the above methods. This method includes: acquiring patient data, wherein the patient data includes the patient's static characteristics and historical medication data, the static characteristics including age, gender, and disease duration, and the historical medication data including levodopa preparation medication records; and performing virtual patient modeling based on the patient data to obtain a virtual patient model, wherein the virtual patient model is used to simulate... The pharmacokinetic behavior of levodopa in patients is analyzed and motor symptom scores are generated. Pre-defined medication behavior, motor symptom scores, and patient data are input into a state transition model to obtain the predicted symptom changes after a target time. Multi-stage reinforcement learning is performed on the state transition model based on patient data, medication behavior, and the predicted symptom changes after the target time to obtain a personalized levodopa prescription strategy. The personalized levodopa prescription strategy is used to output levodopa prescription recommendations based on the input patient data. The levodopa prescription recommendations include the dosage and decision interval of the levodopa preparation.
[0108] In another aspect, the present invention also provides a non-transitory computer-readable storage medium storing a computer program thereon, which, when executed by a processor, implements a method for generating personalized levodopa prescription recommendations for Parkinson's disease provided by the methods described above. This method includes: acquiring patient data, wherein the patient data includes the patient's static characteristics and historical medication data, the static characteristics including age, gender, and disease duration, and the historical medication data including levodopa medication records; and performing virtual patient modeling based on the patient data to obtain a virtual patient model, wherein the virtual patient model is used to simulate the drug effect of levodopa in the patient's body. The system generates a motor symptom score based on the patient's kinetic behavior. Pre-defined medication behavior, motor symptom scores, and patient data are input into a state transition model to obtain a predicted symptom change after a target time. Multi-stage reinforcement learning is then performed on the state transition model based on patient data, medication behavior, and the predicted symptom change after the target time to obtain a personalized levodopa prescription strategy. This personalized levodopa prescription strategy is used to output a levodopa prescription recommendation based on the input patient data, including the dosage and decision interval of the levodopa preparation.
[0109] The device embodiments described above are merely illustrative. The units described as separate components may or may not be physically separate, and the components shown as units may or may not be physical units; that is, they may be located in one place or distributed across multiple network units. Some or all of the modules can be selected to achieve the purpose of this embodiment according to actual needs. Those skilled in the art can understand and implement this without any creative effort.
[0110] Through the above description of the embodiments, those skilled in the art can clearly understand that each embodiment can be implemented by means of software plus necessary general-purpose hardware platforms, and of course, it can also be implemented by hardware. Based on this understanding, the above technical solutions, in essence or the part that contributes to the prior art, can be embodied in the form of a software product. This computer software product can be stored in a computer-readable storage medium, such as ROM / RAM, magnetic disk, optical disk, etc., and includes several instructions to cause a computer device (which may be a personal computer, server, or network device, etc.) to execute the methods described in the various embodiments or some parts of the embodiments.
[0111] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention, and not to limit them; although the present invention has been described in detail with reference to the foregoing embodiments, those skilled in the art should understand that modifications can still be made to the technical solutions described in the foregoing embodiments, or equivalent substitutions can be made to some of the technical features; and these modifications or substitutions do not cause the essence of the corresponding technical solutions to deviate from the spirit and scope of the technical solutions of the embodiments of the present invention.
Claims
1. A method for generating personalized levodopa prescription recommendations for Parkinson's disease, characterized in that, include: Acquire patient data, wherein the patient data includes the patient's static characteristics and historical medication data, the static characteristics include age, gender and disease course, and the historical medication data includes medication records of levodopa preparations; Based on the patient data, a virtual patient model is generated to simulate the pharmacokinetic behavior of the levodopa preparation in the patient and generate a motor symptom score. The preset medication behavior, the motor symptom score, and the patient data are input into the state transition model to obtain the predicted symptom change after the target time output by the state transition model. The state transition model performs multi-stage reinforcement learning based on the patient data, medication behavior, and the predicted changes in symptoms after the target time, to obtain a personalized levodopa prescription strategy. The personalized levodopa prescription strategy is used to output a levodopa prescription recommendation based on the input patient data. The levodopa prescription recommendation includes the dosage and decision interval of the levodopa preparation.
2. The method for generating personalized levodopa prescription recommendations for Parkinson's disease according to claim 1, characterized in that, The process of creating a virtual patient model based on the patient data includes: Pharmacokinetic parameters of different types of levodopa formulations are obtained, wherein the pharmacokinetic parameters include time to peak concentration, peak concentration and half-life; Based on the historical medication data and the pharmacokinetic parameters, a pharmacokinetic model of levodopa was performed to obtain the total equivalent concentration of levodopa in the blood. Based on the total equivalent concentration of levodopa and the static characteristics, a blood-nerve central transport process of the drug is modeled to obtain the levodopa concentration in the central nervous system. Based on the levodopa concentration in the central nervous system, the patient's motor function response is modeled to obtain a motor symptom score, which includes a bradykinesia score and a dyskinesia score.
3. The method for generating personalized levodopa prescription recommendations for Parkinson's disease according to claim 1, characterized in that, The state transition model includes: a multilayer perceptron and a gated loop unit; The process of inputting the preset medication behavior, the motor symptom score, and the patient data into the state transition model to obtain the predicted symptom change after the target time output by the state transition model includes: The multilayer perceptron is used to perform a nonlinear mapping on the preset medication behavior, the motor symptom score, and the patient data to obtain static features; The time-dependent modeling of the preset medication behavior, the motor symptom score, and the patient data is performed by the gated loop unit to obtain time-series features; The temporal features and static features are concatenated and fused, and the features are interacted and mapped through the multilayer perceptron to obtain the predicted result of the symptom change after the target time.
4. The method for generating personalized levodopa prescription recommendations for Parkinson's disease according to claim 1, characterized in that, The action space of the reinforcement learning model includes: the dosage and decision interval of different types of levodopa formulations; The state space of the reinforcement learning model includes: the patient data, the medication behavior, and the predicted results of the symptom changes after the target time. The reward function of the reinforcement learning model includes: a penalty for irregular medication use, a reward for improved motor function, and a reward for maintaining motor function.
5. The method for generating personalized levodopa prescription recommendations for Parkinson's disease according to claim 4, characterized in that, The penalty for irregular medication use is expressed by the following formula: The reward for improved motor function is expressed by the following formula: The motor function maintenance reward is represented by the following formula: in, This represents the penalty value for irregular medication use. Indicates the penalty coefficient. Indicates time in hours. Indicates the current medication time. Indicates the date of the last medication administration; This indicates the reward value for improved motor function. The reward coefficient representing improvement in bradykinesia. The bradykinesia score indicates the current duration of medication administration. The bradykinesia score indicates the time since the last medication administration. The reward coefficient representing the improvement in dyskinesia. Dyskinesia score indicating the current medication duration. Dyskinesia score indicating the time since the last medication administration; This indicates the reward value for maintaining motor function. This represents the penalty coefficient for maintaining bradykinesia. This represents the penalty coefficient for maintaining dyskinesia.
6. The method for generating personalized levodopa prescription recommendations for Parkinson's disease according to claim 1, characterized in that, The process involves multi-stage reinforcement learning using the state transition model based on the patient data, medication behavior, and the predicted symptom changes after the target time, to obtain a reinforcement learning model, including: Based on the state transition model, an offline pre-training process is performed to generate an initial drug delivery strategy. The initial dosing strategy is deployed in the environment of the virtual patient model and fine-tuned using model-free reinforcement learning to obtain a personalized levodopa prescription strategy.
7. A device for generating personalized levodopa prescription recommendations for Parkinson's disease, characterized in that, include: The acquisition module is used to acquire patient data, wherein the patient data includes the patient's static characteristics and historical medication data. The static characteristics include age, gender and disease course, and the historical medication data includes medication records of levodopa preparations. The modeling module is used to perform virtual patient modeling based on the patient data to obtain a virtual patient model, wherein the virtual patient model is used to simulate the pharmacokinetic behavior of the levodopa preparation in the patient and generate a motor symptom score. The prediction module is used to input the preset medication behavior, the motor symptom score and the patient data into the state transition model to obtain the prediction result of the symptom change after the target time output by the state transition model. The strategy module is used to perform multi-stage reinforcement learning based on the patient data, the medication behavior, and the predicted changes in symptoms after the target time through the state transition model, to obtain a personalized levodopa prescription strategy for the reinforcement learning model. The personalized levodopa prescription strategy is used to output a levodopa prescription recommendation based on the input patient data. The levodopa prescription recommendation includes the dosage and decision interval of the levodopa preparation.
8. An electronic device comprising a memory, a processor, and a computer program stored in the memory and running on the processor, characterized in that, When the processor executes the computer program, it implements the method for generating personalized levodopa prescription recommendations for Parkinson's disease as described in any one of claims 1 to 6.
9. A non-transitory computer-readable storage medium having a computer program stored thereon, characterized in that, When the computer program is executed by a processor, it implements the method for generating personalized levodopa prescription recommendations for Parkinson's disease as described in any one of claims 1 to 6.
10. A computer program product, comprising a computer program, characterized in that, When the computer program is executed by a processor, it implements the method for generating personalized levodopa prescription recommendations for Parkinson's disease as described in any one of claims 1 to 6.
Citation Information
Patent Citations
Parkinson's disease drug recommendation model based on multi-source data fusion
CN111640481A
Traditional Chinese medicine dynamic diagnosis and treatment scheme optimization method and system based on deep reinforcement learning
CN114783571A