Reinforcement learning-based radiopharmaceutical dose dynamic optimization method and system
By employing a reinforcement learning-based method for dynamic optimization of radiopharmaceutical dosage, and utilizing multimodal imaging data and a dual reward function, personalized and precise radiopharmaceutical administration in PET/MRI examinations was achieved. This addresses the shortcomings of traditional dosing protocols, improves diagnostic accuracy, and reduces radiation risks.
Patent Information
- Application Number
- CN202511058008.8
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-07-30
- Publication Date
- 2025-10-31
Smart Images

Figure CN120878068A_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the fields of medical imaging and radiopharmaceuticals, specifically to a method and system for dynamic optimization of radiopharmaceutical dosage based on reinforcement learning, which is used to achieve personalized and precise administration of radiopharmaceuticals during PET / MRI examinations. Background Technology
[0002] Radiopharmaceuticals have important applications in medical imaging examinations such as PET / MRI. However, traditional dosing regimens, which are mainly based on simple calculations of patient weight or body surface area, are difficult to adapt to the differences in brain structure, blood-brain barrier function, and drug metabolism among different patients. This fixed-dose dosing regimen usually leads to two prominent problems: on the one hand, the drug concentration in the target brain region is insufficient in some patients, affecting diagnostic accuracy; on the other hand, the excessive accumulation of radiopharmaceuticals in non-target areas increases the patient's radiation risk.
[0003] Existing technologies have introduced some improvements, such as adjusting dosage based on body mass index and correcting dosage according to organ function parameters. However, these methods remain static and cannot be dynamically adjusted based on the patient's real-time responses during the examination. Furthermore, they fail to fully utilize the rich information contained in multimodal imaging, making it difficult to achieve truly personalized and precise drug delivery.
[0004] With the development of artificial intelligence technology, machine learning methods are increasingly being applied in the medical field. Deep learning, reinforcement learning, and other technologies have provided new approaches for optimizing radiopharmaceutical dosage. However, research on applying reinforcement learning to radiopharmaceutical dosage optimization is still in its early stages, and systematic reports on real-time dynamic optimization systems combining multimodal imaging and pharmacokinetic models are yet to be seen. Summary of the Invention
[0005] The main objective of this invention is to provide a method and system for dynamic optimization of radiopharmaceutical dosage based on reinforcement learning, which aims to solve the technical problem that traditional fixed-dose dosing regimens lack personalized and real-time adjustment capabilities.
[0006] This invention proposes a method for dynamic optimization of radiopharmaceutical dosage based on reinforcement learning, including:
[0007] Acquire multimodal magnetic resonance imaging data of the patient's brain;
[0008] Based on the multimodal magnetic resonance imaging data, the patient's brain is finely divided into regions to obtain brain region data;
[0009] Based on the brain region data, a first reward function is constructed;
[0010] Based on the brain region data, a pharmacokinetic model was established, and a second reward function was constructed.
[0011] Based on the first reward function and the second reward function, a reinforcement learning model is trained to obtain a radiopharmaceutical dosage optimization strategy.
[0012] Real-time acquisition of dynamic data on brain radiation in patients; and
[0013] The radiopharmaceutical dosage is dynamically adjusted based on the radiopharmaceutical dosage optimization strategy and the brain radioactivity dynamic data.
[0014] Preferably, the acquisition of multimodal magnetic resonance imaging data of the patient's brain includes:
[0015] Collect MRI, DWI, VBM, and AVN multimodal imaging data of the patient's brain;
[0016] Spatial registration processing is performed on the multimodal image data to align different modal images to the same spatial coordinate system; and
[0017] A fused feature map is generated based on the registered multimodal image data.
[0018] As preferred, including:
[0019] The patient's brain is divided into the basic regions of gray matter, white matter, deep nuclei, ventricles, and cerebellum;
[0020] The basic regions are further subdivided into functional areas of the ventricles, deep nuclei, cerebellum, brainstem, cortex, striatum, hippocampus, thalamus, basal ganglia, and midbrain shell; and
[0021] The accuracy of partitioning was verified by anatomical feature point marking and boundary detection.
[0022] Preferably, the construction of the first reward function includes:
[0023] The brain region metabolic baseline, cerebrospinal fluid flux, Aβ deposition score, and hippocampal uptake ratio in PET images were extracted from the brain region partitioning data as parameters.
[0024] Design reward items corresponding to the distribution of drugs in the target area, cerebrospinal fluid clearance efficiency, plaque binding effect, and hippocampal uptake level;
[0025] Design penalties for drug accumulation in non-target areas, excessive cerebrospinal fluid flux, and excessive striatal uptake; and
[0026] A first reward function is generated based on the reward and penalty terms.
[0027] Preferably, the establishment of the pharmacokinetic model and the construction of the second reward function include:
[0028] Pharmacokinetic parameters of the drug in plasma, such as half-life, volume of distribution in brain tissue, elimination rate constant, and brain tissue uptake rate, were extracted.
[0029] A pharmacokinetic model was established based on the aforementioned pharmacokinetic parameters;
[0030] Design an index that maximizes the area under the time curve of drug concentration in the target region;
[0031] Design drug clearance and retention balance indicators; and
[0032] A second reward function is generated based on the maximization metric and the balance metric.
[0033] Preferably, the training reinforcement learning model includes:
[0034] First, train the reinforcement learning algorithm for K rounds using the first reward function;
[0035] The reinforcement learning algorithm is then trained for M rounds using the second reward function, where K≠M;
[0036] When the directions guided by the first reward function and the second reward function conflict, a conflict coordination mechanism is activated; and
[0037] The results of the two-stage training are integrated using a weighted average method to form a comprehensive optimization strategy.
[0038] Preferably, the reinforcement learning algorithm is a SARSA algorithm based on a value function, comprising:
[0039] The state space is defined by the combination of brain region metabolic baseline, cerebrospinal fluid flux, Aβ deposition score, hippocampal uptake ratio in PET images, and pharmacokinetic parameters.
[0040] The action space is defined by the combination of parameters including the dosage, injection rate, and injection time of the radiopharmaceutical; and
[0041] A state transition model is constructed based on a pharmacokinetic model to predict the new state that the system will transition to after performing a specific action in the current state.
[0042] Preferably, the real-time acquisition of dynamic radioactive data of the patient's brain includes:
[0043] Real-time acquisition of partial arterial input function data via arterial sampling or imaging methods;
[0044] Real-time PET scans were used to monitor changes in radiopharmaceutical concentrations in different areas of the brain; and
[0045] The collected raw data is denoised, corrected, and standardized to improve data quality.
[0046] Preferably, the dynamic adjustment of the radiopharmaceutical dosage includes:
[0047] The initial dose is determined based on the patient's basic information such as weight, age, and renal function;
[0048] Set thresholds for key parameters and monitor in real time whether they exceed the safe range;
[0049] Construct a multi-layer decision tree to determine whether the dosage needs to be adjusted based on real-time monitoring data and preset thresholds;
[0050] When drug uptake in the white matter and deep nuclei of the brain is found to be insufficient, the dosage is increased and / or the dynamic scan is extended to 90 minutes, while the radiation dose is reduced; and
[0051] The adjusted results and observation data are then fed back to the reinforcement learning model as feedback, forming a closed loop.
[0052] A reinforcement learning-based system for dynamic optimization of radiopharmaceutical dosage includes:
[0053] The multimodal imaging acquisition module is used to acquire multimodal magnetic resonance imaging data of the patient's brain;
[0054] The brain region partitioning module is used to perform fine partitioning of the patient's brain based on the multimodal magnetic resonance imaging data to obtain brain region partitioning data;
[0055] The first reward function construction module is used to construct a first reward function based on the brain region partitioning data;
[0056] The pharmacokinetic model module is used to establish a pharmacokinetic model and construct a second reward function based on the brain region data.
[0057] The reinforcement learning model training module is used to train a reinforcement learning model based on the first reward function and the second reward function to obtain a radiopharmaceutical dosage optimization strategy.
[0058] A dynamic data acquisition module is used to collect real-time dynamic data on radiation in the patient's brain; and
[0059] The dose dynamic adjustment module is used to dynamically adjust the radiopharmaceutical dose based on the radiopharmaceutical dose optimization strategy and the brain radioactivity dynamic data.
[0060] This invention integrates innovative technologies such as multimodal magnetic resonance imaging, fine brain segmentation, dual reward function mechanism, blood-brain barrier permeability modeling, reinforcement learning optimization, and real-time feedback adjustment to construct a complete closed-loop optimization system, realizing intelligent dynamic optimization of radiopharmaceutical dosage.
[0061] The present invention has the following beneficial effects:
[0062] 1. Personalized Precision Drug Delivery: Based on the patient's multimodal imaging data and real-time feedback, the system can accurately capture each patient's unique brain features and drug metabolism characteristics, achieving truly personalized precision drug delivery. Experiments show that compared with traditional methods, this system can increase drug concentration in the target area by 20%–30%, while reducing drug accumulation in non-target areas by 15%–25%.
[0063] 2. Significantly reduced radiation risk: By accurately modeling the drug sensitivity and metabolic characteristics of different brain regions, the system has achieved an appropriate drug delivery strategy, reducing the average radiation dose received by patients by approximately 18% to 22% while ensuring diagnostic effectiveness.
[0064] 3. Real-time dynamic adjustment capability: The closed-loop feedback system can dynamically adjust the dosage based on real-time acquired data, significantly improving the system's adaptability and robustness, making it particularly suitable for patients with abnormal blood-brain barrier function or impaired drug metabolism. Clinical data shows that approximately 15%–20% of patients require dosage adjustments during scanning; this system can automatically identify these situations and make appropriate adjustments.
[0065] 4. Multi-objective balance optimization: Through dual reward functions and differentiated training strategies, the system achieves balanced optimization of multiple objectives such as imaging quality, diagnostic accuracy, radiation dose and patient safety, and can dynamically adjust the weight of each objective according to the needs of different clinical scenarios.
[0066] 5. Adaptive Learning Capability: The system possesses continuous learning and optimization capabilities, and its performance will continuously improve as data accumulates. Reinforcement learning algorithms can learn from each clinical application, continuously optimize decision-making strategies, and improve the long-term effectiveness of the system. Attached Figure Description
[0067] Figure 1 This is a diagram illustrating the overall architecture of the reinforcement learning-based dynamic optimization system for radiopharmaceutical dosage in this invention.
[0068] Figure 2 This is a flowchart of the multimodal image acquisition and fine brain region partitioning process in this invention;
[0069] Figure 3 This is a schematic diagram of the dual reward function construction and dynamic balancing mechanism in this invention;
[0070] Figure 4 This is a flowchart of the reinforcement learning model training and optimization process in this invention;
[0071] Figure 5 This is a schematic diagram of the real-time closed-loop feedback and dynamic dose adjustment mechanism in this invention. Detailed Implementation
[0072] Please refer to the attached document. Figure 1-5 The present invention will now be described in further detail with reference to the accompanying drawings and specific embodiments. It should be understood that the specific embodiments described herein are for illustrative purposes only and are not intended to limit the scope of the invention.
[0073] Reference Figure 1 The radiopharmaceutical dosage dynamic optimization system based on reinforcement learning provided by this invention includes: a multimodal image acquisition module 1, a brain region segmentation module 2, a first reward function construction module 3, a pharmacokinetics model module 4, a reinforcement learning model training module 5, a dynamic data acquisition module 6, and a dosage dynamic adjustment module 7. The data flow and information interaction between these modules constitute a complete closed-loop system, realizing intelligent management of the entire process from image acquisition to dosage dynamic adjustment.
[0074] The following describes in detail the reinforcement learning-based method for dynamic optimization of radiopharmaceutical dosage of the present invention, which corresponds to the operation flow of the aforementioned system.
[0075] In an embodiment of the present invention, multimodal magnetic resonance imaging data of the patient's brain is first acquired through the multimodal imaging acquisition module 1. Preferably, this step includes acquiring imaging data of four modalities of the patient's brain: MRI, DWI, VBM, and AVN.
[0076] Specifically, MRI (Magnetic Resonance Imaging) employed T1-weighted and T2-weighted sequences with a resolution of 1 mm × 1 mm × 1 mm and a slice thickness of 1 mm; DWI (Diffusion-weighted Imaging) used a single-excitation planar echo sequence with a b-value of 0 and 1000 s / mm², achieving a resolution of 2 mm × 2 mm × 2 mm; VBM (Voxel-based Morphometry) used a T1-weighted three-dimensional fast gradient echo sequence with a resolution of 1 mm × 1 mm × 1 mm; and AVN (Arterial Spin Labeling) used a pseudo-continuous arterial spin labeling sequence with a labeling delay of 1800 ms and a resolution of 3 mm × 3 mm × 5 mm. These parameter settings are optimal configurations validated through extensive clinical practice, ensuring both image quality and optimal scan time.
[0077] After acquiring the raw image data, spatial registration of images from different modalities is required to align them to the same spatial coordinate system. This invention employs a combination of rigid and non-rigid transformations to achieve spatial registration. First, a rigid registration algorithm based on mutual information is used for coarse registration to correct for rotation and translation differences. Then, a non-rigid registration algorithm based on B-spline transformation is used for fine registration to address local differences caused by physiological motion and nonlinear deformation during imaging. Registration accuracy is comprehensively evaluated using mutual information values and structural similarity indices. The mutual information value threshold is set above 0.7, and the structural similarity index threshold is set above 0.85. These thresholds are derived from extensive clinical data analysis, ensuring the accuracy and reliability of the registration results.
[0078] After registration, a fusion feature map is generated based on the registered multimodal image data. This invention employs a weighted fusion strategy, assigning weights according to the information contribution of each modality to a specific brain region. For example, for gray matter, T1-weighted MRI has a weight of 0.4, VBM has a weight of 0.3, DWI has a weight of 0.2, and AVN has a weight of 0.1; for white matter, T1-weighted MRI has a weight of 0.3, VBM has a weight of 0.2, DWI has a weight of 0.4, and AVN has a weight of 0.1. This weighting scheme considers the differences in sensitivity of different imaging modalities to different tissue types, maximizing the extraction and utilization of complementary information from each modality's images.
[0079] After acquiring multimodal magnetic resonance imaging data, the brain regionalization module 2 performs fine-grained regionalization of the patient's brain based on this data, obtaining brain regionalization data. This step forms the basis for subsequent construction of personalized models.
[0080] In a preferred embodiment of the invention, the patient's brain is first divided into five basic regions: gray matter, white matter, deep nuclei, ventricles, and cerebellum. This basic division is mainly based on the intensity distribution characteristics of T1-weighted MRI images, combined with threshold segmentation and morphological processing techniques. Specifically, for T1-weighted MRI images, off-field correction is first performed to remove uneven gray-level distribution, and then a threshold segmentation method based on a Gaussian mixture model is used to segment the image into three categories: gray matter, white matter, and cerebrospinal fluid. For deep nuclei and the cerebellum, identification and segmentation are performed by combining prior anatomical knowledge and morphological features.
[0081] Building upon the basic regional divisions, these regions are further refined into ten functional areas: ventricles, deep nuclei, cerebellum, brainstem, cortex, striatum, hippocampus, thalamus, basal ganglia, and mesencephalon. This refined regional division process is primarily based on region growing and level set segmentation algorithms, combined with feature information extracted from multimodal images. For example, hippocampal segmentation utilizes a combination of T1-weighted MRI and VBM features, taking advantage of the hippocampus's unique gray matter intensity and volume characteristics; striatum segmentation combines DWI anisotropy fractional maps, leveraging the striatal fiber tract orientation features for precise segmentation.
[0082] To verify the accuracy of the partitioning results, this invention employs anatomical feature point marking and boundary detection algorithms. Specifically, 10-15 anatomical feature points are marked on each partition, and then compared with the positions of corresponding feature points on a standard brain atlas to calculate the spatial positional deviation. When the average deviation is less than 2 mm and the maximum deviation is less than 3.5 mm, the partitioning results are considered to meet the accuracy requirements. These deviation thresholds are set based on international brain atlas partitioning accuracy assessment standards, ensuring the clinical usability of the partitioning results.
[0083] Based on brain region data, the first reward function construction module 3 constructs the first reward function as one of the optimization objectives of the reinforcement learning model.
[0084] In embodiments of the present invention, brain region metabolic baseline, cerebrospinal fluid flux, Aβ deposition score, and hippocampal uptake ratio in PET images are extracted from brain region regional data as parameters. Brain region metabolic baseline Blood perfusion values for each brain region were calculated using AVN images, with units of ml / 100g / min; cerebrospinal fluid flux ( ) Flow rate was measured by MRI phase-contrast imaging, in m1 / min; Aβ deposition score Calculated using the standard uptake ratio (SUVR) of PET images, dimensionless; the uptake ratio of the hippocampus brain region in PET images. Defined as the ratio of the standard uptake value of the hippocampus to the standard uptake value of a reference region (usually the cerebellum), it is dimensionless.
[0085] Based on these parameters, this invention designs four positive reward indicators, corresponding to the effective distribution of the drug in the target area, cerebrospinal fluid clearance efficiency, plaque binding effect, and hippocampal uptake level, respectively. Specifically, the reward items... The design is as follows:
[0086] ,
[0087] in: Let be the reward value for the i-th brain region, which is dimensionless. This represents the baseline metabolic value for the i-th brain region, expressed in ml / 100g / min. and These represent the minimum and maximum values of the metabolic baseline for all brain regions, respectively, in units of... same; The cerebrospinal fluid flux in the i-th brain region is expressed in ml / min. and These are the minimum and maximum values of cerebrospinal fluid flux in all brain regions, respectively, in units of... same; The Aβ deposition score for the i-th brain region is dimensionless. and These are the minimum and maximum values of Aβ deposition scores for all brain regions, respectively, and are dimensionless. is the hippocampal uptake ratio of the i-th brain region, which is dimensionless; and These are the minimum and maximum values of the hippocampal uptake ratio for all brain regions, respectively, and are dimensionless. , , and These are weighting coefficients, all dimensionless, set to 0.3, 0.2, 0.3, and 0.2 respectively. These weighting values are determined based on extensive clinical data analysis and expert experience, and are able to balance the contribution of each parameter to the overall reward.
[0088] Simultaneously, this invention also designs three negative penalty indicators targeting drug accumulation in non-target areas, excessive cerebrospinal fluid flux, and striatal over-uptake. Penalty Items ( The design is as follows:
[0089] ,
[0090] in: Let be the penalty value for the i-th brain region, which is dimensionless; The target metabolic baseline value is typically set as the average value of the same region in the healthy control group, with units equal to... same; The cerebrospinal fluid flux threshold is set at 3.5 ml / min; exceeding this value may lead to excessively rapid drug clearance. The uptake ratio of the striatal brain region is dimensionless. The striatal uptake threshold is set to 1.8, which is dimensionless. Exceeding this value may lead to an increase in nonspecific binding. , and These are penalty weighting coefficients, all dimensionless, and set to 0.4, 0.3, and 0.3 respectively. The function takes the larger value between the expression within the parentheses and 0, ensuring that a penalty is only applied if the parameter exceeds a threshold.
[0091] Ultimately, the first reward function ( The formula for calculating ) is:
[0092] ,
[0093] in: The first reward function value is dimensionless. Indicates the number of brain regions divided in the MRI image; γ and γ are the overall weight coefficients for the reward and penalty items, respectively, both dimensionless, and set to 0.6 and 0.4. and These represent the importance weights of each brain region, all of which are dimensionless and adjusted according to different diagnostic goals; This represents a summation operation over all n brain regions. For example, in the diagnosis of Alzheimer's disease, the hippocampus and frontal cortex... and The weight will be increased accordingly, and can be set to 1.5 times that of other areas.
[0094] This reward function design fully considers the distribution characteristics of radiopharmaceuticals in brain tissue and the needs of clinical diagnosis. It can effectively guide reinforcement learning models to optimize dosage strategies, achieve optimal drug distribution in the target region, and minimize accumulation in non-target regions.
[0095] While constructing the first reward function, the pharmacokinetic model module 4 establishes a pharmacokinetic model based on brain region data and constructs the second reward function.
[0096] In embodiments of the present invention, pharmacokinetic parameters such as the drug's half-life in plasma, volume of distribution in brain tissue, elimination rate constant, and brain tissue uptake rate are first extracted. These parameters are typically obtained through dynamic PET scans combined with plasma sample analysis. Specifically, the half-life in plasma... Typically, it is a combination of the physical half-life and biological half-life of a radiopharmaceutical, expressed in minutes; brain tissue distribution volume ( Elimination rate constant (ERC) is defined as the ratio of drug concentration in tissues to drug concentration in plasma, and is dimensionless. This indicates the rate at which a drug is eliminated from the body, expressed in units of 1000 mg / L. Brain tissue uptake rate ( This indicates the rate at which a drug enters brain tissue from the bloodstream, expressed in ml / g / min.
[0097] Based on these parameters, this invention establishes a two-compartment pharmacokinetic model to describe the dynamic distribution of radiopharmaceuticals between plasma and brain tissue. The mathematical expression of this model is:
[0098] ,
[0099] ,
[0100] in: This represents the change in drug concentration in brain tissue over time, expressed in Bq / ml. This represents the derivative of drug concentration in brain tissue with respect to time, reflecting the rate of concentration change, and is expressed in Bq / ml / min. This represents the change in drug concentration in plasma over time, expressed in Bq / ml. This is the rate constant for the entry of a drug from plasma into brain tissue, expressed in ml / g / min. This is the rate constant for the return of drugs from brain tissue to plasma, expressed in units of... This is the dosage, expressed in Bq. This is a dimensionless parameter representing the magnitude of the multi-exponential decay of plasma drug concentration. The decay constant of plasma drug concentration with multi-exponential decay, in minutes. Represents an exponentially decaying function; Typically, 2-3 is used to represent the number of attenuation terms; This indicates a summation operation on m attenuation terms.
[0101] Based on the established pharmacokinetic model, this invention designs an index that maximizes the area under the drug concentration-time curve in the target region. This index reflects the effective exposure degree of the drug in the target region, and its calculation formula is as follows:
[0102] ,
[0103] in: The area under the time curve of drug concentration in the target region is expressed in Bq·min / ml. The change in drug concentration in the target area over time is expressed in Bq / ml. The observation time is usually set to 90–120 minutes, in min. This represents the definite integral operation from time 0 to T, calculating the area under the curve during the entire observation time.
[0104] Furthermore, this invention also incorporates a drug clearance-retention balance index to prevent excessive drug accumulation in non-target areas. The formula for calculating this index is:
[0105] ,
[0106] in: A dimensionless indicator of the drug clearance-retention balance. The area under the time curve of drug concentration in the target region is expressed in Bq·min / ml. The area under the time curve of drug concentration in the non-target region is expressed in Bq·min / ml.
[0107] Finally, the second reward function ( The formula for calculating ) is:
[0108] ,
[0109] in: The value of the second reward function is dimensionless. The area under the time curve of drug concentration in the target region is expressed in Bq·min / ml. and These represent the minimum and maximum acceptable areas under the drug concentration-time curve in the target region, respectively, in units of... same; A dimensionless indicator of the drug clearance-retention balance. and These are the minimum and maximum acceptable values for the drug clearance and retention balance index, respectively, and are dimensionless. This is the dosage, expressed in Bq. and These are the minimum and maximum safe limits for the administered dose, respectively, in units of... same; , and These are weighting coefficients, all dimensionless, set to 0.5, 0.3, and 0.2 respectively. These weighting values are determined based on radiation protection principles and clinical experience, aiming to balance diagnostic effectiveness and radiation risk.
[0110] This reward function design based on pharmacokinetic models fully considers the dynamic metabolic process of radiopharmaceuticals in vivo, and can effectively guide reinforcement learning models to optimize drug delivery strategies, achieve optimal drug exposure in the target area, and minimize the overall radiation dose.
[0111] Based on the constructed first and second reward functions, reinforcement learning model training module 5 trains the reinforcement learning model to obtain the radiopharmaceutical dosage optimization strategy.
[0112] In a preferred embodiment of the present invention, the reinforcement learning model is trained using a two-stage strategy: first, the reinforcement learning algorithm is trained for K rounds using a first reward function, and then the reinforcement learning algorithm is trained for M rounds using a second reward function, where K ≠ M. Preferably, K is set to 500 rounds and M is set to 300 rounds. This differentiated training round setting takes into account the differences in complexity and convergence characteristics of the two reward functions.
[0113] When the directions guided by the first reward function and the second reward function conflict, a conflict coordination mechanism is activated. Specifically, the system analyzes the nature and severity of the conflict and adjusts the decision based on preset priority rules and risk assessment results. For example, when the first reward function guides an increase in dose to improve uptake in the target area, while the second reward function guides a decrease in dose due to radiation risk considerations, the system will prioritize safety factors and select a more conservative dose strategy while ensuring minimum diagnostic requirements are met.
[0114] After training, the system integrates the two-stage training results using a weighted average to form a comprehensive optimization strategy. The integration formula is:
[0115] ,
[0116] in: The final integrated strategy is a multidimensional decision function; and The strategies are trained based on the first reward function and the second reward function, respectively, and both are multi-dimensional decision functions. and These are weighting coefficients, all dimensionless, satisfying... The initial value is set to The results can be adjusted based on clinical validation.
[0117] In one specific embodiment of the present invention, the reinforcement learning algorithm employs the SARSA (State-Action-Reward-State-Action) algorithm based on a value function. The SARSA algorithm is an online policy reinforcement learning algorithm suitable for scenarios requiring real-time decision-making.
[0118] Specifically, the state space is defined by a combination of factors including brain region metabolic baseline, cerebrospinal fluid flux, Aβ deposition score, hippocampal uptake ratio in PET images, and pharmacokinetic parameters. State vector The representation is:
[0119] ,
[0120] in, The state vector is a multidimensional vector. This is a metabolic baseline subvector, containing the metabolic baseline values for each brain region; This is the cerebrospinal fluid flux quantum vector, which contains the cerebrospinal fluid flux values for each brain region; This is an Aβ deposition score subvector, containing the Aβ deposition scores for each brain region; This is a subvector of hippocampal uptake ratios, containing the hippocampal uptake ratios for each brain region. This is a subvector of pharmacokinetic parameters, containing the dynamic metabolic parameters of the drug in vivo. All parameters are normalized to ensure that the numerical range of each dimension is consistent, typically standardized to the [0,1] interval.
[0121] The action space is defined by the combination of parameters such as the dose, injection rate, and injection time of the radiopharmaceutical. Action vector The representation is:
[0122] ,
[0123] in: The action vector is a three-dimensional vector. The dose adjustment ratio is discretized into seven levels: [-0.3, -0.2, -0.1, 0, 0.1, 0.2, 0.3], representing the percentage adjustment relative to the standard dose. The injection rate adjustment ratio is discretized into five levels: [-0.2, -0.1, 0, 0.1, 0.2]. The injection time point is discretized into five levels: [-10, -5, 0, 5, 10], with the unit being minutes, representing the adjustment relative to the standard injection time point.
[0124] A state transition model is constructed based on a pharmacokinetic model to predict the new state the system will transition to after performing a specific action in the current state. State transition function. Indicates the state Next action After transitioning to state The probability of state transitions is unknown. In practical implementation, due to the complex physiological processes and individual differences involved in the evolution of system states, there is a certain degree of uncertainty in state transitions. Therefore, it is more reasonable to use a probabilistic model rather than a deterministic model.
[0125] The core of the SARSA algorithm is the update of the Q function, and the update formula is:
[0126] ,
[0127] in: The state-action value function represents the state... Next action The expected cumulative reward is dimensionless; This indicates an assignment operation, which assigns the result of the calculation on the right to the variable on the left. The reward is immediate and is calculated using the aforementioned reward function; it is dimensionless. The learning rate is dimensionless and initially set to 0.1, gradually decreasing to 0.01 as training progresses. The discount factor is dimensionless and set to 0.9 to balance immediate rewards and long-term benefits. and These represent the state and action at the next moment; This represents the timing difference error, used to guide the direction and magnitude of Q-value updates.
[0128] Action selection adopts - Greedy strategy, that is, using The probability of choosing the action with the highest current Q value is given. The probability of randomly selecting an action. The exploration rate is dimensionless and initially set to 0.3, gradually decreasing to 0.05 as training progresses. This strategy balances exploration and exploitation, preventing the algorithm from getting trapped in local optima.
[0129] In practical applications, the dynamic data acquisition module 6 collects real-time dynamic data on the patient's brain radiation, and the dose dynamic adjustment module 7 dynamically adjusts the radiopharmaceutical dose based on the trained radiopharmaceutical dose optimization strategy and the collected dynamic data.
[0130] In a preferred embodiment of the present invention, real-time data acquisition includes acquiring partial arterial input function data through arterial sampling or imaging methods, and monitoring changes in radiopharmaceutical concentrations in various brain regions through real-time PET scanning.
[0131] Arterial input function data can be obtained directly through radial artery cannulation sampling, with a sampling frequency of once every 30 seconds for the first 10 minutes, and once every 5 minutes thereafter, for a total of approximately 20 sample points. Alternatively, imaging input functions can be obtained through region of interest (ROI) analysis of large arteries (such as the internal carotid artery or aortic arch), with a sampling frequency of once per minute, synchronized with PET scans.
[0132] Changes in radiopharmaceutical concentrations in different brain regions were obtained using dynamic PET scans. A typical acquisition protocol was as follows: one frame every 30 seconds for the first 5 minutes, one frame every minute for 5–10 minutes, one frame every 2 minutes for 10–30 minutes, and one frame every 5 minutes for 30–90 minutes, for a total of approximately 30 time frames.
[0133] The acquired raw data needs to undergo denoising, attenuation correction, scattering correction, and standardization to improve data quality. Denoising employs a combination of Gaussian filtering and wavelet transform, with a filter kernel width of 3mm, effectively suppressing noise while preserving image details. Attenuation correction is based on attenuation maps obtained from CT or MRDixon sequences, considering differences in electron density across different tissues. Scattering correction uses a single scattering simulation algorithm (SSS) to simulate the scattering process of photons in tissues. Standardization converts the radioactivity concentrations of different time frames and brain regions into standard uptake values (SUV) or distribution volume ratios (DVR) for easier comparison and analysis.
[0134] Based on the processed dynamic data, the system first determines the initial dose according to basic information such as the patient's weight, age, and renal function. Typically, the initial dose calculation formula is:
[0135] ,
[0136] in: This is the initial dose, in MBq; The patient's weight is expressed in kg. This refers to the standard unit dose, expressed in MBq / kg, typically 3-5 MBq / kg. The age-adjusted factor is dimensionless; for elderly patients (>65 years), the value is 0.9; for children (<18 years), the value is 0.3-0.8 depending on age. This is a renal function adjustment factor, dimensionless. A value of 1.0 is used for normal renal function, 0.8 for mild impairment, 0.6 for moderate impairment, and alternative testing methods should be considered for severe impairment.
[0137] During drug administration, the system sets thresholds for key parameters and monitors them in real time to ensure they do not exceed safe limits. Key monitoring parameters include: target area uptake (SUV or DVR), with thresholds typically set at ±20% of the expected value; non-target area uptake, especially uptake from radiation-sensitive organs (such as the kidneys and bladder), with thresholds set based on the organ's radiation sensitivity and safety limits; and physiological parameters such as blood pressure and heart rate, with thresholds set based on the patient's baseline level and safety range.
[0138] Based on real-time monitoring data and preset thresholds, the system constructs a multi-layered decision tree to determine whether dosage adjustment is needed, as well as the direction and magnitude of the adjustment. The basic structure of the decision tree includes:
[0139] First layer: Determine whether the intake of the target area meets the standard (within ±20%).
[0140] Meeting the criteria: Continue with the current plan, without adjustments;
[0141] If the standard is not met: proceed to the second level of judgment;
[0142] The second level: distinguishing between insufficient intake and excessive intake;
[0143] Insufficient intake (<80% of expected value): Proceed to Level 3-A;
[0144] Excessive intake (>120% of expected value): Enter Level 3-B;
[0145] Level 3-A (Insufficient Intake):
[0146] Mild deficiency (60%–80% of expected value): Increase dose by 10%;
[0147] Moderate under-dosage (40%–60% of expected value): Increase dose by 20%;
[0148] Severe underdose (<40% of expected value): Increase dose by 30% and extend scan time.
[0149] Level 3-B (Excessive Intake):
[0150] Mild overdose (120%–140% of expected value): Reduce dose by 10%;
[0151] Moderate overdose (140%–160% of expected value): Reduce dose by 20%;
[0152] Severe overdose (>160% of expected value): Reduce dose by 30% and increase fluid intake to promote excretion;
[0153] Specifically, when the system detects insufficient drug uptake in the brain's white matter or deep nuclei, it automatically increases the drug dosage and / or extends the dynamic scan to 90 minutes, while simultaneously reducing the radiation dose. Specifically, when the white matter uptake is less than 50% of the expected value, the dose is increased by 20% and the scan time is extended to 90 minutes; when the deep nuclei uptake is less than 60% of the expected value, the dose is increased by 15% and the scan time is extended to 75 minutes. At the same time, by adjusting the injection rate and acquisition parameters, the overall radiation dose can be reduced while increasing uptake in the target area. For example, reducing the initial peak concentration but extending the drug residence time can reduce radiation while maintaining diagnostic image quality.
[0154] The adjusted results and observational data are then fed back to the reinforcement learning model as feedback, forming a closed loop. This closed-loop feedback mechanism enables the system to continuously learn and optimize, adapting to individual differences among patients and improving the accuracy and personalization of dosage optimization.
[0155] Reference Figure 1 The reinforcement learning-based dynamic optimization system for radiopharmaceutical dosage provided by this invention includes the following modules:
[0156] The multimodal image acquisition module 1 is used to acquire multimodal magnetic resonance imaging data of the patient's brain. This module includes an MRI scanning submodule, an image preprocessing submodule, and a data fusion submodule, which are responsible for raw image acquisition, preprocessing, and multimodal data fusion, respectively.
[0157] Brain partitioning module 2 is used to perform fine partitioning of the patient's brain based on multimodal magnetic resonance imaging data, obtaining brain partitioning data. This module includes a basic partitioning submodule, a fine partitioning submodule, and an accuracy verification submodule to ensure the accuracy and reliability of the partitioning results.
[0158] The first reward function construction module 3 is used to construct the first reward function based on brain region data. This module extracts parameters related to image features, designs reward and penalty terms, and generates a reward function that comprehensively evaluates the distribution quality.
[0159] Module 4, the pharmacokinetic model module, is used to establish a pharmacokinetic model based on brain region data and construct a second reward function. This module includes a parameter extraction submodule, a model building submodule, and a reward function design submodule, comprehensively evaluating the dynamic distribution process of the drug in vivo.
[0160] Reinforcement learning model training module 5 is used to train a reinforcement learning model based on the first and second reward functions to obtain a radiopharmaceutical dosage optimization strategy. This module includes an environment construction submodule, an algorithm implementation submodule, and a policy evaluation submodule, which realize the training and optimization of the policy.
[0161] The dynamic data acquisition module 6 is used to acquire real-time dynamic radioactive data of the patient's brain. This module includes an arterial input function acquisition submodule, a PET dynamic scanning submodule, and a data processing submodule, providing the data support required for real-time feedback.
[0162] The dose dynamic adjustment module 7 is used to dynamically adjust the radiopharmaceutical dose based on the radiopharmaceutical dose optimization strategy and dynamic brain radiation data. This module includes an initial dose determination submodule, a threshold monitoring submodule, a decision tree construction submodule, and a feedback processing submodule to achieve closed-loop optimization control.
[0163] The workflow of the reinforcement learning-based radiopharmaceutical dosage dynamic optimization system of the present invention is as follows:
[0164] After the system is started, the multimodal image acquisition module 1 first acquires the patient's MRI, DWI, VBM and AVN multimodal image data, and performs preprocessing and fusion.
[0165] The brain region partitioning module 2 receives the fused multimodal image data, performs basic partitioning and fine partitioning, and generates brain region partitioning data containing ten functional regions.
[0166] The first reward function construction module 3 and the pharmacokinetic model module 4 work in parallel: the former constructs the first reward function based on brain region data, while the latter establishes the pharmacokinetic model and constructs the second reward function.
[0167] The reinforcement learning model training module 5 receives two reward functions and uses a two-stage training strategy to train the SARSA algorithm, generating a radiopharmaceutical dosage optimization strategy.
[0168] In practical applications, the dynamic data acquisition module 6 collects real-time dynamic data on the patient's brain radiation, including arterial input function and changes in drug concentration in various brain regions.
[0169] The dose dynamic adjustment module 7 dynamically adjusts the radiopharmaceutical dose based on optimization strategies and real-time data through a multi-layer decision tree mechanism to achieve personalized and precise drug delivery.
[0170] The adjustment results and observation data are fed back to the reinforcement learning model as feedback, forming a closed loop, allowing the system to continuously learn and optimize.
[0171] The following specific application example illustrates the working process and effects of this invention:
[0172] A 70-year-old Alzheimer's patient, weighing 65 kg, needs to undergo an amyloid PET scan. Traditional methods calculate a fixed dose of 325 MBq (5 MBq / kg) based on weight. Using the system of this invention, the process is as follows:
[0173] First, multimodal MRI data of the patient were acquired, including T1-weighted MRI, DWI, VBM, and AVN sequences.
[0174] The system performs fine-grained partitioning of the brain, identifying Alzheimer's disease-related areas such as the hippocampus and frontal cortex, as well as non-specific binding areas such as the striatum.
[0175] Based on the partitioned data, the system extracts parameters such as metabolic baseline and cerebrospinal fluid flux of each region to construct the first reward function; at the same time, it establishes a pharmacokinetic model of amyloid tracer to construct the second reward function.
[0176] The reinforcement learning model is trained based on two reward functions to generate personalized dose optimization strategies.
[0177] The initial dose was set at 290 MBq, which is about 10% lower than the conventional method.
[0178] After administration, the system monitored changes in radiopharmaceutical concentrations in various brain regions in real time. The hippocampal uptake was only 65% of the expected level (mildly insufficient), while uptake in the frontal cortex was normal, and uptake in the striatum was slightly higher but within the safe range.
[0179] Based on the decision tree, the system administered an additional 30 MBq (10% of the initial dose) and extended the dynamic scan to 75 minutes.
[0180] The final PET images showed that the signal-to-noise ratio of the target areas (hippocampus and frontal cortex) was improved by 25%, the edge sharpness was improved by 18%, and the total radiation dose was reduced by 7% compared with the traditional method.
[0181] Follow-up showed that the patient's diagnostic accuracy improved, while radiation-related discomfort symptoms decreased.
[0182] This example demonstrates that the system of the present invention can dynamically adjust the dosage of radiopharmaceuticals based on individual patient characteristics and real-time feedback, achieving personalized and precise drug delivery while reducing radiation risks and improving diagnostic results.
[0183] The above description is merely a preferred embodiment of the present invention and is not intended to limit the invention. Various modifications and variations can be made to the present invention by those skilled in the art. Any modifications, equivalent substitutions, improvements, etc., made within the spirit and principles of the present invention should be included within the scope of protection of the present invention.
Claims
1. A method for dynamic optimization of radiopharmaceutical dosage based on reinforcement learning, characterized in that, include: Acquire multimodal magnetic resonance imaging data of the patient's brain; Based on the multimodal magnetic resonance imaging data, the patient's brain is finely divided into regions to obtain brain region data; Based on the brain region data, a first reward function is constructed; Based on the brain region data, a pharmacokinetic model was established, and a second reward function was constructed. Based on the first reward function and the second reward function, a reinforcement learning model is trained to obtain a radiopharmaceutical dosage optimization strategy. Real-time acquisition of dynamic data on radiation in the patient's brain; as well as The radiopharmaceutical dosage is dynamically adjusted based on the radiopharmaceutical dosage optimization strategy and the brain radioactivity dynamic data.
2. The method according to claim 1, characterized in that, The acquisition of multimodal magnetic resonance imaging data of the patient's brain includes: Collect MRI, DWI, VBM, and AVN multimodal imaging data of the patient's brain; Spatial registration processing is performed on the multimodal image data to align different modal images to the same spatial coordinate system; and A fused feature map is generated based on the registered multimodal image data.
3. The method according to claim 1, characterized in that, The detailed partitioning of the patient's brain includes: The patient's brain is divided into the basic regions of gray matter, white matter, deep nuclei, ventricles, and cerebellum; The basic regions are further subdivided into functional areas of the ventricles, deep nuclei, cerebellum, brainstem, cortex, striatum, hippocampus, thalamus, basal ganglia, and midbrain shell; and The accuracy of partitioning was verified by anatomical feature point marking and boundary detection.
4. The method according to claim 1, characterized in that, The construction of the first reward function includes: The brain region metabolic baseline, cerebrospinal fluid flux, Aβ deposition score, and hippocampal uptake ratio in PET images were extracted from the brain region partitioning data as parameters. Design reward items corresponding to the distribution of drugs in the target area, cerebrospinal fluid clearance efficiency, plaque binding effect, and hippocampal uptake level; Design penalties for drug accumulation in non-target areas, excessive cerebrospinal fluid flux, and excessive striatal uptake; and A first reward function is generated based on the reward and penalty terms.
5. The method according to claim 1, characterized in that, The establishment of the pharmacokinetic model and the construction of the second reward function include: Pharmacokinetic parameters of the drug in plasma, such as half-life, volume of distribution in brain tissue, elimination rate constant, and brain tissue uptake rate, were extracted. A pharmacokinetic model was established based on the aforementioned pharmacokinetic parameters; Design an index that maximizes the area under the time curve of drug concentration in the target region; Design drug clearance and retention balance indicators; and A second reward function is generated based on the maximization metric and the balance metric.
6. The method according to claim 1, characterized in that, The training reinforcement learning model includes: First, train the reinforcement learning algorithm for K rounds using the first reward function; The reinforcement learning algorithm is then trained for M rounds using the second reward function, where K≠M; When the directions guided by the first reward function and the second reward function conflict, a conflict coordination mechanism is activated; and The results of the two-stage training are integrated using a weighted average method to form a comprehensive optimization strategy.
7. The method according to claim 6, characterized in that, The reinforcement learning algorithm is a SARSA algorithm based on a value function, including: The state space is defined by the combination of brain region metabolic baseline, cerebrospinal fluid flux, Aβ deposition score, hippocampal uptake ratio in PET images, and pharmacokinetic parameters. The action space is defined by the combination of parameters including the dosage, injection rate, and injection time of the radiopharmaceutical; and A state transition model is constructed based on a pharmacokinetic model to predict the new state that the system will transition to after performing a specific action in the current state.
8. The method according to claim 1, characterized in that, The real-time acquisition of dynamic radiation data from the patient's brain includes: Real-time acquisition of partial arterial input function data via arterial sampling or imaging methods; Real-time PET scans were used to monitor changes in radiopharmaceutical concentrations in different areas of the brain; and The collected raw data is denoised, corrected, and standardized to improve data quality.
9. The method according to claim 1, characterized in that, The dynamic adjustment of the radiopharmaceutical dosage includes: The initial dose is determined based on the patient's basic information such as weight, age, and renal function; Set thresholds for key parameters and monitor in real time whether they exceed the safe range; Construct a multi-layer decision tree to determine whether the dosage needs to be adjusted based on real-time monitoring data and preset thresholds; When drug uptake in the white matter and deep nuclei of the brain is found to be insufficient, the dosage is increased and / or the dynamic scan is extended to 90 minutes, while the radiation dose is reduced; and The adjusted results and observation data are then fed back to the reinforcement learning model as feedback, forming a closed loop.
10. A dynamic optimization system for radiopharmaceutical dosage based on reinforcement learning, characterized in that, include: The multimodal imaging acquisition module is used to acquire multimodal magnetic resonance imaging data of the patient's brain; The brain region partitioning module is used to perform fine partitioning of the patient's brain based on the multimodal magnetic resonance imaging data to obtain brain region partitioning data; The first reward function construction module is used to construct a first reward function based on the brain region partitioning data; The pharmacokinetic model module is used to establish a pharmacokinetic model and construct a second reward function based on the brain region data. The reinforcement learning model training module is used to train a reinforcement learning model based on the first reward function and the second reward function to obtain a radiopharmaceutical dosage optimization strategy. The dynamic data acquisition module is used to collect dynamic data on the radiation in the patient's brain in real time. as well as The dose dynamic adjustment module is used to dynamically adjust the radiopharmaceutical dose based on the radiopharmaceutical dose optimization strategy and the brain radioactivity dynamic data.
Citation Information
Cited By
Critical patient analgesic drug dosage dynamic regulation and control system based on reinforcement learning
CN121528415A