Adaptive optimization treatment planning system and reinforcement learning method
Through deep reinforcement learning methods, the parameters and dose limitations of the radiotherapy planning system are automatically optimized, which solves the problems of inefficient and local optimal design of traditional radiotherapy planning, and achieves efficient and precise individualized treatment planning generation.
Patent Information
- Application Number
- CN202510194812.2
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-02-21
- Publication Date
- 2025-05-30
AI Technical Summary
Traditional radiotherapy plans are inefficient in design, easily fall into local optimality, and it is difficult to effectively deal with complex individual differences in patients.
Deep reinforcement learning method is adopted to train deep Q networks, and the parameters and dose limit conditions of the treatment planning system are automatically optimized to generate an individualized treatment plan that meets clinical needs.
It improves the efficiency and accuracy of radiotherapy planning design, reduces the workload of manual adjustment, and can automatically generate optimization results that meet clinical requirements.
Smart Images

Figure CN120072200A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of radiotherapy, and in particular to an adaptive optimization treatment planning system and a reinforcement learning method. Background Art
[0002] Radiotherapy is one of the important means for treating cancer. Precise radiotherapy plan design is crucial for improving the treatment effect and reducing side effects. Traditional radiotherapy plan design usually relies on doctors' experience and manual adjustment and optimization methods. When facing complex patient individual differences and large-scale optimization problems, there are often problems of low efficiency and local optimality. With the development of medical imaging technology and the improvement of computing power, computer-aided automated radiotherapy plan design has gradually become a research hotspot.
[0003] Deep Learning (DL) and Reinforcement Learning (RL) are two very active research directions in the field of artificial intelligence. Deep learning is based on neural networks and can process a large amount of data and automatically learn complex patterns, while reinforcement learning is a method whose goal is to enable a machine to learn through interaction with the environment so as to make better decisions. Deep Reinforcement Learning (DRL) combines the perception ability of deep learning and the decision-making ability of reinforcement learning. By using a deep neural network to predict the optimal action, an agent can learn strategies in a complex environment. Applying it to a radiotherapy planning system can effectively optimize the treatment dose distribution, reduce the damage to critical organs, and improve the efficiency and accuracy of plan design.
[0004] Traditional radiotherapy plan design usually relies on doctors' experience and manual adjustment and optimization methods. When facing complex patient individual differences and large-scale optimization problems, there are often problems of low efficiency and local optimality.
[0005] Therefore, we propose an adaptive optimization treatment planning system and a reinforcement learning method. Summary of the Invention
[0006] In view of the above-mentioned shortcomings in the existing production technologies, the present applicant provides an adaptively optimized treatment planning system and a reinforcement learning method. By adopting the deep reinforcement learning method, a trained model is used to automatically optimize the parameters of the system and set the dose limitation conditions for the target area and organs at risk according to the completely contoured target area CT images, so as to be used as the input for the optimization system to obtain a treatment plan that meets clinical requirements. When a physicist makes a treatment plan, the optimized limitation conditions can be automatically generated according to the input patient images, target areas, and dose requirements of organs at risk, reducing the manual adjustment workload during the planning process. After being input into the planning system, an optimized result that meets clinical requirements can be obtained in one step.
[0007] The technical solution adopted by the present invention is as follows:
[0008] An adaptively optimized treatment planning system includes the following steps:
[0009] A physicist makes training cases, records the operations at each time and the current state of the treatment planning system software, and saves the completed treatment plans.
[0010] Integrate the state parameters of the treatment planning system software as a data set and define it as the state space of reinforcement learning; integrate the operations performed by the physicist and define it as the action space; integrate the parameters affecting the quality of the treatment plan and define it as the reward space.
[0011] Take the state space as the input and the action space as the output to train the DQN network. The network trains the best actions in different states, which interacts with the treatment planning system software. At the same time, the reinforcement learning intelligent Agent also participates in the process, recording the returned states and rewards.
[0012] After the Agent is trained, place the model into the treatment planning system software. When a new patient case is input, the software converts the case into an initial state. The Agent makes a plan according to the trained actions and continuously modifies and adjusts according to the feedback rewards to obtain a result with the optimal reward, which is a treatment plan that meets clinical specifications.
[0013] Its further features are as follows:
[0014] The state parameters of the treatment planning system software include the spatial position of the treatment node and the shape of the irradiation field, the positions, shapes, and sizes of the tumor and organs at risk, and the dose distribution.
[0015] The parameters affecting the quality of the treatment plan include the increase or decrease in the dose coverage rate of the target area and the increase or decrease in the doses of the organs at risk and normal tissue organs.
[0016] The present invention also provides a reinforcement learning method for an adaptively optimized treatment planning system, including the following steps:
[0017] Define the state space S to contain the configuration information of the radiotherapy plan;
[0018] Define the action space A to contain the optimization functions for adding and adjusting tumors and organs at risk, adjusting the MU weights, and adding and reducing the spatial targets for constraints;
[0019] Define the reward mechanism R to provide feedback based on the results generated by the intelligent agent's different actions;
[0020] Use the Deep Q-learning method to combine the neural network with reinforcement learning, separating the target policy from the action policy. Q is an action-value function, and the Agent is the intelligence of reinforcement learning. At each time step t, the Agent selects an action a t , and obtains a reward r t , enters a new state s t+1 , and updates the Q value, that is:
[0021] Q(s t , a t ) ← Q(s t , a t ) + α · [r t + γ max Q(s t+1 , a t ) - Q(s t , a t )];
[0022] Among them, α is the learning rate, and γ is the discount factor;
[0023] Extract the completed case data from the treatment planning system software to construct the training dataset; perform feature extraction on the CT images, and convert the original image data into a state vector through convolution and pooling operations; the state vector represents different processing stages of the case; define a discrete action space, in the treatment planning system software, each action corresponds to a decision and operation;
[0024] Action-value function: Input the state vector into the deep Q-network, and the network outputs a discrete action-value function Q(s, a, θ), where s represents the state, a represents the action, and θ represents the network weights;
[0025] Selection of Q value: After obtaining the Q values of all possible actions, use the ε-greedy strategy to select an action to balance exploration and exploitation, select the maximum Q value, denoted as Q_max, which represents the expected reward for performing the optimal action in the current state;
[0026] Target network and target value calculation: Use the state at t+1 to input the target network to obtain the target Q value, which serves as the target value during the training process; the weights of the target network are updated regularly to stabilize the learning process;
[0027] Loss function calculation: Calculate the loss function L by comparing the Q value of the current state and the Q value of the target state; the loss function uses the temporal difference error, i.e., the difference between the predicted Q value and the target Q value;
[0028] Network weight update: The network continuously adjusts the weights θ through the gradient descent algorithm to minimize the loss function L, thereby learning the network weights suitable for the current task;
[0029] At the same time, an experience replay module is introduced. The experiences collected by the Agent are stored in an array and repeatedly used to train the Agent during training.
[0030] The configuration information including the radiotherapy plan includes radiotherapy equipment information, the input patient images, the positions, shapes, and sizes of tumors and organs at risk, and the dose distribution.
[0031] Define the reward according to the clinical goal: Give a positive reward when the tumor coverage rate increases and the dose of the organ at risk decreases; give a negative reward when the tumor coverage rate decreases and the dose of the organ at risk increases; give a penalty when the dose of the organ at risk exceeds the limit and the dose of the target area is too low.
[0032] The completed case data includes radiotherapy equipment information, CT images, tumor target area and organ at risk information, and dose distribution information.
[0033] Store the trajectory of the Agent in an array. The array is specified with a certain size for training; when the array exceeds the specified size, the oldest data is deleted; this module can improve the learning efficiency and stability.
[0034] The beneficial effects of the present invention are as follows:
[0035] When making a treatment plan according to the present invention, in order to obtain a suitable target area distribution, the physicist needs to adjust the limiting conditions and optimization system parameters based on his own experience multiple times. When there are multiple target areas, more adjustments are required. This method is an automated method and can automatically make an individualized treatment plan that meets clinical requirements based on the input patient images.
[0036] At the same time, the present invention also has the following advantages:
[0037] (1) Using the method of reinforcement learning, the Agent can continuously learn during the process of making a case. After the training is completed, the Agent can also continuously optimize the weight parameters according to new cases and the plans newly made by the physicist to obtain a model suitable for the current treatment plan system software.
[0038] (2) Combine the neural network with reinforcement learning to integrate the high-dimensional state space, i.e., the complex case conditions of the patient, and the high-dimensional action space, i.e., various adjustment operations of the physicist, into a trainable data set respectively, facilitating the training of the reinforcement learning intelligence. BRIEF DESCRIPTION OF THE DRAWINGS
[0039] Figure 1 It is a schematic diagram of the treatment planning system of the present invention.
[0040] Figure 2 It is a schematic diagram of the neural network training of the present invention.
[0041] Figure 3 It is a schematic diagram of the training input and output of the present invention. DETAILED DESCRIPTION OF THE INVENTION
[0042] The following combines the drawings to illustrate the detailed implementation manners of the present invention.
[0043] As Figure 1 shown, an adaptive optimization treatment planning system includes the following steps:
[0044] The physicist makes training cases, records the operations at each time and the current state of the treatment planning system software, and saves the completed treatment plan.
[0045] Integrate the state parameters of the treatment planning system software, i.e., the spatial position of the treatment nodes, the shape of the irradiation field, the positions, shapes and sizes of the tumors and critical organs, and the dose distribution, as a data set, and define it as the state space of the reinforcement learning; integrate the operations performed by the physicist and define it as the action space; integrate the parameters affecting the quality of the treatment plan, such as the rise and fall of the dose coverage rate, and the rise and fall of the doses of critical organs and normal tissue organs, and define it as the reward space.
[0046] Use the state space as the input and the action space as the output to train the Deep Q Neural network (DQN network). The network trains the best actions in different states. This process interacts with the treatment planning system software, and at the same time, the reinforcement learning intelligent Agent also participates in the process, recording the returned states and rewards.
[0047] After the Agent is trained, place the model into the treatment planning system software. When a new patient case is input, the software converts the case into an initial state. The Agent makes a plan according to the trained actions and continuously modifies and adjusts according to the feedback rewards to obtain a result with the optimal reward, i.e., a treatment plan that meets the clinical specifications.
[0048] As Figures 2 - 3As shown in the figure, a reinforcement learning method for an adaptive optimization treatment planning system includes the following steps:
[0049] Define the state space S (state), which includes the configuration information of the radiotherapy plan: radiotherapy equipment information, input patient images, the positions, shapes and sizes of tumors and organs at risk, and dose distribution;
[0050] Define the action space A (action), which includes adding and adjusting the optimization functions for tumors and organs at risk, adjusting the MU weights, and adding or reducing the spatial targets for limitation;
[0051] Define the reward mechanism R (reward), and give feedback according to the results generated by the intelligent agent taking different actions; define the reward according to the clinical goals: give a positive reward when the tumor coverage rate increases or the dose of the organ at risk decreases; give a negative reward when the tumor coverage rate decreases or the dose of the organ at risk increases; give a penalty when the dose of the organ at risk exceeds the limit value and the target area dose is too low;
[0052] Use the Deep Q-learning method to combine the neural network with reinforcement learning, separate the target policy and the action policy. Q is an action value function, and Agent is the intelligence of reinforcement learning. At each time t, the Agent selects an action a t , and obtains a reward r t , enters a new state s t+1 , and updates the Q value, that is:
[0053] Q(s t , a t ) ← Q(s t , a t ) + α · [r t + γ max Q(s t+1 , a t ) - Q(s t , a t )]
[0054] where α is the learning rate and γ is the discount factor;
[0055] Input part: Extract the completed case data from the treatment planning system software, including radiotherapy equipment information, CT images, tumor target area and organ-at-risk information, and dose distribution information; these data are used to construct the training dataset; perform feature extraction on the CT images, and convert the original image data into a state vector through convolution and pooling operations, that is, the state; these state vectors represent different processing stages of the cases; define a discrete action space, and in the treatment planning system software, each action corresponds to a decision or operation.
[0056] Output section: Action value function: The state vector is input into the Deep Q-Network (DQN), and the network outputs a discrete action value function Q(s, a; θ), where s represents the state, a represents the action, and θ represents the network weights. The function outputs a vector, and each element in the vector corresponds to the expected future reward for performing a specific action in the current state.
[0057] Q-value selection: After obtaining the Q-values of all possible actions, the ε-greedy strategy is used to select actions to balance exploration and exploitation. The maximum Q-value is selected, denoted as Q_max, which represents the expected reward for performing the optimal action in the current state.
[0058] Target network and target value calculation: The t+1 state (i.e., the next state) is input into the target network to obtain the target Q-value, which will be used as the target value during the training process. The weights of the target network are updated periodically to stabilize the learning process.
[0059] Loss function calculation: By comparing the Q-value of the current state (current value) and the Q-value of the target state (target value), the loss function L is calculated. The loss function uses the temporal difference error (TD-error), which is the difference between the predicted Q-value and the target Q-value.
[0060] Network weight update: The network continuously adjusts the weights θ through the gradient descent algorithm to minimize the loss function L, thereby learning the network weights suitable for the current task.
[0061] At the same time, an experience replay module (Replay memory) is introduced. The experiences collected by the Agent are stored in an array and repeatedly used to train the Agent during training. The trajectory of the Agent (s t , a t , r t , s t+1 ) is stored in the array. The array has a specified size for training. When the array exceeds the specified size, the oldest data is deleted. This module can improve learning efficiency and stability.
[0062] When making a treatment plan, in order to obtain a suitable target area distribution, the physicist needs to adjust the optimization constraints and optimization system parameters based on their own experience multiple times. When there are multiple target areas, more adjustments are required. This method is an automated method that can automatically make an individualized treatment plan that meets clinical requirements based on the input patient images.
[0063] Using the method of reinforcement learning, the Agent can continuously learn during the process of making cases. After the training is completed, the Agent can also continuously optimize the weight parameters according to new cases and the plans newly made by the physicist to obtain a model suitable for the current treatment planning system software.
[0064] Combining a neural network with reinforcement learning enables the high-dimensional state space, i.e., the complex case situation of the patient, and the high-dimensional action space, i.e., various adjustment operations of the physicist, to be combined separately to form a trainable data set, facilitating the training of the reinforcement learning intelligence.
[0065] The above description is an explanation of the present invention, not a limitation thereof. The scope defined by the present invention is to be seen in the claims, and any form of modification may be made within the protection scope of the present invention.
Claims
1. An adaptive optimization treatment planning system, characterized in that: The steps include: The physiotherapist creates training cases, records the operation at each time and the current status of the treatment plan system software, and saves the completed treatment plan; The state parameters of the treatment planning system software are integrated as a data set and defined as the state space of reinforcement learning; the operations performed by the physicist are integrated and defined as the action space; the parameters that affect the quality of the treatment plan are integrated and defined as the reward space; The state space is used as input and the action space is used as output to train the DQN network. The network trains the best actions under different states. This process interacts with the treatment planning system software. At the same time, the reinforcement learning intelligent agent also participates in the process and records the returned state and reward. After the agent training is completed, the model is placed in the treatment planning system software. When a new patient case is input, the software converts the case to the initial state. The agent makes a plan based on the trained actions and continuously modifies and adjusts it based on the feedback rewards to obtain an optimal reward result, which is a treatment plan that meets clinical standards.
2. The adaptive optimization treatment planning system according to claim 1, characterized in that: The status parameters of the treatment planning system software include the spatial position of the treatment node and the shape of the irradiation field, the position, shape and size of the tumor and organs at risk, and the dose distribution.
3. The adaptive optimization treatment planning system according to claim 2, characterized in that: The parameters that affect the quality of the treatment plan include the increase or decrease of the target area dose coverage, and the increase or decrease of the dose of organs at risk and normal tissues and organs.
4. A reinforcement learning method for an adaptively optimized treatment planning system, characterized in that: The steps include: Define the state space S to contain the configuration information of the radiotherapy plan; Defining the action space A includes adding and adjusting the optimization functions of tumors and organs at risk, adjusting MU weights, and adding and reducing space targets for restrictions; Define a reward mechanism R to provide feedback based on the results of different intelligent behaviors; Using the Deep Q-learning method, neural networks are combined with reinforcement learning. The target strategy is separated from the action strategy. Q is an action value function. The agent is the intelligence of reinforcement learning. At each time t, the agent selects an action a. t , get a reward r t , enter the new state s t+1 , and update the Q value, that is: Q(s t ,a t )←Q(s t ,a t )+α·[r t +γmaxQ(s t+1 ,a t )-Q(s t ,a t )]; Among them, α is the learning rate and γ is the discount factor; Extract completed case data from the treatment planning system software to build a training data set; perform feature extraction on CT images and convert raw image data into state vectors through convolution and pooling operations; the state vectors represent different processing stages of the case; define a discrete action space, where each action in the treatment planning system software corresponds to a decision, operation; Action value function: The state vector is input into the deep Q network, and the network outputs a discrete action value function Q(s, a, θ), where s represents the state, a represents the action, and θ represents the network weight; Q value selection: After obtaining the Q values of all possible actions, use the ε-greedy strategy to select actions to balance exploration and exploitation, and select the largest Q value, denoted as Q_max, which represents the expected reward for performing the optimal action in the current state; Target network and target value calculation: Use the t+1 state to input the target network to obtain the target Q value, which is used as the target value in the training process; the weights of the target network are updated regularly to stabilize the learning process; Loss function calculation: By comparing the Q value of the current state with the Q value of the target state, the loss function L is calculated; the loss function uses the temporal difference error, that is, the difference between the predicted Q value and the target Q value; Network weight update: The network continuously adjusts the weight θ through the gradient descent algorithm to minimize the loss function L, thereby learning the network weight suitable for the current task; At the same time, the experience replay module is introduced to store the experience collected by the agent into an array, and the experience is repeatedly used to train the agent during training.
5. The reinforcement learning method for an adaptively optimized treatment planning system according to claim 4, characterized in that: The configuration information containing the radiotherapy plan includes radiotherapy equipment information, input patient images, the location, shape and size of tumors and organs at risk, and dose distribution.
6. The reinforcement learning method for an adaptively optimized treatment planning system according to claim 5, characterized in that: Rewards are defined according to clinical objectives: positive rewards are given if tumor coverage increases and dose to organs at risk decreases; negative rewards are given if tumor coverage decreases and dose to organs at risk increases; penalties are given if dose to organs at risk exceeds the limit and the dose to the target area is too low.
7. The reinforcement learning method for an adaptively optimized treatment planning system according to claim 6, characterized in that: The completed case data includes radiotherapy equipment information, CT images, tumor target area and organ at risk information, and dose distribution information.
8. The reinforcement learning method for an adaptively optimized treatment planning system according to claim 7, characterized in that: The Agent's trajectory is stored in an array with a certain size for training. When the array exceeds the specified size, the oldest data is deleted. This module can improve learning efficiency and stability.