A method for adjusting dosage parameters
By combining feature transformation of the DVH curve with a reinforcement learning model, the problems of time-consuming parameter adjustment and reliance on human experience in traditional radiotherapy planning are solved, achieving efficient and accurate automated adjustment of dose parameters and improving the quality and consistency of radiotherapy planning.
Patent Information
- Application Number
- CN202311418913.0
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2023-10-30
- Publication Date
- 2025-10-28
- Estimated Expiration
- 2043-10-30
AI Technical Summary
Existing radiotherapy planning methods suffer from problems such as time-consuming parameter tuning, reliance on human experience, lack of consistency, and low model learning efficiency. In particular, in IMRT, traditional TPS systems require manual adjustment, machine learning models fail to fully explore the features of the DVH curve, and the discrete form of reinforcement learning actions limits the model's expressive power.
By performing feature transformation on the DVH curve, using basis expansion and integration to generate feature vectors as input to the reinforcement learning model, and combining this with neural networks to generate the distribution of adjustment actions, a reinforcement learning model is established to automatically adjust the dose parameters.
It improves the accuracy and automation of dose adjustment, reduces manual intervention, and enhances the quality and efficiency of radiotherapy planning.
Smart Images

Figure CN117482411B_ABST
Abstract
Description
Technical Field
[0001] This invention relates to a method for adjusting dosage parameters. Background Technology
[0002] Malignant tumors (cancer) are the second leading cause of death worldwide and one of the most serious public health problems facing the Chinese population. China has a high incidence of cancer, and the incidence and mortality rates of cancer have been on the rise for the past decade. Among the many cancer treatments, radiotherapy is one of the most commonly used methods. The design of the radiotherapy plan directly affects the efficiency and efficacy of treatment, playing a decisive role in the treatment outcome. High-precision radiotherapy plans help reduce the incidence of complications and the risk of cancer recurrence.
[0003] Intensity-modulated radiation therapy (IMRT), which was proposed and developed in the 1990s, is one of the most commonly used radiotherapy techniques. By optimizing the radiotherapy plan design, it can protect the normal tissues near the target area as much as possible while meeting the dose requirements of the target area, thereby reducing damage to normal tissues.
[0004] Unlike other radiotherapy techniques, IMRT typically employs a reverse optimization approach during treatment planning. Specifically, the physician first determines the prescription dose for the target area and the dose limits for organs at risk. Then, based on the determined objective function and field parameters, the system optimization module derives the corresponding parameter values to ensure that the dose distribution for each structure meets clinical requirements. Intensity-modulated radiotherapy (IMRT) adjusts the intensity of the beams within multiple irradiation fields and uses dynamically variable beams to irradiate the tumor region, achieving a dose distribution that conforms to the shape of the tumor and is relatively uniform within the region. This avoids applying high doses to normal tissues, thereby improving treatment efficacy. Simultaneously, to reduce risks during treatment and ensure the quality and efficiency of the treatment plan, rigorous quality control is required throughout the entire process from patient registration to the end of treatment.
[0005] However, due to the numerous parameters and complex logical relationships among them involved in the planning process, physicists in traditional manual planning need to spend a significant amount of time during the parameter tuning phase to determine reasonable objective functions, constraints, and dose limit weights. Furthermore, the quality of planning design largely depends on the physician's clinical experience, parameter tuning skills, and understanding of radiation physics and the internal optimization algorithms of the planning system. These factors lead to inconsistent and uneven quality in planning designs.
[0006] Therefore, an increasing number of scholars are focusing on the automation and standardization of radiotherapy planning. Numerous studies have already conducted feasibility studies on automated intensity-modulated radiotherapy (IMRT) planning based on the automated planning module of a treatment planning system (TPS). The methods for automating IRT planning can be mainly divided into the following three categories:
[0007] Radiotherapy planning based on Treatment Planning System (TPS). This method primarily uses the patient's medical imaging data (such as CT scans) to assist in planning and dose calculation. The plan is then evaluated and optimized by the physician, and finally validated to ensure its accuracy and feasibility. Traditional treatment planning systems provide physicians with a wealth of functions and tools to help them develop personalized radiotherapy plans for each patient, ensuring precise dose delivery and maximizing the protection of surrounding critical organs and healthy tissues. Currently, in most hospitals, physicians typically rely on traditional treatment planning systems when developing radiotherapy plans. However, this method cannot achieve true automation; it often requires further manual adjustments and optimizations, consuming significant manpower. Furthermore, due to human interference, the final plan is also influenced by the physicist's experience, introducing a degree of uncertainty.
[0008] Machine learning-based radiotherapy planning. This method primarily introduces machine learning algorithms to train and optimize the model, learning potential patterns in samples. Compared to traditional radiotherapy planning, it can improve the efficiency, accuracy, and personalization of radiotherapy planning, supporting decision-making. However, most researchers directly use the obtained dose volume histogram (DVH) image or the coordinates of the DVH curve as input to the machine learning model. This approach does not fully reflect the essential characteristics of the DVH curve, leading to low model learning efficiency.
[0009] This method employs reinforcement learning-based radiotherapy planning. It builds a deep learning model based on patient image information or dose distribution data related to the radiotherapy plan to predict the patient's dose distribution after the plan is implemented. Based on these dose distributions, it determines the necessary adjustment parameters, thus achieving automated parameter adjustment. However, current research often designs the next action in reinforcement learning as discrete, modeling action prediction as a classification problem. This contradicts the reality that dose is typically adjusted continuously by physicists, thus limiting the model's expressive power.
[0010] The three methods mentioned above are commonly used in clinical or research IMRT protocols, and each has its shortcomings. Therefore, designing a protocol that can overcome the deficiencies of all three methods is an urgent problem to be solved. Summary of the Invention
[0011] The purpose of this invention is to overcome the shortcomings of the prior art and provide a method for adjusting dose parameters by first performing feature transformation on the DVH curve and then using the transformed vector as the input of a reinforcement learning model. This method can fully extract the feature information in the DVH curve and improve the accuracy of dose adjustment.
[0012] The objective of this invention is achieved through the following technical solution: a method for adjusting a dosage parameter, comprising the following steps:
[0013] Step 1: Obtain initial DVH curves for the target area and organs at risk using TPS: First, input the patient's computed tomography (CT) dataset, structural segmentation dataset, and processing plan parameters (including number of irradiations, beam angle, and radiation mode, etc.) to calculate the metrology matrix d. ij Then, based on user-defined constraints and objective functions, the reverse plan is optimized to generate results describing the cumulative dose value after each voxel is irradiated; finally, the dose values of all voxels are statistically analyzed by volume fraction to obtain a dose-volume histogram (DVH) for evaluating the quality of the plan.
[0014] Step 2: Perform a function-type transformation on the obtained DVH curve: First, perform feature transformation on the DVH curve, and then use the transformed vector as the input to the reinforcement learning model; the feature transformation consists of the following two steps:
[0015] Step 21: Reduce the dimensionality of functional data using basis expansion: Basis expansion is the process of representing a function as a linear combination of a finite number of basis functions. The expression for the basis expansion of the function f(X) is as follows:
[0016]
[0017] Where β m h is the basis expansion coefficient. m (X) represents the basis functions, and M represents the number of basis functions;
[0018] Since DVH is functional data, meaning it represents the dose distribution information of each structure as a curve, M basis functions are selected to perform the basis expansion of the n DVH curves as shown above:
[0019]
[0020] Where, β i =(β) i1 ,β i2 ,…,β im ) represents the coefficients corresponding to the basis functions, i = 1, 2, ..., n; h m (t) is a basis function;
[0021] Step 22: Integrate the dimensionality-reduced data: After obtaining the expression for the basis expansion, select k basis functions to perform approximate integration ∫ on the expansion of each DVH curve. τ φ k (t)X n The transformed set of feature vectors (t)dt will serve as the input to the subsequent reinforcement learning model; among them, the feature vector obtained by approximate integration of the nth DVH curve is:
[0022]
[0023] Where φ1(t), φ2(t), …, φ k (t) is a basis function;
[0024] Step 3: Learn the distribution of actions through a neural network: Use a neural network encoder to generate and adjust the distribution of actions; specifically, the neural network uses the feature vector (input) obtained in the previous step. n The input is a compressed feature x; based on feature x, the distribution q(a|x) of actions is further learned and adjusted, and then an action is randomly selected from the action distribution q(a|x). As the next adjustment step;
[0025] Step 4: Establish a reinforcement learning model: Use the obtained feature vector as the input of the reinforcement learning model, and define the behavior, reward, and loss function in reinforcement learning in combination with the actual clinical situation. Use the adjustment action generated in Step 3 as the action of the reinforcement learning model. Train the reinforcement learning model through iterative interaction with the environment TPS, and finally complete the decision-making process to obtain a series of output actions, which are the optimal adjustment of the dose parameters of various structures of the patient in the current state.
[0026] The beneficial effects of this invention are: this invention first performs feature transformation on the DVH curve, and then uses the transformed vector as the input of the reinforcement learning model. This method can fully extract the feature information in the DVH curve, thereby improving the accuracy of dose adjustment. Attached Figure Description
[0027] Figure 1 This is a framework diagram of the adjustment model of the present invention;
[0028] Figure 2 This is a flowchart of the dosage parameter adjustment method of the present invention. Detailed Implementation
[0029] The technical solution of the present invention will be further described below with reference to the accompanying drawings.
[0030] Overall framework of the adjustment model of the present invention Figure 1 As shown, it is divided into three modules: a functional embedding module for feature extraction, a neural network module for generating and adjusting the distribution of actions, and an environment module.
[0031] like Figure 2 As shown, the method for adjusting the dosage parameters of the present invention includes the following steps:
[0032] Step 1: Obtain DVH curves for the target area and organs at risk using TPS: First, input the patient's computed tomography (CT) dataset, structural segmentation dataset, and processing plan parameters (including number of irradiations, beam angle, and radiation mode, etc.) to calculate the metrology matrix d. ij Then, based on user-defined constraints and objective functions, the reverse plan is optimized to generate results describing the cumulative dose value after each voxel is irradiated; finally, the dose values of all voxels are statistically analyzed by volume fraction to obtain a dose-volume histogram (DVH) for evaluating the quality of the plan.
[0033] Step 2: Perform a function-type transformation on the obtained DVH curve: First, perform feature transformation on the DVH curve, and then use the transformed vector as the input to the reinforcement learning model; the feature transformation consists of the following two steps:
[0034] Step 21: Reduce the dimensionality of functional data using basis expansion: Basis expansion is the process of representing a function as a linear combination of a finite number of basis functions. The expression for the basis expansion of the function f(X) is as follows:
[0035]
[0036] Where β m h is the basis expansion coefficient. m (X) represents the basis functions, and M represents the number of basis functions;
[0037] Since DVH is functional data, meaning it represents the dose distribution information of each structure as a curve, M basis functions are selected to perform the basis expansion of the n DVH curves as shown above:
[0038]
[0039] Where, β i =(β) i1 ,β i2 ,…,β im ) represents the coefficients corresponding to the basis functions, i = 1, 2, ..., n; h m (t) is a basis function;
[0040] Step 22: Integrate the dimensionality-reduced data: After obtaining the expression for the basis expansion, select k basis functions to perform approximate integration ∫ on the expansion of each DVH curve. τ φ k (t)X n The transformed set of feature vectors (t)dt will serve as the input to the subsequent reinforcement learning model; among them, the feature vector obtained by approximate integration of the nth DVH curve is:
[0041]
[0042] Where φ1(t), φ2(t), …, φ k (t) is a basis function;
[0043] Step 3: Learn the distribution of actions through a neural network: The distribution of actions is generated and adjusted using a neural network encoder process. Specifically, the neural network takes the functional embedding processed in the previous step as input and outputs compressed features x. Based on features x, the distribution of adjusted actions q(a|x) is further learned. It is assumed to be a Gaussian distribution, i.e., q(a|x) ~ N(μ,σ). 2 The parameter μ is the expectation of the action, which is a function of x, i.e., μ:=μ(x). In this invention, it is designed as a linear function, expressed as μ(x)=e T x+b, where e is the high-dimensional regression parameter and b is the one-dimensional bias term; σ 2 It is the variance of the action, also designed as a function of x, i.e., σ. 2 :=σ 2 (x), in this embodiment σ 2 (x) is designed as a positive function, such as logσ 2 =c T x+d, where c are high-dimensional regression parameters and d is a one-dimensional bias term; μ and σ 2The two functions only need to ensure that the range of the first function is the set of real numbers and the range of the second function is the set of positive numbers. Their forms can be chosen by the user and are not limited to the function forms listed in this embodiment. Introducing continuous adjustment actions makes the model output more consistent with the actual process of dosage adjustment by physicists; these are all generated by an encoder, and then an action is randomly selected from the action distribution q(a|x). As the next adjustment step;
[0044] Step 4: Establish a reinforcement learning model: Use the obtained feature vector as the input of the reinforcement learning model, and define the behavior, reward, and loss function in reinforcement learning in combination with the actual clinical situation. Use the adjustment action generated in Step 3 as the action of the reinforcement learning model. Train the reinforcement learning model through the TPS iterative interaction between the agent and the environment to complete the decision-making process and obtain a series of output actions, which are the optimal adjustment of the dose parameters of various structures of the patient in the current state.
[0045] Those skilled in the art will recognize that the embodiments described herein are intended to help the reader understand the principles of the invention, and should be understood that the scope of protection of the invention is not limited to such specific statements and embodiments. Those skilled in the art can make various other specific modifications and combinations based on the technical teachings disclosed in this invention without departing from the spirit of the invention, and these modifications and combinations are still within the scope of protection of this invention.
Claims
1. A method for adjusting a dosage parameter, characterized in that, Includes the following steps: Step 1: Obtain initial DVH curves for the target area and organs at risk using TPS: First, input the patient's computed tomography dataset, structural segmentation dataset, and processing plan parameters to calculate the metrology matrix d. ij Then, based on user-defined constraints and objective functions, the reverse planning is optimized to generate results describing the cumulative dose value after each voxel is irradiated; finally, the dose values of all voxels are statistically analyzed by volume fraction to obtain the dose-volume histogram (DVH) used to evaluate the quality of the plan. Step 2: Perform a function-type transformation on the obtained DVH curve: First, perform feature transformation on the DVH curve, and then use the transformed vector as the input to the reinforcement learning model; the feature transformation consists of the following two steps: Step 21: Reduce the dimensionality of the functional data using basis expansion. Select M basis functions to perform basis expansion on n DVH curves: Where, β i =(β) i1 ,β i2 ,…,β im ) represents the coefficients corresponding to the basis functions, i = 1, 2, ..., n; h m (t) represents the basis functions; M is the number of basis functions; Step 22: Integrate the dimensionality-reduced data: After obtaining the expression for the basis expansion, select k basis functions to perform approximate integration ∫ on the expansion of each DVH curve. t φ k (t)X n The transformed set of feature vectors (t)dt will serve as the input to the subsequent reinforcement learning model; among them, the feature vector obtained by approximate integration of the nth DVH curve is: Where φ1(t), φ2(t), …, φ k (t) is a basis function; Step 3: Learn the distribution of actions through a neural network: Generate and adjust the distribution of actions using a neural network encoder; specifically, the neural network takes the feature vector obtained in the previous step as input and outputs compressed features x; based on features x, further learn and adjust the distribution q(a|x), and then randomly select an action from the action distribution q(a|x). As the next adjustment step; Step 4: Establish a reinforcement learning model: Use the obtained feature vector as the input of the reinforcement learning model, and define the behavior, reward, and loss function in reinforcement learning in combination with the actual clinical situation. Use the adjustment action generated in Step 3 as the action of the reinforcement learning model. Train the reinforcement learning model through iterative interaction with the environment TPS, and finally complete the decision-making process to obtain a series of output actions, which are the optimal adjustment of the dose parameters of various structures of the patient in the current state.
Citation Information
Patent Citations
Method and system for measuring and calculating radiotherapy radiation dose distribution and dose objective functions
CN110404184A
Convex inverse planning method
CN110603075A