Dynamic trajectory intensity modulated radiotherapy plan optimization method and system
By transforming dynamic trajectory intensity-modulated radiotherapy (IMRT) planning into an autonomous driving problem, and employing motion planning and reinforcement learning to optimize accelerator parameters, the problem of insufficient optimization speed and quality in existing technologies is solved, achieving more efficient radiotherapy planning optimization.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-04-01
- Publication Date
- 2026-03-24
AI Technical Summary
Existing dynamic trajectory intensity-modulated radiotherapy (IMRT) planning optimization methods are insufficient to meet clinical requirements in terms of optimization speed and quality, thus limiting the improvement of radiotherapy planning quality.
The dynamic trajectory intensity-modulated radiotherapy planning optimization problem is transformed into an autonomous driving problem in the control field. Motion planning algorithm and reinforcement learning network are used for optimization. Inverse optimization is performed using state transition equation and objective function. Combined with nonlinear programming and limit position constraints, the accelerator parameters are optimized.
It significantly improves the speed and quality of radiotherapy planning optimization, reduces the difficulty of optimization, and can improve the treatment accuracy and efficiency of radiotherapy planning without hardware modification.
Smart Images

Figure CN120242337B_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of radiotherapy equipment technology, specifically to a method and system for optimizing dynamic trajectory intensity-modulated radiotherapy plans. Background Technology
[0002] Dynamic Trajectory Radiation Treatment (DTRT) is an emerging radiotherapy technique. Its core principle is to expand the treatment space and achieve higher quality radiotherapy through the synchronized movement of the treatment bed, gantry, and collimator. Compared to traditional Volumetric Modulated Arc Treatment (VMAT) and Intensity Modulated Radiation Treatment (IMRT), DTRT significantly improves the quality of treatment planning. In 2011, Yang Yingli et al. from the Sloan Kettering Cancer Institute in the United States proposed that synchronizing the movement of the collimator with the treatment bed can significantly improve the quality of radiotherapy planning and enhance the therapeutic effect of radiotherapy. In the same year, Gregory et al. proposed that synchronizing the movement of the gantry with the treatment bed for arc therapy can improve the quality of radiotherapy planning for various cancers, including breast cancer, brain tumors, and prostate cancer, enhancing the precision and effectiveness of treatment. With the continuous development and improvement of DTRT technology, an increasing number of clinical studies have confirmed that this technology can effectively improve the quality of radiotherapy planning and improve patient prognosis without reducing treatment efficiency, making it an innovative technology with broad application prospects in the field of radiotherapy. DTRT is expected to become a major breakthrough in the field of cancer radiotherapy, significantly improving the precision and effectiveness of tumor treatment and driving radiotherapy technology to a higher level.
[0003] However, due to the large solution space and complex coupling relationships, existing dynamic trajectory intensity-modulated radiotherapy (IMRT) planning optimization methods are difficult to meet clinical requirements in terms of optimization speed and quality. This greatly hinders the application of IMRT in clinical treatment and limits the further improvement of radiotherapy planning quality.
[0004] Existing dynamic trajectory enhancement methods can be mainly divided into two categories: step-by-step methods and synchronous methods. Step-by-step methods first search for the firing field trajectory, and then optimize the intensity distribution based on the determined trajectory. The advantage of this type of algorithm is that it only adds the function of firing field trajectory search to the traditional VMAT plan optimization algorithm, thus making it relatively easy to implement. This type of algorithm was first proposed by Yang in 2011. This method uses a geometric heuristic to optimize the firing field trajectory, and then optimizes the intensity distribution based on the obtained trajectory. In 2015, Wild established a firing field selection anchor point based on non-coplanar IMRT, and a VMAT firing field trajectory optimization algorithm that transitions to complex trajectories by connecting anchor points. Smyth proposed a local search nVMAT plan optimization algorithm (FBLS) based on injection optimization, which attempts to solve the problem that geometric heuristic algorithms cannot reflect local trajectory adjustments. In 2018, Michael proposed using the A* search algorithm in the path planning field to optimize the firing field path. In 2024, Gian proposed a column-generation method to determine the radiation field path, which showed good planning quality for patients with nasopharyngeal carcinoma, breast cancer, and esophageal cancer. It can be seen that a step-by-step approach can achieve dynamic trajectory intensity-modulated radiotherapy (IMRT) planning, but because this method cannot consider the relationship between the radiation field trajectory and intensity distribution, it is difficult to obtain the optimal plan. Synchronous methods can fully consider the coupling relationship between the radiation field trajectory and intensity distribution during optimization, theoretically achieving optimal plan quality. In 2018, Lyu proposed considering the coupling relationship between the radiation field trajectory and intensity distribution through alternating optimization. In the same year, Dong proposed a Monte Carlo tree search-based method. With a sufficient number of searches, this method can theoretically achieve optimal plan quality, but it involves high computational cost and low optimization efficiency. Mullins proposed a column-generation-based method in 2024, which improves optimization efficiency while maintaining plan quality, but the plan execution time is long. Overall, synchronous methods are still in the exploratory stage due to their difficulty. Although some breakthroughs have been achieved, problems such as long optimization and plan execution times and the need to improve optimization quality remain. Summary of the Invention
[0005] The purpose of this invention is to provide a method and system for optimizing dynamic trajectory intensity-modulated radiotherapy plans, so as to solve at least one of the technical problems existing in the background art.
[0006] To achieve the above objectives, the present invention adopts the following technical solution:
[0007] In a first aspect, the present invention provides a method for optimizing dynamic trajectory intensity-modulated radiotherapy (IMRT) plans, comprising:
[0008] Obtain basic patient information and imaging data; determine the region of interest for radiotherapy (i.e., the tumor target area requiring radiotherapy and the normal tissues and organs to be protected) and the clinical prescription dose requirements for the region of interest; calculate the intensity-modulated radiation physical dose based on conventional radiotherapy planning parameters or dynamic trajectory intensity-modulated radiotherapy planning parameters, and perform inverse optimization of the dynamic trajectory intensity-modulated radiotherapy dose.
[0009] As a further limitation of the first aspect of the present invention, the state quantity of the system at time t is set as the dose distribution D in the patient's body. t The control parameters include the bed corner C. t Frame angle G t MLC blade motion trajectory M t and dose rate u t, The state transition equation of the system is then expressed as:
[0010] D t+Δt =D t +u t ·P(C t G t G t )·Δt
[0011] Wherein, P is the dose calculation engine, which gives the dose at the calculation point when calculating the unit intensity based on the bed angle, frame angle, and MLC blade position.
[0012] As a further definition of the first aspect of this invention, the target conditions related to the tumor target area are defined as T(*), the limiting conditions related to organs at risk are defined as R(*), the requirement for treatment efficiency is defined as E(*), and the mechanical constraints on the gantry, treatment bed, and MLC are set as U(*). Then, the model optimized using the motion planning method is expressed as follows:
[0013] [C t+Δt ',G t+Δt ',M t+Δt ',u t+Δt ']=argmin(L(T(D t+Δt ),E(u t+Δt ),R(D t+Δt )))
[0014] stU(C′ t+Δt ,G′ t+Δt ,M′ t+Δt ,u′ t+Δt )<0
[0015] Where L(*) is the objective function used by the motion planning algorithm.
[0016] As a further limitation of the first aspect of the present invention, optionally, nonlinear programming is used to predict the optimal N control point parameters for the future, and the first predicted control point parameter is selected for execution and dose distribution update. If the target area dose distribution reaches the optimization termination condition at this time, the optimization is stopped and the final plan is output; if not, the execution time is increased and optimization continues. The plan optimization model is expressed as:
[0017] [C t+Δt ',G t+Δt ',M t+Δt ',u t+Δt ']=argmin(NLP(T(D t+Δt ),E(u t+Δt ),R(D t+Δt ),N))
[0018] stU(C′ t+Δt ,G′ t+Δt ,M′ t+Δt ,u′ t+Δt )<0
[0019] Here, NLP represents the nonlinear programming optimization engine used in the model predictive control algorithm, and N represents the need to predict N future control points and calculate the objective function.
[0020] As a further limitation of the first aspect of the present invention, optionally, the planned optimization model for optimizing accelerator parameters using a reinforcement learning network is expressed as follows:
[0021] [C t+Δt ',G t+Δt ',M t+Δt ',u t+Δt ']=argmax(Q(T(D t+Δt ),E(u t+Δt ),R(D t+Δt )))
[0022] stU(C′ t+Δt ,G′ t+Δt ,M′ t+Δt ,u′ t+Δt )<0
[0023] Where Q() represents the Q function used in reinforcement learning;
[0024] A limit position constraint layer was added to the reinforcement learning network:
[0025] l8 = [Sigmoid(a·(AA)] min ))+Sigmoid(a·(A max -A))]·l7
[0026] Here, A is a vector, and each element of A represents a machine parameter of the accelerator; A max and A min These represent the maximum and minimum allowed values for the action machine parameter A, respectively. Sigmoid is the Sigmoid function.
[0027] Set the corresponding action based on the maximum rate of change of each parameter:
[0028]
[0029] in This represents the maximum rate of change for the k-th parameter.
[0030] In a second aspect, the present invention provides a dynamic trajectory intensity-modulated radiotherapy planning optimization system, comprising:
[0031] The patient data management module is used to obtain basic patient information and patient imaging data;
[0032] The region of interest delineation module is used to determine the region of interest for radiotherapy and the clinical prescription dose requirements for the region of interest;
[0033] The conventional radiotherapy planning module is used to calculate the physical radiation dose based on conventional radiotherapy plan parameters and to perform inverse optimization of the conventional radiotherapy plan dose.
[0034] The dynamic trajectory intensity-modulated radiotherapy (IMRT) planning module is used to calculate the intensity-modulated radiation dose based on the IMRT planning parameters and to perform inverse optimization of the IMRT dose.
[0035] As a further limitation of the second aspect of the present invention, the state quantity of the system at time t is set as the dose distribution D in the patient's body. t The control parameters include the bed corner C. t Frame angle G t MLC blade motion trajectory M t and dose rate u t, The state transition equation of the system is then expressed as:
[0036] D t+Δt =D t +u t ·P(C t G t G t )·Δt
[0037] Wherein, P is the dose calculation engine, which gives the dose at the calculation point when calculating the unit intensity based on the bed angle, frame angle, and MLC blade position.
[0038] As a further limitation of the second aspect of the present invention, the target conditions related to the tumor target area are defined as T(*), the limiting conditions related to organs at risk are defined as R(*), the requirement for treatment efficiency is defined as E(*), and the mechanical constraints on the gantry, treatment bed, and MLC are set as U(*). Then, the model optimized using the motion planning method is expressed as follows:
[0039] [C t+Δt ',G t+Δt ',M t+Δt ',u t+Δt ']=argmin(L(T(D t+Δt ),E(u t+Δt ),R(D t+Δt )))
[0040] stU(C′ t+Δt ,G′ t+Δt ,M′ t+Δt ,u′ t+Δt )<0
[0041] Where L(*) is the objective function used by the motion planning algorithm.
[0042] As a further limitation of the second aspect of the present invention, optionally, nonlinear programming is used to predict the optimal N control point parameters for the future, and the first predicted control point parameter is selected for execution and dose distribution update. If the target area dose distribution reaches the optimization termination condition at this time, the optimization is stopped and the final plan is output; if not, the execution time is increased and optimization continues. The plan optimization model is expressed as:
[0043] [C t+Δt ',G t+Δt ',M t+Δt ',u t+Δt ']=argmin(NLP(T(D t+Δt ),E(u t+Δt ),R(D t+Δt ),N))
[0044] stU(C′ t+Δt ,G′ t+Δt ,M′ t+Δt ,u′ t+Δt )<0
[0045] Here, NLP represents the nonlinear programming optimization engine used in the model predictive control algorithm, and N represents the need to predict N future control points and calculate the objective function.
[0046] As a further limitation of the second aspect of the present invention, optionally, the planned optimization model for optimizing accelerator parameters using a reinforcement learning network is expressed as follows:
[0047] [C t+Δt ',G t+Δt ',M t+Δt ',u t+Δt ']=argmax(Q(T(D t+Δt ),E(u t+Δt ),R(D t+Δt )))
[0048] stU(C′ t+Δt ,G′ t+Δt ,M′ t+Δt ,u′ t+Δt )<0
[0049] Where Q() represents the Q function used in reinforcement learning;
[0050] A limit position constraint layer was added to the reinforcement learning network:
[0051] l8 = [Sigmoid(a·(AA)] min ))+Sigmoid(a·(A max -A))]·l7
[0052] Here, A is a vector, and each element of A represents a machine parameter of the accelerator; A max and A min These represent the maximum and minimum allowed values for the action machine parameter A, respectively. Sigmoid is the Sigmoid function.
[0053] Set the corresponding action based on the maximum rate of change of each parameter:
[0054]
[0055] in This represents the maximum rate of change for the k-th parameter.
[0056] Thirdly, the present invention provides a non-transitory computer-readable storage medium for storing computer instructions, which, when executed by a processor, implement the dynamic trajectory intensity-modulated radiotherapy planning optimization method as described in the first aspect.
[0057] Fourthly, the present invention provides a computer device including a memory and a processor, wherein the processor and the memory communicate with each other, the memory stores program instructions that can be executed by the processor, and the processor calls the program instructions to execute the dynamic trajectory intensity-modulated radiotherapy planning optimization method as described in the first aspect.
[0058] Fifthly, the present invention provides an electronic device, comprising: a processor, a memory, and a computer program; wherein the processor is connected to the memory, the computer program is stored in the memory, and when the electronic device is running, the processor executes the computer program stored in the memory to cause the electronic device to execute instructions for implementing the dynamic trajectory intensity-modulated radiotherapy planning optimization method as described in the first aspect.
[0059] The beneficial effects of this invention are: improving the speed and quality of dynamic trajectory intensity-modulated radiotherapy (IMRT) planning optimization by utilizing advanced motion planning algorithms from autonomous driving. IMRT not only encompasses current research hotspots but also boasts low implementation difficulty, requiring no hardware modifications to enhance radiotherapy planning quality. The provided technical specifications not only offer a comprehensive and flexible framework and basic methods for IMRT technology but also describe specific implementation details, such as methods for designing IMRT plans using different motion planning algorithms. The methods and systems provided by this invention fill a gap in IMRT technology, which is of great significance for the translational application of IMRT technology and improving cancer treatment levels.
[0060] The advantages of additional aspects of the invention will be set forth more clearly in the following description or will be learned by practice of the invention. Attached Figure Description
[0061] To more clearly illustrate the technical solutions of the embodiments of the present invention, the drawings used in the description of the embodiments will be briefly introduced below. Obviously, the drawings described below are only some embodiments of the present invention. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.
[0062] Figure 1 This is a schematic diagram of the dynamic trajectory intensity-modulated radiotherapy planning system architecture according to an embodiment of the present invention.
[0063] Figure 2 This is a flowchart of the dynamic trajectory intensity-modulated radiotherapy planning optimization based on model predictive control as described in an embodiment of the present invention.
[0064] Figure 3 This is a flowchart illustrating the dynamic trajectory intensity-modulated radiotherapy (IMRT) planning optimization based on reinforcement learning, as described in an embodiment of the present invention. Detailed Implementation
[0065] Embodiments of the present invention are described in detail below, examples of which are shown in the accompanying drawings, wherein the same or similar reference numerals denote the same or similar elements or elements having the same or similar functions throughout. The embodiments described below with reference to the accompanying drawings are exemplary and are only used to explain the present invention, and should not be construed as limiting the present invention.
[0066] It will be understood by those skilled in the art that, unless otherwise defined, all terms used herein (including technical and scientific terms) have the same meaning as commonly understood by one of ordinary skill in the art to which this invention pertains.
[0067] It should also be understood that terms such as those defined in general dictionaries should be understood to have meanings consistent with their meanings in the context of the prior art, and should not be interpreted in an idealized or overly formal sense unless defined as here.
[0068] Those skilled in the art will understand that, unless specifically stated otherwise, the singular forms “a,” “an,” “the,” and “the” used herein may also include the plural forms. It should be further understood that the term “comprising” as used in this specification means the presence of the stated features, integers, steps, operations, elements, and / or components, but does not exclude the presence or addition of one or more other features, integers, steps, operations, elements, and / or groups thereof.
[0069] In the description of this specification, references to terms such as "one embodiment," "some embodiments," "example," "specific example," or "some examples," etc., indicate that a specific feature, structure, material, or characteristic described in connection with that embodiment or example is included in at least one embodiment or example of the present invention. Furthermore, the specific features, structures, materials, or characteristics described may be combined in any suitable manner in one or more embodiments or examples. Without contradiction, those skilled in the art can combine and integrate the different embodiments or examples described in this specification, as well as the features of those different embodiments or examples.
[0070] To facilitate understanding of the present invention, the present invention will be further explained and described below with reference to the accompanying drawings and specific embodiments. However, the specific embodiments do not constitute a limitation on the embodiments of the present invention.
[0071] Those skilled in the art should understand that the accompanying drawings are merely schematic diagrams of embodiments, and the components in the drawings are not necessarily essential for implementing the present invention.
[0072] This invention provides a method and system for optimizing dynamic trajectory intensity-modulated radiotherapy (DTRT) plans. It transforms the DTRT planning problem into an autonomous driving problem in the control field, thus establishing a novel DTRT planning design and system. This invention takes a completely new perspective, analogizing the DTRT planning optimization problem to an autonomous driving problem in the control field. Leveraging motion planning algorithms, and utilizing their efficiency and adaptability in complex decision-making and dynamic optimization, an innovative automatic DTRT planning optimization method is designed. This method fully utilizes the advantages of motion planning algorithms in rapidly calculating and processing high-dimensional complex optimization problems, significantly improving optimization efficiency, shortening plan generation time, and achieving breakthroughs in global optimization quality. Based on this optimization method, this invention also provides a DTRT optimization planning system. The method and system of this invention provide a novel and efficient DTRT optimization solution for clinical practice.
[0073] This invention proposes a method for designing dynamic trajectory intensity-modulated radiotherapy (IMRT) plans using an autonomous driving approach from the control domain. This design method is a reverse optimization-based design approach. Because the solution space for dynamic trajectory IMRT plans is large and the solution parameters have complex coupling relationships, traditional radiotherapy planning optimization algorithms struggle to find the optimal radiotherapy plan quickly and effectively. Therefore, this embodiment analogizes dynamic trajectory IMRT planning to an autonomous driving problem, as shown in Table 1 below.
[0074] Table 1. Conceptual Analogy between Dynamic Trajectory Intensity-Modulated Radiotherapy Planning Optimization and Autonomous Driving Optimization
[0075]
[0076]
[0077] In this way, the planning optimization problem can be transformed into an autonomous driving optimization problem. Autonomous driving problems often involve more complex and unpredictable situations than planning optimization, and current domestic autonomous driving technology is already capable of handling these situations well. Therefore, transforming the dynamic trajectory intensity-modulated radiotherapy planning problem into an autonomous driving problem, and then using advanced autonomous driving algorithms to solve it, can quickly and effectively find the optimal solution.
[0078] Example 1
[0079] In this embodiment 1, a method for optimizing dynamic trajectory intensity-modulated radiotherapy (IMRT) plans is first provided, including: initializing accelerator parameters according to the dynamic trajectory IMRT plan parameters; calculating the intensity-modulated radiation physical dose and performing inverse optimization of the dynamic trajectory IMRT dose.
[0080] The initialization of accelerator parameters includes setting the initial bed angle and frame angle based on the principle of proximity. For MLC blades, the initial values are the blade positions after conforming the target area according to the initial bed angle and frame angle.
[0081] In this embodiment, the dynamic trajectory intensity-modulated radiotherapy (IMRT) planning problem is transformed into an autonomous driving problem in the control field, thereby establishing a novel dynamic trajectory IMRT planning method and system.
[0082] This embodiment proposes a dynamic trajectory intensity-modulated radiotherapy (IMRT) planning system. This system adds the functionality of dynamic trajectory IMRT planning to existing conventional radiotherapy planning systems. It includes a user-friendly interface for dynamic trajectory IMRT planning and anti-collision features considering accelerator-treatment bed linkage to ensure treatment safety. Finally, the system can output either a dynamic trajectory IMRT plan or a conventional plan using the standard Dicom RT planning data format.
[0083] In this embodiment, the aforementioned planning system relies on a motion planning-based dynamic trajectory intensity-modulated radiography (IMRT) planning optimization model. This model, in practice, transforms the one-time optimization of all accelerator parameters into a step-by-step optimization of accelerator parameters according to treatment time, thereby reducing the optimization difficulty. The model sets the system's state variable at time t as the dose distribution D within the patient's body. t The control parameters include the bed corner C. t Frame angle G t MLC blade motion trajectory M t and dose rate u t The state transition equation of the system can then be expressed as:
[0084] D t+Δt =D t +u t ·P(C t G t G t )·Δt
[0085] Where P is the dose calculation engine, such as Collapsed Cone Convolution (CCC), which gives the dose at the calculation point when calculating the unit intensity based on the bed angle, frame angle, and MLC blade position.
[0086] Next, the target conditions related to the tumor target area are defined as T(), the limiting conditions related to organs at risk are defined as R(), the requirement for treatment efficiency is defined as E(), and the mechanical constraints on the gantry, treatment bed, and MLC are set as U(). Then, the model optimized using the motion planning method can be expressed as follows:
[0087] [Ct+Δt ',G t+Δt ',M t+Δt ',u t+Δt ']=argmin(L(T(D t+Δt ),E(u t+Δt ),R(D t+Δt )))
[0088] stU(C′ t+Δt ,G′ t+Δt ,M′ t+Δt ,u′ t+Δt )<0
[0089] Where L() is the objective function used by the motion planning algorithm. Because D t At time t, argmin() represents the parameters C, G, M, and u calculated using an advanced path planning optimization algorithm that minimize the objective function L. It can be seen that at each time step, all parameters related to the trajectory and intensity distribution of the firing field are optimized, and a local optimum is obtained in the time domain. It should be noted that state-of-the-art motion planning algorithms predict the state variables after time n*Δt when solving for the action at time t+Δt, to avoid the impact of temporal locality on the performance of the final solution set.
[0090] The planning system architecture proposed in this embodiment is as follows: Figure 1 As can be seen, similar to conventional dose-rate planning systems, this planning system includes a patient data management module to manage basic patient information, patient imaging data, and patient planning data; the region of interest (ROI) delineation module allows physicians or physicists to manually or automatically delineate the target area, organs at risk, and other contours. Then, the patient's planning CT scans and the delineated ROI information can be used for plan design.
[0091] Based on clinical needs, this planning system can design two types of plans: 1) A conventional planning module includes three sub-modules: conventional plan parameter input, physical dose calculation, and conventional plan inverse optimization. This module can design traditional conventional dose IMRT or VMAT radiotherapy plans; 2) A dynamic trajectory IMRT planning module includes three sub-modules: parameter input, physical dose calculation, and dynamic trajectory IMRT plan inverse optimization. The parameter input sub-module is mainly used to input the initial bed angle, gantry angle, dose optimization parameters, etc. The physical dose calculation module is consistent with the conventional planning module. The dynamic trajectory IMRT plan inverse optimization function automatically calculates and generates the optimal dynamic trajectory IMRT plan parameters based on the set plan parameters.
[0092] This planning system can transmit both conventional and dynamic trajectory intensity-modulated radiotherapy (IMRT) plans to other systems in the standard Dicom RT format via the plan parameter output module. Using this system, one can design IMRT plans and freely choose between conventional VMAT or IMRT plans to meet various clinical needs.
[0093] In this embodiment, as Figure 1 , Figure 2 As shown, two dynamic trajectory intensity-modulated radiotherapy planning optimization models based on different motion planning algorithms are proposed.
[0094] For the first model, such as Figure 1 As shown, this is a dynamic trajectory intensity-modulated radiotherapy (IMRT) planning optimization model based on model predictive control. First, the accelerator parameters are initialized. Initial bed angles and gantry angles are set according to the principle of proximity field placement. For MLC blades, the blade position after conformal adjustment of the target area based on the initial bed angles and gantry angles is used as the initial value.
[0095] Next, a model predictive control algorithm (MMC) is used to optimize the dynamic trajectory intensity-modulated radiotherapy (IMRT) plan. This algorithm is one of the most commonly used motion planning algorithms in the field of autonomous driving. The algorithm uses nonlinear programming to predict the optimal parameters for N control points in the future, and selects the first predicted control point parameter to execute and update the dose distribution. If the target dose distribution reaches the optimization termination condition, the optimization stops and the final plan is output. If not, the execution time is increased and optimization continues. The plan optimization model can be represented as:
[0096] [C t+Δt ',G t+Δt ',M t+Δt ',u t+Δt ']=argmin(NLP(T(D t+Δt ),E(u t+Δt ),R(D t+Δt ),N))
[0097] stU(C′ t+Δt ,G′ t+Δt ,M′ t+Δt ,u′ t+Δt )<0
[0098] Here, NLP refers to the nonlinear programming optimization engine used in Model Predictive Control (MMCC) algorithms. N represents the number of future control points to predict and the objective function to be calculated. It's important to note that N is a hyperparameter that needs to be manually set. If N is too large, the optimization time may be too long; if N is too small, the final solution may be suboptimal.
[0099] For the second model, such as Figure 2As shown, the accelerator parameters are first initialized. The initialization method is the same as the first method. Then, a reinforcement learning network is used to optimize the accelerator parameters. The planned optimization model can then be expressed as:
[0100] [C t+Δt ',G t+Δt ',M t+Δt ',u t+Δt ']=argmax(Q(T(D t+Δt ),E(u t+Δt ),R(D t+Δt )))
[0101] stU(C′ t+Δt ,G′ t+Δt ,M′ t+Δt ,u′ t+Δt )<0
[0102] Here, Q() represents the Q (Quality) function used in reinforcement learning. In each optimization, the control point parameters are adjusted to maximize the Q function, thereby optimizing the dynamic trajectory intensity-modulated radiotherapy plan. Furthermore, unlike traditional optimization algorithms, this embodiment incorporates a limit position constraint layer into the reinforcement learning network.
[0103] l8 = [Sigmoid(a·(AA)] min ))+Sigmoid(a·(A max -A))]·l7
[0104] Where A is a vector, and each element of A represents a machine parameter of the accelerator. max and A min These represent the maximum and minimum allowed values for the motion machine parameter A, respectively. These values are determined by the accelerometer's performance. 'a' is a large scaling factor. 'Sigmoid' refers to the Sigmoid function. Using the above formula, the scores for actions that might cause the accelerometer parameters to exceed their range can be forced to zero. Finally, based on the maximum rate of change of each parameter, the corresponding action is set:
[0105]
[0106] in This represents the maximum rate of change for the k-th parameter. This method avoids the problem of parameters changing too rapidly.
[0107] In this embodiment, based on the above method, the following is achieved: Figure 3 The dynamic trajectory intensity-modulated radiotherapy planning optimization system shown includes:
[0108] The patient data management module is used to acquire basic patient information and patient imaging data; the region of interest delineation module determines the region of interest for radiotherapy (i.e., the tumor target area requiring radiotherapy and the normal tissues and organs that need to be protected); the conventional radiotherapy planning module determines the clinical prescription dose requirements for the region of interest and calculates the physical radiation dose based on the conventional radiotherapy planning parameters, performing reverse optimization of the conventional radiotherapy plan; the dynamic trajectory intensity-modulated radiotherapy (IMRT) planning module determines the clinical prescription dose requirements for the region of interest and calculates the intensity-modulated radiation dose based on the IMRT planning parameters, performing reverse optimization of the IMRT plan.
[0109] Example 2
[0110] This embodiment 2 provides a non-transitory computer-readable storage medium for storing computer instructions. When executed by a processor, the computer instructions implement the dynamic trajectory intensity-modulated radiotherapy planning optimization method described above. The method includes:
[0111] Obtain basic patient information and imaging data; determine the region of interest for radiotherapy (i.e., the tumor target area requiring radiotherapy and the normal tissues and organs to be protected) and the clinical prescription dose requirements for the region of interest; calculate the intensity-modulated radiation physical dose based on conventional radiotherapy planning parameters or dynamic trajectory intensity-modulated radiotherapy planning parameters, and perform inverse optimization of the dynamic trajectory intensity-modulated radiotherapy dose.
[0112] Example 3
[0113] This embodiment 3 provides a computer device, including a memory and a processor, wherein the processor and the memory communicate with each other, and the memory stores program instructions that can be executed by the processor. The processor calls the program instructions to execute the dynamic trajectory intensity-modulated radiotherapy planning optimization method described above, the method including:
[0114] Obtain basic patient information and imaging data; determine the region of interest for radiotherapy (i.e., the tumor target area requiring radiotherapy and the normal tissues and organs to be protected) and the clinical prescription dose requirements for the region of interest; calculate the intensity-modulated radiation physical dose based on conventional radiotherapy planning parameters or dynamic trajectory intensity-modulated radiotherapy planning parameters, and perform inverse optimization of the dynamic trajectory intensity-modulated radiotherapy dose.
[0115] Example 4
[0116] This embodiment 4 provides an electronic device, including: a processor, a memory, and a computer program; wherein, the processor is connected to the memory, and the computer program is stored in the memory. When the electronic device is running, the processor executes the computer program stored in the memory to cause the electronic device to execute instructions to implement the dynamic trajectory intensity-modulated radiotherapy planning optimization method as described above, the method including:
[0117] Obtain basic patient information and imaging data; determine the region of interest (ROI) for radiotherapy (i.e., the tumor target area requiring radiotherapy and the normal tissues and organs to be protected) and the clinical prescription dose requirements for the RIO; calculate the intensity-modulated radiation dose based on conventional radiotherapy planning parameters or dynamic trajectory intensity-modulated radiotherapy (IMRT) planning parameters, and perform inverse optimization of the IMRT dose.
[0118] Those skilled in the art will understand that embodiments of the present invention can be provided as methods, systems, or computer program products. Therefore, the present invention can take the form of a completely hardware embodiment, a completely software embodiment, or an embodiment combining software and hardware aspects. Furthermore, the present invention can take the form of a computer program product embodied on one or more computer-usable storage media (including, but not limited to, disk storage, CD-ROM, optical storage, etc.) containing computer-usable program code.
[0119] This invention is described with reference to flowchart illustrations and / or block diagrams of methods, apparatus (systems), and computer program products according to embodiments of the invention. It will be understood that each block of the flowchart illustrations and / or block diagrams, and combinations of blocks in the flowchart illustrations and / or block diagrams, can be implemented by computer program instructions. These computer program instructions can be provided to a processor of a general-purpose computer, special-purpose computer, embedded processor, or other programmable data processing apparatus to produce a machine, such that the instructions, which execute via the processor of the computer or other programmable data processing apparatus, generate instructions for implementing the flowchart illustrations and / or block diagrams. Figure 1 One or more processes and / or boxes Figure 1 A device that provides the functions specified in one or more boxes.
[0120] These computer program instructions may also be stored in a computer-readable storage medium that can direct a computer or other programmable data processing device to function in a particular manner, such that the instructions stored in the computer-readable storage medium produce an article of manufacture including instruction means, which are implemented in a process Figure 1 One or more processes and / or boxes Figure 1 The function specified in one or more boxes.
[0121] These computer program instructions may also be loaded onto a computer or other programmable data processing equipment, whereby a series of operational steps are performed to produce a computer-implemented process, thereby providing instructions that execute on the computer or other programmable equipment for implementing the process. Figure 1 One or more processes and / or boxes Figure 1 The steps of the function specified in one or more boxes.
[0122] While the specific embodiments of the present invention have been described above in conjunction with the accompanying drawings, this is not intended to limit the scope of protection of the present invention. Those skilled in the art should understand that, based on the technical solutions disclosed in the present invention, various modifications or variations that can be made by those skilled in the art without creative effort should be included within the scope of protection of the present invention.
Claims
1. A dynamic trajectory intensity-modulated radiotherapy planning optimization system, characterized in that, include: The patient data management module is used to obtain basic patient information and patient imaging data; The region of interest delineation module is used to determine the region of interest for radiotherapy and the clinical prescription dose requirements for the region of interest; The conventional radiotherapy planning module is used to calculate the physical radiation dose based on conventional radiotherapy plan parameters and to perform inverse optimization of the conventional radiotherapy plan dose. The dynamic trajectory intensity-modulated radiotherapy (IMRT) planning module is used to calculate the intensity-modulated radiation dose based on the IMRT planning parameters and to perform inverse optimization of the IMRT dose. Here, the state variable of the system at time t is set as the dose distribution in the patient's body. D t The control quantity includes the bed corner C t Frame corner G t MLC blade motion trajectory M t and dose rate u t , The state transition equation of the system is then expressed as: ; Wherein, P is the dose calculation engine, which gives the dose at the calculation point when calculating the unit intensity based on the bed angle, frame angle, and MLC blade position; The target conditions related to the tumor target area are defined as T( The limits for organ-risk substances are defined as R( The required treatment efficiency is defined as E( ); The mechanical constraints on the gantry, treatment bed, and MLC are set to U( If ), then the model optimized using motion planning methods is represented as: ; ; Among them, L( ) is the objective function used by the motion planning algorithm; Nonlinear programming is used to predict the optimal parameters for N control points in the future. The first predicted control point parameter is selected for execution, and the dose distribution is updated. If the target dose distribution reaches the optimization termination condition at this point, the optimization stops and the final plan is output. If the condition is not met, the execution time is increased and optimization continues. The plan optimization model is expressed as: ; ; Here, NLP represents the nonlinear programming optimization engine used in model predictive control algorithms, and N represents the number of future variables that need to be considered. N Predict the values of each control point and calculate the objective function.
2. The dynamic trajectory intensity-modulated radiotherapy (IMRT) planning optimization system according to claim 1, wherein the planning optimization model for optimizing accelerator parameters using a reinforcement learning network is expressed as follows: ; ; in, Q() represents the Q function used in reinforcement learning; A limit position constraint layer was added to the reinforcement learning network: ; Here, A is a vector, where each element represents a machine parameter of the accelerator; A max and A min These represent the maximum and minimum allowed values for the action machine parameter A, respectively. Sigmoid is the Sigmoid function. Set the corresponding action based on the maximum rate of change of each parameter: ; in Indicates the first k The maximum rate of change of each parameter.
3. A method for optimizing dynamic trajectory intensity-modulated radiotherapy (IMRT) plans based on the system described in claim 1 or 2, characterized in that, include: Obtain basic patient information and patient imaging data; Determine the region of interest for radiotherapy and the clinical prescription dose requirements for the region of interest; the region of interest for radiotherapy is the tumor target area that needs to be radiotreated and the normal tissues and organs that need to be protected. Based on conventional radiotherapy planning parameters or dynamic trajectory intensity-modulated radiotherapy planning parameters, calculate the intensity-modulated radiation physical dose and perform inverse optimization of the dynamic trajectory intensity-modulated radiotherapy dose.
4. The dynamic trajectory intensity-modulated radiotherapy planning optimization method according to claim 3, characterized in that, The state variables of the system at time t are set as the dose distribution within the patient's body. D t The control quantity includes the bed corner C t Frame corner G t MLC blade motion trajectory M t and dose rate u t, The state transition equation of the system is then expressed as: ; Wherein, P is the dose calculation engine, which gives the dose at the calculation point when calculating the unit intensity based on the bed angle, frame angle, and MLC blade position.
5. The dynamic trajectory intensity-modulated radiotherapy planning optimization method according to claim 3, characterized in that, The target conditions related to the tumor target area are defined as T( The limits for organ-risk substances are defined as R( The required treatment efficiency is defined as E( ); The mechanical constraints on the gantry, treatment bed, and MLC are set to U( If ), then the model optimized using motion planning methods is represented as: ; ; Among them, L( ) is the objective function used by the motion planning algorithm.
6. The dynamic trajectory intensity-modulated radiotherapy planning optimization method according to claim 3, characterized in that, Nonlinear programming is used to predict the optimal parameters for N control points in the future. The first predicted control point parameter is selected for execution, and the dose distribution is updated. If the target dose distribution reaches the optimization termination condition at this point, the optimization stops and the final plan is output. If the condition is not met, the execution time is increased and optimization continues. The plan optimization model is expressed as: ; ; Here, NLP represents the nonlinear programming optimization engine used in model predictive control algorithms, and N represents the number of future variables that need to be considered. N Predict the values of each control point and calculate the objective function.
7. The dynamic trajectory intensity-modulated radiotherapy planning optimization method according to claim 3, characterized in that, The planned optimization model for optimizing accelerator parameters using reinforcement learning networks is expressed as follows: ; ; Where Q() represents the Q function used in reinforcement learning; A limit position constraint layer was added to the reinforcement learning network: ; Here, A is a vector, and each element of A represents a machine parameter of the accelerator; A max and A min These represent the maximum and minimum allowed values for the action machine parameter A, respectively. Sigmoid is the Sigmoid function. Set the corresponding action based on the maximum rate of change of each parameter: ; in Indicates the first k The maximum rate of change of each parameter.
Citation Information
Patent Citations
Dynamic multi-axes trajectory optimization and delivery method for radiation treatment
US20130142310A1