Dynamic trajectory intensity modulated radiotherapy plan optimization method and system

By transforming dynamic trajectory intensity-modulation radiotherapy planning into autonomous driving control problems, using motion planning and reinforcement learning optimization algorithms, a dynamic trajectory intensity-modulation radiotherapy planning optimization system was designed, which solved the problem of insufficient optimization speed and quality in the existing technology, and achieved more efficient radiotherapy planning optimization and cancer treatment effects.

CN120242337AActive Publication Date: 2025-07-04CANCER INST & HOSPITAL CHINESE ACADEMY OF MEDICAL SCI

Patent Information

Application Number
CN202510400082.7
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-04-01
Publication Date
2025-07-04
Estimated Expiration
2045-04-01

AI Technical Summary

Technical Problem

The existing dynamic trajectory intensity-modulated radiotherapy plan optimization methods are difficult to meet clinical requirements in terms of optimization speed and quality, which hinders the application and quality improvement of this technology in radiotherapy.

Method used

The dynamic trajectory intensity-modulated radiotherapy plan optimization problem is transformed into control problems in autonomous driving, and the motion planning algorithm and reinforcement learning network are used for optimization. A dynamic trajectory intensity-modulated radiotherapy plan optimization system is designed, including patient data management, area of ​​interest outline, routine planning design and dynamic trajectory intensity-modulated radiotherapy plan design modules to calculate the radio physical dose through reverse optimization.

Benefits of technology

It significantly improves the optimization speed and quality of dynamic trajectory intensity-modulation radiotherapy plans, without hardware transformation, and improves the accuracy and effectiveness of cancer treatment.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120242337A_ABST
    Figure CN120242337A_ABST
Patent Text Reader

Abstract

The invention provides a dynamic trajectory intensity modulated radiotherapy plan optimization method and system, and belongs to the technical field of radiotherapy equipment. A patient data management module is used for acquiring basic information, patient image data and patient radiotherapy plan data of a patient; the region-of-interest sketching module is used for determining a radiotherapy region-of-interest; the conventional plan design module is used for calculating a radiation physical dose according to the conventional plan parameters and performing reverse optimization of a conventional dose rate; and the dynamic track intensity-modulated radiotherapy plan design module is used for calculating the intensity-modulated radiation physical dose according to the dynamic track intensity-modulated radiotherapy plan parameters, and performing dynamic track intensity-modulated radiotherapy dose reverse optimization. According to the method, the optimization speed and quality of the dynamic trajectory intensity modulated radiation therapy plan are improved by using an advanced motion planning algorithm in automatic driving, and the quality of the radiation therapy plan can be improved without hardware modification; the blank of the dynamic trajectory intensity modulated radiotherapy technology is filled, and the cancer treatment level is improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of radiotherapy equipment, and particularly relates to a method and system for optimizing a dynamic trajectory intensity modulated radiotherapy plan. Background Art

[0002] Dynamic Trajectory Radiation Treatment (DTRT) is a new radiotherapy technology. Its core principle is to expand the treatment space and achieve higher-quality radiotherapy through the synchronous movement of the treatment couch, gantry, and collimator. Compared with traditional Volumetric Modulated Arc Treatment (VMAT) and Intensity Modulated Radiation Treatment (IMRT), DTRT can significantly improve the quality of treatment plans. In 2011, Yang Yingli et al. from Memorial Sloan-Kettering Cancer Center in the United States proposed that by synchronously moving the collimator and the treatment couch, the quality of radiotherapy plans can be significantly improved, and the treatment effect of radiotherapy can be enhanced. In the same year, Gregory et al. proposed that by performing arc therapy with the synchronous movement of the gantry and the treatment couch, the quality of radiotherapy plans for various cancers such as breast cancer, brain tumors, and prostate cancer can be improved, and the accuracy and effect of treatment can be enhanced. With the continuous development and improvement of DTRT technology, more and more clinical studies have confirmed that this technology can effectively improve the quality of radiotherapy plans and improve the treatment prognosis of patients without reducing the treatment efficiency, becoming an innovative technology with broad application prospects in the field of radiotherapy. DTRT is expected to become an important breakthrough in the field of cancer radiotherapy, significantly improving the accuracy and effect of tumor treatment and promoting the development of radiotherapy technology to a higher level.

[0003] However, due to the huge solution space and complex coupling relationships, existing methods for optimizing dynamic trajectory intensity modulated radiotherapy plans are difficult to meet clinical requirements in terms of optimization speed and quality, which greatly hinders the application of dynamic trajectory intensity modulated radiotherapy in clinical treatment and also limits the further improvement of the quality of radiotherapy plans.

[0004] The existing dynamic trajectory intensity modulation methods can be mainly divided into two categories: step-by-step methods and synchronous methods. The step-by-step methods first search for beam trajectories and then optimize the intensity distribution according to the determined trajectories. The advantage of this type of algorithm is that it only adds the function of beam trajectory search on the basis of the traditional VMAT plan optimization algorithm, so the implementation difficulty is relatively low. This type of algorithm was first proposed by Yang in 2011. This method uses a geometric heuristic method to optimize the beam trajectory and then optimizes the intensity distribution according to the obtained trajectory. Wild established the beam selection anchor points (Anchor point) based on non-coplanar IMRT in 2015, and proposed a VMAT beam trajectory optimization algorithm that transitions from the anchor point connection to complex trajectories. Smyth proposed a local search nVMAT plan optimization algorithm (FBLS) based on fluence optimization, which attempts to solve the problem that the geometric heuristic algorithm cannot reflect the local adjustment of the trajectory. In 2018, Michael proposed to use the A* search algorithm in the field of path planning to optimize the beam path. In 2024, Gian proposed to use the column generation method to determine the beam path, and the results showed that this method can achieve better plan quality for patients with nasopharyngeal cancer, breast cancer and esophageal cancer. It can be seen that the dynamic trajectory intensity modulation radiotherapy plan design can be realized by using a step-by-step method. However, since this type of method cannot consider the relationship between the beam trajectory and the intensity distribution, it is difficult to obtain the optimal plan. The synchronous method can fully consider the coupling relationship between the beam trajectory and the intensity distribution during the optimization process, and theoretically can achieve the optimal plan quality. In 2018, Lyu proposed to consider the coupling relationship between the two by alternately optimizing the beam trajectory and the intensity distribution. In the same year, Dong proposed a method based on Monte Carlo tree search. When the number of searches is large enough, this method can theoretically obtain the optimal plan quality, but the computational cost is large and the optimization efficiency is low. Mullins proposed a method based on column generation in 2024, which improved the optimization efficiency under the condition of ensuring the plan quality, but the plan execution time is long. Generally speaking, due to the great difficulty, the synchronous method is still in the exploration stage. Although certain breakthroughs have been made, there are still problems such as long optimization time and plan execution time, and the optimization quality needs to be improved. Summary of the Invention

[0005] The purpose of the present invention is to provide a dynamic trajectory intensity modulation radiotherapy plan optimization method and system to solve at least one of the technical problems in the above background technology.

[0006] To achieve the above purpose, the present invention adopts the following technical solutions:

[0007] In the first aspect, the present invention provides a dynamic trajectory intensity modulation radiotherapy plan optimization method, including:

[0008] Obtain the basic information of the patient and the patient's imaging data; determine the radiotherapy region of interest (i.e., the tumor target area to be radiated and the normal tissue organs to be protected) and the clinical prescription dose requirements for the region of interest; calculate the intensity-modulated radiotherapy dose according to the conventional radiotherapy plan parameters or the dynamic trajectory intensity-modulated radiotherapy plan parameters, and perform inverse optimization of the dynamic trajectory intensity-modulated radiotherapy dose.

[0009] As a further limitation of the first aspect of the present invention, the state quantity of the system at time t is set as the dose distribution D in the patient's body t , and the control quantity includes the couch angle C t , the gantry angle G t , the MLC leaf motion trajectory M t and the dose rate u t, Then the state transition equation of the system is expressed as:

[0010] D t+Δt = D t + u t · P(C t , G t , G t ) · Δt

[0011] where P is the dose calculation engine, which calculates the dose given to the calculation point per unit intensity according to the couch angle, gantry angle, and MLC leaf position.

[0012] As a further limitation of the first aspect of the present invention, the target condition related to the tumor target area is defined as T(*), the limit condition related to the organs at risk is positioned as R(*); the requirement for the treatment efficiency is defined as E(*); the mechanical limit conditions for the gantry, treatment couch, and MLC are set as U(*), then the model optimized by the motion planning method is expressed as:

[0013] [C t+Δt ', G t+Δt ', M t+Δt ', u t+Δt '] = argmin(L(T(D t+Δt ), E(u t+Δt ), R(D t+Δt )))

[0014] s.t. U(C′ t+Δt , G′ t+Δt , M′ t+Δt , u′ t+Δt ) < 0

[0015] where L(*) is the objective function used by the motion planning algorithm.

[0016] As a further limitation of the first aspect of the present invention, optionally, the future optimal N control point parameters are predicted using nonlinear programming, and the first predicted control point parameter is selected to execute and update the dose distribution. If the target area dose distribution reaches the optimization termination condition at this time, the optimization is stopped and the final plan is output. If not, the execution time is increased and the optimization continues. The plan optimization model is expressed as:

[0017] [C t+Δt ',G t+Δt ',M t+Δt ',u t+Δt ']=argmin(NLP(T(D t+Δt ),E(u t+Δt ),R(D t+Δt ),N))

[0018] s.t.U(C′ t+Δt ,G′ t+Δt ,M′ t+Δt ,u′ t+Δt )<0

[0019] Among them, NLP represents the nonlinear programming optimization engine used in the model predictive control algorithm, and N represents the need to predict the future N control points and calculate the objective function.

[0020] As a further limitation of the first aspect of the present invention, optionally, the plan optimization model for optimizing the accelerator parameters using a reinforcement learning network is expressed as:

[0021] [C t+Δt ',G t+Δt ',M t+Δt ',u t+Δt ']=argmax(Q(T(D t+Δt ),E(u t+Δt ),R(D t+Δt )))

[0022] s.t.U(C′ t+Δt ,G′ t+Δt ,M′ t+Δt ,u′ t+Δt )<0

[0023] Among them, Q() represents the Q function used in reinforcement learning;

[0024] A limit position constraint layer is added to the reinforcement learning network:

[0025] l8=[Sigmoid(a·(A - A min )) + Sigmoid(a·(A max - A))]·l7

[0026] Among them, A is a vector, and each element of A represents a machine parameter of the accelerator; A max and A min respectively represent the maximum and minimum values allowed for the action machine parameter A, and Sigmoid is the Sigmoid function;

[0027] Set corresponding actions according to the maximum change speed of each parameter:

[0028]

[0029] Among them represents the maximum change speed of the k-th parameter.

[0030] In a second aspect, the present invention provides a dynamic trajectory intensity-modulated radiotherapy plan optimization system, including:

[0031] A patient data management module for obtaining the basic information of the patient and the patient's image data;

[0032] A region of interest delineation module for determining the radiotherapy region of interest and the clinical prescription dose requirements for the region of interest;

[0033] A conventional radiotherapy plan design module for calculating the radiotherapy physical dose according to the conventional radiotherapy plan parameters and performing inverse optimization of the conventional radiotherapy plan dose;

[0034] A dynamic trajectory intensity-modulated radiotherapy plan design module for calculating the intensity-modulated radiotherapy physical dose according to the dynamic trajectory intensity-modulated radiotherapy plan parameters and performing inverse optimization of the dynamic trajectory intensity-modulated radiotherapy dose.

[0035] As a further limitation of the second aspect of the present invention, the state quantity of the system at time t is set as the dose distribution D t in the patient's body, and the control quantity includes the gantry angle C t , the couch angle G t , the MLC leaf motion trajectory M t and the dose rate u t, Then the state transition equation of the system is expressed as:

[0036] D t+Δt = D t + u t · P(C t , G t , G t ) · Δt

[0037] Among them, P is a dose calculation engine that calculates the dose given to the calculation point at unit intensity according to the gantry angle, couch angle, and MLC leaf position.

[0038] As a further limitation of the second aspect of the present invention, the target conditions related to the tumor target area are defined as T(*), and the dose limit conditions related to the organs at risk are defined as R(*); the requirement for treatment efficiency is defined as E(*); the mechanical limit conditions for the gantry, treatment couch, and MLC are set as U(*). Then, the model optimized using the motion planning method is expressed as:

[0039] [C t+Δt ',G t+Δt ',M t+Δt ',u t+Δt ']=argmin(L(T(D t+Δt ),E(u t+Δt ),R(D t+Δt )))

[0040] s.t.U(C′ t+Δt ,G′ t+Δt ,M′ t+Δt ,u′ t+Δt )<0

[0041] Among them, L(*) is the objective function used in the motion planning algorithm.

[0042] As a further limitation of the second aspect of the present invention, optionally, the future optimal N control point parameters are predicted using nonlinear programming, and the first predicted control point parameter is selected to execute and update the dose distribution. If the dose distribution in the target area reaches the optimization termination condition at this time, the optimization is stopped and the final plan is output. If not, the execution time is increased and the optimization continues. The plan optimization model is expressed as:

[0043] [C t+Δt ',G t+Δt ',M t+Δt ',u t+Δt ']=argmin(NLP(T(D t+Δt ),E(u t+Δt ),R(D t+Δt ),N))

[0044] s.t.U(C′ t+Δt ,G′ t+Δt ,M′ t+Δt ,u′ t+Δt )<0

[0045] Among them, NLP represents the nonlinear programming optimization engine used in the model predictive control algorithm, and N represents the need to predict the future N control points and calculate the objective function.

[0046] As a further limitation of the second aspect of the present invention, optionally, the plan optimization model for optimizing the accelerator parameters using a reinforcement learning network is expressed as:

[0047] [C t+Δt ',G t+Δt ',M t+Δt ',u t+Δt '] = argmax(Q(T(D t+Δt ), E(u t+Δt ), R(D t+Δt )))

[0048] s.t. U(C′ t+Δt , G′ t+Δt , M′ t+Δt , u′ t+Δt ) < 0

[0049] Among them, Q() represents the Q function used in reinforcement learning;

[0050] A limit position constraint layer is added to the reinforcement learning network:

[0051] l8 = [Sigmoid(a·(A - A min )) + Sigmoid(a·(A max - A))]·l7

[0052] Among them, A is a vector, and each element of A represents a machine parameter of the accelerator; A max and A min respectively represent the maximum and minimum values allowed for the action machine parameter A, and Sigmoid is the Sigmoid function;

[0053] According to the maximum change speed of each parameter, set the corresponding action:

[0054]

[0055] Among them represents the maximum change speed of the kth parameter.

[0056] In a third aspect, the present invention provides a non-transitory computer-readable storage medium, which is used to store computer instructions. When the computer instructions are executed by a processor, the dynamic trajectory intensity-modulated radiotherapy plan optimization method described in the first aspect is implemented.

[0057] In a fourth aspect, the present invention provides a computer device, including a memory and a processor. The processor and the memory communicate with each other. The memory stores program instructions executable by the processor, and the processor calls the program instructions to execute the dynamic trajectory intensity-modulated radiotherapy plan optimization method described in the first aspect.

[0058] Fifth aspect, the present invention provides an electronic device, including: a processor, a memory, and a computer program; wherein, the processor is connected to the memory, the computer program is stored in the memory, and when the electronic device runs, the processor executes the computer program stored in the memory, so that the electronic device executes instructions for implementing the dynamic trajectory intensity modulated radiotherapy plan optimization method as described in the first aspect.

[0059] Advantages of the present invention: By using advanced motion planning algorithms in autonomous driving, the optimization speed and quality of dynamic trajectory intensity modulated radiotherapy plans are improved. Dynamic trajectory intensity modulated radiotherapy is not only a current research hotspot, but also has low implementation difficulty and can improve the quality of radiotherapy plans without hardware modification. The provided technical conditions not only give a comprehensive and flexible technical framework and basic methods for dynamic trajectory intensity modulated radiotherapy, but also describe specific implementation details, such as methods for designing dynamic trajectory intensity modulated radiotherapy plans using different motion planning algorithms. The method and system provided by the present invention can fill the gap in dynamic trajectory intensity modulated radiotherapy technology, which is of great significance for the translational application of dynamic trajectory intensity modulated radiotherapy technology and improving the level of cancer treatment.

[0060] The advantages of the additional aspects of the present invention will be more clearly given in the following description part, or understood through the practice of the present invention. BRIEF DESCRIPTION OF THE DRAWINGS

[0061] In order to more clearly illustrate the technical solutions of the embodiments of the present invention, the following will briefly introduce the drawings required for the description of the embodiments. Obviously, the following drawings are only some embodiments of the present invention. For those of ordinary skill in the art, other drawings can be obtained based on these drawings without creative efforts.

[0062] Figure 1 It is a schematic diagram of the system architecture of the dynamic trajectory intensity modulated radiotherapy plan according to the embodiment of the present invention.

[0063] Figure 2 It is a flowchart of the dynamic trajectory intensity modulated radiotherapy plan optimization based on model predictive control according to the embodiment of the present invention.

[0064] Figure 3 It is a flowchart of the dynamic trajectory intensity modulated radiotherapy plan optimization based on reinforcement learning according to the embodiment of the present invention. DETAILED DESCRIPTION OF THE EMBODIMENTS

[0065] The following details the embodiments of the present invention. The examples of the embodiments are shown in the drawings, where the same or similar reference numerals represent the same or similar elements or elements with the same or similar functions throughout. The embodiments described below with reference to the drawings are exemplary and are only used to explain the present invention and should not be construed as a limitation of the present invention.

[0066] Those skilled in the art can understand that, unless otherwise defined, all terms used herein (including technical and scientific terms) have the same meaning as commonly understood by one of ordinary skill in the art to which this invention belongs.

[0067] It should also be understood that terms such as those defined in a general dictionary should be understood as having a meaning consistent with their meaning in the context of the prior art, and will not be interpreted in an idealized or overly formal sense unless defined as here.

[0068] Those skilled in the art can understand that, unless specifically stated otherwise, the singular forms "a", "an", "the", and "said" used herein may also include the plural forms. It should be further understood that the term "comprising" used in the description of the present invention means the presence of the stated features, integers, steps, operations, elements, and / or components, but does not exclude the presence or addition of one or more other features, integers, steps, operations, elements, and / or groups thereof.

[0069] In the description of this specification, the description with reference to terms such as "one embodiment", "some embodiments", "example", "specific example", or "some examples", etc. means that the specific features, structures, materials, or characteristics described in connection with the embodiment or example are included in at least one embodiment or example of the present invention. Moreover, the specific features, structures, materials, or characteristics described may be combined in a suitable manner in any one or more embodiments or examples. Without contradiction, those skilled in the art can combine and combine the different embodiments or examples described in this specification and the features of different embodiments or examples.

[0070] For the convenience of understanding the present invention, the following will further explain the present invention with specific embodiments in conjunction with the accompanying drawings, and the specific embodiments do not constitute a limitation to the embodiments of the present invention.

[0071] Those skilled in the art should understand that the drawings are only schematic diagrams of the embodiments, and the components in the drawings are not necessarily essential for implementing the present invention.

[0072] The present invention provides a method and system for optimizing dynamic trajectory intensity-modulated radiotherapy (DTRT) plans. The present invention transforms the problem of designing DTRT plans into an autonomous driving problem in the field of control, thereby establishing a new type of DTRT plan design and system. Starting from a brand-new perspective, the present invention analogizes the DTRT plan optimization problem to an autonomous driving problem in the field of control, and by means of a motion planning algorithm, relying on its high efficiency and adaptability in complex decision-making and dynamic optimization, designs an innovative automatic optimization method for DTRT plans. This method gives full play to the advantages of the motion planning algorithm in quickly calculating and processing high-dimensional complex optimization problems, not only significantly improving the optimization efficiency and shortening the plan generation time, but also achieving a breakthrough in the global optimization quality. Based on this optimization method, the present invention also provides a DTRT optimization plan system. The method and system of the present invention provide a novel and efficient DTRT optimization solution for clinical applications.

[0073] In the present invention, a method for designing DTRT plans by using an autonomous driving method in the field of control is proposed. This plan design method is a reverse optimization type design method. Due to the large solution space of DTRT plans and the complex coupling relationships between solution parameters, it is difficult for traditional radiotherapy plan optimization algorithms to quickly and effectively find the optimal radiotherapy plan. Therefore, in this embodiment, the DTRT plan is analogized to an autonomous driving problem, as shown in Table 1 below.

[0074] Table 1. Conceptual analogy between DTRT plan optimization and autonomous driving optimization

[0075]

[0076]

[0077] In this way, the plan optimization problem can be transformed into an autonomous driving optimization problem. Since in the autonomous driving problem, more complex and unexpected situations often need to be faced than in plan optimization. And the current domestic autonomous driving technology has been able to handle this situation well. Therefore, after transforming the DTRT plan problem into an autonomous driving problem and then using advanced autonomous driving algorithms to solve it, the optimal solution can be quickly and effectively found.

[0078] Example 1

[0079] In this Example 1, first, a method for optimizing DTRT plans is provided, including: initializing accelerator parameters according to DTRT plan parameters; calculating the radiophysical dose after intensity modulation, and performing inverse optimization of the DTRT dose.

[0080] Among them, the initialization of accelerator parameters is performed, including: setting the initial couch angle and gantry angle according to the principle of adjacent field arrangement. For MLC leaves, the leaf positions after conformal shaping of the target area based on the initial couch angle and gantry angle are used as the initial values.

[0081] In this embodiment, the problem of dynamic trajectory intensity-modulated radiotherapy (IMRT) plan design is transformed into an autonomous driving problem in the control field, thereby establishing a new method and system for dynamic trajectory IMRT plan design.

[0082] In this embodiment, a dynamic trajectory IMRT plan system is proposed. Based on the existing conventional radiotherapy plan system, this plan system adds the function of dynamic trajectory IMRT plan design. A human-computer interaction interface suitable for dynamic trajectory IMRT plan design and a collision prevention function considering the linkage between the accelerator and the treatment couch are designed to ensure treatment safety. Finally, this plan system can output dynamic trajectory IMRT plans or conventional plans in the standard Dicom RT plan data format.

[0083] In this embodiment, the above-mentioned plan system depends on a dynamic trajectory IMRT plan optimization model based on motion planning. In the use of this model, the one-time optimization of all accelerator parameters is transformed into the step-by-step optimization of accelerator parameters according to the treatment time, thereby reducing the optimization difficulty. This model sets the state quantity of the system at time t as the dose distribution D in the patient's body. t , and the control quantities include the couch angle C t , the gantry angle G t , the MLC leaf motion trajectory M t and the dose rate u t . Then the state transition equation of the system can be expressed as:

[0084] D t+Δt = D t + u t ·P(C t , G t , G t )·Δt

[0085] Among them, P is a dose calculation engine, such as Collapsed Cone Convolution (CCC), which calculates the dose given to the calculation point per unit intensity according to the couch angle, gantry angle, and MLC leaf position.

[0086] After that, the target condition related to the tumor target area is defined as T(), the limit condition related to the organs at risk is positioned as R(); the requirement for treatment efficiency is defined as E(); the mechanical limit conditions for the gantry, treatment couch, and MLC are set as U(). Then the model optimized by using the motion planning method can be expressed as the following formula:

[0087] [Ct+Δt ',G t+Δt ',M t+Δt ',u t+Δt '] = argmin(L(T(D t+Δt ), E(u t+Δt ), R(D t+Δt )))

[0088] s.t. U(C′ t+Δt , G′ t+Δt , M′ t+Δt , u′ t+Δt ) < 0

[0089] where L() is the objective function used in the motion planning algorithm. Since D t is a constant value at time t, argmin() here represents C, G, M, and u that can minimize the objective function L calculated by an advanced path planning optimization algorithm. It can be seen that at each moment, all parameters related to the beam trajectory and intensity distribution are optimized, and the local optimal value is taken in the time domain. It should be noted that for the existing state-of-the-art motion planning algorithms, when calculating the actions at time t+Δt, the state quantities after n*Δt are predicted to avoid the influence of the time-domain locality of each step of the solution on the performance of the final solution set.

[0090] The planning system architecture proposed in this embodiment is as Figure 1 . It can be seen that similar to the conventional dose rate planning system, this planning system includes a patient data management module to manage patient basic information, patient image data, patient planning data, etc.; the region of interest (ROI) delineation module allows doctors or physicists to delineate target areas, organs at risk, and other contours manually or automatically. After that, the planning design can be carried out using information such as the patient's planning CT and the delineated regions of interest.

[0091] According to clinical requirements, this planning system can design two types of plans: 1) The conventional plan design module includes three sub-modules: conventional plan parameter input, physical dose calculation, and conventional plan inverse optimization. This module can implement the design of traditional conventional dose IMRT or VMAT radiotherapy plans; 2) The dynamic trajectory intensity-modulated radiotherapy (IMRT) plan design module contains three sub-modules: parameter input, physical dose calculation, and dynamic trajectory IMRT plan inverse optimization. Among them, the parameter input sub-module is mainly used to input initial gantry angles, couch angles, dose optimization parameters, etc. The physical dose calculation module is the same as that in the conventional plan design module. The dynamic trajectory IMRT plan inverse optimization function automatically calculates and generates the optimal dynamic trajectory IMRT plan parameters according to the set plan parameters.

[0092] For the obtained conventional plan and dynamic trajectory intensity-modulated radiotherapy (IMRT) plan, this planning system can transmit them to other systems in the standard Dicom RT format through the plan parameter output module. Using this system, dynamic trajectory IMRT plans can be designed, and conventional VMAT or IMRT plans can also be freely selected to meet various clinical needs.

[0093] In this embodiment, as Figure 1 , Figure 2 shown, two optimization models for dynamic trajectory IMRT plans based on different motion planning algorithms are proposed.

[0094] For the first model, as Figure 1 shown, it is an optimization model for dynamic trajectory IMRT plans based on model predictive control. First, the accelerator parameters are initialized. According to the principle of adjacent field arrangement, the initial couch angle and gantry angle are set. For the MLC leaves, the leaf positions after conforming to the target area based on the initial couch angle and gantry angle are used as the initial values.

[0095] Next, the model predictive control algorithm is used to optimize the dynamic trajectory IMRT plan. This algorithm is one of the most commonly used motion planning algorithms in the field of autonomous driving. This algorithm uses nonlinear programming to predict the optimal N control point parameters in the future and selects the first predicted control point parameter to execute and update the dose distribution. If the dose distribution in the target area reaches the optimization termination condition at this time, the optimization is stopped and the final plan is output. If not satisfied, the execution time is increased and the optimization continues. The plan optimization model can be expressed as:

[0096] [C t+Δt ',G t+Δt ',M t+Δt ',u t+Δt ']=argmin(NLP(T(D t+Δt ),E(u t+Δt ),R(D t+Δt ),N))

[0097] s.t.U(C′ t+Δt ,G′ t+Δt ,M′ t+Δt ,u′ t+Δt )<0

[0098] Here, NLP is the nonlinear programming optimization engine used in the model predictive control algorithm. N represents predicting the future N control points and calculating the objective function. It should be noted that N is a hyperparameter that needs to be set manually. If the value of N is too large, it may lead to too long optimization time; if the value of N is too small, it may lead to a non-optimal solution for the final plan.

[0099] For the second model, as Figure 2As shown in the figure, first, the accelerator parameters are initialized. The initialization method is the same as that of the first method. After that, the reinforcement learning network is used to optimize the accelerator parameters. Then, the planned optimization model can be expressed as:

[0100] [C t+Δt ',G t+Δt ',M t+Δt ',u t+Δt ']=argmax(Q(T(D t+Δt ),E(u t+Δt ),R(D t+Δt )))

[0101] s.t.U(C′ t+Δt ,G′ t+Δt ,M′ t+Δt ,u′ t+Δt )<0

[0102] Among them, Q() represents the Q (Quality) function used in reinforcement learning. Then, for each optimization, the control point parameter adjustment action that can maximize the reinforcement learning Q function is selected, so as to realize the optimization of the dynamic trajectory intensity-modulated radiotherapy plan. In addition, different from the traditional optimization algorithm, in this embodiment, a limit position constraint layer is added to the reinforcement learning network:

[0103] l8=[Sigmoid(a·(A - A min )) + Sigmoid(a·(A max - A))]·l7

[0104] Among them, A is a vector, and each element of A represents a machine parameter of the accelerator. A max and A min respectively represent the maximum and minimum values allowed for the action machine parameter A, and these two values are determined according to the performance of the accelerator. a is a relatively large scaling factor. Sigmoid is the Sigmoid function. Using the above formula, the score of the action that may cause the accelerator parameters to exceed the range can be forced to be 0. Finally, according to the maximum change speed of each parameter, the corresponding action is set:

[0105]

[0106] Among them represents the maximum change speed of the k-th parameter. In this way, the problem of too rapid parameter change can be avoided.

[0107] In this embodiment, based on the above method, a dynamic trajectory intensity-modulated radiotherapy plan optimization system as shown in Figure 3 is realized. This system includes:

[0108] A patient data management module for obtaining basic information of a patient and patient imaging data; a region of interest delineation module for determining radiotherapy regions of interest (i.e., tumor target regions to be radiated and normal tissue organs to be protected); a conventional plan design module for determining clinical prescription dose requirements for the regions of interest and calculating radiophysical doses according to conventional radiotherapy plan parameters and performing inverse optimization of the conventional radiotherapy plan; a dynamic trajectory intensity-modulated radiotherapy plan design module for determining clinical prescription dose requirements for the regions of interest and calculating the intensity-modulated radiophysical doses according to dynamic trajectory intensity-modulated radiotherapy plan parameters and performing inverse optimization of the dynamic trajectory intensity-modulated radiotherapy plan.

[0109] Example 2

[0110] This Example 2 provides a non-transitory computer-readable storage medium for storing computer instructions, which when executed by a processor, implement the dynamic trajectory intensity-modulated radiotherapy plan optimization method as described above. The method includes:

[0111] Obtaining basic information of a patient and patient imaging data; determining radiotherapy regions of interest (i.e., tumor target regions to be radiated and normal tissue organs to be protected) and clinical prescription dose requirements for the regions of interest; calculating the intensity-modulated radiophysical doses according to conventional radiotherapy plan parameters or dynamic trajectory intensity-modulated radiotherapy plan parameters and performing inverse optimization of the dynamic trajectory intensity-modulated radiotherapy dose.

[0112] Example 3

[0113] This Example 3 provides a computer device including a memory and a processor, the processor and the memory communicate with each other, the memory stores program instructions executable by the processor, and the processor calls the program instructions to execute the dynamic trajectory intensity-modulated radiotherapy plan optimization method as described above. The method includes:

[0114] Obtaining basic information of a patient and patient imaging data; determining radiotherapy regions of interest (i.e., tumor target regions to be radiated and normal tissue organs to be protected) and clinical prescription dose requirements for the regions of interest; calculating the intensity-modulated radiophysical doses according to conventional radiotherapy plan parameters or dynamic trajectory intensity-modulated radiotherapy plan parameters and performing inverse optimization of the dynamic trajectory intensity-modulated radiotherapy dose.

[0115] Example 4

[0116] This Example 4 provides an electronic device including: a processor, a memory, and a computer program; wherein, the processor is connected to the memory, the computer program is stored in the memory, and when the electronic device runs, the processor executes the computer program stored in the memory so that the electronic device executes instructions for implementing the dynamic trajectory intensity-modulated radiotherapy plan optimization method as described above. The method includes:

[0117] Obtain the basic information of the patient and the patient's imaging data; determine the radiotherapy target area (i.e., the tumor target area to be radiated and the normal tissue organs to be protected) and the clinical prescription dose requirements for the target area; calculate the intensity-modulated radiotherapy dose according to the conventional radiotherapy plan parameters or the dynamic trajectory intensity-modulated radiotherapy plan parameters, and perform inverse optimization of the dynamic trajectory intensity-modulated radiotherapy dose.

[0118] Those skilled in the art should understand that the embodiments of the present invention can be provided as methods, systems, or computer program products. Therefore, the present invention can take the form of a completely hardware embodiment, a completely software embodiment, or an embodiment combining software and hardware aspects. Moreover, the present invention can take the form of a computer program product implemented on one or more computer-usable storage media (including but not limited to disk storage, CD-ROM, optical storage, etc.) containing computer-usable program code.

[0119] The present invention is described with reference to the flowcharts and / or block diagrams of methods, apparatuses (systems), and computer program products according to the embodiments of the present invention. It should be understood that each process and / or block in the flowchart and / or block diagram can be implemented by computer program instructions, and the combination of the processes and / or blocks in the flowchart and / or block diagram can also be implemented by computer program instructions. These computer program instructions can be provided to the processor of a general-purpose computer, a special-purpose computer, an embedded processor, or other programmable data processing devices to generate a machine, so that the instructions executed by the processor of the computer or other programmable data processing devices generate means for implementing the specified functions in Figure 1 one process or multiple processes and / or blocks Figure 1 one block or multiple blocks.

[0120] These computer program instructions can also be stored in a computer-readable memory that can direct a computer or other programmable data processing device to work in a specific manner, so that the instructions stored in the computer-readable memory generate a manufactured article including instruction means, and the instruction means implement the specified functions in Figure 1 one process or multiple processes and / or blocks Figure 1 one block or multiple blocks.

[0121] These computer program instructions can also be loaded onto a computer or other programmable data processing device, and a series of operation steps are executed on the computer or other programmable device to generate a computer-implemented process, so that the instructions executed on the computer or other programmable device provide steps for implementing the specified functions in Figure 1 one process or multiple processes and / or blocks Figure 1 one block or multiple blocks.

[0122] Although the specific embodiments of the present invention have been described above in conjunction with the accompanying drawings, it is not a limitation on the protection scope of the present invention. Those skilled in the art should understand that, based on the technical solutions disclosed in the present invention, various modifications or deformations that can be made by those skilled in the art without creative efforts should be covered within the protection scope of the present invention.

Claims

1. A method for optimizing a dynamic trajectory intensity-modulated radiotherapy plan, characterized in that, including: Obtain the basic information of the patient and the patient's imaging data; Determine the radiotherapy region of interest and the clinical prescription dose requirements for the region of interest; wherein, the radiotherapy region of interest is the tumor target area to be radiated and the normal tissue organs to be protected According to the conventional radiotherapy plan parameters or the dynamic trajectory intensity modulated radiotherapy plan parameters, calculate the intensity modulated radiotherapy physical dose, and perform dynamic trajectory intensity modulated radiotherapy dose inverse optimization.

2. The dynamic trajectory intensity-modulated radiotherapy plan optimization method according to claim 1, wherein Set the state quantity of the system at time t as the dose distribution D in the patient's body t , and the control quantities include the gantry angle C t , the collimator angle G t , the MLC leaf motion trajectory M t and the dose rate u t, Then the state transition equation of the system is expressed as: D t+Δt = D t + u t · P(C t , G t , G t ) · Δt Wherein, P is the dose calculation engine, which calculates the dose given to the calculation point at unit intensity according to the gantry angle, couch angle, and MLC leaf position.

3. The dynamic trajectory intensity modulated radiotherapy plan optimization method according to claim 2, characterized in that Define the target condition related to the tumor target area as T(*), define the limit condition related to the organ at risk as R(*); define the requirement for treatment efficiency as E(*); set the mechanical limit conditions for the gantry, treatment couch, and MLC as U(*), then the model optimized by the motion planning method is expressed as: [C t+Δt ',G t+Δt ',M t+Δt ',u t+Δt '] = argmin(L(T(D t+Δt ), E(u t+Δt ), R(D t+Δt ))) s.t.U(C′ t+Δt ,G t+Δt ,M′ t+Δt ,u′ t+Δt )<0 Wherein, L(*) is the objective function used by the motion planning algorithm.

4. The dynamic trajectory intensity modulated radiotherapy plan optimization method according to claim 2, wherein Use nonlinear programming to predict the optimal N control point parameters in the future, and select the first predicted control point parameter to execute and update the dose distribution. If the dose distribution in the target area reaches the optimization termination condition at this time, stop the optimization and output the final plan. If not, increase the execution time and continue the optimization. The plan optimization model is expressed as: [C t+Δt ',G t+Δt ',M t+Δt ',u t+Δt '] = argmin(NLP(T(D t+Δt ), E(u t+Δt ), R(D t+Δt ), N)) s.t.U(C′ t+Δt ,G′ t+Δt ,M′ t+Δt ,u′ t+Δt )<0 Wherein, NLP represents the nonlinear programming optimization engine used in the model predictive control algorithm, and N represents the need to predict the future N control points and calculate the objective function.

5. The dynamic trajectory intensity-modulated radiotherapy plan optimization method according to claim 2, wherein The plan optimization model for optimizing the accelerator parameters using the reinforcement learning network is expressed as: [C t+Δt ',G t+Δt ',M t+Δt ',u t+Δt '] = argmax(Q(T(D t+Δt ), E(u t+Δt ), R(D t+Δt ))) S.t.U(C′ t+Δt ,G′ t+Δt ,M′ t+Δt ,u′ t+Δt )<0 Wherein, Q() represents the Q function used in reinforcement learning; Add a limit position constraint layer to the reinforcement learning network: l8 = [Sigmoid(a·(A - A min )) + Sigmoid(a·(A max - A))]·l7 Among them, A is a vector, and each element of A represents a machine parameter of the accelerator; A max and A min respectively represent the maximum value and the minimum value allowed for the action machine parameter A, and Sigmoid is the Sigmoid function; Set the corresponding actions according to the maximum change speed of each parameter: Among them represents the maximum change speed of the k-th parameter.

6. A dynamic trajectory intensity-modulated radiotherapy plan optimization system, characterized in that, including: A patient data management module for obtaining the basic information of the patient and the patient's imaging data; A region of interest delineation module for determining the radiotherapy region of interest and the clinical prescription dose requirements for the region of interest; A conventional radiotherapy plan design module for calculating the radiotherapy physical dose according to the conventional radiotherapy plan parameters and performing conventional radiotherapy plan dose inverse optimization; A dynamic trajectory intensity modulated radiotherapy plan design module for calculating the intensity modulated radiotherapy physical dose according to the dynamic trajectory intensity modulated radiotherapy plan parameters and performing dynamic trajectory intensity modulated radiotherapy dose inverse optimization.

7. The dynamic trajectory intensity-modulated radiotherapy plan optimization system according to claim 6, characterized in that, Set the state quantity of the system at time t as the dose distribution D in the patient's body t , and the control quantity includes the gantry angle C t , the collimator angle G t , the MLC leaf movement trajectory M t and the dose rate u t, Then the state transition equation of the system is expressed as: D t+Δt = D t + u t · P(C t , G t , G t )· Δt Wherein, P is the dose calculation engine, which calculates the dose given to the calculation point at unit intensity according to the gantry angle, couch angle, and MLC leaf position.

8. The dynamic trajectory intensity modulated radiotherapy plan optimization system according to claim 7, characterized in that Define the target condition related to the tumor target area as T(*), define the limit condition related to the organ at risk as R(*); define the requirement for treatment efficiency as E(*); set the mechanical limit conditions for the gantry, treatment couch, and MLC as U(*), then the model optimized by the motion planning method is expressed as: [C t+Δt ',G t+Δt ',M t+Δt ',u t+Δt '] = argmin(L(T(D t+Δt ), E(u t+Δt ), R(D t+Δt ))) s.t.U(C′ t+Δt ,G′ t+Δt ,M′ t+Δt ,u′ t+Δt )<0 Wherein, L(*) is the objective function used by the motion planning algorithm.

9. The dynamic trajectory intensity modulated radiotherapy plan optimization system according to claim 7, wherein Predict the future optimal parameters of N control points using nonlinear programming, and select the first predicted control point parameter to execute and update the dose distribution. If the dose distribution in the target area reaches the optimization termination condition at this time, stop the optimization and output the final plan. If not, increase the execution time and continue the optimization. The plan optimization model is expressed as: [C t+Δt ',G t+Δt ',M t+Δt ',u t+Δt '] = argmin(NLP(T(D t+Δt ), E(u t+Δt ), R(D t+Δt ), N)) s.t.U(C′ t+Δt ,G′ t+Δt ,M′ t+Δt ,u′ t+Δt )<0 Among them, NLP represents the nonlinear programming optimization engine used in the model predictive control algorithm, and N represents the need to predict the future N control points and calculate the objective function.

10. The plan optimization model for optimizing accelerator parameters using a reinforcement learning network in the dynamic trajectory intensity-modulated radiotherapy plan optimization system according to claim 7 is expressed as: [C t+Δt ',G t+Δt ',M t+Δt ',u t+Δt '] = argmax(Q(T(D t+Δt ), E(u t+Δt ), R(D t+Δt ))) s.t.U(C′ t+Δt ,G′ t+Δt ,M′ t+Δt ,u′ t+Δt )<0 Among them, Q() represents the Q function used in reinforcement learning; A limit position constraint layer is added to the reinforcement learning network: l8 = [Sigmoid(a·(A - A min )) + Sigmoid(a·(A max - A))]·l7 where A is a vector, and each element represents a machine parameter of the accelerator; A max and A min respectively represent the maximum and minimum values allowed for the action machine parameter A, and Sigmoid is the Sigmoid function; Set corresponding actions according to the maximum change speed of each parameter: where represents the maximum rate of change with respect to the k-th parameter.

Citation Information

Patent Citations

  • Radiotherapy system and method based on deep learning

    CN107715314A

  • Motion control method and device

    CN108785877A

  • Intensity modulated radiation therapy plan optimizing method based on predicted dosage distribution guidance and application

    CN110124214A

  • Method and device for determining intensity modulated radiation therapy plan evaluation parameter matrix, and medium

    CN114420252A

  • High-dose-rate brachytherapy dose optimization algorithm based on reinforcement learning

    CN115019934A

Cited By

  • Ultrahigh dose rate radiotherapy plan optimization system based on compensator modulation

    CN121266025A