System and method for automatic radiotherapy treatment plan generation using reinforcement learning
By combining reinforcement learning and dose prediction machine learning models, the cost function objective and weights of radiotherapy treatment plans are adjusted, solving the problems of time-consuming and resource-intensive generation of radiotherapy treatment plans in existing technologies, and achieving more efficient generation of personalized treatment plans.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-11-07
- Publication Date
- 2026-05-12
AI Technical Summary
The existing radiotherapy treatment plan generation process is time-consuming and resource-intensive, especially in the trade-off between optimizing target coverage and preserving at-risk organs, resulting in inefficiency in generating personalized plans.
By combining reinforcement learning models with dose prediction machine learning models, a predicted three-dimensional dose distribution is generated by receiving patient treatment attributes. The reinforcement learning agent is used to adjust the objective and weights of the cost function to optimize the radiotherapy treatment plan, reducing the number of iterations and processing resources.
It improves the efficiency and accuracy of radiotherapy treatment plan generation, reduces processing time and resource requirements, and generates personalized radiotherapy treatment plans.
Smart Images

Figure CN122006144A_ABST
Abstract
Description
Technical Field
[0001] This application generally involves using reinforcement learning to generate radiotherapy treatment plans. Background Technology
[0002] Radiation therapy treatment planning (RTTP) is a complex process involving specific guidelines, protocols, and instructions adopted by various medical professionals, such as clinicians and medical device manufacturers. Typically, the identification and application of guidelines to achieve radiation therapy is performed by a complex computer model that receives treatment goals from the treating physician and identifies appropriate attributes for the RTTP. For example, the treating physician might identify the form of treatment (e.g., choosing between volumetric modulated arc therapy (VMAT) or intensity-modulated radiation therapy (IMRT)). The treating physician can then input various goals and objectives to be achieved through the treatment, such as dose targets to be achieved for one or more structures of the patient. The software solution can then use various methods to calculate attributes of the patient's treatment, such as determining beam-limiting device angles and radiation emission properties. In the case of IMRT, the beam delivery direction and number of beams are specific relevant variables that must be determined, while for VMAT, the software solution may need to select the number of arcs and their corresponding start and end angles.
[0003] In personalized radiotherapy planning optimization, the trade-off between achieving and / or evaluating target coverage and OAR (organ atrisk) retention largely depends on how the cost function is constructed. Constructing the cost function may require optimizing situation-specific objectives previously unknown to the planner. Therefore, the planner needs to find the objectives through an iterative process. Consequently, even if all components of the plan generation pipeline are fully automated, generating personalized plans can still be a time-consuming and resource-intensive process. Summary of the Invention
[0004] Computer models can be configured to generate radiotherapy treatment plans using a cost function to determine the radiation dose distribution within a patient's structure. Users can input an initial target for the cost function, and the computer model can iteratively adjust the target to identify or determine the optimal dose distribution for treating the patient. This process can involve significant time and computing resources, depending on how close the initial target is to the optimal dose distribution and / or the number of adjustment iterations the computer model performs to identify the optimal dose distribution.
[0005] Computers implementing the systems and methods described herein can use machine learning and reinforcement learning techniques to improve the efficiency of generating radiotherapy treatment plans. The computer can do this using reinforcement learning models and dose prediction machine learning models. For example, the computer can receive patient treatment attributes (e.g., computed tomography (CT) images, field geometry settings, dose prescriptions, etc.) and use these attributes as input to a dose prediction machine learning model. The computer can execute the dose prediction machine learning model to generate a predicted three-dimensional dose distribution. The computer can use the predicted three-dimensional dose distribution to create a cost function with weighted objectives for different patient structures (e.g., organs, bones, tumors, etc.). The computer can perform or apply optimization algorithms on the cost function containing the objectives to generate a first three-dimensional dose distribution that reduces (e.g., minimizes) the cost function to a first-generation value that achieves a first three-dimensional dose distribution (e.g., uses the first three-dimensional dose distribution to treat the patient).
[0006] The computer can implement a reinforcement learning agent to adjust the target value and / or weights of a cost function to identify the optimal target that can be used to generate an optimal plan for treating a patient. For example, the reinforcement learning agent can determine the difference between a first three-dimensional dose distribution and a predicted three-dimensional dose distribution (e.g., using a distance function to determine the distance). The reinforcement learning agent can adjust the target value and / or weights of the cost function based on the difference to, for example, make the three-dimensional dose distribution of the cost function closer to the predicted three-dimensional dose distribution. The reinforcement learning agent can further adjust the target value and / or weights of the cost function according to one or more rules (e.g., constraint rules), which may correspond to generating a target that improves the predicted three-dimensional dose distribution. The computer can apply or execute the adjusted cost function to generate a second three-dimensional dose distribution with a second generation value that reduces the cost function to achieve a second three-dimensional dose distribution. The computer can compare sequentially determined cost values to determine whether the difference between several consecutive cost values meets (e.g., is less than) a threshold or otherwise converges. In response to determining that the difference does not meet the threshold, the computer can use the reinforcement learning agent to repeat the process until the sequentially generated cost values of the cost function meet the threshold or converge. The computer can use the final target in the patient's radiotherapy treatment plan. Using reinforcement learning models in this way, in conjunction with dose prediction machine learning models, can reduce latency and processing resources by starting closer to the optimal goal (e.g., the optimal goal for an individual whose radiation therapy treatment plan is being generated by a reinforcement learning agent) and requiring fewer adjustments and iterations to the cost function than other methods.
[0007] In some cases, reinforcement learning agents can determine whether other criteria are met before determining that the objective has been finalized or converged. For example, a reinforcement learning agent might determine that the average dose level delivered to a defined risk organ is below a threshold, and / or determine whether a target coverage metric (which may not be explicitly included in the cost function) meets the defined objective. If the reinforcement learning agent determines after several iterations that these metrics have not improved, it can determine that the cost function has converged (e.g., instead of determining that the objective of the cost function has converged, or in addition to determining that the objective of the cost function has converged).
[0008] In one embodiment, a method includes: executing a dose prediction machine learning model by a processor using one or more treatment attributes for a patient to generate a predicted three-dimensional dose distribution for the patient; generating a weighted dose-volume target set of a cost function by the processor based on the predicted three-dimensional dose distribution for the patient; determining a first three-dimensional dose distribution that reduces a first-generation value of the cost function by the processor; determining a difference between the first three-dimensional dose distribution and the predicted three-dimensional dose distribution by the processor; adjusting the weighted dose-volume target set of the cost function by the processor using a reinforcement learning agent based on the difference between the first three-dimensional dose distribution and the predicted three-dimensional dose distribution; determining a second three-dimensional dose distribution that reduces a second-generation value by the processor based on the adjusted weighted dose-volume target set of the cost function; and generating a radiotherapy treatment plan for the patient by the processor based on the adjusted weighted dose-volume target set in response to determining that the difference between the first-generation value and the second-generation value meets a threshold.
[0009] The processor can train a reinforcement learning agent based on training data from multiple different patients. For example, for an individual patient, the method may further include: the processor executing a dose prediction machine learning model using one or more second treatment attributes for a second patient to generate a second predicted three-dimensional dose distribution for the second patient; the processor generating a second weighted dose-volume target set for a second cost function based on the second predicted three-dimensional dose distribution for the second patient; the processor determining a third three-dimensional dose distribution that reduces the third cost value of the second cost function; the processor determining a second difference between the third three-dimensional dose distribution and the second predicted three-dimensional dose distribution; the processor determining a reward value based at least on the difference between the third three-dimensional dose distribution and the second predicted three-dimensional dose distribution; and the processor training a reinforcement learning agent based on the reward value.
[0010] The method may further include: receiving a third or more treatment attributes of a third radiotherapy treatment plan for a third patient by a processor; generating a third weighted dose-volume target set by the processor based on the third or more treatment attributes for the third patient; executing a trained reinforcement learning agent by the processor to adjust the third weighted dose-volume target set; and generating a third radiotherapy treatment plan for the third patient by the processor based on the adjusted third weighted dose-volume target set.
[0011] In some cases, determining the reward value involves the processor comparing the difference between a third three-dimensional dose distribution and a second threshold.
[0012] In some cases, determining the reward value involves: the processor applying a set of criteria to the third three-dimensional dose distribution; and the processor determining the reward value based on the application of a set of criteria to the third three-dimensional dose distribution.
[0013] In some cases, determining the reward value includes: in response to determining that the third three-dimensional dose distribution is within a threshold of the second predicted three-dimensional dose distribution, the processor applies a set of criteria to the third three-dimensional dose distribution; and the processor determines the reward value based on the application of the set of criteria to the third three-dimensional dose distribution.
[0014] In some cases, the weighted dose-volume target set for generating the cost function includes one or more weights assigned by the processor based on a stored weight template, which indicates the weights to be applied to different structures of the patient.
[0015] In some cases, the weighted dose-volume target set that generates the cost function includes one or more weights assigned by the processor based on a stored target ranking list for the radiotherapy treatment plan.
[0016] In some cases, adjusting the weighted dose-volume target set involves the processor using a reinforcement learning agent to insert one or more secondary targets and their corresponding weights into the weighted dose-volume target set, where the one or more secondary targets correspond to different structures of the patient.
[0017] In some cases, each target in the weighted dose-volume target set corresponds to a different structure within the patient and a different reinforcement learning agent among multiple reinforcement learning agents, and the weighted dose-volume target set for adjusting the cost function includes the processor using multiple reinforcement learning agents to adjust the weighted dose-volume target set.
[0018] In one embodiment, a system includes: one or more processors coupled to a memory including instructions that, when executed by the one or more processors, cause the one or more processors to: execute a dose prediction machine learning model using one or more treatment attributes for a patient to generate a predicted three-dimensional dose distribution for the patient; generate a weighted dose-volume target set of a cost function based on the predicted three-dimensional dose distribution for the patient; determine a first three-dimensional dose distribution that reduces the first generation value of the cost function; determine a difference between the first three-dimensional dose distribution and the predicted three-dimensional dose distribution; adjust the weighted dose-volume target set of the cost function using a reinforcement learning agent based on the difference between the first three-dimensional dose distribution and the predicted three-dimensional dose distribution; determine a second three-dimensional dose distribution that reduces the second generation value based on the adjusted weighted dose-volume target set of the cost function; and generate a radiotherapy treatment plan for the patient based on the adjusted weighted dose-volume target set in response to determining that the difference between the first generation value and the second generation value meets a threshold.
[0019] In some cases, the instruction also causes one or more processors to: execute a dose prediction machine learning model using one or more second treatment attributes for a second patient to generate a second predicted three-dimensional dose distribution for the second patient; generate a second weighted dose-volume target set for a second cost function based on the second predicted three-dimensional dose distribution for the second patient; determine a third three-dimensional dose distribution that reduces the third cost value of the second cost function; determine a second difference between the third three-dimensional dose distribution and the second predicted three-dimensional dose distribution; determine a reward value based at least on the difference between the third three-dimensional dose distribution and the second predicted three-dimensional dose distribution; and train a reinforcement learning agent based on the reward value.
[0020] In some cases, the instruction also causes one or more processors to: receive a third or more treatment attributes of a third radiotherapy treatment plan for a third patient; generate a third weighted dose-volume target set based on the third or more treatment attributes for the third patient; execute a trained reinforcement learning agent to adjust the third weighted dose-volume target set; and generate a third radiotherapy treatment plan for the third patient based on the adjusted third weighted dose-volume target set.
[0021] In some cases, this instruction causes one or more processors to determine the reward value by comparing it with a second threshold based on a third three-dimensional dose distribution.
[0022] In some cases, this instruction enables one or more processors to determine the reward value by: applying a set of criteria to a third three-dimensional dose distribution set of criteria; and determining the reward value based on the application of a set of criteria to the third three-dimensional dose distribution.
[0023] In some cases, the instruction causes one or more processors to determine the reward value by: applying a set of criteria to the third three-dimensional dose distribution in response to determining that the third three-dimensional dose distribution is within a threshold of the second predicted three-dimensional dose distribution; and determining the reward value based on the application of the set of criteria to the third three-dimensional dose distribution.
[0024] In some cases, the instruction causes one or more processors to generate a weighted dose-volume target set of the cost function by assigning one or more weights according to a stored weight template, which indicates the weights to be applied to different structures of the patient.
[0025] In some cases, this instruction causes one or more processors to generate a weighted dose-volume target set for the cost function by assigning one or more weights based on a stored target ranking list for the radiotherapy treatment plan.
[0026] In some cases, the instruction causes one or more processors to adjust the weighted dose-volume target set by using a reinforcement learning agent to insert one or more secondary targets and their corresponding weights into the weighted dose-volume target set, where the one or more secondary targets correspond to different structures of the patient.
[0027] In some cases, each target in the weighted dose-volume target set corresponds to a different structure within the patient and a different reinforcement learning agent among multiple reinforcement learning agents, and the instruction causes one or more processors to adjust the weighted dose-volume target set of the cost function by using multiple reinforcement learning agents to adjust the weighted dose-volume target set. Attached Figure Description
[0028] Non-limiting embodiments of the present disclosure are described by way of example with reference to the accompanying drawings, which are schematic and not intended to be drawn to scale. Unless indicated as representing background art, the drawings represent aspects of the present disclosure.
[0029] Figure 1 Components of a reinforcement learning plan generation system according to one embodiment are shown.
[0030] Figure 2 A flowchart illustrating the process performed in a reinforcement learning plan generation system according to one embodiment is shown.
[0031] Figure 3 A sequence diagram of the operation of a reinforcement learning plan generation system according to one embodiment is shown.
[0032] Figure 4 A flowchart illustrating a process for training a reinforcement learning agent according to one embodiment is shown. Detailed Implementation
[0033] Reference will now be made to the illustrative embodiments depicted in the accompanying drawings, which will be described herein in specific language. However, it will be understood that this is not intended to limit the scope of the claims or this disclosure. Variations and further modifications to the inventive features shown herein, as well as additional applications to the principles of the subject matter shown herein, that will occur to those skilled in the art and those with knowledge of this disclosure, will be considered within the scope of the subject matter disclosed herein. Other embodiments and / or other changes may be used without departing from the spirit or scope of this disclosure. The illustrative embodiments described in the detailed description are not intended to limit the subject matter presented.
[0034] Computer models can use a cost function of the objective to determine the radiation dose distribution among different structures of a patient (e.g., organs, bones, etc.) to generate a radiotherapy treatment plan for the patient. To do this, the computer can use an initial objective as a starting point and iteratively adjust the objective using an optimization algorithm until the optimal dose distribution for the patient is determined. The initial objective can initially be input by the user. However, because the initial objective may be entered "blindly" without considering the patient being treated, determining the optimal objective for the radiotherapy treatment plan may require numerous iterations of adjusting the objective and determining the dose distribution based on the adjusted objective. This iterative process can be very time-consuming and resource-intensive.
[0035] For the aforementioned reasons, a computer model is desired that, when generating therapeutic attributes, begins with an initial set of targets that are closer to the expected achievable goals or values. Starting with such an initial set of targets can significantly reduce the number of processing iterations required to generate the optimal targets for a radiotherapy treatment plan. Furthermore, an improved AI modeling / training technique is desired to train the model to generate the optimal dose distribution in a computationally efficient and cost-effective manner with fewer processing iterations, thereby producing timely results.
[0036] Using the methods and systems discussed in this paper, a processor can be trained and used with reinforcement learning models in conjunction with a dose prediction machine learning model to automatically generate radiotherapy treatment plans. The processor can receive a set of treatment attributes for the patient. Treatment attributes may include one or more computed tomography (CT) images, field-of-view geometry settings, dose prescriptions, etc. The processor can use the treatment attributes to execute a dose prediction machine learning model (e.g., a neural network, random forest, support vector machine, etc.) to generate a predicted three-dimensional dose distribution for the patient. The processor can use the predicted three-dimensional dose distribution to generate a cost function comprising a weighted set of objectives. These objectives may each correspond to different structures of the patient. The processor can use optimization algorithms to reduce (e.g., minimize) the cost function of the objectives to generate a first three-dimensional dose distribution for treating the patient.
[0037] The processor can use a reinforcement learning agent to compare a first three-dimensional dose distribution with a predicted three-dimensional dose distribution generated by a dose prediction machine learning model. The reinforcement learning agent can adjust the objectives and / or weights of the cost function based on the comparison to make the first three-dimensional dose distribution closer to the predicted three-dimensional dose distribution (e.g., having values closer to the predicted three-dimensional dose distribution). In some cases, the reinforcement learning agent can be trained to drive an optimization process to obtain a better dose distribution than the predicted three-dimensional dose distribution. For example, during training, the reinforcement learning agent can receive positive rewards for improving protection of at-risk organs from levels indicated in the predicted three-dimensional dose distribution. When using a reinforcement learning agent trained in this manner, the processor can generate an adjusted set of weighted dose-volume objectives for the cost function based on the predicted three-dimensional dose distribution for the patient. The processor can determine a second three-dimensional dose distribution that reduces (e.g., minimizes) the cost value of the adjusted cost function. The reinforcement learning agent can compare the sequentially determined cost values with each other to determine differences. The reinforcement learning agent can compare the differences with a threshold. In response to determining that the difference exceeds the threshold, the reinforcement learning agent can adjust the cost function a second time. The processor can reuse the reinforcement learning model's process to iteratively adjust the objective of the cost function based on the difference between the predicted 3D dose distribution and the 3D dose distribution generated according to the adjusted cost function, until the cost value of the cost function converges within two or more iterations (e.g., without changing above a threshold). In some cases, the reinforcement learning agent can continuously adjust the objective and / or weights to optimize or improve other plan quality metrics (e.g., treatment length, number of treatments, etc.). The reinforcement learning agent can adjust the objective and / or weights until any number of metrics are determined to satisfy the reinforcement learning agent's criteria or internal strategy.
[0038] The processor can use the objective of the final cost function in the patient's radiotherapy treatment plan. By combining predicted dose predictions generated by a dose prediction machine learning model based on the patient's treatment attributes with a reinforcement learning agent, the processor can start from a starting point that is closer to the optimal set of objectives for the patient's radiotherapy treatment plan, thus reaching the optimal set of objectives faster (e.g., with fewer iterations). Therefore, the processor can generate radiotherapy treatment plans with less latency and using fewer processing resources. Furthermore, by using a predicted dose distribution generated based on the individual patient's attributes as a starting point, the processor can generate personalized radiotherapy treatment plans for individual patients.
[0039] Accordingly, processors implementing the systems and methods described herein can use reinforcement learning models and dose prediction machine learning models to automatically generate optimized radiotherapy treatment plans, thereby reducing the number of iterations required to generate and adjust the radiotherapy treatment plan and reducing the latency and processing resources required to do so.
[0040] For example, Figure 1 Components of a reinforcement learning plan generation system 100 according to one embodiment are shown. System 100 may include an analysis server 110a, a system database 110b, a reinforcement learning agent 111, end-user devices 120a-120d (collectively referred to as end-user device 120), a medical device 150, a medical device computer 152, a database 160, and a dose prediction machine learning model 162. Figure 1 The various components depicted may belong to a radiotherapy treatment clinic, where patients can receive radiotherapy treatment, in some cases via one or more radiotherapy machines (e.g., medical device 150).
[0041] System 100 is not limited to the components described herein, and may include additional or other components that are not shown for brevity but are considered to be within the scope of the embodiments described herein.
[0042] The aforementioned components can be interconnected via network 130. Examples of network 130 may include, but are not limited to, private or public local-area networks (LANs), wireless local-area networks (WLANs), metropolitan-area networks (MANs), wide-area networks (WANs), and the Internet. Network 130 may include wired and / or wireless communications according to one or more standards and / or via one or more transmission media. Communication on network 130 may be performed according to various communication protocols, such as Transmission Control Protocol and Internet Protocol (TCP / IP), User Datagram Protocol (UDP), and IEEE communication protocols. In one example, network 130 may include wireless communications according to the Bluetooth specification set or another standard or proprietary wireless communication protocol. In another example, network 130 may also include communications on cellular networks, including, for example, GSM (Global System for Mobile Communication), CDMA (Code Division Multiple Access), or EDGE (Enhanced Data for Global Evolution) networks.
[0043] Analysis server 110a can generate and display an electronic platform configured to allow users to interact with reinforcement learning agent 111 and to receive patient and / or treatment information or attributes, as well as output the execution results of reinforcement learning agent 111 and / or dose prediction machine learning model 162. The electronic platform may include a graphical user interface (GUI) displayed on each of end-user device 120, medical device 150, and / or medical device computer 152. Examples of electronic platforms generated and hosted by analysis server 110a may be web-based applications or websites configured to be displayed on various electronic devices, such as mobile devices, tablets, personal computers, etc.
[0044] The information displayed by the electronic platform may include, for example, input elements to receive data associated with the patient to be treated (e.g., planning goals or targets), and to display predictions generated by the reinforcement learning agent 111 (e.g., text, images, or videos generated in response to input received through the electronic platform). The analysis server 110a can then display the results to medical professionals and / or directly modify one or more operational properties of the medical device 150. In some embodiments, the medical device 150 may be a diagnostic imaging device or a treatment delivery device.
[0045] Analysis server 110a can be any computing device, including processors and non-transitory machine-readable storage, capable of performing the various tasks and processes described herein. Analysis server 110a can employ various processors, such as central processing units (CPUs) and graphics processing units (GPUs). Non-limiting examples of such computing devices may include workstation computers, laptop computers, server computers, etc. While system 100 includes a single analysis server 110a, analysis server 110a can also include any number of computing devices operating in a distributed computing environment, such as a cloud environment.
[0046] End-user device 120 can be any computing device, including a processor and non-transitory machine-readable storage medium, capable of performing the various tasks and processes described herein. Non-limiting examples of end-user device 120 may include workstation computers, laptop computers, tablet computers, and server computers. In operation, various users can use end-user device 120 to access a GUI operatively managed by analytics server 110a. Specifically, end-user device 120 may include clinic computer 120a, clinic server 120b, and medical professional device 120c. Even though these devices are referred to herein as "end-user" devices, they may not always be operated by end-users. For example, clinic server 120b may not be directly used by end-users. However, results stored on clinic server 120b can be used to populate various GUIs accessed by end-users via medical professional device 120c.
[0047] Medical device 150 may be a radiotherapy machine configured to administer radiotherapy treatment to a patient. Medical device 150 may also communicate with medical device computer 152, which is configured to display various GUIs discussed herein. For example, analysis server 110a may display results generated by reinforcement learning agent 111 on the computing device described herein. In a non-limiting example, the GUI may display the cost function and / or the target three-dimensional dose distribution of the radiotherapy treatment plan generated by reinforcement learning agent 111 at medical device computer 152 and / or end-user device 120.
[0048] Dose prediction machine learning model 162 can be stored in database 160. Database 160 and database 110b can be the same or different databases. Dose prediction machine learning model 162 can be or includes a machine learning model (e.g., neural network, support vector machine, random forest, etc.) configured to process the patient's treatment attributes to generate a predicted three-dimensional dose distribution for the patient. In some cases, dose prediction machine learning model 162 can be or includes a linear regression model (e.g., in the form of a dose-volume histogram (DVH) curve) used to predict two-dimensional dose information. For example, treatment attributes can be or include any type of patient attributes such as height, sex, weight, patient treatment options, machine attributes (e.g., gantry movement, gantry position, etc.), treatment goals, attributes about the tumor (e.g., size or shape), images of the patient or tumor, tumor stage, primary site of treatment, endpoint, whether the tumor has spread, body mass index, blood pressure, medical history (e.g., previous medical treatments received by the patient), etc. The dose prediction machine learning model 162 can be trained using supervised, semi-supervised, or unsupervised training methods to generate a predicted three-dimensional dose distribution for the patient based on this treatment attribute.
[0049] For example, analysis server 110a can train a dose prediction machine learning model 162 by inputting a set of treatment attributes. Analysis server 110a can execute dose prediction machine learning model 162 based on the input to generate a predicted three-dimensional dose distribution. Analysis server 110a can compare the predicted three-dimensional dose distribution with an expected or baseline three-dimensional dose distribution (e.g., by using a loss function) to determine differences. Analysis server 110a can use backpropagation techniques to train dose prediction machine learning model 162 based on these differences (e.g., by adjusting internal weights and / or parameters based on the differences). Analysis server 110a can repeat this process any number of times until dose prediction machine learning model 162 is determined to be accurate to a threshold (e.g., an accuracy threshold). In some cases, analysis server 110a can implement a testing and validation phase before deploying dose prediction machine learning model 162. In response to determining that dose prediction machine learning model 162 is accurate to the threshold, analysis server 110a can deploy (e.g., begin use) dose prediction machine learning model 162 to generate a predicted three-dimensional dose distribution.
[0050] The reinforcement learning agent 111 can be stored in the system database 110b. The reinforcement learning agent 111 can be configured or trained to make decisions and take actions within the environment to generate an optimal three-dimensional dose distribution for the patient's radiotherapy treatment plan, such as by maximizing the cumulative reward value. The reinforcement learning agent 111 can be trained to operate on a state frame and / or iteratively adjust the objective and / or weights of the cost function generated from the objective of the radiotherapy treatment plan.
[0051] Analysis server 110a can train reinforcement learning agent 111. To do this, analysis server 110a can identify or compute reward values based on the corresponding differences between a three-dimensional dose distribution that minimizes the cost value of the cost function and a predicted three-dimensional dose distribution generated by dose prediction machine learning model 162, to train reinforcement learning agent 111. Reinforcement learning agent 111 can update its internal policy (e.g., internal weights and / or parameters) and adjust the weights and / or objectives of the cost function based on the corresponding reward values. Through iterative interaction, the agent can be trained using algorithms such as Q-learning or policy gradient to improve the decision-making ability of reinforcement learning agent 111 over time, to generate objectives for radiotherapy treatment plans for different patients more quickly, and to adjust the objectives and / or weights of the cost function with fewer iterations. By utilizing an exploration and development strategy in this way, reinforcement learning agent 111 can efficiently learn to generate optimal objectives for radiotherapy treatment plans with fewer and fewer iterations of adjusting the cost function, where the cost function is initially generated from a predicted three-dimensional dose distribution output by dose prediction machine learning model 162 based on the individualized treatment attributes of the patient undergoing radiotherapy.
[0052] In one example of training the reinforcement learning agent 111, the analysis server 110a can determine a reward value based on the difference between the predicted 3D dose distribution and the 3D dose distribution that reduces the cost function. The analysis server 110a can determine the reward value as proportional to the difference (e.g., a higher difference results in a higher reward value, or a lower difference results in a lower reward value). In one example, the direction of the difference can be considered during training. For example, the goal of a radiotherapy treatment plan might be to achieve a certain level of risk organ preservation (e.g., an average dose of 15 Gy). Accordingly, if the current optimized dose is much higher, the reward value might be negative or penalized. However, if the current optimized dose is much less than 15 Gy, the reward value might be positive or rewarded. The reinforcement learning agent 111 can use this rule to determine the reward value, and / or determine the reward value as a distance (e.g., a weighted distance) metric between the predicted 3D dose distribution and the optimized 3D dose distribution, with different weights for different regions or structures of the body in some cases. The weights can correspond to the level of uncertainty in the dose prediction model when generating the predicted 3D dose distribution.
[0053] During training, the reinforcement learning agent 111 can adjust the objective and / or weights of the cost function to maximize the reward value, such as by making the cost function more likely to correspond to a 3D dose distribution with a reduced (e.g., minimized) cost value that is closer to the predicted 3D dose distribution.
[0054] Analysis server 110a can train reinforcement learning agent 111 to drive radiotherapy treatment plan optimization towards a better dose distribution than the predicted dose distribution generated by dose prediction machine learning model 162. Analysis server 110a can do this, for example, using step-based reward value determination, where in step-based reward value determination, analysis server 110a initially determines the reward based on the difference or distance across iterations between the predicted 3D dose distribution and the optimized 3D dose distribution, until it is determined that such difference or distance is below a threshold. In response to determining that the difference or distance is below the threshold, analysis server 110a can adjust (e.g., increase or decrease) the reward value using one or more rules from a set of criteria (e.g., if a different aspect of the 3D dose distribution is an improvement on the predicted 3D dose distribution). Accordingly, the reinforcement learning agent can initially generate reward values toward generating a 3D dose distribution similar to the predicted 3D dose distribution. Subsequently, if the reinforcement learning agent can generate a more effective 3D dose distribution than the predicted 3D dose distribution, the reinforcement learning agent can generate a potentially smaller but still positive reward value.
[0055] Analysis server 110a can implement reinforcement learning agent 111 after training reinforcement learning agent 111. For example, analysis server 110a can receive a request from end user device 120 to generate a radiotherapy treatment plan for a patient. Analysis server 110a can receive a request with a patient identifier. In some cases, the request may include one or more treatment attributes for the patient. In some cases, the analysis server can use the patient identifier to query database 110b to identify treatment attributes for the patient. Analysis server 110a can receive, identify, or otherwise obtain the patient's treatment attributes in any way and use the treatment attributes to generate a radiotherapy treatment plan (e.g., in response to a request).
[0056] Analysis server 110a can use the patient's treatment attributes to generate targets (including treatment targets for treating the patient). Treatment targets can correspond to radiation doses (e.g., doses) for different structures used to treat the patient. Analysis server 110a can use the patient's treatment attributes (e.g., using the patient's treatment attributes as input) to execute a dose prediction machine learning model. This execution can cause the dose prediction machine learning model 162 to generate a predicted three-dimensional dose distribution for treating the patient. Analysis server 110a can generate targets for treating the patient from the predicted three-dimensional dose distribution. Analysis server 110a can weight different targets in the cost function based on stored weight templates corresponding to different targets (e.g., structures) or based on stored rankings of different targets and / or structures corresponding to the targets. The weighted set of targets can constitute the cost function generated from the predicted three-dimensional dose distribution.
[0057] Analysis server 110a can implement reinforcement learning agent 111 to adjust the cost function. First, analysis server 110a can apply an optimization algorithm (e.g., gradient descent, gradient-based optimization, model-based optimization, etc.) to the cost function containing the objective to generate a three-dimensional dose distribution that minimizes the cost function (e.g., minimizes or reduces the cost or value of the cost function). Then, reinforcement learning agent 111 can adjust the cost function based on the difference between the generated three-dimensional dose distribution and the predicted three-dimensional dose distribution output by a dose distribution machine learning model. For example, analysis server 110a can apply an optimization algorithm to the cost function to generate a three-dimensional dose distribution that reduces (e.g., minimizes) the cost value of the cost function. Reinforcement learning agent 111 can compare the generated three-dimensional dose distribution with the predicted three-dimensional dose distribution generated by a dose prediction machine learning model to determine the difference between the generated and predicted three-dimensional dose distributions. Reinforcement learning agent 111 can adjust the objective and / or weights of the cost function based on strategies learned by reinforcement learning agent 111 (such as by making the cost function more likely to correspond to a three-dimensional dose distribution with a reduced (e.g., minimized) cost that is closer to the predicted three-dimensional dose distribution). Analysis server 110a can apply an optimization algorithm to the cost function to generate another three-dimensional dose distribution that reduces (e.g., minimizes) the cost value of the cost function. Analysis server 110a can repeat this process any number of times, thereby comparing the sequentially generated cost values for each iteration. Analysis server 110a can compare the changes or differences between cost values with a threshold. Analysis server 110a can stop repeating the process in response to determining that the changes or differences between cost values are below the threshold or that the cost values have converged elsewhere.
[0058] In some cases, the reinforcement learning agent 111 may additionally determine whether other criteria are met before determining that the target has been finalized or has converged. For example, the reinforcement learning agent 111 may determine that the average dose level delivered to the defined risk organ is below a threshold, and / or determine whether the target coverage metric meets the defined target. If, after several iterations, the reinforcement learning agent 111 determines that these metrics have not improved, the reinforcement learning agent 111 may determine that the cost function has converged (e.g., instead of determining that the target of the cost function has converged, or in addition to determining that the target of the cost function has converged).
[0059] Analysis server 110a can generate a radiotherapy treatment plan based on the adjusted objective of the cost function. For example, analysis server 110a can identify the objective of the cost function (e.g., determine that the change or difference between cost values is below a threshold or that the cost values have converged) after reinforcement learning agent 111 stops iteratively updating the cost function. Analysis server 110a can insert the objective into a record containing the radiotherapy treatment plan (e.g., the objective of the radiotherapy treatment plan and / or any other treatment attributes). In doing so, analysis server 110a can include the objective of the cost function in the patient's radiotherapy treatment plan. In doing so, the reinforcement learning agent can improve the initial predicted three-dimensional dose distribution.
[0060] Analysis server 110a can present a radiotherapy treatment plan, including the target adjusted by reinforcement learning agent 111, on the user interface of end user device 120 (e.g., the same end user device 120 that requests a radiotherapy treatment plan for a patient). Users accessing end user device 120 (e.g., patients or medical professionals treating patients) can view and / or implement the radiotherapy treatment plan or analysis server, or end user device 120 can use the radiotherapy treatment plan to control (e.g., automate) the treatment of the patient by medical device 150. In some cases, analysis server 110a can send the radiotherapy treatment plan to medical device 150, and the controller of medical device 150 can use the radiotherapy treatment plan to automate the treatment of the patient by medical device 150.
[0061] Now for reference Figure 2 According to one embodiment, method 200 illustrates an operational workflow performed in a reinforcement learning plan generation system. Method 200 may include steps 202-216. However, other embodiments may include additional or alternative steps, or one or more steps may be omitted entirely. Method 200 is described as being performed by a server (such as...) Figure 1 The analysis server described in [the document] can be used to execute this. However, one or more steps of method 200 can be performed by [the server described in the document]. Figure 1 The distributed computing system described herein can operate on any number of computing devices to execute. For example, one or more computing devices can execute locally. Figure 2 Some or all of the steps described in the document.
[0062] Using method 200, the analysis server can implement a combination of a dose prediction machine learning model and a reinforcement learning agent to generate a radiotherapy treatment plan. To do this, the analysis server can use the patient's treatment attributes to execute the dose prediction machine learning model to generate a predicted three-dimensional dose distribution for treating the patient. The analysis server can generate targets (e.g., dose-volume targets) from the predicted three-dimensional dose distributions corresponding to different structures within the patient. The analysis server can assign weights to the targets. The analysis server can use an optimization algorithm (e.g., applying weights to the cost function) on the cost function of the targets to determine a three-dimensional dose distribution that reduces (e.g., minimizes) the cost value of the cost function. The analysis server can use a reinforcement learning agent to compare the three-dimensional dose distribution with the predicted three-dimensional dose distribution to determine the difference between the two distributions. The reinforcement learning agent can adjust one or more targets and / or weights of the cost function based on the difference. The analysis server can apply the cost function (e.g., adjusting the weights of the cost function using any operation of the cost function) to the adjusted targets and iteratively repeat the process until the cost value converges (e.g., until the reinforcement learning agent determines a change or difference between sequence cost values below a threshold). The analysis server can then use the final target in the patient's radiotherapy treatment plan.
[0063] For example, in step 202, the analysis server can execute a dose prediction machine learning model. The analysis server can execute the dose prediction machine learning model using the patient's treatment attributes (e.g., using the patient's treatment attributes as input). For example, treatment attributes can be or include one or more of any type of patient attributes, such as height, sex, weight, patient treatment options, machine attributes (e.g., gantry movement, gantry position, etc.), treatment goals, tumor-related attributes (e.g., size or shape), images of the patient or tumor, tumor stage, primary site of treatment, endpoint, whether the tumor has spread, body mass index, blood pressure, medical history (e.g., previous medical treatments received by the patient), etc. The dose prediction machine learning model can be a machine learning model (e.g., a neural network, random forest, support vector machine, etc.) configured to process treatment attributes to generate three-dimensional dose predictions for treating patients corresponding to the respective treatment attributes.
[0064] The analysis server can take a patient's treatment attributes as input and execute a dose prediction machine learning model to generate a predicted three-dimensional dose distribution. The three-dimensional dose distribution can be or includes a set of voxels representing different sites on the patient's body in a three-dimensional mesh. Each voxel (e.g., a volume element) can correspond to a dose value in the three-dimensional dose distribution. In some cases, the analysis server can respond to a request from a user or another computing device, inputting the patient's treatment attributes, to generate a radiotherapy treatment plan for the patient.
[0065] In step 204, the analysis server can generate a weighted dose-volume target set of cost functions. The analysis server can do this using a predicted 3D dose distribution and structural contours received or identified by the analysis server for the patient. The analysis server can identify regions of interest (ROIs) within the 3D dose distribution. ROIs can be or correspond to structures of the patient, such as organs, tumors, and / or any other anatomical structures (e.g., treatment-related structures, such as organs, tumors, and / or any target structures used for treatment). The analysis server can identify voxels of the predicted 3D dose distribution corresponding to the ROIs and extract or identify dose values from the identified voxels. In doing so, the analysis server can create or generate a list of dose values for each ROI. The analysis server can sort the extracted dose values in ascending order. The analysis server can group these dose values into discrete dose bins. Bin widths can vary, but several examples of bin widths can include increments of 0.1 Gy or 0.5 Gy. The analysis server can determine the volume of the ROI falling within each dose bin. The analysis server can do this by counting the number of voxels in each dose bin and multiplying by the volume of a single voxel.
[0066] The analysis server can use dose intervals to create one or more dose-volume histogram (DVH) curves. For example, for cumulative DVH (cDVH), the analysis server can sum the volumes from the highest dose interval to the lowest dose interval. In doing so, the analysis server can generate the cumulative volume at least received for a given dose. In another example, for differential DVH (dDVH), the analysis server can directly use the volumes calculated in each dose interval. The analysis server can normalize the DVH. The analysis server can normalize the volume data to the total volume of the region of interest (e.g., to obtain a relative volume (typically expressed as a percentage)) or the volume per voxel (e.g., if processing absolute volumes). The analysis server can plot DVH curves. The analysis server can plot the dose (x-axis) against the normalized cumulative volume (y-axis) for cDVH or against the volume per interval for dDVH.
[0067] The analysis server can use DVH curves or otherwise to create targets (including one or more targets for individual patient structures). To do this, the analysis server can identify the clinical purpose for treating the patient. For example, a clinical purpose can identify the percentage of individual structures (e.g., target tumors and / or organs of risk) targeted to receive doses less than or greater than a specified dose. The analysis server can identify clinical purposes from user input and / or from treatment attributes obtained from a database or radiotherapy treatment plan. For each region of interest (e.g., tumors and / or organs of risk), the analysis server can generate targets representing the DVH curve. Targets may include, for example, maximum dose, minimum dose, average dose, and / or volume constraints.
[0068] The analysis server can assign weights to a cost function on a target. The analysis server can assign weights to a target using stored templates. For example, the analysis server can store one or more stored templates, each corresponding to a set of treatment attributes. The set of treatment attributes can include, for example, any permutation or combination of tumor location, tumor size, patient demographics, and / or treatment attributes. Each stored template can include weights (e.g., defined weights) for the individual target that can be used in the cost function used to generate a radiotherapy treatment plan. The analysis server can retrieve templates corresponding to a set of treatment attributes that match (e.g., exactly match or match multiple treatment attributes exceeding a threshold) the treatment attributes obtained for generating a radiotherapy treatment plan for a patient.
[0069] In another example, the analysis server can assign weights to targets based on a ranking list of individual objectives (e.g., clinical purposes). For instance, the analysis server can assign higher weights to targets associated with high-ranking clinical purposes (e.g., those mapped to high-ranking clinical purposes in memory) than to targets associated with lower-ranking clinical purposes. When multiple targets correspond to the same structure or objective, the analysis server can assign the same or identical weights to the targets. The ranking of the targets can be input or correspond to a stored template that corresponds to a set of therapeutic attributes similar to those described above. The targets and their corresponding weights can constitute a cost function.
[0070] The analysis server can use one or more types of objectives in the cost function to generate a radiotherapy treatment plan. For example, the cost function may include a lower-level objective (e.g., lower dose limit), an upper-level objective (e.g., upper dose limit), a cubic objective, an upper linear objective, and / or a generalized equivalent uniform dose (gEUD) objective. The cost function may include weights for each of the one or more objectives. A lower-level objective may include three parameters: DVH graph volume position. V Target dose levelD target and priority p If the DVH line of the structure does not reach the point in the DVH chart ( D target , V If the structure contains a dose, then the structure has a dose. D(x) <D And for volume V All points contributing to the DVH line below. x This allows the cost to be calculated using the square law: Structures with lower-level targets (e.g., all structures) can be considered as radial target structures. If the volume value is 100%, the arc optimization of the lower-level targets can be performed using special types of calculations, such as those for cubic targets described below.
[0071] The upper-level target can have three parameters: DVH chart volume position. V Target dose level D target and priority p If the DVH line of the structure extends beyond a point in the DVH chart ( D target , V If the structure contains a dose, then the structure has a dose. D(x)>D And for volume V The DVH line within contributes to all voxels. x This allows the cost to be calculated using the square law: If the volume value is 0%, the arc optimization of the upper-level target can use a special type of calculation, such as the calculation of the cubic target described below.
[0072] Cubic targets can be used for volume-modulated arc therapy (VMAT) optimization with the same targets as intensity-modulated radiotherapy (IMRT) optimization to identify significant cold and hot spots due to continuous gantry movement and small MU per gantry angle. It can handle all lower-level targets at 100% volume and upper-level targets at 0% volume, allowing for readjustment of dose limits in each iteration. D spike Among them, the dose limit D spike Define what constitutes cold / hot for this objective. The cost is normalized to reflect the parameters. h The initial parameterization remains continuous. Then, the cost function is computed (using the same logic for the lower-level objective, for the upper-level objective):
[0073] The upper linear objective can be one with common priorities. p DVH chart volume position v Target dose level D target The target dose-volume curve can be constructed using splines. D(v) If the DVH line of the structure extends beyond any point on the spline, then the structure has a dose. D(x) > D(v) And for volume v The DVH line within contributes to all voxels. x This makes the cost follow the square law:
[0074] Optimization can support three target types of generalized equivalent uniform dose (gEUD) values: lower gEUD, upper gEUD, and exact (target) gEUD. The achievement of the target follows the formal system described in reference (7). [The text then abruptly shifts to a different topic:] ...with dose... D(x) Use parameters ( a The volume of ) V The gEUD value can be calculated as follows:
[0075] The cost function described in this article is merely an example. When implementing the systems and methods described in this article, any type of cost function, such as the exponential cost function, can be used.
[0076] In step 206, the analysis server can determine a first three-dimensional dose distribution that reduces (e.g., minimizes) the first-generation value of the cost function. The analysis server can do this using a planning optimization algorithm. For example, the analysis server can use an optimization algorithm on a cost function that includes an objective, which the analysis server generates for the patient using a dose prediction machine learning model. When using the cost function to generate the first three-dimensional dose distribution, the cost function can be penalized based on one or more rules. For example, rules can include penalizing doses exceeding a defined maximum dose, penalizing deviations from a defined average dose, and / or penalizing deviations from defined volume constraints at a specific dose level. In some cases, the penalty values for the rules can be derived from the predicted three-dimensional dose distribution. In some cases, the penalty values can be inputs or determined in another manner. The reinforcement learning agent can iteratively adjust the weights and / or parameters of the cost function until a three-dimensional dose distribution that reduces or minimizes the first cost value of the cost function (e.g., corresponding to the target or control point parameters) is generated.
[0077] In step 208, the analysis server can determine the difference between the first three-dimensional dose distribution and the predicted three-dimensional dose distribution. The analysis server can do this using a reinforcement learning agent. The reinforcement learning agent can be trained to... (The sentence is incomplete and requires more context to translate accurately.) Figure 4 The described internal strategy is used to adjust the objectives and / or weights of the cost function. For example, a reinforcement learning agent may compare an individual dose value for a predicted three-dimensional dose distribution with that corresponding to the same structure or space as the patient. The reinforcement learning agent may use a function such as a distance function to determine the difference or distance between the two three-dimensional dose distributions. In some cases, the distance function may be weighted based on the structure corresponding to the dose of the dose distribution. The weights may be predefined, proportional, or additionally based on the dose confidence (e.g., probability or confidence score (e.g., a maximum score of 100)) of the dose prediction machine learning model in the predicted three-dimensional dose distributions across different structures. The reinforcement learning agent may determine adjustments based on the difference or distance (e.g., adjustments to one or more objectives and / or weights of the cost function), and / or determine adjustments proportionally to the difference or distance (e.g., adjustments to one or more objectives and / or weights of the cost function).
[0078] In some cases, reinforcement learning agents may use other factors to determine the weights of the cost function and / or the target, in addition to predicting the difference between the three-dimensional dose distribution and the first three-dimensional dose distribution, or instead of predicting the difference between the three-dimensional dose distribution and the first three-dimensional dose distribution. For example, a reinforcement learning agent may use a set of rules to determine the adjustment. For example, a reinforcement learning agent may determine the adjustment based on whether an individual structure in the first three-dimensional dose distribution receives at least the target dose, is above the target dose, or is within the target dose range. In another example, a reinforcement learning agent may determine the adjustment based on whether structures marked as not receiving any dose (such as risk organs (e.g., risk organs far from the target structure)) receive any dose. A reinforcement learning agent may determine the adjustment based on or additionally on any number of such determinations and / or using the learned or trained policies of the reinforcement learning agent.
[0079] In step 210, the analysis server may adjust the weighted dose-volume target set. The analysis server may use a reinforcement learning agent to perform this operation. The reinforcement learning agent may adjust the weighted dose-value set based on the adjustments determined in step 208. For example, the reinforcement learning agent may adjust the weighted dose-volume target set by adjusting the targets and / or the weights associated with the targets. In cases where weights are assigned according to a ranking system, the reinforcement learning agent may adjust the weights while maintaining the relative order of the target weights.
[0080] In some cases, reinforcement learning agents can apply policies learned by the agent and add one or more structural objectives to the cost function. The reinforcement learning agent can include weights for the added objectives (e.g., by maintaining the ranking of structures, following stored templates, and / or according to any other rules). By adding structures to the cost function, reinforcement learning agents can better control dose levels in certain spatial regions of the patient.
[0081] In some cases, reinforcement learning agents determine adjustments to the targets and / or weights, where the targets and / or weights are proportional, or alternatively based on the difference between the predicted 3D dose distribution and a first 3D dose distribution. For example, a reinforcement learning agent may make large changes to the weighted targets (e.g., by generating new structures, or by making large changes to values when the difference is small). When the difference is large, the reinforcement learning agent may make smaller changes to the weighted targets. By doing so, the reinforcement learning agent can reduce the number of iterations required to identify the set of weighted targets corresponding to the 3D dose distribution used to treat a patient.
[0082] In step 212, the analysis server can determine a second three-dimensional dose distribution. The analysis server can determine the second three-dimensional dose distribution that reduces (e.g., minimizes) the second-generation value of the cost function based on the adjusted weighted dose-volume target set of the cost function. The analysis server can determine the second three-dimensional dose distribution that minimizes the second-generation value in a similar manner to how the analysis server determines the first three-dimensional dose distribution that minimizes the first-generation value in step 206.
[0083] In step 214, the analysis server can determine whether the difference between the first-generation value and the second-generation value meets a threshold (e.g., a cost threshold or a second threshold). The analysis server can do this using a reinforcement learning agent and / or by comparing the difference to a threshold and determining whether the difference is less than the threshold. For example, in response to determining that the difference is less than the threshold, the data processing system can determine that the difference meets the threshold. In response to determining that the difference exceeds the threshold, the reinforcement learning agent can iteratively repeat steps 210-214.
[0084] In some cases, a reinforcement learning agent can determine whether a difference meets a threshold by determining whether the difference between the values of the generations is converging. The reinforcement learning agent can do this based on a single comparison of the difference between the values of the generations with a threshold, or by iteratively repeating steps 210-214 and identifying at least a threshold number of instances (e.g., sequential instances) where the differences between the sequentially determined values of the generations meet the threshold.
[0085] In some cases, when performing steps 206-214, the analysis server may use multiple reinforcement learning agents. For example, the analysis server may store individual reinforcement learning agents, each trained separately to or additionally corresponding to a target tailored to an individual objective or structure. The reinforcement learning agents may be trained together in a manner that allows them to collectively control the overall plan generation process. Reinforcement learning agents corresponding to the same structural group (e.g., risk organs) or target (e.g., plan volume) may collaborate to achieve the common objective of that group, while reinforcement learning agents corresponding to opposing structural groups may compete (e.g., by preserving as many structures as possible in the risk organ group, or by providing better target coverage in the target group) to identify the optimal trade-off. Each reinforcement learning agent may be trained simultaneously with the generation of radiotherapy treatment plan objectives, allowing the agents to generate radiotherapy treatment plan objectives for future patients more accurately and / or more quickly.
[0086] In response to determining that the difference meets a threshold, in step 216, the analysis server can generate a radiotherapy treatment plan for the patient. The analysis server can generate the radiotherapy treatment plan for the patient based on an adjusted weighted dose-volume target set. For example, the analysis server can identify targets in the adjusted weighted dose-volume target set and include the targets in the radiotherapy treatment plan. The analysis server can determine treatment attributes (e.g., gantry positioning settings) and / or any other treatment attributes that lead to the targets and include the determined treatment attributes in the radiotherapy treatment plan. The analysis server can include a three-dimensional dose distribution (e.g., a second three-dimensional dose distribution) that reduces the cost value of the adjusted weighted dose-volume target set in the radiotherapy treatment plan. Targets can be treatment attributes. The analysis server can store the radiotherapy treatment plan in memory, such as in a database containing records (e.g., files, documents, tables, lists, messages, notifications, etc.) with the patient identifier.
[0087] The analysis server can present the radiotherapy treatment plan (e.g., one or more treatment attributes of the radiotherapy treatment plan, including the identified target) on a user interface displayed on the end-user device. In some cases, the analysis server can present the radiotherapy treatment plan on the computer device that initially sent the request for the patient's radiotherapy treatment plan.
[0088] In some embodiments, the analysis server may send one or more treatment attributes of a radiotherapy treatment plan to a radiotherapy machine for use in delivering treatment to a patient. The radiotherapy machine may operate (e.g., automatically) based on one or more treatment attributes to, for example, deliver a dose identified by the goals of the radiotherapy treatment plan.
[0089] Figure 3A sequence diagram is shown illustrating an operational sequence 300 of a trained reinforcement learning plan generation system according to one embodiment. The operations of sequence 300 can be performed by any one or more computing devices. In one example, one or more operations of sequence 300 can be performed by a reference... Figure 1 The analysis server 110a shown and described performs this action. Execution of sequence 300 can enable the analysis server 110a to generate a radiotherapy treatment plan using less processing power and with less latency.
[0090] In sequence 300, analysis server 110a can identify treatment attributes 302 (e.g., CT images, field of view geometry settings, prescribed dose, etc.) of a patient undergoing radiotherapy. Analysis server 110a can identify treatment attributes from memory and / or in a request to generate a radiotherapy treatment plan. Analysis server 110a can input the treatment attributes into a dose prediction machine learning model 304, which is trained to generate predictions of a three-dimensional dose distribution based on the input treatment attributes of an individual patient. Analysis server 110a can execute dose prediction machine learning model 304 based on the input treatment attributes 302, such that dose prediction machine learning model 304 generates a predicted three-dimensional dose distribution 306. Analysis server 110a can use the predicted three-dimensional dose distribution 306 and the patient's structural profile 312 to determine 308 one or more DVH curves 310. Analysis server 110a can transform 314 the DVH curves 310 into one or more targets corresponding to different structures of the patient (e.g., risk organs and / or target organs or tumors).
[0091] Analysis server 110a can formulate or determine a cost function based on an objective. The cost function may include one or more rules, where penalties are applied when a three-dimensional dose distribution generated using the cost function does not follow one or more rules. The cost function may include one or more weights of the objective. Analysis server 110a may execute an optimizer (e.g., an optimization algorithm) to reduce (e.g., minimize) the cost function; for example, this may include using the cost function to generate a new three-dimensional dose distribution 320 and determining the cost value (e.g., first-generation value) of the three-dimensional dose distribution until a minimum (e.g., a global or local minimum) is reached. Analysis server 110a may use the three-dimensional dose distribution 320 and / or the patient's structural profile 312 to determine one or more DVH curves 324. Analysis server 110a may execute one or more reinforcement learning agents 326 to determine the difference between the three-dimensional dose distribution 320 and the predicted three-dimensional dose distribution 306 (e.g., the difference between DVH curve 324 and DVH curve 310). The reinforcement learning agent 326 may adjust the weights and / or objectives of the cost function 328, which may include increasing or decreasing weights or objectives, or adding one or more new objectives to the patient's additional or auxiliary structures, and / or removing one or more weights and / or objectives. The reinforcement learning agent 326 may do so based on its observations (e.g., the determined dose 320 and DVH curve 324, the predicted three-dimensional dose distribution 306 and its associated DVH curve 310). The reinforcement learning agent 326 may decide which action to take (e.g., which objective to modify and by how much) based on its internal policies and rules. The reinforcement learning agent 326 may do so based on differences, such as proportionally to the differences. The reinforcement learning agent 326 may determine second-generation values based on the adjusted weights and / or objectives.
[0092] The reinforcement learning agent 326 can iteratively repeat 316-328 until a convergent set of weighted objectives is determined. For example, the reinforcement learning agent 326 can do this by comparing the first-generation value and the second-generation value to determine the difference in the generation value. The reinforcement learning agent 326 can compare the difference to a threshold. In response to determining that the difference exceeds or otherwise does not meet the threshold, the reinforcement learning agent 326 can determine another three-dimensional dose distribution and DVH curve (including the contour 330 based on any newly added structure (e.g., around hot and / or cold zones in a three-dimensional space representing the patient)). The reinforcement learning agent 326 can determine the difference between the newly determined three-dimensional dose distribution or DVH curve and the predicted three-dimensional dose distribution 306 or DVH curve 310. The reinforcement learning agent 326 can readjust the weighted objectives of the cost function and determine a third-generation value for the two readjusted weighted objectives. The reinforcement learning agent 326 can repeat this process until it is determined that the difference between two consecutively determined generation values is less than a threshold, or at least a limited number of consecutively determined differences are less than a threshold.
[0093] Analysis server 110a can use the final target set to generate a patient's radiotherapy treatment plan. In some cases, the final target set can constitute the patient's radiotherapy treatment plan. The analysis server can store the radiotherapy treatment plan along with the patient identifier in memory, and / or send the radiotherapy treatment plan to a computing device or radiotherapy machine for treatment.
[0094] Figure 4 A flowchart illustrating a process for training a reinforcement learning agent according to one embodiment is shown. Method 400 can be performed to train a reinforcement learning agent, such as reinforcement learning agent 111. Method 400 may include steps 402-412. However, other embodiments may include additional or alternative steps, or one or more steps may be omitted entirely. Method 400 is described as being performed by a server (such as...) Figure 1 The analysis server described in [the document] can be used to execute this. However, one or more steps of method 200 can be performed by [the server described in the document]. Figure 1 The distributed computing system described herein can operate on any number of computing devices to execute. For example, one or more computing devices can execute locally. Figure 2 Some or all of the steps described in the document.
[0095] For example, in step 402, the analysis server can execute a dose prediction machine learning model. The analysis server can execute the dose prediction machine learning model using treatment attributes for the patient (e.g., using the patient's treatment attributes as input). In one example, the analysis server can execute the dose prediction machine learning model using one or more treatment attributes for a patient (e.g., a second patient) (e.g., a second treatment attribute) to generate a predicted three-dimensional dose distribution for the patient (e.g., a second predicted three-dimensional dose distribution). The analysis server can execute the dose prediction machine learning model in the same or similar manner as described with reference to step 202.
[0096] In step 404, the analysis server may generate a weighted dose-volume target set (e.g., a second weighted dose-volume target set) of a cost function (e.g., a second cost function). The analysis server may do this using a predicted three-dimensional dose distribution and structural contours received or identified by the analysis server for the patient. The analysis server may generate the weighted dose-volume target set in the same or similar manner as described with reference to step 204.
[0097] In step 406, the analysis server can determine a three-dimensional dose distribution (e.g., a third-generation dose distribution) that reduces (e.g., minimizes) the cost function's cost value (e.g., a third-generation value). The analysis server can do this using a planning optimization algorithm. For example, the analysis server can use an optimization algorithm on a cost function that includes a target generated for the patient by the analysis server using a dose prediction machine learning model. The analysis server can use an optimization algorithm on a cost function that includes the target to generate a three-dimensional dose distribution that reduces the cost value of the cost function. The analysis server can determine the three-dimensional dose distribution in the same or similar manner as described with reference to step 206.
[0098] In step 408, the analysis server can determine a reward value. The analysis server can determine the reward value based at least on the difference between the three-dimensional dose distribution and the predicted three-dimensional dose distribution. For example, the analysis server can compare the individual dose values of the three-dimensional dose distribution corresponding to the same structure or space of the patient with the predicted three-dimensional dose distribution. The analysis server can use a function (such as a distance function or a weighted distance function) to determine the difference or distance between the two three-dimensional dose distributions. The analysis server can determine the difference in the same or similar manner as the reinforcement learning agent described with reference to step 208. The analysis server can determine the reward value based on the difference and / or proportionally to the difference. For example, the analysis server can determine a higher reward value when the difference is small, or a higher reward value when the difference is large.
[0099] In some cases, besides predicting the difference between three-dimensional dose distributions, or even alternatively, the analysis server may use other factors to determine the reward value. For example, the analysis server may use a set of rules to adjust the reward value. For instance, the analysis server may determine the reward value based on whether an individual structure in the three-dimensional dose distribution receives at least the target dose, is above the target dose, or is within the target dose range. In another example, the analysis server may determine the reward value based on whether structures marked as not receiving any dose (such as risk organs far from the target structure of the patient) receive any dose. Reinforcement learning agents can use any type of rule to determine the reward value.
[0100] In some cases, the analysis server can use a step-based reward value determination. For example, the analysis server can first determine the difference between the predicted 3D dose distribution and the actual 3D dose distribution. The analysis server can compare the difference to a threshold (e.g., a distance threshold or a difference threshold). In response to determining that the difference exceeds or otherwise does not meet the threshold, the analysis server can determine the reward value using only the difference. However, in response to determining that the difference is less than or otherwise meets the threshold, the analysis server can adjust (e.g., increase or decrease) the reward value using one or more rules from a set of criteria. Accordingly, the analysis server can initially generate a reward value toward producing a 3D dose distribution similar to the predicted 3D dose distribution. Subsequently, if the analysis server can generate a more effective 3D dose distribution (e.g., a 3D dose distribution with a lower total dose or a lower dose in a defined or defined region or structure of the patient), the analysis server can generate a potentially smaller but still positive reward value.
[0101] In step 410, the analysis server can adjust the weighted dose-volume target set. The analysis server can do this using a reinforcement learning agent. The reinforcement learning agent can adjust the weighted dose-value set in a manner similar to that described with reference to step 210. For example, the reinforcement learning agent can adjust the weighted dose-volume target set by adjusting the targets and / or the weights associated with the targets. The reinforcement learning agent can adjust the weighted dose-volume target set according to its internal strategy. In the case of assigning weights based on a ranking system, the reinforcement learning agent can adjust these weights while maintaining the relative order of the target weights. The analysis server can use the reinforcement learning agent to adjust the weighted dose-volume target set in any way.
[0102] In step 412, the analysis server can train a reinforcement learning agent. The analysis server can train the reinforcement learning agent based on the reward value determined in step 408. In some cases, the analysis server can perform step 412 before step 410. For example, the analysis server can train the reinforcement learning agent by updating its current policy and / or learning algorithm (such as using Q-learning, policy gradient, and / or a reward-based actor-critic approach). The analysis server can use any of these methods to update the reinforcement learning agent's policy to reinforce the agent by adjusting its weighted objectives (e.g., actions) to bring the weighted objectives of the radiotherapy treatment plan closer to and / or improve the personalized predicted 3D dose distribution for the patient. In cases where the reward is negative or penalized, the analysis server can update the policy to avoid actions that would lead to adjustments to the weighted objectives: weighted objectives that would result in a worse 3D dose distribution; and / or weighted objectives that would lead to a 3D dose distribution further away from the predicted dose distribution. Therefore, the analysis server can train a reinforcement learning agent to more accurately adjust the weighted target of the radiotherapy treatment plan when iteratively adjusting the weighted target, thereby reducing the number of adjustments and / or iterations that may be needed when generating an optimized radiotherapy treatment plan.
[0103] The analysis server can repeat steps 404-412 any number of times, and / or until the analysis server generates a weighted set of objectives that converges the cost function, as described herein. With each iteration, the analysis server can further train the reinforcement learning agent based on the new reward values, thereby further improving the policy of the reinforcement learning agent. The analysis server can perform method 400 on any number of patients to train the reinforcement learning agent, in some cases, training the reinforcement learning agent until the reinforcement learning agent is accurate to a threshold. In some cases, in response to determining that the reinforcement learning agent is accurate to a threshold, the analysis server can implement the reinforcement learning agent (e.g., in the manner described with reference to method 200).
[0104] The analytics server can similarly utilize reinforcement learning agents and dose prediction machine learning models to generate radiotherapy treatment plans for patients over time. In doing so, the analytics server can progressively train the reinforcement learning agent for each radiotherapy treatment plan based on the reward values generated during the iteration process. Therefore, the analytics server can progressively train the reinforcement learning agent to improve the actions taken by the agent when adjusting the weighted target set, thereby reducing the number of adjustment iterations the agent takes when generating the weighted target set before it converges. Thus, the analytics server can progressively train the reinforcement learning agent to require fewer processing resources and operate with less latency when generating radiotherapy treatment plan targets.
[0105] Advantageously, by implementing the systems and methods described herein, a computer can train and use one or more reinforcement learning agents to generate radiotherapy treatment plans with less latency and using fewer processing resources. The computer can do this by using a dose prediction machine learning model that can generate a predicted personalized baseline three-dimensional dose distribution for the patient. For example, during the training phase, the reinforcement learning agent can be trained based on the difference between the three-dimensional dose distribution and the predicted personalized baseline three-dimensional dose distribution, where the three-dimensional dose distribution is generated based on a weighted objective generated by the reinforcement learning model. The reinforcement learning agent can be rewarded based on this difference, allowing it to identify a three-dimensional dose distribution that follows any defined criteria while remaining close to the predicted three-dimensional dose distribution. In doing so, the reinforcement learning agent can be trained to determine the target of the radiotherapy treatment plan using fewer and fewer iterations during the inference phase, which can encourage the agent to generate accurate targets faster and with fewer computational resources.
[0106] The various illustrative logic blocks, modules, circuits, and algorithm steps described in conjunction with the embodiments disclosed herein can be implemented as electronic hardware, computer software, or a combination of both. To clearly illustrate this interchangeability between hardware and software, various illustrative components, blocks, modules, circuits, and steps have been described above in general terms of their functionality. Whether this functionality is implemented as hardware or software depends on the specific application and design constraints on the overall system. Those skilled in the art can implement the described functionality in different ways for each specific application, but such implementation decisions should not be construed as departing from the scope of this disclosure or the claims.
[0107] Implementations in computer software can be implemented in software, firmware, middleware, microcode, hardware description languages, or any combination thereof. Code segments or machine-executable instructions can represent procedures, functions, subroutines, programs, routines, subroutines, modules, software packages, classes, or any combination of instructions, data structures, or program statements. A code segment can be coupled to another code segment or hardware circuit by passing and / or receiving information, data, arguments, parameters, or memory contents. Information, arguments, parameters, data, etc., can be passed, forwarded, or sent via any suitable means, including memory sharing, messaging, token passing, network transmission, etc.
[0108] The actual software code or dedicated control hardware used to implement these systems and methods does not limit the claimed features or this disclosure. Therefore, the operation and behavior of the systems and methods are described without the understanding that the specific software code could be designed to implement the systems and methods based on the description herein.
[0109] When implemented in software, functions may be stored as one or more instructions or code on a non-transitory computer-readable or processor-readable storage medium. The steps of the methods or algorithms disclosed herein may be embodied in a processor-executable software module that may reside on a computer-readable or processor-readable storage medium. Non-transitory computer-readable or processor-readable media include both computer storage media and tangible storage media that facilitate the transfer of a computer program from one place to another. A non-transitory processor-readable storage medium may be any available medium accessible to a computer. By way of example and not limitation, such a non-transitory processor-readable medium may include RAM, ROM, EEPROM, CD-ROM or other optical disc storage, disk storage or other magnetic storage devices, or any other tangible storage medium that can be used to store desired program code in the form of instructions or data structures and is accessible to a computer or processor. Disks and optical discs as used herein include compact discs (CDs), laser discs, optical discs, digital universal discs (DVDs), floppy disks, and Blu-ray discs, wherein disks typically magnetically reproduce data, while optical discs optically reproduce data using lasers. Combinations of the above should also be included within the scope of computer-readable media. In addition, the operation of a method or algorithm may reside as one of the codes and / or instructions or any combination or set thereof on a non-transitory processor-readable medium and / or computer-readable medium into which a computer program product may be incorporated.
[0110] The foregoing description of the disclosed embodiments is provided to enable any person skilled in the art to make or use the embodiments described herein and variations thereof. Various modifications to these embodiments will readily be apparent to those skilled in the art, and the principles defined herein may be applied to other embodiments without departing from the spirit or scope of the subject matter disclosed herein. Therefore, this disclosure is not intended to be limited to the embodiments shown herein, but is accorded the widest scope consistent with the following claims and the principles and novel features disclosed herein.
[0111] While various aspects and embodiments have been disclosed, other aspects and embodiments are contemplated. The various aspects and embodiments disclosed are for illustrative purposes and not for limitation, wherein the true scope and spirit are indicated by the following claims.
Claims
1. A method comprising: The processor uses one or more treatment attributes specific to the patient to execute a dose prediction machine learning model to generate a predicted three-dimensional dose distribution for the patient. The processor generates a weighted dose-volume target set with a cost function based on the predicted three-dimensional dose distribution for the patient; The processor determines a first three-dimensional dose distribution that reduces the first-generation value of the cost function; The processor determines the difference between the first three-dimensional dose distribution and the predicted three-dimensional dose distribution; The processor uses a reinforcement learning agent to adjust the weighted dose-volume target set of the cost function based on the difference between the first three-dimensional dose distribution and the predicted three-dimensional dose distribution; The processor determines a second three-dimensional dose distribution that reduces the value of the second generation based on the adjusted weighted dose-volume target set according to the cost function; as well as In response to determining that the difference between the first-generation value and the second-generation value meets a threshold, the processor generates a radiotherapy treatment plan for the patient based on the adjusted weighted dose-volume target set.
2. The method according to claim 1, further comprising: The processor executes the dose prediction machine learning model using one or more second treatment attributes for the second patient to generate a second predicted three-dimensional dose distribution for the second patient; The processor generates a second weighted dose-volume target set based on the second predicted three-dimensional dose distribution for the second patient, using the second cost function. The processor determines a third three-dimensional dose distribution that reduces the third cost value of the second cost function; The processor determines a second difference between the third three-dimensional dose distribution and the second predicted three-dimensional dose distribution; The processor determines the reward value based at least on the difference between the third three-dimensional dose distribution and the second predicted three-dimensional dose distribution. as well as The processor trains the reinforcement learning agent based on the reward value.
3. The method according to claim 2, further comprising: The processor receives a third or more treatment attributes of a third radiotherapy treatment plan for a third patient; The processor generates a third weighted dose-volume target set based on the third or more treatment attributes for the third patient; The processor executes the trained reinforcement learning agent to adjust the third weighted dose-volume target set; as well as The processor generates a third radiotherapy treatment plan for the third patient based on the adjusted third weighted dose-volume target set.
4. The method of claim 2, wherein determining the reward value comprises: The processor determines the reward value based on a comparison between the third three-dimensional dose distribution and the second threshold.
5. The method of claim 2, wherein determining the reward value comprises: The processor applies a set of criteria to the third three-dimensional dose distribution; as well as The processor determines the reward value based on the application of the third three-dimensional dose distribution to the set of criteria.
6. The method of claim 2, wherein determining the reward value comprises: In response to determining that the third three-dimensional dose distribution is within the threshold of the second predicted three-dimensional dose distribution, the processor applies a set of criteria to the third three-dimensional dose distribution; as well as The processor determines the reward value based on the application of the third three-dimensional dose distribution to the set of criteria.
7. The method of claim 1, wherein generating the weighted dose-volume target set of the cost function comprises: The processor assigns one or more weights based on a stored weight template, the weight template indicating the weights to be applied to different structures of the patient.
8. The method of claim 1, wherein generating the weighted dose-volume target set of the cost function comprises: The processor assigns one or more weights based on a stored target ranking list for the radiotherapy treatment plan.
9. The method of claim 1, wherein adjusting the weighted dose-volume target set comprises: The processor uses the reinforcement learning agent to insert one or more second targets and corresponding weights into the weighted dose-volume target set, the one or more second targets corresponding to different structures of the patient.
10. The method of claim 1, wherein each target in the weighted dose-volume target set corresponds to a different structure within the patient and a different reinforcement learning agent among a plurality of reinforcement learning agents, and The weighted dose-volume target set for adjusting the cost function includes: The processor uses the plurality of reinforcement learning agents to adjust the weighted dose-volume target set.
11. A system comprising: One or more processors, said one or more processors coupled to memory, said memory including instructions, said one or more processors causing said one or more processors to: A dose prediction machine learning model is executed using one or more treatment attributes for the patient to generate a predicted three-dimensional dose distribution for the patient. A weighted dose-volume target set for the cost function is generated based on the predicted three-dimensional dose distribution for the patient. Determine the first three-dimensional dose distribution that reduces the first-generation value of the cost function; Determine the difference between the first three-dimensional dose distribution and the predicted three-dimensional dose distribution; The weighted dose-volume target set of the cost function is adjusted using a reinforcement learning agent based on the difference between the first three-dimensional dose distribution and the predicted three-dimensional dose distribution; The second three-dimensional dose distribution that reduces the value of the second generation is determined based on the weighted dose-volume target set adjusted according to the cost function; as well as In response to determining that the difference between the first-generation value and the second-generation value meets a threshold, a radiotherapy treatment plan for the patient is generated based on the adjusted weighted dose-volume target set.
12. The system of claim 11, wherein the instructions further cause the one or more processors to: The dose prediction machine learning model is executed using one or more second treatment attributes for the second patient to generate a second predicted three-dimensional dose distribution for the second patient; A second weighted dose-volume target set is generated based on the second predicted three-dimensional dose distribution for the second patient; Determine a third three-dimensional dose distribution that reduces the third cost value of the second cost function; Determine the second difference between the third three-dimensional dose distribution and the second predicted three-dimensional dose distribution; The reward value is determined at least based on the difference between the third three-dimensional dose distribution and the second predicted three-dimensional dose distribution; and The reinforcement learning agent is trained based on the reward value.
13. The system of claim 12, wherein the instructions further cause the one or more processors to: Receive a third or more treatment attributes of a third radiotherapy treatment plan for a third patient; A third weighted dose-volume target set is generated based on the third or more treatment attributes for the third patient. The trained reinforcement learning agent is executed to adjust the third weighted dose-volume target set; as well as A third radiotherapy treatment plan for the third patient is generated based on the adjusted third weighted dose-volume target set.
14. The system of claim 12, wherein the instructions cause the one or more processors to determine the reward value by comparing the difference between the third three-dimensional dose distribution and the second threshold.
15. The system of claim 12, wherein the instructions cause the one or more processors to determine the reward value in the following manner: A set of standards is applied to the third three-dimensional dose distribution; and The reward value is determined based on the application of the third three-dimensional dose distribution to the aforementioned set of standards.
16. The system of claim 12, wherein the instructions cause the one or more processors to determine the reward value in the following manner: In response to determining that the third three-dimensional dose distribution is within the threshold of the second predicted three-dimensional dose distribution, a set of criteria is applied to the third three-dimensional dose distribution; and The reward value is determined based on the application of the third three-dimensional dose distribution to the aforementioned set of standards.
17. The system of claim 11, wherein the instructions cause the one or more processors to generate the weighted dose-volume target set of the cost function by assigning one or more weights according to a stored weight template indicating weights to be applied to different structures of the patient.
18. The system of claim 11, wherein the instructions cause the one or more processors to generate the weighted dose-volume target set of the cost function by assigning one or more weights according to a stored target ranking list for the radiotherapy treatment plan.
19. The system of claim 11, wherein the instructions cause the one or more processors to adjust the weighted dose-volume target set by inserting one or more second targets and corresponding weights into the weighted dose-volume target set using the reinforcement learning agent, the one or more second targets corresponding to different structures of the patient.
20. The system of claim 11, wherein each target in the weighted dose-volume target set corresponds to a different structure within the patient and a different reinforcement learning agent among a plurality of reinforcement learning agents, and The instructions wherein the one or more processors adjust the weighted dose-volume target set of the cost function by using the plurality of reinforcement learning agents to adjust the weighted dose-volume target set.