Automatic radiation therapy treatment planning with deep reinforcement learning guided by dose distribution-based reward function
A dual-machine learning system with a neural network optimization engine and policy network addresses scalability and flexibility issues in radiation therapy planning, producing high-quality treatment plans for complex cancer treatments by iteratively adjusting objectives with a dose distribution-based reward function.
Patent Information
- Application Number
- PCT/US2025/045394
- Authority / Receiving Office
- WO · WO
- Patent Type
- Applications
- Current Assignee / Owner
- Priority Date
- 2024-09-09
- Filing Date
- 2025-09-08
- Publication Date
- 2026-03-12
AI Technical Summary
Existing machine learning-based radiation therapy treatment planning systems face limitations in scalability, flexibility, and generalizability, particularly for complex treatment sites with multiple target volumes and organs-at-risk, leading to suboptimal treatment plans.
A dual-machine learning approach utilizing a neural network optimization engine and a policy network with a dose distribution-based reward function to iteratively adjust treatment objectives, enabling continuous action space optimization and improved organ-at-risk sparing while maintaining target tumor coverage.
Generates treatment plans with improved organ-at-risk sparing and equal or superior target tumor coverage, achieving human-level performance in automating radiation therapy planning for complex cancer treatments.
Smart Images

Figure US2025045394_12032026_PF_FP_ABST
Abstract
Description
Via Patent Center Docket No.024636-771WO1 / 2025-054-2 Filing Date: September 8, 2025 Customer No.39564 AUTOMATIC RADIATION THERAPY TREATMENT PLANNING WITH DEEP REINFORCEMENT LEARNING GUIDED BY DOSE DISTRIBUTION‐ BASED REWARD FUNCTION CROSS-REFERENCE TO RELATED APPLICATION
[0001] This application claims the benefit of US Application No. 63 / 692,418, titled “AUTOMATIC RADIATION THERAPY TREATMENT PLANNING WITH DEEP REINFORCEMENT LEARNING GUIDED BY DOSE DISTRIBUTION‐BASED REWARD FUNCTION,” filed September 9, 2024, the contents of which are incorporated herein in its entirety. TECHNICAL FIELD
[0002] The subject matter disclosed herein generally relates to automatic treatment planning for radiation therapy including a planning system, method, and article of manufacture. BACKGROUND
[0003] Proton pencil beam scanning (PBS) treatment planning for cancer treatment, such as head and neck (H&N) cancers and / or other types of cancers, is a time-consuming and experience-demanding task, where a large number of potentially conflicting planning objectives are involved. Deep reinforcement learning (DRL) has been introduced to the planning processes of certain cancer treatments, such as an externally applied intensity-modulated radiation therapy (IMRT) and brachytherapy for prostate, lung, and cervical cancers. Intensity Modulated Radiation Therapy refers to beam radiation therapy that precisely delivers varying intensities of radiation beams (e.g., targeted pencil beams) to a tumor while minimizing damage to surrounding healthy tissue. And, brachytherapy refers to an internal radiation therapy where radioactive sources are placed inside a patient or near a tumor. SUMMARY
[0004] Disclosed herein are systems, methods, and articles of manufacture for automatically planning of radiation therapy for certain cancers, such as head and neck cancers as well as other types of cancer.
[0005] In some embodiments, there is provided a method of automating radiation therapy treatment planning. The method may include receiving at least an objective for a controlVia Patent Center Docket No.024636-771WO1 / 2025-054-2 Filing Date: September 8, 2025 Customer No.39564 structure associated with a plurality of organs-at-risk; optimizing, using a first machine learning model configured as an optimization engine, a treatment plan including the control structure and the objective; selecting, based on one or more scores, at least a first organ‐at‐risk to enable objective adjustment; adjusting, using a second machine learning model using a continuous action space, one or more objective parameters for the selected first organ‐at‐risk; and re- optimizing, using the first machine learning model configured as the optimization engine, the treatment plan including the adjusted one or more objective parameters for the selected first organ‐at‐risk; and outputting the re-optimized treatment plan including one or more treatment parameters.
[0006] In one or more variations, there may also be provided one or more features as disclosed herein including one or more of the following. The control structure and the objective may be received with a set of rules to deconflict objectives. The objective may be initialized to prioritize optimization of dose coverage of at least one target volume included in the control structure. The objective may include a maximal dose, a minimal dose, and / or a mean dose. The control structure may define at least in part a structure associated with a target volume to be radiated. The control structure may include at least one clinical target volume, at least one planning target volume, at least one gap, and / or at least one ring. The first machine learning model may include a neural network, an optimization engine, and / or a limited memory, Broyden–Fletcher–Goldfarb– Shanno (L-BFGS) optimization algorithm. The second machine learning model may include a policy network using the continuous action space continuous action space. The second machine learning model may further include a feature selector, wherein the selecting of the first organ‐at‐risk is performed by the feature selector. The second machine learning model including the policy network policy may be trained using a dose distribution‐based reward function. The continuous action space may be constrained to a real- valued number, and an agent of the policy network uses one or more states as one or more inputs. BRIEF DESCRIPTION OF DRAWINGS
[0007] The accompanying drawings, which are incorporated in and constitute a part of this specification, illustrate non‐limiting embodiments of the present disclosure and together with the description, explain the principles of the disclosed subject matter. These drawings are schematic and are not necessarily drawn to scale. The disclosed subject matter, described inVia Patent Center Docket No.024636-771WO1 / 2025-054-2 Filing Date: September 8, 2025 Customer No.39564 general terms above, will now be further illustrated by the following detailed description with reference to the accompanying drawings,
[0008] FIG. 1A depicts an example of a workflow for a treatment planning system, in accordance with some embodiments;
[0009] FIG. 1B depicts an example of a system block diagram including a first machine learning model and a second machine learning model for a treatment planning system, in accordance with some embodiments;
[0010] FIG. 2 shows structures created for a target volume, in accordance with some embodiments;
[0011] FIG. 3 illustrates an example of a second machine learning model comprising an actor‐critic agent, in accordance with some embodiments;
[0012] FIG. 4 depicts dose distribution-based reward, in accordance with some embodiments;
[0013] FIG. 5A and FIG. 5B depict examples of processed for a treatment planning system, in accordance with some embodiments; and
[0014] FIG. 6 depicts an example of a system, in accordance with some embodiments. DETAILED DESCRIPTION
[0015] In radiation treatment planning, a goal is to balance target volume coverage (e.g., volume of a tumor being radiated) while sparing surrounding, normal tissue, so that example organs‐at‐risk (OARs) are protected while killing the tumor covered by the target volume. The treatment planning system (TPS) may thus solve a patient‐specific optimization that takes into account among other things the target volume coverage and the organs-at-risk. The quality of a generated plan may be dictated at least in part by the performance of the treatment planning systems (as well as other factors). For example, the treatment planning may include creating one or more control structures (also referred to herein as structures or planning structures) and / or one or more treatment objectives (also referred to herein as objectives and objective parameters) according to one or more clinical requirements. The control structure refers to for example the structure associated with the target volume, and the objectives refer to dosage information, such as minimum dosage, maximum dosage, and / or the like.Via Patent Center Docket No.024636-771WO1 / 2025-054-2 Filing Date: September 8, 2025 Customer No.39564
[0016] Using at least the objectives as inputs, the treatment planning system’s machine learning (ML) model may be configured as an optimization engine to generate (given the input) an output including machine deliverable parameters, such as dwell time of brachytherapy, spot monitor unit (MU) values of proton pencil beam scanning, and / or other radiation treatment parameters. Moreover, dose‐volume histograms (DVHs) may be generated from resultant radiation dose distribution(s) on the target volumes and surrounding tissue to evaluate a quality of generated plans. For example, the treatment planning system’s objectives may involve a set of parameters, such as hyper‐parameters including weight factors, dose limit values, volume constraints, and / or the like. If the generated plan does not satisfy one or more clinical requirements (e.g., constraints, criteria, and / or the like), the control structures and / or objectives (e.g., related to the hyper‐parameters) may be adjusted before an optimization is repeated. This process may be repeated until a treatment plan is generated that satisfies some, if not all, of the patient’s clinical requirements.
[0017] Prior ML-based treatment planning systems may be considered to have certain limitations with respect to for example scalability (e.g., ML network size grows linearly with the number of parameters to be adjusted so prior ML-based may only be suitable for simpler treatment sites); flexibility (e.g., predict actions in a discrete space); and / or generalizability (e.g., rely heavily on weighted combinations of clinical metrics when calculating rewards for predicted actions thus failing to make reasonable trade-offs when a diverse number of target volumes, organs-at-risk, and prescription levels are involved). As such, the prior ML-based approaches may only be effective with relatively simple treatment sites, such as cervical cancer, prostate cancer, lung cancer and / or other simple treatments where only one target volume, one prescription level at a predetermined fixed-value, and / or a limited number of organs-at-risk are involved.
[0018] In some embodiments, there are provided systems, methods, and articles of manufacture for treatment planning. In some embodiments, the treatment planning system includes a first ML model and a second ML model.
[0019] In some embodiments, the objectives are fed (as an input) into the first machine learning (ML) model, such as a neural network, an optimization engine, a limited memory Broyden–Fletcher–Goldfarb– Shanno (L-BFGS) optimization algorithm, and / or the like. TheVia Patent Center Docket No.024636-771WO1 / 2025-054-2 Filing Date: September 8, 2025 Customer No.39564 first ML model generates output parameters, such as spot monitor unit (MU) values that quantify the radiation output to be applied (e.g., spot pencil beam of protons delivered to a specific location within the tumor), and / or other types of treatment parameters.
[0020] In some embodiments, a second ML model, such as a policy network, a proximal policy optimization (PPO) network, a decision-making policy network, a transformer- based proximal policy optimization, and / or the like, iteratively adjusts the objectives. In the case of a decision-making policy network for example, it predicts actions in a continuous action space and guides the treatment planning system to refine the objectives using, in accordance with some embodiments, a dose distribution-based reward function. For example, the second ML model may adjust the objectives for the proton pencil beam scanning (PBS) treatment. These adjusted objectives 1220 are provided back to the first ML model as an input to re-optimize and generate another set of output parameters. This process may repeat until a final set of output parameters are determined.
[0021] In some implementations of the disclosed treatment planning system, there may be generated proton head and neck treatment plans that provide improved organ-at-risk sparing, with equal or superior target tumor coverage, when compared with other treatment plan generation technology including human assisted or generated plans.
[0022] Although some of the examples refer to treatment plans for head and neck, the treatment plan may generate treatment plans for other types of cancer treatment sites as well including liver cancer, prostate, lung, brain, breast, other sites, and / or the like.
[0023] In some embodiments, the second ML model may be configured to adjust the objectives (which may be for a specific set of selected organs-at-risk). The second ML model may be trained by, for example, interacting with a treatment planning system that uses a novel dose distribution-based reward function, in accordance with some embodiments. The dose distribution-based reward function guides the second ML model (e.g., the policy network) in navigating the trade-offs among dose distributions (which are indicated by the objectives) of the target volumes and the organs-at-risk (which are indicated by the control structures).
[0024] In accordance with some embodiments, the ML-based treatment planning system automatically generates patient treatment plans for cancers, such as head and neck cancers, liver, and / or other types of cancer, with human-level performance or better.Via Patent Center Docket No.024636-771WO1 / 2025-054-2 Filing Date: September 8, 2025 Customer No.39564
[0025] FIG. 1A depicts an example of a workflow for a treatment planning system 100, in accordance with some embodiments.
[0026] The inputs 1020 to the treatment planning system 100 may include one or more of the following: CT data 1040A (which indicates the 3D representation of the patient’s interior), organs-at-risk (OAR) and target volumes 1040B (e.g., volume and location data for one or more organs-at-risk and one or more target volumes to be radiated or treated), prescription dose data including fractions 1040C (e.g., 2 Gy per fraction for 30 fractions), field targets 1040D (e.g., the volume within which a specific beam delivers radiation), and beam angles 1040E (e.g., the orientations that the radiation enters the patient).
[0027] The control structures and associated objectives are created, at 1060, and may use a set of empirical rules to deconflict among conflicting objectives. Table 1 below lists examples of the empirical rules. For example, an attending physician may draw (or indicate) the CTV, and the PTV is generated by expanding the CTV by several millimeters. The size of this expansion may depend on, for example the how difficult it is to align the patient during treatment. For a brain tumor for example, the expansion from CTV to PTV is typically 2 to 3 mm. For tumors in the head and neck area, the expansion is typically 3 to 5 mm. For tumors in the thorax, the expansion is around 5 mm. For tumors in the abdomen, the expansion is around 5 to 7 mm. For tumors in the pelvis, the expansion is between 7 to 10 mm. By the same token, the ring is generated by expanding the PTV, but instead of being immediately adjacent to the PTV, there is a gap of 2 mm between the PTV and the ring. This gap allows the optimizer to avoid potential conflicts between the target / PTV’s minimum dose objective and the ring’s maximum dose objective.
[0028] If the various CTV’s and the OARs overlap with one another, the rules, such as the rules of Table 1, may be used to attribute the overlapped region to one of the structures, such that the rules in Table 1 may be used to redefine or adjust the overlapped structures such that there is no overlap in the redefined structural contours.
[0029] The first ML model 1080 may, as noted, comprise an optimization engine, a limited memory Broyden–Fletcher–Goldfarb– Shanno (L-BFGS) optimization algorithm, and / or the like. The first ML model may receive the control structures and / or objectives created at 1060Via Patent Center Docket No.024636-771WO1 / 2025-054-2 Filing Date: September 8, 2025 Customer No.39564 as an input, and then generate (e.g., as an output) an initial set of parameters 1100, such as spot MU values and / or other radiation treatment parameter values.
[0030] The objectives (created at 1060 and used as an input to 1080) and the initial set of treatment parameters 1100 may also be processed by the second ML model 1200. The second ML model 1200 may be configured to adjust the objectives until the set of parameters 1100 are considered a final set of parameters that can be used at 1500 by the treatment planning system. In some embodiments, the objectives of a selected set of organs-at-risk are selected for adjustment by the second machine learning model. At 1450, a quantity of iterations may be set, such as 200 iterations. In this example, the loop repeats until the 200 iterations is reached. The repeated iterations between the optimizing by the first ML model (at 1080) and the objectives adjustment by the second ML model (at 1200) repeats to enhance the quality of the optimization. Although the previous example uses 200 iterations, other iteration quantities may be used as well.
[0031] To illustrate by way of an implementation example, the treatment planning system may model a plurality of all organs-at-risk (which may be found in a patient cohort). A goal of the second ML model 1200 (e.g., the policy network) may be to gradually reduce an objective (e.g., OAR dose) with minimal impact on target volume coverages. This may be achieved by deliberately adjusting the planning objective parameters in continuous action spaces. Here, the policy network is trained with a PPO framework under the guidance of the novel dose distribution-based reward function.
[0032] FIG. 1B depicts a system block diagram of the treatment planning system 100, in accordance with some embodiments. The treatment planning system 100 may receive one or more inputs as noted above with respect to 1020 and 1040A-E. These inputs may then be used to create at 1060 structures and objectives as noted above.
[0033] The first ML model 1080 may receive at least the structures and objectives as an input. The first ML model may output one or more parameters 1100, such as spot MU values, and / or other treatment parameter values. Moreover, empirical rules may be used to deconflict some of the objectives.
[0034] The second ML model 1200 receives (as an input) treatment parameters 1100 as well as other data, such as the structures and objectives (e.g., dose limits including a maximaldose (Dmax), a minimal dose (Dmin), and / or a mean dose (Dmean) for the associated targetVia Patent Center Docket No.024636-771WO1 / 2025-054-2 Filing Date: September 8, 2025 Customer No.39564 structure and / or organ(s)-at-risk). The second ML model may be configured to adjust the objectives until the parameters 1100 may be considered a final set of parameters that can be used at 1500 by the treatment planning system. As noted, the second ML model may adjust a selected set of the objectives, such as the objectives associated with certain OARs. For example, the objective for one of the OARs (e.g., spinal cord) may be adjusted from a weighting of 0.8 to 0.852, if the dose to the spinal cord is above the required clinical value. In this example, the higher weighting has the effect of raising (e.g., emphasizing) the importance of the particular objective relative to other objectives, which pushes the optimization to work on that particular objective. The second ML model 1200 is trained to adjust the objectives (e.g., in the continuous action spaces), while the first ML model 1080 is configured to find an optimum set of treatment parameters with minimal impact on target coverages defined or indicated by the objective for the target volume. In operation, the first ML model operates in an inner loop iteratively determining optimal treatment parameters, while the outer loop of the second ML model’s PPO iteratively adjusts the one or more objectives.
[0035] Referring again to the creation of structures and objectives at 1060, the created structure may comprise for example a 3 millimeter (mm) expansion from a clinical target volume (CTV) to form a planning target volume (PTV), a 2 mm expansion from PTV to form a gap, and a 20 mm expansion from the gap to form a ring. The clinical target volume (CTV) represents a tumor volume plus a margin for microscopic disease spread, while the planning target volume (PTV) is a larger volume that includes the CTV and adds margins to account for uncertainties in patient positioning, organ motion, and treatment delivery. The gap represents an area where no specific optimization objective is specified (thus avoiding potential conflicts), and the ring represents _an area where the dose level needs to start to become lower since all regions beyond the ring are further to the tumor. FIG. 2 graphically depicts the CTV 202, the PTV 204, the gap 206, and the ring 208.
[0036] In an implementation example, 32 organs-at-risk are accounted for in a head and neck model of the treatment planning system 100, and certain critical organs-at-risk, such as spinal cord and brainstem are given higher priority than the target volumes to be radiated. In this example, all the target volume structures and organ-at-risk structures are classified into one of the following three priority groups. For example, the Level 1 OAR may correspond to nervesVia Patent Center Docket No.024636-771WO1 / 2025-054-2 Filing Date: September 8, 2025 Customer No.39564 (which in principle) should be protect at a higher priority than the killing of the tumor. This is because permanently harming a patient’s nerves and damaging important sensory aspects, such as vision, fundamental reflective functions (e.g., brainstem and spinal cord), may defeat the purpose of curing the patient. As a result, Level 1 OARs are of higher priority than the tumor, which is of Level 2 priority. The other OARs are categorized as Level 3, which is at a lower priority than the important Level 1 OARs and the tumor. For example, the levels may correspond to: • Level 1: brainstem, brainstem core, spinal cord, chiasm, optic nerves; • Level 2: clinical target volume (CTV), planning target volume (PTV); and • Level 3: ring, bone mandible, brachial plexus, brain, cavity oral, cochlea, ear intermediate ear (e.g., middle ear), esophagus, gland submandibular, gland thyroid, larynx, lips, lungs, muscle constrict, parotid, spinal canal, skin, eye, oropharynx, gland lacrimal, lens, pituitary, lobe temporal, cornea, retina, vocal cords, nasopharynx, carotid
[0037] In operation, the levels thus define the parameters that define the reward functions used for each group. For example, the reward function (used by for example the ML model) for Level 1 OARs would impose the highest penalty for voxels with dose over the limit value.
[0038] To illustrate further by way of an example, the objectives may be created, at106, by setting dose limits, such as a maximal dose (Dmax), a minimal dose (Dmin), and / or amean dose (Dmean) for the associated target structure and / or organ-at-risk structure. For example, target volumes may require dose limits for both Dmax and Dmin, so that target volume doses are constrained within a certain range. The rings and organs-at-risk may require dose limits, such as Dmaxand / or Dmean, so that normal tissues are appropriately spared during radiation treatment. However, there may be conflicts when clinical target volumes (CTVs) and planning target volumes (PTVs) for different radiation prescription levels overlap and / or when CTVs and PTVs intersect with rings and organs-at-risk. These conflicts among objectives may pose challenges to the ML model’s optimization, which may lead to undesirable results. To address this issue, oneVia Patent Center Docket No.024636-771WO1 / 2025-054-2 Filing Date: September 8, 2025 Customer No.39564 or more empirical rules may be defined and used to deconflict the treatment planning structures and associated objectives. Table 1 below depicts an example set of empirical rules.
[0039] Table 1 1. For all structures, portions outside the “external” or “Body” contour will be removed; 2. If CTV (or PTV) intersects with “Skin,” the overlap should be subtracted from CTV (or PTV); 3. If CTViintersects with CTVjand ^^i > ^^j (^^ denotes prescription dose), the overlap should be subtracted from CTVj; 4. If PTVi intersects with CTVj, the overlap should be subtracted from PTVi as CTVs are always of higher priority than PTVs; 5. If PTVi intersects with PTVj and ^^i > ^^j , the overlap should be subtracted from PTVj; 6. If Ring intersects with CTV (or PTV), the overlap should be subtracted from Ring; 7. If Ringiintersects with Ringj(or any OARs), the overlap belongs to both Ringiand Ringj(or OARs) since there is no conflict between their objectives; 8. When CTVi (or PTVi) intersects with OARj, if the priority of OARj is higher, the overlap should be sub- tracted from CTVi (or PTVi); otherwise, the overlap should be subtracted from OARj; 9. Dmaxobjectives are used for structures related to Ring, Bone Mandible, Brachial Plexus, Brain, Cochlea, Spinal Canal, Skin, Eye, Pituitary, Lobe Temporal, Retina, Vocal Cords, Carotid, Brainstem, Brainstem Core, Spinal Cord, Optic Chiasm, and Optic Nerve, and Dmean objectives are used for the rest of the OARs; 10. Dose limits and weights for the Dmin and Dmax objectives of CTVs are not adjustable and they are fixed at 1.02 ⋅ ^^, 1.03 ⋅ ^^^^^, 2.0, and 2.0, respectively. Likewise, these parameters are fixed at 0.97 ⋅ ^^, 1.02 ⋅ ^^^^^, 1.0, and 1.0 for the Dmin and Dmax objectives of PTVs, respectively; and 0.93 ⋅ ^^ and 1.0 for the Dmax objectives of Rings. ^^^^^ represents the maximal prescription dose amongst all target volumes; 11. Dose limits for OARs’ objectives are initialized to relatively larger values in order to minimize their interference with target coverages in the first iteration of the optimization. Specifically, a Unet transformers (UNETR)-like dose prediction model is trained in-house to estimate the dose distribution for each patient and the initial dose limits for each OAR are randomly sampled to be 10% to 30% larger than the UNETR’s estimation. Weights for the OARs’ objectives are initialized to 1.0 and will be iteratively adjusted by our policy network, together with their associated dose limits.Via Patent Center Docket No.024636-771WO1 / 2025-054-2 Filing Date: September 8, 2025 Customer No.39564
[0040] As noted, the first ML model 1080 may use an optimization to determine the set of parameters 1100 (e.g., spot MU values, and / or the like) given the input comprising the structures and objectives created at 1060. For example, the first ML model’s optimization may be in accordance with the following: ,wherein L(X) is the loss function given the argument X to be optimized, and for proton PBS plans, X is the spot MU values. SDmax , SDmin , and SDmean represent sets of planning structures with objectives on maximum dose, minimum dose, and mean dose, respectively. The corresponding weights for these objectives are denoted as wi, wj, and wk, and the corresponding dose limits are denoted as Dij k max, Dmin, and D mean, respectively.
[0041] In an implementation example, the number of voxels in each of the structures i, j, k in the sets SDmax, SDmin, and SDmeanare represented as Ni, Nj, and Nk, and their corresponding dose influence matrices (e.g., objectives) are denoted as Mi, Mj, and Mk, respectively. Moreover, ‖⋅‖^^and ‖⋅‖^^denote standard l2-norms that compute only the positive and the negativeHere for example, the quasi-Newton method (e.g., L- BFGS) is used to solve the optimization problem. L-BFGS is specifically designed to solve problems involving large numbers of variables, so L-BFGS utilizes second-order derivatives (e.g., a Hessian matrix) to capture the curvature information and performs line search to automatically find appropriate step sizes during optimization. L-BFGS may be considered memory usage efficient and may be considered suitable for large-scale optimization.Via Patent Center Docket No.024636-771WO1 / 2025-054-2 Filing Date: September 8, 2025 Customer No.39564
[0042] Rather than use a manual adjustment of for example planning objective parameters, such as wi, wj, wkand Djj k max, Dmin, D mean, the disclosed treatment planning system 100 uses a reinforcement learning algorithm, so the second ML model is configured to use for example proximal policy optimization (PPO) to predict actions (e.g., adjusting the value for wifrom 0.8 to 0.852) in a continuous action space. The continuous action space refers to selecting actions from a range of real-valued numbers or quantities rather than a finite set of discrete choices. The second ML model’s PPO framework functions as an outer loop, optimizing the objectives of the inner loop associated with the first ML model 1080.
[0043] In some of the examples disclosed herein, the first ML model 1080 is configured to use L-BFGS as an optimizer (given its efficient memory usage and automatic step size determination), although other ML models and optimizers (e.g., nonlinear optimization, interior point methods, and / or the like) may be used for the “inner loop.” In terms of objective functions, maximum, minimum, and mean dose objectives, these are examples as other objectives may be used as well. Other types of objective functions, such as equivalent uniform dose (EUD) and dose-volume criteria, and robust optimization may also be used as objectives by the first ML model at the inner loop.
[0044] With respect to the second ML model configured to use for example proximal policy optimization (PPO), PPO is part of a policy gradient family of algorithms, which alternates between sampling data through interactions with a task-specific environment and training a policy using stochastic gradient ascent. With PPO, it optimizes the policy by computing the gradients of expected rewards with respect to policy parameters. Assuming that a PPO agent receives a state st at time step t, the PPO agent then generates an action at accordingto its current policy ^^. Here, ^ is a policy mapping from states to actions, and is used to denotea policy with parameters ^. In response to at, the environment reaches a new state st+1 after applying at to st and returns a scalar reward rt as the feedback for taking action at. This process continues until a terminal state is reached at time step T . Here, the policy used for data sampling is denoted as ^^′ , and the policy to be updated as ^^. In each episode, M trajectories ^i= {ai, ai, … , ai}, i = 1 ⋯ M, are sampled with the policy ^^′ , and t h e PPO updates ^ multiple times usingthe gradient estimator:Via Patent Center Docket No.024636-771WO1 / 2025-054-2 Filing Date: September 8, 2025 Customer No.39564 wherein(sit, ait) represents the advantage of action ait at state sitin the i-th trajectory, compared to behavior of policy ^^′ . The value function V^^′(sit) is the expected accumulative reward of state situnder policy ^^′ , representing the performance of this policy. The policy function (ait| sit) represents the probability of policy ^^ choosing action aitat state sit. The importance sampling ratio is introduced to allow multiple policy updates per sampling rollout collection,training more data efficient. Intuitively, when aitis a better-than-average action, that is, A^^′(sit, ait) > 0, the gradient term A^^′(sit, ait) ∇log (ait| sit) points in the direction of increasing the probability of yielding action ait. Similarly, if aitis a worse-than-average action, that is, A^^′(sit, ait) < 0, the probability of yielding action ait will be decreased. Using ∇f (x) = f (x)∇log f (x), the objective function of PPO corresponding to the gradient estimator in Equation (2) is given by and ^ canascent.
[0046] The second ML model’s proximal policy optimization (PPO) may be configured in accordance with an actor-critic architecture, where an actor plays the role of policy network ^, and a critic estimates value V^(st) for state st when calculating advantage A^(st, at).
[0047] FIG. 3 depicts an example of a transformer-based actor critic agent. As illustrated in FIG. 3, the agent comprises three components: a transformer encoder 305 that takes states 301A as inputs and produces latent features 301B for subsequent modules, an actor sub-network 310Via Patent Center Docket No.024636-771WO1 / 2025-054-2 Filing Date: September 8, 2025 Customer No.39564 that generates a probability distribution for sampling actions, and a critic sub-network 315 used to evaluate the performance of the actor 310. As noted above, the agent is configured to select actions from a range of real quantities to enable a smooth actions (e.g., infinite smooth actions such as adjusting the value for weight wifrom 0.8 to 0.852 or any other continuous value, as opposed to adjusting the value from 0.8 to either + / -10%, + / -5%, or 0% of 0.8, i.e. adjusting it to one of the five allowed discrete values: 0.88, 0.84, 0.8, 0.76 or 0.72.).
[0048] In the treatment planning system 100, states 301A are composed of dose- volume histogram (DVH) features, which are obtained by sampling N points uniformly along the dose axis of DVH curves. Corresponding percent volumes are taken as N-dimensional DVH features, that is, each planning structure’s DVH information is extracted to a vector of length N. Planning structures’ types, importance levels, scores (as noted above) and prescription doses are also considered as part of the states to provide additional information. Moreover, each planning structure is assigned a label representing its type (e.g., 1 for CTV, 2 for PTV, 3 for ring, and 4 for OARs). This label is appended to its DVH features as shown at FIG. 3. The importance levels noted above are also concatenated with the DVH features, along with the planning structure’s score, and prescription dose ^^.
[0049] For patients with for example multiple levels of prescription doses, the organs- at-risk may take the lowest prescription dose among all target volumes as their prescription dose. The type label, importance level, and Rx are normalized by constants, such as 4, 3, and 100, respectively, in order to ensure that they are within the range of [0, 1]. For each planning structure, its feature vector is a concatenation of its DVH, type, importance, score, and ^^. With the four additional quantities added to the DVH feature vector, the dimension of a planning structure’s feature vector is [N + 4]. Assuming M planning structures are accounted for in each plan, the combination of these M feature vectors forms the features tensor of dimension [M, N + 4].
[0050] Moreover, a learnable parameter of shape [M, N + 4] is randomly initialized as the positional encoding, and another learnable parameter (e.g., a global feature of shape [1, N + 4] with all values) is initialized to zero. The positional encoding is added to the features tensor, and the Global feature (as shown at FIG. 3) is attached on the top to gather information from a global perspective. The positional encoding is a learnable parameter initialized randomly withVia Patent Center Docket No.024636-771WO1 / 2025-054-2 Filing Date: September 8, 2025 Customer No.39564 values drawn from for example a truncated normal distribution of zero mean, standard deviation = 1.0, lower boundary = −2.0, and upper boundary = 2.0. The positional encoding tensor with dimension [M, N + 4] is elementwise added to the features tensor. The resultant tensor is then concatenated with the global feature, producing the state tensor of dimension [M + 1, N + 4] as input 301C. The state tensor is encoded by the transformer encoder 305, and the encoded global feature is then forwarded to the actor subnetwork 310 and the critic subnetwork 310 as latent features 301B for further prediction.
[0051] In the example of FIG. 3, the transformer encoder 305 comprises multiple transformer layers, wherein each layer includes a multi-head, self-attention module, and a fully connected feed forward network (FFN), around which residual connection and layer normalization are embedded. Assuming for example that there are K adjustable planning objective parameters, the actor subnetwork 310 may predict a K-dimensional Gaussian distribution with a diagonal covariance matrix for action sampling, so it predicts the mean of the distribution from the input states. The standard deviation of the distribution (e.g., Logstd at FIG. 3) is randomly initialized and learned alongside other network parameters during training.
[0052] In operation for example, the treatment planning system 100 may include a large number of adjustable parameters. For each OAR objective for example, one will have a weight value indicating its relative importance with all other objectives, and a Djmax value setting the highest dose that this OAR can receive. Given that head and neck treatment plans for example may have around 25 planning structures (including both OARs and targets), one would have about 25 of these weight and dose limit pairs for each planning structure, i.e. more than 50 parameters to adjust simultaneously as part of the planning objectives. The action space is therefore of high dimensionality, and a large amount of data must be sampled from the environment in order to fully explore it. Consequently, network training is computationally expensive. In addition, effective guidance is required to find positive samples in such a large exploration space, posing challenges to maintaining data balance. This “curse of dimensionality” is the main obstacle in applying reinforcement learning to many real-world problems.
[0053] To address this dimensionality issue, the dimensionality of the action space may be reduced by adjusting only the objectives (e.g., planning objective parameters) from a subset of planning structures that are dynamically selected by a feature selection module 350 at FIG. 3.Via Patent Center Docket No.024636-771WO1 / 2025-054-2 Filing Date: September 8, 2025 Customer No.39564 The term dynamically refers to the selection being performed for each iteration. The feature selection module may be based on a score. For example, the feature selection module may directly select planning structures (e.g., an OAR structure) with a low score (e.g., lowest scores per Equation 6 for example where OARs with the lowest scores, i.e. the most unfavorable dose statistics, are selected for objective parameter adjustments in the subsequent iterations). For each iteration for example, eight OARs may be dynamically selected by the feature selection module 350 from a total 38 OARs. In this example, the eight dynamically selected OARs can vary for each iteration, and the PPO network may manage 76 adjustable parameters of 38 OARs, with the dimensionality of the action space at 16 (e.g., two objective parameters for each of the eight dynamically selected OARs). Alternatively, or additionally, the search space of actions may be limited to a certain range of for example 0 to about 0.5 and −0.5 to about 0 for weight factors and relative dose limits, respectively. Given the action at sampled from the predicted distribution at time step t, the adjustable planning objective parameters may be updated with pnew = pold ∗ (1 + at), where poldand pneware planning objective parameters before and after adjustment.
[0054] The following provides some examples with respect to training. To train the second ML model’s PPO network, a loss function LPPO may be configured as follows:wherein Lpolicyis the policy loss designed based on Equation (3), Lvaluedenotes the mean square error (MSE) loss used for updating the critic, and Lentropydenotes the entropy loss used to enable stability and exploration. For example, the operation clip (f (⋅), 1 − ^, 1 + ^) limits the output of the function f (⋅) to the interval [1 − ^, 1 + ^], in order to stabilize the training process byVia Patent Center Docket No.024636-771WO1 / 2025-054-2 Filing Date: September 8, 2025 Customer No.39564 removing the incentive for larger differences between the old policy ^^′and the new policy ^^. The H ((⋅|st)) is the entropy of the current policy ^^ at state st. The idea of entropy is derived from soft actor-critic where the policy is incentivized to explore more unpromising avenues, and capture multiple modes of near-optimal behavior by maximizing both the expected returns and the expected entropy. Experiments showed that learning a stochastic policy with entropy regularization can stabilize the training process. Moreover, a generalized advantage estimation (GAE) for advantageA^(st, at)may be used as follows:trace-decay parameter used to control the compromise between bias and variance of the estimation. During training, planning objectives with weights and dose limits (defined initially by the empirical rules or adjusted later by the policy network) are used for first ML model’s L-BFGS optimization. The spot MU values (which result from the optimization) are used to compute the dose distribution, DVH, and then the reward and the new state. Based on the new state, the second ML model’s policy network predicts the new action (e.g., adjusting the value for one of the OAR’s weight wifrom 0.8 to 0.852), which is then used to adjust the weights and the dose limits for the next iteration of first ML model’s L-BFGS optimization. This process iterates T times to generate a sequence of data st, at, rt, st+1with which the second ML model’s PPO network is updated. The updated PPO network is subsequently used to sample new sets of data sequences for training in a next episode.
[0055] Regarding the dose distribution-based reward function, treatment planning may be considered a process of balancing competing priorities between dose-volume requirements of the target volumes and the OARs. In the disclosed treatment planning system 100, a novel reward function may be used. The reward function disclosed herein is from the perspective of global dose distributions. The reward function is the negative of the sum of penalties imposedVia Patent Center Docket No.024636-771WO1 / 2025-054-2 Filing Date: September 8, 2025 Customer No.39564 onto each OAR and target based on how much their dose distributions deviate from design. In some embodiments, the dose-distribution-based reward examines the entire shape of the histogram, and evaluates the score based on that entire shape of the histogram. In past approaches, only one data point (e.g., median dose) is evaluated on the histogram, so the quality of the dose is based on a single data point.
[0056] As shown in FIG. 4, a “no penalty” is set for the interval [l, h] relative to the prescription dose ^^ for each planning structure, where all voxels within that structure receiving doses in the range of [l ⋅ ^^, h ⋅ ^^] would not receive a penalty. The penalty is only incurred for any voxels receiving doses outside that range. In some experiments, the “no penalty” intervals for CTVs and PTVs are set to [1.0, 1.05 ⋅ ^^^^^ / ^^] and [0.95, 1.03 ⋅ ^^^^ / ^^], respectively.
[0057] For plans with multiple levels of prescription doses, each target volume may have its own ^^ value, ^^^^^ denotes the highest of all prescription doses in the plan, and ^^^^^ denotes the lowest.
[0058] For plans with only one level of prescription dose, the above three values are equal. The “no penalty” interval for rings and OARs (except lenses) may be set to for example [0, 0.5] and [0, 10 / ^^^^^], respectively. The rings may be normalized relative to the prescription of the target volume used to generate it, and the OARs may be normalized using ^^^^^. In an example implementation, the 10 Gy value is chosen for all OARs because voxel doses less than 10 Gy are generally insignificant for most OARs. Further stratification of dose levels for each OAR is straightforward with the framework. As an example, the “no penalty” interval for the lenses to be [0, 0] in order to minimize the non-stochastic radiation effect. The actual dose distribution of a planning structure could include voxels receiving doses outside of the range [l ⋅ ^^, h ⋅ ^^], a penalty would consequently be incurred. The voxels for each planning structure are binned by normalized dose d∕^^ with a bin size of 0.002, and the relative voxel number n / N is used for display in FIG. 4, where N is the total number of voxels within the planning structure. Given the dose diof the i-th bin and the number of voxels ni in this bin, the penalty is calculated for each planning structure as,Via Patent Center Docket No.024636-771WO1 / 2025-054-2 Filing Date: September 8, 2025 Customer No.39564dose less than l ⋅ ^^ and the second sum Σj from bins with dose more than h ⋅ ^^. The function ^(x) = cx is used to control the severity of the penalty, and in an example implementation, is set to 20, 10, 0.5, 2, 1 for CTVs, PTVs, rings, level-1 OARs and level-3 OARs, respectively. The score of a structure (e.g., planning structure) may be defined as a negative of its penalty. The OARs and rings may only have a second sum, and ^^^^^ may be used in place of ^^ for OARs and rings in treatment plans with multiple levels of prescriptions. When there is no penalty, the planning structure will receive its highest score of 0. Moreover, the uniformity (uniform) for target structures may be defined as the relative dose difference between its D2% and D98%, as follows:wherein the uniformity is at its largest value of 0 when dose distribution within the target is completely uniform. The treatment plan’s score may be determined by considering some, if not all, of the planning structures as follows:time step t as follows: wherein S1 is the set of CTVs, PTVs, rings, and level-1 OARs, S2 is the set of level-3 OARs, and S3 is the set of CTVs.Via Patent Center Docket No.024636-771WO1 / 2025-054-2 Filing Date: September 8, 2025 Customer No.39564 Additional Examples
[0059] FIG. 5A is another example of a workflow according to some embodiments.
[0060] At 101, a treatment planning system, such as treatment planning system 100, may receive patient‐specific information as input. The input information may be similar to input 1020 and 1040A-E and include computed tomography (CT) scans of a patient, contours of structures (e.g., contours of various OARs such as spinal cord, brainstem, liver, and / or other organs), target volumes (e.g., various CTVs and PTVs), prescription doses (e.g., 54 Gy to CTV1, 60 Gy to CTV2, and 70 Gy to CTV3), fraction numbers and beam angles (e.g., total 30 fractions of treatments, with a treatment planning using 3 beams at gantry angles 330, 60, and 120 degrees), and field targets for each beam (e.g., CTV1_RAO is the target segment that defines the volume on which the treatment beam named RAO can place proton radiation spots for treatment).
[0061] At 102, the workflow may include creating auxiliary control structures and objectives (step 102) for the subsequent optimization engine. Specifically, a bunch of expansions, as shown in FIG. 2, are created for treatment planning. In some embodiments, an expansion from clinical target volumes (CTVs) 202 is created as planning target volumes (PTV) 204, gap 206, and ring 208. Afterwards, target volumes and organs‐at‐risk are grouped according to their importance. In some embodiments, “Spinal Cord”, “Chiasm”, “Optic Nerves”, “Brainstem” and “Brainstem Core” are within group one and have higher priority than target volumes. CTV and PTV are within group two, and the rest of the organs‐at‐risk are within group three. As noted, the Group 1 OARs (e.g., spinal cord, brainstem, optical chiasm, optics nerves, and / or the like) are considered important nerve structures that need to be preserved, often with higher priority (e.g.,. during processing by a ML model) than treating the tumor, in order to preserve the quality of life of the patient.
[0062] Objectives for the optimization engine may be created by setting limit values on the maximal dose (Dmax), minimal dose (Dmin) and mean dose (Dmean) of the associated structures. Specifically, Dmaxand Dminare used for target volumes so that their doses are constrained within a proper range. Dmean and Dmax are used for organs‐at‐risk so that doses delivered to normal tissues are appropriately suppressed. There will be conflicts when target volumes of different prescription levels overlap with one another, and target volumes intersectVia Patent Center Docket No.024636-771WO1 / 2025-054-2 Filing Date: September 8, 2025 Customer No.39564 with organs‐at‐risk. In some embodiments, creating control structures for the non‐conflict objectives (step 102) may follow one or more empirical rules as noted above with respect to Table 1.
[0063] Next, one or more objectives may be initialized to prioritize dose coverage of target volumes. In a non‐limiting example of head and neck cancers, this may be implemented by setting the dose limit values for Dminand Dmaxobjectives of CTVs to large values like 1.02Rx and 1.03Rx, respectively. Likewise, the dose limit values for Dminand Dmaxobjectives of PTVs are set to 0.97Rx and 1.02Rx, respectively; and the values for the Dmax objectives of rings are 0.93Rx. Here, Rx is the prescription dose of current target volumes. Meanwhile, dose limit values for Dmaxand Dmeanobjectives of organs‐at‐risk are estimated using a dose prediction model and relaxed by increasing the estimated values by a randomly sampled percentage ranging from 10% to 30%. The dose prediction model may be UNet‐like (e.g., a convolutional neural network structure). The weight factors for CTVs’ objectives are set to a relatively large value of 2.0, and for the other objectives the weight factors are initialized to 1.0. Notably, objective parameters for CTVs, PTVs and rings are fixed. By contrast, those for organs‐at‐risk are adjustable. Although some of the examples refer to the objectives and / or structures being created by the treatment planning system 100 for example, the objectives and / or structures may be created by another device, in which case the treatment planning system 100 would received the objectives and / or structures.
[0064] At 103, a ML model configured as an optimization engine may be used. For example, the first ML model 1080 (which may be configured as a Limited‐memory Broyden– Fletcher–Goldfarb– Shanno (L‐BFGS) or other type of ML model) may then generate one or more plans (which provide or include treatment parameters 1100) by taking the above-noted initial objectives and / or structures as inputs.
[0065] After obtaining the optimization results, the treatment planning system 100 may generate, at 104, dose‐volume histogram (DVH) to evaluate the quality of a generated plan. If the plan does not satisfy one or more clinical criteria (105), the objective parameters including weight factors and dose limit values may be iteratively adjusted by a policy network, such as the second ML model including a policy network as noted above (e.g., at 106 and 107).Via Patent Center Docket No.024636-771WO1 / 2025-054-2 Filing Date: September 8, 2025 Customer No.39564
[0066] As illustrated in FIG. 3, the second ML model’s agent (which comprises the transformer encoder 305 including the feature selector 350, feature encoding module 305, the actor sub-network 310, and the critic sub-network 315), wherein the actor plays the role of policy predictor. In some embodiments, feature encoder described above can be convolution network or transformer network that takes state as inputs. The state is DVH features extracted by a treatment planning environment from DVHs mentioned above. The DVH features are obtained by sampling N points uniformly along the dose axis on DVH curves and taking the volume values corresponding to these points as N‐dimensional features. Types, importance levels, scores and prescription doses of control structures are also considered as part of the states to provide additional information. Therefore, assuming that there are M control structures in total for example, the resultant state will be with a shape of [M, N+4]. In addition, a global feature is attached on the top of the state to gather information from a global view, and the encoded global feature will be the latent features for the actor module and the critic module.
[0067] In complex treatment sites like head and neck cancers proton PBS, a large number of adjustable objective parameters are involved, meaning that the action space to be explored is of high dimensionality. For instance, given ^ adjustable parameters, the dimension of action space is also ^, and in some embodiments using non‐deterministic deep reinforcement learning framework, the actor needs to predict a ^‐dimensional Gaussian distribution for the action sampling. Moreover, a feature selection module may be integrated into the agent to dynamically select a portion (e.g., some) of organs‐at‐risk (at 106 of FIG. 5A). Only objective parameters for the selected organs‐at‐risk may be adjusted in each iteration (at 107 of FIG. 5A). Thus, the dimension of action space can be effectively reduced.
[0068] In some embodiments, the feature selection module 350 directly selects organs‐ at‐risk with lowest scores. In another embodiment, organs‐at‐risk with more space to improve are selected. In addition, the action space searching can be constrained into a small range to reduce the chance of sampling negative samples. In some embodiments, the action space searching range for dose limit values and weight factors are constrained to 0-0.5 and ‐0.5-0, respectively. For example, one can limit the adjustment of the value for weight wifrom its original value of 0.8 to any value within the range of 7.5 to 8.5, say 0.784. Given the action ^t predicted by the actor at the time step ^ and the old objective parameters ^old, new objective parameters can beVia Patent Center Docket No.024636-771WO1 / 2025-054-2 Filing Date: September 8, 2025 Customer No.39564 updated with Pnew= Pold* (1 + ^t) (step 107). Afterwards, objectives are updated (at 108 of FIG. 5A) and plan optimization is conducted again with aforementioned optimization engine (at 109 of FIG. 5A).
[0069] FIG. 5B is another example of a process according to some embodiments.
[0070] At 550, the process may include receiving at least an objective and a one control structure associated with a plurality of organs-at-risk, in accordance with some embodiments. For example, the treatment plan system 100 may receive objectives, such as dose limits (e.g., amaximal dose (Dmax), a minimal dose (Dmin), a mean dose (Dmean), and / or other objectives forthe radiation treatment of a target volume). These objectives may be associated with the structure. In other words, each set of objective parameters is associated with a control structure. The structure may include the target volume (e.g., the cancer to be treated) as well as one or more organs-at- risk, such as the location of the spinal cord, brainstem, optical chiasm, optics nerves, and / or the like, which may be considered important nerve structures that need to be preserved when treating the target volume, such as the tumor. Alternatively, or additionally, the treatment planning system 100 may create the at least one objective and at least one control structure associated with a plurality of organs-at-risk. At 550, a plurality of objectives and corresponding control structures may be received as well.
[0071] At 555, the process may include optimizing, using a first machine learning model configured as an optimization engine, a treatment plan including the control structure and the corresponding objective, in accordance with some embodiments. For example, the first ML model may be configured as an optimization engine as noted above at 1080. When this is the case, the first ML model 1080 may receive the at least one objective and at least one control structure associated with a plurality of organs-at-risk and output treatment parameter values 1100. At 555, a plurality of objectives and corresponding control structures may optimized as well.
[0072] At 560, the process may include selecting, based on one or more scores, at least a first organ‐at‐risk to enable objective adjustment, in accordance with some example embodiments. As noted, the feature selection module 350 may select a structure, such as an OAR structure. The structure may be selected based on a score. As noted, an OAR with (“lowest score” instead of “unfavorable dose statistics”) (as determined from a dose histogramVia Patent Center Docket No.024636-771WO1 / 2025-054-2 Filing Date: September 8, 2025 Customer No.39564 indicating too much radiation dosage on the OAR) may be selected at 560 for adjustment. For example, the dose values of each OAR may each be plotted into a histogram as shown in FIG. 4, and these histograms may be used to compute each OAR’s score. Based on their computed score, the one with the lowest score will be selected for adjustment. At 560, a plurality of OARs may be selected based on their scores as well.
[0073] At 565, the process may include adjusting, using a second machine learning model using a continuous action space, one or more objective parameters for the selected first organ‐at‐risk, in accordance with some example embodiments. For example, the organ-at-risk selected at 555 may then be adjusted at 560. Specifically, the objectives for the selected organ- at-risk may be provided as an input to the second ML model 1200 which uses a continuous action space to determine the adjusted objective parameters for that first organ-at-risk. At 565, the adjustment may be for a plurality of selected OARs as well.
[0074] At 570, the process may include re-optimizing, using the first machine learning model configured as the optimization engine, the treatment plan including the adjusted one or more objective parameters for the selected first organ‐at‐risk, in accordance with some example embodiments. For example, the adjusted objective for the first organ-at-risk may be provided to the first ML model 1080, which re-optimizes the treatment plan using at least the new adjusted objective. As noted above, the first and second ML models may repeat the optimization and adjustment process. For example, a predetermined set of iterations of, for example 200 may be set, and after 200 iterations, the output of the first ML model may be provided, at 575, as the output (e.g., a re-optimized treatment plan including one or more treatment parameters).
[0075] In some implementations, the current subject matter may be configured to be implemented in a system 600, as shown in FIG. 6. The system may be used to provide one or more aspects disclosed herein, such as the machine learning models (e.g., the first ML model, the second ML model, etc.), the treatment planning system 100, and / or other operations disclosed herein at for example FIGs. 1 A, 5A, 5B, etc.)
[0076] The system 600 may include a processor 610, a memory 620, a storage device 630, and an input / output device 640. Each of the components (e.g., processor 610, memory 620, storage device 630, and input / output device 640) may be interconnected using a system bus 650. The processor 610 may be configured to process instructions for execution within the systemVia Patent Center Docket No.024636-771WO1 / 2025-054-2 Filing Date: September 8, 2025 Customer No.39564 600. In some implementations, the processor 610 may be a single-threaded processor. In alternate implementations, the processor 610 may be a multi-threaded processor. The processor 610 may include one or more of the following: at least one graphics processing unit (GPU), at least one AI processor, at least one ML model processor, and / or other types of processors.
[0077] The processor 610 may be further configured to process instructions stored in the memory 620 or on the storage device 630, including receiving or sending information through the input / output device 640.
[0078] The memory 620 may store information within the system 600. In some implementations, the memory 620 may be a computer-readable medium. In alternate implementations, the memory 620 may be a volatile memory unit. In yet some implementations, the memory 620 may be a non-volatile memory unit. The storage device 630 may be capable of providing mass storage for the system 600. In some implementations, the storage device 630 may be a computer-readable medium. In alternate implementations, the storage device 630 may be a floppy disk device, a hard disk device, an optical disk device, a tape device, non-volatile solid state memory, or any other type of storage device. The input / output device 640 may be configured to provide input / output operations for the system 600. In some implementations, the input / output device 640 may include a keyboard and / or pointing device. In alternate implementations, the input / output device 640 may include a display unit for displaying graphical user interfaces.
[0079] One or more aspects or features of the subject matter described herein can be realized in digital electronic circuitry, integrated circuitry, specially designed application specific integrated circuits (ASICs), field programmable gate arrays (FPGAs) computer hardware, AI chips, ML chips, GPUs, firmware, software, and / or combinations thereof. These various aspects or features can include implementation in one or more computer programs that are executable and / or interpretable on a programmable system including at least one programmable processor, which can be special or general purpose, coupled to receive data and instructions from, and to transmit data and instructions to, a storage system, at least one input device, and at least one output device. The programmable system or computing system may include clients and servers. A client and server are generally remote from each other and typically interact through a communication network. The relationship of client and server arises by virtue of computerVia Patent Center Docket No.024636-771WO1 / 2025-054-2 Filing Date: September 8, 2025 Customer No.39564 programs running on the respective computers and having a client-server relationship to each other.
[0080] These computer programs, which can also be referred to as programs, software, software applications, applications, components, or code, include machine instructions for a programmable processor, and can be implemented in a high-level procedural and / or object- oriented programming language, and / or in assembly / machine language. As used herein, the term “machine-readable medium” refers to any computer program product, apparatus and / or device, such as for example magnetic discs, optical disks, memory, and Programmable Logic Devices (PLDs), used to provide machine instructions and / or data to a programmable processor, including a machine-readable medium that receives machine instructions as a machine-readable signal. The term “machine-readable signal” refers to any signal used to provide machine instructions and / or data to a programmable processor. The machine-readable medium can store such machine instructions non-transitorily, such as for example as would a non-transient solid- state memory or a magnetic hard drive or any equivalent storage medium. The machine-readable medium can alternatively or additionally store such machine instructions in a transient manner, such as for example as would a processor cache or other random access memory associated with one or more physical processor cores.
[0081] To provide for interaction with a user, one or more aspects or features of the subject matter described herein can be implemented on a computer having a display device, such as for example a cathode ray tube (CRT) or a liquid crystal display (LCD) or a light emitting diode (LED) monitor for displaying information to the user and a keyboard and a pointing device, such as for example a mouse or a trackball, by which the user may provide input to the computer. Other kinds of devices can be used to provide for interaction with a user as well. For example, feedback provided to the user can be any form of sensory feedback, such as for example visual feedback, auditory feedback, or tactile feedback; and input from the user may be received in any form, including, but not limited to, acoustic, speech, or tactile input. Other possible input devices include, but are not limited to, touch screens or other touch-sensitive devices such as single or multi-point resistive or capacitive trackpads, voice recognition hardware and software, optical scanners, optical pointers, digital image capture devices and associated interpretation software, and the like.Via Patent Center Docket No.024636-771WO1 / 2025-054-2 Filing Date: September 8, 2025 Customer No.39564
[0082] The subject matter described herein can be embodied in systems, apparatus, methods, and / or articles depending on the desired configuration. The implementations set forth in the foregoing description do not represent all implementations consistent with the subject matter described herein. Instead, they are merely some examples consistent with aspects related to the described subject matter. Although a few variations have been described in detail above, other modifications or additions are possible. In particular, further features and / or variations can be provided in addition to those set forth herein. For example, the implementations described above can be directed to various combinations and sub-combinations of the disclosed features and / or combinations and sub-combinations of several further features disclosed above. In addition, the logic flows depicted in the accompanying figures and / or described herein do not necessarily require the particular order shown, or sequential order, to achieve desirable results. Other implementations may be within the scope of the following claims.
[0083] In view of the above-described implementations of subject matter this application discloses the following list of examples, wherein one feature of an example in isolation or more than one feature of said example taken in combination and, optionally, in combination with one or more features of one or more further examples are further examples also falling within the disclosure of this application:
[0084] Example 1. A method of automating radiation therapy treatment planning, the method comprising: receiving at least an objective for a control structure associated with a plurality of organs- at-risk; optimizing, using a first machine learning model configured as an optimization engine, a treatment plan including the control structure and the objective; selecting, based on one or more scores, at least a first organ‐at‐risk to enable objective adjustment; adjusting, using a second machine learning model using a continuous action space, one or more objective parameters for the selected first organ‐at‐risk; re-optimizing, using the first machine learning model configured as the optimization engine, the treatment plan including the adjusted one or more objective parameters for the selected first organ‐at‐risk; andVia Patent Center Docket No.024636-771WO1 / 2025-054-2 Filing Date: September 8, 2025 Customer No.39564 outputting the re-optimized treatment plan including one or more treatment parameters.
[0085] Example 2. The method of Example 1, wherein the control structure and the objective are received with a set of rules to deconflict objectives.
[0086] Example 3. The method of any of Examples 1-2 further comprising: initializing the objective to prioritize optimization of dose coverage of at least one target volume included in the control structure.
[0087] Example 4. The method of any of Examples 1-3, wherein the objectivecomprises a maximal dose, a minimal dose, and / or a mean dose, and wherein the controlstructure defines at least in part a structure associated with a target volume to be radiated, and / or wherein the control structure includes at least one clinical target volume, at least one planning target volume, at least one gap, and / or at least one ring.
[0088] Example 5. The method of any of Examples 1-4, wherein the first machine learning model comprises a neural network, an optimization engine, and / or a limited memory, Broyden–Fletcher–Goldfarb– Shanno (L-BFGS) optimization algorithm.
[0089] Example 6. The method of any of Examples 1-5, wherein the second machine learning model comprises a policy network using the continuous action space continuous action space.
[0090] Example 7. The method of any of Examples 1-6, wherein the second machine learning model further includes a feature selector, wherein the selecting of the first organ‐at‐risk is performed by the feature selector.
[0091] Example 8. The method of any of Examples 1-7, wherein the second machine learning model comprising the policy network policy is trained using a dose distribution‐basedVia Patent Center Docket No.024636-771WO1 / 2025-054-2 Filing Date: September 8, 2025 Customer No.39564 reward function, wherein the continuous action space is constrained to a real-valued number, and an agent of the policy network uses one or more states as one or more inputs.
[0092] Example 9. The method of any of Examples 1-8, wherein the one or more states each comprise a dose‐volume histogram feature.
[0093] Example 10. A system for automating radiation therapy treatment planning, the system comprising: at least one processor; and at least one memory including instructions, which when executed by the at least one processor, causes operations comprising: receiving at least an objective for a control structure associated with a plurality of organs- at-risk; optimizing, using a first machine learning model configured as an optimization engine, a treatment plan including the control structure and the objective; selecting, based on one or more scores, at least a first organ‐at‐risk to enable objective adjustment; adjusting, using a second machine learning model using a continuous action space, one or more objective parameters for the selected first organ‐at‐risk; re-optimizing, using the first machine learning model configured as the optimization engine, the treatment plan including the adjusted one or more objective parameters for the selected first organ‐at‐risk; and outputting the re-optimized treatment plan including one or more treatment parameters.
[0094] Example 11. The system of Example 10, wherein the control structure and the objective are received with a set of rules to deconflict objectives.
[0095] Example 12. The system of any of Examples 10-11 further comprising: initializing the objective to prioritize optimization of dose coverage of at least one target volume included in the control structure.Via Patent Center Docket No.024636-771WO1 / 2025-054-2 Filing Date: September 8, 2025 Customer No.39564
[0096] Example 13. The system of any of Examples 10-12, wherein the objectivecomprises a maximal dose, a minimal dose, and / or a mean dose, and wherein the controlstructure defines at least in part a structure associated with a target volume to be radiated, and / or wherein the control structure includes at least one clinical target volume, at least one planning target volume, at least one gap, and / or at least one ring.
[0097] Example 14. The system of any of Examples 10-13, wherein the first machine learning model comprises a neural network, an optimization engine, and / or a limited memory, Broyden–Fletcher–Goldfarb– Shanno (L-BFGS) optimization algorithm.
[0098] Example 15. The system of any of Examples 10-14, wherein the second machine learning model comprises a policy network using the continuous action space continuous action space.
[0099] Example 16. The system of any of Examples 10-15, wherein the second machine learning model further includes a feature selector, wherein the selecting of the first organ‐at‐risk is performed by the feature selector.
[0100] Example 17. The system of any of Examples 10-16, wherein the second machine learning model comprising the policy network policy is trained using a dose distribution‐based reward function, wherein the continuous action space is constrained to a real- valued number, and an agent of the policy network uses one or more states as one or more inputs.
[0101] Example 18. The system of any of Examples 10-17, wherein the one or more states each comprise a dose‐volume histogram feature.
[0102] Example 19. A non-transitory computer-readable storage medium instructions, which when executed by at least one processor, causes operations comprising:Via Patent Center Docket No.024636-771WO1 / 2025-054-2 Filing Date: September 8, 2025 Customer No.39564 receiving at least an objective for a control structure associated with a plurality of organs- at-risk; optimizing, using a first machine learning model configured as an optimization engine, a treatment plan including the control structure and the objective; selecting, based on one or more scores, at least a first organ‐at‐risk to enable objective adjustment; adjusting, using a second machine learning model using a continuous action space, one or more objective parameters for the selected first organ‐at‐risk; re-optimizing, using the first machine learning model configured as the optimization engine, the treatment plan including the adjusted one or more objective parameters for the selected first organ‐at‐risk; and outputting the re-optimized treatment plan including one or more treatment parameters.
[0103] The illustrated methods are exemplary only. Although the methods are illustrated as having a specific operational flow, two or more operations may be combined into a single operation, a single operation may be performed in two or more separate operations, one or more of the illustrated operations may not be present in various implementations, and / or additional operations which are not illustrated may be part of the methods.
Claims
Via Patent Center Docket No.024636-771WO1 / 2025-054-2 Filing Date: September 8, 2025 Customer No.39564 CLAIMS What is claimed is:
1. A method of automating radiation therapy treatment planning, the method comprising: receiving at least an objective for a control structure associated with a plurality of organs- at-risk; optimizing, using a first machine learning model configured as an optimization engine, a treatment plan including the control structure and the objective; selecting, based on one or more scores, at least a first organ‐at‐risk to enable objective adjustment; adjusting, using a second machine learning model using a continuous action space, one or more objective parameters for the selected first organ‐at‐risk; re-optimizing, using the first machine learning model configured as the optimization engine, the treatment plan including the adjusted one or more objective parameters for the selected first organ‐at‐risk; and outputting the re-optimized treatment plan including one or more treatment parameters.
2. The method of claim 1, wherein the control structure and the objective are received with a set of rules to deconflict objectives.
3. The method of claim 2 further comprising: initializing the objective to prioritize optimization of dose coverage of at least one target volume included in the control structure.
4. The method of claim 1, wherein the objective comprises a maximal dose, aminimal dose, and / or a mean dose, and wherein the control structure defines at least in part a structure associated with a target volume to be radiated, and / or wherein the control structure includes at least one clinical target volume, at least one planning target volume, at least one gap, and / or at least one ring.Via Patent Center Docket No.024636-771WO1 / 2025-054-2 Filing Date: September 8, 2025 Customer No.39564 5. The method of claim 1, wherein the first machine learning model comprises a neural network, an optimization engine, and / or a limited memory, Broyden–Fletcher–Goldfarb– Shanno (L-BFGS) optimization algorithm.
6. The method of claim 1, wherein the second machine learning model comprises a policy network using the continuous action space continuous action space.
7. The method of claim 6, wherein the second machine learning model further includes a feature selector, wherein the selecting of the first organ‐at‐risk is performed by the feature selector.
8. The method of claim 6, wherein the second machine learning model comprising the policy network policy is trained using a dose distribution‐based reward function, wherein the continuous action space is constrained to a real-valued number, and an agent of the policy network uses one or more states as one or more inputs.
9. The method of claim 8, wherein the one or more states each comprise a dose‐ volume histogram feature.
10. A system for automating radiation therapy treatment planning, the system comprising: at least one processor; and at least one memory including instructions, which when executed by the at least one processor, causes operations comprising: receiving at least an objective for a control structure associated with a plurality of organs- at-risk; optimizing, using a first machine learning model configured as an optimization engine, a treatment plan including the control structure and the objective;Via Patent Center Docket No.024636-771WO1 / 2025-054-2 Filing Date: September 8, 2025 Customer No.39564 selecting, based on one or more scores, at least a first organ‐at‐risk to enable objective adjustment; adjusting, using a second machine learning model using a continuous action space, one or more objective parameters for the selected first organ‐at‐risk; re-optimizing, using the first machine learning model configured as the optimization engine, the treatment plan including the adjusted one or more objective parameters for the selected first organ‐at‐risk; and outputting the re-optimized treatment plan including one or more treatment parameters.
11. The system of claim 10, wherein the control structure and the objective are received with a set of rules to deconflict objectives.
12. The system of claim 11 further comprising: initializing the objective to prioritize optimization of dose coverage of at least one target volume included in the control structure.
13. The system of claim 10, wherein the objective comprises a maximal dose, aminimal dose, and / or a mean dose, and wherein the control structure defines at least in part a structure associated with a target volume to be radiated, and / or wherein the control structure includes at least one clinical target volume, at least one planning target volume, at least one gap, and / or at least one ring.
14. The system of claim 10, wherein the first machine learning model comprises a neural network, an optimization engine, and / or a limited memory, Broyden–Fletcher–Goldfarb– Shanno (L-BFGS) optimization algorithm.
15. The system of claim 10, wherein the second machine learning model comprises a policy network using the continuous action space continuous action space.Via Patent Center Docket No.024636-771WO1 / 2025-054-2 Filing Date: September 8, 2025 Customer No.39564 16. The system of claim 15, wherein the second machine learning model further includes a feature selector, wherein the selecting of the first organ‐at‐risk is performed by the feature selector.
17. The system of claim 15, wherein the second machine learning model comprising the policy network policy is trained using a dose distribution‐based reward function, wherein the continuous action space is constrained to a real-valued number, and an agent of the policy network uses one or more states as one or more inputs.
18. The system of claim 17, wherein the one or more states each comprise a dose‐ volume histogram feature.
19. A non-transitory computer-readable storage medium instructions, which when executed by at least one processor, causes operations comprising: receiving at least an objective for a control structure associated with a plurality of organs- at-risk; optimizing, using a first machine learning model configured as an optimization engine, a treatment plan including the control structure and the objective; selecting, based on one or more scores, at least a first organ‐at‐risk to enable objective adjustment; adjusting, using a second machine learning model using a continuous action space, one or more objective parameters for the selected first organ‐at‐risk; re-optimizing, using the first machine learning model configured as the optimization engine, the treatment plan including the adjusted one or more objective parameters for the selected first organ‐at‐risk; and outputting the re-optimized treatment plan including one or more treatment parameters.
Citation Information
Patent Citations
Using reinforcement learning in radiation treatment planning optimization to locate dose-volume objectives
US11529531B2
Bed calculation with isotoxic planning
US20230218926A1
Machine-learning-driven auto-planning for radiation treatment
US20230293907A1
Deep learning based dosed prediction for treatment planning and quality assurance in radiation therapy
US20230352135A1