Method and apparatus for controlling processing delivery using reinforcement learning
By using an AI agent trained through reinforcement learning to optimize the parameters of the radiation delivery device in real time, the suboptimal nature of radiation therapy caused by changes in patient movement and anatomical structure has been resolved, achieving more precise radiation dose distribution and protection of healthy tissues.
Patent Information
- Application Number
- CN202180022524.9
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Priority Date
- 2020-03-27
- Filing Date
- 2021-03-22
- Publication Date
- 2026-02-24
- Estimated Expiration
- 2041-03-22
AI Technical Summary
Existing radiation therapy methods are difficult to adapt effectively to changes in patient movement and anatomical structure during treatment, resulting in suboptimal dose distribution and radiation exposure to healthy tissues.
An AI agent trained using reinforcement learning controls the parameters of the radiation delivery device in real time, including the position and orientation of the radiation source and movable components, and optimizes the distribution of radiation dose based on the patient's real-time geometry and motion changes.
It enables precise control of radiation dose delivery to target tissues even with changes in patient movement and anatomical structure, reducing radiation exposure to healthy tissues and improving the accuracy and effectiveness of treatment.
Smart Images

Figure CN115485018B_ABST
Abstract
Description
Technical Field
[0001] This invention relates to radiation treatment. In particular, it relates to methods and apparatus employing an artificial intelligence agent that uses machine learning to configure (train) a radiation delivery device and thereby provide a three-dimensional distribution of radiation dose. Background Technology
[0002] Carefully planned radiation dose delivery can be used to treat a variety of medical conditions. For example, radiation therapy is often used in conjunction with other treatments to manage and control cancer. While delivering an appropriate amount of radiation to certain structures or tissues can be beneficial, radiation generally damages living tissue. Therefore, it may be desirable to target radiation onto a target volume containing the structure or tissue to be irradiated, while minimizing (or keeping) the radiation dose delivered to surrounding tissues at a clinically acceptable level. Intensity modulated radiation therapy (IMRT) is a method that has been used to deliver radiation to a target volume in a living subject while mitigating the amount of radiation absorbed by surrounding tissues.
[0003] IMRT typically involves delivering shaped radiation beams from several different directions. For each direction, the beam cross-section can be shaped to conform to the projection of the target volume into the beam eye view (i.e., a view truncated along the central axis of the radiation beam, abbreviated as BEV). The beam may additionally or alternatively have other cross-sectional shapes. The radiation beams are typically delivered sequentially. Each radiation beam contributes to the desired dose within the target volume.
[0004] Typical radiation delivery devices have a radiation source, such as a linear accelerator, and a rotatable pedestal. The pedestal can be rotated to allow the radiation beam to be incident on the object from various angles. The shape of the incident radiation beam can be modified by a multi-leaf collimator (MLC). An MLC has multiple blades, most of which are opaque to radiation. The MLC blades define an aperture through which radiation can propagate. The position of the blades can be adjusted to change the shape of the aperture, thereby shaping the radiation beam propagating through the MLC. The MLC can also be rotated to different angles (e.g., about the BEV axis).
[0005] Methods for delivering radiation doses disclosed in the prior art involve first planning the dose distribution to be delivered. Such planning may include selecting a series of control points at which radiation doses are delivered, and for each control point, selecting a corresponding set of radiation delivery parameters. The locations of the control points and the parameters of the radiation delivery device (e.g., optimized) can be assigned within the solution space of all candidate solutions to determine the dose delivered at each control point and the trajectory of the corresponding radiation beam. Some of these prior art methods involve selecting control points and / or radiation delivery parameters that optimize the geometric radiation distribution between the planned target volume and the organ at risk (OAR) volume to minimize radiation exposure to critical structures while ensuring adequate radiation delivery to the target tissue. Typically, previously collected patient images are used to plan the dose distribution to assess whether the plan achieves the clinically desired treatment.
[0006] Typically, patient treatment involves administering multiple time-spaced treatment sessions (or, for brevity, sessions). During each treatment session, the radiation delivery device typically operates automatically based on a selected (e.g., planned) sequence of control points and corresponding radiation delivery parameters. At each control point, radiation is delivered to the object using a set of parameters (for the radiation delivery device), which are predefined (e.g., optimized) according to clinical objectives during the planning process. Control points may specify the location where the radiation source should be positioned along the trajectory, and the radiation delivery device parameters may define the characteristics of the beam at each control point. The only interactive action in these types of prior art treatment sessions is usually the premature termination of the planned delivery sequence in the event of an unexpected event, such as a sudden change in the object's position. Motion caused by respiratory cycles or intestinal gas within a treatment session (e.g., within the normal amplitude range) can be considered typical patient motion. This typical patient motion is generally considered (in the prior art) to be within the "expected limits" and does not lead to the termination or modification of the treatment session. However, since treatment plans are typically created and optimized based on static planning images, any spatial deviation (e.g., typical patient motion) will result in a suboptimal dose distribution. We cannot guarantee that this suboptimality is very small.
[0007] Between radiation delivery segments or between imaging and delivery of any segment, the patient's anatomical geometry (e.g., the corresponding geometry between target tissue and healthy tissue) may change. For example, the patient may lose weight, organs may move relative to organs within the patient's abdominal cavity, and the geometry of the target tissue may change, etc.
[0008] Various attempts have been made in the prior art to address changes in patient geometry that may occur during any particular treatment session, between treatment sessions, and / or between imaging and any treatment session. Such strategies include tracking, which attempts to move the patient (e.g., by moving the patient's bed) or beam aperture to attempt to follow a target aperture; gating, which involves completely shutting off the radiation source in response to respiratory cycle monitoring; and adaptive treatment, which allows adjustments between treatment sessions. These prior art attempts mitigate the problems associated with patient geometrical movement to some extent, but each of these techniques has its own drawbacks.
[0009] Despite advancements in the field of radiation therapy, there remains a general expectation for improved radiation treatment methods and devices, as well as radiation treatment planning methods and devices, to provide better control over radiation delivery, such as by delivering an appropriate dose to the target tissue and maintaining the dose delivered to healthy tissue at a clinically acceptable level. This improved control is generally expected to accommodate factors such as patient movement, differences in the geometry of the patient's anatomy, and so on.
[0010] The foregoing examples and related limitations of the related art are intended to be illustrative rather than exclusive. Further limitations of the related art will become apparent to those skilled in the art upon reading the specification and studying the accompanying drawings. Summary of the Invention
[0011] The following embodiments and aspects thereof are described and illustrated in conjunction with systems, tools, and methods intended to be exemplary and illustrative rather than limiting. In various embodiments, one or more of the aforementioned problems have been reduced or eliminated, while other embodiments involve other improvements.
[0012] Aspects of the present invention relate to planning and delivering radiation treatment by moving a radiation source along a trajectory relative to the object while delivering radiation to the object. An artificial intelligence (AI) agent trained using reinforcement learning (and / or some other suitable form of machine learning) is used to control radiation delivery parameters in an effort to achieve the desired delivery of radiation therapy. In some embodiments, the AI agent selects appropriate control steps (e.g., radiation delivery parameters for a specific time step) while taking into account patient motion, multiple differences in patient anatomical geometry, and so on.
[0013] One non-limiting aspect of the invention provides a method for determining a radiation dose to be delivered to a patient using a radiation delivery device. The method includes: providing a radiation delivery device comprising a radiation source and one or more movable elements; defining a machine state including the positions of the one or more movable elements, wherein the machine state defines the characteristics of radiation emitted by the radiation delivery device when the radiation source is activated; for each of a plurality of time steps in the radiation delivery portion: receiving at an artificial intelligence (AI) agent including a processor configured to execute software instructions: an observation set regarding the machine state and the geometry of the patient; determining, by the AI agent, a current treatment state of the patient based at least in part on the observation set; and determining, by the AI agent, a next action including a next machine state for a subsequent time step based at least in part on the current treatment state and an artificial intelligence (AI) strategy, the AI strategy being prior determined using a machine learning process performed on training data, the next machine state for the subsequent time step defining the characteristics of radiation emitted by the radiation delivery device in the subsequent time step.
[0014] Another non-limiting aspect of the invention provides a method for delivering a radiation dose to a patient using a radiation delivery device. The method includes: providing a radiation delivery device comprising a radiation source and one or more movable elements; defining a machine state, the machine state including the positions of the one or more movable elements, and wherein, when the radiation source is activated, the machine state defines the characteristics of radiation emitted by the radiation delivery device; for each of a plurality of time steps in the radiation delivery portion: receiving at an artificial intelligence (AI) agent including a processor configured to execute software instructions: an observation set regarding the machine state and the geometry of the patient; determining, by the AI agent, a current treatment state of the patient based at least in part on the observation set; determining, by the AI agent, a next action including a next machine state for a subsequent time step based at least in part on the current treatment state and an artificial intelligence (AI) strategy, the artificial intelligence (AI) strategy being prior determined using a machine learning process performed on training data, the next machine state for the subsequent time step defining the characteristics of radiation emitted by the radiation delivery device in the subsequent time step; and causing the device to perform the next action, thereby reaching the next machine state in the subsequent time step.
[0015] Another non-limiting aspect of the invention provides a system for determining the radiation dose delivered to a patient using a radiation delivery device. The system includes: a radiation delivery device including a radiation source and one or more movable elements; and an artificial intelligence (AI) agent including a processor configured to execute software instructions. The AI agent is operable via the output of one or more sensors to define a machine state, which includes the position of the one or more movable elements, and wherein, when the radiation source is activated, the machine state defines the characteristics of the radiation emitted by the radiation delivery device. The AI agent is configured, for each of a plurality of time steps in the radiation delivery portion: receive a set of observations regarding the machine state and the geometry of the patient; determine, at least in part, the current treatment state of the patient based on the set of observations; and determine, at least in part, a next action including a next machine state for a subsequent time step based on the current treatment state and an AI strategy, the AI strategy being prior-determined using a machine learning process performed on training data, the next machine state for the subsequent time step defining the characteristics of the radiation emitted by the radiation delivery device in the subsequent time step.
[0016] Another non-limiting aspect of the invention provides a system for delivering a radiation dose to a patient. The system includes: a radiation source; one or more movable elements; and an artificial intelligence (AI) agent including a processor configured to execute software instructions. The AI agent is configured to: define a machine state, the machine state including the position of the one or more movable elements, and wherein, when the radiation source is activated, the machine state defines the characteristics of radiation emitted by the radiation delivery device; and for each of a plurality of time steps in the radiation delivery portion: receive a set of observations regarding the machine state and the geometry of the patient; determine, at least in part, the current treatment state of the patient based on the set of observations; determine, at least in part, the current treatment state and an AI strategy including a next action for a next machine state for a subsequent time step, the AI strategy being prior determined using a machine learning process performed on training data, the next machine state for the subsequent time step defining the characteristics of radiation emitted by the radiation delivery device in the subsequent time step; and cause the device to perform the next action, thereby reaching the next machine state in the subsequent time step.
[0017] Another non-limiting aspect of the invention provides a method for delivering a radiation dose to a patient. The method includes providing a machine comprising: a radiation source; a plurality of movable aperture defining elements, the positions of which, in combination, define an aperture through which radiation from the radiation source can pass; and a movable gantry arm on which the radiation source and the plurality of aperture defining elements are mounted. The method further includes defining a machine state comprising: the intensity of the radiation source; the position of each of the plurality of movable aperture defining elements; and the position of the movable gantry. The method further includes, for each of a plurality of time intervals in a radiation delivery segment: receiving an observation set representing sensed information about the machine state and the patient state at an artificial intelligence (AI) agent including a processor configured to execute software instructions; determining a current processing state by the AI agent and at least in part based on the observation set; determining a next action, including for a next machine state in a subsequent time interval, by the AI agent at least in part based on the current processing state and using an artificial intelligence (AI) strategy determined using previously processed training data; and causing the machine to perform the next action, thereby reaching the next machine state in the subsequent time interval.
[0018] Another non-limiting aspect of the invention provides an apparatus for delivering a radiation dose to a patient. The apparatus includes: a machine comprising: a radiation source; a plurality of movable aperture defining elements, the positions of which are combined to define an aperture through which radiation from the radiation source can pass; and a movable gantry arm on which the radiation source and the plurality of aperture defining elements are mounted. The apparatus also includes an artificial intelligence (AI) agent comprising a processor configured to execute software instructions. The AI agent is configured to define machine states including: the intensity of the radiation source; the position of each of the plurality of movable aperture defining elements; and the position of the movable gantry. The AI agent is also configured to, for each of a plurality of time intervals in the radiation delivery section: receive at the artificial intelligence (AI) agent a set of observations representing sensed information about the machine state and the patient state; determine the current processing state based at least in part on the set of observations; determine the next action, including for the next machine state in the subsequent time interval, based at least in part on the current processing state and using an artificial intelligence (AI) strategy determined using training data from the previous processing; and cause the machine to perform the next action, thereby reaching the next machine state in the subsequent time interval.
[0019] Another non-limiting aspect of the invention provides a method for training an artificial intelligence (AI) agent using machine learning to deliver radiation doses to a patient. The method includes: receiving a set of clinical goals; determining a reward function based at least in part on the set of clinical goals; determining an AI strategy; providing a simulated radiation delivery device including a simulated radiation source and one or more simulated movable elements, wherein a simulated machine state is defined by a simulated intensity of the simulated radiation source and a simulated position of the one or more simulated movable elements, the machine state defining the characteristics of simulated radiation fractionated by the radiation delivery device. The method further includes, for each of a plurality of time steps in the simulated radiation delivery portion: receiving at the AI agent a set of observations representing simulated sensing information about the simulated machine state and the simulated patient state; determining a current simulated processing state based at least in part on the set of observations; updating the AI strategy at least based on a current reward, wherein the current reward is evaluated based on the current simulated processing state and the reward function; determining, by the AI agent at least in part on the current simulated processing state and the AI strategy, a next action including a next simulated machine state for a subsequent time step; and causing the simulated radiation delivery device to perform the next action, thereby reaching the next simulated machine state in the subsequent time step.
[0020] Another non-limiting aspect of the invention provides a method for training an artificial intelligence (AI) agent using machine learning to deliver radiation doses to a patient. The method includes: receiving a set of clinical goals; determining a reward function based at least in part on the set of clinical goals; determining an AI strategy; providing a simulator including: a simulated radiation source; a plurality of simulated movable aperture defining elements, the positions of which, in combination, define an aperture through which radiation from the simulated radiation source can pass; a simulated movable gantry arm on which the simulated radiation source and the plurality of simulated aperture defining elements are mounted; defining a simulated machine state including: the intensity of the simulated radiation source; the position of each of the plurality of simulated movable aperture defining elements; and the position of the simulated movable gantry. The method further includes, for each of a plurality of time steps in the simulated radiation delivery portion: receiving at an AI agent an observation set representing simulated sensing information about the state of the simulated machine and the state of the simulated patient; determining the current simulated processing state based at least in part on the observation set; updating the artificial intelligence policy based at least on the current reward, wherein the current reward is evaluated based on the current simulated processing state and a reward function; determining, by the AI agent at least in part on the current simulated processing state and the AI policy, a next action including a next simulated machine state for a subsequent time step; and causing the simulator to perform the next action, thereby reaching the next simulator state in a subsequent time step.
[0021] One or more movable elements may include one or more of the following: a plurality of aperture defining elements whose positions are combined to define an aperture through which radiation from a radiation source can be guided; and a radiation source moving element whose position defines the orientation of radiation from a radiation source that can be guided.
[0022] Determining the patient's current treatment status by the AI agent based at least in part on the observation set may include: determining that a previous time step exists in the radiation delivery segment, and determining the patient's current treatment status by the AI agent based at least in part on the patient's previous treatment status in the segment determined to be the previous time step. Determining the patient's current treatment status by the AI agent based at least in part on the observation set may include: determining that there is no previous time step in the radiation delivery segment and the patient has already been irradiated in the previous radiation delivery segment, and determining the patient's current treatment status by the AI agent based at least in part on the previous treatment status determined at the end of the previous radiation delivery segment. Determining the patient's current treatment status by the AI agent based at least in part on the observation set may include: determining that there is no previous time step in the radiation delivery segment and the patient has not been irradiated in the previous radiation delivery segment, and determining the patient's current treatment status by the AI agent based at least in part on defining the initial treatment status.
[0023] Defining the initial treatment state may include at least one of the following: defining the patient geometry based on image data obtained from the patient prior to radiation delivery; instantiating a variable in memory representing the cumulative treatment time initialized to zero; and instantiating a variable in memory representing the cumulative delivered dose initialized to zero.
[0024] If it can be determined that there is no previous time step in the radiation delivery phase and the patient has already been irradiated in the previous radiation delivery phase, then the method and system may include determining the patient's current treatment status by an AI agent based at least in part on the patient geometry determined using image data obtained from the patient after the previous radiation delivery phase.
[0025] The patient's current treatment status may include an estimated geometry of voxels of interest within the patient's body, which may be at least partially based on observations of the patient's geometry. Voxels of interest may include at least one of the following: voxels corresponding to the target tissue; and voxels corresponding to one or more organs within the patient's body.
[0026] Methods and systems may include determining an estimated geometry by an AI agent based at least in part on observations of the patient's geometry and one or more patient movement models. The one or more patient movement models may include models of changes in the patient's geometry due to respiration. The one or more patient movement models may include models that predict changes in the voxel geometry inside the patient's body based on changes in the geometry outside the patient's body.
[0027] One or more patient motion models may be based at least in part on multiple images of the patient acquired over a period of time prior to the radiation delivery portion. One or more patient motion models may be based on the correlation between a reference motion model and the multiple images of the patient acquired over the period of time. One or more patient motion models may include population-based models based on a population excluding the patient.
[0028] Determining a patient’s current treatment status may include determining the estimated cumulative dose absorbed by the target tissue during the radiation delivery phase.
[0029] Determining the patient’s current treatment status may include: updating the 3D image reconstruction of the volume of interest within the patient’s body based on observations of the patient’s geometry; converting the cumulative dose absorbed by the target tissue in a previous time step into the updated 3D image reconstruction; estimating the additional dose delivered to the target tissue at the current time step based at least on the updated 3D image reconstruction, observations of the machine’s state at the current time step, and the dose estimation engine; determining the estimated cumulative dose absorbed by the target tissue during the radiation delivery phase may include: determining the sum of the converted cumulative dose and the additional dose from the previous time step.
[0030] Determining the patient's current treatment status may include determining the estimated cumulative dose absorbed by the patient's non-target organs during the radiation delivery phase. Determining the estimated cumulative dose absorbed by the patient's non-target organs during the radiation delivery phase may include determining that the estimated cumulative dose absorbed by the patient's non-target organs during the radiation delivery phase exceeds a threshold for the non-target organ, and the method includes determining a next action including a next machine state for a subsequent time step, to include terminating irradiation of the patient's radiation delivery phase.
[0031] Determining a patient’s current treatment status may include determining the cumulative treatment time during the radiation delivery phase.
[0032] AI strategies may include mappings. Determining the next action, including the next machine state for subsequent time steps, by the AI agent may include using the mapping to create a correspondence from the current processing state to the next action.
[0033] In creating the mapping from the current processing state to the next action, the mapping can be based on maximizing the reward function that will maximize the cumulative reward obtained over all expected subsequent time steps.
[0034] The method or system may include a set of one or more constraints imposed on the AI agent at each of a plurality of time steps in the radiation delivery section, the set of one or more constraints restricting the option space that the AI agent can use to determine the next action. The set of one or more constraints may include any one or more of the following: the maximum distance that the radiation source can travel between the current time step and subsequent time steps; the maximum distance that one or more movable elements can travel between the current time step and subsequent time steps; the maximum change in the intensity of the radiation source between the current time step and subsequent time steps; and the maximum value of the intensity of the radiation source.
[0035] Methods or systems may include AI agents trained using reinforcement-based machine learning processes along with training data to determine AI policies.
[0036] The method or system may include defining an initial processing state, wherein: if a previous time step exists in the radiation delivery portion, the current processing state is defined at least in part based on the previous processing state of the previous time step in the delivery portion; if there is no previous time step in the radiation delivery portion but a previous portion exists, the current processing state is defined at least in part based on the previous processing state of the previous portion; and if there is no previous time step in the radiation delivery portion and no previous portion exists, the current processing state is defined at least in part based on the initial processing state.
[0037] Determining the current treatment status may include determining the cumulative dose absorbed by the target volume during the radiation delivery phase and the cumulative treatment time during the radiation delivery phase.
[0038] Determining the current processing state may include: updating the 3D image reconstruction of the volume of interest; updating the depiction of the volume of interest in the 3D image reconstruction; converting the cumulative dose absorbed by the target volume in the previous time step into the 3D image reconstruction; and calculating the additional dose delivered to the target volume at the current time step based at least on the 3D image reconstruction, an observation set regarding the machine state at the current time step, and a dose calculation engine. The cumulative dose may include the sum of the cumulative dose in the previous time step and the additional dose.
[0039] Converting the cumulative dose into a 3D image for reconstruction and calculating the additional dose delivered to the target volume can each include an updated dose matrix.
[0040] The AI agent may determine the next action based at least in part on an assessment of whether the cumulative dose absorbed by the volume of interest during the radiation delivery portion has exceeded the target portion dose. If the cumulative dose absorbed by the target volume during the radiation delivery portion has exceeded the target portion dose, the next action may include issuing a stop command and causing the machine (device) to perform the next action, which includes setting the intensity of the radiation source to zero.
[0041] The initial processing state may include variables instantiated in memory representing the cumulative processing time initialized to zero and the cumulative delivery dose initialized to zero.
[0042] AI strategies may include mapping, and determining the next action may involve the AI agent using the mapping to create a correspondence from the current processing state to the next action. In creating this correspondence, the mapping may rely in part on a function that maximizes the cumulative reward obtained over all subsequent time steps.
[0043] The method or system may include providing a bed for accommodating a patient during the delivery of radiation dose, the machine status including the position of the bed.
[0044] The method or system may include providing a patient motion model that represents the expected changes in the geometry of a volume of interest within the patient's body. Providing the patient motion model may include analyzing multiple images of the patient acquired prior to the radiation delivery portion. Providing the patient motion model may also include relating a reference motion model to one or more images of the patient acquired prior to the radiation delivery portion.
[0045] The method or system may include: at each of a plurality of time steps in the radiation delivery section, determining an estimated patient motion at the current time step, at least in part based on previous patient motion and patient motion models from one or more previous time steps. The processing state may include data representing previous patient motion from one or more previous time steps.
[0046] The method or system may include assessing, at each of multiple time steps within the radiation delivery section, whether estimated patient motion exceeds a motion management threshold, and if the estimated patient motion is assessed as exceeding the motion management threshold, indicating a requirement for active motion management in one or more of the patient state and the current treatment state. The motion management threshold may be approximately 3 mm.
[0047] The estimated patient motion may include aggregate values representing the overall motion at multiple points within the volume of interest. When determining the aggregate values of the estimated patient motion, a lower weight may be assigned to the motion in the volume of interest in the direction defined by the beam eye view of the radiation source, compared to the motion in the volume of interest in the direction transverse to the beam eye view.
[0048] Patient motion can describe the circulatory motion of the volume of interest caused by the patient's respiratory cycle.
[0049] A method or system may include a set of one or more constraints imposed on an AI agent at each of a plurality of time steps in the radiation delivery portion, the set of one or more constraints limiting the option space that the AI agent can use to determine the next machine state. The set of one or more constraints may include one or more of the following: the maximum distance that the radiation source can travel between the current time step and subsequent time steps; the maximum distance that each of a plurality of movable aperture defining elements can move between the current time step and subsequent time steps; the maximum change in the intensity of the radiation source between the current time step and subsequent time steps; and the maximum value of the intensity of the radiation source.
[0050] A set of one or more constraints imposed on an AI agent can enable the machine to move continuously between the current time step and subsequent time steps.
[0051] In addition to the exemplary aspects and embodiments described above, other aspects and embodiments will become clear by referring to the accompanying drawings and by studying the following detailed description. Attached Figure Description
[0052] Exemplary embodiments are illustrated in the accompanying drawings. The embodiments and drawings disclosed herein are intended to be illustrative rather than restrictive.
[0053] Figure 1 This is a schematic diagram of an exemplary radiation delivery device that can be practiced in conjunction with embodiments of the present invention.
[0054] Figure 1A This is a schematic diagram of another exemplary radiation delivery device that can be practiced in conjunction with embodiments of the present invention.
[0055] Figure 2 This is a schematic diagram of the trajectory of a radiation source according to an example embodiment.
[0056] Figure 3A This is a schematic cross-sectional view of a beam shaping mechanism according to an example embodiment.
[0057] Figure 3B This is a schematic beam eye plan view of a multi-leaf collimator-type beam shaping mechanism according to a specific embodiment.
[0058] Figure 4 This is a flowchart illustrating a method for delivering radiation processing guided by an artificial intelligence agent according to a specific embodiment of the present invention.
[0059] Figure 5 This is a flowchart illustrating a method for training an artificial intelligence agent using reinforcement learning according to a specific embodiment of the present invention.
[0060] Figure 6 This is a flowchart illustrating a method for determining the current processing state according to a specific embodiment of the present invention.
[0061] Figure 7 This is a flowchart illustrating a method for predicting patient movement according to a specific embodiment of the present invention. Detailed Implementation
[0062] Throughout the following description, specific details are set forth in order to provide a more thorough understanding to those skilled in the art. However, well-known elements may not have been shown or described in detail to avoid unnecessarily obscuring this disclosure. Therefore, the specification and drawings are to be considered illustrative rather than restrictive.
[0063] Aspects of the present invention relate to planning and delivering radiation treatment by moving a radiation source along a trajectory relative to the object while delivering radiation to the object. An artificial intelligence (AI) agent trained using reinforcement learning (and / or some other suitable form of machine learning) is used to control radiation delivery parameters in an effort to achieve the desired delivery of radiation therapy. In some embodiments, the AI agent selects appropriate control steps (e.g., radiation delivery parameters for a specific time step) while taking into account patient motion, differences in patient anatomical geometry, and so on.
[0064] Figure 1 An exemplary radiation delivery device 10 is shown, which includes a radiation source 12 capable of generating or otherwise emitting a radiation beam 14. For example, the radiation source 12 may include a linear accelerator. An object S is positioned on a table or “bed” 15, which may be placed in the path of the beam 14. The device 10 includes movable components that allow movement of the position of the radiation source 12 and the orientation of the radiation beam 14 relative to the object S. These components may be collectively referred to as a beam positioning mechanism 13.
[0065] In the illustrated radiation delivery apparatus 10, the beam positioning mechanism 13 includes a platform 16 that supports the radiation source 12 and is rotatable or pivotable about an axis 18. The axis 18 and the beam 14 intersect at an isocenter point 20. In some embodiments, a suitable beam positioning mechanism 13 may include a platform capable of moving the radiation source 12 with additional or alternative degrees of freedom. The beam positioning mechanism 13 of the illustrated embodiment also includes a movable bed 15. In the exemplary radiation delivery apparatus 10, the bed 15 can be positioned in any three orthogonal directions (in... Figure 1 The source 12 can translate in the X, Y, and Z directions and can rotate about axis 22. In some embodiments, bed 15 can additionally or alternatively rotate about one or more of its other axes. The position of source 12 and the orientation of beam 14 (relative to object S) can be changed by moving one or more of the movable parts of beam positioning mechanism 13.
[0066] Each individually controllable component used to move source 12 relative to object S and / or orient beam 14 may be referred to as a "motion axis". In some cases, moving source 12 or beam 14 along a specific trajectory may require the movement of two or more motion axes. Figure 1 In the exemplary radiation delivery device 10 shown, the motion axis includes:
[0067] • Rotation of the stand 16 around axis 18;
[0068] • Translation of bed 15 in any one or more of the X, Y, and Z directions; and
[0069] • The bed 15 rotates around axis 22.
[0070] Other embodiments of the radiation delivery device may include additional or alternative motion axes.
[0071] The radiation delivery device 10 typically includes a control system 23 capable of controlling the mechanical aspects of the radiation delivery device 10 and optionally controlling the intensity of its radiation source 12. As a non-limiting example, the control system 23 can in particular control the movement of the motion axes of the radiation delivery device 10, the intensity of the radiation source 12, the movement of the MLC blades (e.g., position and orientation) (described in more detail below), the movement of the MLC jaws (described in more detail below), and so on. The control system 23 typically includes hardware and / or software components. In the illustrated embodiment, the control system 23 includes a controller 24 capable of executing software instructions. For example, the control system 23 is preferably capable of receiving (as input) a desired set of parameters (e.g., position and / or orientation) of its motion axes and, in response to such input, controllably moving one or more of its motion axes to achieve the desired set of motion axis parameters. Similarly, the control system 23 may also receive instructions as input regarding other radiation delivery parameters. As a non-limiting example, other radiation delivery parameters may include: the intensity of the radiation source 12; one or more parameters of the beam shaping mechanism 33, which is described in more detail below (e.g., parameters corresponding to the movement, position, and / or orientation of the MLC blades and / or MLC jaws); and so on. In response to such inputs, the control system 23 can cause the radiation delivery device 10 to implement the corresponding radiation delivery parameters.
[0072] While radiation delivery device 10 represents a specific type of radiation delivery device that can be implemented in conjunction with the present invention, it should be understood that the present invention can be implemented in different radiation delivery devices, which may include different axes of motion and / or different sets of radiation delivery parameters. Generally, embodiments of the present invention can be implemented using any set of axes of motion that can generate relative movement between the radiation source 12 and the object S from a starting point along a trajectory to an ending point. Generally, embodiments of the present invention can be implemented using any set of radiation delivery parameters that can be used by any such radiation delivery device.
[0073] Figure 1A Another example of a radiation delivery device 10A providing an alternative set of motion axes is shown. In the exemplary device 10A, a source 12 is housed in an annular housing 26. A mechanism 27 allows the source 12 to move around the housing 26 to irradiate an object S from different sides. The object S rests on a stage 28, which can advance through a central aperture 29 in the housing 26. It has similar characteristics to... Figure 1A The device, schematically illustrated, is used to deliver radiation in a manner commonly referred to as "tomography."
[0074] According to a specific embodiment of the invention, the beam positioning mechanism 13 moves the source 12 and / or the beam 14 along a trajectory while a radiation dose is controllably delivered to a target area within the object S. The “trajectory” is a set of movements of one or more movable components of the beam positioning mechanism 13 that result in a change in beam position and orientation from a first position and orientation to a second position and orientation. The first and second positions and orientations are not necessarily different. For example, the trajectory may include a rotation of the gantry 16 from a starting point through a 360° angle about axis 18 to an ending point, in which case the beam position and orientation at the starting and ending points are the same.
[0075] The first and second beam positions and beam orientations can be specified by a first set of motion axis parameters (corresponding to the first beam position and the first beam orientation) and a second set of motion axis parameters (corresponding to the second beam position and the second beam orientation). As discussed above, the control system 23 of the radiation delivery device 10 can controllably move its motion axis between the first and second sets of motion axis parameters. Typically, the trajectory can be described by more than two beam positions and beam orientations. For example, the trajectory can be specified by multiple sets of motion axis parameters, each corresponding to a specific beam position and a specific beam orientation. The control system 23 can then controllably move its motion axis along the trajectory between each set of motion axis parameters. In some embodiments, an artificial intelligence (AI) agent (described in more detail below) can define the trajectory in each time step (e.g., by selecting motion axis parameters at successive control points).
[0076] The set of motion axis parameters may constitute a subset of the radiation delivery parameters corresponding to a control point. More generally, a control point may include (or be otherwise associated with) a set of control parameters (also referred to herein as a set of radiation delivery parameters) for a given time step. For any given control point, as a non-limiting example, such control parameters may include: a set of motion axis parameters, one or more parameters corresponding to the intensity of radiation source 12, one or more parameters of beam shaping mechanism 33 (e.g., parameters corresponding to the movement, position, and / or orientation of MLC blades and / or MLC jaws), which will be described in more detail below; and so on. In some embodiments, during radiation delivery, a trained artificial intelligence (AI) agent (described in more detail below) may select a set of control parameters for consecutive control points corresponding to consecutive time steps. In some embodiments, the trained AI agent may define a trajectory (e.g., the location of consecutive control points) in consecutive time steps. In some embodiments, the trajectory may be provided by some other entity (e.g., a clinician) or otherwise defined for the AI agent. In some embodiments, the location of control points on the trajectory may be defined by the AI agent.
[0077] Generally, the trajectory can be arbitrary and is limited only by the specific radiation delivery device and its specific beam positioning mechanism. Within the constraints imposed by the design of the specific radiation delivery device 10 and its beam positioning mechanism 13, the radiation source 12 and / or the beam 14 can be made to follow an arbitrary trajectory relative to the object S by appropriate combinations of movements that cause the available axes of motion to move.
[0078] Figure 2 The diagram schematically depicts a radiation source 12 traveling relative to an object S along an arbitrary three-dimensional trajectory 30 while delivering a radiation dose to the object S via a radiation beam 14. Figure 2 The control points of the trajectory are schematically shown in a star shape and corresponding beams. The position and orientation of the radiation beam 14 change as the source 12 moves along the trajectory 30. In some embodiments, changes in the position and / or orientation and / or other radiation delivery parameters of the beam 14 may occur substantially continuously (between control points) as the source 12 moves along the trajectory 30. In some embodiments, such changes may occur discretely at each control point. As the source 12 moves along the trajectory 30, a radiation dose may be delivered to the object S continuously (i.e., at all times during the movement of the source 12 along the trajectory 30) or intermittently (i.e., radiation is blocked or turned off at certain times during the movement of the source 12 along the trajectory 30). The source 12 may move continuously along the trajectory 30 or may move intermittently between various locations on the trajectory 30 (e.g., between control points). As discussed above, the trajectory 30 may be specified by an AI agent, which may define the trajectory (e.g., the location of continuous control points) at each time step.
[0079] Although the trajectory 30 can be arbitrarily defined, in some embodiments it may be desirable that the source 12 and / or beam 14 do not have to move back and forth along the same path. Therefore, in some embodiments, the trajectory 30 can be constrained so that it does not overlap with itself (except possibly at the beginning and end of the trajectory 30). In such embodiments, the positions of the motion axes of the radiation delivery device may differ except possibly at the start and end points of the trajectory 30. In such embodiments, processing time can be minimized (or at least reduced) by irradiating the object S only once from each set of motion axis positions.
[0080] In some embodiments, trajectory 30 may be constrained such that the motion axis of the radiation delivery device can move in one direction without having to reverse the direction (i.e., the source 12 and / or beam 14 need not be moved back and forth along the same path). Choosing trajectory 30 involving movement of the motion axis in a single direction can minimize wear on the components of the radiation delivery device. For example, in device 10, it may be desirable to move the stage 16 in one direction because the stage 16 may be relatively large (e.g., greater than 1 ton), and reversing the movement of the stage 16 at different locations on the trajectory could cause strain on the components of the radiation delivery device 10 (e.g., on the drivetrain associated with the movement of the stage 16).
[0081] In some embodiments, trajectory 30 may be constrained such that the motion axis of the radiation delivery device moves substantially continuously (i.e., without stopping). Substantively continuous movement of the motion axis along trajectory 30 may be preferred over discontinuous movement, as stopping and starting the motion axis can cause wear on the components of the radiation delivery device. In other embodiments, the motion axis of the radiation delivery device may be allowed to stop at one or more locations along trajectory 30 (e.g., at discrete control points).
[0082] In some embodiments, trajectory 30 may be constrained to include a single, unidirectional, continuous 360° rotation of the platform 16 about axis 18, such that trajectory 30 overlaps with itself only at its start and end points. In some embodiments, this single, unidirectional, continuous 360° rotation of the platform 16 about axis 18 may be coupled with a corresponding unidirectional, continuous translational or rotational movement of the bed 15, such that trajectory 30 does not overlap at all.
[0083] Radiation delivery devices, such as exemplary device 10 ( Figure 1 ) and 10A ( Figure 1A It typically includes an adjustable beam shaping mechanism 33 located between the source 12 and the object S for shaping the radiation beam 14. Figure 3A A beam-shaping mechanism 33 is schematically depicted between the source 12 and the object S. The beam-shaping mechanism 33 may include a fixed and / or movable metal assembly 31. The assembly 31 may define an aperture 31A through which portions of the radiating beam 14 can pass. The aperture 31A of the beam-shaping mechanism 33 defines a two-dimensional boundary of the radiating beam 14 in a plane perpendicular to the radiation direction from the source 12 to the target volume in the object S (e.g., perpendicular to the BEV). The control system 23 is preferably capable of controlling the configuration of the beam-shaping mechanism 33.
[0084] A non-limiting example of the adjustable beam shaping mechanism 33 includes a multi-leaf collimator (MLC) 35 located between the source 12 and the object S. Figure 3B A suitable MLC 35 is schematically depicted. For example... Figure 3BAs shown, the MLC 35 includes a plurality of blades 36 that can be independently translated into or out of the radiation field to define one or more apertures 38 through which radiation can pass. The blades 36, which may include metal components, can act as radiation blockers. In the illustrated embodiment, the blades 36 are translatable in the direction indicated by the double-headed arrow 41. The size and shape of the aperture(s) 38(s) can be adjusted by selectively positioning each blade 36.
[0085] like Figure 3B As shown in the illustrated embodiment, blades 36 are typically provided in pairs. The MLC 35 is typically mounted such that it can rotate about an axis 37 extending in a plane perpendicular to the blades 36 to different orientations (e.g., in the BEV direction). Figure 3B In the illustrated embodiment, axis 37 extends into and out of the page, and dashed outline 39 shows an example of an alternative orientation of MLC 35 around axis 37.
[0086] In some embodiments, in addition to the MLC blade 36, the beam shaping mechanism 33 may optionally include an MLC jaw 43. The MLC jaw 43 may include a pair of radiation-blocking metal jaws that are relatively movable, which may be thicker than the metal jaws of the MLC blade 36, and which may be positioned to define the outer periphery of an aperture that can be defined by the MLC blade 35. In some embodiments, the MLC jaw 43 may also be movable in direction 41 and / or about axis 37.
[0087] The configuration of MLC 35 can be specified through a set of beam-shaping parameters. As a non-limiting example, the parameter set may include MLC blade position parameters defining the position of each blade 36, orientation parameters defining the orientation of MLC 35 about axis 37, and one or more MLC jaw parameters defining the position and / or orientation of MLC jaw 43. The control system of the radiation delivery device (e.g., the control system 23 of radiation delivery device 10) is typically capable of controlling the position of the blades 36, the orientation of MLC 35 about axis 37, and the position / orientation of MLC jaw 43 in response to receiving a suitable set of beam-shaping parameters. MLCs can vary in design details such as the presence or absence of MLC jaw 43, the mobility of MLC jaw 43, the number of blades 36, the width of the blades 36, the shape of the tips and edges of the blades 36, the range of positions any blade 36 can have, the constraints imposed by the positions of other blades 36 on the position of one blade 36, the mechanical design of the MLC, etc. The invention described herein should be understood to be adaptable to any type of configurable beam shaping device 33, including MLCs with these and other design variations.
[0088] While the radiation source 12 is operating and while the radiation source 12 is moving around trajectory 30, the configuration of MLC 35 can be changed (e.g., by moving MLC jaw 43, rotating MLC jaw 43 about axis 37, moving blades 36, and / or rotating MLC 35 about axis 37), thereby allowing the shape of the aperture(s) 38 to be dynamically changed as radiation is delivered to the target volume in object S. Since MLC 35 can have a large number of blades 36, each blade 36 can be placed in a large number of positions, and MLC 35 can rotate about its axis 37, MLC 35 can have a large number of possible configurations.
[0089] Embodiments of the present invention provide methods and systems for training and / or using appropriately trained artificial intelligence (AI) agents to control radiation treatment of patients. The AI agent may be implemented on the processor portion of a computer or processing device to execute software instructions loaded into memory. The AI agent can observe its environment (e.g., the patient's state and the state of the radiation delivery device), and based on these observations along with its training, the AI agent can act autonomously to deliver radiation treatment in a manner that optimally achieves the desired treatment strategy (which may include a set of treatment targets). A suitable user (e.g., an oncologist and / or a properly trained clinician) can define the set of treatment targets, which can be provided to the AI agent prior to training. One or more constraints can also be provided to the AI agent prior to training. As a non-limiting example, such constraints may include: limitations on the radiation delivery device, limitations on the trajectory and / or which axes of motion can be moved during a portion, hard dose limitations, limitations on the duration of the portion, etc. Such constraints may be known a priori to the AI agent or may be input into the AI agent prior to training. The set of treatment targets and the set of constraints may be collectively referred to herein as a treatment strategy. The AI agent can then develop an artificial intelligence (AI) strategy (also referred to herein as a "treatment strategy" or "strategy" for brevity) by training using reinforcement learning and / or other machine learning methods designed to achieve the treatment strategy.
[0090] Figure 4 A radiation treatment delivery method 100 according to an exemplary embodiment of the present invention is schematically depicted. Method 100 can be used to deliver partial radiation treatment. That is, method 100 can have multiple partial iterations (each generally referred to as a part), which together provide a complete radiation treatment for a particular patient. Radiation treatment delivery method 100 can be executed at least partially by an artificial intelligence (AI) agent 25 and may involve: determining instructions (e.g., radiation delivery (control) parameters) at each time step in a series of time steps; providing those instructions to a radiation delivery device (e.g., ... Figure 1The radiation delivery device 10 shown has a controller 23; and thereby causes the radiation delivery device to deliver the desired radiation dose distribution to the object S. For ease of explanation and without loss of generality, method 100 is described in the remainder of this specification as... Figure 1 The radiation delivery device 10 shown is used in conjunction with it. AI agent 25 can be... Figure 1 This is a portion of the more general processing planning system 25C shown. In the illustrated embodiment, the AI agent 25 includes its own controller 25A, which is configured to execute suitable software 25B. In some embodiments, the control system 23 and the processing planning system 25C (which may include the AI agent 25) may share one or more controllers. As a non-limiting example, the processing planning system 25C and / or the AI agent 25 may be implemented by a suitably configured computer. In some embodiments, as a non-limiting example, the controller 25A may include one or more data processors, along with suitable hardware, including: accessible memory, logic circuitry, drivers, amplifiers, A / D and D / A converters, etc. Such a controller may include, but is not limited to, a microprocessor, an on-chip computer, a computer's CPU, or any other suitable data processor and / or microcontroller. The controller 25A may include multiple data processors.
[0091] Method 100 begins at box 110. In box 110, method 100 involves obtaining one or more treatment objectives 111A to be satisfied by delivering radiation treatment via method 100; and, optionally, one or more constraints 111B associated with radiation delivery. The treatment objectives 111A and / or constraints 111B in box 110 can be obtained in any suitable manner. The treatment objectives 111A and / or constraints 111B obtained in box 110 can be inputs to method 100. For example, such treatment objectives 111A and / or constraints 111B can be determined by a clinician or physician (possibly using some computer-based assistance) and provided as input to method 100. The treatment objectives 111A and / or constraints 111B obtained in box 110 can be determined as part of box 110. For example, treatment planning system 25C may include suitable software that allows a clinician or physician to create and / or estimate the treatment objectives 111A and / or constraints 111B of box 110 using a suitable user interface.
[0092] Box 110, processing target 111A, may include processing targets for a specific part and / or for the entire treatment that are delivered as part of method 100, which may then be divided (e.g., by the treatment planning system 25C and / or by the clinician) into processing targets 111A for a specific part that are delivered as part of method 100. The processing target 111A obtained in box 110 may include, for example, a desired dose distribution, which in turn may include: a plurality of desired radiation dose amounts (e.g., a desired minimum dose amount) to be delivered to voxels(s) in the target volume; and a desired maximum radiation dose amount to be delivered to voxels corresponding to healthy tissues and / or organs. Box 110, processing target 111A, may additionally or alternatively include other processing targets. As a non-limiting example, other treatment objectives 111A may include: the desired uniformity of dose distribution in a voxel corresponding to the target volume; the desired accuracy in which the dose distribution in the target volume should match the desired dose amount; the maximum time required to deliver radiation treatment based on the individual patient's ability to remain still during treatment; the priority or weight of different treatment objectives, etc.
[0093] Some treatment objectives 111A may be based on a so-called dose-volume histogram, which maps the percentage of a specific tissue volume (e.g., corresponding to the volume of a target cancerous or healthy organ) receiving a specific dose (or estimated dose). By specific but non-limiting examples, a dose-volume histogram-based treatment objective 111A may take the form of: “The volume percentage of critical organ X receiving dose Y should be less than Z”; “The volume percentage of target volume A receiving dose B should be greater than C”; “The maximum dose covering Z percentages of critical organ X should be Y”; “The minimum dose covering C percentages of target volume A should be B”; and so on. In general, box 110 treatment objectives 111A can take many different forms. As a non-limiting example, biological models may be used to compute metrics that estimate the probability that a specified dose distribution will control the probability of disease in the subject and / or the probability that a specified dose delivered to non-pathological tissue may cause complications. Such biological models are called radiobiological models. Box 110 treatment objectives 111A may be based in part on one or more radiobiological models.
[0094] Box 110 may also optionally involve obtaining one or more processing constraints 111B—since the AI agent 25 is not yet aware of these constraints. In some embodiments, some processing constraints 111B may be hard-coded into the AI agent 25 and set during the calibration of the AI agent 25 for use by a particular radiation delivery device, etc. Such box 110 processing constraints 111B may include hard constraints (e.g., hard constraints that define or limit the search space of possible actions) and / or soft constraints (e.g., which may themselves be represented as terms in a reward function). Such processing constraints 111B may include, for example, constraints related to the physical limitations of a particular radiation delivery device 10. As a non-limiting example, such processing constraints 111B may include: movement (position, velocity, and / or acceleration) constraints of the gantry 12, bed 15, and / or other components / axis of the radiation processing device 10; the need to maintain the movement of components of the radiation delivery device 10 (e.g., gantry 12); the minimum and / or maximum radiation intensity that the radiation source may output; and so on. In some embodiments, box 110 constraint 111B may be inherent in the radiation delivery device used to deliver radiation and may be predetermined or hard-coded into method 100. In some embodiments, a user (e.g., a clinician) may introduce one or more box 110 constraints 111B. For example, a clinician may specify that only one or more specific movement axes of the radiation delivery device may be used for partial processing.
[0095] Method 100 then proceeds to box 112, where one or more machine learning techniques are used to train the AI agent 25 to develop and determine a processing policy 113,π (also referred to herein as artificial intelligence policy 113,π, or AI policy 113,π), which guides the decision-making process of the AI agent 25 during radiation delivery. The AI agent 25 may be trained in box 112 using reinforcement learning techniques, whereby the AI agent 25 learns to take actions based on some concept of cumulative reward. Box 112 may involve the use of any suitable reinforcement learning technique or algorithm. Non-limiting examples of suitable reinforcement learning algorithms include: Monte Carlo techniques, Q-learning techniques, SARSA (state-action-reward-state-action) techniques, deep Q-network techniques, and so on. In some embodiments, box 112 may additionally or alternatively use one or more other suitable machine learning techniques to train the AI agent 25 and thereby develop a suitable processing policy 113,π.
[0096] In a particular embodiment based on reinforcement learning, box 112 relates to the use of reinforcement learning such that AI agent 25 can develop (“learn”) a processing policy 113,π, wherein AI agent 25 receives as input observations of the AI agent’s environment in the current and / or past time steps to determine a “processing state” s at any given time step (during processing or training), and maps the processing state to a preferred action a in the next time step. For brevity, this description refers to the “action” of AI agent 25. Such a reference to AI agent action may include AI agent 25 providing instructions to a radiation delivery system (e.g., radiation delivery parameters corresponding to the preferred action to be performed by the radiation delivery system), and then the radiation delivery system actually performing the action.
[0097] The patient's treatment state s (including the environment in which the AI agent 25 operates (e.g., observable or estimable variables in the current radiation treatment session)) can be construed as a Markov state, although this is not mandatory. A Markov state s is a state in which all history (e.g., the patient's treatment history) is incorporated into the current state s, or in other words, given the current state s, the future state is independent of historical states. In some embodiments, box 112, the training of the AI agent, can be implemented using reinforcement learning techniques that can be modeled as a Markov decision process (MDP). In an MDP, given the current treatment state s (e.g., in the current time step), the probability P... a (s, s') represents the probability that the processing state transitions to a new state s' after the AI agent 25 takes action a (i.e., provides the radiation delivery parameter set executed by the radiation delivery device 10) (e.g., at the end of the next time step). The transition from processing state s to the new processing state s' can be described as a stochastic process derived from multiple random variables. Non-limiting examples of variables that can be stochastically modeled include: the difference between the amount of radiation delivered from the radiation source and the amount of radiation absorbed by the tissue, the non-uniformity of the dose absorbed by the tissue, the sensitivity of different tissues to radiation, unexpected adverse reactions of the patient to the received radiation dose, patient movement, setup errors, etc. In the case that the processing state observation is modeled as a stochastic process, the AI agent 25 can generate a stochastic processing policy π, which specifies the probability π(a|s) of taking action a in each state s.
[0098] The processing strategy 113 of AI agent 25 can be based on optimization (e.g., maximizing) the accumulated expected reward E[R]. The accumulated expected reward E[R] can also be called the expected reward E[R], where the reward R is defined as the sum of future discounted rewards. Where r t Let be the reward at time step t, and γ∈[0,1] be the discount rate for rewards occurring in the future. The reward r at any time step...t The reward function can be determined based on the reward function. The reward function may include a metric representing the reward gained from the execution of an action based on the current processing state s (including, for example, observations of the current environment) and the expected change in the processing state resulting from that action (including, for example, expected environmental changes). The processing policy 113 π learned by the AI agent 25 in box 112 may involve maximizing the cumulative expected reward E[R]. This reward maximization problem may be subject to one or more constraints (e.g., processing constraint 111B in box 110).
[0099] The reward function may be at least partially based on the processing objective 111A of box 110, and optionally, one or more constraints in constraint 111B of box 110 (e.g., soft constraints may be incorporated into the reward function). The fact that the reward function is based on such a processing objective 111A means that the reward function can reflect how clinicians assess the consistency and success of radiation treatment plans. If the processing objective 111A obtained at box 110 is satisfied by executing action a of the AI agent at a specific time step t, it is attributed to a relatively high reward r. t A smaller, additional positive reward can be given for achieving one or more optimization objectives. Conversely, if the processing objective 111B of box 110 is not satisfied by a specific action a, or if the action produces a state that causes future uncertainty and may lead to the processing objective not being satisfied, a lower reward r can be attributed to. t During training in box 112, oncologists or clinicians may interfere with the training process to provide more accurate guidance, and in doing so, alter the reward function. For example, an oncologist might provide further instructions on which types of dose distributions are beneficial and which are harmful.
[0100] Figure 5 A machine learning method 200, which can be used to train an AI agent 25 in block 112 according to a particular embodiment, is schematically depicted. Figure 5 The machine learning method 200 depicted in a particular exemplary embodiment is a reinforcement learning method 200, but other machine learning techniques may be used in other embodiments. In some embodiments, method 200 may be based wholly or partially on method 100 ( Figure 4 The simulation environment used in method 200 can be based on a simulated representation of the patient to be processed in method 100. For example, the simulation environment used in method 200 can be based on the simulated representation of the patient to be processed in method 100. Figure 4The simulation environment used in training method 200 may incorporate data (e.g., CT scans, other images, ECG measurements, respiratory measurements, etc.) and / or models directly taken from the patients being processed in method 100, with the model parameters based on such patient-specific data. In some embodiments, some or all of the data used to implement training method 200 may be taken from population-based data and / or data and / or models from other individuals, with the model parameters based on such data.
[0101] Some problems that may be addressed using machine learning-based processing include patient motion during processing, changes in patient anatomical geometry (e.g., between parts, between imaging and processing, and / or during processing), and / or other changes in patient geometry. Unless the context otherwise requires, such motion and geometric changes may be collectively referred to herein as changes in patient geometry. In some embodiments, the simulation environment used in machine learning method 200 includes models of how patient geometry changes between parts, between imaging and processing, and / or during processing, etc. Such patient geometry models may include, but are not limited to, models of a specific patient that are entirely or partially based on measurements taken from a specific patient. For example, such a model of respiratory-related motion may be based on respiratory motion captured by 4D-CT. Such patient geometry simulation may be entirely or partially based on, but is not limited to, a library of motion models observed in a patient population. For example, if a patient's weight decreases by 10% between imaging and delivery of a specific part, a motion model may be used to simulate changes in the patient's anatomical geometry. The simulation environment may additionally or alternatively include algorithmic techniques for creating new motion model candidates (e.g., generative neural networks, as a non-limiting example).
[0102] Return now Figure 5 Method 200 begins at box 202, where the simulated treatment state of the patient is initialized. In some embodiments, the treatment state can be implemented as a suitable data structure (e.g., a vector), where each element of the data structure is initialized with a suitable value in box 202. In some embodiments, some parameters of the state initialization in box 202 can be selected based on known (or obtained) data. As a non-example, in training method 200 based on method 100 ( Figure 4 In the case of a specific patient to be processed, box 202 can be initialized based on data from images taken from the patient (e.g., voxels containing target tissues, voxels containing key organs, parameters of a respiratory model, etc.) and other patient-related measurements (e.g., ECG measurements, respiratory measurements, etc.). In some embodiments, some parameters of the initialization state of box 202 can be set to specific values and / or random values.
[0103] Method 200 then proceeds to box 205, which involves determining the simulated processing state. The simulated processing state in box 205 can be based on training data 207. In the first iteration of method 200, the simulated processing state in box 205 can be the same as the initial state in box 202. In subsequent iterations, the simulated processing state in box 205 can be (at least partially) based on a model of the simulated environment, which can be based on training data 207, and observations about the simulated environment that can be based on training data 207, possibly along with actions taken since the last execution of box 205. Method 200 then proceeds to box 210, which involves deciding whether to explore a random next action (box 210 "yes" branch) or to determine the next action based on the existing version of processing policy 113, π (box 210 "no" branch). The decision for Box 210 can be based on a trade-off between exploration (possible new actions – the "yes" branch of Box 210) and exploitation (existing processing strategy 113, π, which will maximize the cumulative expected reward – the "no" branch of Box 210). This trade-off between exploration and exploitation in Box 210 can be implemented using any suitable technique, including, as a non-limiting example, the so-called ε-greedy method, where 0 < ε < 1 are parameters controlling the quantities of exploration and exploitation. For example, the probability of choosing exploitation (the "no" branch of Box 210) in Box 210 can be set to 1 - ε, and the probability of choosing exploration (the "yes" branch of Box 210) in Box 210 can be set to ε. Random numbers in the range (0, 1) can then be generated to make the decision for Box 210.
[0104] In some embodiments, the value of ε can vary during the implementation of method 200. For example, the parameter ε can be set large at the beginning of method 200 (leading to more "yes" branch (exploration) decisions in box 210) and can converge to a relatively small value in subsequent iterations of box 210 (leading to more "no" branch (exploitation) decisions in box 210) as the learned processing strategy 113, π, becomes more developed. The ε-greedy method represents only one possible technique that can be used to make decisions for box 210. Any other suitable technique(s) can be used. Other non-limiting examples of techniques that can be used in box 210 include counter-based exploration and proximity-based exploration. A non-limiting example of an additional or alternative technique for implementing the decision for box 210, along with the steps of boxes 215 or 220, is the so-called Boltzmann distributed exploration technique, where boxes 210, 215, and 220 are replaced by the probability of selecting any action associated with a Boltzmann weight (from which the expected reward is derived).
[0105] If the decision in box 210 is affirmative, then method 200 proceeds to box 220, where action a is randomly selected for the next time step, after which method 200 proceeds to box 225. If the decision in box 210 is negative, then method 200 proceeds to box 215, where action a for the next time step is selected based on the processing policy 113, π learned so far in method 200—that is, based on the cumulative expected reward from the processing policy 113, π learned so far.
[0106] After selecting the next action a (in box 215 or box 220), method 200 proceeds to box 225, which involves implementing the selected action a in a simulation environment. In practice, the next action a may involve specifying a set of radiation delivery parameters that represent the delivery of radiation with specific characteristics to a patient in the simulation environment in which method 200 is trained. The dose delivered to the simulated patient in the simulation environment can be estimated in box 225 according to known dose estimation techniques, which may include one or more models of the patient's body, one or more models of the patient's geometry, including variations in the patient's geometry (e.g., changes in the anatomical geometry of the patient's movement and / or organs and / or target tissues).
[0107] Once the action in box 225 has been implemented, method 200 proceeds to box 230, which involves determining the reward r for the action in box 225 that was just implemented. t As discussed in this article, the reward r t This can include a metric on how well the action of response box 225 has performed relative to achieving processing goal 111A. It should be noted that in some implementations and / or scenarios, the reward r may not be known in every state. t —Or it may only be known at the end of the treatment—or at the point when a sufficient dose has been delivered to the target. In such a case, the determination of box 230 can be made when it is possible to do so. Furthermore, after the action of box 225 is performed, method 200 executes the steps of box 235, which involves observing the treatment state s (including various observables about the simulated environment and the simulated patient) at the next time step (i.e., after the action of implanting box 225). As a non-limiting example, such observables may include information about the patient's location, radiation delivery parameters of the radiation delivery device, estimated doses of various voxels delivered to the patient, etc. The treatment state of box 235 may be based on training data 207. In the currently preferred embodiment, any observations about the environment (including but not limited to the patient) that can be used in a real treatment situation (e.g. Figure 4 Method 100) may have a simulated counterpart in the simulation environment used during training (e.g., Figure 5 Method 200).
[0108] Method 200 then proceeds to box 240, which involves collecting interaction data. The interaction data collected in box 240 may include: the processing state of box 205 (i.e., the processing state before executing the action of box 225); the action of box 225; the reward of box 230; and the processing state of box 235 (i.e., the processing state after executing the action of box 225). After collecting this interaction data in box 240, method 200 proceeds to box 245, where the interaction data of box 240 is used to update processing policy 113, π. Processing policy 113, π can be updated in box 245 in the direction that maximizes the accumulated expected reward. As a non-limiting example, where method 200 implements a so-called Q-learning technique, processing policy 113, π involves selecting the expected reward r... t Maximize action a, and update the expected reward r for the selected action a in state s. t To complete the actual learning. That is, if in the current state s, action a actually results in a larger (smaller) reward than expected, then the expected reward r for action a in state s is... t It will increase (decrease). More precisely, it's not just the immediate reward, but also the observed new state s' (and the expected reward associated with the new state s'). Therefore, if an action a in state s actually leads to a new state s' with a high expected reward, then the expected reward r of action a in state s will increase (decrease). t It will also increase. It's important to note that the expected reward of state s can be understood as the maximum expected reward for any action taken while in state s.
[0109] Method 200 then proceeds to query 250. In box 250, if it is determined that the learning process should end (box 250 "Yes" branch), then method 200 ends. Otherwise (box 250 "No" branch), method 200 returns to box 205 for another iteration from box 205 to 250.
[0110] Return now Figure 4 Following training in box 112, method 100 proceeds to the radiation delivery process 114, which utilizes the machine learning-based processing strategy 113 of box 112 to deliver radiation processing to object S over multiple time steps. Figure 4In the illustrated embodiment, the radiation delivery process 114 begins at optional box 120, which involves obtaining an initial processing plan 123 for the portion delivered in method 100. The initial processing plan 123 can be obtained as input to method 100 and can be determined using any suitable radiation delivery planning technique known in the art. An exemplary and non-limiting radiation delivery planning technique is described in Patent Cooperation Treaty Application No. PCT / CA2006 / 001225, which is incorporated herein by reference. In some embodiments, the initial processing plan 123 may be provided by a processing planning system 25C (… Figure 1 While this can be used for development, it is not necessary. In some embodiments, an explicit initial processing plan 123 is not required. For example, if the initial processing state s of the object is known and it is assumed that the object being processed has not moved at all, then the processing strategy 113 of application box 112, π, will specify the processing action a (e.g., a set of radiation delivery parameters) that will take effect at each time step. Such a series of processing actions can be understood as the initial processing plan 123, even if such an initial processing plan may not be explicitly defined.
[0111] The initial treatment plan 123 may specify a series of actions a (radiation delivery parameters) over a series of time steps, which, when provided to the radiation delivery device 10, cause the radiation delivery device 10 to deliver an initially desired partial dose to the patient over the series of time steps. However, the initial treatment plan 123 can be understood as fixed. Advantageously, the radiation delivery process 114 of method 100 can adapt to unpredictable changes in the patient's treatment state s (e.g., patient movement, etc.) and the radiation delivery actions a performed in each time step can be controllably adjusted to adapt to such changes based on a time-based treatment strategy 113, π developed in box 112 by machine learning (e.g., reinforcement learning).
[0112] exist Figure 4 In the schematic diagram, the radiation delivery process 114 includes an iterative loop between boxes 125 and 140. Each iteration of the delivery process 114 and its constituent steps may involve AI agent 25 selecting and executing a specific action a for a corresponding time step (e.g., providing radiation delivery parameters to radiation delivery device 10 to cause radiation delivery device 10 to perform the action). The combination of iterations of box 114 may provide a treatment portion for a specific patient, wherein the treatment plan is adapted to inter- and / or intra-part variations in patient geometry based on the treatment strategy 113, π learned in box 112.
[0113] In the illustrated embodiment of method 100, the processing delivery procedure 114 involves iteratively controlling a radiation delivery device, such as exemplary device 10, to deliver radiation to the object S at multiple time steps according to the processing strategy 113 of block 112. For any given time step, in block 130, the AI agent 25 may determine a decision on the action a to be taken for the next time step based at least in part on an evaluation of empirical observations of the current time step (and / or past time steps). Such empirical observations may include, for example, observations related to the patient's location and / or movement (e.g., patient geometry as discussed above). In block 125, the AI agent 25 may determine a processing state 127 for the current time step using observations at the current time step (and / or past time steps). The processing state 127 at the current time step may therefore include information about the patient's geometry. In block 130, the AI agent 25 may determine a decision on the action 133 to be taken for the next time step based at least in part on an evaluation of the patient's processing state 127 for the current time step in block 125. Given a processing state 127 for the current time step, and a processing strategy 113 for box 112, π can map the current processing state 127 to a specific action 133(a). As discussed above, the output of AI agent 25 is "action" 133(a), which includes manipulating or updating the radiation delivery parameters for the next time step. The updated radiation delivery parameters are then provided to radiation delivery device 10 so that radiation delivery device 10 delivers radiation to the patient in accordance with the radiation delivery parameters and in a manner that further satisfies the radiation treatment plan (e.g., approaching the radiation treatment target of box 110).
[0114] In some embodiments, the AI agent 25 used in block 112 and in the delivery process 114 can be implemented by the same means, such as... Figure 1 The case of AI agent 25 shown in the processing plan system 25C of the embodiment. This is generally not required. In some embodiments, the AI agent executing block 112 is embodied and trained separately on a suitable computing device different from the computing device used to execute the processing delivery process 114. The processing strategy 113, π of block 112 can be imported into the computing device executing the processing delivery process 114. The training process of block 112 can generally be implemented on a computing device different from the computing device executing the processing delivery process 114. The processing strategy 113, π of block 112 can be embodied by any suitable data structure that can be stored in memory and executed by a data processor. Some non-limiting examples that can emulate the processing strategy 113, π include hash tables, vector arrays, tuples, and multidimensional tensors.
[0115] After optionally obtaining the initial processing plan 123 in box 120, at any given time step, the processing delivery process 114 begins in box 125, where the AI agent determines the current state of the processing, or the "processing state" 127(s) for the current time step. t The processing state of box 125 is 127(s). t This typically includes all information known to AI agent 25 up to the current time step. As a non-limiting example, this is incorporated into processing state 127(s). t The information in the process may include: information about observations of the environment made at the current time step and / or past time steps (including observations about the patient (e.g., patient geometry), observations about the radiation delivery device 10, etc.); dose information delivered to the patient at the current time step and / or past time steps (e.g., dose estimation); etc. In some embodiments, processing state 127(s) t () can be represented as a vector, array or other similar data structure.
[0116] As a specific and non-restrictive example, merged into processing state 127(s) t The information in the document may include any one or more of the following: cumulative treatment time (in the current section, in (multiple) previous sections, and / or total treatment time); estimates of doses delivered to various tissues (e.g., target volume and OAR volume) or characteristics of the delivered doses during the current and / or past sections (e.g., 3D dose distribution reconstruction, mean dose, mean dose for a specific volume, fluence map, dose-volume histogram, dose-volume information, volume-dose type aggregated dose information, geometrically equivalent dose, dose homogeneity index, etc.); indices and / or measures relating to the temporal characteristics of the doses delivered in the current and / or past sections (e.g., indices relating to how the fractionation and / or dose delivery rate affects the biological effects of the delivered dose); observed or otherwise known patient / tissue motion characteristics (e.g., direct measurements of patient motion; past measurements of patient respiratory activity); and observed or otherwise known characteristics of the radiation delivery device 10 (e.g., MLC blade velocity and position, MLC jaw velocity and / or position, gantry velocity and / or position, bed velocity and / or position).
[0117] Figure 6 A non-limiting example embodiment of method 300 is shown, which can be used to determine suitable methods in block 125 ( Figure 4 Processing state 127 (s) used in ) t ). Figure 6It begins at decision box 310, which evaluates whether a previous time step exists in the current processing section. If the evaluation in box 310 is positive and a previous time step t-1 exists, then the processing state s at the previous time step is observed / determined. t-1 (Referred to in this document as the previous processing state 312) can be used as input to method 300 to determine the current processing state 127 (s) at the current time step t. t ).
[0118] If the evaluation of box 310 is negative, then method 300 proceeds to decision box 314, which evaluates whether a previous processing portion exists. If the evaluation of decision box 314 is positive, then the processing state at the end of the previous portion at its last time step T is s. T (Referred to in this document as the previous partial end state 316) can be used as input to method 300 to determine the current processing state 127(s). t In box 318, the previous part's ending state 316 can be updated. As a non-limiting example, the update of box 318 may include: modeling the passage of time between the end of the previous part and the current part; updating the patient image and / or other detected / sensed information about the patient, including patient geometry; updating the patient model (which may be based on the updated patient image and / or other detected / sensed information about the patient, including patient geometry); and so on. The update of box 318 may be based in part on the sensed information input to AI agent 25, the patient image, the patient-specific model, the population-based model, and so on. The latest sensed information, patient image, and / or model used in the update of box 318 can be advantageously used to take into account the inter-part variability of the patient's physiology (e.g., patient geometry) between processing parts—that is, based on (the (…applied in box 130)…) Figure 4 The processing strategy of machine learning may take into account the latest available information about the patient's current condition.
[0119] If the evaluation of box 314 is negative, this indicates that there is no previous processing portion or time step scenario (i.e., the process of starting radiation processing according to the processing plan—the first iteration of process 114 of method 100). Figure 4 (where the initial processing state is determined for the first time is 127). In this case, the initial processing state s can be determined in box 320. t=0 This initial processing state s t=0This can be the input to methods 100 and 300. As a non-limiting example, the initial processing state of box 320 can be based on one or more suitable patient images, other types of detected / sensed information about the patient, including patient geometry, patient-specific models, population-based models, known or determinable information about the radiation delivery device 10, information input by a clinician, and so on. In a non-limiting example embodiment, box 320 can involve AI agent 25 instantiating a vector (or other suitable data structure) containing dose estimation variables in memory, where the values of the dose estimation variables are set to 0.
[0120] Based on various possible past (or initial) processing states from previous time steps (respectively processing state 312(s)... t-1 ), from the processing state 316(s) of the previous part T ), and / or initial processing state 320(s) t=0 After defining the initial processing state, method 300 proceeds to process 322. Process 322 involves obtaining a set of environmental observations. As a non-limiting example, such environmental observations may include: observation 324, which includes information about the current state of the radiation delivery device 10; observation 326, which includes information about the current state of the patient (e.g., patient geometry), and so on. Such environmental observations may be based on, but are not limited to, appropriate patient images, other types of detected / sensed information about the patient (including information about patient geometry), patient-specific models, population-based models, information from clinician input, and so on.
[0121] Box 324 may relate to observing (e.g., sensing) information about the radiation processing delivery device 10. The information determined in box 324 may be referred to herein as machine status. As a non-limiting example, the machine status information in box 324 may include current measurements, operating parameters, and / or otherwise observable or known information about the radiation source's beam intensity, MLC blade position, MLC orientation angle, etc. Information such as table angle and bed position. In some embodiments, some observable quantities constituting the machine state are sensed by sensors external to the radiation delivery device, such as images or image streams captured from one or more cameras and / or other suitable sensors. In some embodiments, some observable quantities constituting the machine state are sensed by sensors, systems, and / or components that are part of the radiation processing system 10. In some embodiments, some of these observable quantities may be otherwise known operational parameters or information of the control system 23 of the radiation delivery device 10. Figure 1For example, the radiation delivery parameters discussed above can be controlled by the control system 23 of the radiation delivery device 10. In this case, the control system 23 is able to retrieve machine state parameters corresponding to such radiation delivery parameters (e.g., from internal memory or otherwise) and provide them to the AI agent 25 or any other software process of the execution block 324.
[0122] Box 326 relates to observing (e.g., sensing) information about the patient's current state. The information determined in box 326 may be referred to herein as the current patient state. As a non-limiting example, the current patient state information in box 326 may include observations about the patient's geometry. Non-limiting examples of how patient geometry can be obtained include referencing from a marked box and / or the patient's relative positioning to a reference mark. Patient geometry can also be obtained using captured images, such as images from a surface camera, 2D kV X-ray images, 2D MV inlet images, ultrasound images, magnetic resonance images, etc. In some embodiments, 3D images of the patient and the tissue of interest can be reconstructed from images obtained through any suitable imaging modality. For example, in some embodiments, 2D kV X-ray images can be used to reconstruct 3D cone-beam CT images. Such 3D images can be automatically segmented during the execution of box 326 to depict the target volume and OAR volume.
[0123] The current patient status information in box 326 may additionally or alternatively include observations of the dose distribution of the target volume and the OAR volume. This dose distribution can be measured and / or estimated using appropriate patient-specific, population-based, and / or empirical models. Several possible dosimetry instruments and methods can be employed to quantitatively assess the dose distribution of the target tissue based on the characteristics of the delivered radiation. As an illustrative example of a dosimetry method that can be employed at box 326, the COMPASS dose calculation algorithm, combined with sensed ionization data from the detector array, can be used to calculate the dose distribution of the target volume within the patient's body during IMRT. Other possible techniques include placing an Exradin A1SL type ionization chamber at an appropriate location close to the volume of interest to measure the delivered radiation dose.
[0124] The exemplary dosimetry method described above describes point dose measurement, which correlates the radiation dose delivered in a single delivery with the absorbed dose in tissue. In some embodiments, box 326 observes the resulting cumulative dose distribution information. In some embodiments, Software (and / or similar software) can be with A 3D diode array detector (and / or similar device) is used in combination to generate a dose-volume histogram (DVH) that can assess the progress of the current treatment plan without further processing.
[0125] In some embodiments, the patient status of box 326 may additionally or alternatively be based on other types of detected / sensed information about the patient (including information about the patient geometry), patient-specific models, population-based models, information input by clinicians, and so on.
[0126] At the end of process 322, after obtaining observations representing sensed information about the current machine state and the current patient state at boxes 324 and 326 respectively, method 300 can proceed to optional box 328 to predict expected patient motion (patient geometry). Figure 7 A non-limiting example of a method 400 for predicting a patient geometry suitable for use in box 328 is shown. Method 400 can estimate the expected displacement (e.g., updated geometry) of the patient's volume of interest in the next time step, such that this expected patient displacement (geometry) information can be incorporated into the radiation delivery parameters for the next time step. In this way, radiation processing delivery in the next time step can accurately target(multiple) target volumes and avoid(multiple) OAR volumes. In method 400 ( Figure 7 In the illustrated embodiment, a representative patient motion model 410 is supplied to method 400. (Box 124) Figure 4 This can optionally be performed as part of method 100 to generate a patient motion model 410. It may be desirable for the patient motion model 410 to capture anticipated changes in the patient geometry within and around the volume of interest during a radiation therapy treatment plan. For example, in the case of treating a tumor in a patient's brain, it may be desirable for the patient motion model 410 to correlate respiratory movements with the contribution of these movements to the movement of the brain tumor and surrounding tissues within the patient's skull.
[0127] In the execution box 124 ( Figure 4In an exemplary embodiment, a representative motion model for the patient based on respiratory movements is obtained by analyzing a set of multiple static 3D images acquired at consecutive time intervals. In another exemplary embodiment, a 4D image set is acquired and analyzed to obtain a patient motion model. Such a 4D image set allows for the reconstruction of multiple discrete 3D scans that can demonstrate the range of tissue motion and allows for the acquisition of tissue path data during respiratory intervention phases. Analysis of the image set may include applying image segmentation algorithms to identify the location and motion of the volume of interest between images and / or image frames. Any suitable imaging modality can be used to acquire the image set to be analyzed to determine the patient motion model. In some embodiments, computed tomography (CT), ultrasound imaging, magnetic resonance imaging modalities, etc., are employed to acquire image sets capturing temporal changes in the location of internal tissues. In some embodiments, observed motion of the patient's external surfaces may be used additionally or alternatively with a model that correlates external motion with the motion of internal volumes (and / or may be used as an alternative). Other sources of anticipated voluntary and / or involuntary patient motion and their contribution to the motion of the volume of interest may also be modeled individually or in combination with each other, such as motion caused by intestinal gas, overall patient movement (i.e., movement of parts of the patient's body), etc. The resulting models can be combined in any appropriate manner so that the resulting patient motion model in box 124 can accurately account for the various different expected patient motions during treatment.
[0128] In some embodiments, the patient motion model generated at box 124 may be specific to the patient receiving radiation treatment (e.g., via...). Figure 4 (Process delivery procedure 114). In such embodiments, prior to delivery of radiation processing via process 114, an image set capturing the patient's motion can be acquired and analyzed to provide a basis for creating a patient motion model as described above. In some embodiments, determining the patient motion model at block 124 includes employing motion models of a reference patient and / or a reference patient population. The reference patient motion model can be selected from a database containing motion models from various patient profiles (e.g., based on characteristics such as age, weight, sex, or height), which contains motion models that correlate patient movement with movement of different tissues of interest. A subset of patient profiles most closely related to the current patient and motion models associated with the volume of interest can be selected from the database as the reference motion model. In some embodiments, the reference patient motion model is based at least in part on a population-based statistical model—e.g., based on the average of a reference patient population.
[0129] In some embodiments, the reference motion model may include 4D data that can be correlated with static 3D images of a specific patient to perform box 124 and thereby generate a patient motion model 410. This correlation can be achieved using any suitable numerical method or algorithm, including techniques such as linear regression and machine learning. In scenarios where cost and / or time are considerations, using a reference motion model rather than custom modeling the motion for each individual patient may be preferred, as acquiring a set of 3D or 4D images can be a resource-intensive process. Radiation dose delivered to the patient can also be minimized due to the extensive CT or other imaging required in the motion model based solely on images of a specific patient. In an example embodiment, the patient motion model 410 is represented by deformation and displacement matrices and / or tensors with key-value pairs representing the motion of one or more internal volumes at different points in the patient's respiratory cycle.
[0130] In conjunction with supplying the patient motion model 410 as input to the method 400, a previous patient geometry 412 may also be supplied as input to the method 400. The previous patient geometry 412 may include information about the geometry of the patient himself or her and / or patient tissues at one or more previous time steps. For example, the previous patient geometry 412 may have a historical length of l = 3, including the geometry of patient tissues (and / or the patient) from three previous time steps. In some embodiments, for example, patient geometry data p at each such time step... t-3 p t-2 and p t-1 This can be represented as a matrix of coordinates for discrete tissue / patient points (tissue / patient markers). In some embodiments, the previous patient geometry p t-3 p t-2 and p t-1 This includes the spatial coordinates of multiple points on the tissue of interest within the patient at each of these previous time steps. For example, previous patient geometry data p from previous time steps. t-3 p t-2 and p t-1 Past processing states can be embodied in the previous patient geometry 412 and included in boxes 312, 316 and 320 of method 300. t-1 s T and s t=0 In addition, at the current time step p t The current patient geometry can be determined by observing the current patient state at box 326 in method 300.
[0131] At box 414 of method 400, the subsequent time step can be estimated based on the patient motion model 410 and / or the previous patient geometry 412. The geometry of the volume of interest (e.g., typically target tissue, healthy tissue, OAR tissue, and / or patient) is considered. In one example embodiment, patient motion (changes in patient geometry) can be predicted in box 414 by describing cyclical motion (such as respiratory cycles) in consideration of patient motion model 410. Previous patient geometry 412 can be used to track the current phase of the motion cycle, so the prediction in box 414 can assume that the cycle will continue based on patient motion model 410. In some embodiments, multiple motion models 410 can be used to describe different possible time-related motions. Previous patient geometry 412 can then be used to determine which motion model (if any) is relevant to the observed nearby history, and the prediction in box 414 can include following the most highly relevant motion model.
[0132] After obtaining an estimate of patient motion (changes in patient geometry) at box 414, method 400 proceeds to decision box 416, which assesses whether the estimated patient motion (changes in patient geometry) over the next time step exceeds a motion management threshold. The motion management threshold at box 416 may represent the maximum acceptable displacement before desired intervention or modification of treatment parameters (e.g., manual intervention, manual modification, automatic cessation of treatment portions, etc.). In some embodiments, box 414 may determine aggregated values representing the overall motion in voxels corresponding to the tissue of interest, derived from the various voxels from which motion estimates were obtained.
[0133] As an illustrative example, the aggregated value may include the arithmetic sum of estimated displacements at each voxel of interest, or the aggregated value may include a weighted sum, where weights are assigned to voxels based on some suitable metric. In some embodiments, each voxel may be weighted (at least in part) based on its proximity to the center of the volume of interest (e.g., the volume of the target cancer or the OAR). In some embodiments, weights may be assigned to voxels representing motion of different tissues. For example, a voxel representing motion of the OAR volume may be given a different weight than a voxel representing motion of the target volume. In some embodiments, different weights may be assigned to different directions based on the orientation of the radiation source (e.g., radiation source 12) relative to the patient. For example, movement in a direction more closely aligned with the direction defined by the unit vector of the beam eye view (BEV) may be given a lower weight because patient movement in such a direction may involve a corresponding small amount of reshaping of the aperture (e.g., defined by the shaping of the MLC blade 35) through which radiation is delivered. Conversely, movement in directions more closely aligned with the lateral (e.g., orthogonal) direction of the BEV can be given higher weight, because tissue movement in such directions may result in a correspondingly larger adjustment to the aperture to deliver the same radiation distribution. In some embodiments, the motion management threshold for box 416 is 3 mm. In some embodiments, the threshold is 2 mm. In some embodiments, box 416 is not required. In some such embodiments, box 328 ( Figure 6 It can be assumed that active motion management is not required under any circumstances and that the processing strategy of machine learning 113,π is sufficient to handle the full range of expected motion in patients.
[0134] If the evaluation of decision box 416 is positive, this may indicate that active motion management is expected in the current state and method 400 terminates. If the evaluation of decision box 416 is negative, then there is no indication of active motion management in the current state and method 400 terminates.
[0135] Return to method 300 ( Figure 6 The execution of box 328 of method 400 described above results in a current observation of the predicted patient movement and a determination of whether proactive movement management based on the predicted movement is desired. In some embodiments, multiple thresholds are considered at decision box 416. For example, different proactive management actions may be taken based on the estimated patient movement exceeding an increased threshold. In some embodiments, the estimated movement exceeding a maximum threshold indicates that the patient is moving beyond expected limits and that radiation treatment automatically delivered by the AI agent should be terminated (at least temporarily terminated) until the patient's movement can be resolved.
[0136] In some embodiments, only observations about the predicted patient movement are generated at box 328, without assessing whether active movement management is desired. In other words, the assessment of decision box 416 of method 400 can be omitted, and the output of box 328 (method 400) can include the predicted target movement in box 414. In some such embodiments, consideration of whether active movement management should be prescribed can be deferred to method 100 (method 400). Figure 4 Box 130, where predicted patient movement can be incorporated into machine learning-based processing strategies 113, π and box 125, which process states s to perform action a, are included in the expected reward. In some such embodiments, processing strategies 113, π specify action a that takes into account such estimated movement.
[0137] Box 328 is optional. In some embodiments, method 300 (box 125) does not involve predicting motion, but instead generates the current processing state in box 330 based on machine state 324 and patient state 326, and patient motion is incorporated into a machine (e.g., reinforcement) learning-based processing policy 113, π.
[0138] Regardless of whether patient movement is predicted in optional box 328, method 300 ( Figure 6 The process proceeds to box 330, which includes determining the current processing state 127(s) based on predicted patient movement based on machine state 324, patient state 326, and optional box 328. t Current processing status: 127 (s) t It also indicates that by method 100 ( Figure 4 The current processing state 127 (s) is caused by box 125. t Based on the observations and inputs from previous steps in method 300, the processing state 127(s) is determined at box 330. t ) may include and / or be based on some or all of the following components:
[0139] Cumulative observations representing the progress of radiation treatment up to the current time step for a specific patient include:
[0140] ○ Cumulative processing time, which may include the total elapsed time in the current processing section, the total elapsed time in the current and previous processing sections, and the elapsed processing time of each previous processing section;
[0141] ○ Cumulative delivered dose, which may include and / or based on 3D dose distribution reconstruction (e.g., patient images based on MLC, gantry, and radiation source parameters and tissue types that may define various voxels), fluence maps of the MLC and gantry configuration of delivered radiation in previous time steps, dose-volume histograms, dose-volume information, and volume-dose type aggregated dose information; and / or
[0142] ○ A measure of the physiological effects of delivered radiation based on the dose delivered in the current and / or previous sections (e.g., a model-based measure).
[0143] Current observations of the machine's current status, including measurements or parameters regarding the intensity of the radiation source, MLC blade position, MLC orientation, MLC jaw position, bench angle, bed position and orientation, etc.
[0144] Observation of the current patient status, including measurements and / or sensor outputs regarding patient position and / or movement (patient geometry) and dose distribution information and / or physiological effects information regarding the target volume and OAR volume.
[0145] The estimated patient motion (changes in patient geometry) and indications of whether active motion management is desired at subsequent time steps.
[0146] In some example embodiments, various additional components of the processing state can be updated at block 330 to reflect the current processing state 127(s) by performing some or all of the following steps. t ):
[0147] 1. Update the 3D CT image reconstruction of the volume of interest;
[0148] 2. Update 3D target volume and OAR volume depiction / segmentation;
[0149] 3. By updating the dose matrix and / or physiological effect matrix, the previous cumulative delivered dose distribution is converted into the currently reconstructed 3D CT image; and
[0150] 4. Using the updated 3D CT images, current machine status measurements, and dose calculation engine, calculate the sum of the dose delivered at the current time step and the dose matrix from step 3.
[0151] In method 300 ( Figure 4Following the conclusion of box 125 of method 100, method 100 proceeds to box 130, which involves an AI agent using a strategy 113,π determined by machine learning in box 112 to determine the next action 133(a). The action 133(a) selected by the AI agent may include software instructions readable by the radiation delivery device to manipulate or update the available set of radiation delivery parameters. The action 133(a) selected by the AI agent may include the updated radiation delivery parameters themselves. As a non-limiting example, the set of radiation delivery parameters may include the beam intensity of the radiation source (which includes the possibility of having an intensity value of 0 or off), beam on-time, MLC blade position, MLC orientation, bench angle, bed position, etc. Machine learning (e.g., reinforcement learning) processes strategy 113,π to operate on recommending action 133(a) within an action space, which represents all manipulable radiation delivery parameters for a particular radiation delivery device, such as those discussed herein. As discussed earlier in this paper, the action 133(a) selected by the AI agent in box 130 can typically include the action 133(a) that maximizes the cumulative reward for all future states.
[0152] If the current processing state of box 125 is 127(s) t If the indication is that active movement management is needed, then appropriate actions can be determined as part of box 130. There are many non-limiting possible scenarios regarding the estimated patient movement. A first potential scenario includes estimated patient movement being less than a movement management threshold (such as in decision box 416). Figure 7 The threshold used in )). In this scenario, there is no need to change the current processing state 127(s). t And can be based on the processing strategy 113 of box 112, π, and the current processing state 127 in box 125. t To select the appropriate action 133(a). A second potential scenario includes the estimated patient movement being greater than the movement management threshold (e.g., in decision box 416(a)). Figure 7 The threshold used in )). In this scenario, the current processing state 127(s) can be optionally changed. t One or more components of the CT image can be used to reflect this movement. For example, a new or updated transformation or image reconstruction can be applied to one or more planning images and / or reconstructed 3D CT images to reflect an estimated change in the patient geometry of the volume of interest. By doing so, the action 133(a) determined at box 130 can appropriately account for the impact on the processing state 127(s) due to substantial patient movement. t The expected changes in the dose distribution delivery are to achieve a more optimized dose delivery. As discussed above, the use of active motion management is optional. In some embodiments, the action 133(a) selected in block 125 is based on the processing strategy 113 of block 112, π, and the current processing state 127(s).t The determination was made without considering active exercise management.
[0153] In some embodiments, the constraints limiting “illegal” actions are known to the processing planning system 25C (e.g., through user input, through predefined operating parameters, etc.), thereby representing undesirable parameter changes given the current state to limit the available action space from which the AI agent can select actions. Examples of possible constraints limiting illegal actions in a given state include:
[0154] Radiation source 12 cannot travel beyond the maximum distance between consecutive time steps. This can be achieved, either entirely or partially, by imposing maximum variation on any motion axis between consecutive time steps. Individual constraints can be provided for each motion axis. For example, a maximum angular variation can be specified for the platform angle, a maximum displacement variation can be provided for the bed translation, etc.
[0155] The parameters affecting the beam shape must not vary beyond a specified amount over consecutive time steps. For example, a maximum value can be specified for the positional variation of the blade 36 of the MLC 35 or the variation of the rotational orientation of the MLC 35.
[0156] For each unit change in the motion axis, the parameters affecting the beam shape cannot change by more than a specified amount. For example, for each degree of rotation of the test stand 16 around axis 18, a maximum value can be specified for the positional change of the blades 36 of the MLC 35.
[0157] The source intensity variation between time steps must not exceed a specified amount.
[0158] For each unit change in motion axis, the source intensity must not change by more than a specified amount.
[0159] The source intensity cannot exceed a certain level.
[0160] As discussed elsewhere in this paper, constraints can be obtained in box 110 and can be incorporated into the optimization process as soft constraints (e.g., part of the reward function, objective function, etc.) or hard constraints. In addition to being incorporated into the action selection process in box 130 (or as an alternative), some or all of these constraints can be incorporated into the training (machine learning) process in box 112.
[0161] Imposing constraints can help reduce overall processing time by taking into account the physical limitations of a particular radiation delivery device. For example, if a particular radiation delivery device has a maximum radiation output rate, and the AI agent selects actions that result in a radiation intensity exceeding that maximum radiation output rate, then the movement rate of the radiation delivery device's axis of motion will have to slow down to deliver the specified intensity. Therefore, a constraint imposed on the maximum source intensity can force a solution where the specified intensity is within the capabilities of the radiation delivery device, and the movement axis of the radiation delivery device does not have to slow down. Since the movement axis does not have to slow down, such a solution can be delivered to the object S relatively quickly, resulting in a corresponding reduction in overall processing time. Those skilled in the art will appreciate that other constraints can be used to take into account other limitations of a particular radiation delivery device and can be used to reduce overall processing time.
[0162] A non-restrictive example of how such constraints can be limited is that the following parameters should not change by more than a specified amount between any two consecutive time steps:
[0163] Strength – 10%;
[0164] MLC blade position – 5mm;
[0165] MLC Orientation Angle —5%;
[0166] The angle of the test stand is 1 degree; and
[0167] Bed position - 3mm.
[0168] Once the appropriate action 133(a) is determined in block 130, process 114 proceeds to block 135, where action 133(a) is performed by an appropriate processing delivery device (such as radiation delivery device 10). Figure 1 This can be achieved by an AI agent (such as AI agent 25). As discussed elsewhere in this paper, action 133(a) can be performed by an AI agent (such as AI agent 25). Figure 1 This is achieved by providing a set of radiation delivery parameters to the radiation delivery device, thereby enabling the radiation delivery device to perform actions defined by the radiation delivery parameters.
[0169] In some embodiments, the radiation delivery device is configured such that its radiation source moves intermittently between time steps. As an illustrative and non-limiting example, in an embodiment where the radiation source moves intermittently, action 133(a), selected in block 130 and implemented in block 135, may specify the bench angle, MLC blade position, MLC orientation angle, MLC jaw position, etc. The radiation delivery device is operated to move according to these parameters, and upon completion, the radiation source can deliver radiation at the intensity and beam-on time specified by action 133. In some embodiments, other additional or alternative radiation delivery parameters may form part of action 133(a).
[0170] In some embodiments, the radiation delivery device is configured such that its radiation source moves continuously along a trajectory. In such embodiments, the radiation intensity specified by action 133(a) is typically not delivered to the object from a static gantry angle, but rather continuously delivered to the entire portion of trajectory 30 as the radiation source moves according to the current action 133(a). The radiation output rate of radiation source 12 can be adjusted by radiation delivery device 10 and control system 23 such that the radiation dose for the target volume at that time step conforms to that specified action 133(a). In some embodiments, the radiation output rate of the radiation source can be one of the radiation delivery parameters provided to the radiation delivery device by the AI agent as part of the action 133(a) to be implemented. The radiation output rate is typically determined by the amount of time required for the position of radiation source 12 and the shape of the radiation beam to change between intervention time steps. In some embodiments, processing strategy 113, π can be constrained to select action 133(a) having parameters that allow the movement of the radiation delivery device to be continuous between adjacent time steps.
[0171] In some embodiments, action 133(a), determined at block 130 and implemented at block 135, takes into account the inherent latency in any software operations performed by the processing planning system 25C and / or the radiation delivery device 10. For example, when the AI agent performs a state update according to method 300, there may be latency in data acquisition from various sensors of the radiation delivery device 10, latency in 3D image reconstruction, and latency in updating the dose distribution matrix. In some embodiments, such latency may be taken into account as part of the training of the AI agent 25 at block 112.
[0172] After performing the prescribed action 133(a) at box 135, method 100( Figure 4The process proceeds to decision box 140, where it assesses whether the treatment objective of box 110 has been met. If the assessment at box 140 is positive and the treatment objective has been met, then method 100 terminates. If the assessment at box 140 is negative because the treatment objective of box 110 has not been met, then method 100 performs another iteration of process 114 by determining the updated treatment state at box 125 in a subsequent time step. In some embodiments, decision box 140 assesses whether the treatment objective of the current treatment portion has been met within clinically acceptable tolerability. In such embodiments, the treatment state 125 of the preceding iteration represents the termination state of the current treatment portion and can be used as the initial state in subsequent treatment portions (e.g., treatment state 316). Figure 6 )).
[0173] Those skilled in the art can choose from a number of possible training (machine learning) algorithms at block 112 of method 100 for training the AI agent to develop the treatment strategy 113,π. Certain forms of learning may have relevance and specific application in the field of radiation treatment (more specifically, in the field of IMRT) to the task of training the AI agent. For example, transfer learning is a machine learning method in which discoveries gained in learning to perform a task can help accelerate training and improve learning in related but different tasks. In some embodiments, transfer learning can be employed at block 112 to provide a pre-existing AI agent and treatment strategy as a starting point for training the AI agent for a specific patient. In such embodiments, a pre-existing model, for example, with close correspondences in treatment factors such as target volume, OAR volume, and desired dose distribution, is preferably used. In other embodiments, a general AI agent can be trained for a large population of patients receiving similar radiation treatments. Such an agent can be used to deliver radiation therapy to patients without training specific to any particular patient and may involve only optional patient-specific validation.
[0174] A common feature of the various training algorithms discussed in this paper is the provision of a simulated processing environment. This simulated processing environment includes digital representations of the radiation delivery device and the patient (which can be a general patient or a specific patient). In some embodiments, the digital representation of the radiation delivery device is merely a validator that checks the feasibility of the requested machine state from the kinematic equations modeled over various components of the processing device (e.g., gantry 16, bed 15, MLC 35, etc.). In such embodiments, simulated observations of the machine state (e.g., in block 324 of method 300) can include current axis values of various processing device components. In some embodiments, the simulated processing device represents an idealized model of the apparatus used during processing. In other embodiments, real-world inaccuracies inherent in the machine can be incorporated into the simulated processing device, such as anticipated inaccuracies from backlash or finite acceleration constraints.
[0175] The simulation processing apparatus may additionally include a simulated radiation source to deliver radiation doses to a simulated patient. Various software packages and algorithms for simulating medical linear accelerators are known in the art. For example, EGSnrc, BEAM, and GEANT4 are exemplary software toolkits used to perform Monte Carlo simulations of ionizing radiation transport through matter. To simulate the aperture defined by the blades of the MLC, software libraries and toolkits can be used to further calculate phase-space data of the radiation beam passing through the MLC aperture, such as data provided by the PRIMO project.
[0176] A digital representation of the patient can be implemented in a variety of possible ways. In some embodiments, a 3D CT image set (such as that obtained in box 124) and a patient motion model can form a digital representation of the patient. In other embodiments, the digital representation of the patient may include a 4D CT image set to represent temporal changes in patient motion. Simulated observation of the patient's state (e.g., in box 324 of method 300) may include applying patient movement from the patient motion model to the 3D CT images using any suitable technique (such as deformable reconstruction). Other sensor readings embodied in the observation of the patient's state can be obtained by simulating sensor readings and / or appropriate processing. Exemplary examples of this include a reconstructed projection depicting a target volume or a simulated location of a marker box. Several other simulated observations suitable for a particular simulation processing scenario are possible and should be apparent to those skilled in the art. Any suitable software tool can be used to simulate the absorbed dose during the simulation processing. For example, the EGSnrc toolkit mentioned above contains functionality for calculating absorbed dose.
[0177] Terminology Explanation
[0178] Unless the context explicitly requires otherwise, throughout the specification and claims:
[0179] • "Including" and so on should be interpreted as inclusive, not exclusive or exhaustive; that is, it means "including but not limited to".
[0180] • “Connection,” “coupling,” or any variation thereof, means any direct or indirect connection or coupling between two or more elements; the coupling or connection between elements can be physical, logical, or a combination thereof.
[0181] • When used to describe this specification, the words “here,” “this article,” “above,” “below,” and similar terms should be used to refer to this specification as a whole, and not to any particular part of this specification;
[0182] • When referring to a list containing two or more items, “or” encompasses all of the following interpretations of the word: any item in the list, all items in the list, and any combination of items in the list;
[0183] • The singular forms “one,” “a,” and “the” also include the meaning of any appropriate plural form.
[0184] The directional terms used in this specification and any appended claims (if any), such as “vertical,” “lateral,” “horizontal,” “upward,” “downward,” “forward,” “backward,” “inward,” “outward,” “vertical,” “lateral,” “left,” “right,” “front,” “backward,” “top,” “bottom,” “lower,” “above,” “below,” etc., depend on the specific orientation of the described and illustrated device. The subject matter described herein can take various alternative orientations. Therefore, these directional terms are not strictly defined and should not be interpreted narrowly.
[0185] Embodiments of the present invention may be implemented using specially designed hardware, configurable hardware, or a programmable data processor configured with software (which may optionally include "firmware") capable of execution on a data processor, a special-purpose computer, or a combination of two or more of these, specifically programmed, configured, or constructed to perform one or more steps of the methods explained in detail herein. Examples of specially designed hardware are: logic circuits, application-specific integrated circuits ("ASICs"), large-scale integrated circuits ("LSIs"), very large-scale integrated circuits ("VLSIs"), etc. Examples of configurable hardware are: one or more programmable logic devices, such as programmable array logic ("PALs"), programmable logic arrays ("PLAs"), and field-programmable gate arrays ("FPGAs"). Examples of programmable data processors are: microprocessors, digital signal processors ("DSPs"), embedded processors, graphics processors, math coprocessors, general-purpose computers, server computers, cloud computers, mainframe computers, computer workstations, etc. For example, one or more data processors in the control circuitry of a device may implement the methods described herein by executing software instructions in a processor-accessible program memory.
[0186] Processing can be centralized or distributed. In distributed processing, information, including software and / or data, can be stored centrally or distributed. Such information can be exchanged between different functional units via communication networks, such as local area networks (LANs), wide area networks (WANs) or the Internet, wired or wireless data links, electromagnetic signals, or other data communication channels.
[0187] For example, while procedures or boxes may be presented in a given order, alternative examples may execute routines with steps, or employ systems with boxes in different orders, and some procedures or boxes may be deleted, moved, added, subdivided, combined, and / or modified to provide alternatives or sub-combinations. Each of these procedures or boxes may be implemented in a variety of different ways. Furthermore, while procedures or boxes are sometimes shown to be executed sequentially, these procedures or boxes may alternatively be executed in parallel, or may be executed at different times.
[0188] Furthermore, although the elements are sometimes shown to be executed sequentially, they may alternatively be executed simultaneously or in a different order. Therefore, it is intended that the following claims be interpreted to include all such variations within their intended scope.
[0189] Software and other modules may reside on servers, workstations, personal computers, tablet computers, image data encoders, image data decoders, PDAs, color grading tools, video projectors, audiovisual receivers, displays (such as televisions), digital cinema projectors, media players, and other devices suitable for the purposes described herein. Those skilled in the art will appreciate that aspects of the system can be implemented using other communication, data processing, or computer system configurations, including: internet devices, handheld devices (including personal digital assistants (PDAs)), wearable computers, various cellular or mobile phones, multiprocessor systems, microprocessor-based or programmable consumer electronics (e.g., video projectors, audiovisual receivers, displays such as televisions, etc.), set-top boxes, color grading tools, network PCs, microcomputers, mainframe computers, etc.
[0190] This invention can also be provided in the form of a program product. The program product may include any non-transitory medium carrying a set of computer-readable instructions that, when executed by a data processor, cause the data processor to perform the methods of this invention. The program product according to the invention can take many forms. The program product may include, for example, non-transitory media such as magnetic data storage media including floppy disks, hard disk drives, optical data storage media including CD-ROMs and DVDs, electronic data storage media including ROMs, flash RAMs, EPROMs, hard-wired or pre-programmed chips (e.g., EEPROM semiconductor chips), nanotechnology memories, and the like. Computer-readable signals on the program product may optionally be compressed or encrypted.
[0191] In some embodiments, the present invention may be implemented in software. For clarity, "software" includes any instructions that execute on a processor and may include (but is not limited to) firmware, resident software, microcode, etc. As those skilled in the art will appreciate, both the processing hardware and software may be wholly or partially centralized or distributed (or a combination thereof). For example, the software and other modules may be accessed via local memory, via a network, via a browser or other applications in a distributed computing context, or via other means suitable for the purposes described above.
[0192] In the context of the foregoing reference to components (e.g., software modules, processors, accessories, devices, circuits, etc.), unless otherwise stated, reference to such component (including reference to “means”) should be interpreted as including any component that performs the function of the described component as an equivalent of such component (i.e., a functional equivalent), including components that are not structurally equivalent to components that perform the functions of the disclosed structures illustrated in the exemplary embodiments of the present invention.
[0193] In the context of references to records, fields, entries, and / or other elements of a database, unless otherwise stated, such references should be interpreted as including multiple records, fields, entries, and / or other elements, as the case may be. Such references should also be interpreted as including a portion of one or more records, fields, entries, and / or other elements, as the case may be. For example, for the purposes of the foregoing description and the following claims, multiple “physical” records in a database (i.e., records encoded in the database structure) may be considered as a single “logical” record, even if the multiple physical records include information excluded from the logical record.
[0194] For illustrative purposes, specific examples of systems, methods, and apparatus have been described herein. These are merely examples. The techniques provided herein can be applied to systems other than those illustrated above. Many changes, modifications, additions, omissions, and substitutions are possible in the practice of this invention. This invention includes variations of the described embodiments that are apparent to those skilled in the art, including variations obtained by: replacing features, elements, and / or actions with equivalent features, elements, and / or actions; mixing and matching features, elements, and / or actions from different embodiments; combining features, elements, and / or actions from the embodiments described herein with features, elements, and / or actions from other technologies; and / or omitting combined features, elements, and / or actions from the described embodiments.
[0195] Various features are described herein as existing in “some embodiments.” Such features are not mandatory and may not be present in all embodiments. Embodiments of the invention may include zero, any, or any combination of two or more such features. This is limited to the extent that some of such features are incompatible with the others among these features, i.e., it is impossible for someone of ordinary skill in the art to construct an actual embodiment combining these incompatible features. Therefore, the description of “some embodiments” having feature A and “some embodiments” having feature B should be interpreted as explicitly indicating that the inventors also contemplate embodiments combining features A and B (unless the description otherwise states or features A and B are not compatible at all).
[0196] Therefore, it is intended that the appended claims and the claims described below be interpreted to include all such modifications, substitutions, additions, omissions, and sub-combinations that can be reasonably inferred. The scope of the claims should not be limited to the preferred embodiments set forth in the examples, but should be given the broadest interpretation consistent with the entire description.
Claims
1. A system for delivering a radiation dose from a radiation source to a patient, the system comprising: One or more movable elements; as well as Artificial intelligence (AI) agent, which is configured as follows: Define a machine state, which includes the position of the one or more movable elements, and wherein, when the radiation source is activated, the machine state defines the characteristics of the radiation emitted by the radiation delivery system; as well as For each of the multiple time steps in the radiation delivery section: Receive a set of observations regarding the machine's state and the patient's geometry; The patient's current treatment status is determined at least in part based on the observation set; The next action, including the next machine state for the next time step, is determined at least in part based on the current processing state and an artificial intelligence (AI) strategy, which is determined a priori using a machine learning process performed on training data. The next machine state for the next time step is defined by the characteristics of the radiation to be emitted by the radiation delivery system in the subsequent time step. as well as The system is made to perform the next action, thereby reaching the next machine state in the subsequent time step; The AI agent is configured to determine the patient's current treatment status based at least in part on the observation set by one of the following methods: The AI agent determines the current treatment state of the patient based at least in part on the previous treatment state of the patient in the portion identified as the previous time step; It is determined that there is no previous time step in the radiation delivery segment and that the patient has already been irradiated in the previous radiation delivery segment, and the current treatment status of the patient is determined by the AI agent based at least in part on the previous treatment status determined at the end of the previous radiation delivery segment; as well as It is determined that there is no previous time step in the radiation delivery segment and the patient has not been irradiated in the previous radiation delivery segment, and the current treatment state of the patient is determined by the AI agent at least in part based on the defined initial treatment state.
2. The system of claim 1, wherein the one or more movable elements comprise one or more of the following: a plurality of aperture defining elements, the positions of the plurality of aperture defining elements collectively defining an aperture through which radiation from the radiation source can be guided; and a radiation source moving element, the position of the radiation source moving element defining a directional orientation of radiation from the radiation source.
3. The system of claim 1, wherein the AI agent is configured to determine the current processing state of the patient based at least in part on patient geometry, the patient geometry being determined using image data obtained from the patient after the previous radiation delivery portion.
4. The system according to claim 1, 2, or 3, wherein the patient's current treatment status includes: The estimated geometry of a voxel of interest within the patient's body, the estimated geometry being at least in part based on the observations of the patient's geometry.
5. The system of claim 4, wherein the AI agent is configured to determine the estimated geometry based at least in part on the observations of the geometry of the patient and one or more patient movement models.
6. The system of claim 5, wherein the one or more patient movement models include a model of changes in patient geometry due to respiration.
7. The system of claim 5, wherein the one or more patient movement models include a model that predicts changes in the geometry of voxels inside the patient's body based on changes in the geometry of the external geometry of the patient's body.
8. The system of claim 5, wherein the one or more patient movement models are based at least in part on multiple images of the patient obtained over a period of time prior to the radiation delivery portion.
9. The system according to any one of claims 1 to 3, wherein the AI agent is configured to determine the current treatment status of the patient by determining an estimated cumulative dose absorbed by the target tissue during the radiation delivery portion.
10. The system of claim 9, wherein determining the current treatment status of the patient comprises: Based on the observations regarding the patient's geometry, update the 3D image reconstruction of the volume of interest within the patient's body; The cumulative dose absorbed by the target tissue in the previous time step is converted into an updated 3D image reconstruction; as well as The additional dose delivered to the target tissue at the current time step is estimated based at least on the updated 3D image reconstruction, the observation of the machine state at the current time step, and the dose estimation engine. Determining the estimated cumulative dose absorbed by the target tissue during the radiation delivery phase includes: determining the sum of the converted cumulative dose from the previous time step and the additional dose.
11. The system according to any one of claims 1 to 3, wherein the AI agent is configured to determine the current treatment status of the patient by determining an estimated cumulative dose absorbed by the patient's non-target organs during the radiation delivery portion.
12. The system of claim 9, wherein the AI agent is configured to determine the current treatment status of the patient by determining the cumulative treatment time during the radiation delivery portion.
13. The system of any one of claims 1 to 3, wherein the AI strategy includes mapping, and wherein determining the next action, including the next machine state for the subsequent time step, by the AI agent includes: The mapping is used to create a correspondence from the current processing state to the next action.
14. The system of claim 13, wherein in creating the correspondence from the current processing state to the next action, the mapping is based on maximizing a reward function that maximizes the cumulative reward obtained over all expected subsequent time steps.
15. The system according to any one of claims 1 to 3, configured to impose a set of one or more constraints on the AI agent at each of the plurality of time steps in the radiation delivery portion, the set of one or more constraints restricting the option space available to the AI agent for determining the next action.
16. The system of claim 15, wherein the set of one or more constraints comprises one or more of the following: The maximum distance that the radiation source can travel between the current time step and the subsequent time step; The maximum distance that the one or more movable elements can travel between the current time step and the subsequent time step; The maximum change in the intensity of the radiation source between the current time step and the subsequent time step; as well as The maximum value of the intensity of the radiation source.
17. The system according to any one of claims 1 to 3, wherein the AI agent is trained using a reinforcement-based machine learning process together with the training data to determine the AI policy.
18. A method for determining a radiation dose delivered to a patient using a radiation delivery device, the method comprising: A radiation delivery device is provided, the radiation delivery device comprising a radiation source and one or more movable elements; The machine state is defined, including the position of the one or more movable elements, and wherein, when the radiation source is activated, the machine state defines the characteristics of the radiation emitted by the radiation delivery device; For each of the multiple time steps in the radiation delivery section: Received at an artificial intelligence (AI) agent, which includes a processor configured to execute software instructions, a set of observations regarding the machine's state and the patient's geometry; The AI agent determines the patient's current treatment status based at least in part on the observation set; The AI agent determines the next action, including for the next machine state for the next time step, based at least in part on the current processing state and an artificial intelligence (AI) strategy, which is determined a priori using a machine learning process performed on training data. The next machine state for the next time step is defined by the characteristics of the radiation to be emitted by the radiation delivery device in the next time step. The determination of the patient's current treatment status by the AI agent, at least in part, based on the observation set, includes one of the following: The AI agent determines the current treatment state of the patient based at least in part on the previous treatment state of the patient in the portion identified as the previous time step; It is determined that there is no previous time step in the radiation delivery segment and that the patient has already been irradiated in the previous radiation delivery segment, and the current treatment status of the patient is determined by the AI agent based at least in part on the previous treatment status determined at the end of the previous radiation delivery segment; as well as It is determined that there is no previous time step in the radiation delivery segment and the patient has not been irradiated in the previous radiation delivery segment, and the current treatment state of the patient is determined by the AI agent at least in part based on the defined initial treatment state.
19. A method for determining a radiation dose to be delivered to a patient using a radiation delivery device, the method comprising: A radiation delivery device is provided, the radiation delivery device comprising a radiation source and one or more movable elements; The machine state is defined, including the position of the one or more movable elements, and wherein, when the radiation source is activated, the machine state defines the characteristics of the radiation emitted by the radiation delivery device; For each of the multiple time steps in the radiation delivery section: Received at an artificial intelligence (AI) agent, which includes a processor configured to execute software instructions, a set of observations regarding the machine's state and the patient's geometry; The AI agent determines the patient's current treatment status based at least in part on the observation set; The AI agent determines the next action, including for the next machine state for the next time step, based at least in part on the current processing state and an artificial intelligence (AI) strategy, which is determined a priori using a machine learning process performed on training data. The next machine state for the next time step is defined by the characteristics of the radiation to be emitted by the radiation delivery device in the next time step. as well as The device is made to perform the next action, thereby reaching the next machine state in the subsequent time step; The determination of the patient's current treatment status by the AI agent, at least in part, based on the observation set, includes one of the following: The AI agent determines the current treatment state of the patient based at least in part on the previous treatment state of the patient in the portion identified as the previous time step; It is determined that there is no previous time step in the radiation delivery segment and that the patient has already been irradiated in the previous radiation delivery segment, and the current treatment status of the patient is determined by the AI agent based at least in part on the previous treatment status determined at the end of the previous radiation delivery segment; as well as It is determined that there is no previous time step in the radiation delivery segment and the patient has not been irradiated in the previous radiation delivery segment, and the current treatment state of the patient is determined by the AI agent at least in part based on the defined initial treatment state.
20. The method according to claim 18 or 19, comprising: Operate the system according to any one of claims 1 to 17.
Citation Information
Patent Citations
Methods and Apparatus For the Planning and Delivery of Radiation Treatments
US20080226030A1
Radiotherapy treatment plan modeling using generative adversarial networks
US20190333623A1