4d motion tracking of anatomical structures

The 4D modeling using DNeRF and deformation models addresses limitations in 2D tracking by providing accurate, patient-agnostic motion tracking of anatomical structures, enhancing radiation therapy precision and adaptability.

WO2025163055A1PCT designated stage Publication Date: 2025-08-07SIEMENS HEALTHINEERS INTERNATIONAL AG
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
PCT/EP2025/052374
Authority / Receiving Office
WO · WO
Patent Type
Applications
Current Assignee / Owner
Priority Date
2024-01-30
Filing Date
2025-01-30
Publication Date
2025-08-07

AI Technical Summary

Technical Problem

Existing two-dimensional motion tracking methods for anatomical structures in radiation therapy are limited by high intensity structures obscuring targets, low visibility, high noise, and patient dependency, lacking robustness and accuracy in tracking moving tumors or organs at risk.

Method used

A 4D modeling approach using Dynamic Neural Radiance Fields (DNeRF) with a deformation model to generate a 4D model of anatomical structures by determining a movement vector based on a static 3D model, incorporating attenuation coefficients from offline projection images and surrogate signals for precise motion tracking.

Benefits of technology

Enables accurate, robust, and patient-agnostic motion tracking of anatomical structures, allowing for adaptive radiation therapy plans by generating a 4D model that accounts for tissue type and movement, improving dose deposition assessment and therapeutic precision.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure EP2025052374_07082025_PF_FP_ABST
    Figure EP2025052374_07082025_PF_FP_ABST
Patent Text Reader

Abstract

An imaging method (7000) for tracking motion of an anatomical structure in a subject, the method (7000) comprising: determining (7020) a static 3D model (450) of a 3D volume in the subject, the 3D volume including the anatomical structure, each coordinate of the static 3D model (450) associated with an attenuation coefficient (µ) representing a radiodensity of the coordinate, wherein the attenuation coefficient (µ) of each coordinate is obtained from offline projection images of the subject; receiving (7030) an input 3D model of the 3D volume constructed from acquired projection images of the 3D volume in the subject; determining (7040), by a deformation model (410), a movement vector describing a translation of coordinates in the input 3D model with respect to corresponding coordinates in the static 3D model (450); adding a time dimension to the static 3D model (450) by applying the movement vector to the static 3D model (450) to generate a 4D model.
Need to check novelty before this filing date? Find Prior Art

Description

4D MOTION TRACKING OF ANATOMICAL STRUCTURESTECHNICAL FIELD

[0001] This invention relates to techniques for tracking the motion of human or animal anatomical structures of interest using radioimaging, and in particular, to the use of 4D modelling implemented by machine learning techniques for motion tracking.BACKGROUND

[0002] Radiation therapy is a localised treatment for a specific target tissue (a planning target volume), such as a cancerous tumour. Ideally, radiation therapy is performed on the planning target volume that spares the surrounding normal tissue from receiving doses above specified tolerances, thereby minimising risk of damage to healthy tissue. Prior to the delivery of radiation therapy, an imaging system is typically employed to provide a three- dimensional image of the target tissue and surrounding area. From such imaging, the size and mass of the target tissue can be estimated, a planning target volume determined, and an appropriate treatment plan generated.

[0003] So that the prescribed dose is correctly supplied to the planning target volume (i.e. , the target tissue) during radiation therapy, the patient should be correctly positioned relative to the linear accelerator that provides the radiation therapy. Typically, dosimetric and geometric data are checked before and during the treatment, to ensure correct patient placement and that the administered radiotherapy treatment matches the previously planned treatment. This process is referred to as image guided radiation therapy (IGRT), and involves the use of an imaging system to view target tissues immediately before or while radiation treatment is delivered to the planning target volume. IGRT incorporates imaging coordinates from the treatment plan to ensure the patient is properly aligned for treatment in the radiation therapy device.

[0004] Tracking the motion of anatomical structures is useful for many purposes. For example, during treatment, motion tracking is useful so that therapeutic radiation may be applied only when the target volume is within a certain distance of the position the target volume was in during the planning stage. Such gating of therapeutic radiation is important to avoid irradiating healthy tissues when the tumour moves during therapy. Another usefulapplication of motion tracking is to assess where the dose was actually deposited during the treatment. That is, instead of assuming that the dose was deposited at the tumour in accordance with the treatment plan, motion tracking provides information of where the dose was deposited given movement of the target and given movement of the organs at risk. In general, the delivery of the treatment plan may be modified using motion tracking of the target in order to adapt the plan to the moving tumour or the moving organs at risk. However, it is noted that the motion tracking of the present disclosure may be performed in itself, and does not necessarily have to be performed during treatment of a tumour.

[0005] However, most prior art technology developed to track a tumour is based on two- dimensional approaches, such as template matching using X-rays acquired e.g. during the treatment. For example, a cross corellation between (i) a template containing a projection of a tumour and (ii) a real-time acquired 2D projection, may be used to establish the area in the 2D real-time projection that is most similar to the template. This can then be assumed to be the new position of the moving tumour.

[0006] Two-dimensional approaches such as these are limited by the lack of information in 2D projections. High intensity structures, such as bones, may obscurate or occlude the target. Further, low visibility and high noise prohibit the robustness of 2D approaches. Also, these approaches usually are patient dependent, meaning that the model is fitted to each patient treatment, and knowledge at the patient population level is not used.SUMMARY OF THE INVENTION

[0007] In accordance with a first aspect of the invention, there is provided an imaging method for tracking motion of an anatomical structure in a subject, as defined by claim 1 .

[0008] In accordance with a second aspect of the invention, there is provided a computer system as defined by claim 36.

[0009] In accordance with a third aspect of the invention, there is provided a computer program product as defined by claim 37.

[0010] In accordance with the first aspect of the invention, there is provided an imaging method for tracking motion of an anatomical structure in a subject, the method comprising: determining a static 3D model of a 3D volume in the subject, the 3D volume including the anatomical structure, each coordinate of the static 3D model associated with an attenuationcoefficient representing a radiodensity of the coordinate, wherein the attenuation coefficient of each coordinate is obtained from offline projection images of the subject; receiving an input 3D model of the 3D volume constructed from acquired projection images of the 3D volume in the subject; determining, by a deformation model, a movement vector describing a translation of coordinates in the input 3D model with respect to corresponding coordinates in the static 3D model; adding a time dimension to the static 3D model by applying the movement vector to the static 3D model to generate a 4D model.

[0011] Advantageously, the present disclosure provides a method to produce a 4D model of a moving anatomical structure inside a subject. This presents an improvement in traditional techniques that use 2D projection data in conjunction with 2D template matching to identify the position of a tumour or organ in the 2D projection data.

[0012] The 4D model of the present disclosure may, in one example, be based on a Dynamic Neural Radiance field DNeRF. In particular, the static 3D model may be generated using a canonical network of a DNeRF while the movement vector may be generated using a deformation network of the DNeRF.

[0013] The movement vector may be based on a deformation of an input 3D model with respect to the static 3D model. The input 3D model may be generated using newly acquired 2D projection images, by for example, standard reconstruction techniques.

[0014] Deformation

[0015] Optionally, determining the movement vector comprises a non-rigid deformation process. In particular, the deformation process may involve stretching, bending, twisting or compressing in a flexible and non-linear manner.

[0016] Optionally, the deformation model is a first machine learning model, such as a deep learning neural network. For example, the deformation model may comprise a convolutional neural network or a multilayer perceptron. The deformation model may be trained using data from similar past subjects (e.g. with similar tumour locations and / or tumour shape) or using past data obtained from a current subject. For example, data from a previous fraction may be used to train the deformation model and then newly acquired imaging data may be used to update the training of the deformation model neural network.

[0017] Optionally, the deformation model is determined using 3D Gaussian splatting. 3D Gaussian splatting is an imaging technique for rendering 3D models by representing the contents of the model with a set of 3D Gaussian primitives. A 3D Gaussian primitive is a mathematical representation of a volumetric entity in 3D space, modelled as a Gaussian distribution. 3D Gaussian splatting provides an alternative way of implementing the present 4D modelling that avoids the need for neural networks.

[0018] Optionally, the deformation model receives, as input, one or more surrogate signals, wherein the one or more surrogate signals provide further information for calculating the deformation such as tidal volume, tidal airflow, breathing phase and / or respiratory motion signals.

[0019] Advantageously, a surrogate signal provides additional information to the deformation model that aids the deformation model in determining the movement vector. Any signal that aids in determining the temporal dimension for combination with the 3D static model may be termed a surrogate signal.

[0020] Optionally, the surrogate signal is used to uniquely identify each deformation. The unique identification of each deformation is important in accurate construction of the 4D model. For example, it is important to associate each deformation with the correct time frame. A surrogate signal, such as tidal volume (see below), may be used to provide this association.

[0021] Optionally, the imaging method further comprises adding a random code to the surrogate signal to uniquely identify the deformation. A random code is an alternative way to associate each deformation with the correct time frame, by uniquely identifying each deformation.

[0022] Optionally, the random code comprises a random latent code and the method further comprises optimising the random latent code used to identify the deformation using Generative Latent Optimisation techniques. Generative Latent Optimisation is a machine learning technique for generative modelling by optimising latent space representations to reconstruct data.

[0023] Optionally, at least one surrogate signal is provided by a neural network encoder; the neural network encoder takes as input: (i) a real-time 2D projection image of the 3D volume, and (ii) meta-data of the real-time 2D projection image comprising a time stamp;and the neural network encoder outputs the surrogate signal in dependence on the realtime 2D projection image and associated meta-data. This provides an alternative method to determine a surrogate signal. The alternative method comprises using a time stamp or a newly acquired 2D projection image as data that provides temporal information useful in determining a movement vector.

[0024] Static 3D model

[0025] Optionally, the static 3D model is determined using a second machine learning model, such as a deep learning network. For example, the static 3D model may comprise a convolutional neural network or a multilayer perceptron. The static 3D model may be trained using data from similar past subjects (e.g. with similar tumour locations and / or similar tumour shapes) or using past data obtained from a current subject. For example, data from a previous fraction may be used to train the static 3D model.

[0026] Optionally, the static 3D model is determined using 3D Gaussian splatting. The 3D Gaussian splatting may be similar to the 3D Gaussian splatting used by the deformation model, as discussed above.

[0027] Optionally, the static 3D model maps input spatial coordinates to an attenuation coefficient at each spatial coordinate. The attenuation coefficient at a particular coordinate depends on the extent to which a hypothetical imaging beam passing through tissue at the coordinate is attenuated by the tissue at that coordinate. The extent to which the imaging beam is attenuated, in turn, depends on the radiodensity of the tissue at the coordinate. Thus, the attenuation coefficient can be used to infer the tissue type at the coordinate. Upon determining a 4D model of attenuation coefficient, it is possible to segment the 4D model to identify tissue types associated with the attenuation coefficients, thereby making it possible to track the motion of the anatomical structure in the 4D model.

[0028] Optionally, the imaging method further comprises applying positional encoding to the inputs of the static 3D model to map low-dimensional input coordinates into a higher dimensional space to model higher-frequency content in the static 3D model, such as sharp edges or fine texture.

[0029] As explained in more detail below, the static 3D model may comprise a canonical network of a D-NeRF. In general, NeRF networks are not suited to learn high-frequency details in input data. Positional encoding overcomes this limitation by mapping the low-dimensional input coordinates (e.g. 3D spatial coordinates x,y,z) into a higher-dimensional space with both low and high frequency components. The input coordinates are transformed into a set of periodic functions using sine and cosine functions with varying frequencies. In the present disclosure, use may be made of six exponentially increasing frequencies, although a different number of frequencies may be used.

[0030] Dynamic Neural Radiance Fields

[0031] Optionally, the combination of the static 3D model and the deformation model comprises a (modified) dynamic neural radiance field neural network, D-NeRF. While conventional D-NeRFs map 3D spatial coordinates, viewing angle and time to colour and volumetric density at each coordinate, the modified version of the D-NeRF used in the present invention simply maps 3D spatial coordinates to attenuation coefficients.

[0032] In particular, although D-NeRF aims to model viewing angle, colour and volumetric density, these are not used in the present modified D-NeRF network. This is because in 4D motion tracking it is not colour and volumetric density that is in question, but rather attenuation coefficient. Colour and volumetric density cannot be used to track motion of an anatomical structure but attenuation coefficient, as discussed above, is effective in tracking motion of anatomical structures. Further, as attenuation coefficient does not depend on viewing angle, there is no need to model viewing angle either.

[0033] Optionally, determining a movement vector is performed by a deformation network of the D-nERF; and said attenuation coefficient of each coordinate of the static 3D model is obtained by means of a canonical network of the D-nERF that maps input spatial coordinates to attenuation coefficients corresponding to each input spatial coordinate.

[0034] Typical D-NeRF architectures use a combination of a (i) radiance field that maps a 3D coordinate and viewing direction with colour and density at a location, and (ii) a dynamic component that adds a time component to the radiance field. The static 3D model of the present disclosure may be compared to the radiance field of a D-NeRF while the deformation model of the present disclosure may be compared to the time component of the D-NeRF.

[0035] Using past patient data

[0036] Optionally, the deformation model is determined using first data from one or more past subjects. In particular, the training dataset used to develop the deformation model may include data from past subjects with similar anatomies (such as tumour location) to that of a current subject.

[0037] Optionally, the first data is processed using a convolutional neural network encoder to extract features from the first data prior to determining the deformation model. Such a neural network encoder may be configured to transform high-dimensional input data into a compact, lower-dimensional representation such as a latent vector. The encoder may be used to extract the key features from the training data prior to using the data to train the deformation model.

[0038] Optionally, the convolutional neural network encoder for extracting features from the first data is a 4D ResNet34 network. A ResNet34 network comprises a deep residual neural network with 34 layers using residual connections. A residual connection skips one or more layers in the network and adds the input of those layers directly to their output. This aids computation by focussing on the learning of differences, or residuals, between the input and output, instead of learning the entire transformation. The 4D ResNet34 network is an adaptation of the ResNet34 network to process 4D data.

[0039] Optionally, the 3D static model is determined using second data from one or more past subjects. In particular, the training dataset used to develop the 3D static model may include data from past subjects with similar anatomies (such as tumour location) to that of a current subject.

[0040] Optionally, the second data is processed using a convolutional neural network encoder to extract features from the second data prior to determining the 3D static model. Such a neural network encoder may be configured to transform high-dimensional input data into a compact, lower-dimensional representation such as a latent vector. The encoder may be used to extract the key features from the training data prior to using the data to train the 3D static model.

[0041] Optionally, the convolutional neural network encoder for extracting features from the second data is a 3D ResNet34 network. A 3D ResNet34 network is an adaptation of the ResNet34 network discussed above.

[0042] Optionally, the output of each encoder is a pixel-aligned feature grid, wherein the pixel-aligned feature grid comprises a 3D grid of coordinates, each coordinate having features aligned with corresponding 2D image pixels. A feature grid is a 3D voxel grid that stores feature vectors at each grid point; here, the feature grid is pixel aligned, meaning that the 3D features are directly associated with 2D pixel features extracted from input images. The alignment ensures that features in the 3D space correspond accurately to 2D image data.

[0043] Tracking using attenuation coefficient

[0044] The attenuation coefficient associated with each coordinate in the static 3D model depends on the tissue represented by the coordinate. Different tissue types have different radiodensities, which results in different tissue types attenuating an imaging beam at different levels of attenuation. This means that each coordinate, which is associated with a different tissue type, is associated with a different radiodensity and therefore is associated with a different attenuation coefficient.

[0045] Optionally, the imaging method further comprises tracking the anatomical structure in the subject by identifying an attenuation coefficient associated with tissue of the anatomical structure and labelling coordinates having the identified attenuation coefficient as belonging to the anatomical structure. In other words, the attenuation coefficient is used to segment the anatomical structure in the 4D image, by associating different attenuation coefficients with different tissue types. In particular, the anatomical structure being tracked will have a particular attenuation coefficient and by knowing the value of this attenuation coefficient, the motion of the anatomical structure may be tracked using the 4D model.

[0046] Density correction

[0047] Optionally, the imaging method further comprises: identifying a tissue type associated with a coordinate of the static 3D model; correcting the attenuation coefficient of the coordinate based on the identified tissue type. In other words, in some cases the attenuation coefficient may need to be corrected for certain tissue types and this may be realised by including a density correction module applied to the output of the static 3D model. The density correction model takes as input the attenuation coefficient, segmentation information, and surrogate signals in order to determine any necessary correction to the attenuation coefficient.

[0048] Optionally, said tissue type is identified by a neural network. That is, segmentation information and tissue type information may be determined by a neural network that takes attenuation coefficient as input and outputs a tissue type associated with the attenuation coefficient.

[0049] Training

[0050] Optionally, the static 3D model and the deformation model are trained using 2D projections acquired before or during tracking. In particular, the 4D model may be used to generate 2D projection images by computer simulation, and these generated 2D projection images may be used in conjunction with ground truth 2D projection images to train the static 3D model and / or the deformation model.

[0051] Optionally, the training comprises: using x-ray physics to acquire a simulated 2D projection image from the 4D model; acquiring a real 2D projection image; determining a loss between the simulated and real 2D projection images using a loss function, and using the loss to train the static 3D model and the deformation model using backpropagation. For example, the loss function may be used to determine or modify the weights of neural connections in the static 3D model and / or the deformation model, wherein the modification is selected to reduce the value of the loss function. Minimising the value of the loss function leads to an optimal set of weights, an optimal static 3D model and an optimal deformation model.

[0052] Optionally, the training comprises using existing projection data generated for one or more past subjects as training datasets. For example, 4D models generated for past subjects with a similar anatomy or a similar tumour may be used to generate simulated 2D projection data. These may be compared with 2D projection images generated using the current 4D model and any differences may be quantified in a loss function. The loss function can then be used to modify the weights of neural networks associated with the 3D static model and the deformation model, using backpropagation.

[0053] In some cases, the training is performed based on the data of a population of past patients but where the data is patient-agnostic. Such data may be processed by a neural network encoder prior to training of or application to the deformation model and / or the static 3D model.

[0054] Optionally, the training comprises using a past 4D model generated for a current subject or past subjects, as ground truth. For example, instead of comparing 2D projection images from a past 4D model with simulated 2D projection images from a current 4D model, the 4D models may be compared directly, in order to calculate a loss function for use in training the neural networks using backpropagation. The simulated 2D projections may not be necessary when a past 4D model is available for direct comparison with the current 4D model.

[0055] Grouping into phases

[0056] Optionally, the imaging method further comprises grouping temporal stages of the 4D model into phases, such as breathing phases. Advantageously, different time intervals of the 4D model may be classified as belonging to different stages of a physiological cycle such as the breathing cycle. For example, this may classify different time intervals of the 4D model as relating to a resting phase, forced inspiration phase, forced expiration phase and transition phases.

[0057] Optionally, the imaging method further comprises using a surrogate signal to map a time index to a phase, such as breathing phases, in order to perform said grouping of temporal stages of the 4D model into phases. Such surrogate signals may include tidal volume or tidal airflow.

[0058] Optionally, the canonical network learns a tumour mask and the deformation network computes a movement vector of the tumour mask to track motion of the anatomical structure. Such a tumour mask effectively segments and distinguishes the tumour from the rest of the 4D model, allowing motion tracking of the tumour. Here, the canonical network provides two functions - the first is to map spatial coordinates to attenuation coefficients, while the second is to segment the 3D static model using the values of the attenuation coefficients. The deformation model may then use the masked 3D volume in computing the movement vector of the tumour.

[0059] In some cases, in addition to comprising attenuation coefficients, the static 3D model may comprise a reference 3D mask that includes segmentation information. The DVF or the movement vector may then be applied to the reference 3D mask. During real-time imaging, the static 3D model is frozen such that only the movement vector is updated.

[0060] Optionally, the output of the deformation model is transformed to comprise a low rank matrix. For example, if the output of the deformation network comprises a matrix (e.g. the deformation network output comprises a value of a movement vector at each coordinate, wherein each coordinate represents an element of the matrix), then the matrix may be transformed into one of low rank (the rank of a matrix is the number of linearly independent rows or columns in the matrix). For example, the deformation network may be split into N low-rank matrices that each produces an independent output. The outputs of the N low- rank matrices may then be combined, for example in a weighted sum, to produce the DVF or movement vector. The weights may be computed using a time-to-motion embedding paradigm, in which a neural network encoder may encode temporal dynamics or motion related features of the moving anatomical structure into a latent space. Advantageously, this may simplify the development of the deformation network.

[0061] After transformation of the output of the deformation model to a low rank matrix, temporal components of the model may be computed using surrogate signals and / or an encoder.

[0062] Optionally, the output of the deformation model is decomposed into spatially variant components and temporally variant components, wherein temporal components are predicted by one or more surrogate signals optionally provided by an encoder. That is, the transformation into a low rank matrix is performed by said decomposition. Advantageously, this simplifies the generation of the movement vector by the deformation model.

[0063] Other aspects of the invention

[0064] In accordance with the second aspect of the invention, there is provided a computer system comprising a memory and a processor, the computer system configured to perform the imaging method as defined above.

[0065] In accordance with a third aspect of the invention, there is provided a computer program product comprising a non-transitory, computer readable storage medium encoded with instructions operable for execution by a processor to perform the imaging method as defined above.

[0066] In accordance with a fourth aspect of the invention, there is provided means for tracking motion of an anatomical structure in a subject, the means configured to perform a method comprising: determining a static 3D model of a 3D volume in the subject, the 3Dvolume including the anatomical structure, each coordinate of the static 3D model associated with an attenuation coefficient representing a radiodensity of the coordinate, wherein the attenuation coefficient of each coordinate is obtained from offline projection images of the subject; receiving an input 3D model of the 3D volume constructed from acquired projection images of the 3D volume in the subject; determining, by a deformation model, a movement vector describing a translation of coordinates in the input 3D model with respect to corresponding coordinates in the static 3D model; adding a time dimension to the static 3D model by applying the movement vector to the static 3D model to generate a 4D model.BRIEF DESCRIPTION OF THE DRAWINGS

[0067] For a better understanding of the present disclosure, some embodiments will now be described by way of example with reference to the accompanying drawings.

[0068] FIG. 1 is a perspective view of a radiation therapy system that can beneficially implement various aspects of the present disclosure.

[0069] FIG. 2 schematically illustrates a drive stand and gantry of the radiation therapy system of FIG. 1, according to various embodiments.

[0070] FIG. 3 schematically illustrates a drive stand and a gantry of the radiation therapy system of FIG. 1, according to various embodiments.

[0071] FIGS. 4A-4D illustrate various exemplary arrangements that use a deformation network and a canonical network to produce a 4D model.

[0072] FIG. 4E illustrates an exemplary arrangement of a deformation network.

[0073] FIG. 5 illustrates density correction of the attenuation coefficient.

[0074] FIGS. 6A-6C illustrate various arrangements to train the 4D model using loss functions and backpropagation.

[0075] FIG. 7 is a flowchart showing an exemplary method of the present disclosure.

[0076] FIG. 8 is an illustration of a computing device configured to perform various embodiments of the present disclosure.

[0077] FIG. 9 is a block diagram of an illustrative embodiment of a computer program product for implementing one or more embodiments of the present disclosure.DETAILED DESCRIPTION

[0078] In the following detailed description, reference is made to the accompanying drawings, which form a part hereof. In the drawings, similar symbols typically identify similar components, unless context dictates otherwise. The illustrative embodiments described in the detailed description, drawings, and claims are not meant to be limiting. Other embodiments may be utilised, and other changes may be made, without departing from the scope of the subject matter presented here. It will be readily understood that the aspects of the disclosure, as generally described herein, and illustrated in the figures, can be arranged, substituted, combined, and designed in a wide variety of different configurations, all of which are explicitly contemplated and make part of this disclosure.

[0079] FIG. 1 is a perspective view of a radiation therapy system 100 that can beneficially implement various aspects of the present disclosure. Radiation therapy (RT) system 100 is a radiation system configured to track motion of anatomical structures in real time or near- real time using X-ray imaging techniques. RT system 100 is configured to provide stereotactic radiosurgery and precision radiotherapy for lesions, tumours, and conditions anywhere in the body where radiation treatment is indicated. As such, RT system 100 can include one or more of a linear accelerator (LINAC) that generates a megavolt (MV) treatment beam of high energy X-rays, one or more kilovolt (kV) X-ray sources, one or more X-ray imagers, and, in some embodiments, an MV electronic portal imaging device (EPID). By way of example, radiation therapy system 100 is described herein configured with a circular gantry. In other embodiments, radiation therapy system 100 can be configured with a C-gantry capable of infinite rotation via a slip ring connection.

[0080] Generally, RT system 100 is capable of motion tracking of a target volume during application of an MV treatment beam, so that an IGRT and / or an intensity-modulated radiation therapy (IMRT) process can be performed using X-ray imaging. However, the motion tracking may be performed in itself, and not necessarily during treatment. RT system 100 may include one or more touchscreens 101 , couch motion controls 102, a bore 103, a base positioning assembly 105, a couch 107 disposed on base positioning assembly 105, and an image acquisition and treatment control computer 106, all of which are disposed within a treatment room. RT system 100 further includes a remote control console 110,which is disposed outside the treatment room and enables treatment delivery and patient monitoring from a remote location. Base positioning assembly 105 is configured to precisely position couch 107 with respect to bore 103, and motion controls 102 include input devices, such as button and / or switches, that enable a user to operate base positioning assembly 105 to automatically and precisely position couch 107 to a predetermined location with respect to bore 103. Motion controls 102 also enable a user to manually position couch 107 to a predetermined location.

[0081] FIG. 2 schematically illustrates a drive stand 200 and gantry 210 of RT system 100, according to various embodiments. Covers, base positioning assembly 105, couch 107, and other components of RT system 100 are omitted in FIG. 2 for clarity. Drive stand 200 is a fixed support structure for components of RT treatment system 100, including gantry 210 and a drive system 201 for rotatably moving gantry 210. Drive stand 200 rests on and / or is fixed to a support surface that is external to RT treatment system 100, such as a floor of an RT treatment facility. Gantry 210 is rotationally coupled to drive stand 200 and is a support structure on which various components of RT system 100 are mounted, including a linear accelerator (LINAC) 204, an MV electronic portal imaging device (EPID) 205, an imaging X-ray source 206, and an X-ray imager 207. During operation of RT treatment system 100, gantry 210 rotates about bore 103 when actuated by drive system 201.

[0082] Drive system 201 rotationally actuates gantry 210. In some embodiments, drive system 201 includes a linear motor that can be fixed to drive stand 200 and interacts with a magnetic track (not shown) mounted on gantry 210. In other embodiments, drive system 201 includes another suitable drive mechanism for precisely rotating gantry 210 about bore 103. LINAC 204 generates an MV treatment beam 230 of high energy X-rays (or in some embodiments electrons, protons, and / or other heavy charged particles, ultra-high dose rate X-rays (e.g., for FLASH radiotherapy) or microbeams for microbeam radiation therapy) and EPID 205 is configured to acquire X-ray images with treatment beam 230. Imaging X-ray source 206 is configured to direct a conical beam of X-rays, referred to herein as imaging X-rays 231 , through an isocenter 203 of RT system 100 to X-ray imager 207, and isocenter 203 typically corresponds to the location of a target volume 209 to be treated. In the embodiment illustrated in FIG. 2, X-ray imager 207 is depicted as a planar device, whereas in other embodiments, X-ray imager 207 can have a curved configuration.

[0083] X-ray imager 207 receives imaging X-rays 231 and generates suitable projection images therefrom. According to certain embodiments, such projection images can then be employed to construct or update portions of imaging data for a digital volume that corresponds to a three-dimensional (3D) region that includes target volume 209. That is, a 3D image of such a 3D region is reconstructed from the projection images. In some embodiments, cone-beam computed tomography (CBCT) and / or digital tomosynthesis (DTS) can be used to process the projection images generated by X-ray imager 207. CBCT is typically employed to acquire projection images over a relatively long acquisition arc, for example over a rotation of 180° or more of gantry 210. As a result, a high-quality 3D reconstruction of the imaged volume can be generated. CBCT is often employed at the beginning of a radiation therapy session to generate a set-up 3D reconstruction. For example, CBCT may be employed immediately prior to application of treatment beam 230 to generate a 3D reconstruction confirming that target volume 209 has not moved or changed shape.

[0084] Such imaging may be performed during an IGRT or IMRT process in order to perform the 4D motion tracking of the present disclosure.

[0085] In the embodiment illustrated in FIG. 2, RT system 100 includes a single X-ray imager and a single corresponding imaging X-ray source. In other embodiments, RT system 100 can include two or more X-ray imagers, each with a corresponding imaging X- ray source. One such embodiment is illustrated in FIG. 3.

[0086] FIG. 3 schematically illustrates a drive stand 300 and gantry 310 of RT system 100, according to various embodiments. Drive stand 300 and gantry 310 are substantially similar in configuration to drive stand 200 and gantry 210 in FIG. 2, except that the components of RT system 100 that are mounted on gantry 310 include a first imaging X-ray source 306, a first X-ray imager 307, a second imaging X-ray source 308, and a second X-ray imager 309. In such embodiments, the inclusion of multiple X-ray imagers in RT system 100 facilitates the generation of projection images (for motion tracking of anatomical structures such as the target volume) over a shorter image acquisition arc. For instance, when RT system 100 includes two X-ray imagers and corresponding X-ray sources, an image acquisition arc for acquiring projection images of a certain image quality can be approximately half that for acquiring projection images of a similar image quality with a single X-ray imager and X-ray source.

[0087] The projection images generated by X-ray imager 207 (or by first x-ray imager 307 and second X-ray imager 309) are used to perform the 4D motion tracking of the present disclosure.

[0088] Image guided radiation therapy (IGRT) is used to treat tumours in areas of the body that are subject to voluntary movement, such as the lungs, or involuntary movement, such as organs affected by peristalsis, gas motion, muscle contraction and the like. IGRT involves the use of an imaging system to view target tissues (also referred to as the “target volume”) immediately before or while radiation treatment is delivered thereto. In IGRT, image-based coordinates of the target volume from a previously determined treatment plan are compared to image-based coordinates of the target volume determined immediately before or during the application of the treatment beam. In this way, changes in the surrounding organs at risk and / or motion or deformation of the target volume relative to the radiation therapy system can be detected. Consequently, dose limits to organs at risk are accurately enforced based on the daily position and shape, and the patient's position and / or the treatment beam can be adjusted to more precisely target the radiation dose to the tumour. For example, in pancreatic tumour treatments, organs at risk include the duodenum and stomach. The shape and relative position of these organs at risk with respect to the target volume can vary significantly from day-to-day. Thus, accurate adaptation to the shape and relative position of such organs at risk enables escalation of the dose to the target volume and better therapeutic results.4D motion tracking

[0089] The present disclosure is directed to the generation of a 4D model of the patient, wherein “4D” refers to three spatial dimensions and a time dimension. The 4D model can be learned from 2D views such as X-rays acquired during the treatment, as well as 4D-CT images which can be acquired before the treatment. This model can be used for multiple tasks, including tracking the tumour per se, visualising the 4DCT of the patient e.g. during the treatment, and determining the actual (rather than planned) dose deposition that happened during the treatment. The motion tracking of the present disclosure may be performed in itself, and actual treatment of a tracked tumour is optional.

[0090] The method trains a 3D static model that maps an input space representing spatial location and possible additional parameters to an output space representing the information necessary to render 2D views from any location and time. The method is based on amodified version of the neural architecture of Neural Radiance Fields (NeRF) and Dynamic Neural Radiance Fields (DNeRF).

[0091] NeRF is a method for representing scenes as neural radiance fields for view synthesis. In NeRF, every coordinate or point in a 3D model is mapped to a parameter such as colour. In simulation, rays may be passed through the 3D model to obtain a 2D view.

[0092] D-NeRF is also a method for synthesising novel 2D views, but at an arbitrary point in time, of dynamic scenes with complex non-rigid geometries. D-NeRF extends NeRF by adding an additional dimension to the input space to represent time. This allows D-NeRF to model dynamic scenes, where the scene can change over time.

[0093] In the case of natural scenes modelled by a NeRF and D-NeRF, the output space corresponds to the density and colour, however in the modified NeRF and D-NeRF of the present disclosure, rather than outputting density or colour, the output of the D-NeRF is the attenuation coefficient of Computed Tomography (“CT”) or X-rays. The attenuation coefficient corresponds to the radiodensity at a particular coordinate in the output space, which can be used to identify an anatomical structure, and which in turn can be used to track the motion of the anatomical structure.

[0094] To do this, D-NeRF learns two functions: 1) one that maps every coordinate in an input 3D space to an attenuation coefficient; this is referred to as the canonical model of the tracked anatomical structure, and 2) another that maps from the canonical model to a realtime 3D model at a particular time; this is referred to as a deformation model or network. The present model adapts the D-NeRF modelling technique to X-ray physics, that is, instead of predicting a density and a colour (RGB vector) for each time value, 3D position and view angle as in traditional D-NeRF, the present model determines an attenuation coefficient, i.e. , how much an X-ray beam is attenuated (reduced in intensity) as it passes through a medium with different material properties. Further, as the attenuation coefficient is independent of the view angle, there is no need for the view angle input parameters that are used in traditional D-NeRF.Network arrangements

[0095] FIG. 4A illustrates an exemplary apparatus for determining a 4D model of a moving anatomical structure in a human subject. The apparatus comprises two principal modules,illustrated here as a deformation network 410 and a canonical network 450. The canonical network is configured to determine a 3D static model of a volume in the subject comprising the moving anatomical structure. Being a static model, the 3D static model does not change over time. It may therefore be referred to as producing a “reference” model.

[0096] The purpose of the canonical network 450 is to map a set of input spatial coordinates (x, y, z) to an attenuation coefficient p. That is, by means of the canonical network 450, every coordinate is assigned a particular value comprising an attenuation coefficient. Such a mapping may be learned by a machine learning model such as a neural network.

[0097] For example, the canonical network 450 may be compared to a Neural Radiance Field neural network, which typically uses multi-layer perceptron networks. However, the present scheme uses a modified Neural Radiance Field. The inputs to a conventional Neural Radiance Field comprises a set of input spatial coordinates (x, y, z) and viewing direction whereas the outputs of a conventional Neural Radiance Field comprises color and volumetric density. In contrast, the inputs of the Neural Radiance Fields as used in the present disclosure comprises only a set of input spatial coordinates (x, y, z) and the outputs comprises attenuation coefficient.

[0098] The modified Neural Radiance Field of the present invention does not need viewing direction as an input because the attenuation coefficient is independent of viewing angle.

[0099] The output of the present modified Neural Radiance Field is attenuation coefficient because attenuation coefficient may be used to track the movement of an anatomical structure within the eventual 4D model. There is no need for outputting the colour or volumetric density of a coordinate because these cannot be used in tracking the anatomical structure.

[0100] The outputted attenuation coefficient of a particular coordinate is a measure of the extent to which a hypothetical imaging beam would be attenuated if it were to pass through the coordinate, which in turn depends on the radiodensity of the coordinate. Further, the tissue type of an anatomical structure being tracked may be associated with a certain known radiodensity, which in turn will be associated with a certain attenuation coefficient. Thus, by labelling coordinates of the 3D static model with an associated attenuation coefficient, it is possible to determine the location of the anatomical structure in the 3D static model. By determining the location of the anatomical structure in the 3D staticmodel over time, a 4D model of the anatomical structure is produced which can track the motion of the anatomical structure.

[0101] To produce a 4D model from the 3D static model provided by the canonical network 450, a deformation network 410 is introduced. The purpose of the deformation network 410 is to calculate a deformation vector field (DVF) or movement vector that reflects the deformation of the static spatial coordinates (x, y, z) of the 3D static model to a current spatial coordinate (x + Ax, y + Ay, z + Az). This is achieved by the deformation network 410 calculating the DVF or the movement vector (Ax, Ay, Az) associated with each coordinate in the 3D spatial coordinates. Such a DVF or movement vector may be learned by a machine learning model such as a neural network.

[0102] For example, the combination of the deformation network 410 and the canonical network 450 may be compared to a Dynamic Neural Radiance Field neural network, and the deformation network typically comprises a multi-layer perceptron network. The DVF or movement vector produced by the deformation network can be combined with the 3D static model produced by the canonical network 450 in order to produce a 4D dynamic model of the moving anatomical structure. That is, the 3D static model provides a “snapshot” of the anatomical structure at a fixed point in time, while the DVF provides a degree to which each spatial coordinate of an input 3D model moves with respect to the spatial coordinates of the 3D static model. The “snapshot” and the DVF or (movement vector) can be combined to produce the 4D model. As every spatial coordinate in the 4D model is associated with an attenuation coefficient, and a tracked anatomical structure is associated with a tissue type having a particular radiodensity that is proportional to the attenuation coefficient, it is possible to segment the 4D model so that the moving anatomical structure is discernible when the 4D model is displayed, for example on a computer screen.

[0103] To summarise FIG. 4A, the deformation network 410 provides a deformation vector field or movement vector that reflects the transform of a set of spatial coordinates at a first point in time to a set of spatial coordinates at a second point in time; the canonical network 450 provides a 3D static attenuation coefficient model of a volume in a subject comprising an anatomical structure to be tracked; and the DVF or movement vector field is combined with the 3D static model to provide a 4D model. The canonical network 450 comprises a modified Neural Radiance Field while the combination of the canonical network 450 and the deformation network 410 comprises a modified Dynamic Neural Radiance Field.

[0104] FIG. 4B shows the deformation network 410 and the canonical network 450 under an alternative arrangement. The arrangement of FIG. 4B is generally the same as FIG. 4A, except that the deformation network 410 and canonical network 450 of FIG. 4B may be trained using patient data from a population, rather than the current patient data from a single patient. In particular, the system of FIG. 4B includes patient databases 420, 422 from which multiple instances of patient specific data are extracted. The data from the patient database may be used to train the 4D model including the canonical network and the deformation network. In particular, the 4D model may be trained using 2D projection data obtained for past patients similar to the patient in question. The 4D model may then be fine-tuned or “overfitted” to the current patient during motion tracking.

[0105] The past patient data used to train the deformation network 410 comprises patient database 420 while the past patient data used to train the canonical network 450 comprises patient database 422. The past patient data may not be used directly to train the deformation network 410 and the canonical network 450; rather, the past patient data for the deformation network may first be processed using neural network encoder 412 while the past patient data for training the canonical network 450 may first be processed using neural network encoder 414. In this manner, the past patient training data is transformed into a compact, lower-dimensional representation or encoding of the data that captures the essential features or latent information present in the data.

[0106] The neural network encoders 412 and 414 may each comprise a fully convolutive neural network. For example, neural network encoder 412 may comprise a 4D Residual Network or 4D ResNet, while the neural network encoder 414 may comprise the first layers (e.g., the first four layers) of a 3D Residual Network with 34 layers, or 3D ResNet34. Using the 4DRes Net network, features of the past patient data may be extracted in the layers prior to a pooling layer, then upsampled and processed to be of a suitable format to be used to train the deformation network 410. The 3D ResNet34 network may be used in the same way as the 4DResNet.

[0107] Both in use and in training, the features or latent codes generated by the neural network encoders 412 and 414 may be concatenated with the other inputs of the deformation network 410 and the canonical network 450. Additionally or alternatively, the features or latent codes generated by the neural network encoders 412 and 414 may be incorporated as a residual at each layer of deformation network 410 and canonical network 450, respectively.

[0108] FIG. 40 shows the deformation network 410 and the canonical network 450 under a further alternative arrangement. The arrangement of FIG. 4C is generally the same as FIG. 4B, except for the use of a further neural network encoder 416 that may be used to provide temporal information to the deformation network 410. The neural network encoder 416 takes a 2D projection image 430 as input, and may also receive other input comprising meta-data of the 2D projection image. For example, suitable meta-data of the 2D projection image may comprise projection angle or the time at which the 2D projection image was taken.

[0109] The output of the neural network encoder 416 may comprise temporal information, referred to as a surrogate signal, that the deformation network may use to compute the DVF or the movement vector. The neural network encoder 416 may be a convolutional neural network ending with pooling operations or a fully connected layer configured to produce a scalar output usable in computing the DVF or movement vector. The neural network encoder 416 may be trained using 2D projection images simulated from 4DCTs and ground truth surrogate signals using a loss function between the surrogate signals of the neural network encoder 416 and the ground truth surrogate signals.

[0110] FIG. 4D shows the deformation network 410 and the canonical network 450 under a still further alternative arrangement. The arrangement of FIG. 4D is generally the same as FIG. 4A, except that the canonical network 450 is also capable of providing a segmented 3D static model, referred to as a 3D tumour mask 472. This is achieved by training the canonical network 450 using tumour segmentation done on the 3D / 4DCT acquired prior to treatment.

[0111] The deformation network 410 provides a DVF or a movement vector to both parts of the canonical network 450, wherein the first part of the canonical model 450 provides an attenuation coefficient while the second part of the canonical model 474 provides tumour segmentation (or segmentation of any anatomical structure the movement of which is being tracked). As discussed, by applying the DVF or movement vector to the static 3D model of attenuation coefficient, a 4D model of changing attenuation coefficients is obtained. However, this does not explicitly show the movement of the anatomical structure being tracked.

[0112] In the arrangement of FIG. 4D, by applying the DVF or movement vector to the segmented 3D static model 474, the movement of anatomical structures as segmentedin the model 472 may be visualised. This explicitly shows the movement of the anatomical structure being tracked.

[0113] FIG. 4E shows an alternative arrangement of the deformation network 410, which may be applied to any of the deformation network arrangements of FIGS. 4A-4D. The arrangement within which the deformation network 410 of FIG. 4E is used is generally similar to the arrangement of FIG. 4G, except that the arrangement of FIG. 4E includes a “Low rank linear”, or LR, module 480 to transform the output of the deformation network 410 to comprise a low rank matrix. For example, if the output of the deformation network comprises a matrix (e.g. the deformation network output comprises a value of a movement vector at each coordinate, wherein each coordinate represents an element of the matrix), then the low rank linear module 480 will transform this matrix into one of low rank (the rank of a matrix is the number of linearly independent rows or columns in the matrix).

[0114] For example, the deformation network 410 output may be split into N low- rank matrices that each comprises an independent output. The values of the N low-rank matrices may then be combined, for example in a weighted sum, to produce the DVF or movement vector. The weights may be computed using a time-to-motion embedding paradigm, in which a neural network encoder may encode temporal dynamics or motion related features of the moving anatomical structure into a latent space. Advantageously, this may simplify the development of the deformation network 410.Surrogate signals

[0115] In the field of Dynamic Neural Radiance Fields, surrogate signals are an additional source of information regarding the motion, deformation or temporal dynamics of a 3D volume. They serve as an additional input to the deformation network 410, enabling the deformation network to handle dynamic environments more effectively. For example, the surrogate signal may act as a “proxy” for motion-related information and any information that assists in determining the temporal changes in a 3D static model may be referred to as a surrogate signal.

[0116] That is, in a dynamic neural radiance field, generally time is an input of the deformation model 410. However, in the present disclosure, additional inputs, or surrogate signals, may be used as further inputs to the deformation model 410. Exemplary surrogate signals include tidal volume, tidal airflow, and respiratory motion signals. Tidal volume represents the amount of air moving into or out of the lungs during a normal, or resting,respiratory cycle. Tidal airflow refers to the flow of air into and out of the lungs during a normal, resting respiratory cycle. The surrogate signal may be any signal that can be used to infer the motion state of the motion tracked anatomical structure.

[0117] FIG. 4G discussed above depicts an arrangement of the deformation network 410 that uses a neural network encoder 416 to generate a surrogate signal. In general, the surrogate signals allow the connection of different views of the moving tracked anatomical structure with the knowledge that they are in a same or a similar motion state, and thus that their deformation vector field should also be similar. Thus, the surrogate signals assist the deformation network 410 in performing its role to apply the motion to a canonical state by predicting the deformation field, which transforms the anatomy or volume from the motion state to the one defined by the canonical network.

[0118] The plurality of surrogate signals may be considered to be a single surrogate signal with multiple dimensions, each dimension corresponding to a different type of surrogate signal. The combination of surrogate signal dimensions may be used to uniquely identify a motion state. A random latent code or the current time may be added or concatenated with the surrogate signal to achieve uniqueness in encoding the state of the deformation at a given time frame. The technique known as Generative Latent Optimisation (GLO) may be used to optimise this latent “deformation” code along with network weights during training in order to learn common motion related structures in the data not captured by other surrogate signals.Density correction

[0119] As discussed above, the output of the canonical network 450 comprises the attenuation coefficient at a particular spatial coordinate. However, it is possible that the attenuation coefficient at a particular coordinate requires correction in dependence on the anatomical structure that it represents. For example, the radiodensity of a particular tissue type at a particular coordinate may be used to correct the attenuation coefficient. This may be achieved by segmenting the output of the canonical network to identify the tissue type that each coordinate of the 3D static model represents. Then, in dependence on the radiodensity of the tissue type that a particular spatial coordinate represents, the outputted attenuation coefficient of the canonical model may be corrected.

[0120] FIG. 5 illustrates the use of a density correction module 500. The density correction module 500 takes as input the output of the canonical network 450, which is theattenuation coefficient at a particular spatial coordinate. The density correction module 500 also takes as input the segmentation information s of each spatial coordinate as well as some surrogate signals. The segmentation information s identifies the tissue type at each spatial coordinate while the surrogate signals provide information regarding the movement of the tissue. Using information regarding the radiodensity of the relevant tissue, the output of the canonical network 450 may be corrected by the density correction module 500 to provide a corrected attenuation coefficient output, p’.Training

[0121] FIG. 6A illustrates an arrangement of components that may be used to train the deformation network 410 and the canonical network 450. In general, the combination of the deformation network 410 and the canonical network 450 may be referred to as a 4D model 620. Using the principles of X-ray physics, 2D projection images are rendered 630 from the 4D model 620. That is, the transmission of simulated X-ray imaging beams from a simulated X-ray source, through a hypothetical imaged object model, to eventually impinge on a simulated X-ray imaging detector, may be modelled 630. This results in the generation of a set of simulated 2D projection images, such as a 4D model generated 2D projection image 650.

[0122] The simulated 2D projection images 650 may then be compared to actual acquired 2D X-ray projection images 600 of the tracked anatomical structure. This comparison is used to compute an error, such as a loss function 670, which is then used to train the 4D model using backpropagation. In other words, the error 670 between the simulated 2D projection images 650 and the actual 2D projection images 600 is used to update the 4D model 620 with the new information contained in the newly acquired 2D X- ray projection images 600.

[0123] Alternatively or additionally, the training process can be extended in order to use 2D, 3D or 4D images acquired during one or more pre-treatment sessions. For example, instead of using the actual acquired 2D projection images as the ground truth, the training process may use the images of previous treatment sessions as ground truth.

[0124] In some arrangements, the training process may use training data obtained from a population of past patients to train the 4D model. Such an arrangement is described above in the context of FIG. 4B.

[0125] FIG. 6B illustrates the training of the 4D model 620 where the deformation network 410 and the canonical network 450 are capable of providing a segmented 4D model, in similar fashion as described in the context of FIG. 4D. Similar to FIG. 6A, using the principles of X-ray physics, 2D CT views or 2D projection views are rendered 632 by simulating an X-ray imaging beam passing through a hypothetical object and being imaged on a simulated X-ray imager. Further, in FIG. 6B, again using the principles of X-ray physics, 2D mask views are rendered or simulated 634 by simulating modelled X-rays passing through a hypothetical imaged object model, whilst incorporating attenuation coefficient information to provide the segmentation.

[0126] Then, actual acquired 2D projection views are obtained and actual segmented 2D projection views are obtained, for example using the 2D projection images being used to track the motion of an anatomical structure. The simulated 2D projection views 650 and the simulated segmented 2D projection views 652 are then compared with the actual 2D projection views and the actual segmented 2D projection views in order to compute an error, such as a loss function 670. The loss function 670 may then be used to train the deformation network 410 and the canonical network 450, to enable the deformation network 410 and the canonical network 450 to provide an accurate 4D model and an accurate segmented 4D model.

[0127] FIG. 60 illustrates the training of the 4D model 620 without using the simulated 2D projection views of FIG. 6A and FIG. 6B. This is possible when ground truth 4DCT images already exist, for example generated from similar past patients, or from the current patient at an earlier time such as during an earlier treatment fraction. In this case, the 4D model 620 can be used to generate a 4DCT image, which can then be compared with the ground truth 4DCT image 624 to obtain a loss function 670, which can then be used to train the 4D model 620 using backpropagation.Overall method

[0128] FIG. 7 is a flowchart showing an exemplary method of the present disclosure.

[0129] In step 7020, a canonical network 450 of a modified Dynamic NeuralRadiance Field is used to determine a 3D static model of a 3D volume in a patient. The 3D volume comprises an anatomical structure the motion of which is being tracked. Establishing the 3D static model of the 3D volume involves mapping a set of input spatial coordinates (x, y, z) to an attenuation coefficient at each coordinate. The attenuationcoefficient represents the radiodensity of the 3D volume at a particular coordinate which is in turn representative of the tissue type at the coordinate.

[0130] In step 7030, an input 3D model is obtained from new 2D projection images acquired of the 3D volume in the patient. These new 2D projection images and the input 3D model include the anatomical structure being tracked.

[0131] In step 7040, a deformation network 410 of a modified Dynamic Neural Radiance Field is used to determine the deformation vector field or movement vector describing the translation of coordinates in the input 3D model with respect to corresponding coordinates in the static 3D model.

[0132] In step 7060, the 3D static model is combined with the deformation vector field or movement vector in order to generate a 4D model, which can be used to track the motion of the anatomical structure in the 3D volume of the patient.

[0133] In summary, the present invention provides a scheme to track the motion of anatomical structures in a patient by the use of 4D modelling based on the paradigm of Dynamic Neural Radiance Fields. A static 3D model that maps a set of input spatial coordinates to attenuation coefficients at each coordinate is established, and the 3D static model may be compared to a Neural Radiance Field. Then, a deformation network is used to determine a deformation vector field or movement vector describing the deformation of an input 3D model with respect to the static 3D model. The combination of the deformation network and 3D static model may correspond to a Dynamic Neural Radiance Field. The deformation vector field is applied to the static 3D model to generate a 4D model which shows the changing attenuation coefficients of the 4D model over time. By segmenting the 4D model in dependence on the attenuation coefficients, the anatomical structure may be identified in the dynamic 4D model and the movement of the anatomical structure may be tracked.

[0134] That is, the problem of motion tracking is approached using a global scene reconstruction paradigm as may be used in Dynamic Neural Radiance Fields.Computing device

[0135] FIG. 8 is an illustration of computing device 700 configured to perform various embodiments of the present disclosure. Computing device 700 may be a desktopcomputer, a laptop computer, a smart phone, or any other type of computing device suitable for practicing one or more embodiments of the present disclosure. It is noted that the computing device described herein is illustrative and that any other technically feasible configurations fall within the scope of the present disclosure.

[0136] As shown, computing device 700 includes, without limitation, an interconnect (bus) 740 that connects a processing unit 750, an input / output (I / O) device interface 760 coupled to input / output (I / O) devices 780, memory 710, a storage 730, and a network interface 770. Processing unit 750 may be any suitable processor implemented as a central processing unit (CPU), a graphics processing unit (GPU), an application-specific integrated circuit (ASIC), a field programmable gate array (FPGA), any other type of processing unit, or a combination of different processing units, such as a CPU configured to operate in conjunction with a GPU or digital signal processor (DSP). In general, processing unit 750 may be any technically feasible hardware unit capable of processing data and / or executing software applications, including the method of the present disclosure as illustrated by method 7000.

[0137] I / O devices 780 may include devices capable of providing input, such as a keyboard, a mouse, a touch-sensitive screen, and so forth, as well as devices capable of providing output, such as a display device and the like. Additionally, I / O devices 780 may include devices capable of both receiving input and providing output, such as a touchscreen, a universal serial bus (USB) port, and so forth. I / O devices 780 may be configured to receive various types of input from an end-user of computing device 700, and to also provide various types of output to the end-user of computing device 700, such as displayed digital images or digital videos. In some embodiments, one or more of I / O devices 780 are configured to couple computing device 700 to a network.

[0138] Memory 710 may include a random access memory (RAM) module, a flash memory unit, or any other type of memory unit or combination thereof. Processing unit 750, I / O device interface 760, and network interface 770 are configured to read data from and write data to memory 710. Memory 710 includes various software programs that can be executed by processor 750 and application data associated with said software programs, to perform the method of the present disclosure such as method 7000.

[0139] FIG. 9 is a block diagram of an illustrative embodiment of a computer program product 800 for implementing a method for motion tracking of anatomicalstructures, according to one or more embodiments of the present disclosure. Computer program product 800 may include a signal bearing medium 804. Signal bearing medium 804 may include one or more sets of executable instructions 802 that, when executed by, for example, a processor of a computing device, may provide at least the functionality described above with respect to FIGS. 4A to FIG. 7.

[0140] In some implementations, signal bearing medium 804 may encompass a non-transitory computer readable medium 808, such as, but not limited to, a hard disk drive, a Compact Disc (CD), a Digital Video Disk (DVD), a digital tape, memory, etc. In some implementations, signal bearing medium 804 may encompass a recordable medium 810, such as, but not limited to, memory, read / write (R / W) CDs, R / W DVDs, etc. In some implementations, signal bearing medium 804 may encompass a communications medium 806, such as, but not limited to, a digital and / or an analog communication medium (e.g., a fiber optic cable, a waveguide, a wired communications link, a wireless communication link, etc.). Computer program product 800 may be recorded on non-transitory computer readable medium 808 or another similar recordable medium 810.

[0141] The descriptions of the various embodiments have been presented for purposes of illustration, but are not intended to be exhaustive or limited to the embodiments disclosed. Many modifications and variations will be apparent to those of ordinary skill in the art without departing from the scope of the described embodiments.

[0142] Aspects of the present embodiments may be embodied as a computer system, method or computer program product. Accordingly, aspects of the present disclosure may take the form of an entirely hardware embodiment, an entirely software embodiment (including firmware, resident software, micro-code, etc.) or an embodiment combining software and hardware aspects. Furthermore, aspects of the present disclosure may take the form of a computer program product embodied in one or more computer readable medium(s) having computer readable program code embodied thereon.

[0143] Any combination of one or more computer readable medium(s) may be utilised. The computer readable medium may be a computer readable signal medium or a computer readable storage medium. A computer readable storage medium may be, for example, but not limited to, an electronic, magnetic, optical, electromagnetic, infrared, or semiconductor system, apparatus, or device, or any suitable combination of the foregoing. More specific examples (a non-exhaustive list) of the computer readable storage mediumwould include the following: an electrical connection having one or more wires, a portable computer diskette, a hard disk, a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM or Flash memory), an optical fibre, a portable compact disc read-only memory (CD-ROM), an optical storage device, a magnetic storage device, or any suitable combination of the foregoing. In the context of this document, a computer readable storage medium may be any tangible medium that can contain, or store a program for use by or in connection with an instruction execution system, apparatus, or device.

[0144] While various aspects and embodiments have been disclosed herein, other aspects and embodiments will be apparent to those skilled in the art. The various aspects and embodiments disclosed herein are for purposes of illustration and are not intended to be limiting, with the true scope being indicated by the following claims.

Claims

CLAIMS1. An imaging method for tracking motion of an anatomical structure in a subject, the method comprising: determining a static 3D model of a 3D volume in the subject, the 3D volume including the anatomical structure, each coordinate of the static 3D model associated with an attenuation coefficient representing a radiodensity of the coordinate, wherein the attenuation coefficient of each coordinate is obtained from offline projection images of the subject; receiving an input 3D model of the 3D volume constructed from acquired projection images of the 3D volume in the subject; determining, by a deformation model, a movement vector describing a translation of coordinates in the input 3D model with respect to corresponding coordinates in the static 3D model; adding a time dimension to the static 3D model by applying the movement vector to the static 3D model to generate a 4D model.

2. The imaging method of claim 1 wherein determining the movement vector comprises a non-rigid deformation process.

3. The imaging method of claim 1 or claim 2 wherein the deformation model is a first machine learning model, such as a deep learning neural network.

4. The imaging method of any preceding claim wherein the deformation model is determined using 3D Gaussian splatting.

5. The imaging method of any preceding claim wherein the deformation model receives, as input, one or more surrogate signals, wherein the one or more surrogate signals provide further information for calculating the deformation such as tidal volume, tidal airflow, breathing phase and / or respiratory motion signals.

6. The imaging method of claim 5 wherein the surrogate signal is used to uniquely identify each deformation.

7. The imaging method of claim 5 or claim 6 further comprising adding a random code to the surrogate signal to uniquely identify the deformation.

8. The imaging method of claim 7 wherein the random code comprises a random latent code and the method further comprises optimising the random latent code used to identify the deformation using Generative Latent Optimisation techniques.

9. The imaging method of any of claim 5 to claim 8 wherein: at least one surrogate signal is provided by a neural network encoder; the neural network encoder takes as input: (i) a real-time 2D projection image of the 3D volume, and (ii) meta-data of the real-time 2D projection image comprising a time stamp; and the neural network encoder outputs the surrogate signal in dependence on the realtime 2D projection image and associated meta-data.

10. The imaging method of any preceding claim wherein the static 3D model is determined using a second machine learning model, such as a deep learning network.

11. The imaging method of any preceding claim wherein the static 3D model is determined using 3D Gaussian splatting.

12. The imaging method of any preceding claim wherein the static 3D model maps input spatial coordinates to an attenuation coefficient at each spatial coordinate.

13. The imaging method of any preceding claim further comprising applying positional encoding to the inputs of the static 3D model to map low-dimensional input coordinates into a higher dimensional space to model higher-frequency content in the static 3D model, such as sharp edges or fine texture.

14. The imaging method of any preceding claim wherein the combination of the static 3D model and the deformation model comprises a dynamic neural radiance field neural network, D-NeRF.

15. The imaging method of claim 14 wherein:said determining a movement vector is performed by a deformation network of the D-nERF; and said attenuation coefficient of each coordinate of the static 3D model is obtained by means of a canonical network of the D-nERF that maps input spatial coordinates to attenuation coefficients corresponding to each input spatial coordinate.

16. The imaging method of any preceding claim, wherein the deformation model is determined using first data from one or more past subjects.

17. The imaging method of claim 16 wherein the first data is processed using a convolutional neural network encoder to extract features from the first data prior to determining the deformation model.

18. The imaging method of claim 17 wherein the convolutional neural network encoder is a 4D ResNet34 network.

19. The imaging method of any preceding claim, wherein the 3D static model is determined using second data from one or more past subjects.

20. The imaging method of claim 19 wherein the second data is processed using a convolutional neural network encoder to extract features from the second data prior to determining the 3D static model.

21. The imaging method of claim 20 wherein the convolutional neural network encoder is a 3D ResNet34 network.

22. The imaging method of any of claim 17, 18, 20 or 21 wherein the output of each encoder is a pixel-aligned feature grid, wherein the pixel-aligned feature grid comprises a 3D grid of coordinates, each coordinate having features aligned with corresponding 2D image pixels.

23. The imaging method of any preceding claim wherein the attenuation coefficient associated with each coordinate in the static 3D model depends on the tissue represented by the coordinate.

24. The imaging method of any preceding claim further comprising tracking the anatomical structure in the subject by identifying an attenuation coefficient associated with tissue of the anatomical structure and labelling coordinates having the identified attenuation coefficient as belonging to the anatomical structure.

25. The imaging method of any preceding claim further comprising: identifying a tissue type associated with a coordinate of the static 3D model; correcting the attenuation coefficient of the coordinate based on the identified tissue type.

26. The imaging method of claim 25 wherein said tissue type is identified by a neural network.

27. The imaging method of any preceding claim wherein the static 3D model and the deformation model are trained using 2D projections acquired before or during tracking.

28. The imaging method of claim 27 wherein the training comprises: using x-ray physics to acquire a simulated 2D projection image from the 4D model; acquiring a real 2D projection image; determining a loss between the simulated and real 2D projection images using a loss function, and using the loss to train the static 3D model and the deformation model using backpropagation.

29. The imaging method of claim 27 or claim 28 wherein the training comprises using existing projection data generated for one or more past subjects as training datasets.

30. The imaging method of claim 27 or claim 28 wherein the training comprises using a past 4D model generated for a current subject or past subjects as ground truth.

31. The imaging method of any preceding claim further comprising grouping temporal stages of the 4D model into phases, such as breathing phases.

32. The imaging method of claim 31 further comprising using a surrogate signal to map a time index to a phase, such as breathing phases, in order to perform said grouping of temporal stages of the 4D model into phases.

33. The imaging method of claim 15 wherein the canonical network learns a tumour mask and the deformation network computes a movement vector of the tumour mask to track motion of the anatomical structure.

34. The imaging method of any preceding claim wherein the output of the deformation model is transformed to comprise a low rank matrix.

35. The imaging method of claim 34 wherein the output of the deformation model is decomposed into spatially variant components and temporally variant components, wherein temporal components are predicted by one or more surrogate signals optionally provided by an encoder.

36. A computer system comprising a memory and a processor, the computer system configured to perform the imaging method as defined in any one of claims 1-35.

37. A computer program product comprising a non-transitory, computer readable storage medium encoded with instructions operable for execution by a processor to perform the imaging method as defined in any one of claims 1-35.

Citation Information

Patent Citations

  • Determining a Four-Dimensional CT Image Based on Three-Dimensional CT Data and Four-Dimensional Model Data

    US20150302608A1

  • Methods and systems for reconstructing a 3D anatomical structure undergoing non-rigid motion

    US20230129194A1