System and method for real-time multi-modal deformable image registration for image-guided intervention

By training MR-US and US-US deformation models, the warp field is updated in real time to register pre-interventional MRI and real-time US images, solving the challenge of image fusion during intervention and realizing real-time visualization of anatomical structures and precise placement of interventional devices.

CN120997261APending Publication Date: 2025-11-21GE PRECISION HEALTHCARE LLC +2
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510610516.6
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Priority Date
2024-05-20
Filing Date
2025-05-13
Publication Date
2025-11-21

AI Technical Summary

Technical Problem

Existing technologies struggle to effectively integrate MRI and real-time US images during interventional procedures, leading to movement and deformation due to the dynamic nature of anatomical structures, which affects the precise placement of interventional devices.

Method used

A deep learning-based deformation model is used to register pre-interventional MRI, CT, or PET images with real-time US images. By training MR-US and US-US deformation models, the warp field is updated in real time to maintain image alignment, combining the high contrast information of MRI with the real-time imaging capabilities of US.

Benefits of technology

It enables real-time visualization of anatomical structures during intervention, improves the accuracy and efficiency of interventional device placement, reduces computational burden, and adapts to patients' physiological changes.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120997261A_ABST
    Figure CN120997261A_ABST
Patent Text Reader

Abstract

Systems and methods are provided for real-time multi-modal deformable image registration for image-guided intervention. Pre-interventional three-dimensional (3D) magnetic resonance imaging (MRI) images are acquired for a patient and a plurality of 3D ultrasound (US) images are captured for various breathing states and gestures. The MRI image is registered to each US image using the trained MR-US deformation model (1000), thereby producing a deformed MRI image. During an intervention, an interventional 3D US image is acquired and registered to a pre-interventional US image using a trained US-US deformation model (1200), thereby determining a warpage field. The warpage field is applied to a corresponding deformed MRI image, producing a registered 3D MRI image for visualizing annotated tissue features from pre-interventional MRI on a live interventional US image. The disclosed approach makes full use of multi-modal imaging and deep learning models to enhance visualization during interventions by combining the superior soft tissue contrast of MRI with real-time US imaging capabilities.
Need to check novelty before this filing date? Find Prior Art

Description

[0001] Government support

[0002] This invention was made with government support under license number 1R01CA266879 granted by the National Institutes of Health. The government owns certain rights to the invention. Technical Field

[0003] This disclosure relates generally to medical imaging, and more specifically to systems and methods for real-time multimodal deformable image registration for image-guided interventions. Background Technology

[0004] Image-guided interventions are medical procedures that utilize imaging techniques to facilitate precise targeting of pathological tissues. These interventions typically employ imaging modalities such as computed tomography (CT), magnetic resonance imaging (MRI), and ultrasound (US) to visualize anatomical structures and guide the placement of interventional devices. A significant challenge in these procedures is the dynamic nature of internal anatomy, which can undergo movement and deformation due to patient movement, respiratory and cardiac cycles, and interactions with interventional instruments.

[0005] Ultrasound imaging is frequently chosen due to its real-time imaging capabilities and the flexibility it offers in accessing a wide range of anatomical regions. However, its limitations in soft tissue visualization and lesion saliency restrict its usefulness in guiding interventional procedures. While MRI offers superior soft tissue contrast and lesion detection without ionizing radiation, its integration into real-time interventional guidance is complex due to the computational demands of fusing MRI with real-time ultrasound images.

[0006] CT fluoroscopy guidance is another approach for abdominal interventions. However, this method is hampered by a slow workflow and limited tissue contrast, which is only slightly improved by the use of contrast agents. The use of contrast agents is further limited by patient dose constraints, thus reducing their utility in environments where optimized tissue contrast is required rapidly. In addition, the manual handling involved in some CT-US fusion-guided procedures is highly dependent on operator skill and often results in alignment inaccuracies due to the use of rigid registration methods.

[0007] While MRI provides superior lesion contrast, especially with contrast agents, fully utilizing this information for interventional guidance is not straightforward. The computational burden of fusing MRI with real-time US images during applicator insertion for guidance is detrimental to real-time applications. Existing workflows also rely heavily on manual intervention, with performance varying based on operator expertise. Alternative registration methods, such as analytical deformable registration, introduce significant computational time incompatible with the real-time requirements of interventional procedures. Summary of the Invention

[0008] In one embodiment, this disclosure provides a method for enhancing visualization of tissue features during an intervention by leveraging the complementary advantages of pre-interventional imaging modalities (such as MRI, CT, or positron emission tomography (PET)) and real-time US imaging. The method involves acquiring pre-interventional three-dimensional (3D) images of the imaging subject using one of the pre-interventional imaging modalities (MRI, CT, or PET). During the intervention, a first interventional 3D US image of the imaging subject is acquired. The pre-interventional 3D image is then registered with the first interventional 3D US image using a trained deformable model specific to the pre-interventional imaging modality and US (such as a trained MR-US deformable model for pre-interventional MRI images). This registration process produces a first deformed 3D image in the pre-interventional imaging modality aligned with the first interventional 3D US image. As the intervention progresses, subsequent interventional 3D US images are acquired. To maintain alignment between the pre-interventional imaging data and the real-time US imaging data, the first interventional 3D US image is registered with subsequent interventional 3D US images using a trained US-US deformable model. This registration defines a warp field that describes the transformation required to align the first interventional 3D US image with the subsequent interventional 3D US image, thus taking into account any anatomical movement or deformation that may have occurred between the acquisition of the two US images. This warp field is then applied to the first deformed 3D image in the pre-interventional imaging modality, producing a 3D image registered to the subsequent interventional 3D US image in the pre-interventional imaging modality. This registered 3D image enables real-time visualization of tissue features annotated on the pre-interventional 3D image on the subsequent interventional 3D US image during intervention, thereby fully utilizing the superior soft tissue contrast and lesion saliency of the pre-interventional imaging modality while maintaining the real-time imaging capabilities of the US. Compared to directly registering the pre-interventional 3D image to the subsequent US image, registering the first deformed pre-interventional 3D image to the subsequent US image using the warp field determined between the first and subsequent interventional 3D US images is computationally more efficient. This approach makes full use of the deformation field calculated between two interventional US images acquired at different time points, avoiding the computationally more expensive registration between the pre-interventional imaging modality (e.g., MRI) and the interventional US imaging modality, which typically exhibit significant differences in appearance and spatial characteristics.

[0009] The above-described advantages, as well as other advantages and features of this disclosure, will become apparent from the following detailed description when considered alone or in conjunction with the accompanying drawings. It should be understood that the above summary is provided to present a series of concepts further described in the detailed description in a simplified form. This is not intended to identify key or essential features of the claimed subject matter, the scope of which is uniquely defined by the claims following the detailed description. Furthermore, the claimed subject matter is not limited to specific implementations that address any of the shortcomings mentioned above or in any part of this disclosure. Attached Figure Description

[0010] This disclosure will be better understood by referring to the following description of non-limiting embodiments, in which:

[0011] Figure 1 This is a schematic diagram illustrating an imaging system for acquiring and processing 3D MRI and 3D US images according to an embodiment of the present disclosure;

[0012] Figure 2 This is a block diagram illustrating an image processing apparatus according to an embodiment of the present disclosure;

[0013] Figure 3 This is a schematic diagram illustrating an imaging subject undergoing pre-interventional MRI and US imaging according to an embodiment of the present disclosure to generate multiple 3D MRI images and multiple 3D US images, which can be used to fine-tune the MR-US deformation model and the US-US deformation model.

[0014] Figure 4 This is a flowchart illustrating a method for acquiring pre-interventional 3D MRI and US images according to an embodiment of the present disclosure;

[0015] Figure 5 This is a schematic diagram illustrating an imaging subject undergoing interventional 3D US imaging according to an embodiment of the present disclosure, showing the use of a trained US-US deformation model to visualize annotated pre-interventional MRI features on interventional 3D US images.

[0016] Figure 6 This is a flowchart illustrating a method for registering pre-interventional 3D MRI images to interventional 3D US images using a trained US-US deformable model, according to an embodiment of the present disclosure.

[0017] Figure 7 This is a flowchart illustrating an embodiment of the present disclosure of a method for acquiring 3D MRI and US images at selected locations and under respiratory conditions and storing these images in a training dataset for MR-US and US-US deformable models;

[0018] Figure 8 This is a flowchart illustrating a method for acquiring 3D MRI images for use in training an MR-US deformable model according to an embodiment of the present disclosure;

[0019] Figure 9 This is a flowchart illustrating a method for acquiring 3D US images for use in training MR-US and US-US deformable models according to an embodiment of the present disclosure;

[0020] Figure 10 This is a schematic diagram illustrating a process for training an MR-US deformable model according to an embodiment of the present disclosure;

[0021] Figure 11 This is a flowchart illustrating a method for training an MR-US deformable model according to an embodiment of the present disclosure;

[0022] Figure 12 This is a schematic diagram illustrating a process for training a US-US deformable model according to an embodiment of the present disclosure;

[0023] Figure 13 This is a flowchart illustrating a method for training a US-US deformable model according to an embodiment of the present disclosure;

[0024] Figure 14 Pre-interventional MRI images and interventional US images according to embodiments of the present disclosure are shown, wherein visualization of pre-interventional MRI features is overlaid on the interventional 3D US images; and

[0025] Figure 15 A moving US image, a fixed US image, and a registered image generated by registering the moving US image to the fixed US image are shown according to an embodiment of the present disclosure. Detailed Implementation

[0026] This disclosure provides systems and methods for real-time multimodal deformable image registration tailored for image-guided interventions, particularly addressing the challenges associated with dynamic internal anatomy during medical procedures. Image-guided interventions are medical procedures that leverage imaging techniques to precisely target pathological tissue, typically employing imaging modalities such as computed tomography (CT), magnetic resonance imaging (MRI), and ultrasound (US) to visualize anatomical structures and guide the placement of interventional devices. A significant challenge in these procedures is the dynamic nature of internal anatomy, which can undergo movement and deformation due to patient movement, respiratory and cardiac cycles, and interactions with interventional instruments.

[0027] The current disclosure addresses these challenges by providing an improved image registration approach that leverages multimodal imaging and artificial intelligence to enhance the visualization of annotated organ or tumor boundaries during intervention. This approach involves registering pre-interventional three-dimensional (3D) images acquired using modalities such as magnetic resonance imaging (MRI), computed tomography (CT), or positron emission tomography (PET) with real-time three-dimensional (3D) ultrasound (US) images. This enables real-time visualization of organ or tumor boundaries annotated on the pre-interventional images onto the US images. This approach overcomes the computational limitations of conventional rigid or analytical deformation calculations, which are unfavorable for real-time applications due to their large computational requirements.

[0028] The disclosed systems and methods utilize pose-specific, deep learning-based deformable models to fuse pre-interventional imaging data from modalities such as MRI, CT, or PET with real-time 3D US for guiding interventional device placement. In one example, the patient undergoes pre-interventional imaging of one modality (MRI, CT, or PET) and US, with imaging performed on the day of intervention. Pre-interventional imaging captures gated or breath-hold 3D images from the selected modality (MRI, CT, or PET) and multiple 3D US images representing different respiratory states for multiple poses, ranging from supine to lateral, depending on the anatomy of interest. Deformable models specific to the pre-interventional imaging modality and US (such as an MR-US deformable model for pre-interventional MRI data) are trained to register the pre-interventional images to the US images. Additionally, US-US deformable models are trained to compute transformations between pairs of US images for the same patient, thereby capturing anatomical movements due to breathing and interventional device placement.

[0029] The disclosed approach involves acquiring pre-interventional 3D images of the imaging subject using various imaging modalities (such as MRI, CT, or PET) and four-dimensional (4D) US images. These pre-interventional images are then used to train a deformable base model. The model is designed for incremental training, where network model weights are initialized from previous training sessions using data from the same imaging subject. This training mechanism offers the advantage of starting with data from a single patient and iteratively retraining the model using additional patient data, thereby enhancing the model's robustness in scenarios involving patient repositioning and internal physiological changes between pre-interventional and interventional imaging sessions. During the intervention, the pre-interventional 3D images acquired via MRI, CT, or PET are registered with real-time US images acquired during the intervention, using a pre-interventional imaging modality and US-specific trained deformable model. This registration process is performed substantially in real-time, enabling visualization of annotated lesions or other anatomical features on the pre-interventional images in conjunction with the real-time US images, facilitating interventional device placement and procedural confirmation.

[0030] In one implementation scheme Figure 1 The MRI device 10 shown can acquire high-resolution pre-interventional 3D MRI images of the imaging subject. The MRI device 10 is designed to operate in conjunction with a US imaging system to capture pre-interventional US images of the imaging subject in various respiratory states and postures. The image processing system 200 can use the acquired pre-interventional 3D MRI images and pre-interventional 3D US images to perform deformable model fine-tuning, and during the intervention, the image processing system 200 can use fine-tuned deformable models (such as a trained MR-US deformable model 208 and a trained US-US deformable model 210) to align and integrate the pre-interventional 3D MRI images with the interventional 3D US images.

[0031] In one example, the image processing system can utilize... Figure 3 and Figure 4 The procedures and methods shown herein acquire pre-interventional US and MRI image data from the same imaging subject for fine-tuning of the US-US deformity model and the MR-US deformity model. Figure 3 A block diagram is presented of a procedure 300 for acquiring high-resolution pre-interventional 3D MRI images and multiple pre-interventional 3D US images (capturing multiple respiratory states and postures) from an imaging subject for which an interventional procedure has been planned. These pre-interventional MRI and US images are used to fine-tune the MR-US deformable model 314 to register the pre-interventional 3D MRI images with the pre-interventional US images, and similarly, the US-US deformable model can be fine-tuned by registering multiple pre-interventional 3D US images with each other. Figure 4 This is a flowchart of a method 400 for acquiring imaging data for fine-tuning the MR-US deformed model 314 and the US-US deformed model.

[0032] Figure 5 and Figure 6 The registration process during the intervention phase is illustrated. Figure 5 This is a schematic diagram illustrating imaging subject 502 undergoing interventional 3DUS imaging, showing the use of a trained US-US morphology model 508 to visualize annotated pre-interventional MRI features on interventional 3D US images. During the intervention, the US imaging device 504 acquires one or more interventional 3DUS images of the imaging subject 502, such as a first interventional 3D US image 512 and an Nth interventional 3D US image 506. To fully utilize the superior soft tissue contrast and lesion salience of the pre-interventional MRI data, the model uses (using according to...) Figure 3 and Figure 4The acquired data-fine-tuned MR-US deformable model registers the pre-interventional 3D MRI image 516 of the imaging subject 502 with the interventional 3D US image. A trained US-US deformable model 508 is used to register the first interventional 3D US image 512 with the Nth interventional 3D US image 506, determining a second warp field 510 describing the transformation required to align the two US images. This second warp field 510 is then applied to the 3D MRI image registered with the first interventional 3D US image 522 (which was previously generated using the trained MR-US deformable model 514), thereby producing a 3D MRI image registered to the Nth interventional 3D US image 534. This registered 3D MRI image enables real-time visualization of tissue features annotated on the pre-interventional 3D MRI image 516 on the Nth interventional 3D US image 506 during the intervention.

[0033] Figure 6 This is a flowchart illustrating a method 600 for real-time visualization of tissue features during an intervention using registered MRI and ultrasound images. The method involves acquiring a first interventional 3D US image 602 and determining a first warp field 604 to deformably map points from the pre-interventional 3D MRI image onto the first interventional 3D US image using a trained MR-US deformation model. This first warp field is applied to the pre-interventional 3D MRI image to produce a 3D MRI image 606 registered to the first interventional 3D US image. As the intervention progresses, an Nth interventional 3D US image 608 is acquired, and a second warp field 610 is determined to deformably map points from the first interventional 3D US image onto the Nth interventional 3D US image using a trained US-US deformation model. This second warp field is applied to the 3D MRI image registered to the first interventional 3D US image to produce a 3D MRI image 612 registered to the Nth interventional 3D US image. Finally, the tissue features annotated on the pre-interventional 3D MRI images are visualized in real time on the Nth interventional 3D US image using the registered 3D MRI image 614.

[0034] Figure 7 , Figure 8 and Figure 9 Flowcharts for methods 700, 800, and 900 for acquiring MRI and US images under various respiratory states and postures are provided, and these images are stored in training datasets used for MR-US deformable models and US-US deformable models. Figure 10 and Figure 11 The training process for the MR-US deformable model is illustrated. Figure 10 This is a block diagram of the training process 1000, which includes feature extraction and deformation field generation steps. Figure 11This is a flowchart of training method 1100, which details the steps for training an MR-US deformable model using a similarity metric 1020 to update the MR-US deformable model parameters. Similarly, Figure 12 and Figure 13 The training process for the US-US deformable model is described. Figure 12 It is a flowchart of the training process 1200, and Figure 13 This is a flowchart of training method 1300, both of which focus on mapping the US image to a warped field and adjusting the parameters of the US-US deformation model based on the similarity metric 1214.

[0035] Figure 14 The image illustrates an example of the effectiveness of the currently disclosed approach for registering high-resolution MRI images, providing visualization of pre-interventional MRI features superimposed on interventional 3D US images 1406. Figure 15 An example is given of the currently disclosed approach for registering a moving US image 1502 to a fixed US image 1504 to produce a registered image 1506.

[0036] First refer to Figure 1 The image shows an MRI apparatus 10, which includes a static magnetic field magnet unit 12, a gradient coil unit 13, an RF coil unit 14, an RF body coil unit 15 (e.g., an image coil unit), a transmit / receive (T / R) switch 20, an RF driver unit 22, a gradient coil driver unit 23, a data acquisition unit 24, a controller unit 25, a patient bed or examination table 26, a data processing unit 31, a scan control device 32, and a display unit 33. In some embodiments, the RF coil unit 14 is a surface coil, which is a local coil typically placed near the anatomical structures of interest in the subject 16. Here, the RF body coil unit 15 is a transmitting coil that transmits RF signals, and the local surface of the RF coil unit 14 receives MR signals. Therefore, the transmitting body coil (e.g., the RF body coil unit 15) and the surface receiving coil (e.g., the RF coil unit 14) are independent but electromagnetically coupled components. The MRI apparatus 10 sends electromagnetic pulse signals to a subject 16 placed in an imaging space 18 that forms a static magnetic field to perform a scan to obtain magnetic resonance signals from the subject 16. One or more images of the subject 16 can be reconstructed based on the magnetic resonance signals obtained by the scan.

[0037] The static magnetic field magnet unit 12 includes, for example, a toroidal superconducting magnet mounted within a toroidal vacuum container. The magnet defines a cylindrical space surrounding the subject 16 and generates a constant main static magnetic field B0.

[0038] The MRI apparatus 10 also includes a gradient coil unit 13 that forms a gradient magnetic field in the imaging space 18 to provide three-dimensional positional information for the magnetic resonance signals received by the RF coil array. The gradient coil unit 13 includes three gradient coil systems, each generating a gradient magnetic field along one of three spatial axes perpendicular to each other, and generating gradient fields in each of the frequency encoding direction, phase encoding direction, and slice selection direction, depending on the imaging conditions. More specifically, the gradient coil unit 13 applies a gradient field in the slice selection direction (or scan direction) of the subject 16 to select slices; and the RF body coil unit 15 or the local RF coil array can transmit RF pulses to the selected slices of the subject 16. The gradient coil unit 13 also applies a gradient field in the phase encoding direction of the subject 16 to perform phase encoding of the magnetic resonance signals from the slices excited by the RF pulses. Then, the gradient coil unit 13 applies a gradient field in the frequency encoding direction of the subject 16 to perform frequency encoding of the magnetic resonance signals from the slices excited by the RF pulses.

[0039] RF coil unit 14 is configured, for example, to surround the imaging region of subject 16. In some examples, RF coil unit 14 may be referred to as a surface coil or receiving coil. In the static magnetic field space or imaging space 18 in which a static magnetic field B0 is formed by static magnetic field magnet unit 12, RF coil unit 15 sends RF pulses as electromagnetic waves to subject 16 based on control signals from controller unit 25, thereby generating a high-frequency magnetic field B1. This excites proton spins in the slice to be imaged in subject 16. RF coil unit 14 receives electromagnetic waves generated as magnetic resonance signals when the proton spins thus excited in the slice to be imaged in subject 16 return to alignment with the initial magnetization vector. In some embodiments, RF coil unit 14 may transmit RF pulses and receive MR signals. In other embodiments, RF coil unit 14 may be used only to receive MR signals without transmitting RF pulses.

[0040] The RF body coil unit 15 is configured, for example, to surround the imaging space 18 and generate RF magnetic field pulses within the imaging space 18 that are orthogonal to the main magnetic field B0 generated by the static magnetic field magnet unit 12 to excite the nucleus. The RF body coil unit 15 is fixedly attached to and connected to the MRI apparatus 10, unlike the RF coil unit 14, which can be disconnected from the MRI apparatus 10 and replaced by another RF coil unit. Furthermore, while local coils (such as the RF coil unit 14) can only send or receive signals to or from a local area of ​​the subject 16, the RF body coil unit 15 typically has a larger coverage area. For example, the RF body coil unit 15 can be used to send or receive signals to or from the whole body of the subject 16. Using only receiving local coils and transmitting body coils provides uniform RF excitation and good image homogeneity, at the cost of higher RF power deposited in the subject. For transmit-receive local coils, the local coil provides RF excitation to the region of interest and receives the MR signal, thereby reducing the RF power deposited in the subject. It should be understood that the specific use of the RF coil unit 14 and / or the RF body coil unit 15 depends on the imaging application.

[0041] When operating in receive mode, T / R switch 20 selectively connects RF body coil unit 15 to data acquisition unit 24, and when operating in transmit mode, it selectively connects the RF body coil unit to RF driver unit 22. Similarly, when RF coil unit 14 operates in receive mode, T / R switch 20 selectively connects RF coil unit 14 to data acquisition unit 24, and when the RF coil unit operates in transmit mode, it selectively connects the RF coil unit to RF driver unit 22. When both RF coil unit 14 and RF body coil unit 15 are used for a single scan, for example, if RF coil unit 14 is configured to receive MR signals and RF body coil unit 15 is configured to transmit RF signals, T / R switch 20 can direct control signals from RF driver unit 22 to RF body coil unit 15 while simultaneously directing the received MR signals from RF coil unit 14 to data acquisition unit 24. The coil of RF coil unit 15 can be configured to operate in transmit-only mode or transmit-receive mode. The coil of RF coil unit 14 can be configured to operate in transmit-receive mode or receive-only mode.

[0042] The RF driver unit 22 includes a gate modulator (not shown), an RF power amplifier (not shown), and an RF oscillator (not shown), which drive an RF coil (e.g., RF body coil unit 15) and generate a high-frequency magnetic field in the imaging space 18. Based on a control signal from the controller unit 25 and using the gate modulator, the RF driver unit 22 modulates the RF signal received from the RF oscillator into a signal with a predetermined timing and a predetermined envelope. The RF signal modulated by the gate modulator is amplified by the RF power amplifier and then output to the RF body coil unit 15.

[0043] The gradient coil driver unit 23 drives the gradient coil unit 13 based on control signals from the controller unit 25, thereby generating a gradient magnetic field in the imaging space 18. The gradient coil driver unit 23 includes three systems (not shown) of driver circuits corresponding to the three gradient coil systems included in the gradient coil unit 13.

[0044] The data acquisition unit 24 includes a preamplifier (not shown), a phase detector (not shown), and an analog-to-digital converter (not shown) for acquiring the magnetic resonance signal received by the RF coil unit 14. In the data acquisition unit 24, the phase detector uses the output of the RF oscillator from the RF driver unit 22 as a reference signal to perform phase detection on the magnetic resonance signal received from the RF coil unit 14 and amplified by the preamplifier. The phase-detected analog magnetic resonance signal is then output to the analog-to-digital converter for conversion into a digital signal. The resulting digital signal is then output to the data processing unit 31.

[0045] The MRI apparatus 10 includes an examination table 26 for placing a subject 16 thereon. The subject 16 can be moved inside and outside the imaging space 18 by moving the examination table 26 based on control signals from the controller unit 25.

[0046] The controller unit 25 includes a computer and a non-transitory computer-readable storage medium thereon storing computer-executable instructions. When executed by the computer, these instructions cause various components of the device 10 to perform operations corresponding to a predetermined scanning protocol. The non-transitory computer-readable storage medium may include one or more of a solid-state drive (SSD), hard disk drive (HDD), hybrid drive, optical disc (e.g., CD, DVD, Blu-ray), flash memory device, random access memory (RAM) device, read-only memory (ROM) device, or any other suitable non-transitory storage medium. In some embodiments, the non-transitory computer-readable storage medium may be a cloud-based or network-attached storage system accessible by the controller unit 25 via a wired or wireless network connection.

[0047] The controller unit 25 is connected to the scanning control device 32 and processes the operation signals input to the scanning control device 32. It also controls the inspection table 26, the RF driver unit 22, the gradient coil driver unit 23, and the data acquisition unit 24 by outputting control signals to them. The controller unit 25 also controls the data processing unit 31 and the display unit 33 based on the operation signals received from the scanning control device 32 to obtain the desired image.

[0048] The scanning control device 32 includes user input devices such as a touchscreen, keyboard, and mouse. The operator uses the scanning control device 32 to, for example, input data as an imaging scheme and set the area to be processed in the imaging sequence. Data regarding the imaging scheme and the area to be processed in the imaging sequence is output to the controller unit 25.

[0049] The data processing unit 31 includes a computer and a recording medium on which a program to be executed by the computer to perform predetermined data processing is recorded. The data processing unit 31 is connected to the controller unit 25 and performs data processing based on control signals received from the controller unit 25. The data processing unit 31 is also connected to the data acquisition unit 24 and generates spectral data by applying various image processing operations to the magnetic resonance signal output from the data acquisition unit 24.

[0050] Display unit 33 includes a display device and displays images on the display screen of the display device based on control signals received from controller unit 25. Display unit 33 displays images, for example, of input items regarding operation data input by the operator from scanning control device 32. Display unit 33 also displays two-dimensional (2D) slice images or three-dimensional (3D) images of the subject 16 generated by data processing unit 31.

[0051] During an MRI scan performed using the MRI apparatus 10, a subject can be positioned within the imaging space 18, and an acquisition protocol can be executed to acquire the subject's MR signal. The acquisition protocol may include multiple pulse sequences, wherein in each pulse sequence, one or more RF pulses applied via the RF body coil unit 15 are used to prepare contrast, and the gradient coil unit 13 is controlled to spatially encode the resulting MR signal. The spatially encoded MR signal received by the RF coil unit 14 is digitized and stored in k-space. Therefore, k-space data or k-space datasets can refer to the raw MR signal before processing into an image. In some examples, the raw MR signal may be filled with a line in k-space for each pulse sequence (also called a repetition time). In other examples, the raw MR signal may be filled with a line in k-space for each echo, wherein more than one echo is generated for each pulse sequence / repetition time. k-space data may also be referred to herein as imaging data or MR data.

[0052] refer to Figure 2 According to an exemplary embodiment, an image processing system 200 for deformable registration of multimodal medical images is disclosed. The system is configured to facilitate the registration of MRI and US images during image-guided interventions. The system utilizes computational models and algorithms to enhance the accuracy and efficiency of medical procedures by providing visualization of tissue features and anatomical structures.

[0053] The image processing system 200 may include an image processing device 202, a user input device 250, a display device 230, an MRI imaging device 240, and a US imaging device 260. Each component is configured to operate in concert to deliver imaging capabilities for successful image-guided interventions.

[0054] Image processing apparatus 202 includes processor 204, non-transitory memory 206, and various modules and data stored in memory 206. Processor 204 is configured to execute machine-readable instructions that control the operation of image processing system 200. Processor 204 may include individual components distributed across two or more devices, which may be located remotely and / or configured for coordinated processing.

[0055] Non-transitory memory 206 stores components such as trained MR-US deformable model 208, trained US-US deformable model 210, training module 212, and image data 214. The trained MR-US deformable model 208 and the trained US-US deformable model 210 are deep learning models trained to perform deformable registration of MRI and US images. These models handle changes and movements of internal anatomical structures during intervention, thereby providing accuracy in image registration.

[0056] Training module 212 is configured to train and fine-tune the deformable model using new data as it becomes available. This module can utilize network model weights initialized from previous training sessions, allowing for incremental learning and adaptation to new patient data or different anatomical features. Image data 214 includes pre- and interventional MRI and US images, which are used to train the model and for guidance during medical procedures.

[0057] User input device 250 (such as a keyboard, mouse, or touchscreen) enables medical professionals to interact with image processing system 200. It allows users to enter commands, adjust settings, and manipulate image data during procedures.

[0058] Display device 230 is configured to present registered images and other relevant information to medical professionals during intervention. It can show the overlay of MRI features on US images, thereby enhancing the visibility of key tissue features and anatomical structures. Display device 230 includes features such as high resolution and real-time response capabilities to ensure clarity and immediacy of visual information.

[0059] MRI imaging device 240 and US imaging device 260 are configured to acquire medical images used by image processing system 200. MRI imaging device 240 captures high-resolution 3D images of the patient's anatomy, thereby providing detailed information for accurate registration and interventional planning. US imaging device 260 provides real-time imaging capabilities, allowing for substantially real-time visualization of anatomy and interventional outcomes.

[0060] In some implementations, the MRI imaging device 240 and the US imaging device 260 can be integrated into a simultaneous MR and ultrasound imaging system with a 3D US probe. This integration facilitates the acquisition of aligned MRI and US images, thereby reducing the complexity of subsequent image registration processes.

[0061] In an alternative implementation, the image processing system 200 may include additional components, such as a network interface for connecting to a hospital information system, thereby facilitating the integration of image data with patient records and other medical data. Additionally, the system may support remote access capabilities, enabling experts to participate in procedures or view images remotely.

[0062] refer to Figure 3 The process 300 for acquiring and preparing training data for fine-tuning the MR-US deformable model 314 is illustrated, thereby registering pre-interventional MRI images with ultrasound images. An MRI imaging device 304 is used to acquire pre-interventional 3D MRI images 306 of the imaging subject 302. The 3D MRI images 306 provide high-resolution anatomical details of the region of interest within the imaging subject 302, offering superior soft tissue contrast and lesion salience compared to other imaging modalities.

[0063] In one embodiment, gating or breath-hold imaging techniques are used to acquire 3D MRI images 306 to minimize artifacts caused by respiratory movements or other patient movement during image acquisition. In another embodiment, 3D MRI images 306 are acquired with the administration of a contrast agent to further enhance visualization of specific tissues or anatomical structures of interest.

[0064] Once the 3D MRI image 306 is acquired, tissue features of interest, such as organ boundaries, tumor boundaries, or other relevant anatomical structures, are defined as indicated by annotation 316 on the 3D MRI image 306, thereby producing an annotated 3D MRI image 318. In one implementation, tissue features are manually annotated by a radiologist or other medical professional using specialized software tools. The annotation process may involve segmenting or labeling the boundaries of tissue features on individual 2D slices of the 3D MRI image 306, or using semi-automatic or automatic segmentation algorithms. In an alternative implementation, machine learning or deep learning techniques trained on a labeled dataset of MRI images are used to automatically annotate tissue features. These techniques may involve convolutional neural networks or other deep learning architectures capable of accurately detecting and segmenting various tissue types and anatomical structures within the 3D MRI image 306.

[0065] In parallel with the acquisition and annotation of 3D MRI images 306, a US imaging device 308 is used to acquire multiple pre-interventional 3D US images 310, 312, thereby capturing multiple respiratory states and postures of the imaging subject 302. These pre-interventional 3D US images 310, 312 can be acquired using a 3D ultrasound probe. In one embodiment, multiple pre-interventional 3D US images 310, 312 are acquired over a predetermined duration at a predetermined imaging frequency (e.g., 3 to 4 images per second), thereby capturing anatomical motion and deformation of the region of interest due to changes in breathing and patient positioning. In another embodiment, multiple pre-interventional 3D US images 310, 312 are acquired while the imaging subject 302 is instructed to hold their breath in different respiratory states (e.g., full inspiration, full expiration, and intermediate states). Additionally, multiple pre-interventional 3D US images 310, 312 can be acquired while the imaging subject 302 is in different postures or positions (such as supine, lateral, or other orientations associated with the planned interventional procedure and the anatomical region of interest).

[0066] Annotated 3D MRI images 318 and multiple pre-interventional 3D US images 310, 312 are used to fine-tune the MR-US deformability model 314 for imaging object 302. The MR-US deformability model 314 is a deep learning-based model trained to accurately align and register MRI image data with ultrasound image data, taking into account the inherent differences in image appearance and characteristics between the two imaging modalities. In one embodiment, the MR-US deformability model 314 is based on a CNN architecture that takes the annotated 3D MRI image 318 and the pre-interventional 3D US image (e.g., 310) as input during training and learns an output that maps the annotated 3D MRI image 318 to a deformable field of the corresponding pre-interventional 3D US image 310. This deformable field is then applied to the annotated 3D MRI image 318 using a spatial transformer to produce a first deformed 3D MRI image 320 aligned with the pre-interventional 3D US image 310. In one implementation, the MR-US deformation model 314 is an unsupervised deep learning model trained without requiring a ground-based deformation field. Instead, a similarity metric (such as normalized cross-correlation or mutual information) is used to train the model to maximize the alignment between the deformed 3D MRI image 320 and the corresponding pre-interventional 3D US image 310. This unsupervised approach allows the model to learn the complex, nonlinear deformations required for registering MRI and ultrasound data without the need for manually annotated ground-based deformation fields.

[0067] The MR-US deformation model 314 is trained on each of the pre-intervention 3D US images 310, 312 to generate multiple deformed 3D MRI images 320, 322, each deformed 3D MRI image corresponding to a specific respiratory state or posture of the imaging subject 302 captured by the corresponding pre-intervention 3D US image 310, 312.

[0068] refer to Figure 4 This paper illustrates a flowchart of a method 400 for fine-tuning a trained MR-US deformable model using pre-interventional 3D MRI and pre-interventional 3D US images of the same imaging subject. Method 400 outlines the steps involved in acquiring and processing the necessary imaging data to fine-tune the MR-US deformable model, which is a deep learning-based model designed to accurately align and register the MRI and US data.

[0069] Method 400 begins with operation 402, in which pre-interventional three-dimensional (3D) MRI images of the imaging subject are acquired. In one implementation, an MRI device (such as...) is used. Figure 1The MRI device 10 shown acquires 3D MRI images before intervention. The MRI device 10 can employ various imaging techniques, including gated or breath-hold imaging, to capture high-resolution 3D images of the anatomical structures of the subject while minimizing artifacts caused by respiratory movements or other patient movements.

[0070] At operation 404, tissue features of interest are annotated on the pre-interventional 3D MRI images acquired in operation 402. In one implementation, tissue features are manually annotated by a radiologist or other medical professional using specialized software tools. The annotation process may involve segmenting or marking the boundaries of tissue features on individual two-dimensional (2D) slices of the 3D MRI image, or using semi-automatic or automatic segmentation algorithms. Annotated tissue features may include organ boundaries, tumor boundaries, or other relevant anatomical structures of interest for the planned interventional procedure. In an alternative implementation, tissue features are automatically annotated using machine learning or deep learning techniques trained on a labeled dataset of MRI images. These techniques may involve convolutional neural networks or other deep learning architectures capable of accurately detecting and segmenting various tissue types and anatomical structures within 3D MRI images.

[0071] At operation 406, multiple pre-interventional 3D US images of the same imaging subject are acquired, thereby capturing multiple respiratory states and postures. In one embodiment, a 3D ultrasound probe or transducer is used to acquire multiple pre-interventional 3D US images. Pre-interventional 3D US images can be acquired at a predetermined imaging frequency over a predetermined duration, thereby capturing anatomical movements and deformations of the region of interest due to changes in respiration and patient positioning. In another embodiment, multiple pre-interventional 3D US images are acquired while the imaging subject is instructed to hold their breath in different respiratory states, such as full inspiration, full expiration, and intermediate states. Additionally, multiple pre-interventional 3D US images can be acquired while the imaging subject is in different postures or positions, such as supine, lateral, or other orientations associated with the planned interventional procedure and the anatomical region of interest.

[0072] At operation 408, the MR-US deformability model is used to register the pre-interventional 3D MRI image acquired in operation 402 with each of the multiple pre-interventional 3D US images acquired in operation 406. The MR-US deformability model is a deep learning-based model trained to accurately align and register MRI and ultrasound image data, taking into account the inherent differences in image appearance and characteristics between the two imaging modalities. In one implementation, the MR-US deformability model is based on a convolutional neural network (CNN) architecture that takes the pre-interventional 3D MRI image and the pre-interventional 3D US image as input during training. The model learns to output a deformable field that maps the pre-interventional 3D MRI image to the corresponding pre-interventional 3D US image. This deformable field is then applied to the pre-interventional 3D MRI image using a spatial transformer to produce a deformed 3D MRI image aligned with the pre-interventional 3D US image. The registration process performed at operation 408 produces multiple deformed 3D MRI images, each deformed 3D MRI image corresponding to a specific pre-interventional 3D US image, and represents the pre-interventional 3D MRI image deformed to align with that specific US image.

[0073] At operation 410, each of the multiple deformed 3D MRI images generated in operation 408 is stored in non-transitory memory (such as a hard disk drive or solid-state drive) associated with the corresponding pre-interventional 3D US image. In one embodiment, the deformed 3D MRI images and their associated pre-interventional 3D US images are stored in a structured database or file system, where each MRI image is linked to a corresponding set of US images acquired under the same respiratory state and posture. This facilitates efficient retrieval and processing of imaging data during the interventional phase when the deformed 3D MRI images are used in conjunction with the interventional US images for real-time visualization of tissue features. In an alternative embodiment, the deformed 3D MRI images and their associated pre-interventional 3D US images are stored together with metadata describing imaging parameters, patient information, respiratory status annotations, and other relevant details. This metadata can be used for data management, quality control, and potential future refinement or expansion of the MR-US deformed model.

[0074] Following operation 410, method 400 concludes, having acquired and processed the pre-interventional imaging data used for fine-tuning the MR-US deformable model. The fine-tuned MR-US deformable model, along with the stored deformed 3D MRI images and their associated pre-interventional 3D US images, will be used during the interventional phase to achieve real-time visualization of tissue features annotated on the pre-interventional MRI data within the live interventional US images, thereby enhancing the guidance and accuracy of the interventional procedure.

[0075] refer to Figure 5 This paper illustrates an interventional image registration process 500 for registering pre-interventional MRI images to interventional 3D ultrasound images. This process enables real-time visualization of tissue features annotated on the pre-interventional MRI images on the interventional 3D ultrasound images, thereby enhancing the guidance and accuracy of the interventional procedure.

[0076] Procedure 500 begins when the imaging subject 502 undergoes an intervention. During the intervention, the US imaging device 504 acquires one or more interventional 3D US images of the imaging subject 502, such as a first interventional 3D US image 512 and an Nth interventional 3D US image 506. These interventional 3D US images capture anatomical regions of interest in real time, providing up-to-date information on the subject's anatomy and the location of any interventional devices or instruments.

[0077] To fully utilize the superior soft tissue contrast and lesion salience of pre-interventional MRI data, a two-step process involving a trained MR-US deformable model 514 and a trained US-US deformable model 508 is used to register the pre-interventional 3D MRI image 516 of the imaging subject 502 with the interventional 3D US image.

[0078] In the first step, a trained MR-US deformable model 514 is used to register the pre-interventional 3D MRI image 516 with the first interventional 3D US image 512. The trained MR-US deformable model 514 is a deep learning-based model that utilizes features derived from both the pre-interventional 3D MRI image 516 and the first interventional 3D US image 512, including radiomics features extracted from the images and Gaussian mixture model tissue class probabilities computed for each voxel in the images. Radiomics features may include descriptors of intensity, shape, texture, and other quantitative properties within the images. The Gaussian mixture model tissue class probabilities provide a voxel-by-voxel estimate of the probability that each voxel belongs to a different tissue class (such as fat, muscle, organ tissue, etc.). By fully utilizing these derived features, the trained MR-US deformable model 514 can account for the inherent differences in image appearance and characteristics between MRI and ultrasound data. In one implementation, the trained MR-US deformation model 514 is based on a convolutional neural network (CNN) architecture that takes as input a pre-interventional 3D MRI image 516, a first interventional 3D US image 512, and derived radiomics and tissue probability features, and outputs a first warping field 518 that encodes the transformation required to align the pre-interventional 3D MRI image 516 with the first interventional 3D US image 512.

[0079] The first warped field 518 is then applied to the pre-interventional 3D MRI image 516 using a first spatial transformer 520. The first spatial transformer 520 is a differentiable module that uses a sampling kernel (such as bilinear or trilinear interpolation) to apply the first warped field 518 to the pre-interventional 3D MRI image 516 to produce a 3D MRI image registered to the first interventional 3D US image 522.

[0080] In the second step, a trained US-US deformability model 508 is used to further register the 3D MRI image registered to the first interventional 3D US image 522 with the Nth interventional 3D US image 506. Similar to the MR-US deformability model, the trained US-US deformability model 508 utilizes radiomics features derived from the first interventional 3D US image 512 and the Nth interventional 3D US image 506, along with Gaussian mixture model tissue category probabilities, to account for differences in appearance between the two US images caused by factors such as anatomical motion and deformability. The trained US-US deformability model 508 is based on a CNN architecture that takes the first interventional 3D US image 512, the Nth interventional 3D US image 506, and the derived radiomics and tissue probability features as input, and outputs a second warped field 510 that describes the transformation required to align the first interventional 3D US image 512 with the Nth interventional 3D US image 506.

[0081] The second warped field 510 is then applied using a second spatial transformer 532 to a 3D MRI image registered to the first interventional 3D US image 522. The second spatial transformer 532 is a differentiable module that uses a sampling kernel to apply the second warped field 510 to the 3D MRI image registered to the first interventional 3D US image 522, thereby generating a 3D MRI image registered to the Nth interventional 3D US image 534.

[0082] In an alternative implementation, the second spatial transformer 532 uses a non-rigid registration algorithm (such as B-spline or freeform deformation) to apply the second warping field 510 to the 3D MRI image registered to the first interventional 3D US image 522, thereby producing a 3D MRI image registered to the Nth interventional 3D US image 534.

[0083] The 3D MRI image registered to the Nth interventional 3D US image 534 is then displayed together with the Nth interventional 3D US image 506 on the display device 536. By overlaying the registered 3D MRI image onto the Nth interventional 3D US image 506, tissue features annotated on the pre-interventional MRI image can be visualized in real time on the Nth interventional 3D US image 506, thereby providing valuable guidance for the interventional procedure.

[0084] In one embodiment, the display device 536 is a high-resolution monitor or display panel capable of simultaneously displaying both the Nth interventional 3D US image 506 and the 3D MRI image registered to the Nth interventional 3D US image 534. The display device 536 can support various visualization modes (such as side-by-side or overlay views) to facilitate comparison and integration of the two image modalities.

[0085] In another implementation, process 500 can incorporate additional pre-interventional imaging modalities or data sources, such as computed tomography (CT) or positron emission tomography (PET), by training additional deformable models and integrating them into the registration pipeline. This flexibility allows the system to fully leverage the advantages of various imaging modalities and provide comprehensive multimodal visualization for interventional guidance.

[0086] The interventional image registration process 500 overcomes the computational limitations of conventional rigid or analytical deformation calculations, which are unfavorable for real-time applications due to their large computational requirements. By fully utilizing the trained US-US deformation model 508 and the pre-calculated 3D MRI image registered to the first interventional 3D US image 522, the process 500 enables real-time visualization of tissue features during intervention, thereby improving the accuracy and efficiency of the procedure.

[0087] refer to Figure 6 A flowchart of a method 600 for real-time visualization of tissue features during an intervention using registered MRI and ultrasound images is shown. Method 600 integrates high-resolution anatomical information from pre-interventional MRI data with real-time ultrasound imaging data, thereby allowing enhanced visualization of tissue features during the interventional procedure.

[0088] Method 600 begins with operation 602, wherein a first interventional 3D US image is acquired during the intervention. In one embodiment, an interventional 3D US image is captured during the intervention using a 3D US probe or transducer positioned near the anatomical region of interest. The 3D US probe is capable of acquiring real-time 3D US images, thereby enabling real-time visualization of the anatomical region and any interventional device or instrument. The acquisition of interventional 3D US images can be manually triggered by the operator or automatically triggered based on predefined criteria, such as detection of the interventional device within the imaging field of view.

[0089] At operation 604, a trained MR-US warp model is used to determine a first warp field that deformably maps points from the pre-interventional 3D MRI image to corresponding points in the interventional 3D US image acquired in operation 602. In one embodiment, the trained MR-US warp model is a deep learning-based model trained to accurately align and register MRI and ultrasound image data, taking into account the inherent differences in image appearance and characteristics between the two imaging modalities. The trained MR-US warp model can take the pre-interventional 3D MRI image and the interventional 3D US image as input and output a first warp field representing the warp or displacement field required to align the pre-interventional 3D MRI image with the interventional 3D US image. The first warp field can be a dense warp field, where each voxel in the pre-interventional 3D MRI image is associated with a displacement vector, or the first warp field can be a sparse warp field, where displacement is defined only at a subset of control points, and the displacements of the remaining voxels are interpolated from the control point displacements.

[0090] At operation 606, a first spatial transformer is used to apply the first warped field determined in operation 604 to the pre-interventional 3D MRI image to produce a 3D MRI image registered to the first interventional 3DUS image acquired in operation 602. In one embodiment, the first spatial transformer is a differentiable module that applies the first warped field to the pre-interventional 3D MRI image using a sampling kernel (such as bilinear or trilinear interpolation). This operation effectively warps or deforms the pre-interventional 3D MRI image to align it with the interventional 3D US image, thus taking into account any anatomical movement or deformation that may have occurred between the acquisition of the two images.

[0091] Method 600 then proceeds to operation 608, in which the Nth interventional 3D US image is acquired during the intervention. This operation can be repeated multiple times throughout the intervention, capturing the anatomical region of interest at different time points or in different deformed states due to factors such as patient movement, respiratory motion, or interaction with the interventional device.

[0092] At operation 610, a trained US-US deformation model is used to determine a second warping field that allows points in the first interventional 3D US image acquired in operation 602 to be deformably mapped to corresponding points in the Nth interventional 3D US image acquired in operation 608. The trained US-US deformation model is a deep learning-based model trained to compute transformations between pairs of US images representing different respiratory states or postures of the same patient. In one implementation, the trained US-US deformation model takes the first interventional 3D US image acquired in operation 602 and the Nth interventional 3D US image acquired in operation 608 as input and outputs a second warping field. This second warping field represents the deformation or displacement field required to align the first interventional 3D US image acquired in operation 602 with the Nth interventional 3D US image acquired in operation 608, taking into account any anatomical movements or deformations that may have occurred between the acquisition of the two US images.

[0093] At operation 612, a second spatial transformer is used to apply the second warping field determined in operation 610 to a 3D MRI image registered to the first interventional 3D US image acquired in operation 602, to produce a 3D MRI image registered to the Nth interventional 3D US image acquired in operation 608. In one embodiment, the second spatial transformer is a differentiable module that uses a sampling kernel (such as bilinear or trilinear interpolation) to apply the second warping field to the 3D MRI image registered to the first interventional 3D US image acquired in operation 602. Operation 612 effectively warps or deforms the 3D MRI image registered to the first interventional 3D US image acquired in operation 602 to align it with the Nth interventional 3D US image acquired in operation 608. By applying a second warping field, the resulting 3D MRI image is registered to the most recent interventional 3D US image, thereby taking into account any anatomical movement or deformation that may have occurred between the acquisition of the first interventional 3D US image acquired in operation 602 and the acquisition of the Nth interventional 3D US image acquired in operation 608.

[0094] At operation 614, using a 3D MRI image registered to the Nth interventional 3D US image generated in operation 612, tissue features annotated on the pre-interventional 3D MRI image are visualized in real time on the Nth interventional 3D US image acquired in operation 608. In one embodiment, using a 3D MRI image registered to the Nth interventional 3D US image, tissue features (such as organ boundaries or tumor boundaries) annotated on the pre-interventional 3D MRI image are overlaid onto the Nth interventional 3D US image. This overlay process may involve techniques such as alpha mixing, where a weighted summation is used to combine the intensities of the MRI and US images, thereby allowing tissue features annotated on the pre-interventional MRI image to be visible on the interventional US image while preserving the real-time imaging capability of the US modality. Alternatively, the tissue features annotated on the pre-interventional 3D MRI image can be segmented and rendered as a 3D surface or mesh, which can then be overlaid onto the Nth interventional 3D US image using appropriate rendering techniques.

[0095] Real-time visualization of tissue features annotated on pre-interventional 3D MRI images on the Nth interventional 3D US image provides visual guidance during the intervention. For example, in the case of ablation procedures, visualization of tumor boundaries on real-time US images can help accurately locate the ablation device and monitor the progress of the ablation process. Similarly, in the case of biopsy procedures, visualization of lesion boundaries can help accurately target the biopsy needle and ensure the acquisition of the desired tissue sample.

[0096] Method 600 overcomes the computational limitations of conventional rigid or analytical deformation calculations, which are unsuitable for real-time applications due to their large computational requirements. By fully utilizing trained MR-US deformation models and trained US-US deformation models, Method 600 enables real-time visualization of tissue features during intervention, thereby improving the accuracy and efficiency of the procedure.

[0097] In alternative implementations, procedures 608, 610, 612, and 614 can be repeated multiple times throughout the intervention, with each iteration using the most recently acquired interventional 3D US image as input for procedure 608. This iterative approach allows for continuous updates to the registered 3D MRI images, ensuring that tissue features annotated on the pre-interventional 3D MRI images are accurately overlaid on the interventional 3D US images, even when the anatomical region of interest undergoes motion or deformation due to factors such as patient movement, respiratory motion, or interaction with the interventional device.

[0098] In summary, Method 600 provides a robust and efficient approach for real-time visualization of tissue features during interventional procedures, leveraging the complementary advantages of MRI and ultrasound imaging modalities. By combining high-resolution anatomical information from pre-interventional MRI data with the real-time imaging capabilities of ultrasound, Method 600 enhances the guidance and precision of interventional procedures, ultimately leading to improved patient outcomes.

[0099] refer to Figure 7 A flowchart of method 700 for acquiring and storing imaging data for training MR-US and US-US deformable models is shown. Method 700 provides a systematic approach for acquiring and organizing the imaging data necessary for self-supervised training of deformable models.

[0100] Method 700 begins with operation 702, in which three-dimensional 3D MRI images of the imaging subject are acquired at a selected location. The selected location may correspond to a specific pose or orientation of the imaging subject, such as supine, lateral, or any other location associated with the planned interventional procedure and the anatomical region of interest.

[0101] In one implementation, gating or breath-hold imaging techniques are used to acquire 3D MRI images to minimize artifacts caused by respiratory motion or other patient movement during image acquisition. This ensures that the acquired 3D MRI images accurately represent anatomical structures in the selected location without significant distortion or blurring due to motion. In another implementation, 3D MRI images are acquired by applying a contrast agent to enhance visualization of specific tissues or anatomical structures of interest.

[0102] At operation 704, multiple 3DUS images of the imaging subject are acquired at selected locations from operation 702, thereby capturing multiple respiratory states. These 3D US images can be acquired using a 3D ultrasound probe or transducer.

[0103] In one implementation, multiple 3DUS images are acquired at a predetermined imaging frequency over a predetermined duration, thereby capturing anatomical motion and deformation of the region of interest due to respiration. This approach allows for the collection of a comprehensive set of US images representing the entire range of respiratory states, enabling the training of a robust deformation model capable of taking into account anatomical deformations caused by respiration.

[0104] In another implementation, multiple 3D US images are acquired while the imaging subject is instructed to hold its breath in different breathing states (e.g., full inspiration, full exhalation, and intermediate states). This approach ensures that the acquired US images accurately represent specific breathing phases, allowing the deformation model to learn the transformations associated with each breathing state.

[0105] At operation 706, the 3D MRI images acquired in operation 702 and multiple 3D US images of the imaging object acquired in operation 704 are stored in a first training dataset for the MR-US deformity model. This training dataset will be used to train the MR-US deformity model to accurately register and align the MRI images with the corresponding US images, taking into account anatomical deformities and pose changes captured in the US images. In one embodiment, the 3D MRI images and multiple 3D US images are stored in a structured database or file system, where each MRI image is associated with a corresponding set of US images acquired at the same selected location and respiratory state. This organization facilitates efficient retrieval and processing of training data during model training. In another embodiment, the 3D MRI images and multiple 3D US images are stored together with metadata describing imaging parameters, patient information, and other relevant details. This metadata can be used for data management, quality control, and potential future refinement or expansion of the deformity model.

[0106] At operation 708, multiple 3D US images of the imaging object acquired in operation 704 are stored in a second training dataset for the US-US deformable model. This training dataset will be used to train the US-US deformable model to compute transformations between pairs of US images representing different respiratory states of the same patient. In one embodiment, multiple 3D US images are stored in a structured database or file system, where each set of US images acquired at a specific selected location and respiratory state is organized together. This organization facilitates efficient retrieval and processing of training data during the model training process of the US-US deformable model. In another embodiment, multiple 3D US images are stored together with metadata describing imaging parameters, patient information, respiratory state annotations, and other relevant details. This metadata can be used for data management, quality control, and potential future refinement or expansion of the US-US deformable model.

[0107] By performing the operations of Method 700, imaging data for training both the MR-US and US-US deformable models are acquired and organized in a structured manner. Method 700 ensures that the deformable models are trained on a comprehensive ensemble of imaging data, thereby capturing various patient positions, respiratory states, and anatomical deformities, resulting in accurate and robust performance during real-time multimodal deformable image registration for image-guided interventions.

[0108] refer to Figure 8A flowchart of a method 800 for acquiring and storing 3D MRI images for training an MR-US deformable model is shown. Method 800 outlines the steps involved in acquiring and storing 3D MRI images of an imaging object that will be used as part of the training data for the MR-US deformable model.

[0109] Method 800 begins with operation 802, in which imaging system parameters for MRI acquisition are initialized. In one embodiment, this operation involves configuring the MRI imaging device 240 ( Figure 2 (As shown) 3D MRI images are acquired using specific imaging parameters tailored to the anatomical region of interest and the object being imaged. These parameters may include, but are not limited to, field of view (FOV), spatial resolution, slice thickness, imaging sequence (e.g., T1-weighted, T2-weighted, or other sequences), and contrast agent administration (if applicable).

[0110] At operation 804, the location of the imaging subject is selected. The selection of the imaging subject's location determines the orientation and pose of the anatomical structures captured in the 3D MRI image. In one embodiment, the imaging subject is positioned in a supine position. In another embodiment, the imaging subject may be positioned in a lateral decubitus position, which may be preferred for certain anatomical regions or interventional procedures.

[0111] At operation 806, gating or breath-holding techniques are used to acquire 3D MRI images at a selected location. Gating or breath-holding techniques are employed to minimize artifacts caused by respiratory motion or other patient movement during image acquisition. In one embodiment, gating techniques are used to acquire 3D MRI images, where MRI data acquisition is synchronized with the respiratory cycle of the imaging subject. This technique involves acquiring MRI data at specific phases of the respiratory cycle (such as end-expiration or end-inspiration) to capture anatomical structures in a consistent state of motion or deformation.

[0112] In another implementation, a breath-hold technique is used to acquire 3D MRI images, where the subject is instructed to hold their breath for a short period (typically 10 to 20 seconds) during MRI data acquisition. This technique eliminates respiratory motion artifacts by capturing static anatomical structures during the breath-hold period.

[0113] At operation 808, the acquired 3D MRI image is stored in a non-transitory memory, such as the non-transitory memory 206 of the image processing system 200 (e.g., ...). Figure 2 (As shown). 3D MRI images are indexed and associated with metadata, including the location of the imaged object and other relevant information, such as the identifier of the imaged object, the anatomical region of interest, and the imaging parameters used during acquisition.

[0114] In one implementation, 3D MRI images are stored in a hierarchical file structure, where each 3D MRI image is stored in a subdirectory or folder associated with an identifier of the imaging object and the specific location or pose in which the image was acquired. This file structure can be organized based on patient identifiers, imaging session timestamps, or other relevant metadata to facilitate efficient retrieval and management of 3D MRI images.

[0115] In another implementation, the 3D MRI images are stored in a database or file system accessible by the image processing system responsible for training the MR-US deformable model. The association between the 3D MRI images and the location of the imaging object can be implemented using data structures such as linked lists, hash tables, or other suitable data organization techniques.

[0116] After operation 808, method 800 can end. The acquired and stored 3D MRI images, along with their associated metadata, will be used as part of the training data for an MR-US deformable model responsible for registering or aligning the MRI images with corresponding ultrasound images during image-guided interventional procedures.

[0117] refer to Figure 9 A flowchart of a method 900 for acquiring and storing 3D US images of an imaging object in various respiratory states for training a US-US deformable model is shown. Method 900 is designed to capture the anatomical motion and deformation of the internal anatomy of the imaging object due to respiration, which is used to train the US-US deformable model to accurately register US images acquired during different respiratory states.

[0118] Method 900 begins with operation 902, in which imaging system parameters for 3D US imaging are initialized. In one embodiment, this operation involves configuring the US imaging device 260 ( Figure 2 (As shown in the diagram) to acquire 3D US images. This may include setting the imaging mode to 3D or 4D (real-time 3D) mode, adjusting the imaging depth and field of view to cover the anatomical region of interest, and optimizing imaging parameters (such as gain, dynamic range, and frequency) to ensure high-quality image acquisition. In another embodiment, operation 902 may involve initializing a simultaneous MR and ultrasound imaging system using a 3D US probe. This integrated imaging system allows for the simultaneous acquisition of MRI and US data, which can facilitate subsequent registration of the two modalities during the training process.

[0119] At operation 904, the location or position of the imaging subject is selected. In one embodiment, the imaging subject may be positioned in a supine position, which is a common starting point for many interventional procedures. In another embodiment, the imaging subject may be positioned in a lateral decubitus position, which may be more suitable for certain anatomical regions or interventional approaches. The selection of the imaging subject's location at operation 904 can be based on specific anatomical regions of interest and the planned interventional procedure. For example, if the region of interest is the liver or abdomen, the imaging subject may be positioned in a left or right lateral decubitus position to provide better access to and visualization of the target anatomy.

[0120] At operation 906, 3D US images of the imaging subject are acquired at a selected location, thereby capturing multiple respiratory states. In one embodiment, 3D US images are acquired continuously at a predetermined imaging frequency for a predetermined duration, allowing the capture of multiple respiratory cycles and associated anatomical movements and deformities. In another embodiment, 3D US images are acquired while the imaging subject is instructed to hold its breath in different respiratory states, such as full inspiration, full expiration, and intermediate states. This approach can provide a more controlled and reproducible set of respiratory states, which can be beneficial for training US-US deformable models. During the acquisition of 3D US images at operation 906, various techniques, such as breathing bellows or optical tracking systems, can be used to monitor the respiratory state of the imaging subject. This information can be used to associate each acquired 3D US image with a specific respiratory state, thereby enabling accurate labeling of the training data used for US-US deformable models.

[0121] At operation 908, the acquired 3D US images are stored in non-transitory memory, indexed by the location and respiratory state of the imaging object. In one embodiment, the 3D US images are stored in a hierarchical file structure, where each location or pose of the imaging object is represented by a separate directory or folder. Within each location directory, the 3D US images are organized and labeled according to the corresponding respiratory state (such as "inhalation," "exhalation," or a specific lung imaging measurement). In another embodiment, the 3D US images are stored in a database or data management system, where each image is associated with metadata describing the imaging object, its location, and its respiratory state. This metadata can be used to efficiently retrieve and organize training data for the US-US deformable model, as well as for subsequent analysis or quality control purposes. The storage of 3D US images at operation 908 may also involve preprocessing steps, such as image enhancement, noise reduction, or data compression, to optimize the storage and retrieval of training data. Additionally, the stored 3D US images can be used in conjunction with other imaging modalities (such as MRI or CT) to create multimodal training datasets for the US-US deformable model or other deformable models used in image registration pipelines.

[0122] By acquiring and storing 3D US images of various respiratory states and locations, Method 900 provides a comprehensive dataset for training a US-US deformation model. This model then plays a role in the real-time registration of interventional US images with pre-interventional US and MRI data, thereby enabling accurate visualization of tissue features and enhancing the guidance and accuracy of image-guided interventions.

[0123] refer to Figure 10 The diagram illustrates a block diagram of an MR-US deformable model training process 1000, which exemplifies the training of a deep learning-based deformable model for registering MRI images to US images. The MR-US deformable model training process 1000 is designed to learn the complex nonlinear deformabilities required for aligning MRI and US data, thus taking into account the inherent differences in image appearance and characteristics between the two imaging modalities.

[0124] The training process 1000 begins with an MRI image 1002, which is a 3D or 4D MRI image of the object being imaged. In one embodiment, the MRI image 1002 is generated using an MRI imaging device (such as...) Figure 1 The MRI device 10) shown in the image acquired before the interventional procedure yielded a high-resolution 3D MRI image. The MRI image 1002 provides detailed anatomical information.

[0125] MRI image 1002 is input into MRI feature extractor 1004, which is responsible for extracting relevant features from MRI image 1002. In one implementation, MRI feature extractor 1004 is a CNN that applies a series of convolutional, pooling, and nonlinear activation operations to MRI image 1002 to produce feature maps that capture significant anatomical and structural information present in MRI image 1002. MRI feature extractor 1004 can be pre-trained on a large dataset of MRI images to learn a robust feature set effective for various medical imaging tasks.

[0126] In parallel, the training process 1000 also receives a US image 1006, which is a 3D or 4D US image of the same imaging object as the MRI image 1002. In one embodiment, the US image 1006 is obtained using a US imaging device (such as...) Figure 2 The 3D US image 1006 is acquired by the US imaging device 260 shown. The US image 1006 captures the anatomical region of interest in different modalities, thereby providing complementary information to the MRI image 1002.

[0127] US image 1006 is input into US feature extractor 1008, which is responsible for extracting relevant features from US image 1006. Similar to MRI feature extractor 1004, US feature extractor 1008 can be a CNN, which applies a series of convolution, pooling, and nonlinear activation operations to US image 1006 to produce feature maps that capture significant anatomical and structural information present in US image 1006. US feature extractor 1008 can be pre-trained on a large dataset of US images to learn a robust feature set specific to the US imaging modality.

[0128] The feature maps generated by the MRI feature extractor 1004 and the US feature extractor 1008 are cascaded to form a cascaded MRI and US feature map 1010. In one embodiment, the cascading operation involves stacking the feature maps along the channel dimension to create a single feature map that combines information from both the MRI and US modalities.

[0129] The cascaded MRI and US feature maps 1010 are then input into a deformation field generator 1012, which is responsible for predicting the deformation field required to align the MRI image 1002 with the US image 1006. In one embodiment, the deformation field generator 1012 is a CNN or a fully connected neural network that takes the cascaded MRI and US feature maps 1010 as input and outputs a dense deformation field, also known as a warped field 1014.

[0130] The warp field 1014 represents the displacement or deformation vector that maps each voxel (3D pixel) in the MRI image 1002 to its corresponding location in the US image 1006. The warp field 1014 captures the complex nonlinear deformations required to account for differences in image appearance, patient positioning, and anatomical deformations between the MRI and US modalities.

[0131] The warped field 1014 is then applied to the MRI image 1002 using a spatial transformer 1016, which warps or deforms differentiable modules of the input image based on the provided warping field. In one embodiment, the spatial transformer 1016 applies the warped field 1014 to the MRI image 1002 using a sampling kernel (such as bilinear or trilinear interpolation) to produce a deformed MRI image 1018 aligned with the US image 1006.

[0132] The similarity measure 1020 is then used to compare the deformed MRI image 1018 and the US image 1006. This similarity measure is a measure of the alignment or registration quality between the two images. In one embodiment, the similarity measure 1020 is a normalized cross-correlation or mutual information measure that evaluates the similarity between the deformed MRI image 1018 and the US image 1006 without requiring a ground-based deformable field or manual annotation.

[0133] Similarity metric 1020 is used to update the parameters of MRI feature extractor 1004, US feature extractor 1008, and deformation field generator 1012 via backpropagation, with the goal of maximizing the similarity between the deformed MRI image 1018 and the US image 1006. This iterative training process allows the MR-US deformation model to learn the complex deformations required for registering MRI and US data without manually annotated ground-based deformation fields, which can be time-consuming and difficult to obtain.

[0134] In alternative implementations, the training process 1000 may incorporate additional components or modules to enhance the performance of the MR-US deformation model. For example, the training process 1000 may include a regularization module that applies constraints or penalties to the deformation field generator 1012, thereby encouraging the model to learn smooth and physically plausible deformations. Additionally, the training process 1000 may include a multi-scale or pyramidal approach, where MRI and US images are processed at multiple resolutions, allowing the model to capture both coarse-grained and fine-grained deformations.

[0135] The MR-US deformable model training process 1000 is designed to be incremental and adaptive, allowing the model to be trained on new data as it becomes available. In one implementation, the training process 1000 initializes the weights of the MRI feature extractor 1004, the US feature extractor 1008, and the deformable field generator 1012 using pre-trained weights from previous training sessions or publicly available models. When new MRI and US image pairs are acquired, the training process 1000 can be re-executed to fine-tune the model parameters to suit the specific characteristics of the new data. Overall, the MR-US deformable model training process 1000 provides a robust and efficient approach to learning the complex deformabilities required for registering MRI and US data, thereby enabling accurate and real-time visualization of tissue features during image-guided interventions.

[0136] refer to Figure 11 A flowchart of method 1100 for training an MR-US deformable model to deformably register 3D MRI images with 3D US images from the same imaging subject is shown. Method 1100 is designed to train a deep learning-based deformable model that can accurately align and register MRI and ultrasound image data, taking into account the inherent differences in image appearance and characteristics between the two imaging modalities. Furthermore, method 1100 employs a novel self-supervised training approach that bypasses the need for training on labeled data, thus enabling training on large amounts of data—a recognized bottleneck in machine learning.

[0137] Method 1100 begins with operation 1102, in which MRI images and US images from the same imaging subject are selected from a training dataset. In one implementation, the training dataset comprises a collection of pre-interventional 3D MRI images and corresponding 4D US images acquired for a set of patients or imaging subjects. The 4D US images capture anatomical regions of interest in multiple respiratory states and postures, thus providing a comprehensive representation of deformities and variations that can occur during the interventional procedure. In an alternative implementation, the training dataset can be initialized with data from a single patient or imaging subject and incrementally expanded as additional patient data becomes available. This approach allows for iterative training and refinement of the MR-US deformability model, fully utilizing model weights from previous training sessions as a starting point for subsequent training using new data.

[0138] At operation 1104, an MRI feature extractor is used to extract features from the MRI images to generate an MRI feature map. The MRI feature extractor is a component of the MR-US deformable model, designed to capture relevant anatomical and structural information from MRI images, encoding it into a compact feature representation or feature map. In one implementation, the MRI feature extractor can be pre-trained on a large dataset of MRI images, allowing it to learn general features and patterns specific to the MRI data. During training of the MR-US deformable model, the weights of the MRI feature extractor can be fine-tuned to adapt to specific anatomical regions and imaging characteristics of the training dataset.

[0139] At operation 1106, a US feature extractor is used to extract features from the US images to generate a US feature map. Similar to an MRI feature extractor, the US feature extractor is a component of the MR-US deformable model, designed to capture relevant anatomical and structural information from US images, encoding them into compact feature representations or feature maps. In one implementation, the US feature extractor can be pre-trained on a large dataset of ultrasound images, allowing it to learn general features and patterns specific to ultrasound data. During training of the MR-US deformable model, the weights of the US feature extractor can be fine-tuned to suit specific anatomical regions and imaging characteristics of the training dataset.

[0140] At operation 1108, the MRI feature map and the US feature map are cascaded to produce cascaded MRI and US feature maps. This cascaded operation combines the feature representations extracted from the MRI and US images, enabling subsequent components of the MR-US deformable model to learn the relationship and correspondence between the two modalities. In an alternative implementation, instead of cascading feature maps, other fusion techniques can be used to combine the MRI and US feature maps, such as element-wise addition, element-wise multiplication, or more complex fusion operations implemented as separate neural network layers.

[0141] At operation 1110, a deformation field generator is used to map the cascaded MRI and US feature maps to a warped field. The deformation field generator takes the cascaded feature maps as input and outputs a warped field representing the deformation or displacement field required to align the MRI image with the corresponding US image. In one embodiment, the deformation field generator can be designed to output a dense warped field, where each voxel in the MRI image is associated with a displacement vector. In another embodiment, the deformation field generator can output a sparse warped field, where displacement is defined only at a subset of control points, and the displacements of the remaining voxels are interpolated based on the control point displacements.

[0142] At operation 1112, a spatial transformer is used to apply the warped field generated by the deformation field generator to the MRI image to produce a deformed MRI image. The spatial transformer is a differentiable module that applies the warped field to the MRI image using a sampling kernel (such as bilinear or trilinear interpolation). This operation, based on the deformation field predicted by the deformation field generator, effectively warps or deforms the MRI image to align it with the corresponding US image. In an alternative implementation, instead of using a spatial transformer, a non-rigid registration algorithm (such as B-spline or free-form deformation) can be used to apply the warped field to the MRI image. This approach can provide additional flexibility and control over the deformation process, but may also introduce additional computational complexity.

[0143] At operation 1114, the loss is determined based on the similarity between the deformed MRI image and the US image. The loss function quantifies the degree of misalignment or dissimilarity between the deformed MRI image and the corresponding US image. In one embodiment, the loss function may be based on a similarity metric, such as normalized cross-correlation or mutual information, which measures the degree of alignment between the deformed MRI image and the US image. In another embodiment, the loss function may be based on a combination of a similarity metric and a regularization term, such as a smoothing constraint or sparsity constraint on the deformed field.

[0144] At operation 1116, the parameters of the spatial transformer, deformation field generator, US feature extractor, and MRI feature extractor are updated based on the loss determined in operation 1114. This update operation can be performed using an optimization algorithm, such as stochastic gradient descent or one of its variants, which adjusts the parameters of the various components of the MR-US deformation model in the direction that minimizes the loss function. In one implementation, the update operation can be performed using backpropagation, where the gradient of the loss function with respect to the parameters of the various components is computed and used to update the parameters in the opposite direction of the gradient, scaled by the learning rate. In another implementation, more advanced optimization techniques can be used to perform the update operation, such as momentum-based methods or adaptive learning rate methods, which can improve the convergence speed and stability of the training process.

[0145] Using different pairs of MRI and US images from the training dataset, operations 1102 to 1116 are repeated for multiple iterations. This iterative process allows the MR-US deformable model to progressively learn the complex relationships and deformities required for accurate registration of MRI and ultrasound data by minimizing the loss function over different sets of training examples. In one implementation, the training process can be performed in batch mode, where mini-batch pairs of MRI and US images are processed at each iteration, and the gradient of the loss function is accumulated on this mini-batch before updating the parameters of the MR-US deformable model. In another implementation, the training process can be performed in online mode, where the parameters of the MR-US deformable model are updated after processing each individual pair of MRI and US images, which potentially leads to faster convergence but also increases computational overhead.

[0146] Method 1100 for training MR-US deformable models is designed to fully leverage the capabilities of deep learning and unsupervised learning techniques to address the challenging problem of multimodal deformable image registration. By learning directly from the data in the ground-based deformable field without manual annotation, the MR-US deformable model can adapt to specific characteristics and variations present in the training dataset, thereby achieving accurate and robust registration of MRI and ultrasound data for image-guided interventions.

[0147] refer to Figure 12 The diagram illustrates a block diagram of the US-US deformability model training process 1200. This process is designed to train a deep learning-based deformability model that can accurately compute the transformations between pairs of US images representing different respiratory states or postures of the same patient. The trained US-US deformability model achieves registration between pre-interventional US images and interventional US images acquired during the procedure.

[0148] The training process 1200 takes a moving US image 1202 and a fixed US image 1204 as input. These images are typically 3D US images acquired during the pre-intervention phase, capturing anatomical regions of interest in different respiratory states or postures. The moving US image 1202 represents the image that needs to be deformed or warped to align with the fixed US image 1204, which serves as a reference or target image.

[0149] Moving US image 1202 and fixed US image 1204 are fed into deformation field generator 1206, which is a deep learning model responsible for predicting the deformation field required to align the moving US image 1202 with the fixed US image 1204. Deformation field generator 1206 can be based on various deep learning architectures, such as CNNs, recurrent neural networks (RNNs), or transformer models, depending on the specific requirements and characteristics of the US image data. In one implementation, deformation field generator 1206 is a CNN-based model that takes the moving US image 1202 and the fixed US image 1204 as input and outputs a dense deformation field, wherein each voxel in the moving US image 1202 is assigned a displacement vector that maps it to a corresponding position in the fixed US image 1204.

[0150] In another embodiment, the deformation field generator 1206 is a transformer-based model that leverages self-attention mechanisms to capture long-range dependencies and global context in the US image. This approach can be particularly useful for handling large deformations and complex anatomical movements that may occur between different breathing states or postures. The deformation field generator 1206 outputs a warped field 1208, which represents the predicted deformation field that maps the voxels of the moving US image 1202 to their corresponding locations in the stationary US image 1204.

[0151] The warped field 1208 is then applied to the moving US image 1202 via a spatial transformer 1210, which is a differentiable module that performs the warping or deformation operation using techniques such as bilinear or trilinear interpolation. The spatial transformer 1210 produces a deformed moving US image 1212, which is an approximation of the moving US image 1202 after deformation to align with the fixed US image 1204.

[0152] The deformed moving US image 1212 is then compared with the fixed US image 1204 using a similarity metric 1214, which measures the degree of alignment or similarity between the two images. In one embodiment, the similarity metric 1214 is based on normalized cross-correlation, which measures the similarity between the intensity patterns of the two images. A higher cross-correlation value indicates better alignment between the deformed moving US image 1212 and the fixed US image 1204. In another embodiment, the similarity metric 1214 is based on mutual information, which measures the statistical dependence between the intensity distributions of the two images. A higher mutual information value indicates better alignment between the deformed moving US image 1212 and the fixed US image 1204 and that they share similar intensity patterns.

[0153] The training process 1200 for the US-US deformable model includes an unsupervised training process that employs a similarity metric 1214 to evaluate the alignment between US image pairs (such as a moving US image 1202 and a stationary US image 1204). The similarity metric 1214 is computed without using labeled ground truth data or manually annotated deformable fields. Instead, the similarity metric 1214 measures the degree of alignment or similarity between the deformed moving US image 1212 and the stationary US image 1204 after applying a warped field predicted by the deformable field generator 1206. The similarity metric 1214 is used as a loss function or objective function during the training process, where the parameters of the deformable field generator 1206 are tuned to maximize the similarity between the deformed moving US image 1212 and the stationary US image 1204.

[0154] The training process 1200 is typically performed using a large dataset of US image pairs representing different respiratory states and postures of multiple patients. Various data augmentation techniques, such as random cropping, rotation, or intensity transformation, can be used to augment the dataset to improve the generalization ability of the trained US-US deformable model.

[0155] Once training process 1200 is complete, the trained US-US deformation model, including deformation field generator 1206 and spatial transformer 1210, can be used during the intervention phase to register pre-intervention US images with real-time acquired interventional US images. This registration process enables visualization of tissue features annotated on the pre-intervention US images within the context of the interventional US images, thereby providing valuable guidance for the interventional procedure.

[0156] refer to Figure 13 A flowchart of method 1300 for training a US-US deformability model to deformably register US images from the same imaging subject in multiple deformable states is shown. Method 1300 can be adopted by an image processing system to train the US-US deformability model using a training dataset comprising pairs of US images representing different respiratory states or postures of the same patient.

[0157] At operation 1302, the image processing system selects moving and fixed US images from the same imaging subject from a training dataset. The training dataset comprises multiple pairs of US images, each pair including a moving and a fixed US image. The moving and fixed US images represent the same anatomical region of the imaging subject but are captured in different deformed states, such as different respiratory phases or different patient postures. In one implementation, the moving and fixed US images are selected from a set of pre-interventional 4D US images acquired for the imaging subject. The 4D US images capture the anatomical region of interest over time, allowing for the extraction of individual 3D US images representing different respiratory states or postures. The moving and fixed US images can be selected from different time points within the 4D US images, ensuring that they represent different deformed states.

[0158] At operation 1304, the image processing system uses a warp field generator to map the moving US image and the fixed US image to a warp field. The warp field generator is a component of the US-US warp model, responsible for estimating the warp or transformation required to align the moving US image with the fixed US image. In one implementation, the warp field generator is based on a deep learning model, such as a CNN or a visual transformer. The warp field generator takes the moving US image and the fixed US image as input and outputs a warp field that describes the warp or displacement field required to map the voxels (3D pixels) of the moving US image to their corresponding positions in the fixed US image.

[0159] At operation 1306, the image processing system uses a spatial transformer to apply a warped field to the moving US image to produce a deformed moving US image. The spatial transformer is a component of the US-US deformable model, responsible for applying the warped field estimated by the deformable field generator to the moving US image. In one implementation, the spatial transformer is a differentiable module that applies the warped field to the moving US image using a sampling kernel (such as bilinear or trilinear interpolation). This allows the spatial transformer to be integrated into a deep learning pipeline and trained end-to-end with other components of the US-US deformable model. In another implementation, the spatial transformer applies the warped field to the moving US image using a non-rigid registration algorithm (such as B-spline or free-form deformation). This method can provide more accurate deformation but may be computationally more expensive than the differentiable spatial transformer approach.

[0160] At operation 1308, the image processing system determines the loss based on the similarity between the deformed moving US image and the fixed US image. The loss represents a measure of dissimilarity or misalignment between the deformed moving US image and the fixed US image and is used to guide the training of the US-US deformable model. In one implementation, a similarity measure (such as normalized cross-correlation, mutual information, or structural similarity index) is used to compute the loss. The similarity measure is applied to both the deformed moving US image and the fixed US image, and the resulting value is used as the loss. Higher values ​​of the similarity measure indicate better alignment between the two images and therefore a lower loss. In another implementation, a combination of multiple similarity measures is used to compute the loss, where each measure is weighted based on its importance or relevance to a specific application or anatomical region of interest. For example, measures such as normalized cross-correlation may be weighted more heavily in regions with high contrast and well-defined boundaries, while measures such as mutual information may be given higher weight in regions with low contrast or blurred boundaries.

[0161] At operation 1310, the image processing system updates the parameters of the spatial transformer and the deformation field generator based on the loss. This operation is part of the training process of the US-US deformation model, where the model's parameters are tuned to minimize the loss and improve the alignment between the deformed moving US image and the fixed US image. In one implementation, the parameter update is performed using a gradient-based optimization algorithm such as stochastic gradient descent (SGD) or Adam. Backpropagation is used to compute the gradient of the loss with respect to the parameters of the spatial transformer and the deformation field generator, and the parameters are updated in the direction that minimizes the loss. In another implementation, a meta-learning approach is used to perform the parameter update, where the US-US deformation model is trained to learn how to update its own parameters based on the loss and the input data. This approach can lead to faster convergence and better generalization performance, especially when the training data is limited or exhibits significant variability.

[0162] Method 1300 can be iterated multiple times using different pairs of moving and fixed US images from the training dataset. This iterative process allows the US-US deformation model to learn the complex deformations and transformations required to align US images representing different respiratory states or postures of the same patient. After training, the US-US deformation model can be combined with a trained MR-US deformation model to achieve real-time visualization of tissue features annotated on pre-interventional MRI images during interventional procedures guided by real-time US imaging. The US-US deformation model is responsible for aligning the pre-interventional US images with the interventional US images, while the MR-US deformation model is responsible for aligning the pre-interventional MRI images with the pre-interventional US images.

[0163] Figure 14 The effectiveness of the disclosed systems and methods for real-time visualization of tissue features annotated on pre-interventional MRI images during US imaging-guided interventional procedures is illustrated. Figure 14 Three images are presented: pre-interventional MRI image 1402, interventional 3D US image 1404, and visualization of pre-interventional MRI features superimposed on the interventional 3D US image 1406.

[0164] Pre-interventional MRI image 1402 is a high-resolution 3D image acquired prior to the interventional procedure. This image provides superior soft tissue contrast and lesion salience, enabling accurate labeling and annotation of tissue features of interest, such as organ boundaries or tumor boundaries. Figure 14 In the example shown, the pre-interventional MRI image 1402 depicts a cross-sectional view of an anatomical region with well-defined lesions or tumors within the tissue.

[0165] Prior to the intervention, the pre-interventional MRI images 1402 underwent an annotation process, in which tissue features of interest (in this case, lesion boundaries or tumor boundaries) were manually or automatically segmented and labeled. These annotated tissue features were then registered with multiple pre-interventional 3D US images using a trained MR-US deformable model, thereby capturing anatomical regions in various respiratory states and postures. This registration process yielded a series of deformable 3D MRI images, each corresponding to a specific pre-interventional 3D US image.

[0166] During the interventional procedure, 3D US images 1404 are acquired in real time using a 3D US probe or transducer. These 3D US images 1404 provide a live view of the anatomical regions, enabling real-time guidance and monitoring of the interventional procedure. However, due to inherent limitations of US imaging, such as poor soft tissue contrast and lesion salience, lesion or tumor boundaries may not be clearly visible in the 3D US images 1404.

[0167] To overcome this limitation and enhance visualization of tissue features during intervention, the disclosed system and method employ a trained US-US deformation model to register the interventional 3D US image 1404 with corresponding pre-interventional 3D US images from multiple pre-interventional 3D US images. This registration process determines a warpage field that describes the transformation required to align the pre-interventional 3D US image with the interventional 3D US image 1404, thus taking into account any anatomical movement or deformation that may have occurred between the acquisition of the two images.

[0168] The warped field is then applied to the deformed 3D MRI image corresponding to the registered pre-interventional 3D US image, thereby effectively aligning the pre-interventional MRI data with the interventional US data. This alignment process produces a visualization 1406 in which tissue features (such as lesion boundaries or tumor boundaries) annotated on the pre-interventional MRI image 1402 are superimposed onto the interventional 3D US image 1404.

[0169] In visualization 1406, annotated tissue features from pre-interventional MRI image 1402 are integrated with real-time interventional 3D US image 1404, providing clinicians with a comprehensive view of anatomical regions and tissue features of interest. This visualization enables precise guidance and targeting during interventional procedures because clinicians can clearly identify lesion or tumor boundaries even in cases of poor soft tissue contrast or lesion prominence in US images alone.

[0170] The overlay of pre-interventional MRI features onto the interventional 3D US images 1406 is achieved in real time, overcoming the computational limitations of conventional rigid or analytical deformation calculations. By fully utilizing a trained US-US deformation model and pre-computed deformed 3D MRI images, the disclosed system and method enable real-time visualization of tissue features during intervention, thereby improving the accuracy and efficiency of the procedure.

[0171] Figure 15 This illustrates the effect of a currently disclosed approach for registering a moving US image 1502 to a fixed US image 1504 to produce a registered image 1506. The moving US image 1502 represents an image to be deformed or warped to align with the fixed US image 1504, which serves as a reference or target image. In the context of this disclosure, the moving US image 1502 may correspond to a pre-interventional 3D US image acquired prior to intervention, while the fixed US image 1504 may represent an interventional 3D US image acquired during intervention.

[0172] Moving US image 1502 and fixed US image 1504 depict the same anatomical region of interest, but may exhibit differences due to factors such as patient movement, respiratory motion, or variations in the positioning of the imaging probe. These differences can lead to misalignment between the two images, making direct comparison or integration of the information they contain challenging.

[0173] To address this challenge, this disclosure employs a trained US-US deformation model, as described in previous sections, to register a moving US image 1502 to a stationary US image 1504. The US-US deformation model is a deep learning-based model that has been trained to compute the transformations required to align US image pairs representing different respiratory states or postures of the same patient.

[0174] During the registration process, the trained US-US deformation model takes a moving US image 1502 and a fixed US image 1504 as inputs and outputs a warp field that describes the deformation or displacement field required to align the moving US image 1502 with the fixed US image 1504. This warp field takes into account any anatomical movement or deformation that may have occurred between the acquisition of the two images, such as changes due to breathing, patient movement, or placement of interventional devices.

[0175] The warped field is then applied to the moving US image 1502 using a spatial transformer or a non-rigid registration algorithm (such as B-spline or freeform deformation). This process deforms or warps the moving US image 1502, thereby effectively aligning it with the fixed US image 1504, resulting in a registered image 1506.

[0176] The registered image 1506 represents the moving US image 1502 after it has been deformed or warped to match the fixed US image 1504. In the registered image 1506, the anatomical structures and features present in the moving US image 1502 are now aligned with their corresponding structures and features in the fixed US image 1504, as highlighted by white circles in each of the three images, thereby enabling direct comparison and integration of information between the two images.

[0177] It is important to note that, although Figure 15 The registration of US images is described, but this disclosure is not limited to this particular imaging modality. By training an appropriate deformable model tailored to the specific modality involved, the principles and techniques described herein can be extended to other imaging modalities, such as magnetic resonance imaging (MRI) or computed tomography (CT).

[0178] This disclosure also provides support for a method comprising: acquiring a pre-interventional three-dimensional (3D) image of an imaging subject in a first imaging modality, wherein the first imaging modality is one of magnetic resonance imaging (MRI), computed tomography (CT), and positron emission tomography (PET); acquiring a first interventional 3D ultrasound (US) image of the imaging subject during an intervention; registering the pre-interventional 3D image in the first imaging modality with the first interventional 3D US image using a first trained deformation model to generate a first deformed 3D image in the first imaging modality; acquiring a subsequent interventional 3D US image during the intervention; registering the first 3D US image with the subsequent interventional 3D US image using a trained US-US deformation model to determine a warping field; applying the warping field to the first deformed 3D image in the first imaging modality to generate a 3D image in the first imaging modality registered to the subsequent interventional 3D US image; and using the 3D image in the first imaging modality registered to the subsequent interventional 3D US image in the subsequent interventional 3D US image. Real-time visualization of tissue features annotated on pre-interventional 3D images on US images. In a first example of the method, the first imaging modality is MRI, and the first trained deformable model is a trained MR-US deformable model. In a second example of the method (optionally including the first example), the trained MR-US deformable model is trained by: initializing the weights of the MR-US deformable model with training data from a single patient, and iteratively retraining the MR-US deformable model with additional patient data, wherein each subsequent training session utilizes weights from previous training sessions, and wherein the training data includes gated or breath-hold 3D MRI images and multiple 3D US images representing different respiratory states and postures. In a third example of the method (optionally including one or both of the first and second examples), visualization of tissue features annotated on pre-interventional 3D images in the first imaging modality on subsequent interventional 3D US images includes: overlaying the 3D image in the first imaging modality, registered to the subsequent interventional 3D US image, onto the subsequent interventional 3D US image. In a fourth example of the method (optionally including one or more or each of the first to third examples), a first trained deformable model is used to register a pre-interventional 3D image in a first imaging modality with a first interventional 3D US image, or a trained US-US deformable model is used to register the first interventional 3D US image with a subsequent interventional 3D US image. The method also includes utilizing features derived from the pre-interventional 3D image in the first imaging modality and the first interventional 3D US image, including radiomics features and Gaussian mixture model tissue category probabilities.In a fifth example of the method (optionally including one or more, or each of, the first to fourth examples), a trained US-US deformable model is trained using pre-interventional 3DUS images captured from the imaging subject at a predetermined imaging frequency over a predetermined duration to capture multiple respiratory states. In a sixth example of the method (optionally including one or more, or each of, the first to fifth examples), a synchronous imaging system capable of simultaneously acquiring images in a first imaging modality and ultrasound images is used to perform the acquisition of pre-interventional 3D images of the imaging subject in the first imaging modality and the acquisition of first-interventional 3D US images of the imaging subject.

[0179] This disclosure also provides support for an image processing system comprising: a display device, a non-transitory memory including instructions, and a processor, wherein, when the instructions are executed, the processor causes the image processing system to: acquire pre-interventional three-dimensional (3D) magnetic resonance imaging (MRI) images of an imaging subject and capture multiple pre-interventional 3D ultrasound (US) images of the imaging subject with multiple respiratory states and postures; register the pre-interventional 3D MRI images with each of the multiple pre-interventional 3D US images using a trained MR-US deformable model to generate multiple deformed 3D MRI images; acquire an interventional 3D US image during the intervention; register the pre-interventional 3D US image from the multiple pre-interventional 3D US images with the interventional 3D US image using a trained US-US deformable model to determine a warping field; apply the warping field to the deformed 3D MRI image from the multiple deformed 3D MRI images corresponding to the pre-interventional 3D US image to generate a 3D MRI image registered to the interventional 3D US image; and, via the display device, use the image registered to the interventional 3D US image during the interventional 3D US image. 3D MRI images of US images, displaying tissue features annotated on the pre-interventional 3D MRI images on the interventional 3D US images. In a first example of the system, the system also includes a simultaneous MR and ultrasound imaging system with a 3D US probe configured to acquire pre-interventional 3D MRI images and multiple pre-interventional 3D US images. In a second example of the system (optionally including the first example), a trained MR-US deformable model and a trained US-US deformable model are incrementally trained using network model weights initialized from a previous training session with data from the imaging object. In a third example of the system (optionally including one or both of the first and second examples), the processor further enables the image processing system to annotate tissue features on the pre-interventional 3D MRI images, wherein the tissue features include organ boundaries or tumor boundaries. In a fourth example of the system (optionally including one or more, or each, of the first through third examples), the trained MR-US deformable model is trained by: initializing the weights of the MR-US deformable model with training data from a single patient, and iteratively retraining the MR-US deformable model with additional patient data, wherein each subsequent training session utilizes the weights from previous training sessions, and wherein the training data includes gated or breath-hold 3D MRI images and multiple 3DUS images representing different respiratory states and postures. In a fifth example of the system (optionally including one or more, or each, of the first through fourth examples), the trained US-US deformable model is trained using pre-interventional 3D US images captured at a predetermined imaging frequency over a predetermined duration to capture multiple respiratory states.In a sixth example of the system (optionally including one or more or each of the first to fifth examples), the processor further enables the image processing system to visualize tissue features annotated on the pre-interventional 3D MRI image on the interventional 3D US image by overlaying the registered 3D MRI image onto the interventional 3D US image.

[0180] This disclosure also provides support for a method for training a deformable model for multimodal image registration during image-guided interventions, the method comprising: acquiring pre-interventional three-dimensional (3D) magnetic resonance imaging (MRI) images and multiple pre-interventional four-dimensional (4D) ultrasound (US) images representing different respiratory states of the patient; initializing model weights for an MR-US deformable model and a US-US deformable model using data from at least one previous patient scan; training the MR-US deformable model to register the MR images to the US images using the pre-interventional 3D MRI images and multiple pre-interventional 4D US images, wherein the multiple pre-interventional 4D US images capture patient-specific anatomical movements due to breathing and placement of interventional devices; training the US-US deformable model to compute transformations between pairs of US images representing different respiratory states of the patient; and registering the pre-interventional MRI images with real-time US images acquired during the intervention using the trained MR-US deformable model and the trained US-US deformable model. In a first example of the method, training the MR-US deformable model includes an unsupervised training process that employs a similarity metric to evaluate the alignment between MR and US images, calculated without using labeled ground-based data. In a second example of the method (optionally including the first example), the similarity metric includes one or more of normalized cross-correlation, mutual information, or structural similarity indices, and the unsupervised training process adjusts the parameters of the MR-US deformable model based on a similarity metric determined across multiple pre-interventional 4D US images and pre-interventional 3D MRI images. In a third example of the method (optionally including one or both of the first and second examples), training the US-US deformable model includes an unsupervised training process that employs a similarity metric to evaluate the alignment between US image pairs, calculated without using labeled ground-based data. In a fourth example of the method (optionally including one or more, or each, of the first to third examples), the similarity measure includes one or more of normalized cross-correlation, mutual information, or structural similarity indices, and wherein the unsupervised training process adjusts the parameters of the US-US deformable model based on the similarity measure determined across multiple pairs of pre-interventional 4D US images. In a fifth example of the method (optionally including one or more, or each, of the first to fourth examples), the method further includes: annotating tissue features (including organ boundaries or tumor boundaries) on pre-interventional MRI images, and transferring the annotated organ boundaries or tumor boundaries to real-time US images acquired during the intervention using a trained MR-US deformable model and a trained US-US deformable model.

[0181] When describing elements of various embodiments of this disclosure, the articles “a,” “an,” and “the” are intended to indicate the presence of one or more such elements. The terms “first,” “second,” etc., do not indicate any order, quantity, or importance, but are used to distinguish one element from another. The terms “comprising,” “including,” and “having” are intended to be inclusive and indicate that additional elements may exist in addition to the listed elements. As used herein, the terms “connected to,” “coupled to,” etc., indicate that an object (e.g., a material, element, structure, component, etc.) may be connected to or coupled to another object, regardless of whether the one object is directly connected to or coupled to the other object, or whether one or more intervening objects exist between the one object and the other object. Furthermore, it should be understood that references to “an embodiment” or “an embodiment” of this disclosure are not intended to be construed as excluding the existence of additional embodiments also incorporating the referenced features.

[0182] In addition to any modifications previously indicated, many other variations and alternative arrangements can be devised by those skilled in the art without departing from the spirit and scope of this specification, and the appended claims are intended to cover such modifications and arrangements. Therefore, although the information has been described in particular and in detail above in conjunction with what is now considered to be the most practical and preferred aspects, it will be apparent to those skilled in the art that many modifications can be made without departing from the principles and concepts set forth herein, including but not limited to changes in form, function, mode of operation, and purpose. Likewise, as used herein, in all respects, embodiments and implementations are intended to be illustrative only and should not be construed as limiting in any way.

Claims

1. A method, the method comprising: Acquire a pre-intervention three-dimensional (3D) image (402) of the imaging object using a first imaging modality, wherein the first imaging modality is one of magnetic resonance imaging (MRI), computed tomography (CT), and positron emission tomography (PET); During the intervention, a first interventional 3D ultrasound (US) image (602) of the imaging object is acquired; The first trained deformable model (1000) is used to register the pre-intervention 3D image in the first imaging modality with the first intervention 3D US image, thereby generating a first deformable 3D image (604, 606) in the first imaging modality. During the intervention, subsequent interventional 3D US images (608) were acquired; The first interventional 3D US image is registered with the subsequent interventional 3D US image using a trained US-US deformation model (1200) to determine the warp field (610); The warping field is applied to the first deformed 3D image in the first imaging modality to generate a 3D image in the first imaging modality registered to the subsequently intervened 3D US image (612); and Using the 3D image registered to the subsequent interventional 3D US image in the first imaging modality, tissue features annotated on the pre-interventional 3D image are visualized in real time on the subsequent interventional 3D US image (614).

2. The method according to claim 1, wherein, The first imaging modality is MRI, and the first trained deformable model is a trained MR-US deformable model (1000).

3. The method according to claim 2, wherein, The trained MR-US deformable model (1000) is trained by the following process: initializing the weights of the MR-US deformable model (1000) with training data from a single patient, and iteratively retraining the MR-US deformable model (1000) with additional patient data, wherein each subsequent training session utilizes weights from previous training sessions, and wherein the training data includes gated or breath-hold 3D MRI images and multiple 3D US images representing different respiratory states and postures.

4. The method according to claim 1, wherein, Visualizing tissue features annotated on the pre-intervention 3D image in the first imaging modality on the subsequent intervention 3D US image (614) includes: superimposing the 3D image in the first imaging modality registered to the subsequent intervention 3D US image onto the subsequent intervention 3D US image.

5. The method according to claim 1, wherein, The method further includes registering the pre-intervention 3D image in the first imaging modality with the first interventional 3D US image using the first trained deformable model (1000) (604, 606), or registering the first interventional 3D US image with the subsequent interventional 3D US image using the trained US-US deformable model (610, 1200), the method further including utilizing features derived from the pre-intervention 3D image in the first imaging modality and the first interventional 3D US image, the features including radiomics features and Gaussian mixture model tissue category probabilities.

6. The method according to claim 1, wherein, The trained US-US deformation model (1200) is trained using pre-intervention 3D US images captured from the imaging subject at a predetermined imaging frequency over a predetermined duration to capture multiple respiratory states.

7. The method according to claim 1, wherein, The acquisition of the pre-intervention 3D image (402) of the imaging object and the acquisition of the first intervention 3D US image (602) of the imaging object in the first imaging mode are performed using a synchronous imaging system capable of simultaneously acquiring images in the first imaging modality and ultrasound images in the first imaging modality.

8. The method according to claim 1, wherein, The first trained deformable model (1000) and the trained US-US deformable model (1200) are incrementally trained using network model weights initialized from a previous training session with data from the imaged object.

9. The method according to claim 1, wherein, The first trained deformation model (1000) and the trained US-US deformation model (1200) are trained using an unsupervised learning approach that employs a similarity metric to evaluate alignment between images, wherein the similarity metric is calculated without using labeled ground-based deformation data.

10. A method for training a deformable model of multimodal image registration during image-guided intervention, the method comprising: Acquired pre-interventional three-dimensional (3D) magnetic resonance imaging (MRI) images (702) and multiple pre-interventional four-dimensional (4D) ultrasound (US) images representing different respiratory states of the patient (704); The model weights for the MR-US deformed model (1000) and the US-US deformed model (1200) were initialized using data from patients with at least one previous scan; The MR-US deformable model (1000) is trained using the pre-intervention 3D MRI images and the plurality of pre-intervention 4D US images to register the MR images to the US images, wherein the plurality of pre-intervention 4D US images capture patient-specific anatomical movements caused by breathing and placement of interventional devices (1100). The US-US deformation model (1200) is trained to compute the transformation between pairs of US images representing different respiratory states of the patient (1300); and The trained MR-US deformable model (1000) and the trained US-US deformable model (1200) are used to register the pre-interventional MRI images with real-time US images acquired during the intervention (600).

11. The method according to claim 10, wherein, Training the MR-US deformable model (1000, 1100) includes an unsupervised training process that uses a similarity metric to evaluate the alignment of the MR image with the US image. This similarity metric is calculated without using labeled ground-based data.

12. The method according to claim 11, wherein, The similarity measure includes one or more of normalized cross-correlation, mutual information, or structural similarity index, and wherein the unsupervised training process adjusts the parameters of the MR-US deformable model (1000) based on the similarity measure determined across the plurality of pre-interventional 4D US images and the pre-interventional 3D MRI images.

13. The method according to claim 10, wherein, Training the US-US deformable model (1200, 1300) includes an unsupervised training process that uses a similarity metric to evaluate the alignment between US image pairs, the similarity metric being calculated without using labeled ground truth data.

14. The method according to claim 13, wherein, The similarity measure includes one or more of normalized cross-correlation, mutual information, or structural similarity index, and wherein the unsupervised training process adjusts the parameters of the US-US deformation model (1200) based on the similarity measure determined across multiple pairs of the pre-intervention 4D US images.

15. The method according to claim 10, further comprising: Tissue features including organ or tumor boundaries are annotated on the pre-interventional MRI images (316, 404), and the annotated organ or tumor boundaries are transferred to the real-time US images acquired during the intervention (614) using the trained MR-US deformable model (1000) and the trained US-US deformable model (1200).