Multi-modal medical image registration and associated devices, systems, and methods

By defining a reference coordinate system in the ultrasound imaging system and utilizing deep learning prediction technology, efficient co-registration of ultrasound images with other modal images was achieved, solving the problem of the lack of an absolute 3D reference system in the ultrasound imaging system and improving the accuracy and efficiency of image fusion.

CN115210761BActive Publication Date: 2026-01-20KONINKLIJKE PHILIPS NV
View PDF 1 Cites 0 Cited by

Patent Information

Application Number
CN202180018558.0
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Priority Date
2020-03-05
Filing Date
2021-02-23
Publication Date
2026-01-20
Estimated Expiration
2041-02-23

AI Technical Summary

Technical Problem

Existing ultrasound imaging systems lack an absolute 3D reference frame, making it difficult to achieve efficient co-registration with other medical imaging modalities such as CT or MRI, resulting in high cost and error-prone image fusion.

Method used

A pose-based multimodal image co-registration technique is adopted. By defining reference coordinate systems in different imaging modalities and using deep learning prediction techniques to determine the image pose, spatial transformation and co-registration of ultrasound images with other modal images are achieved.

Benefits of technology

It improves the registration accuracy and efficiency of ultrasound images with other modal images, reduces hardware and time costs, and assists in medical examinations and interventional procedures.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN115210761B_ABST
    Figure CN115210761B_ABST
Patent Text Reader

Abstract

Devices, systems, and methods for multimodal medical image registration and association are provided. For example, a medical imaging method may include: receiving a first image of a patient's anatomical structure in a first imaging modality; receiving a second image of the patient's anatomical structure in a different second imaging modality; determining a first pose of the first image relative to a reference coordinate system of the patient's anatomical structure; determining a second pose of the second image relative to the reference coordinate system; determining co-registration data between the first image and the second image based on the first pose and the second pose; and outputting the first image, co-registered with the second image based on the co-registration data, to a display.
Need to check novelty before this filing date? Find Prior Art

Description

TECHNICAL FIELD

[0001] The present disclosure relates generally to ultrasound imaging. In particular, multi-modality medical image registration includes determining a position and / or orientation of an ultrasound image of an anatomical structure of interest and an anatomical structure image from a different medical imaging modality (e.g., magnetic resonance or MR, computed tomography or CT) relative to a reference or standardized local coordinate system of the anatomical structure, and spatially transforming the different images. BACKGROUND

[0002] Ultrasound imaging systems are widely used for medical imaging. For example, a medical ultrasound system can include an ultrasound transducer probe coupled to a processing system and one or more display devices. The ultrasound transducer probe can include an array of ultrasound transducer elements that emit sound waves into a patient and record sound waves reflected from internal anatomical structures within the patient, which can include tissue, blood vessels, and internal organs. The operations of emitting sound waves and / or receiving reflected sound waves or echo responses can be performed by the same set of ultrasound transducer elements or by different sets of ultrasound transducer elements. The processing system can apply beamforming, signal processing, and / or imaging processing to the received echo responses to create an image of the internal anatomical structures of the patient. The image can be presented to a clinician in the form of a brightness mode (B-mode) image, in which each pixel of the image is represented by a brightness level or intensity level corresponding to the echo intensity.

[0003] While ultrasound imaging is a safe and useful tool for diagnostic exams, interventions, and / or procedures, ultrasound imaging is based on the motion and positioning of a handheld ultrasound probe and thus lacks an absolute 3-dimensional (3D) frame of reference, whereas the anatomical context of other imaging modalities, such as computed tomography (CT) or magnetic resonance imaging (MRI), can provide an absolute 3-dimensional (3D) frame of reference. Co-registering and / or fusing a two-dimensional (2D) ultrasound image or a 3D ultrasound image with other modalities, such as CT or MRI, can require additional hardware, setup time, and thus can be costly. Additionally, there can be certain limitations on how to use the ultrasound probe to perform imaging to co-register and / or fuse the ultrasound image with other modalities. Co-registration between ultrasound images and another imaging modality is typically performed by identifying common fiducial points, common anatomical landmarks, and / or similarity measures based on image content. Feature-based or image content-based image registration can be time-consuming and can be prone to errors. SUMMARY

[0004] There remains a clinical need for improved systems and techniques for providing multi-modality image co-registration for medical imaging. Embodiments of the present disclosure provide techniques for multi-modality medical image co-registration. Disclosed embodiments define a reference or standardized local coordinate system in an anatomical structure of interest. The reference coordinate system can be represented in one form in a first imaging space of a first imaging modality and in another form in a second imaging space of a second imaging modality different from the first imaging modality for multi-modality image co-registration. For example, the first imaging modality can be two-dimensional (2D) or three-dimensional (3D) ultrasound imaging, while the second imaging modality can be 3D magnetic resonance (MR) imaging. Disclosed embodiments utilize a pose-based multi-modality image co-registration technique to register a first image of an anatomical structure in a first imaging modality with a second image of the anatomical structure in a second imaging modality. In this regard, a medical imaging system can acquire a first image of an anatomical structure in a first imaging modality using a first imaging system (e.g., an ultrasound imaging system) and a second image of the anatomical structure in a second imaging modality using a second imaging system (e.g., an MR imaging system). The medical imaging system determines a first pose of the first image relative to a reference coordinate system in an imaging space of the first imaging modality. The medical imaging system determines a second pose of the second image relative to the reference coordinate system in an imaging space of the second imaging modality. The medical imaging system determines a spatial transformation based on the first image pose and the second image pose. The medical system co-registers the first image of the first imaging modality with the second image of the second imaging modality by applying the spatial transformation to the first image or the second image. The co-registered or combined first and second images can be displayed to assist in a medical imaging examination and / or a medical intervention procedure. In some aspects, the present disclosure can use deep learning prediction techniques for image pose regression in a local reference coordinate system of an anatomical structure. Disclosed embodiments can be applied to co-registration of images of any suitable anatomical structure in two or more imaging modalities.

[0005] In some instances, a system for medical imaging includes a processor circuit in communication with a first imaging system of a first imaging modality and a second imaging system of a second imaging modality different from the first imaging modality, wherein the processor circuit is configured to receive a first image of an anatomical structure of a patient in the first imaging modality from the first imaging system, receive a second image of the anatomical structure of the patient in the second imaging modality from the second imaging system, determine a first pose of the first image relative to a reference coordinate system of the anatomical structure of the patient, determine a second pose of the second image relative to the reference coordinate system, determine co-registration data between the first image and the second image based on the first pose and the second pose, and output the first image co-registered with the second image based on the co-registration data to a display in communication with the processor circuit.

[0006] In some instances, a method of medical imaging includes receiving a first image of an anatomical structure of a patient in a first imaging modality at a processor circuit in communication with a first imaging system of the first imaging modality, receiving a second image of the anatomical structure of the patient in a second imaging modality at the processor circuit in communication with a second imaging system of the second imaging modality, determining a first pose of the first image relative to a reference coordinate system of the anatomical structure of the patient at the processor circuit, determining a second pose of the second image relative to the reference coordinate system at the processor circuit, determining co-registration data between the first image and the second image based on the first pose and the second pose at the processor circuit, and outputting the first image co-registered with the second image based on the co-registration data to a display in communication with the processor circuit.

[0007] Additional aspects, features, and advantages of the disclosure will become apparent from the following detailed description. BRIEF DESCRIPTION OF DRAWINGS

[0008] Illustrative embodiments of the disclosure will be described with reference to the accompanying drawings, of which:

[0009] Figure 1 is a schematic diagram of an ultrasound imaging system in accordance with aspects of the present disclosure.

[0010] Figure 2 is a schematic diagram of a multi-modality imaging system in accordance with aspects of the present disclosure.

[0011] Figure 3 is a schematic diagram of a multi-modality imaging co-registration scheme in accordance with aspects of the present disclosure.

[0012] Figure 4Aillustrates a three-dimensional (3D) image volume in an ultrasound imaging space, in accordance with various aspects of the present disclosure.

[0013] Figure 4B illustrates a 3D image volume in a magnetic resonance (MR) imaging space, in accordance with various aspects of the present disclosure.

[0014] Figure 4C illustrates a two-dimensional (2D) ultrasound image slice, in accordance with various aspects of the present disclosure.

[0015] Figure 4D illustrates a 2D MR image slice, in accordance with various aspects of the present disclosure.

[0016] Figure 5 is a schematic diagram of a deep learning network configuration, in accordance with various aspects of the present disclosure.

[0017] Figure 6 is a schematic diagram of a deep learning network training scheme, in accordance with various aspects of the present disclosure.

[0018] Figure 7 is a schematic diagram of a multi-modality imaging co-registration scheme, in accordance with various aspects of the present disclosure.

[0019] Figure 8 is a schematic diagram of a multi-modality imaging co-registration scheme, in accordance with various aspects of the present disclosure.

[0020] Figure 9 is a schematic diagram of a user interface of a medical system for providing multi-modality image registration, in accordance with various aspects of the present disclosure.

[0021] Figure 10 is a schematic diagram of a processor circuit, in accordance with embodiments of the present disclosure.

[0022] Figure 11 is a flowchart of a medical imaging method with multi-modality image co- registration, in accordance with various aspects of the present disclosure. DETAILED DESCRIPTION

[0023] To facilitate an understanding of the principles of the present disclosure, reference will now be made to the embodiments illustrated in the drawings and described below using specific language. It will nevertheless be understood that no limitation of the scope of the disclosure is intended. Alterations and further modifications of the described devices, systems, and methods, and any further applications of the principles of the disclosure are fully contemplated and included within the present disclosure as would normally be understood by those skilled in the art to which the present disclosure pertains. In particular, it is fully contemplated that features, components, and / or steps described with respect to one embodiment can be combined with features, components, and / or steps described with respect to other embodiments of the present disclosure. However, for the sake of brevity, these numerous iterations of the combinations will not be described separately.

[0024] Figure 1 is a schematic diagram of an ultrasound imaging system 100 in accordance with aspects of the present disclosure. The system 100 is used to scan a region or volume of a patient's body. The system 100 includes an ultrasound imaging probe 110 that communicates with a host 130 through a communication interface or link 120. The probe 110 includes a transducer array 112, a beamformer 114, a processor circuit 116, and a communication interface 118. The host 130 includes a display 132, a processor circuit 134, and a communication interface 136.

[0025] In an example embodiment, the probe 110 is an external ultrasound imaging device that includes a housing configured for handheld operation by a user. The transducer array 112 can be configured to obtain ultrasound data when the user grasps the housing of the probe 110 such that the transducer array 112 is positioned adjacent to and / or in contact with the skin of a patient. The probe 110 is configured to obtain ultrasound data of an anatomical structure within a patient when the probe 110 is positioned outside of the patient's body. In some embodiments, the probe 110 can be an external ultrasound probe suitable for abdominal exams, such as for diagnosing appendicitis or intussusception.

[0026] The transducer array 112 emits ultrasound signals toward an anatomical target 105 of a patient and receives echo signals reflected from the target 105 back to the transducer array 112. The ultrasound transducer array 112 can include any suitable number of acoustic elements, including one or more acoustic elements and / or multiple acoustic elements. In some instances, the transducer array 112 includes a single acoustic element. In some instances, the transducer array 112 can include an array of acoustic elements having any number of acoustic elements in any suitable configuration. For example, the transducer array 112 can include 1 acoustic element to 10,000 acoustic elements, including, for example, 2 acoustic elements, 4 acoustic elements, 36 acoustic elements, 64 acoustic elements, 128 acoustic elements, 500 acoustic elements, 812 acoustic elements, 1,000 acoustic elements, 3,000 acoustic elements, 8,000 acoustic elements, and / or other larger or smaller number values of acoustic elements. In some instances, the transducer array 112 can include an array of acoustic elements having any number of acoustic elements in any suitable configuration, such as a linear array, a planar array, a curved array, a curvilinear array, a peripheral array, a ring array, a phased array, a matrix array, a one-dimensional (ID) array, a 1.x-dimensional array (e.g., a 1.5D array), or a two-dimensional (2D) array. The array of transducer elements (e.g., one or more rows, one or more columns, and / or one or more orientations) can be controlled and activated collectively or independently. The transducer array 112 can be configured to obtain ID images, 2D images, and / or 3D images of the patient’s anatomy. In some embodiments, the transducer array 112 can include piezoelectric micromachined ultrasound transducers (PMUTs), capacitive micromachined ultrasound transducers (CMUTs), single crystals, lead zirconate titanate (PZT), PZT composites, other suitable transducer types, and / or combinations thereof.

[0027] The target 105 can include any anatomical structure of a patient suitable for ultrasound imaging examination, e.g., a blood vessel, a nerve fiber, an airway, a mitral valve leaflet, a cardiac structure, a prostate, an abdominal tissue structure, an appendix, a large intestine (or colon), a small intestine, a kidney, and / or a liver. In some aspects, the target 105 can include at least portions of a patient’s large intestine, small intestine, cecum pouch, appendix, ileum, liver, upper abdomen, and / or psoas muscle. The present disclosure can be implemented in the context of any number of anatomical locations and tissue types, including but not limited to organs (including liver, heart, kidney, gall bladder, pancreas, lung), ducts, intestines, nervous system structures (including brain, dural sac, spinal cord, and peripheral nerves); urinary tract, and valves within blood vessels, blood, chambers, or other portions of the heart, abdominal organs, and / or other systems of the body. In some embodiments, the target 105 can include a malignant tumor, e.g., a tumor, cyst, lesion, hemorrhage, or blood pool within any portion of the human anatomy. The anatomical structure can be a blood vessel, either an artery or a vein, as part of the patient’s vasculature (including the cardiac vasculature, peripheral vasculature, neural vasculature, renal vasculature, and / or any other suitable lumen within the body). In addition to native structures, the present disclosure can be implemented in the context of artificial structures such as, but not limited to, heart valves, stents, shunts, filters, implants, and other devices.

[0028] A beamformer 114 is coupled to the transducer array 112. The beamformer 114 controls the transducer array 112, e.g., for transmission of ultrasound signals and reception of ultrasound echo signals. The beamformer 114 provides image signals to the processor circuit 116 based on responses of the received ultrasound echo signals. The beamformer 114 can include multi-stage beamforming. Beamforming can reduce the number of signal lines for coupling to the processor circuit 116. In some embodiments, the transducer array 112 in combination with the beamformer 114 can be referred to as an ultrasound imaging component.

[0029] The processor circuit 116 is coupled to the beamformer 114. The processor circuit 116 can include a central processing unit (CPU), a graphics processing unit (GPU), a digital signal processor (DSP), an application-specific integrated circuit (ASIC), a controller, a field programmable gate array (FPGA) device, another hardware device, a firmware device, or any combination thereof configured to perform the operations described herein. The processor circuit 134 can also be implemented as a combination of computing devices, e.g., a combination of a DSP and a microprocessor, a plurality of microprocessors, one or more microprocessors in conjunction with a DSP core, or any other such configuration. The processor circuit 116 is configured to process the beamformed image signal. For example, the processor circuit 116 can perform filtering and / or quadrature demodulation to condition the image signal. The processor circuit 116 and / or 134 can be configured to control the array 112 to obtain ultrasound data associated with the target 105.

[0030] The communication interface 118 is coupled to the processor circuit 116. The communication interface 118 can include one or more transmitters, one or more receivers, one or more transceivers, and / or circuitry for transmitting and / or receiving communication signals. The communication interface 118 can include hardware and / or software components that implement a specific communication protocol suitable for transmitting signals to the host 130 over the communication link 120. The communication interface 118 can be referred to as a communication device or a communication interface module.

[0031] The communication link 120 can be any suitable communication link. For example, the communication link 120 can be a wired link, e.g., a universal serial bus (USB) link or an Ethernet link. Alternatively, the communication link 120 can be a wireless link, e.g., an ultra-wideband (UWB) link, an Institute of Electrical and Electronics Engineers (IEEE) 802.11 WiFi link, or a Bluetooth link.

[0032] At the host 130, the communication interface 136 can receive the image signal. The communication interface 136 can be substantially similar to the communication interface 118. The host 130 can be any suitable computing and display device, e.g., a workstation, a personal computer (PC), a laptop, a tablet, or a mobile phone.

[0033] The processor circuit 134 is coupled to the communication interface 136. The processor circuit 134 can be implemented as a combination of software components and hardware components. The processor circuit 134 can include a central processing unit (CPU), a graphics processing unit (GPU), a digital signal processor(s) (DSP), an application-specific integrated circuit (ASIC), a controller, a FPGA device, another hardware device, a firmware device, or any combination thereof configured to perform the operations described herein. The processor circuit 134 can also be implemented as a combination of computing devices, e.g., a combination of a DSP and a microprocessor, a plurality of microprocessors, one or more microprocessors in conjunction with a DSP core, or any other such configuration. The processor circuit 134 can be configured to generate image data from image signals received from the probe 110. The processor circuit 134 can apply advanced signal processing and / or image processing techniques to the image signals. In some embodiments, the processor circuit 134 can form a three-dimensional (3D) volumetric image from the image data. In some embodiments, the processor circuit 134 can perform real-time processing on the image data to provide a streaming video of ultrasound images of the target 105.

[0034] The display 132 is coupled to the processor circuit 134. The display 132 can be a monitor or any suitable display. The display 132 is configured to display ultrasound images of the target 105, image videos, and / or any imaging information.

[0035] In some aspects, the processor circuit 134 can implement one or more deep learning based prediction networks trained to predict an orientation of an input ultrasound image with respect to a coordinate system to assist an ultrasound physician in interpreting the ultrasound image and / or to provide co-registration information with another imaging modality, e.g., computed tomography (CT) or magnetic resonance imaging (MRI), which are described in more detail herein.

[0036] In some aspects, the system 100 can be used to collect ultrasound images to form a training dataset for training of a deep learning network. For example, the host 130 can include a memory 138, which can be any suitable storage device, such as a cache (e.g., of the processor circuit 134), random access memory (RAM), magneto resistive RAM (MRAM), read-only memory (ROM), programmable read-only memory (PROM), erasable programmable read-only memory (EPROM), electrically erasable programmable read-only memory (EEPROM), flash memory, solid-state memory device, hard disk drive, solid-state drive, other form of volatile or non-volatile memory, or a combination of different types of memory. The memory 138 can be configured to store an image dataset 140 to train a deep learning network to predict an image pose with respect to a certain reference coordinate system for multi-modality imaging co-registration, which are described in more detail herein.

[0037] As discussed above, ultrasound imaging is based on motion and positioning of a handheld ultrasound probe, and thus lacks an absolute 3-dimensional (3D) reference frame, whereas the anatomical context of other imaging modalities (e.g., CT or MRI) can provide an absolute 3-dimensional (3D) reference frame. Thus, it can be helpful to provide an ultrasound physician with co-registration information between an ultrasound image and an image of another imaging modality (e.g., CT or MRI). For example, an ultrasound image can be overlaid on a MR 3D image volume based on the co-registration information to assist the ultrasound physician in interpreting the ultrasound image, e.g., to determine an imaging view of the ultrasound image relative to the anatomical structure being imaged.

[0038] Figure 2 is a schematic diagram of a multi-modality imaging system 200 in accordance with aspects of the present disclosure. The system 200 is used to image anatomical structures of a patient using multiple imaging modalities (e.g., ultrasound, MR, CT, positron emission tomography (PET), single photon emission tomography (SPECT), cone beam CT (CBCT), and / or hybrid x-ray systems), and perform image co-registration between the multiple imaging modalities. To simplify illustration and discussion, Figure 2 The illustrated system 200 includes two imaging systems: an imaging system 210 of a first imaging modality and another imaging system 220 of a second imaging modality. However, the system 200 can include any suitable number (e.g., about 3 or 4 or more) of imaging systems having different imaging modalities, and can perform co-registration between images of different imaging modalities.

[0039] In some aspects, the first imaging modality can be associated with static images and the second imaging modality can be associated with mobile images. In some other aspects, the first imaging modality can be associated with mobile images and the second imaging modality can be associated with static images. In yet some other aspects, each of the first imaging modality and the second imaging modality can be associated with either mobile images or static images. In some aspects, the first imaging modality can be associated with static 3D imaging and the second imaging modality can be associated with mobile 3D imaging. In some other aspects, the first imaging modality can be associated with 3D mobile imaging and the second imaging modality can be associated with static 3D imaging. In some aspects, the first imaging modality can be associated with 3D imaging and the second imaging modality can be associated with 2D imaging. In some other aspects, the first imaging modality can be associated with 2D imaging and the second imaging modality can be associated with 3D imaging. In some aspects, the first imaging modality is one of: ultrasound, MR, CT, PET, SPECT, CBCT, or hybrid X-ray, and the second imaging modality is a different one of: ultrasound, MR, CT, PET, SPECT, CBCT, or hybrid X-ray.

[0040] The system 200 also includes a host 230, which is substantially similar to the host 130. In this regard, the host 230 can include a communication interface 236, a processor circuit 234, a display 232, and a memory 238, which are substantially similar to the communication interface 136, the processor circuit 134, the display 132, and the memory 138, respectively. The host 230 is communicatively coupled to the imaging systems 210 and 220 via the communication interface 236.

[0041] The imaging system 210 is configured to scan and acquire images 212 of an anatomical structure 205 of a patient in a first imaging modality. The imaging system 220 is configured to scan and acquire images 222 of the anatomical structure 205 of the patient in a second imaging modality. The anatomical structure 205 of the patient can be substantially similar to the target 105. The anatomical structure 205 of the patient can include any anatomical structure, such as, for example, a blood vessel, a nerve fiber, an airway, a mitral valve leaflet, a heart structure, a prostate, an abdominal tissue structure, an appendix, a large intestine (or colon), a small intestine, a kidney, a liver, and / or any organ or anatomical structure suitable for imaging in the first imaging modality and the second imaging modality. In some aspects, the images 212 are 3D image volumes and the images 222 are 3D image volumes. In some aspects, the images 212 are 3D image volumes and the images 222 are 2D image slices. In some aspects, the images 212 are 2D image slices and the images 222 are 3D image volumes.

[0042] In some aspects, the imaging system 210 is an ultrasound imaging system similar to the system 100, and the second imaging system 220 is an MR imaging system. Thus, the imaging system 210 can acquire and generate images 212 of the anatomical structure 205 by emitting ultrasound or sound waves toward the anatomical structure 205 and recording echoes reflected from the anatomical structure 205, as discussed above with reference to the system 100. The imaging system 220 can acquire and generate images 222 of the anatomical structure 205 by applying a magnetic field to force protons of the anatomical structure 205 to align with the magnetic field, applying a radio frequency current to excite the protons, stopping the radio frequency current and detecting energy released as the protons realign. The different scanning and / or image generation mechanisms used in ultrasound imaging and MR imaging can cause the images 212 and 222 to represent the same portion of the anatomical structure 205 in different perspectives or different views (shown in FIG. 2). Figures 4A-4D

[0043] Thus, the present disclosure provides techniques for performing image co- registration between images in different imaging modalities based on a pose (e.g., a position and / or an orientation) of an image of a patient’s anatomical structure (e.g., an organ) relative to a local reference coordinate system of the patient’s anatomical structure. Since the reference coordinate system is a coordinate system of the anatomical structure, the reference coordinate system is independent of any imaging modality. In some aspects, the present disclosure can use deep learning prediction techniques to regress a position and / or an orientation of a cross-sectional 2D imaging plane or 2D imaging slice of the anatomical structure in the local reference coordinate system of the anatomical structure.

[0044] For example, the processor circuit 234 is configured to receive the images 212 in the first imaging modality from the imaging system 210 and the images 222 in the second imaging modality from the imaging system 220. The processor circuit 234 is configured to determine a first pose of the images 212 relative to a reference coordinate system of the patient’s anatomical structure 205, determine a second pose of the images 222 relative to the reference coordinate system of the patient’s anatomical structure 205, and determine a co- registration between the images 212 and 222 based on the first pose and the second pose. The processor circuit 234 is further configured to output the images 212 co-registered with the images 222 based on the co- registration to the display 232 for display.

[0045] ​In some aspects, the processor circuit 234 is configured to determine the first pose and the second pose using deep learning prediction techniques. In this regard, the memory 238 is configured to store a deep learning network 240 and a deep learning network 250. The deep learning network 240 can be trained to regress an image pose of an input image (e.g., image 212) in the first imaging modality relative to a reference coordinate system. The deep learning network 250 can be trained to regress an image pose of an input image (e.g., image 222) in the second imaging modality relative to the reference coordinate system. The processor circuit 234 is configured to determine co- registration data by applying the deep learning network 240 and the deep learning network 250, as discussed in more detail below in Figure 3 and Figures 4A-4D .

[0046] Figure 3 are discussed with respect to Figures 4A-4D to illustrate multi-modality image co-registration based on regression of image poses in different imaging modalities in an anatomical coordinate system. Figure 3 is a schematic diagram of a multi-modality image co-registration scheme 300 in accordance with aspects of the present disclosure. The scheme 300 is implemented by the system 200. In particular, the processor circuit 234 can implement multi-modality image co-registration as shown in the scheme 300. The scheme 300 includes two prediction paths, one path including a deep learning network 240 trained to perform pose regression for a first imaging modality 306 of the imaging system 210 (shown in the top path), and another path including a deep learning network 250 trained to perform pose regression for a second imaging modality 308 of the imaging system 220 (shown in the bottom path). In Figure 3 the example shown, the imaging modality 306 can be ultrasound imaging, and the imaging modality 308 can be MR imaging. The scheme 300 further includes a multi-modality image co- registration controller 330 coupled to the deep learning network 240 and the deep learning network 250. The multi-modality image co-registration controller 330 can be similar to the processor circuits 134 and 234, and can include hardware components and / or software components.

[0047] The scheme 300 defines a common reference coordinate system for the anatomical structure 205 used for image co-registration between different imaging modalities. For example, the deep learning network 240 is trained to receive an input image 212 in an imaging modality 306 and output an image pose 310 of the input image 212 relative to the common reference coordinate system. Similarly, the deep learning network 250 is trained to receive an input image 222 in an imaging modality 308 and output an image pose 320 of the input image 222 relative to the common reference coordinate system. The image pose 310 can include a spatial transformation that includes at least one of a rotation component or a translation component that transforms the image 212 from a coordinate system of an imaging space in the first imaging modality to the reference coordinate system. The image pose 320 can include a spatial transformation that includes at least one of a rotation component or a translation component that transforms the image 222 from a coordinate system of an imaging space in the second imaging modality to the reference coordinate system. In some aspects, each of the image pose 310 and the image pose 320 includes a 6 degrees of freedom (6DOF) transformation matrix that includes 3 rotation components (e.g., indicating orientation) and 3 translation components (e.g., indicating position). Reference is made to the following discussion of different coordinate systems and transformations. Figures 4A-4D Different coordinate systems and transformations are discussed.

[0048] Figures 4A-4D The use of a reference coordinate system defined to a particular feature or portion of the patient’s anatomy (e.g., the prostate) used for multi-modal image co-registration of images of the prostate acquired from ultrasound imaging and MR imaging is illustrated. To simplify the illustration and discussion, Figures 4A-4D Multi-modal image co-registration of images of a prostate of a patient acquired from ultrasound imaging and MR imaging is illustrated. However, the scheme 300 can be applied to co-registration of images of any anatomical structure (e.g., heart, liver, lung, blood vessels,...) acquired using any suitable imaging modality using similar coordinate system transformations discussed below.

[0049] Figure 4A A 3D image volume 410 in an ultrasound imaging space is illustrated in accordance with aspects of the present disclosure. The 3D image volume 410 includes an image of a prostate 430. The prostate 430 can correspond to the anatomical structure 205. A reference coordinate system 414 is defined for the prostate 430 (referred to as organ US ) in the ultrasound imaging space. In some aspects, the reference coordinate system 414 can be defined based on a centroid of the prostate 430. For example, an origin of the reference coordinate system 414 can correspond to the centroid of the prostate 430. Thus, the reference coordinate system 414 is a local coordinate system for the prostate 430.

[0050] Figure 4BThe illustration shows a 3D image volume 420 in MR imaging space according to various aspects of this disclosure. The 3D image volume 420 includes an image of the same prostate 430. The prostate 430 (referred to as organ) is shown in MR imaging space. MRI A reference coordinate system 424 is defined. Reference coordinate system 424 in MR imaging space is the same as reference coordinate system 414 in ultrasound imaging space. For example, the origin of reference coordinate system 424 corresponds to the centroid of the prostate. In some instances, reference coordinate systems 414 and 424 are identical because they are both defined by the patient's anatomy (e.g., the origin is defined based on the organ of interest). Therefore, the origins of reference coordinate systems 414 and 424 can have the same location (e.g., the centroid in the prostate anatomy), and the x, y, and z axes can be oriented in the same direction (e.g., the x-axis along the first principal axis of the prostate, and the y-axis in the direction of the supine patient). Figures 4A-4B The diagram depicts reference coordinate systems 414 and 424, derived from transrectal ultrasound viewing of the prostate from below, while the MRI view is obtained from above (or side view). In other instances, the orientation of the x, y, and z axes in reference coordinate system 424 may differ from that in reference coordinate system 414. By using different orientations, translation matrices can be used to align the x, y, and z axes of reference coordinate systems 414 and 424.

[0051] Figure 4C The illustration shows a 2D ultrasound image slice 412 (referred to as a plane) of a 3D image volume 410 in ultrasound imaging space according to various aspects of the present disclosure. US The cross-sectional slice 412 is defined in coordinate system 416 of the ultrasound imaging space.

[0052] Figure 4D The illustration shows a 2D MR image slice 422 (referred to as a plane) of a 3D image volume 420 in MR imaging space according to various aspects of this disclosure. MRI Cross-sectional slice 422 is in coordinate system 426 of the MR imaging space.

[0053] To explain Figure 3 In the multimodal image co-registration, image 212 in scheme 300 can correspond to 2D image slice 412, and image 222 in scheme 300 can correspond to 2D image slice 422. Deep learning network 240 is trained to regress the pose 310 of image 212 in the local coordinate system (of imaging modality 306) of an organ (e.g., prostate 430). In other words, deep learning network 240 predicts (plane...) USThe ultrasound imaging spatial coordinate system 416 and the local organ reference coordinate system 414organ US The transformation between 402 and 402. This can be achieved through... To represent transformation 402. It can be achieved through... This represents the pose 310 predicted or estimated by the deep learning network 240.

[0054] Similarly, deep learning network 250 is trained to regress the pose 320 of image 222 in the local coordinate system (imaging modality 308) of an organ (e.g., prostate 430). In other words, deep learning network 250 predicts (plane MRI The MR imaging spatial coordinate system 426 and the local organ reference coordinate system 424organ MRI The transformation between them is 404. It can be achieved through... This can be used to represent transformation 404. It can be done through... This represents the pose 320 predicted or estimated by the deep learning network 250.

[0055] The multimodal image co-registration controller 330 is configured to receive pose 310 from the deep learning network 240 and pose 320 from the deep learning network 250. The multimodal image co-registration controller 330 is configured to compute a multimodal registration matrix (e.g., a spatial transformation matrix) as follows:

[0056]

[0057] in, mri T us This represents the multimodal registration matrix that transforms coordinate system 416 in ultrasound imaging space into coordinate system 426 in MR imaging space. This indicates the transformation of ultrasound image slice 412 or image 212 into ultrasound coordinate system 416. This indicates the transformation of MR slice 422 or image 222 to MR coordinate system 426, and This indicates that from coordinate system 414organ US Transform to coordinate system 424organ MRI The transformation. Since coordinate system 414 and coordinate system 424 refer to the same local anatomical structure coordinate system, therefore It is an identity matrix.

[0058] To register image 212 (e.g., ultrasound image slice 412) with image 222 (e.g., MR image slice 422), the multimodal image co-registration controller 330 is also configured to register the image 212 (e.g., ultrasound image slice 412) with image 222 (e.g., MR image slice 422), the multimodal image co-registration controller 330 is further configured to register the image 212 (e.g., ultrasound image slice 412) with the image 222 (e.g., MR image slice 422) using the transformation matrix in formula (1). mri T usis applied to the image 212 to perform a spatial transformation on the image 212. The co-registered images 212 and 222 (shown in FIG. 13B) can be displayed on a display (e.g., the displays 132 and 232) to assist a clinician in performing an imaging and / or medical procedure (e.g., a biopsy and / or a medical treatment). Figure 9

[0059] In some other aspects, the image 212 can be a 3D moving image volume of the imaging modality 306 similar to the ultrasound 3D volume 410, and the image 222 can be a 3D static image volume of the imaging modality 308 similar to the MR 3D volume 420. The multi-modality image co- registration controller 330 is configured to define or select an arbitrary 2D slice in the ultrasound image volume 410, define or select an arbitrary 2D slice in the MR image volume 420, determine a multi-modality registration matrix to co-register each ultrasound image slice with the MR image slice as shown in equation (1) above.

[0060] In some aspects, the scheme 300 can be applied to co-register images of a patient’s heart obtained from different modalities. To co-register images of a heart, a local organ reference coordinate system (e.g., the reference coordinate systems 414 and 424) can be defined by placing an origin at the center of the left ventricle of the heart and defining x-y axes to be coplanar with a plane defined by the left ventricle center, the left atrium center, and the right ventricle center, where the x-axis points from the left ventricle to the right ventricle, the y-axis points from the left ventricle to the left atrium, and the z-axis is collinear with the normal of the plane. This imaging plane is commonly referred to as the apical 4-chamber view of the heart.

[0061] While the scheme 300 is described in the context of performing co-registration between two imaging modalities, the scheme 300 can also be applied to perform co-registration between any suitable number (e.g., about 3, 4, or more) of imaging modalities using substantially similar mechanisms. In general, for each imaging modality, an image pose of an input image in the imaging modality can be determined relative to a reference coordinate system of an anatomical structure in the imaging space of the imaging modality, and the multi-modality image co-registration controller 330 selects a reference image of a primary imaging modality, determines a spatial transformation matrix (as shown in equation (1)) to co-register the images of each imaging modality with the reference image.

[0062] Figure 5 is a schematic diagram of a deep learning network configuration 500 in accordance with aspects of the present disclosure. The configuration 500 can be implemented by a deep learning network (e.g., the deep learning networks 240 and / or 250). The configuration 500 includes a deep learning network 510 that includes one or more convolutional neural networks (CNNs) 512. To simplify illustration and discussion, Figure 5 ​A CNN 512 is illustrated. However, embodiments can be scaled to include any suitable number (e.g., about 2, 3, or more) of CNNs 512. The configuration 500 can be trained to regress image poses in a local organ coordinate system (e.g., the reference coordinate systems 414 and 424) for a particular imaging modality, as described in more detail below.

[0063] The CNN 512 can include a set of N convolutional layers 520 followed by a set of K fully connected layers 530, where N and K can be any positive integer. The convolutional layers 520 are shown as 520 (1) through 520 (N) . The fully connected layers 530 are shown as 530 (1) through 530 (K) . Each convolutional layer 520 can include a set of filters 522 configured to extract features from the input 502 (e.g., the images 212, 222, 412, and / or 422). The values N and K, as well as the size of the filters 522, can vary according to embodiments. In some instances, the convolutional layers 520 (1) through 520 (N) and the fully connected layers 530 (1) through 530 (K-1) may utilize a non-linear activation function (e.g., ReLU - rectified linear unit) and / or batch normalization and / or dropout and / or merge. The fully connected layers 530 can be non-linear and can progressively shrink a high-dimensional output to the dimension of a prediction result (e.g., the output 540).

[0064] The output 540 can correspond to the poses 310 and / or 320 discussed above with reference to Figure 3 . The output 540 can be a transformation matrix including a rotation component and / or a translation component that can transform the input image 502 from an imaging space (of the imaging modality used to acquire the input image 502) into a local organ coordinate system.

[0065] Figure 6is a schematic illustration of a deep learning network training scheme 600 in accordance with various aspects of the present disclosure. The scheme 600 can be implemented by the systems 100 and / or 200. In particular, the scheme 600 can be implemented to train multiple deep learning networks to regress image poses in a reference or organ coordinate system (e.g., the reference coordinate systems 414 and 424). Each deep learning network can be trained separately for a particular imaging modality. For example, for co-registration between MR images and ultrasound images, one deep learning network can be trained on ultrasound images and another network can be trained on MR images. To simplify illustration and discussion, the scheme 600 is discussed in the context of training a deep learning network 240 based on ultrasound images and training a deep learning network 250 based on MR images, where the deep learning networks 240 and 250 are configured as shown in Figure 5 However, the training scheme 600 can be applied to trained deep learning networks of any network architecture and for any imaging modalities for multi-modality image co-registration.

[0066] In the illustrative example of Figure 6 , the deep learning network 240 is trained to regress poses of ultrasound images of the prostate 430 in a local reference coordinate system 414 of the prostate 430 (shown in the upper half of Figure 6 ). The deep learning network 250 is trained to regress poses of MR images of the prostate 430 in a local reference coordinate system 424 of the prostate 430 (shown in the lower half of Figure 6 ). As discussed above, the local reference coordinate system 414 and the local reference coordinate system 424 locally correspond to the same reference coordinate system at the prostate 430.

[0067] To train the network 240, a collection of 2D cross-sectional planes or image slices (referred to as I US ) generated from 3D ultrasound imaging of the prostate 430 are collected. In this regard, the 3D ultrasound imaging is used to acquire a 3D imaging volume 410 of the prostate 430. The 2D cross-sectional planes (e.g., 2D ultrasound image slices 412) can be randomly selected from the 3D imaging volume 410. The 2D cross-sectional planes are defined by a 6DOF transformation matrix T US ∈ SE(3) (describing the translation and rotation of the plane) in the local organ coordinate system 414. Each 2D image I US is labeled with the transformation matrix T US . Alternatively, 2D images of the prostate 430 can be acquired using 2D ultrasound imaging with tracking and the pose of each image in the local organ coordinate system is determined based on the tracking. 2D ultrasound imaging can provide higher resolution 2D images compared to the 2D cross-sectional planes obtained by slicing the 3D imaging volume 410.

[0068] A training data set 602 can be generated from 2D cross-sectional image slices and corresponding transformations to form ultrasound image-transformation pairs. Each pair includes a 2D ultrasound image slice I US and a corresponding transformation matrix T US , shown for example as (I US , T US ). For example, the training data set 602 can include 2D ultrasound images 603 annotated or labeled with a corresponding transformation describing the translation and rotation of the image 603 in the local organ coordinate system 414. The labeled images 603 are the inputs used to train the deep learning network 240. The labeled transformations T US serve as ground truth for training the deep learning network 240.

[0069] The deep learning network 240 can be applied (e.g., using forward propagation) to each image 603 in the data set 602 to obtain an output 604 for the input image 603. The training component 610 adjusts the coefficients of the filters 522 in the convolutional layers 520 and the weights in the fully connected layers 530, for example, by using backpropagation to minimize a prediction error (e.g., a difference between the ground truth T US and the predicted result 604). The predicted result 604 can include a transformation matrix The transformation matrix is used to transform the input image 603 into the local reference coordinate system 414 of the prostate 430. In some instances, the training component 610 adjusts the coefficients of the filters 522 in the convolutional layers 520 and the weights in the fully connected layers 530 for each input image to minimize a prediction error (e.g., a difference between T US and ). In other instances, the training component 610 applies a batch training procedure to adjust the coefficients of the filters 522 in the convolutional layers 520 and the weights in the fully connected layers 530 based on prediction errors obtained from a set of input images.

[0070] The network 250 can be trained using a mechanism that is substantially similar to the mechanism discussed above for the network 240. For example, the network 250 can be trained on a training data set 606 that includes 2D MR image slices 607 (e.g., 2D MR image slices 422) labeled with a corresponding transformation matrix T MR . The 2D cross-sectional MR image slices 607 can be obtained by randomly selecting a 3D cross-sectional plane (multiplanar reformatting). The 2D cross-sectional MR image slices 607 are defined by a 6DOF transformation matrix T MR ∈ SE(3) (describing the translation and rotation of the plane) in the local organ coordinate system 424.

[0071] The deep learning network 250 can be applied (e.g., using forward propagation) to each image 607 in the dataset 606 to obtain an output 608 for the input image 607. The training component 620 adjusts coefficients of the filters 522 in the convolutional layers 520 and weights in the fully connected layers 530, e.g., by using backpropagation to minimize a prediction error (e.g., a difference between a predicted result 604 and a true result T MR The predicted result 604 can include a transformation matrix T The transformation matrix T for transforming the input image 607 into the local reference coordinate system 424 of the prostate 430. The training component 620 can adjust coefficients of the filters 522 in the convolutional layers 520 and weights in the fully connected layers 530 for each input image or for each batch of input images.

[0072] In some aspects, in addition to translation and rotation, the transformation matrix T US and each of the transformation matrices T MR may include a shear component and a scaling component. Thus, the co-registration between ultrasound and MR can be an affine co-registration rather than a rigid co-registration.

[0073] After the deep learning networks 240 and 250 are trained, the scheme 300 can be applied in an application or inference phase for medical examination and / or guidance. In some aspects, the scheme 300 can be applied to co-register two 3D image volumes of different imaging modalities (e.g., MR and ultrasound), e.g., by co-registering 2D image slices of a 3D image volume in one imaging modality with 2D image slices of another 3D image volume in another imaging modality. In some other aspects, instead of this case of using two 3D volumes as input in the application / inference phase, one of the modalities can be a 2D imaging modality, and images of the 2D imaging modality can be provided for real-time inference of registration with other 3D imaging volumes of the 3D modality and for real-time co-display (as shown in FIG. 3) with the 3D imaging volumes. Figure 7

[0074] Figure 7 ​is a schematic illustration of a multi-modality imaging co-registration scheme 700 in accordance with aspects of the present disclosure. The scheme 700 is implemented by the system 200. In particular, the system 200 can provide real-time co-registration of 2D images of a 2D imaging modality with a 3D image volume of a 3D imaging modality to, for example, provide imaging guidance as illustrated in the scheme 700. To simplify discussion and illustration, the scheme 700 is described in the context of providing real-time co-registration of 2D ultrasound images with a 3D MR image volume. However, the scheme 700 can be applied to co-registration of 2D images of any 2D imaging modality with a 3D image volume of any 3D imaging modality.

[0075] In the scheme 700, 2D ultrasound image slices 702 are acquired in real-time, for example, by using the imaging system 210 and / or 100 with the probe 110 in an arbitrary pose with respect to the target organ (e.g., the prostate 430) in a freehand manner, but within the range of poses extracted from the corresponding 3D ultrasound volume during the training (in the scheme 600). In some other aspects, instead of extracting a large number of cross-sectional slices from the 3D ultrasound volume in an arbitrary manner during the training phase, the poses of the extracted slices can be tailored to cover the range of desired poses encountered during real-time scanning in the application phase. The scheme 700 also acquires a 3D MR image volume 704 of the organ, for example, using the imaging system 220 with a MR scanner.

[0076] The scheme 700 applies the trained deep learning network 240 to the 2D ultrasound images in real-time to estimate the pose of the 2D ultrasound images 702 in the organ coordinate system 414. Similarly, the scheme 700 applies the trained deep learning network 250 to the 3D MR image volume 704 to estimate the transformation of the 3D MR image volume 704 from the MR imaging space to the organ space. At this point, the 3D MR image volume 704 can be acquired prior to real-time ultrasound imaging. Thus, the transformation of the 3D MR image volume 704 from the MR imaging space to the organ space can be performed after the 3D MR image volume 704 is acquired and used for co-registration during real-time 2D ultrasound imaging. At this point, the scheme 700 applies the multi-modality image co-registration controller 330 to the pose estimation from 2D ultrasound imaging and 3D MR imaging to provide real-time estimation of the pose of the 2D ultrasound images 702 with respect to the pre-acquired MR image volume 704. The multi-modality image co-registration controller 330 can apply the above equation (1) to determine the transformation from the ultrasound imaging space to the MR imaging space and perform co-registration based on the transformation as discussed above with reference to the scheme 600. Figure 3

[0077] ​The scheme 700 can co-display the 2D ultrasound image 702 with the 3D MR image volume 704 on a display (e.g., the display 132 or 232). For example, the 2D ultrasound image 702 can be overlaid on top of the 3D MR image volume 704 (as shown), providing the clinician with real-time 2D imaging location information of the acquired 2D ultrasound image 702 with respect to the organ being imaged. The location information can assist the clinician in steering the probe to reach a target imaging view for the ultrasound examination or assist the clinician in performing a medical procedure (e.g., a biopsy). Figure 9

[0078] In some aspects, the above-discussed pose-based multi-modal image registration can be used in conjunction with a feature-based or image content-based multi-modal image registration or any other multi-modal image registration to provide co-registration with high accuracy and robustness. The accuracy of the co-registration can depend on the initial pose distance between the images to be registered. For example, a feature-based or image content-based multi-modal image registration algorithm typically has a "capture range" of initial pose distances within which the algorithm tends to converge to the correct solution, and if the initial pose distance is outside the capture range, the algorithm can fail to converge or can converge to an incorrect local minimum. Thus, the above-discussed pose-based multi-modal image registration can be used to align two images of different imaging modalities to be in close alignment, e.g., to satisfy the capture range of the feature-based or image content-based multi-modal image registration algorithm, before applying the feature-based or image content-based multi-modal image registration algorithm.

[0079] Figure 8 is a schematic diagram of a multi-modal imaging co-registration scheme 800 in accordance with aspects of the present disclosure. The scheme 800 is implemented by the system 200. In particular, the system 200 can apply a pose-based multi-modal image registration to align two images of different imaging modalities to be in close alignment, and then apply a multi-modal image registration refinement as shown in the scheme 800 to provide co-registration with high accuracy.

[0080] As shown, the scheme 800 applies a pose-based multi-modal image registration 810 to the image 212 of the imaging modality 306 and the image 222 of the imaging modality 308. The pose-based multi-modal image registration 810 can implement the above-discussed pose-based multi-modal image registration 410. The pose-based multi-modal image registration 810 can align the image 212 of the imaging modality 306 and the image 222 of the imaging modality 308 to be in close alignment, e.g., to satisfy the capture range of a feature-based or image content-based multi-modal image registration algorithm. Figure 3 ​The discussed scheme 300. For example, the pose of image 212 (e.g., of the prostate 430) is determined relative to a local organ reference coordinate system (e.g., reference coordinate system 414) in the imaging space of imaging modality 306. Similarly, the pose of image 222 (e.g., of the prostate 430) is determined relative to a local organ reference coordinate system (e.g., reference coordinate system 424) in the imaging space of imaging modality 308. The pose-based multi-modality image registration 810 aligns image 212 and image 222 based on the determined poses for image 212 and image 222, e.g., by performing a spatial transformation to provide a co-registered estimate 812. In some instances, after the spatial transformation, image 212 can be aligned with image 222 with a translational misalignment of less than about 30 mm and / or a rotational misalignment of less than about 30 degrees.

[0081] After performing the pose-based multi-modality image registration 810, the scheme 800 applies a multi-modality image registration refinement 820 to the co-registered images (e.g., co-registered estimate 812). In some aspects, the multi-modality image registration refinement 820 can implement a feature-based or image content-based multi-modality image registration, where the registration can be based on a similarity measure (of anatomical features or landmarks) between image 212 and image 222.

[0082] In some other aspects, the multi-modality image registration refinement 820 can implement another deep learning-based image co-registration algorithm. For example, automated multi-modality image registration in fusion-guided interventions can be based on iterative predictions according to a stacked deep learning network. In some aspects, to train the stacked deep learning network, the pose-based multi-modality image registration 810 can be applied to a training dataset for the stacked deep learning network such that the image poses of the training dataset are within a certain alignment range prior to training. In some aspects, the prediction error from the pose-based multi-modality image registration 810 can be computed, e.g., by comparing the predicted registration to the ground truth registration. The range of pose errors can be modeled with a parametric distribution (e.g., a uniform distribution with minimum and maximum pose parameter error values, or a Gaussian distribution with an expected pose parameter mean and standard deviation). The pose parameters can be used to generate a training dataset with artificially created misaligned registrations between modality 306 and modality 308. The training dataset can be used to train the stacked deep learning network.

[0083] Figure 9 is a schematic diagram of a user interface 900 for a medical system for providing multi-modality image registration in accordance with aspects of the present disclosure. The user interface 900 can be implemented by the system 200. In particular, the system 200 can implement the user interface 900 to provide a multi-modality image registration in accordance with the above regarding the scheme 300. The user interface 900 can be implemented by the system 200 to provide a multi-modality image registration in accordance with the above regarding the scheme 300. Figure 3、 Figure 7 and / or Figure 8 The discussed schemes 300, 700, and / or 800 determine a multi-modality image registration. The user interface 900 can be displayed on the display 232.

[0084] As shown, the user interface 900 includes an ultrasound image 910 and an MR image 920 of the same patient’s anatomy. The ultrasound image 910 and the MR image 920 can be displayed based on a co-registration performed using the schemes 300, 700, and / or 800. The user interface 900 also displays an indicator 912 in the image 910 and an indicator 922 in the image 920 according to the co-registration. The indicator 912 can correspond to the indicator 922, but each indicator is displayed in the corresponding image according to the co-registration to indicate the same portion of the anatomy in each image 910, 920.

[0085] In some other aspects, the user interface 900 can display the images 910 and 920 as color-coded images or a checkerboard overlay. For color-coded images, the display can color-code different portions of the anatomy and use the same color to represent the same portion on the images 910 and 920. For a checkerboard overlay, the user interface 900 can display sub-images that overlay the images 910 and 920.

[0086] Figure 10 is a schematic diagram of a processor circuit 1000 according to embodiments of the present disclosure. The processor circuit 1000 can be implemented in Figure 1 a probe 110 and / or a host 130 of the system 100, Figure 2 a host 230 and / or Figure 3 a multi-modality image co-registration controller 330 of the system 230. In an example, the processor circuit 1000 can be in communication with multiple imaging scanners of different imaging modalities (e.g., a transducer array 112 in a probe 110, an MR image scanner). As shown, the processor circuit 1000 can include a processor 1060, a memory 1064, and a communication module 1068. These elements can be in direct communication with one another or in indirect communication (e.g., via one or more buses).

[0087] The processor 1060 can include a CPU, a GPU, a DSP, an application-specific integrated circuit (ASIC), a controller, a FPGA, another hardware device, a firmware device, or any combination thereof configured to perform the operations described herein (e.g., Figures 2-9 and Figure 11 various aspects of the system 100). The processor 1060 can also be implemented as a combination of computing devices, e.g., a combination of a DSP and a microprocessor, a plurality of microprocessors, one or more microprocessors in conjunction with a DSP core, or any other such configurations.

[0088] Memory 1064 can include cache memory (e.g., of processor 1060), random access memory (RAM), magnetoresistive RAM (MRAM), read-only memory (ROM), programmable read-only memory (PROM), erasable programmable read-only memory (EPROM), electrically erasable programmable read-only memory (EEPROM), flash memory, solid-state memory device, hard disk drive, other forms of volatile and non-volatile memory, or a combination of different types of memory. In an embodiment, memory 1064 includes a non-transitory computer-readable medium. Memory 1064 can store instructions 1066. Instructions 1066 can include instructions that, when executed by processor 1060, enable processor 1060 to perform operations described herein and with reference to Figure 2 image systems 210 and 220, host 230, and / or Figure 3 multimodal image co-registration controller 330 of Figures 2-9 and Figure 11 various aspects of the disclosure). Instructions 1066 can also be referred to as code. The terms “instructions” and “code” should be interpreted broadly to include any type of computer-readable statement(s). For example, the terms “instructions” and “code” can refer to one or more programs, routines, sub-routines, functions, procedures, etc. “Instructions” and “code” can include a single computer-readable statement or many computer-readable statements.

[0089] Communication module 1068 can include any electronic and / or logical circuitry to facilitate direct or indirect data communication between processor circuitry 1000, Figure 2 image systems 210 and 220, Figure 2 host 230, and / or Figure 3 multimodal image co-registration controller 330 of In this regard, communication module 1068 can be an input / output (I / O) device. In some instances, communication module 1068 facilitates direct or indirect communication between various elements of processor circuitry 1000 and / or Figure 2 image systems 210 and 220, host 230, and / or Figure 3 multimodal image co-registration controller 330 of

[0090] Figure 11is a flowchart of a medical imaging method 1100 with multi-modal image co- registration according to aspects of the present disclosure. The method 1100 is implemented by the system 200, e.g., by the processor circuitry (e.g., the processor circuitry 1000) and / or other suitable components (e.g., the host 230, the processor circuitry 234, and / or the multi-modal image co-registration controller 330). In some examples, the system 200 can include a computer- readable medium having program code recorded thereon, the program code including code to cause the system 200 to perform the various steps of the method 1100. The method 1100 can employ similar mechanisms as those described with respect to the systems 100 and / or 200, respectively, the schemes 300, 600, 700, and / or 800 described with respect to Figure 1 and Figure 2 the configurations 500 described with respect to Figure 2 , Figure 6 , Figure 7 and / or Figure 8 the user interfaces 900 described with respect to Figure 5 and / or the user interfaces 900 described with respect to Figure 9 respectively. As shown, the method 1100 includes a number of enumerated steps, but embodiments of the method 1100 can include additional steps performed before, after, and in between the enumerated steps. In some embodiments, one or more of the enumerated steps can be omitted, and one or more of the enumerated steps can be performed in a different order.

[0091] At step 1110, the method 1100 includes receiving, at a processor circuit (e.g., the processor circuitry 1000 and 234) in communication with a first imaging system (e.g., the imaging system 210) of a first imaging modality (e.g., the imaging modality 306), a first image of an anatomical structure of a patient in the first imaging modality.

[0092] At step 1120, the method 1100 includes receiving, at a processor circuit in communication with a second imaging system (e.g., the imaging system 220) of a second imaging modality (e.g., the imaging modality 308), a second image of the anatomical structure of the patient in the second imaging modality, the second imaging modality being different from the first imaging modality.

[0093] At step 1130, the method 1100 includes determining, at the processor circuit, a first pose (e.g., the pose 310) of the first image relative to a reference coordinate system of the anatomical structure of the patient.

[0094] At step 1140, the method 1100 includes determining, at the processor circuit, a second pose (e.g., the pose 320) of the second image relative to the reference coordinate system.

[0095] At step 1150, the method 1100 includes determining, at the processor circuit, co- registration data between the first image and the second image based on the first pose and the second pose.

[0096] At step 1160, the method 1100 includes outputting the first image co-registered with the second image based on the co-registration data to a display (e.g., the display 132 and / or 232) in communication with the processor circuit.

[0097] In some aspects, the anatomical structure of the patient includes an organ, and the reference coordinate system is associated with a centroid of the organ. The reference coordinate system can also be associated with a centroid, a vessel bifurcation, a tip or a boundary of the organ, a portion of a ligament, and / or any other aspect that can be reproducibly identified across a large population of medical images.

[0098] In some aspects, step 1130 includes applying a first prediction network (e.g., the deep learning network 240) to the first image, the first prediction network trained based on a set of images of the first imaging modality and corresponding poses relative to a reference coordinate system in an imaging space of the first imaging modality. Step 1140 includes applying a second prediction network (e.g., the deep learning network 250) to the second image, the second prediction network trained based on a set of images of the second imaging modality and corresponding poses relative to a reference coordinate system in an imaging space of the second imaging modality.

[0099] In some aspects, the first pose includes a first transformation including at least one of a translation or a rotation, and the second pose includes a second transformation including at least one of a translation or a rotation. Step 1150 includes determining a co-registration transformation based on the first transformation and the second transformation. Step 1150 further includes applying the co-registration transformation to the first image to transform the first image into the coordinate system in the imaging space of the second imaging modality. In some aspects, step 1150 further includes determining the co-registration data further based on the co-registration transformation and a secondary multi-modality co-registration (e.g., the multi-modality image registration refinement 820) between the first image and the second image, wherein the secondary multi-modality co-registration is based on at least one of an image feature similarity measure or an image pose prediction.

[0100] In some aspects, the first image is a 2D image slice or a first 3D image volume, and wherein the second image is a second 3D image volume. In some aspects, the method 1100 includes determining a first 2D image slice from the first 3D image volume and a second 2D image slice from the second 3D image volume. Step 1130 includes determining a first pose of the first 2D image slice relative to the reference coordinate system. Step 1140 includes determining a second pose of the second 2D image slice relative to the reference coordinate system.

[0101] In some aspects, the method 1100 includes displaying, at a display, a first image with a first indicator (e.g., indicator 912) and a second image with a second indicator, the first and second indicators (e.g., indicator 922) indicating a same portion of the patient’s anatomy based on the co-registration data.

[0102] Various aspects of the present disclosure can provide several benefits. For example, image pose-based multi-modal image registration can be less challenging and less error-prone compared to feature-based multi-modal image registration that relies on feature recognition and similarity metrics. Using a deep learning-based framework for image pose regression in a local reference coordinate system at the anatomy of interest can provide accurate co-registration results without relying on the particular imaging modality being used. Using deep learning can also provide a less costly and less time-consuming systematic solution compared to feature-based image registration. Additionally, using image pose-based multi-modal image registration to co-register 2D ultrasound images with 3D imaging volumes of 3D imaging modalities (e.g., MR or CT) in real-time can automatically provide spatial location information of the ultrasound probe being used without using an external tracking system. Using image pose-based multi-modal image registration in real-time in conjunction with 2D ultrasound imaging can also automatically identify anatomical information associated with 2D ultrasound image frames from 3D imaging volumes. The disclosed embodiments can provide clinical benefits such as, for example, improved diagnostic confidence, better guidance of interventional procedures, and / or better ability to document findings. In this regard, the ability to compare annotations from pre-operative MRI with results and findings from intra-operative ultrasound can enhance the final report and / or increase confidence in the final diagnosis.

[0103] Those skilled in the art will recognize that the above-described apparatus, systems, and methods can be modified in various ways. Accordingly, those of ordinary skill in the art will recognize that the embodiments encompassed by the present disclosure are not limited to the particular exemplary embodiments described above. In this regard, while the exemplary embodiments have been presented in the foregoing disclosure, a wide variety of modifications, changes and substitutes are contemplated by the foregoing disclosure. It should be understood that various modifications, adaptations, and variations can be made to the foregoing without departing from the scope of the present disclosure. Accordingly, the claims and the specification governing the scope of the present disclosure are to be interpreted in the broadest sense allowable by law.

Claims

1. A system for medical imaging, comprising: A processor circuit that communicates with a first imaging system of a first imaging mode and a second imaging system of a second imaging mode different from the first imaging mode, wherein the processor circuit is configured to: Receive a first image of the patient's anatomical structure in the first imaging modality from the first imaging system; Receive a second image of the patient's anatomical structure in the second imaging modality from the second imaging system; Determine the first pose of the first image relative to a reference coordinate system of the patient's anatomical structure; Determine the second pose of the second image relative to the reference coordinate system; Based on the first pose and the second pose, co-registration data between the first image and the second image is determined; and The first image, which is co-registered with the second image based on the co-registration data, is output to a display that communicates with the processor circuitry; The processor is configured as follows: (i) The first pose is determined by applying a first prediction network to the first image, the first prediction network being trained based on a set of images of the first imaging modality and the corresponding pose relative to the reference coordinate system in the imaging space of the first imaging modality; and (ii) The second pose is determined by applying a second prediction network to the second image, the second prediction network being trained based on a set of images of the second imaging modality and the corresponding pose relative to the reference coordinate system in the imaging space of the second imaging modality.

2. The system according to claim 1, wherein, The patient's anatomical structures include organs, and the reference coordinate system is associated with the centroid of the organ.

3. The system according to claim 1, wherein, The first pose includes a first transformation, the first transformation including at least one of translation or rotation, wherein the second pose includes a second transformation, the second transformation including at least one of translation or rotation, and wherein the processor circuitry configured to determine the co-registration data is configured to: The co-registration transformation is determined based on the first transformation and the second transformation; and The co-registration transformation is applied to the first image to transform the first image into a coordinate system in the imaging space of the second imaging modality.

4. The system according to claim 3, wherein, The processor circuit configured to determine the co-registration data is configured as follows: The co-registration data is also determined based on the co-registration transform and the secondary multimodal co-registration between the first image and the second image, wherein the secondary multimodal co-registration is based on at least one of image feature similarity measurement or image pose prediction.

5. The system according to claim 1, wherein, The first imaging mode is ultrasound.

6. The system according to claim 1, wherein, The first imaging modality is one of the following: ultrasound, magnetic resonance imaging, computed tomography, X-ray, positron emission tomography, single-photon emission tomography-CT, or cone-beam CT, and wherein the second imaging modality is a different one of the following: ultrasound, magnetic resonance imaging, computed tomography, X-ray, positron emission tomography, single-photon emission tomography-CT, or cone-beam CT.

7. The system according to claim 1, further comprising the first imaging system and the second imaging system.

8. The system according to claim 1, wherein, The first image is a two-dimensional (2D) image slice, and the second image is a three-dimensional (3D) image volume.

9. The system according to claim 1, wherein, The first image is a first three-dimensional image volume, and the second image is a second three-dimensional image volume.

10. The system according to claim 9, wherein: The processor circuit is configured as follows: The first two-dimensional image slice is determined based on the volume of the first three-dimensional image. The second two-dimensional image slice is determined based on the volume of the second three-dimensional image; The processor circuit configured to determine the first attitude is configured as follows: Determine the first pose of the first two-dimensional image slice relative to the reference coordinate system; and The processor circuit configured to determine the second attitude is configured as follows: Determine the second orientation of the second two-dimensional image slice relative to the reference coordinate system.

11. The system according to claim 1, further comprising: The display is configured to show a first image with a first indicator and a second image with a second indicator, the first and second indicators indicating the same parts of the patient's anatomy based on the co-registration data.

12. A medical imaging method, comprising: A first image of the patient's anatomical structure in the first imaging modality is received at a processor circuit that communicates with a first imaging system in the first imaging modality; The processor circuit receives a second image of the patient's anatomical structure in the second imaging modality, which is different from the first imaging modality, at the processor circuit in communication with the second imaging system of the second imaging modality; At the processor circuit, a first pose of the first image relative to a reference coordinate system of the patient's anatomical structure is determined. The second pose of the second image relative to the reference coordinate system is determined at the processor circuit. At the processor circuit, co-registration data between the first image and the second image is determined based on the first pose and the second pose; and The first image, which is co-registered with the second image based on the co-registration data, is output to a display that communicates with the processor circuitry; Determining the first pose includes: applying a first prediction network to the first image, the first prediction network being trained based on a set of images of the first imaging modality and the corresponding pose relative to the reference coordinate system in the imaging space of the first imaging modality; and Determining the second pose includes applying a second prediction network to the second image, the second prediction network being trained based on a set of images of the second imaging modality and the corresponding pose relative to the reference coordinate system in the imaging space of the second imaging modality.

13. The method according to claim 12, wherein, The patient's anatomical structures include organs, and the reference coordinate system is associated with the centroid of the organ.

14. The method according to claim 12, wherein, The first pose includes a first transformation, which includes at least one of translation or rotation; the second pose includes a second transformation, which includes at least one of translation or rotation; and determining the co-registration data includes: The co-registration transformation is determined based on the first transformation and the second transformation; and The co-registration transformation is applied to the first image to transform the first image into a coordinate system in the imaging space of the second imaging modality.

15. The method according to claim 14, wherein, Determining the co-registration data includes: The co-registration data is also determined based on the co-registration transform and the secondary multimodal co-registration between the first image and the second image, wherein the secondary multimodal co-registration is based on at least one of image feature similarity measurement or image pose prediction.

16. The method according to claim 12, wherein, The first image is a two-dimensional (2D) image slice or a first three-dimensional image volume, and the second image is a second three-dimensional image volume.

17. The method of claim 16, further comprising: The first two-dimensional image slice is determined based on the volume of the first three-dimensional image. The second two-dimensional image slice is determined based on the volume of the second three-dimensional image. Determining the first posture includes: Determine the first pose of the first two-dimensional image slice relative to the reference coordinate system, and Determining the second posture includes: Determine the second orientation of the second two-dimensional image slice relative to the reference coordinate system.

18. The method of claim 12, further comprising: A first image with a first indicator and a second image with a second indicator are displayed on the monitor, the first indicator and the second indicator indicating the same part of the patient's anatomical structure based on the co-registration data.

Citation Information

Patent Citations

  • Methods and systems for model driven multi-modal medical imaging

    CN108720807A