Segmentation and view guidance in ultrasound imaging and associated devices, systems and methods
Through time-aware deep learning network segmentation and imaging guidance technology, the problem of identifying thin, non-rigid or mobile objects in traditional ultrasound imaging is solved, stable segmentation and real-time imaging guidance are achieved, and the accuracy and efficiency of data interpretation are improved.
Patent Information
- Application Number
- CN202080026799.5
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Priority Date
- 2020-01-23
- Filing Date
- 2020-03-30
- Publication Date
- 2025-08-12
- Estimated Expiration
- 2040-03-30
AI Technical Summary
Traditional ultrasound imaging techniques are difficult to effectively identify and segment thin, non-rigid or mobile medical devices and anatomical structures. Especially in 3D and 4D imaging, data interpretation is complex and user-dependent, and traditional algorithms perform poorly when identifying mobile objects.
The time-aware deep learning network is adopted, combining recursive prediction networks and convolutional encoding-decoding layer, using the time continuity information in 3D and 4D ultrasound data, segmenting moving objects and providing imaging guidance, and identifying and predicting the position and movement of medical devices and anatomical structures by training deep learning networks.
The stable segmentation of mobile objects and real-time imaging guidance are achieved, reducing the time for clinicians to find ideal imaging views, and improving the accuracy and efficiency of data interpretation.
Smart Images

Figure CN113678167B_ABST
Abstract
Description
[0001] CROSS-REFERENCE TO RELATED APPLICATIONS
[0002] This application claims priority to and the benefit of U.S. Provisional Patent Application No. 62 / 828,185, filed April 2, 2019, and U.S. Provisional Patent Application No. 62 / 964,715, filed January 23, 2020, which are hereby incorporated by reference in their entireties as if fully set forth below and for all applicable purposes. Technical Field
[0003] The present disclosure relates generally to ultrasound imaging, and more particularly to providing segmentation of moving objects and guidance for locating optimal imaging views. Background Art
[0004] Ultrasound can provide non-radiative, safe and real-time dynamic imaging of anatomical structures and / or medical devices during medical procedures (e.g., diagnosis, intervention and / or treatment). Traditionally, clinicians have relied on two-dimensional (2D) ultrasound imaging to provide guidance in diagnosis and / or navigation of medical devices through the patient's body during medical procedures. However, in some cases, medical devices and / or anatomical structures may be thin, non-rigid and / or mobile, making them difficult to identify in 2D ultrasound images. Similarly, anatomical structures may be thin, tortuous and, in some cases, may be in constant motion (e.g., due to breathing, heart and / or arterial pulses).
[0005] The recent development and availability of three-dimensional (3D) ultrasound makes it possible to observe 3D volumes rather than 2D image slices. The ability to visualize 3D volumes can be valuable in medical procedures. For example, due to foreshortening, the tip of a medical device may be indeterminate in a 2D image slice, but may be clear when viewed in a 3D volume. Operations such as locating the optimal imaging plane in a 3D volume can significantly benefit from four-dimensional (4D) imaging (e.g., 3D imaging over time). Examples of clinical areas that can benefit from 3D and / or 4D imaging can include the diagnosis and / or treatment of peripheral vascular disease (PVD) and structural heart disease (SHD).
[0006] While 3D and / or 4D imaging can provide valuable visualization and / or guidance for medical procedures, the interpretation of 3D and / or 4D imaging data can be complex and challenging due to the high volume, high dimensionality, low resolution, and / or low frame rate of the data. For example, accurate interpretation of 3D and / or 4D imaging data may require a user or clinician with extensive training and expertise. Additionally, the interpretation of the data may be user-dependent. Typically, during ultrasound-guided procedures, clinicians may spend a significant portion of their time searching for an ideal imaging view of the patient's anatomy and / or medical device.
[0007] Computers are generally better at interpreting high-volume, high-dimensional data. For example, algorithmic models can be applied to assist in interpreting 3D and / or 4D imaging data and / or locating optimal imaging views. However, conventional algorithms may perform poorly in identifying and / or segmenting thin and / or moving objects in ultrasound images due to, for example, low signal-to-noise ratio (SNR), ultrasound artifacts, occlusion by devices positioned in confusing poses (such as along a blood vessel wall), and / or high-intensity artifacts that may resemble moving objects. Summary of the Invention
[0008] There remains a clinical need for improved systems and techniques for image segmentation and imaging guidance. Embodiments of the present disclosure provide a deep learning network that utilizes temporal continuity information in three-dimensional (3D) ultrasound data and / or four-dimensional (4D) ultrasound data to segment moving objects and / or provide imaging guidance. 3D ultrasound data may refer to a time series of 2D images obtained from 2D ultrasound imaging across time. 4D ultrasound data may refer to a time series of 3D volumes obtained from 3D ultrasound imaging across time. A time-aware deep learning network includes a recursive component (e.g., a recursive neural network (RNN)) coupled to multiple convolutional encoding-decoding layers operating at multiple different spatial resolutions. The deep learning network is applied to a time series of 2D or 3D ultrasound imaging frames that include moving objects and / or medical devices. The recursive component passes the deep learning network's prediction of the current image frame as an auxiliary input to the prediction of the next image frame.
[0009] In an embodiment, a deep learning network is trained to distinguish flexible, slender, thin medical devices (e.g., catheters, guidewires, needles, therapeutic devices, and / or treatment devices) passing through anatomical structures (e.g., the heart, lungs, and / or blood vessels) from the anatomical structures and to predict the position and / or motion of the medical devices based on temporal continuity information in ultrasound image frames. In an embodiment, the deep learning network is trained to distinguish moving portions of the anatomical structure caused by cardiac motion, respiratory motion, and / or arterial pulses from static portions of the anatomical structure and to predict the motion of the moving portions based on temporal continuity information in ultrasound image frames. In an embodiment, the deep learning network is trained to predict a target imaging plane for the anatomical structure. The predictions of the deep learning network can be used to generate control signals and / or instructions (e.g., rotation and / or translation) to automatically steer the ultrasound beam to image the target imaging plane. Alternatively, the predictions of the deep learning network can be used to provide instructions to the user for navigating the ultrasound imaging device toward the target imaging plane. The deep learning network can be applied in real time during 3D and / or 4D imaging to provide dynamic segmentation and imaging guidance.
[0010] In one embodiment, an ultrasound imaging system includes a processor circuit that communicates with an ultrasound imaging device, the processor circuit being configured to: receive a sequence of input image frames of a moving object within a time period from the ultrasound imaging device, wherein the moving object includes at least one of a patient's anatomy or a medical device passing through the patient's anatomy, and wherein a portion of the moving object is at least partially invisible in a first input image frame in the sequence of input image frames; apply a recursive prediction network associated with image segmentation to the sequence of input image frames to generate segmentation data; and output a sequence of output image frames to a display that communicates with the processor circuit based on the segmentation data, wherein the portion of the moving object is fully visible in a first output image frame in the sequence of output image frames, and the first output image frame and the first input image frame are associated with the same moment in the time period.
[0011] In some embodiments, the processor circuit configured to apply the recursive prediction network is further configured to: generate previous segmentation data based on a previous input image frame in the sequence of input image frames, the previous input image frame being received before the first input image frame; and generate first segmentation data based on the first input image frame and the previous segmentation data. In some embodiments, the processor circuit configured to generate the previous segmentation data is configured to apply a convolutional encoder and a recursive neural network to the previous input image frame; the processor circuit configured to generate the first segmentation data is configured to: apply the convolutional encoder to the first input image frame to generate encoded data; and apply the recursive neural network to the encoded data and the previous segmentation data; and the processor circuit configured to apply the recursive prediction network is further configured to apply a convolutional decoder to the first segmentation data and the previous segmentation data. In some embodiments, the convolutional encoder, the recursive neural network, and the convolutional decoder operate at multiple spatial resolutions. In some embodiments, the moving object comprises the medical device passing through the patient's anatomy, and the convolutional encoder, the recurrent neural network, and the convolutional decoder are trained to identify the medical device from the patient's anatomy and predict the motion associated with the medical device passing through the patient's anatomy. In some embodiments, the moving object comprises the patient's anatomy having at least one of cardiac motion, respiratory motion, or an arterial pulse, and the convolutional encoder, the recurrent neural network, and the convolutional decoder are trained to identify moving portions of the patient's anatomy from static portions of the patient's anatomy and predict the motion associated with the moving portions. In some embodiments, the moving object comprises the medical device passing through the patient's anatomy, and the system comprises the medical device. In some embodiments, the medical device comprises at least one of a needle, a guidewire, a catheter, a guide catheter, a therapeutic device, or an interventional device. In some embodiments, the input image frames comprise at least one of two-dimensional image frames or three-dimensional image frames. In some embodiments, the processor circuit is further configured to apply a spline fit to the sequence of input image frames based on the segmentation data. In some embodiments, the system further comprises the ultrasound imaging device, and wherein the ultrasound imaging device comprises an ultrasound transducer array configured to obtain the sequence of input image frames.
[0012] In one embodiment, an ultrasound imaging system includes a processor circuit that communicates with an ultrasound imaging device, the processor circuit being configured to receive a sequence of image frames representing an anatomical structure of a patient over a time period from the ultrasound imaging device; apply a recursive prediction network associated with image acquisition to the sequence of image frames to generate imaging plane data associated with a clinical property of the patient's anatomical structure; and output at least one of a target imaging plane of the patient's anatomical structure or instructions for repositioning the ultrasound imaging device toward the target imaging plane based on the imaging plane data to a display in communication with the processor circuit.
[0013] In some embodiments, the processor circuit configured to apply the recursive prediction network is further configured to generate first imaging plane data based on a first image frame in the sequence of image frames; and generate second imaging plane data based on a second image frame in the sequence of image frames and the first imaging plane data, the second image frame being received after the first image frame. In some embodiments, the processor circuit configured to generate the first imaging plane data is configured to apply a convolutional encoder and a recursive neural network to the first image frame; the processor circuit configured to generate the second imaging plane data is configured to apply the convolutional encoder to the first image frame to generate encoded data; and apply the recursive neural network to the encoded data and the first imaging plane data; and the processor circuit configured to apply the recursive prediction network is further configured to apply a convolutional decoder to the first imaging plane data and the second imaging plane data. In some embodiments, the convolutional encoder, the recursive neural network, and the convolutional decoder operate at multiple spatial resolutions, and the convolutional encoder, the recursive neural network, and the convolutional decoder are trained to predict the target imaging plane for imaging the clinical attribute of the patient's anatomy. In some embodiments, the image frames include at least one of a two-dimensional image frame or a three-dimensional image frame of the patient's anatomical structure. In some embodiments, the processor circuit is configured to output the target imaging plane, the target imaging plane including at least one of a cross-sectional image slice, an orthogonal image slice, or a multi-planar reconstruction (MPR) image slice of the patient's anatomical structure, the patient's anatomical structure including the clinical attribute. In some embodiments, the system further includes the ultrasound imaging device, and wherein the ultrasound imaging device includes an ultrasound transducer array, the ultrasound transducer array configured to obtain the sequence of image frames. In some embodiments, the processor circuit is further configured to generate an ultrasound beam steering control signal based on the imaging plane data; and output the ultrasound beam steering control signal to the ultrasound imaging device. In some embodiments, the processor circuit is configured to output the instructions including at least one of a rotation or a translation of the ultrasound imaging device.
[0014] Other aspects, features, and advantages of the present disclosure will become apparent from the following detailed description. BRIEF DESCRIPTION OF THE DRAWINGS
[0015] Illustrative embodiments of the present disclosure will be described with reference to the accompanying drawings, in which:
[0016] Figure 1 is a schematic diagram of an ultrasound imaging system according to some aspects of the present disclosure.
[0017] Figure 2 is a schematic diagram of a deep learning-based image segmentation scheme according to aspects of the present invention.
[0018] Figure 3 is a schematic diagram illustrating a configuration for a time-aware deep learning network according to aspects of the present invention.
[0019] Figure 4 is a schematic diagram illustrating a configuration for a time-aware deep learning network according to aspects of the present invention.
[0020] Figure 5 A scenario of an ultrasound-guided procedure according to aspects of the present disclosure is illustrated.
[0021] Figure 6 A scenario of an ultrasound-guided procedure according to aspects of the present disclosure is illustrated.
[0022] Figure 7 A scenario of an ultrasound-guided procedure according to aspects of the present disclosure is illustrated.
[0023] Figure 8 A scenario of an ultrasound-guided procedure according to aspects of the present disclosure is illustrated.
[0024] Figure 9 is a schematic diagram of a deep learning-based image segmentation scheme with spline fitting according to aspects of the present invention.
[0025] Figure 10 is a schematic diagram of a deep learning-based imaging guidance scheme according to aspects of the present disclosure.
[0026] Figure 11 An ultrasound image obtained from an ultrasound-guided procedure according to aspects of the present disclosure is illustrated.
[0027] Figure 12 is a schematic diagram of a processor circuit according to an embodiment of the present disclosure.
[0028] Figure 13 is a flowchart of a deep learning-based ultrasound imaging method according to aspects of the present disclosure.
[0029] Figure 14 is a flowchart of a deep learning-based ultrasound imaging method according to aspects of the present disclosure. DETAILED DESCRIPTION
[0030] In order to promote the understanding of the principles of the present disclosure, reference will now be made to the embodiments illustrated in the accompanying drawings, and specific language will be used to describe these embodiments. However, it should be understood that it is not intended to limit the scope of the present disclosure. As would be generally expected by those skilled in the art to which the present disclosure relates, any changes and further modifications to the described apparatus, systems, and methods, as well as any further applications of the principles of the present disclosure, are fully anticipated and included in the present disclosure. In particular, it is fully anticipated that the features, components, and / or steps described with respect to one embodiment may be combined with the features, components, and / or steps described with respect to other embodiments of the present disclosure. However, for the sake of brevity, the numerous iterative forms of these combinations will not be described separately.
[0031] Figure 1 FIG1 is a schematic diagram of an ultrasound imaging system 100 according to aspects of the present disclosure. System 100 is used to scan a region or volume of a patient's body. System 100 includes an ultrasound imaging probe 110 that communicates with a host computer 130 via a communication interface or link 120. Probe 110 includes a transducer array 112, a beamformer 114, a processing component 116, and a communication interface 118. Host computer 130 includes a display 132, a processing component 134, and a communication interface 136.
[0032] In an exemplary embodiment, the probe 110 is an external ultrasound imaging device that includes a housing configured to be handheld by a user. The transducer array 112 can be configured to obtain ultrasound data when the user grasps the housing of the probe 110 so that the transducer array 112 is positioned adjacent to and / or in contact with the patient's skin. The probe 110 is configured to obtain ultrasound data of anatomical structures within the patient's body when the probe 110 is positioned external to the patient. In some embodiments, the probe 110 is a transthoracic ultrasound (TTE) probe. In some other embodiments, the probe 110 can be a transesophageal ultrasound (TEE) probe.
[0033] The transducer array 112 transmits ultrasound signals toward the patient's anatomical object 105 and receives echo signals reflected from the object 105 and returned to the transducer array 112. The ultrasound transducer array 112 can include any suitable number of acoustic elements, including one or more acoustic elements and / or a plurality of acoustic elements. In some cases, the transducer array 112 includes a single acoustic element. In some cases, the transducer array 112 can include an array of acoustic elements having any number of acoustic elements in any suitable configuration. For example, the transducer array 112 can include 1 to 10,000 acoustic elements, including numbers such as 2 acoustic elements, 4 acoustic elements, 36 acoustic elements, 64 acoustic elements, 128 acoustic elements, 500 acoustic elements, 812 acoustic elements, 1,000 acoustic elements, 3,000 acoustic elements, 8,000 acoustic elements, and / or other greater and lesser values. In some cases, the transducer array 112 can include an array of acoustic elements having any number of acoustic elements in any appropriate configuration, for example, a linear array, a planar array, a curved array, a curvilinear array, a circular array, an annular array, a phased array, a matrix array, a one-dimensional (1D) array, a 1.x-dimensional array (e.g., a 1.5D array), or a two-dimensional (2D) array. The array of acoustic elements can be controlled and activated uniformly or independently (e.g., one or more rows, one or more columns, and / or one or more orientations). The transducer array 112 can be configured to obtain one-dimensional, two-dimensional, and / or three-dimensional images of the patient's anatomical structure. In some embodiments, the transducer array 112 can include piezoelectric micromachined ultrasonic transducers (PMUTs), capacitive micromachined ultrasonic transducers (CMUTs), single crystals, lead zirconate titanate (PZT), PZT composites, other suitable transducer types, and / or combinations thereof.
[0034] Object 105 can include any anatomical structure suitable for ultrasound imaging, such as a patient's blood vessels, nerve fibers, airways, mitral valve leaflets, kidneys, and / or liver. In some embodiments, object 105 can include at least a portion of a patient's heart, lungs, and / or skin. In some embodiments, object 105 can be in constant motion, such as caused by respiration, cardiac activity, and / or arterial pulses. The motion can be regular or cyclical, such as the movement of the heart, associated blood vessels, and / or lungs in the context of a cardiac cycle or heartbeat cycle. The present disclosure can be implemented in the context of any number of anatomical locations and tissue types, including, but not limited to, organs, including the liver, heart, kidneys, gallbladder, pancreas, lungs; ducts; intestines; nervous system structures, including the brain, thecal sac, spinal cord, and peripheral nerves; the urinary tract; and valves within blood vessels, blood vessels, chambers, or other parts of the heart and / or other systems of the body. The anatomical structure can be a blood vessel, such as an artery or vein of the patient's vascular system, including the cardiac vasculature, peripheral vasculature, neural vasculature, renal vasculature, and / or any other suitable lumen within the body. In addition to natural structures, the present disclosure can be implemented in the context of artificial structures, such as, but not limited to, heart valves, stents, shunts, filters, implants, and other devices.
[0035] In some embodiments, system 100 is used to guide a clinician during a medical procedure (e.g., treatment, diagnosis, therapy, and / or intervention). For example, a clinician can insert a medical device 108 into an anatomical object 105. In some examples, the medical device 108 can include an elongated flexible member having a thin geometry. In some examples, the medical device 108 can be a guidewire, a catheter, a guide catheter, a needle, an intravascular ultrasound (IVUS) device, a diagnostic device, a treatment / therapy device, an interventional device, and / or an intraductal imaging device. In some examples, the medical device 108 can be any imaging device suitable for imaging the patient's anatomical structure and can have any suitable imaging modality, such as optical tomography (OCT) and / or endoscopy. In some examples, the medical device 108 can include a sheath, an imaging device, and / or an implantable device. In some examples, the medical device 108 can be a treatment / therapy device including a balloon, a stent, and / or a rotary cutting device. In some examples, the medical device 108 can have a diameter smaller than the diameter of the blood vessel. In some examples, medical device 108 can have a diameter or thickness of approximately 0.5 millimeters (mm) or less. In some examples, medical device 108 can be a guidewire having a diameter of approximately 0.035 inches. In such embodiments, transducer array 112 can generate ultrasound echoes that are reflected by object 105 and medical device 108.
[0036] The beamformer 114 is coupled to the transducer array 112. For example, the beamformer 114 controls the transducer array 112 for transmitting ultrasound signals and receiving ultrasound echo signals. The beamformer 114 provides an image signal to the processing component 116 based on the responses or received ultrasound echo signals. The beamformer 114 may include multiple stages of beamforming. Beamforming can reduce the number of signal lines used to couple to the processing component 116. In some embodiments, the transducer array 112 combined with the beamformer 114 may be referred to as an ultrasound imaging component.
[0037] The processing component 116 is coupled to the beamformer 114. The processing component 116 may include a central processing unit (CPU), a digital signal processor (DSP), an application-specific integrated circuit (ASIC), a controller, a field-programmable gate array (FPGA) device, another hardware device, a firmware device, or any combination thereof configured to perform the operations described herein. The processing component 134 may also be implemented as a combination of computing devices (e.g., a DSP and a microprocessor, multiple microprocessors, a combination of one or more microprocessors in conjunction with a DSP core, or any other such configuration). The processing component 116 is configured to process the beamformed image signals. For example, the processing component 116 may perform filtering and / or quadrature demodulation to manipulate the image signals. The processing components 116 and / or 134 may be configured to control the array 112 to obtain ultrasound data associated with the subject 105 and / or the medical device 108.
[0038] The communication interface 118 is coupled to the processing component 116. The communication interface 118 may include one or more transmitters, one or more receivers, one or more transceivers, and / or circuitry for transmitting and / or receiving communication signals. The communication interface 118 may include hardware components and / or software components that implement a specific communication protocol suitable for transmitting signals to the host 130 via the communication link 120. The communication interface 118 may be referred to as a communication device or a communication interface module.
[0039] The communication link 120 may be any suitable communication link. For example, the communication link 120 may be a wired link, such as a Universal Serial Bus (USB) link or an Ethernet link. Alternatively, the communication link 120 may be a wireless link, such as an Ultra Wideband (UWB) link, an Institute of Electrical and Electronics Engineers (IEEE) 802.11 WiFi link, or a Bluetooth link.
[0040] At host 130, image signals may be received by communication interface 136. Communication interface 136 may be substantially similar to communication interface 118. Host 130 may be any suitable computing and display device, such as a workstation, personal computer (PC), laptop, tablet, or mobile phone.
[0041] The processing component 134 is coupled to the communication interface 136. The processing component 134 can be implemented as a combination of software components and hardware components. The processing component 134 may include a central processing unit (CPU), a graphics processing unit (GPU), a digital signal processor (DSP), an application-specific integrated circuit (ASIC), a controller, an FPGA device, another hardware device, a firmware device, or any combination thereof configured to perform the operations described herein. The processing component 134 can also be implemented as a combination of computing devices (e.g., a DSP and a microprocessor, a plurality of microprocessors, a combination of one or more microprocessors combined with a DSP core, or any other such configuration). The processing component 134 can be configured to generate image data based on the image signal received from the probe 110. The processing component 134 can apply advanced signal processing and / or image processing techniques to the image signal. In some embodiments, the processing component 134 can form a three-dimensional (3D) volume image based on the image data. In some embodiments, the processing component 134 can perform real-time processing on the image data to provide a streaming video of an ultrasound image of the object 105 and / or the medical device 108.
[0042] The display 132 is coupled to the processing component 134. The display 132 can be a monitor or any suitable display. The display 132 is configured to display ultrasound images, image videos, and / or any imaging information of the subject 105 and / or the medical device 108.
[0043] As described above, system 100 can be used to provide guidance to clinicians during medical procedures. In an example, as medical device 108 passes through object 105, system 100 can capture a sequence of ultrasound images of object 105 and medical device 108. The sequence of ultrasound images can be 2D or 3D. In some examples, system 100 can be configured to perform bi-plane imaging or multi-plane imaging to provide the sequence of ultrasound images as bi-plane images or multi-plane images, respectively. In some cases, due to the motion of medical device 108 and / or the thin geometry of medical device 108, it may be difficult for clinicians to identify and / or distinguish medical device 108 from object 105 based on the captured images. For example, medical device 108 may appear to jump from one frame to another without temporal continuity. To improve the visualization, stability, and / or temporal continuity of device 108 as it moves through object 105, processing component 134 can apply a time-aware deep learning network trained for segmentation to the sequence of images. The deep learning network identifies and / or distinguishes the medical device 108 from the anatomical object 105 and uses the temporal information carried in the sequence of images captured across time to predict the motion and / or position of the medical device 108. The processing component 134 can incorporate the predictions into the captured 2D and / or 3D image frames to provide a time series of output images with a stabilized view of the moving medical device 108 frame by frame.
[0044] In some examples, the sequence of ultrasound images input to the deep learning network can be a 3D volume, and the output prediction can be a 2D image, a bi-plane image, and / or a multi-plane image. In some examples, the medical device 108 can be a 2D ultrasound imaging probe, and the deep learning network can be configured to predict a volumetric 3D segmentation, wherein the sequence of ultrasound images input to the deep learning network can be a 2D image, a bi-plane image, and / or a multi-plane image, and the output prediction can be a 3D volume.
[0045] In some examples, it may be difficult to identify anatomical structures (e.g., object 105) under 2D and / or 3D imaging due to the geometry and / or motion of the anatomical structures. For example, tortuous vessels in distal peripheral anatomical structures and / or small structures close to the heart may be affected by arterial and / or cardiac motion. Depending on the cardiac phase, the mitral valve leaflets and / or other structures may move in and out of the ultrasound imaging view over a period of time. In another example, due to the patient's respiratory motion, blood vessels, airways, and tumors may move in and out of the ultrasound imaging view during endobronchial ultrasound imaging. Similarly, to improve the visualization, stability, and / or temporal continuity of the motion of anatomical structures, the processing component 134 may apply a time-aware deep learning network trained for segmentation to a series of 2D and / or 3D images of the object 105 captured across time. The deep learning network identifies and / or distinguishes moving parts (e.g., foreground) from relatively more static parts (e.g., background) of the object 105, and uses the temporal information carried in the sequence of images captured across time to predict the motion and / or position of the moving parts. For example, in cardiac imaging, the moving portion may correspond to the mitral valve leaflets, and the stationary portion may correspond to the cardiac chambers, which may include relatively slower motion than the valve. In peripheral vascular imaging, the moving portion may correspond to the pulsating arteries, and the stationary portion may correspond to the surrounding tissue. In lung imaging, the moving portion may correspond to the lung chambers and airways, and the stationary portion may correspond to the surrounding cavities and tissue. The processing component 134 may incorporate the predictions into the captured image frames to provide a series of output images with a stable view of the moving anatomical structure frame by frame. The present invention describes in more detail the mechanisms for providing a stable view of a moving object (e.g., a medical device 108 and / or a subject 105) using a time-aware deep learning model.
[0046] In one embodiment, the system 100 can be used to assist a clinician in finding the optimal imaging view of a patient for a particular clinical attribute or clinical examination. For example, the processing component 134 can utilize a time-aware deep learning network trained for image acquisition to predict the optimal imaging view or image slice of the object 105 from the captured 2D and / or 3D images for a particular clinical attribute. For example, the system 100 can be configured for cardiac imaging to assist a clinician in measuring ventricular volumes, determining the presence of arrhythmias, performing a transseptal puncture, and / or providing visualization of the mitral valve for repair and / or replacement. The cardiac imaging can be configured to provide a four-chamber view, a three-chamber view, and / or a two-chamber view. In an example, cardiac imaging can be used to visualize the left ventricular outflow tract (LVOT), which can be critical for mitral valve clips and valves in mitral valve replacement. In one example, cardiac imaging can be used to visualize the mitral annulus for any procedure involving annuloplasty. In an example, cardiac imaging can be used to visualize the left atrial appendage during a transseptal puncture (TSP) to prevent aortization. For endobronchial ultrasound imaging, the clinical attribute can be the presence and location of a suspected tumor and can be obtained from a transverse or sagittal ultrasound view in which the ultrasound transducer is aligned with the tumor and the adjacent airway. In some examples, the processing component 134 can provide instructions (e.g., rotation and / or translation) to the clinician to manipulate the probe 110 from one position to another or from one imaging plane to another to obtain the best imaging view of the clinical attribute based on the prediction output by the deep learning network. In some examples, the processing component 134 can automate the process of reaching the best imaging view. For example, the processing component 134 is configured to automatically steer the 2D or X-plane beam generated by the transducer array 112 to the best imaging position based on the prediction results output by the deep learning network. The X-plane can include a cross-sectional plane and a longitudinal plane. The mechanism for reaching the best imaging view using a deep learning model is described in more detail herein.
[0047] In some embodiments, the system 100 can be configured for various stages of a biopsy prediction and guidance process based on ultrasound imaging. In one embodiment, the system 100 can be used to collect ultrasound images to form a training dataset for deep learning network training. For example, the host 130 may include a memory 138, which can be any suitable storage device, such as a cache memory (e.g., a cache memory of the processing component 134), a random access memory (RAM), a magnetoresistive RAM (MRAM), a read-only memory (ROM), a programmable read-only memory (PROM), an erasable programmable read-only memory (EPROM), an electrically erasable programmable read-only memory (EEPROM), a flash memory, a solid-state storage device, a hard drive, a solid-state drive, other forms of volatile and non-volatile memory, or a combination of different types of memory. The memory 138 can be configured to store an image dataset 140 to train a time-aware deep learning network for image segmentation and / or imaging view guidance. The mechanism for training a time-aware deep learning network is described in more detail herein.
[0048] Figure 2-4 Together, we illustrate the mechanism of image segmentation using a time-aware multi-layer deep learning network. Figure 2 2 is a schematic diagram of a deep learning-based image segmentation scheme 200 according to aspects of the present disclosure. Scheme 200 is implemented by system 100. Scheme 200 utilizes a time-aware multi-layer deep learning network 210 to provide segmentation of moving objects in ultrasound images. In some examples, the moving object can be a medical device (e.g., a guidewire, catheter, guide catheter, needle, or therapeutic device similar to device 108 and / or 212) moving within a patient's anatomical structure (e.g., the heart, lungs, blood vessels, and / or skin similar to subject 105). In some examples, the moving object can be an anatomical structure (e.g., subject 105) with cardiac motion, respiratory motion, and / or arterial pulse. At a high level, multi-layer deep learning network 210 receives a sequence of ultrasound image frames 202 of devices and / or anatomical structures. Each image frame 202 passes through time-aware multi-layer deep learning network 210. The deep learning network 210's prediction for the current image frame 202 is passed as input for predicting the next image frame 202. In other words, the deep learning network 210 includes a recursive component that utilizes temporal continuity in the sequence of ultrasound image frames 202 for prediction. Therefore, the deep learning network 210 is also referred to as a recursive prediction network.
[0049] The sequence of image frames 202 is captured across a time period (e.g., from time T0 to time Tn). The image frames 202 can be captured using the system 100. For example, the sequence of image frames 202 is reconstructed from ultrasound echoes collected by the transducer array 112, beamformed by the beamformer 114, filtered and / or conditioned by the processing components 116 and / or 134, and reconstructed by the processing component 134. The sequence of image frames 202 is input into the deep learning network 210. Although Figure 2 The image frames 202 are illustrated as 3D volumes, but the scheme 200 can be similarly applied to a sequence of 2D input image frames captured across time to provide segmentation. In some examples, the sequence of 3D image frames 202 across time can be referred to as a continuous 4D (e.g., 3D volume and time) ultrasound sequence.
[0050] Deep learning network 210 includes convolutional encoder 220, time-aware RNN 230, and convolutional decoder 240. Convolutional encoder 220 includes multiple convolutional encoding layers 222. Convolutional decoder 240 includes multiple convolutional decoding layers 242. In some examples, the number of convolutional encoding layers 222 and the number of convolutional decoding layers 242 can be the same. In some examples, the number of convolutional encoding layers 222 and the number of convolutional decoding layers 242 can be different. For simplicity of illustration and discussion, Figure 2 The four convolutional coding layers 222 in the convolutional encoder 220 are shown. K0 , 222 K1 , 222 K2 and 222 K3 and four convolutional decoding layers 242 in the convolutional decoder 240 L0 , 224 L1 , 242 L2 and 242 L3 , however, it should be appreciated that embodiments of the present disclosure can be scaled to include any suitable number of convolutional encoding layers 222 (e.g., approximately 2, 3, 5, 6, or more) and any suitable number of convolutional decoding layers 242 (e.g., approximately 2, 3, 5, 6, or more). The subscripts K0, K1, K2, and K3 represent the layer indices of the convolutional encoding layers 222. The subscripts L0, L1, L2, and L3 represent the layer indices of the convolutional decoding layers 242.
[0051] Each convolutional encoding layer 222 and each convolutional decoding layer 242 can include a convolution filter or kernel. The convolution kernel can be a 2D kernel or a 3D kernel, depending on whether the deep learning network 210 is configured to operate on a 2D image or a 3D volume. For example, when the image frame 202 is a 2D image, the convolution kernel is a 2D filter kernel. Alternatively, when the image frame 202 is a 3D volume, the convolution kernel is a 3D filter kernel. The filter coefficients of the convolution kernel are trained to learn to segment moving objects, as described in more detail herein.
[0052] In some embodiments, the convolutional encoding layer 222 and the convolutional decoding layer 242 can operate at multiple different spatial resolutions. In such embodiments, each convolutional encoding layer 222 can be followed by a downsampling layer. Each convolutional decoding layer 242 can be preceded by an upsampling layer. The downsampling and upsampling can be by any suitable factor. In some examples, the downsampling factor at each downsampling layer and the upsampling factor at each upsampling layer can be approximately 2. The convolutional encoding layer 222 and the convolutional decoding layer 242 can be trained to extract features from a sequence of image frames 202 at different spatial resolutions.
[0053] The RNN 230 is positioned between the convolutional encoding layer 222 and the convolutional decoding layer 242. The RNN 230 is configured to capture temporal information (e.g., temporal continuity) from the sequence of input image frames 202 for segmenting moving objects. The RNN 230 may include multiple time-aware recurrent components (e.g., Figure 3 and 4 For example, the RNN 230 passes the prediction for the current image frame 202 (captured at time T0) back to the RNN 230 as an auxiliary input for the prediction for the next image frame 202 (captured at time T1), as shown by arrow 204. Figure 3 and 4 The use of temporal information at different spatial resolutions for segmenting moving objects is described in more detail.
[0054] Figure 3 is a schematic diagram illustrating a configuration 300 for a time-aware deep learning network 210 according to aspects of the present invention. Figure 3 A more detailed view of the use of time information at the deep learning network 210 is provided. For simplicity of illustration and discussion, Figure 3 The operation of the network 210 at two times T0 and T1 is illustrated. However, similar operations can be propagated to subsequent times T2, T3, ..., and Tn. In addition, for simplicity, the convolutional encoding layer 222 is shown without the layer index subscripts K0, k1, K2, and K3, and the convolutional decoding layer 242 is shown without the layer index subscripts L0, l1, L2, and L3. Figure 3 Subscripts T0 and T1 are used to denote time indexes.
[0055] At time T0, the system 100 captures image frame 202 T0 Image frame 202 T0 is input into the deep learning network 210. The image frame 202 is processed by each convolutional coding layer 222. The convolutional coding layer 222 generates the coded features 304 T0 , encoding feature 304T0 Features at different spatial resolutions may be included, as described in more detail below.
[0056] RNN 230 may include multiple recurrent components 232, each recurrent component operating at one of the spatial resolutions. In some examples, recurrent component 232 may be a long short-term memory (LSTM) unit. In some examples, recurrent component 232 may be a gated recurrent unit (GRU). Each recurrent component 232 is applied to the encoded features 304 of the corresponding spatial resolution. T0 to produce output 306 T0 Output 306 T0 is stored in a memory (eg, memory 138). In some examples, recursive component 232 may include a single convolution operation for each feature channel.
[0057] Output 306 T0 This is then processed by each convolutional decoding layer 242 to produce a confidence map 308 T0 Confidence map 308 T0 Predict whether the pixels of the image include a moving object. In this example, the confidence map 308 T0 The confidence map 308 may include values between about 0 and about 1 that represent the likelihood that a pixel includes a moving object, where values closer to 1 represent pixels that are likely to include a moving object and values closer to 0 represent pixels that are less likely to include a moving object. Alternatively, values closer to 1 may represent pixels that are less likely to include a moving object and values closer to 0 represent pixels that are likely to include a moving object. In general, for each pixel, the confidence map 308 may include a value between about 0 and about 1 that represents the likelihood that a pixel includes a moving object. T0 The probability or confidence level of a pixel comprising a moving object may be indicated. In other words, the confidence map 308 T0 A prediction of the position and / or motion of the moving object in each image frame 202 in the sequence may be provided.
[0058] At time T1, the system 100 captures image frame 202 T1 The deep learning network 210 can be combined with the image frame 202 T0 The same operation is applied to the image frame 202 T1 However, the encoded features 304 generated by each convolutional coding layer 222 T1 Before being passed to the convolutional decoding layer 242, the output 306 from the previous time T0 is T0 Cascade (as shown by arrow 301). Output 306 of the transfer is performed at each spatial resolution layer T0 and the current encoding feature 304 T1 The previous output 306 at time T0 T0and the current encoding features 304 at each spatial resolution layer T0 The cascade allows the recursive part of the network 210 to be able to learn the input image frame 202 at the current time T1. T1 Before making a prediction, the features at every past time point and every spatial resolution level (e.g., from coarse to fine) are fully exposed. Figure 4 Capturing temporal information at each spatial resolution layer is described in more detail.
[0059] Figure 4 is a schematic diagram illustrating a configuration 400 for a time-aware deep learning network 210 according to aspects of the present invention. Figure 4 A more detailed view of the internal operations at the deep learning network 210 is provided. For simplicity of discussion and illustration, Figure 4 The operation of the deep learning network 210 on a single input image frame 202 is illustrated (e.g., at time T1). However, similar operations can be applied to each image frame 202 in the sequence. Additionally, operations are shown for four different spatial resolutions 410, 412, 414, and 416. However, similar operations can be applied to any suitable number of spatial resolutions (e.g., approximately 2, 3, 5, 6, or more). Figure 4 An expanded diagram of the RNN 230 is provided. As shown, the RNN 230 includes a recurrent component 232 at each spatial resolution 410, 412, 414, and 416 to capture temporal information at each spatial resolution 410, 412, 414, and 416. For the spatial resolutions 410, 412, 414, and 416, the recurrent component 232 is shown as 232, respectively. R0 , 232 R1 , 232 R2 and 232 R3 In addition, each convolutional encoding layer 222 is followed by a downsampling layer 422 , and each convolutional decoding layer 242 is preceded by an upsampling layer 442 .
[0060] At time T1, image frame 202 T1 is captured and fed into the deep learning network 210. Image frame 202 T1 After convolutional coding layer 222 K0 , 222 K1 , 222 K2 and 222 K3 Each of the image frames 202 T1 It may have a spatial resolution 410. As shown, the image frame 202 T1 With convolutional coding layer 222 K0 Convolution is performed to output the encoded features 304 at a spatial resolution 410 T1,K0(e.g., in the form of a tensor). Convolutional coding layer 222 K0 The output of the downsampling layer 422 D0 Downsampling to produce a tensor 402 at a spatial resolution 412 D0 Tensor 402 D0 With convolutional coding layer 222 K1 Convolution is performed to output the encoded features 304 at a spatial resolution of 412 T1,K1 Convolutional coding layer 222 K1 The output of is downsampled by the downsampling layer 422DI to produce a tensor 402 with a spatial resolution 414 D1 Tensor 402 D1 With convolutional coding layer 222 K2 Convolution is performed to output the encoded features 304 at a spatial resolution 414 T1,K2 Convolutional coding layer 222 K2 The output of the downsampling layer 422 D2 Downsample to produce a tensor 402 at a spatial resolution 416 D2 Tensor 402 D2 With convolutional coding layer 222 K3 Convolution is performed to output the encoded features 304 at a spatial resolution 416 T1,K3 .
[0061] Temporal continuity information is captured at each of the spatial resolutions 410, 412, 414, and 416. At the spatial resolution 410, the encoded feature 304 T1,K0 and the recursive component 232 obtained at the previous time T0 R0 The output of 306 T0,K0 Concatenation for convolutional coding layer 222 K0 For example, the previous output 306 T0,K0 Stored in memory (e.g., memory 138) at time T0 and retrieved from memory for processing at time T1. Retrieve previous recursive component output from memory 306 T0,K0 Illustrated by blank filled arrows. Recursive component 232 R0 Applied to encoding feature 304 T1,K0 and output 306 T0,K0 cascaded to produce output 306 T1,K0 In some examples, output 306 may be T1,K0 Downsampling is performed so that the output is 306 T1,K0 May have the same encoding characteristics as 304 T1,K0 Same dimensions. Output 306 T1,K0 is stored in memory (shown by the hatched arrows) and can be retrieved at the next time T2 for a similar cascade.
[0062] Similarly, at spatial resolution 412, the encoded feature 304 T1,K1 and the recursive component 232 obtained at the previous time T0 R1 The output of 306 T0,K1 Cascade. Recursive component 232 R1 Applied to encoding feature 304 T1,K1 and output 306 T0,K1 cascaded to produce output 306 T1,K1 Output 306 T1,K1 is stored in memory (shown by the pattern-filled arrows) for use in a similar cascade at the next time T2.
[0063] At spatial resolution 414, the encoded features 304 T1,K2 and the recursive component 232 obtained at the previous time T0 R2 The output of 306 T0,K2 Cascade. Recursive component 232 R2 Applied to encoding feature 304 T1,K2 and output 306 T0,K2 cascaded to produce output 306 T1,K2 Output 306 T1,K2 is stored in memory (shown by the pattern-filled arrows) for a similar cascade at the next time T2.
[0064] At the final spatial resolution 416, the encoded features 304 T1,K3 and the recursive component 232 obtained at the previous time T0 R3 The output of 306 T0,K2 Connection. Recursive component 232 R3 Applied to encoding feature 304 T1,K3 and output 306 T0,K3 cascaded to produce output 306 T1,K3 Output 306 T1,K3 is stored in memory (shown by the pattern-filled arrows) for a similar cascade at the next time T2.
[0065] Output 306 T1,K3 、306 T1,K2 、306 T1,K1 and 306 T1,K0 are passed to the convolutional decoding layer 242 respectively L0 , 242 L1 and 242 L2 For example, output 306 T1,K3 By upsampling layer 442 U0 Upsample to produce tensor 408 U0(eg, including extracted features). Tensor 408 U0 and output 306 T1,K2 With convolutional decoding layer 242 L0 Convolution is performed and upsampled by upsampling layer 442U1 to produce tensor 408 U1 Tensor 408 U1 and output 306 T1,K1 With convolutional decoding layer 242 L1 Convolution is performed and the upsampling layer 442 U2 Upsample to produce tensor 408 U2 Tensor 408 U2 and output 306 T1,K0 With convolutional decoding layer 242 L2 Convolution is performed to produce a confidence map 308 T1 .Although Figure 4 Four encoding layers 222 and three decoding layers 242 are illustrated, but the network 210 may alternatively be configured to include four decoding layers 242 to provide similar prediction. Figure 4 The number of encoding layers 222 can be determined based on the size of the input volume and the receptive field of the network 210. The depth of the network 210 can be varied based on how large the input image is and its effect on the learned features (i.e., by controlling the receptive field of the network 210). Thus, the network 210 may not have a decoder / upsampling layer corresponding to the innermost layer. Figure 4 (shown on the right side of network 210 in ) takes features from lower resolution feature maps and assembles them while upsampling towards the original output size.
[0066] As can be observed, the deep learning network 210 performs a prediction for the current image frame 202 (at time Tn) based on features extracted from the current image frame 202 and the previous image frame 202 (at time Tn-1) rather than based on a single image frame captured at a single point in time. The deep learning network 210 can infer motion and / or position information associated with a moving object based on past information. Temporal continuity information (e.g., provided by temporal concatenation) can provide an additional dimension of information. The use of temporal information can be particularly useful in segmenting thin objects because thin objects can typically be represented by a relatively smaller number of pixels in an imaging frame than thicker objects. Therefore, the present disclosure can improve the visualization and / or stability of ultrasound images and / or videos of mobile medical devices and / or anatomical structures that include moving parts.
[0067] The downsampling layers 422 may perform downsampling by any suitable downsampling factor. In one example, each downsampling layer 422 may perform downsampling by a factor of 2. For example, the input image frame 202 T1 The input image frame 202 has a resolution of 200×200×200 voxels (eg, spatial resolution 410). T1 Downsampled by 2 to produce tensor 402 at a resolution of 100×100×100 voxels (eg, spatial resolution 412 ) D0 Tensor 402 D0 Tensor 402 is downsampled by 2 to produce a resolution of 50×50×50 voxels (eg, spatial resolution 414 ). D1 Tensor 402 D1 Tensor 402 is downsampled by 2 to produce a resolution of 25×25×25 voxels (eg, spatial resolution 416) D2 The upsampling layers 442 may invert the downsampling. For example, each of the upsampling layers 442 may perform upsampling by a factor of 2. In some other examples, the downsampling layers 422 may perform downsampling by different downsampling factors, and the upsampling layers 442 may perform upsampling using a factor that matches the downsampling factor. For example, the downsampling layers 422 may perform downsampling by different downsampling factors. D0 、422 D1 and 422 D2 Downsampling can be performed by 2, 4 and 8 respectively, and upsampling layer 442 U0 , 442 U1 and 442 U2 Upsampling can be performed by 8, 4, and 2 respectively.
[0068] Convolutional encoding layer 222 and convolutional decoding layer 242 may include convolution kernels of any size. In some examples, the kernel size may depend on the size of input image frame 202 and may be selected to limit network 210 to a particular complexity. In some examples, each convolutional encoding layer 222 and each convolutional decoding layer 242 may include 5×5×5 convolution kernels. In an example, convolutional encoding layer 222 K0 Approximately one feature (e.g., feature 304) may be provided at a spatial resolution 410. T1,K0 Convolutional coding layer 222 K1 Approximately two features (e.g., feature 304) may be provided at a spatial resolution 412. T1,K1 Convolutional coding layer 222 K2 Approximately four features (e.g., feature 304) may be provided at a spatial resolution 414. T1,K2 Convolutional coding layer 222 K3Approximately eight features (e.g., feature 304) may be provided at a spatial resolution 416. T1,K3 with a size of 8). In general, the number of features can increase as the spatial resolution decreases.
[0069] In some embodiments, the convolution at the convolutional encoding layer 222 and / or the convolutional decoding layer 242 may be repeated. For example, the convolutional encoding layer 222 K0 The convolution at can be repeated twice, the convolutional coding layer 222 K1 The convolution at can be performed once, the convolutional coding layer 222 K2 The convolution at can be repeated twice, and the convolutional coding layer 222 K3 The convolution at can be repeated twice.
[0070] In some embodiments, each convolutional encoding layer 222 and / or each convolutional decoding layer 242 may include a nonlinear function (e.g., a parametric rectified linear unit (PReLu)).
[0071] In some examples, each recurrent component 232 can include a convolutional gated recurrent unit (convGRU). In some examples, each recurrent component 232 can include a convolutional long short-term memory (convLSTM).
[0072] Although Figure 4 The propagation of time information over two time points (eg, from T0 to T1 or from T1 to T2) is illustrated, but in some examples, the time information may be propagated over a greater number of time points (eg, approximately 3 or 4).
[0073] Return to Figure 2 , the deep learning network 210 can output a confidence map 308 for each image frame 202. As described above, for each pixel in the image frame 202, the corresponding confidence map 308 can include a probability or confidence level that the pixel includes a moving object. The sequence of output image frames 206 can be generated based on the sequence of input image frames 202 and the corresponding confidence map 308. In some examples, time-aware inference can interpolate or otherwise predict missing image information of the moving object based on the confidence map 308. In some examples, the inference, interpolation, and / or prediction can be implemented external to the deep learning network 210. In some examples, the interpolation and / or reconstruction can be implemented as part of the deep learning network 210. In other words, the learning and training of the deep learning network 210 can include inferring, interpolating, and / or predicting missing imaging information.
[0074] In an embodiment, a deep learning network 210 can be trained to distinguish an elongated, flexible, thin, mobile medical device (e.g., a guidewire, a guide catheter, a catheter, a needle, a therapeutic device, and / or a treatment device) from an anatomical structure. For example, a training dataset (e.g., an image dataset 140) can be created for training using the system 100. The training dataset can include input-output pairs. For each input-output pair, the input can include a sequence of image frames (e.g., 2D or 3D) of a medical device (e.g., device 108) traversing an anatomical structure (e.g., object 105) over time, and the output can include ground truth or annotations of the position of the medical device within each image frame in the sequence. In an example, the ground truth position of the medical device can be obtained by attaching an ultrasound sensor to the medical device during imaging (e.g., at the ends of the medical device) and then fitting a curve or spline to the captured image using at least the ends as endpoint constraints for the spline. After fitting the curve to the ultrasound image, the image can be annotated or labeled with the ground truth for training. During training, the deep learning network 210 can be applied to the sequence of image frames using forward propagation to generate the output. Backpropagation can be used to adjust the coefficients of the convolution kernels at the convolutional encoding layer 222, the recursive component 232, and / or the convolutional decoding layer 242 to minimize the error between the output of the device and the ground truth position. The training process can be repeated for each input-output pair in the training dataset.
[0075] In another embodiment, a training dataset (e.g., image dataset 140) can be used to train the deep learning network 210 to distinguish moving parts of anatomical structures from static parts of anatomical structures. For example, a training dataset (e.g., image dataset 140) can be created for training using system 100. The training dataset can include input-output pairs. For each input-output pair, the input can include a sequence of image frames (e.g., 2D or 3D) of an anatomical structure with motion (e.g., associated with the heart, breathing, and / or arterial pulse), and the output can include ground truth or annotations of various moving and / or static parts of the anatomical structure. The ground truth and / or annotations can be obtained from various annotated datasets available in the medical community. Alternatively, the sequence of image frames can be manually annotated with ground truth. After obtaining the training dataset, a similar mechanism as described above (e.g., for moving objects) can be used to train the deep learning network 210 to segment moving anatomical structures.
[0076] Figure 5-8 Various clinical use case scenarios are illustrated where the time-aware deep learning network 210 can be used to provide improved segmentation based on a series of observations over time.
[0077] Figure 5Scenario 500 of an ultrasound-guided procedure according to aspects of the present disclosure is illustrated. Scenario 500 may correspond to a scenario when system 100 is used to capture ultrasound images of a thin guidewire 510 (e.g., medical device 108) passing through a vessel lumen 504 having a vessel wall 502 including an occluded area 520 (e.g., plaque and / or calcification). For example, a sequence of ultrasound images is captured at times T0, T1, T2, T3, and T4. Figure 5 The column to the right of includes a check mark and a cross mark. A check mark indicates that the guide wire 510 is completely visible in the corresponding image frame. A cross mark indicates that the guide wire 510 is not completely visible in the corresponding image frame.
[0078] At time T0, the guidewire 510 enters the lumen 504. At time T1, the beginning portion 512a of the guidewire 510 (shown by a dotted line) enters the occluded region 520. At time T2, the guidewire 510 continues to pass through the lumen 504, wherein the middle portion 512b of the guidewire 510 (shown by a dotted line) is within the occluded region 520. At time T3, the guidewire 510 continues to pass through the lumen 504, wherein the end portion 512c of the guidewire 510 (shown by a dotted line) is within the occluded region 520. At time T, the guidewire 510 exits the occluded region 520.
[0079] Conventional 3D segmentation without utilizing temporal information may be unable to segment portions 512a, 512b, and 512c within the occluded region 520 at times T1, T2, and T3, respectively. Consequently, image frames acquired at times T1, T2, and T3 without temporal information may each include missing segments, sections, or portions of the guidewire 510 corresponding to portions 512a, 512b, and 512c, respectively, within the occluded region 520. Consequently, crosses are shown for times T1, T2, and T3 under the column for segmentation without utilizing temporal information.
[0080] The time-aware deep learning network 210 is designed to interpolate missing information based on previous image frames, and thus the system 100 can apply the deep learning network 210 to infer missing portions 512a, 512b, and 512c in the image. Therefore, check marks are shown for times T1, T2, and T3 under the column for segmentation utilizing time information.
[0081] In some examples, scenario 500 can resemble a peripheral vascular intervention procedure, where occlusion region 520 can correspond to a chronic total occlusion (CTO) crossing a peripheral vascular structure. In some examples, scenario 500 can resemble a clinical procedure in which a tracking device passes through an air gap, calcification, or shadowed area (e.g., occlusion region 520).
[0082] Figure 6Scenario 600 of an ultrasound-guided procedure according to aspects of the present disclosure is illustrated. Scenario 600 may correspond to a scenario when the system 100 is used to capture ultrasound images of a guidewire 610 (e.g., a medical device 108) passing through a vessel lumen 604 having a vessel wall 602, where the guidewire 610 may slide along the vessel wall 602 for a period of time. For example, a sequence of ultrasound images is captured at times T0, T1, T2, T3, and T4. Figure 6 The column to the right of includes a check mark and a cross mark. A check mark indicates that the guide wire 610 is completely visible in the corresponding image frame. A cross mark indicates that the guide wire 610 is not completely visible in the corresponding image frame.
[0083] At time T0, guidewire 610 initially enters lumen 604 near the center of lumen 604. At time T1, portion 612a of guidewire 610 (shown by dashed lines) slides against vessel wall 602. Guidewire 610 continues to slide against vessel wall 602. As shown, at time T2, portion 612b of guidewire 610 (shown by dashed lines) is adjacent to vessel wall 602. At time T3, portion 612c of guidewire 610 (shown by dashed lines) is adjacent to vessel wall 602. At time T4, portion 612d of guidewire 610 (shown by dashed lines) is adjacent to vessel wall 602.
[0084] The guidewire 610 may reflect similarly to the vessel wall 602, and therefore, conventional 3D segmentation without utilizing time information may be unable to segment portions 612a, 612b, 612c, and 612d that are close to the vessel wall 602 at times T1, T2, T3, and T4, respectively. Consequently, image frames acquired at times T1, T2, T3, and T4 without time information may each include missing segments, sections, or portions of the guidewire 610 corresponding to portions 612a, 612b, 612c, and 612d, respectively. Consequently, crosses are shown for times T1, T2, T3, and T4 under the column for segmentation without utilizing time information.
[0085] The time-aware deep learning network 210 is exposed to the entire sequence of ultrasound image frames or video frames across time and can therefore be applied to the sequence of images to predict the position and / or motion of portions 612a, 612b, 612c, and 612d near the vessel wall 602 at times T1, T2, T3, and T4, respectively. Therefore, check marks are shown for times T1, T2, T3, and T4 under the column for segmentation utilizing time information.
[0086] In some examples, scenario 600 can resemble a cardiac imaging procedure, where a medical device or guidewire is slid along the wall of a heart chamber. In some examples, scenario 600 can resemble a peripheral vascular intervention procedure, where a device or guidewire is purposefully guided subintimally into the adventitia of a vessel wall to bypass an occlusion.
[0087] Figure 7 Scenario 700 of an ultrasound-guided procedure according to aspects of the present disclosure is illustrated. Scenario 700 may correspond to a scenario when system 100 is used to capture ultrasound images of a guidewire 710 (e.g., medical device 108) passing through a vessel lumen 704 having a vessel wall 702, wherein acoustic coupling is lost for a period of time. For example, a sequence of ultrasound images is captured at times T0, T1, T2, T3, and T4. Figure 7 The column to the right of includes a check mark and a cross mark. A check mark indicates that the guidewire 710 is completely visible in the corresponding image frame. A cross mark indicates that the guidewire 710 is not completely visible in the corresponding image frame.
[0088] At time T0, guidewire 710 enters lumen 704. Acoustic coupling is lost at times T1 and T2. Acoustic coupling is regained at time T3. When acoustic coupling is lost, conventional 3D imaging that does not utilize temporal information may lose all knowledge of the position of guidewire 610. Thus, without temporal information, guidewire 710 may not be visible in image frames acquired at times T1 and T2. Therefore, crosses are shown for times T1 and T2 under the column for segmentation that does not utilize temporal information.
[0089] The time-aware deep learning network 210 has the ability to remember the position of the guidewire 710 for at least several frames and can therefore be applied to a sequence of images to predict the position of the guidewire 710 at time T1 and time T2. Therefore, check marks are shown for times T1 and T2 under the column for segmentation using temporal information. If acoustic coupling is lost for an extended period of time, the time-aware deep learning network 210 is less likely to produce incorrect segmentation results.
[0090] Scenario 700 can occur whenever acoustic coupling is lost. Maintaining acoustic coupling throughout imaging can be difficult. Therefore, time-aware deep learning-based segmentation can improve visualization of various devices and / or anatomical structures in ultrasound images, particularly when automation is involved, such as during automated beam steering, sensor tracking using image-based constraints, and / or robotic control of ultrasound imaging equipment. In other scenarios, acoustic coupling can be lost for short periods during cardiac imaging, for example, due to cardiac motion. Therefore, time-aware deep learning-based segmentation can improve visualization in cardiac imaging.
[0091] Figure 8Scenario 800 of an ultrasound-guided procedure according to aspects of the present disclosure is illustrated. Scenario 800 may correspond to a scenario when system 100 is used to capture ultrasound images of a guidewire 810 (e.g., medical device 108) passing through a vessel lumen 804 having a vessel wall 802, wherein the guidewire 810 may move in and out of planes during imaging. For example, a sequence of ultrasound images may be captured at times T0, T1, T2, T3, and T4. Figure 8 The column to the right of includes a check mark and a cross mark. A check mark indicates that the guide wire 810 is completely visible in the corresponding image frame. A cross mark indicates that the guide wire 810 is not completely visible in the corresponding image frame.
[0092] At time T0, the guide wire 810 enters the lumen 804 and is in the plane under imaging. At time T1, the guide wire 810 begins to drift out of the plane (e.g., partially out of plane). At time T2, the guide wire 810 is completely out of plane. At time T3, the guide wire 810 continues to drift and is partially out of plane. At time T4, the guide wire 810 moves back into the plane. General 3D imaging that does not utilize time information may not be able to detect any structure out of plane. Therefore, in the absence of time information, the guide wire 810 may not be completely visible in the image frames obtained at time T1, T2 and T3. Therefore, a cross is shown for time T1, T2 and T3 under the column for segmentation that does not utilize time information.
[0093] The time-aware deep learning network 210 is able to predict out-of-plane device positions to provide full visibility of the device and can therefore be applied to a sequence of images to predict the position of the guidewire 810. Therefore, check marks are shown for times T1, T2, and T3 under the column for segmentation that does not utilize time information.
[0094] In some examples, scenario 800 may occur in an ultrasound-guided procedure where a non-volumetric imaging mode (e.g., 2D imaging) is used. In some examples, scenario 800 may occur in real-time 3D imaging where a relatively small 3D volume is acquired in the transverse direction in order to maintain a sufficiently high frame rate. In some examples, scenario 800 may occur in cardiac imaging where the motion of the heart may cause specific portions of the heart to enter and leave the imaging plane.
[0095] Although scenarios 500-800 illustrate the use of a time-aware deep learning network 210 for providing segmentation of a moving guidewire (e.g., guidewires 510, 610, 710, and / or 810), similar time-aware deep learning-based segmentation mechanisms can be applied to any elongated, flexible, thin, moving device (e.g., a catheter, a guide catheter, a needle, an IVUS device, and / or a therapeutic device) and / or anatomical structure with moving parts. In general, time-aware deep learning-based segmentation can be used to improve visualization and / or stability of moving devices and / or anatomical structures with motion under imaging. In other words, time-aware deep learning-based segmentation can minimize or remove discontinuities in the motion of a moving device and / or moving anatomical structure.
[0096] Figure 9 FIG2 is a schematic diagram of a deep learning-based image segmentation scheme 900 utilizing spline fitting according to aspects of the present invention. Scheme 900 is implemented by system 100. Scheme 900 is substantially similar to scheme 200. For example, scheme 900 utilizes a time-aware multi-layer deep learning network 210 to provide segmentation of moving objects in ultrasound images. In addition, scheme 900 includes a spline fitting component 910 coupled to the output of deep learning network 210. Spline fitting component 910 may be implemented by processing component 134 at system 100.
[0097] The spline fitting component 910 is configured to apply a spline fitting function to the confidence map 308 output by the deep learning network 210. An expanded view of the confidence map 308 for the image frames 202 in the sequence is shown as a heat map 902. As shown, the deep learning network 210 predicts a moving object, as shown by curve 930. However, curve 930 is discontinuous and includes a gap 932. The spline fitting component 910 is configured to fit a spline 934 to smooth the discontinuity of curve 930 at the gap 932. The spline fitting component 910 can perform spline fitting by considering device parameters 904 associated with the moving object being imaged. Device parameters 904 can include the shape of the device, the position of the end of the device, and / or other dimensional and / or geometric information of the device. Therefore, using spline fitting as a post-processing improvement to the prediction based on temporal deep learning can further improve the visualization and / or stability of moving objects under imaging.
[0098] Figure 10FIG2 is a schematic diagram of a deep learning-based imaging guidance scheme 1000 according to aspects of the present disclosure. Scheme 1000 is implemented by system 100. Scheme 1000 is substantially similar to scheme 200. For example, scheme 1000 utilizes a time-aware multi-layer deep learning network 1010 to provide imaging guidance for ultrasound imaging. Deep learning network 1010 can have a substantially similar architecture to deep learning network 210. For example, deep learning network 1010 includes a convolutional encoder 1020, a time-aware RNN 1030, and a convolutional decoder 2100. Convolutional encoder 1020 includes a plurality of convolutional encoding layers 1022. Convolutional decoder 1040 includes a plurality of convolutional decoding layers 1042. The convolutional encoding layer 1022, the convolutional decoding layer 1042, and the RNN 1030 are substantially similar to the convolutional encoding layer 222, the convolutional decoding layer 242, and the RNN 230, respectively, and can operate at a plurality of different spatial resolutions (e.g., spatial resolutions 410, 412, 414, and 416), as shown in configuration 400. However, the convolutional encoding layer 1022, the convolutional decoding layer 1042, and the RNN 1030 are trained to predict the optimal imaging plane for imaging the target anatomical structure (e.g., including a particular clinical attribute of interest). The optimal imaging plane can be a 2D plane, an X-plane (e.g., including a cross-sectional plane and an orthogonal imaging plane), an MPR, or any suitable imaging plane.
[0099] For example, a sequence of image frames 1002 is captured over a time period (e.g., from time T0 to time Tn). The system 100 can be used to capture the image frames 202. The deep learning network 1010 can be applied to the sequence of image frames 1002 to predict the optimal imaging plane. As an example, a sequence of input image frames 1002 is captured as a medical device 1050 (e.g., medical device 108) passes through a vessel lumen 1052 having a vessel wall 1054 (e.g., subject 105). The output of the deep learning network 1010 provides the optimal long-axis slice 1006 and short-axis slice 1008. Similar to the scheme 200, each image frame 1002 is processed by each convolutional encoding layer 1022 and each convolutional decoding layer 1042. The RNN 1030 passes the prediction of the current image frame 1002 (captured at time T0) back to the RNN 1030 as an auxiliary input for the prediction of the next image frame 1002 (captured at time T1), as shown by arrow 1004.
[0100] In a first example, the system 100 can automatically steer the ultrasound beam to an optimal position using the predictions output by the deep learning network 1010. For example, the processing component 116 and / or 134 can be configured to control or steer the ultrasound beam generated by the transducer array 112 based on the predictions.
[0101] In a second example, the deep learning network 1010 may predict that the optimal imaging plane is an inclined plane. The deep learning network 1010 may provide the user with navigation instructions for manipulating (e.g., rotating and / or translating) the ultrasound probe 110 to align the axis of the probe 110 with the predicted optimal plane. In some examples, the navigation instructions may be displayed on a display similar to the display 132. In some examples, the navigation instructions may be displayed using a graphical representation (e.g., a rotation symbol or a translation symbol). After the user repositions the probe 110 to the suggested position, the imaging plane may be in a non-tilted plane. Thus, the deep learning network 1010 may transition to provide a prediction as described in the first example and may communicate with the processing component 116 and / or 134 to steer the beam generated by the transducer array 112.
[0102] Figure 11 Ultrasound images 1110, 1120, and 1130 obtained from an ultrasound-guided procedure according to aspects of the present disclosure are illustrated. Image 1110 is a 3D image captured during a PVD examination using a system similar to system 100. Image 1110 shows a thin guidewire 1112 (e.g., medical device 108 and / or 1050) passing through a vessel lumen 1114 surrounded by a vessel wall 1116 (e.g., object 105). Device 1112 is passed through the vessel along the x-axis. As device 1112 passes through the vessel, the system can capture a series of 3D images similar to image 1110. As described above, the movement of device 1112 can cause device 1112 to move in and out of the imaging view. In addition, the thin geometry of device 1112 can cause challenges in distinguishing device 1112 from anatomical structures (e.g., vessel lumen 1114 and / or vessel wall 1116).
[0103] To improve visualization, a time-aware deep learning network trained for segmentation and / or imaging guidance can be applied to a series of 3D images (including image 1110). The predictions generated by deep learning network 1010 are used to automatically set an MPR that passes through the end of device 1112 and is aligned with the long axis (e.g., x-axis and y-axis) of device 1112. Images 1120 and 1130 are generated based on the deep learning segmentation. Image 1120 shows a longitudinal MPR (along the zx plane) constructed from image 1110 based on the predictions output by the deep learning network. Image 1130 shows a transverse MPR (along the yz plane) constructed from image 1110 based on the predictions. Orthogonal MPR planes (e.g., images 1120 and 1130) are generated based on the predicted segmentation. In this case, images 1120 and 1130 correspond to the longitudinal and sagittal planes, respectively, passing through the end of the segmented device 1112, but similar mechanisms can also be used to generate other MPR planes.
[0104] In some cases, the device 1112 may be positioned in close proximity to an anatomical structure (e.g., a vessel wall) and may be similarly reflective as the anatomical structure. Consequently, it may be difficult for the clinician to visualize the device 1112 from the captured images. To further improve visualization, the images 1120 and 1130 may be color-coded. For example, the anatomical structure may be shown in grayscale, and the device 1112 may be shown in red or any other suitable color.
[0105] Figure 12 is a schematic diagram of a processor circuit 1200 according to an embodiment of the present disclosure. The processor circuit 1200 may be Figure 1 The processor circuit 1200 may be implemented in the probe 110 and / or the host 130. As shown, the processor circuit 1200 may include a processor 1260, a memory 1264, and a communication module 1268. These elements may communicate with each other directly or indirectly, for example, via one or more buses.
[0106] The processor 1260 may include a processor configured to perform the operations described herein (e.g., Figure 1-11 and aspects 13-15) of the present invention (a CPU, a DSP, an application specific integrated circuit (ASIC), a controller, an FPGA, another hardware device, a firmware device, or any combination thereof. The processor 1260 may also be implemented as a combination of computing devices, for example, a combination of a DSP and a microprocessor, a plurality of microprocessors, one or more microprocessors in conjunction with a DSP core, or any other such configuration.
[0107] Memory 1264 may include cache memory (e.g., cache memory of processor 1260), random access memory (RAM), magnetoresistive RAM (MRAM), read-only memory (ROM), programmable read-only memory (PROM), erasable programmable read-only memory (EPROM), electrically erasable programmable read-only memory (EEPROM), flash memory, solid-state memory devices, hard disk drives, solid-state drives, other forms of volatile and non-volatile memory, or a combination of different types of memory. In one embodiment, memory 1264 includes non-transitory computer-readable media. Memory 1264 may store instructions 1266. Instructions 1266 may include instructions that, when executed by processor 1260, cause processor 1260 to perform the operations described herein (e.g., Figure 1-11 and 13-15 and with respect to the probe 110 and / or the host 130 ( Figure 1) aspects of the present invention). Instructions 1266 may also be referred to as code. The terms "instructions" and "code" should be interpreted broadly to include any type of computer-readable representation(s). For example, the terms "instructions" and "code" may refer to one or more programs, routines, subroutines, functions, processes, etc. The terms "instructions" and "code" may include a single computer-readable representation or many computer-readable representations.
[0108] The communication module 1268 may include any electronic circuitry and / or logic circuitry to facilitate direct or indirect data communication between the processor circuit 1200, the probe 110, and / or the display 132. In this regard, the communication module 1268 may be an input / output (I / O) device. In some cases, the communication module 1268 facilitates communication between the processor circuit 1200 and / or the probe 110 ( Figure 1 ) and / or host 130( Figure 1 ) direct or indirect communication between various elements.
[0109] Figure 13 1300 is a flow chart of a deep learning-based ultrasound imaging method 1300 according to aspects of the present disclosure. The method 1300 is implemented by the system 100 (e.g., by a processor circuit (such as the processor circuit 1200) and / or other suitable components (such as the probe 110, the processing component 114, the host 130 and / or the processing component 134)). In some examples, the system 100 may include a computer-readable medium having program code recorded thereon, the program code including code for causing the system 100 to perform the steps of the method 1300. The method 1300 may be implemented in the manner described in connection with the respective embodiments of the present disclosure. Figure 2 、 9 , 10 described in the scheme 200, 900 and / or 1000, respectively, about Figure 3 and 4 The configurations 300 and / or 400 described, and / or respectively with respect to Figure 5 、 6 , 7 and / or 8. As shown, method 1300 includes several enumerated steps, but embodiments of method 1300 may include additional steps before, after, and between the enumerated steps. In some embodiments, one or more of the enumerated steps may be omitted or performed in a different order.
[0110] At step 1310, method 1300 includes receiving, by a processor circuit (e.g., processing component 116 and / or 134 and / or processor circuit 1200), a sequence of input image frames (e.g., image frame 202) of a moving object over a time period (e.g., spanning times T0, T1, T2, ..., Tn) from an ultrasound imaging device (e.g., probe 110). The moving object includes at least one of an anatomical structure of a patient or a medical device passing through the anatomical structure of the patient, and a portion of the moving object is at least partially invisible in a first input image frame in the sequence of input image frames. The first input image frame can be any image frame in the sequence of input image frames. The anatomical structure can be similar to object 105 and can include the patient's heart, lungs, blood vessels (e.g., vessel lumens 504, 604, 705, and / or 804 and vessel walls 502, 602, 702, and / or 802), nerve fibers, and / or any other suitable anatomical structure of the patient. The medical device is similar to medical device 108 and / or guidewires 510 , 610 , 710 , and / or 810 .
[0111] At step 1320 , method 1300 includes applying, by a processor circuit, a recursive prediction network associated with image segmentation (e.g., deep learning network 210 ) to the sequence of input image frames to generate segmentation data.
[0112] At step 1330, the method includes outputting a sequence of output image frames (e.g., image frames 206 and / or 906) based on the segmentation data to a display (e.g., display 132) in communication with the processor circuit. The portion of the moving object is fully visible in a first output image frame in the sequence of output image frames, wherein the first output image frame and the first input image frame are associated with the same moment in time within the time period.
[0113] In some examples, portions of the moving object may be within an occlusion region (e.g., occlusion region 520), e.g., as described above with respect to Figure 5 500. In some examples, the portion of the mobile object can be against the patient's anatomy (e.g., vessel wall 605, 602, 702, and / or 802), for example, as described above with respect to Figure 6 In some examples, when acoustic coupling is low or lost, the portion of the moving object can be captured, for example, as described above with respect to Figure 7 In some examples, when the first input image frame is captured, the portion of the moving object may be out of plane, for example, as described above with respect to Figure 8 Scenario 800 is depicted.
[0114] In an embodiment, applying the recursive prediction network includes generating previous segmentation data based on a previous input image frame in a sequence of input image frames, wherein the previous input image frame is received before the first input image frame, and generating the first segmentation data based on the first input image frame and the previous segmentation data. The previous input image frame can be any image frame in the sequence received before the first input image frame, or an image frame in the sequence immediately preceding the first input image frame. For example, the first input image frame corresponds to the input image frame 202 received at the current time T1. T1 , the first segmented data corresponds to output 306 T1 , the previous input image frame corresponds to the input image frame 202 received at the previous time T0 T0 , and the previously segmented data corresponds to output 306 T0 , such as about Figure 3 The configuration 300 is depicted.
[0115] In an embodiment, generating the previous segmentation data includes applying a convolutional encoder (e.g., convolutional encoder 220) and a recurrent neural network (e.g., RNN 230) to a previous input image frame. Generating the first segmentation data includes applying the convolutional encoder to the first input image frame to generate encoded data, and applying the recurrent neural network to the encoded data and the previous segmentation data. Applying the recursive prediction network also includes applying a convolutional decoder (e.g., convolutional decoder 240) to the first segmentation data and the previous segmentation data. In an embodiment, the convolutional encoder, recurrent neural network, and convolutional decoder operate at multiple spatial resolutions (e.g., spatial resolutions 410, 412, 414, and 416).
[0116] In an embodiment, the moving object comprises a medical device that passes through the patient's anatomy. In such an embodiment, the convolutional encoder, the recurrent neural network, and the convolutional decoder are trained to recognize the medical device from the patient's anatomy and predict motion associated with the medical device that passes through the patient's anatomy.
[0117] In an embodiment, the moving object includes an anatomical structure of a patient having at least one of cardiac motion, respiratory motion, or an arterial pulse. In such an embodiment, the convolutional encoder, the recurrent network, and the convolutional decoder are trained to identify moving portions of the patient's anatomical structure from static portions of the patient's anatomical structure and predict motion associated with the moving portions.
[0118] In an embodiment, the moving object comprises a medical device that passes through the patient's anatomy, and the system comprises the medical device.In an embodiment, the medical device comprises at least one of a needle, a guidewire, a catheter, a guide catheter, a therapeutic device, or an interventional device.
[0119] In an embodiment, the input image frames include 3D image frames, and a recursive prediction network is trained based on temporal information for 4D image segmentation. In an embodiment, the sequence of input image frames includes 2D image frames, and a recursive prediction network is trained based on temporal information for 3D image segmentation.
[0120] In an embodiment, the method 1300 further includes applying spline fitting (e.g., spline fitting component 920) to the sequence of input image frames based on the segmentation data. The spline fitting can utilize spatial and temporal information in the sequence of input image frames and predictions by the recursive prediction network.
[0121] Figure 14 is a flowchart of a deep learning-based ultrasound imaging method according to aspects of the present disclosure. Figure 14 is a flow chart of a deep learning-based ultrasound imaging method according to aspects of the present disclosure. Method 1400 is implemented by system 100 (e.g., by a processor circuit (such as processor circuit 1200) and / or other suitable components (such as probe 110, processing component 114, host 130 and / or processing component 134)). In some examples, system 100 may include a computer-readable medium having program code recorded thereon, the program code including code for causing system 100 to perform the steps of method 1400. Method 1400 may be implemented in the manner described with respect to Figure 10 The described scheme 1000 is respectively about Figure 3 and 4 Similar mechanisms are described in configurations 300 and 400. As shown, method 1400 includes a plurality of enumerated steps, but embodiments of method 1400 may include additional steps before, after, and between the enumerated steps. In some embodiments, one or more of the enumerated steps may be omitted or performed in a different order.
[0122] At step 1410, method 1400 includes receiving, by a processor circuit (e.g., processing components 116 and / or 134 and / or processor circuit 1200), a sequence of image frames (e.g., image frames 1002 and / or 1110) representing an anatomical structure of a patient over a time period (e.g., spanning times T0, T1, T2, ... Tn) from an ultrasound imaging device (e.g., probe 110). The anatomical structure may be similar to object 105 and may include the patient's heart, lungs, and / or any anatomical structure.
[0123] At step 1420, method 1400 includes applying a recursive prediction network associated with the image acquisition (e.g., deep learning network 1010) to the sequence of image frames to generate imaging plane data associated with a clinical attribute of the patient's anatomy. The clinical attribute can be associated with a cardiac condition, a pulmonary condition, and / or any other clinical condition.
[0124] At step 1430, method 1400 includes outputting a target imaging plane (e.g., a cross-sectional plane, a longitudinal plane, or an MPR plane) of the patient's anatomical structure or at least one of instructions for repositioning the ultrasound imaging device toward the target imaging plane based on the imaging plane data to a display (e.g., display 132) in communication with the processor circuit.
[0125] In an embodiment, applying the recursive prediction network includes generating first imaging plane data based on a first image frame of the sequence of images, and generating second imaging plane data based on a second image frame in the sequence of image frames and the first imaging plane data, the second image frame being received after the first image frame. For example, the first image frame corresponds to the input image frame 1002 received at a previous time T0, the first imaging plane data corresponds to the output of the RNN 1030 at time T0, and the second image frame corresponds to the input image frame 1002 received at a current time T1. T1 , and the second imaging plane data corresponds to the output of RNN1030 at time T1, as described with respect to Figure 10 The scheme 1000 is depicted.
[0126] In an embodiment, generating first imaging plane data includes applying a convolutional encoder (e.g., convolutional encoder 1020) and a recurrent neural network (e.g., RNN 1030) to a first image frame. Generating second imaging plane data includes applying a convolutional encoder to a second image frame to generate encoded data, and applying a recurrent neural network to the encoded data and the first imaging plane data. Applying the recursive prediction network also includes applying a convolutional decoder (e.g., convolutional decoder 1040) to the first imaging plane data and the second imaging plane data. In an embodiment, the convolutional encoder, recurrent neural network, and convolutional decoder operate at multiple spatial resolutions (e.g., spatial resolutions 410, 412, 414, and 416). In an embodiment, the convolutional encoder, recurrent network, and convolutional decoder are trained to predict a target imaging plane for imaging clinical properties of an anatomical structure of a patient.
[0127] In an embodiment, the input image frames include 3D image frames, and the recursive prediction network is trained based on temporal information for 3D image acquisition. In an embodiment, the sequence of input image frames includes 2D image frames, and the recursive prediction network is trained based on temporal information for 2D image acquisition.
[0128] In an embodiment, method 1400 outputs a target imaging plane comprising at least one of: a cross-sectional image slice (e.g., image slice 1006 and / or 1120), an orthogonal image slice (e.g., image slice 1008 and / or 1130), or a multi-planar MPR image slice of an anatomical structure of a patient having clinical attributes.
[0129] In an embodiment, the method 1400 includes generating an ultrasound beam steering control signal based on the imaging plane data and outputting the ultrasound beam steering control signal to the ultrasound imaging device. For example, the ultrasound beam steering control signal may steer an ultrasound beam generated by a transducer array (e.g., the transducer array 112) of the ultrasound imaging device.
[0130] In an embodiment, the processor circuit outputs instructions including at least one of rotation or translation of the ultrasound imaging device. The instructions may provide guidance to a user to steer the ultrasound imaging device to an optimal imaging position (e.g., a target imaging plane) in order to obtain a target image view of the patient's anatomy.
[0131] Various aspects of the present disclosure can provide several benefits. For example, the use of temporal continuity information in a deep learning network (e.g., deep learning networks 210 and 1010) allows the deep learning network to learn and predict based on a series of observations over time rather than at a single point in time. Temporal continuity information provides an additional dimension of information that can improve the segmentation of moving objects of elongated, flexible, thin shapes that may otherwise be difficult to segment. Thus, the disclosed embodiments can provide a stable view of the motion of a moving object under 2D and / or 3D imaging. The use of spline fitting as an improvement to the output of the deep learning network can further provide smooth transitions in motion associated with the moving object under imaging. The use of temporal continuity information can also provide automatic view finding when reaching a target imaging view, for example, including beam steering control and / or imaging guidance instructions.
[0132] Those skilled in the art will recognize that the above-described devices, systems, and methods can be modified in various ways. Therefore, those skilled in the art will appreciate that the embodiments encompassed by the present disclosure are not limited to the specific exemplary embodiments described above. In this regard, although exemplary embodiments have been shown and described, various modifications, changes, and substitutions are contemplated in the foregoing disclosure. It should be understood that such changes may be made to the foregoing without departing from the scope of the present disclosure. Therefore, it is appropriate to interpret the claims broadly in a manner consistent with the present disclosure.
Claims
1. An ultrasound imaging system, comprising: a processor circuit in communication with the ultrasound imaging device, the processor circuit being configured to: receiving, from the ultrasound imaging device, a temporal sequence of input image frames of a moving object over a time period, wherein the moving object comprises at least one of an anatomy of a patient or a medical device passing through the anatomy of the patient, and wherein a portion of the moving object is at least partially not visible in a first input image frame in the sequence of input image frames; applying a recursive prediction network associated with image segmentation to the sequence of input image frames to generate segmentation data, wherein the recursive prediction network is adapted to predict the motion and / or position of the moving object based on temporal information carried in the sequence of input image frames, and wherein the recursive prediction network comprises a deep learning network adapted to pass a prediction for a current image frame as input for a prediction of a next image frame; and Outputting a sequence of output image frames to a display in communication with the processor circuit based on the segmentation data, wherein the portion of the moving object is fully visible in a first output image frame in the sequence of output image frames, the first output image frame and the first input image frame being associated with a same moment in the time period.
2. The system according to claim 1, wherein: The processor circuit configured to apply the recursive prediction network is further configured to: generating previous segmentation data based on a previous input image frame in the sequence of input image frames, the previous input image frame being received before the first input image frame; and First segmentation data is generated based on the first input image frame and the previous segmentation data.
3. The system of claim 2, wherein: The processor circuit configured to generate the previously segmented data is configured to: applying a convolutional encoder and a recurrent neural network to the previous input image frame; The processor circuit configured to generate the first segmented data is configured to: applying the convolutional encoder to the first input image frame to generate encoded data; and applying the recurrent neural network to the encoded data and the previously segmented data; and The processor circuit configured to apply the recursive prediction network is further configured to: A convolutional decoder is applied to the first segmented data and the previous segmented data.
4. The system according to claim 3, wherein: The convolutional encoder, the recurrent neural network, and the convolutional decoder operate at multiple spatial resolutions.
5. The system according to claim 3, wherein: The moving object includes the medical device passing through the patient's anatomy, and wherein the convolutional encoder, the recurrent neural network, and the convolutional decoder are trained to recognize the medical device from the patient's anatomy and predict motion associated with the medical device passing through the patient's anatomy.
6. The system according to claim 3, wherein: The moving object includes an anatomical structure of the patient having at least one of cardiac motion, respiratory motion, or an arterial pulse, and wherein the convolutional encoder, the recurrent neural network, and the convolutional decoder are trained to identify moving portions of the patient's anatomical structure from static portions of the patient's anatomical structure and predict motion associated with the moving portions.
7. The system according to claim 1, wherein: The moving object comprises the medical device passing through the patient's anatomy, and wherein the system comprises the medical device.
8. The system according to claim 7, wherein: The medical device includes at least one of a needle, a guidewire, a catheter, a guide catheter, a therapeutic device, or an interventional device.
9. The system according to claim 1, wherein: The input image frame includes at least one of a two-dimensional image frame and a three-dimensional image frame.
10. The system according to claim 1, wherein: The processor circuit is further configured to: A spline fit is applied to the sequence of input image frames based on the segmentation data.
11. The system of claim 1 , further comprising the ultrasound imaging device, and wherein: The ultrasound imaging device comprises an ultrasound transducer array configured to obtain the sequence of input image frames.
12. A method for processing an ultrasound image, the method comprising the following steps: receiving a time sequence of image frames representing anatomy of a patient over a time period from an ultrasound imaging device; applying a recursive prediction network associated with image acquisition to the sequence of image frames to generate imaging plane data associated with clinical properties of the patient's anatomical structure, wherein the recursive prediction network is adapted to predict motion and / or position of the patient's anatomical structure based on temporal information carried in the sequence of image frames, and wherein the recursive prediction network comprises a deep learning network adapted to pass a prediction for a current image frame as input for a prediction for a next image frame; and At least one of: a target imaging plane of the patient's anatomy, or instructions for repositioning the ultrasound imaging device toward the target imaging plane is output to a display.
13. The method according to claim 12, wherein: The steps for applying a recursive prediction network include the following: generating first imaging plane data based on a first image frame in the sequence of image frames; and Second imaging plane data is generated based on a second image frame in the sequence of image frames and the first imaging plane data, the second image frame being received after the first image frame.
14. The method according to claim 13, in, The step of generating first imaging plane data includes: applying a convolutional encoder and a recurrent neural network to the first image frame, The step of generating the second imaging plane data includes: applying the convolutional encoder to the first image frame to generate encoded data; and applying the recurrent neural network to the encoded data and the first imaging plane data; And wherein the step of applying the recursive prediction network includes: A convolutional decoder is applied to the first imaging plane data and the second imaging plane data.
15. A non-transitory computer-readable storage medium having stored thereon a computer program, the computer program comprising instructions for configuring a processor circuit to control an ultrasound imaging device, wherein The instructions, when executed by the processor circuit, cause the processor circuit to: receiving a time sequence of image frames representing anatomy of a patient over a time period from the ultrasound imaging device; applying a recursive prediction network associated with image acquisition to the sequence of image frames to generate imaging plane data associated with clinical properties of the patient's anatomical structure, wherein the recursive prediction network is adapted to predict motion and / or position of the patient's anatomical structure based on temporal information carried in the sequence of image frames, and wherein the recursive prediction network comprises a deep learning network adapted to pass a prediction for a current image frame as input for a prediction for a next image frame; and At least one of: a target imaging plane of the patient's anatomy, or instructions for repositioning the ultrasound imaging device toward the target imaging plane is output to a display.
Citation Information
Patent Citations
Systems and methods for navigation to targeted anatomical objects in medical imaging-based procedures
JP2019508072A