Elevational resolution-enhanced 3D ultrasound localization imaging with linear array transducers
By employing spatiotemporal clutter filtering and deep-learning models, the method enhances elevational resolution and shortens data acquisition time in ultrasound imaging, addressing limitations of linear array transducers and enabling high-resolution 3D imaging for real-time clinical use.
Patent Information
- Application Number
- PCT/US2025/013494
- Authority / Receiving Office
- WO · WO
- Patent Type
- Applications
- Current Assignee / Owner
- Priority Date
- 2024-01-31
- Filing Date
- 2025-01-29
- Publication Date
- 2025-08-07
AI Technical Summary
Existing ultrasound imaging techniques, particularly those using linear array transducers, face limitations in elevational resolution and require long data acquisition times, hindering their application in real-time clinical scenarios.
A method involving spatiotemporal clutter filtering and deep-learning models, such as multi-stream GAN and U-Net, is used to enhance elevational resolution and reduce data acquisition time by accurately tracking microbubble trajectories and separating in-plane and out-of-plane signals, allowing for super-elevational resolution ULM imaging.
This approach significantly improves elevational resolution and reduces data acquisition time, enabling high-resolution 3D imaging without additional hardware, facilitating real-time clinical applications.
Smart Images

Figure US2025013494_07082025_PF_FP_ABST
Abstract
Description
Elevational Resolution-Enhanced 3D Ultrasound Localization Imaging with Linear Array TransducersPRIORITY AND RELATED APPLICATION
[0001] This application claims priority to U.S. Provisional Application No.63 / 627,343, filed January 31, 2024, which is incorporated by reference in its entirety. TECHNICAL FIELD
[0002] This disclosure relates to medical ultrasound imaging, particularly to the enhancement of super-resolution ultrasound localization imaging with linear array transducers. BACKGROUND
[0003] Ultrasound imaging is a widely deployed medical imaging in clinics due to its safety, low cost and real-time capability. Furthermore, ultrasound localization microscopy (ULM) has been developed to achieve higher resolution (e.g., axial resolution or lateral resolution). However, there are some issues / problems hindering its usage for a variety of clinical applications. For one example, the elevational resolution associated with a linear array transducer limits the spatial resolution of these imaging techniques. For another example, a long data acquisition time of ULM (e.g., several minutes per frame) hampers its real-time implementation in clinical applications. BRIEF DESCRIPTION OF THE DRAWINGS
[0004] The system, device, product, and / or method described below may be better understood with reference to the following drawings and description of non-limiting and non-exhaustive embodiments. The components in the drawings are not necessarily to scale. Emphasis instead is placed upon illustrating the principles of the present disclosure. The patent or application file contains at least one drawing executed in color.
[0005] FIG. 1A shows an exemplary embodiment for acquiring and processing ultrasound signals.
[0006] FIG. 1B shows a diagram of an exemplary flow for processing ultrasound signals.
[0007] FIG. 2 shows a schematic diagram of another exemplary embodiment for enhancing ultrasound localization microscopy (ULM) image with a deep learning model.
[0008] FIG. 3 shows a flow diagram of an embodiment of a method in the present disclosure.
[0009] FIG. 4 shows a computer system that may be used to implement various components in an apparatus / device or various steps in a method described in the present disclosure.
[0010] FIG. 5 shows an overview of three dimension (3D) dual modal ULM / photoacoustic (PA) imaging.
[0011] FIG.6 shows elevational resolution improvement of ULM.
[0012] FIG.7 shows the quality of dynamic ULM using different thresholds.
[0013] FIG. 8 shows performance of dynamic ULM with different microbubble flow speed.
[0014] FIG.9 shows performance of dynamic ULM in complex vascular structures.
[0015] FIG.10 shows in vivo performance of dynamic ULM.
[0016] FIG.11 shows validation of deep learning model in 2D.
[0017] FIG.12 shows validation of deep learning model in simulated brain datasets.
[0018] FIG.13 shows an exemplary embodiment with deep learning model to enable 3D super-resolution imaging.
[0019] FIG.14 shows that 3D dual-modal 3D PA / ULM reveals molecular and structural information of tumor.
[0020] FIG. 15 shows that 3D dual-modal PA / ULM reveals structural and functional information in brain.
[0021] FIG. 16 shows that 3D dual-modal PA / ULM detects blood-brain barrier disruption.
[0022] FIG. 17 shows that 3D dual-modal PA / ULM reveals structural and anatomy information in hindlimb.
[0023] FIG.18 shows validation of EGFR antibody targeting.
[0024] FIG.19 shows biodistribution study in tumor bearing mice.DETAILED DESCRIPTION
[0025] The disclosed systems, devices, and methods will now be described in detail hereinafter with reference to the accompanied drawings that form a part of the present application and show, by way of illustration, examples of specific embodiments. The described systems and methods may, however, be embodied in a variety of different forms and, therefore, the claimed subject matter covered by this disclosure is intended to be construed as not being limited to any of the embodiments. This disclosure may be embodied as methods, devices, components, or systems. Accordingly, embodiments of the disclosed system and methods may, for example, take the form of hardware, software, firmware or any combination thereof.
[0026] Throughout the specification and claims, terms may have nuanced meanings suggested or implied in context beyond an explicitly stated meaning. Likewise, the phrase “in one embodiment” or “in some embodiments” as used herein does not necessarily refer to the same embodiment and the phrase “in another embodiment” or “in other embodiments” as used herein does not necessarily refer to a different embodiment. It is intended, for example, that claimed subject matter may include combinations of exemplary embodiments in whole or in part. Moreover, the phrase “in one implementation”, “in another implementation”, “in some implementations”, or “in some other implementations” as used herein does not necessarily refer to the same implementation(s) or different implementation(s). It is intended, for example, that claimed subject matter may include combinations of the disclosed features from the implementations in whole or in part.
[0027] In general, terminology may be understood at least in part from usage in context. For example, terms, such as “and”, “or”, or “and / or,” as used herein may include a variety of meanings that may depend at least in part upon the context in which such terms are used. In addition, the term “one or more” or “at least one” as used herein, depending at least in part upon context, may be used to describe any feature, structure, or characteristic in a singular sense or may be used to describe combinations of features, structures or characteristics in a plural sense. Similarly, terms, such as “a”, “an”, or “the”, again, may be understood to convey a singular usage or to convey a plural usage, depending at least in part upon context. In addition, the term “based on” or “determined by” may be understood as not necessarily intended to convey an exclusive set of factors and may, instead, allow for existence of additionalfactors not necessarily expressly described, again, depending at least in part on context.
[0028] The present disclosure describes various embodiments for enhancing super- resolution ultrasound localization imaging (ULM) with linear array transducers.
[0029] Super-resolution ultrasound localization imaging offers a significant resolution improvement from conventional ultrasound imaging. Rather than leaning on tissue reflections, it tracks and pinpoints the positions of the contrast agents introduced into the bloodstream. This novel method, by aggregating the positions of these contrast agents over sequential frames, fabricates images with detail, predominantly of the microvascular architecture. This prowess in capturing microscopic vascular intricacies is poised to advance several medical sectors, including oncology, nephrology, cardiology, and neonatal neurology.
[0030] Linear array ultrasound transducers are currently the most popular ultrasound imaging probes in clinical applications. However, the linear array transducers face limitations: they often produce subpar elevational resolution in 3D imaging reconstruction. This limitation stems from the intrinsic single-element arrangement in the elevational plane and the diffraction limitation of the acoustic lens for elevational focusing. For one example, using linear array transducers for 3D super-resolution often yields inadequate elevational resolution, as pinpointing the exact location of contrast agents within the ultrasound's slice-thickness proves elusive. For another example, a long data acquisition time (typically several minutes per frame) is needed for obtaining 3D super-resolution ULM.
[0031] The present disclosure describes various embodiments for enhancing super- resolution ULM with linear array transducers, addressing at least one of the issues / problems discussed above, improving elevational resolution of 3D super- resolution ULM, and / or decreasing requirement data acquisition time for 3D super- resolution ULM. This refinement in resolution, achieved without the need for specialized or additional hardware, and / or the shortened acquisition time may revolutionize diagnostic and therapeutic strategies across multiple medical disciplines.
[0032] In some embodiments for enhancing the elevational resolution of the 3D super- resolution ULM with transducers, a method may begin with introducing ultrasound contrast agents into the circulatory system. Following a succession of ultrasound captures, a spatiotemporal clutter filter comes into play, effectively sidelining tissue signals to highlight the microbubbles. Subsequently, a blend of contrast agentlocalization and tracking algorithms determines each agent's precise trajectory through the series of refined images. The pivotal innovation lies in optimizing the elevational resolution. By harnessing ultrafast ultrasound, the invention continually tracks the time-evolving scattering amplitude of each agent within the transducer's designated "slice-thickness." Analyzing this backscattering amplitude, which varies depending on the contrast agent's relative position within the focusing slice, allows for the accurate determination of each agent's position relative to the transducer's center, thus enhancing the elevational resolution.
[0033] In some embodiments for enhancing the 3D super-resolution ULM with a deep- leaning approach, a method includes using a multi-conditional deep-learning model to reconstruct a complete super-resolution vasculature map from an incomplete data collection. In addition, the method may also explore the microbubble backscattered signal and enhance elevational resolution by separating in-plane and out-of-plane microbubble information in each 2D ULM. The deep learning model may include at least one of the following: multi-stream generative adversarial networks (GAN) model, multi-stream U-Net, GAN, and / or U-Net.
[0034] Referring now to FIG.1A, an example embodiment (e.g., a device or system) 100 for enhancing the elevational resolution of the 3D super-resolution ULM with transducers is shown. The example embodiment 100 may include a portion or all of the following: a capture subsystem 110 for acquiring signals from a sample 101 and / or a processing subsystem 120 for processing the acquired signals to form images.
[0035] In the example, the capture subsystem includes a transducer 114. The transducer may be a linear array transducer that is disposed on a XYZ motorized linear stage; and acquires ultrasound signals from the sample. The capture system sends raw data corresponding to the ultrasound signal to the processing subsystem for further processing. In some implementations, to acquire photoacoustic (PA) signal, the capture subsystem may include a laser 112, wherein a customized bifurcated fiber bundle may be used to deliver light from the laser source to the sample, and the transducer may acquire a hybrid signal including both PA signal and normal ultrasound signal. In some implementations, the fiber bundle may be integrated to the transducer for the side-illumination.
[0036] The processing subsystem 120 may include a memory 122 and a hardware- based processor 124. The memory 122 may store instructions and raw data from the ultrasound signal from the transducer. The memory may also include intermediate andfinal image data. In some implementations, the memory may store a neural network (or other machine learning protocol) to generate a complete ULM image based on a sparse (partial) ULM image. In some distributed implementations, not shown here, the processing subsystem 120 (or portions thereof) may remotely communicate with the capture subsystem 110 (for example, cloud computing). Accordingly, the processing subsystem 120 may further include networking hardware that may receive raw ultrasound data in a remotely captured and / or partially remotely-pre-processed form. The processor 124 may execute instructions stored on the memory to perform a portion or all of the steps / methods described in the present disclosure.
[0037] Referring to FIG.1B, in some implementations with a PA / ULM hybrid system, hybrid signal acquired by a transducer from a sample may include a PA signal 132 and an ultrasound signal 142. The PA signal may be processed to generate a PA frame 134.
[0038] In some implementations with either a PA / ULM hybrid system or an ULM system, the ultrasound signal 142 may undergo a reconstruction process 144 to form ultrasound images 146, wherein the acquired ultrasound data may be beamformed using a delay and sum algorithm to reconstruct ultrasound images. By applying a singular value decomposition (SVD) based spatiotemporal filtering 148 to the reconstructed ultrasound cineloop, doppler signal over time 150 and flowing microbubble signal 160 may be obtained.
[0039] In some implementations for the obtained doppler signal, a rigid motion model may be used to correct motions of the sample (152), wherein the doppler signal is corrected to compensate for sample motion using translation and rotation information. The power doppler image 154 may be calculated as the power of the doppler signal.
[0040] In some implementations for the obtained microbubble signal, the microbubble may be localized via a localization process 162, wherein a system points spread function (PSF) is estimated using Gaussian fitting of one isolated microbubble signal, and then the cross-correlation between PSF and microbubble signals over frames is calculated to localize microbubbles. In some implementations, motion compensation 164 may be applied to correct tissue motion, which is similar to the motion compensation for the doppler signal.
[0041] In some implementations, after correcting all microbubble locations over frames, moving microbubbles are tracked by a tracking algorithm (166). For example, a Hungarian tracking algorithm may be implemented to track moving microbubbles.The algorithm calculates all distances between all microbubbles in the current frame and all microbubbles in the next frame. Then the algorithm minimizes the total distance and finds the optimal pairs of microbubbles in adjacent frames. Applying the algorithm to all frames can provide the collections of a series of microbubble tracks over time. In some implementations, a super resolution ultrasound image may be obtained by accumulating these positions over frames; and / or velocities of each microbubble may be calculated by differentiating each position in a trajectory according to the time vector to generate a flow velocity map.
[0042] In some implementations, dynamic localization 170 may be applied to each microbubble trajectory and discarded out-of-plane microbubbles to obtain a super- elevational resolution ULM image 172 (or called as elevational-improved ULM image) and / or velocity map image 174. The dynamic localization may include a portion or all of the following: for each microbubble trajectory corresponding to a microbubble: retrieving backscattered amplitudes of the microbubble by identifying intensities associated with microbubble positions in the ultrasound image; determining an amplitude threshold for the microbubble based on the backscattered amplitudes of the microbubble; in response to a backscattered amplitude of the microbubble being below the amplitude threshold, determining microbubble signal corresponding to the backscattered amplitude of the microbubble as out-of-plan microbubble signal; and / or in response to a backscattered amplitude of the microbubble being above or equal to the amplitude threshold, determining microbubble signal corresponding to the backscattered amplitude of the microbubble as in-plan microbubble signal. In some implementations, the dynamic localization may further include a portion or all of the following: fitting the microbubble trajectory with a Gaussian curve; and recovering the microbubble position for the microbubble within the microbubble trajectory in a three- dimensional (3D) domain based on the Gaussian curve. By setting the threshold to each trajectory, unwanted out-of-plane signals can be removed and thus elevational resolution can be achieved.
[0043] Referring now to FIG.2, an example embodiment for enhancing a 3D super- resolution ULM includes a deep-leaning model (neural network). The deep-learning model may include a generator 230, which is a pre-trained neural network model. In some implementations, the generator may include a U-net, which takes a sparse ULM image as an input, wherein, for example, the sparse ULM image corresponds toultrasound signal acquired within a short time duration (e.g., 0.2 or 1 second). The generator may output a generated ULM image 235.
[0044] In some implementations, the generator may include a two-stream U-net, which has two inputs: a first input is the sparse ULM image 210 and a second input is a power doppler image, wherein the sparse ULM image and the power doppler image correspond to ultrasound signal acquired within a short time duration (e.g., 0.2 or 1 second). The generator may output a generated ULM image 235. An example of the generator’s architecture is described in later paragraphs of the present disclosure.
[0045] In some implementations, the generator 230 may be pre-trained with a discriminator 240 and a loss function 250 based on a plurality of training sample sets. When the generator is the two-stream U-net, each training sample set may include a sparse (partial) ULM image, a power doppler image (partial), and a dense (complete) ULM image 245. The sparse (partial) ULM image and the power doppler image (partial) correspond to ultrasound signal of a short time period of the total acquisition duration, for non-limiting example, the total acquisition duration may be 8 or 12 second and the short time period may be 0.2 or 1 second. The dense (complete) ULM image corresponds to ultrasound signal of the total (or substantially total) acquisition duration.
[0046] During training, a generated ULM image from the generator and a corresponding dense (complete) ULM image are input into the discriminator 240, and a loss function value is calculated based on the loss function 250. Based on the loss function value, the generator and / or the discriminator is trained to optimize the loss function with stochastic gradient descent.
[0047] After training of the neural network model is completed (e.g., when the loss function value or its change between iterations satisfies a pre-defined condition), testing sample sets may be generated from either a new sparse ULM image or acquired from a new sample to test the neural network model’s generality. The testing sample sets are fed into the trained neural network, and a loss function is calculated based on the generated ULM image with its corresponding complete ULM image.
[0048] Referring to FIG.3, the present disclosure describes various embodiments of a method 300 for obtaining a super-elevational resolution ultrasound localization microscopy (ULM) image. The method 300 may be performed by an electronic device (e.g., a computer as shown in FIG. 4) comprising a memory and a processor in communication with the memory, wherein the memory stores instructions, and the processor executes the instructions to perform a portion or all the steps included in themethod. The method 300 may include a portion or all of the following steps: step 310, obtaining ultrasound signal acquired from a sample containing microbubbles by a transducer; step 320, reconstructing the ultrasound signal to form ultrasound images; step 330, filtering the ultrasound images to obtain microbubble signal with a filter; step 340, localizing and tracking the microbubbles based on the microbubble signal to obtain microbubble trajectories; step 350, performing dynamic localization by selecting in-plane microbubble signal and eliminating out-of-plane microbubble signal based on backscattered amplitudes of the microbubbles; and / or step 360, generating a super- elevational resolution ULM image based on microbubble positions corresponding to the in-plane microbubble signal.
[0049] In some implementations, the transducer comprises a linear array transducer.
[0050] In some implementations, the filter comprises a singular value decomposition (SVD) based spatiotemporal clutter filter.
[0051] In some implementations, the localizing and tracking the microbubbles based on the microbubble signal to obtain the microbubble trajectories comprises applying a localization algorithm based on the microbubble signal to obtain positions of the microbubbles; and / or tracking the microbubbles with a Hungarian tracking algorithm based on the positions of the microbubbles to obtain the microbubble trajectories.
[0052] In some implementations, the localization algorithm comprises a point spread function (PSF) cross correlation localization algorithm; and / or the tracking algorithm comprises a Hungarian tracking algorithm.
[0053] In some implementations, the performing the dynamic localization comprises a portion or all of the following: for each microbubble trajectory corresponding to a microbubble: retrieving backscattered amplitudes of the microbubble by identifying intensities associated with microbubble positions in the ultrasound image; determining an amplitude threshold for the microbubble based on the backscattered amplitudes of the microbubble; in response to a backscattered amplitude of the microbubble being below the amplitude threshold, determining microbubble signal corresponding to the backscattered amplitude of the microbubble as out-of-plane microbubble signal; and in response to a backscattered amplitude of the microbubble being above or equal to the amplitude threshold, determining microbubble signal corresponding to the backscattered amplitude of the microbubble as in-plane microbubble signal.
[0054] In some implementations, the performing the dynamic localization further comprises fitting the microbubble trajectory with a Gaussian curve; and / or recoveringthe microbubble position for the microbubble within the microbubble trajectory in a three-dimensional (3D) domain based on the Gaussian curve.
[0055] In some implementations, the method further comprises: generating a velocity image based on the microbubble positions corresponding to the in-plane microbubble signal; and / or generating a 3D super-elevational resolution ULM image based on a plurality of super-elevational resolution ULM images.
[0056] In some implementations, the method may further include inputting the super- elevational resolution ULM image into a neural network for training to generate a complete super-elevational resolution ULM image.
[0057] In some implementations, the pre-trained neural network comprises a generative adversarial network (GAN), and / or the GAN comprises a generator neural network and a discriminator neural network.
[0058] In some implementations, the method may further include generating a power doppler image based on the ultrasound signal; and / or inputting the power doppler image into the pre-trained neural network for generating the complete super- elevational resolution ULM image, wherein the pre-trained neural network comprises a multi-stream GAN.
[0059] Referring to FIG. 4, a computer system (electronic device) 400 may include communication interfaces 402, system circuitry 404, input / output (I / O) interfaces 406, storage 409, and display circuitry 408 that generates machine interfaces 410 locally or for remote display, e.g., in a web browser running on a local or remote machine. The machine interfaces 410 and the I / O interfaces 406 may include GUIs, touch sensitive displays, voice or facial recognition inputs, buttons, switches, speakers and other user interface elements.
[0060] The machine interfaces 410 and the I / O interfaces 406 may further include communication interfaces with modulators, sensors, and / or detectors. The communication between the computer system 400 and the sensors and detector may include wired communication or wireless communication. The communication may include but not limited to, a serial communication, a parallel communication; an Ethernet communication, a USB communication, and a general purpose interface bus (GPIB) communication. Additional examples of the I / O interfaces 406 include microphones, video and still image cameras, headset and microphone input / output jacks, Universal Serial Bus (USB) connectors, memory card slots, and other types of inputs. The I / O interfaces 406 may further include magnetic or optical media interfaces(e.g., a CDROM or DVD drive), serial and parallel bus interfaces, and keyboard and mouse interfaces.
[0061] The communication interfaces 402 may include wireless transmitters and receivers ("transceivers") 412 and any antennas 414 used by the transmitting and receiving circuitry of the transceivers 412. The transceivers 412 and antennas 414 may support Wi-Fi network communications, for instance, under any version of IEEE 802.11, e.g., 802.11n or 802.11ac. The communication interfaces 402 may also include wireline transceivers 416. The wireline transceivers 416 may provide physical layer interfaces for any of a wide range of communication protocols, such as any type of Ethernet, data over cable service interface specification (DOCSIS), digital subscriber line (DSL), Synchronous Optical Network (SONET), or other protocol. In another implementation, the communication interfaces 402 may further include communication interfaces with the modulators, sensors, transducers, and / or detectors.
[0062] The storage 409 may be used to store various initial, intermediate, or final data. In one implementation, the storage 409 of the computer system 400 may be integral with a database server. The storage 409 may be centralized or distributed, and may be local or remote to the computer system 400. For example, the storage 409 may be hosted remotely by a cloud computing service provider.
[0063] The system circuitry 404 may include hardware, software, firmware, or other circuitry in any combination. The system circuitry 404 may be implemented, for example, with one or more systems on a chip (SoC), application specific integrated circuits (ASIC), microprocessors, discrete analog and digital circuits, and other circuitry. For example, the system circuitry 404 may include one or more instruction processors 421 and memories 422. The memories 422 store, for example, control instructions 426 and an operating system 424. In one implementation, the instruction processors 421 execute the control instructions 426 and the operating system 424 to carry out any desired functionality related to the controller.
[0064] The electronic device in the present disclosure and / or the methods described in the present disclosure may be implemented by a portion or all of the computer system 400 as described above.
[0065] The present disclosure describes various non-limiting embodiments for enhancing ULM. The embodiments and / or example implementations below are intended to be illustrative embodiments and / or examples of the techniques andarchitectures discussed above. The example implementations are not intended to constrain the above techniques and architectures to particular features and / or examples but rather demonstrate real world implementations of the above techniques and architectures. Further, the features discussed in conjunction with the various example implementations below may be individually (or in virtually any grouping) incorporated into various implementations of the techniques and architectures discussed above with or without others of the features present in the various example implementations below. Embodiment: 3D Dual Modal Ultrasound Localization Microscopy and Photoacoustic Imaging Using a Linear Array Transducer
[0066] Noninvasive biomedical imaging technologies have become an integral part of medical diagnostics. Each imaging technology has its strengths and weaknesses; while certain technologies primarily monitor the detailed anatomical abnormalities, others focus on detecting changes in physiological signs or disease-associated biomarkers. With the recent advancement in biology, many medical conditions, such as cancer, are complex and cannot be characterized by a single or a few hallmarks. To date, no particular diagnostic imaging technique can reveal all the hallmarks. To overcome the limitation of a particular imaging technique, since the 1990s, the possibility of combining more than one imaging technology for offering complementary information has been actively investigated. In the late 1990s, the first positron- emission tomography (PET) / computed tomography (CT) was realized, which revolutionized diagnostic imaging. Recently, new hybrid imaging modalities such as magnetic resonance imaging (MRI) / PET, fluorescent / MRI, and fluorescent / CT have become a new frontier of imaging research. The technical challenge of MRI / PET includes placing nuclear medicine detectors inside the MR magnet, which drastically increases technical complexity. Other hybrid imaging such as fluorescent / MRI and fluorescent / CT imaging require aligning different detectors of each imaging and co- register the images acquired from different detectors that could differ significantly in the resolution and imaging depth.
[0067] Ultrasound imaging is the most widely deployed medical imaging in clinics due to its safety, low cost and real-time capability. Photoacoustic tomography, a variant of ultrasound imaging, produces ultrasound signals using ultrashort laser pulses through optical absorption of imaging targets. This technique can provide images of thevasculature, hemoglobin concentration, oxygen saturation of tissue, metabolic rate of oxygen, and melanin concentration by measuring the absorption of light from native chromophores. It is also possible to image the distribution of molecular targets and gene expression using optically absorbing targeted contrast agents. Because photoacoustic and ultrasound imaging can share the same ultrasound detector and data-acquisition system, the photoacoustic images can be seamlessly fused with conventional ultrasound images to provide hybrid images of function (molecular markers) / structure.
[0068] Photoacoustic / ultrasound imaging has been explored for a variety of clinical applications, including breast imaging, melanoma detection, sentinel lymph node analysis, endoscopic applications, tumor monitoring and even brain imaging. As powerful as this imaging combination is, however, the acoustic diffraction limit bounds the spatial resolution of both imaging techniques. Further enhancing the resolution by increasing the ultrasound frequency will significantly decrease the imaging depth due to frequency-dependent ultrasound attenuation. Thus, detailed structures below hundreds of microns are difficult to image, which sometimes misses the critical hallmarks for diagnosis, for instance, tumor microvasculature. In addition, while photoacoustic imaging can produce high contrast vasculature imaging without exogenous labeling, a complex imaging setup such as full view tomography and customized designed ultrasound detectors are typically needed. The artifacts due to limited view and limited bandwidth significantly reduce the imaging quality when using the clinical ultrasound transducers, such as linear array transducers. It is still technically challenging to produce artifact free photoacoustic vascular image with a linear array transducer.
[0069] Ultrasound localization microscopy (ULM) may be used to break the acoustic diffraction limit. ULM detects the centroids of isolated microbubbles flowing in the bloodstreams. The sparse points of the microbubble positions from multiple ultrasound frames are then collected to construct a super-resolution vascular image. Because the microbubble localization is free from the diffraction limit as long as the microbubbles can be isolated from each other, ULM can enhance the spatial resolution of traditional ultrasound imaging by up-to ten-fold, paving a new way for detailed vasculature images in deep tissue. ULM can also produce a precise velocity map of the bloodstreams by tracing the microbubble flows. However, some information offered by photoacoustic imaging, such as oxygen saturation of tissue and molecular targetsexpression, is currently missing in ULM. The hybrid of photoacoustic imaging and ULM (PA / ULM), thus, can provide complementary information for many clinical applications. For instance, angiogenesis is a well-known cancer hallmark, hypoxia has been recognized as one of the fundamentally important features of solid tumors, and the expression of many cancer-associated biomarkers is used for diagnosis or prognosis of cancers. The hybrid PA / ULM in combination with imaging agents can potentially provide all this information for cancer diagnosis. Because of the immediate implementation of ultrasound system, ULM can be directly used in PA / US system for 2D imaging.
[0070] However, the long data acquisition time of ULM (typically several minutes per frame) hampers the real-time implementation of dual PA / ULM, where the frame rate of PA is in the range of several Hz to hundreds of Hz. The major reason of the slow acquisition of ULM is the injection of low concentrated microbubbles. Typically, even with the recently developed ultrafast ultrasound imaging technique that produces 500- 1000 images per second, data acquisition still requires minutes to create one frame of the 2D super-resolved vascular image using a linear array transducer. To overcome this issue, the first-of-its-kind dual PA / fast ULM imaging system accelerated by sparsity-constraint optimization has been proposed, successfully demonstrating the applicability and feasibility of the dual modality in kidney and hindlimb.
[0071] Due to the use of high concentrated microbubbles, it may be difficult for the dual imaging system to capture large-view 3D imaging with mechanical scanning (i.e. whole tumor of a mouse). the maximum imaging duration of ULM will be limited by the maximum allowable dosage of microbubbles. While increasing the injection concentration speeds up the frame rate, it adversely reduces the maximum imaging volume and makes it less favorable for 3D imaging. In addition, the out-of-plane information in each 2D ultrasound frame, due to the elevational geometry of the linear transducer, hinders the attainment of super-resolution in the elevational direction through mechanical scanning-based 3D ULM or stacks of 2D frames. One feasible approach for achieving isotropic resolution in 3D ULM involves the utilization of a 2D array. Although the cases of 3D acquisitions based on matrix probes are growing, 2D arrays are very expensive, require complex electronics and have relatively small imaging view, while 1D arrays are widely used in both preclinical and clinical applications.
[0072] To solve these technical difficulties, the present disclosure describes a new deep-learning approach without the requirement of high concentration microbubbles to enable large-view 3D imaging. Very different from previous approaches, instead of differentiating overlapping high-density microbubbles, various embodiments in the present disclosure may reconstruct a complete super-resolution vasculature map from an incomplete data collection through multi-conditional deep-learning modeling. In addition, various embodiments explore the microbubble backscattered signal and enhance elevational resolution by separating in-plane and out-of-plane microbubble information in each 2D ULM. Enabled with deep learning, various embodiments achieve the development of 3D PA / ULM and demonstrate that this new hybrid imaging can be used to reveal tumor angiogenesis, tumor hypoxia, and epidermal growth factor receptor expression on a mouse skin cancer model. It can also produce multifunctional images on important organs, such as the brain and lymphatic nodes. Dual modal imaging system
[0073] The dual modal PA / ULM releases the hardware burden to build dual imaging system, and dual modal photoacoustic / ultrasound (PA / US) system is enough to produce dual PA / ULM images with the customization of acquisition imaging sequence. In various embodiments, Verasonics ultrasound research system and OPOTEK laser system with the repetition rate of 10 Hz is used is to perform US and PA imaging. MS200 transducer (central frequency 15 MHz) is combined with bifurcated fiber bundle using a customized adapter. The transducer is mounted on the XYZ linear stage for mechanical scanning to collect 3D data. One dual modal PA / ULM imaging sequence is shown in part B of FIG.5. Data acquisition for single dual PA / ULM frame includes repeated sections of multiwavelength PA (MwPA) data acquisition and ultrafast US data acquisition. The acquisition time of single dual PA / ULM is defined as the total acquisition time to generate one functional PA frame and one ULM frame. The frame rate of US imaging is set to 500 Hz and the frame rate of PA imaging is 10 Hz. To optimize the acquisition time of dual imaging, PA imaging sequence may be interleaved with US imaging sequence. The time interval between two adjacent frames (either PA / US or US / US) is 2 ms (or 500 Hz). In this case, the acquisition time of dual PA / ULM is determined by the acquisition time of ULM.
[0074] FIG. 5 shows an overview of 3D dual modal ULM / PA imaging. (a) The representative 3D tumor vasculature and oxygenation image, generated by stacking aseries of 2D ULM / PA images (indicated by green patches) based on mechanical scanning. The scanning is along y axis. (b) during the data acquisition, at each imaging position, interleaved acquisition of multiwavelength PA (MwPA) and plane-wave US is performed. The frame rate of US imaging is 500 Hz and the acquisition time of US between two neighbor PA sequence is 100 msec. The frame rate of dual imaging is determined by the total acquisition of US imaging or ULM imaging. (c) The collected 2D MwPA images and ultrasound images at each position are processed with linear spectrum unmixing and multi-stream deep learning separately to generate molecular / functional PA image and super-resolution ultrasound vascular image. The detailed processing of ULM images are shown in (d)-(f). (d) The schematic of dynamic localization for elevational resolution improvement. During a series of time (i.e. t1-t4), microbubbles are localized and tracked after SVD filtering of ultrasound images. Based on the curve of the backscattered amplitude of microbubbles as function of acquisition time in each trajectory, threshold (red-line) is applied to remove out-of- plane microbubble signals. The comparison of ULM images with or without dynamic localization is shown in (e). (f) The processing of multi-stream deep learning model. The model includes a two-stream U-net generator and discriminator. During the training, SVD filtering is applied to retrieve microbubble / blood signals from a series of 2D ultrasound images. sparse ULM and power Doppler images are generated by dynamic localization and power summation separately and are fed into the network. The deep learning output is compared to the ground truth which is complete ULM images via two loss functions: super resolution loss and cGAN discriminator loss. The super resolution loss measures the difference between the ground truth and the output of the network using L1 norm loss and MS-SSIM loss. The cGAN discriminator loss evaluates the probability that the output of the generator can be classified as the real image.
[0075] The pipeline to generate ULM and functional PA images is shown in panel c of FIG.5. The raw data is first reconstructed using Verasonics’ beamforming to generate a series of US and MwPA images slice by slice. To extract functional information from PA images, linear regression-based spectrum unmixing is used to extract the signal of either tissue chromophores or contrast agents based on the prior absorption spectrum of these photoacoustic contrast agents. For ULM processing, the SVD- based spatiotemporal clutter filtering is first used to extract the microbubble signal. Cross-correlation based localization algorithm is used to identify microbubblelocalizations. After that, the trained GAN model provides the full microvascular structure by feeding the input of sparse microbubble positions and a power Doppler image. The volume ULM and functional PA image is the stack of 2D slices of ULM and functional images with mechanical scanning. Enhancing elevational resolution of 3D ULM imaging using backscattered microbubble signals
[0076] It may be demonstrated that ULM may enhance spatial resolution in the axial and lateral directions by utilizing a linear array. However, the presence of out-of-plane information in each 2D ultrasound frame, resulting from elevational resolution, hinders the attainment of super-resolution in the elevational direction through mechanical scanning-based 3D ULM or stacks of 2D frames. To address the challenge of enhancing elevational resolution in a linear array setup, various embodiments may leverage the backscattering information from microbubbles.
[0077] The backscattered energy emitted by individual microbubbles is solely determined by the amplitude of the pressure wave that interacts with them. Therefore, the backscattered amplitude directly correlates with the position of each microbubble within the transmit pressure field. In addition, the backscattering amplitude of these microbubbles varies depending on their position (whether in-plane or out-of-plane) and microbubbles have higher amplitude when they are in-plane than that of out-of-plane. Leveraging a microbubble tracking algorithm (such as Hungarian tracking algorithm), we can effectively trace the trajectory of individual microbubbles and employ the disparity in backscattering to distinguish between in-plane and out-of-plane microbubbles in 2D frames. One example trajectory of microbubbles is shown in panel d of FIG.6, wherein two microbubbles (indicated by blue square and blue circle) are localized in imaging plane. The amplitude of both microbubbles’ changes with different time, meaning the distance between microbubble and imaging plane varies at different time. When the amplitude is plotted as the function of travelling time (panel e of FIG. 6), it can be seen that the amplitude follows Gaussian distribution, which means microbubbles flow from out-of-plane to in-plane and then go back to out-of-plane and the microbubble signal with peak backscattered amplitude means the microbubble has the closest distance with the imaging plane during the travelling. Therefore, backscattered signal of microbubbles in a trajectory may be used to map ultrasound elevational response. Considering the width of elevational response corresponds tothe elevational resolution of ultrasound system, the microbubbles with the amplitude of FWHM of backscattered curve has the elevation location ± half of elevation resolution. By setting the threshold to each trajectory, unwanted out-of-plane signals may be removed and microbubble locations within the desired elevational resolution may be obtained. By accumulating these microbubble locations together, elevational- improved ULM can be obtained.
[0078] Panel a of FIG.6 shows the simulated beam profile along elevational and depth direction (yz plane). The central frequency of the ultrasound wave is 15 MHz and the elevational focus is at 15 mm depth. The FWHM of the profile at 15 mm is 0.46 mm, as shown in panel b of FIG. 6, which is the elevational resolution of the array. Compared to the lateral resolution (around single wavelength), the elevational resolution is much poorer. To evaluate the 3D resolution of ULM, two parallel tubes in xy plane and with 45oto xz plane is simulated. The diameter of the tube is 10 micrometer (um). The distance between two tubes are 85 um. Microbubbles flow in the tubes at the speed of 10 mm / s.3D ultrasound images are recorded by scanning along y direction with the step size of 20 um. The frame rate of ultrasound imaging is 500 Hz.
[0079] Microbubble signals in one xz plane in different time is shown in panel d of FIG. 6, where two microbubbles flow over xz plane, localized by circle and square points. When plotting the amplitude of microbubbles as the function of travelling time as shown in panel e of FIG. 6, it can be seen that the amplitude follows Gaussian distribution, which means microbubbles flow from out-of-plane to in-plane and then go back to out-of-plane and the microbubble signal with peak backscattered amplitude means the microbubble has the closest distance with the imaging plane during the travelling.
[0080] The reconstructed MIP doppler image is shown in panel f of FIG.6, where two tubes are overlapped in xy plane. Panel g of FIG. 6 shows the tubes are also undistinguishable in xz plane indicated by red arrow. Even though microbubble localization is used to capture the microbubble positions, the reconstructed tube image is still blur and the tube profiles or the cross-section of two tubes can not be revealed due to the exist of out-of-plane noise, ss shown in panels f and g of FIG.6. It can be appreciated that the resolution improvement of ULM in xz plane compared to that of Doppler is shown in panel g of FIG.6. Using dynamic localization and set threshold toxx, dynamic ULM in xy and xz plane can be obtained. Compared to ULM, dynamic ULM can identify two tube profiles.
[0081] FIG. 6 shows elevational resolution improvement of ULM. (a) Pressure distribution in elevational and depth direction. (b) Beam profile at 15 mm imaging depth. (c) The simulation sketch of two parallel tubes with 45oto xz plane; (d) ultrasound images of one scanning slice at different acquisition time. The position of the xy imaging plane is indicated by red arrows in (g). Two microbubbles are identified and indicated by circle and square points. (e) Microbubble backscattered amplitude as the function of time. (f) MIP images of Doppler, ULM and dynamic ULM in xy plane. The threshold for dynamic ULM is 0.96. (g) Doppler, ULM and dynamic ULM frame at the scanning slice indicated by red arrows in (f).
[0082] The present disclosure describes the reconstructed ULM images using different thresholds. The scanning step is 20 um and the images are reconstructed using 12600 ultrasound frames. Described in other paragraphs in the present disclosure, thethreshold may be linked to the desired elevational resolution: the threshold is0.5^^బ / ^భ^మ, where a0 is the desired elevational resolution and a1 is the originalelevational resolution, which can be measured by FWHM of the beam profile. From this relationship, when there is a higher threshold, higher elevational resolution may be obtained. Because the 3D ULM image is generated by mechanical scanning, the elevational resolution is limited by the step size. When the threshold is higher than the step size, reconstructed images could suffer from under-sampling.
[0083] FIG.7 shows that (a) MIP images of dynamic ULM in xy plane with different separation range from 10 to 300 um. (b) The quantitative analysis of saturation and localization error as function of separation threshold when the transducer scanning step size is 50 um.
[0084] The flow speed of microbubbles on the effect of dynamic ULM may be investigated. The present disclosure describes one backscattered profile generated by microbubble velocity at 40 mm / s. Because of faster speed, fewer locations are captured in each trajectory. Because the threshold is based on the Gaussian fitted curve, to get Gaussian fitted backscattered amplitude curve, at least three microbubble signals in each trajectory may be needed. When it is assumed that the elevational beam width is ^^ um, and the imaging frame rate is ^^ Hz, maximum elevational velocityis௪ிଷ um / s. In addition, the localization error under different velocity may be compared. The results show that the higher speed could introduce more localization error.
[0085] FIG.8 shows that (a) MIP images of dynamic ULM in xy plane with different microbubble flow speed. The separation range is 50 um and the scanning step size is 50 um. (b) The comparison of localization error in x,y,z dimension for different flow speed.
[0086] The performance of dynamic localization method may be evaluated in a simulated complex vascular structure. FIG. 9 shows that (a) The positions of microbubbles reconstructed using regular localization methods are shown as blue circles, while the actual microbubble positions are represented by yellow stars. The imaging plane is indicated by gray window The regular localization method fails to recover elevational (z-axis) information. (b) The imaging plane and actual microbubble positions are projected onto the xy, xz, and yz planes for visualization.(c) The backscattered amplitude of a microbubble is plotted as a function of time, with the smoothed amplitude represented by the red curve.(d) The recovered microbubble positions in the xz and yz planes, obtained using dynamic localization, demonstrate the improved accuracy in the elevational dimension.(e) A 3D visualization of the microbubble positions reconstructed with dynamic localization highlights the method's ability to resolve spatial positions in all three dimensions.(f) The localization error for dynamic localization is quantified in the x, y, and z dimensions, demonstrating its accuracy and reliability.
[0087] The in vivo performance of dynamic ULM may be evaluated with a mouse brain as target. To minimize the acquisition time of 3D ULM, the brain may be scanned with the step size of 100 um. Each position is scanned with 10 sec of data acquisition. By stacking 2D frames together, 3D Doppler and 3D ULM may be generated, as shown in panel a of FIG. 10. These MIP images in xz plane are encoded with scanning distance. Both ULM and dynamic ULM have sharper and finer vasculature than Doppler. The resolution of ULM is better than that of Doppler. To evaluate the elevational resolution, MIP images of ULM, dynamic ULM and 3D Doppler are compared with 1 mm thickness, as shown in panel b of FIG. 10. Although ULM improves resolution in xz plane, it still suffers from low elevation resolution (blue line in panel c of FIG. 10). However, dynamic ULM improves resolution in xy plane, as indicated by the line profile in panel c of FIG.10.
[0088] FIG.10 shows performance of dynamic ULM in vivo. (a) The comparison of in vivo 3D Doppler, ULM and dynamic ULM images of a tumor in xz plane. The acquisition time is 10 sec per frame. (b) The comparison of Doppler, ULM, dynamic ULM and deep learning generated ULM in yz plane (thickness is 1 mm). (c) The resolution comparison of Doppler, ULM, dynamic ULM and deep learning generated ULM. Interpolated profiles are plotted along the lines marked in (b). Multi-stream model massively speeds up 3D super-resolution imaging
[0089] Although dynamic ULM could improve elevational resolution, due to the restricted circulation time of microbubbles, it is not feasible to generate large-view super-resolution ultrasound images using a small step size (i.e.50 μm or less), such as whole brain or whole tumor. Whether DL model can speed up the acquisition while keeping the high elevational resolution is investigated.
[0090] The 2D performance of trained multi-stream model may be first validated using in vivo brain data. A power Doppler image is generated from a set of ultrasound images acquired within 4 seconds, as shown in panel a of FIG.12. While the power Doppler provides an overview of brain vasculature, it fails to capture details of small vessels. Using the microbubble-localization algorithm, the sparse ULM may be generated with the same data acquisition time, as shown in panel b of FIG.12. Although the sparse ULM depicts finer vasculature, it suffers from incomplete vessel reconstruction. By feeding the sparse ULM and the power Doppler into the multi-stream model, one can appreciate the completeness and fineness of vasculature in DL-ULM, as shown in panel d of FIG.11. Compared to the complete ULM acquired over 100 seconds, as shown in panel c of FIG. 11, most of the micro-vessels in the DL-ULM were in accordance with ULM, as confirmed by the merged images in panel g of FIG.11. When zooming into the region indicated by the white box, a gradual improvement in image quality can be seen with increasing acquisition time (from 1 seconds to 6 seconds) for both DL-ULM (A+LR) and ULM, as shown in panel h of FIG. 11. The results demonstrate that the standard ULM only reconstructs sparse and hardly discernible vascular structures even within 6 seconds of acquisition, one the other hand, the DL- ULM image achieves almost complete reconstruction of the main vascular structures within 4 seconds.
[0091] FIG.11 shows validation of a deep learning model. (a) 2D power Doppler image of a mouse brain. (b) 2D sparse ULM image with 4 sec of data acquisition. (c) 2Ddense ULM image with 100 sec of data acquisition. (d) 2D DL-ULM image reconstructed from the 2D power Doppler image (a) and 2D sparse ULM image (b). (e) 2D DL-ULM image reconstructed from the 2D power Doppler image (a) only. (f) 2D DL-ULM image reconstructed from the 2D sparse ULM image (b) only. (g) Merged zoom-in images comparing DL-ULM reconstructions from different inputs (d), (e), (f) with 2D dense ULM image (c). DL-ULM images are shown in green and dense ULM are shown in red. (h) Gradual improvement of image quality of DL-ULM and ULM in zoom-in region in (b) (d) (e) (f) with increasing acquisition time.
[0092] The GAN model may be trained / tested with different input (i.e. only power Doppler (namely LR) input, only sparse ULM (namely A) input). The results are as shown in panels e-f of FIG.11. Only using power Doppler, although DL-ULM gains better quality than power Doppler, some detailed vessel structure is not recognizable compared to the complete ULM, as indicated in zoom-region of panels e and g of FIG. 11. With only sparse ULM information, due to the unbalanced sparse of different vasculature, the model is hard to depict the overall structure, but only complete the partial region, as indicated in zoom-region of panels f and g of FIG.11. Although with only one input (either power Doppler or sparse ULM), the deep learning results are worse than the results from two inputs, the zoom-in regions in panel h of FIG.11 show that both DL-ULM with LR only and with A-only images have faster reconstruction than standard ULM within 6 seconds.
[0093] Additionally, the deep learning model with LR and A inputs may be evaluated with simulated brain datasets. FIG.12 shows the performance of deep learning model in simulated brain datasets. It demonstrates that (a) The ground truth (GT) coronal image of brain vasculature with enhanced elevational resolution. (b) DL-ULM coronal image demonstrates high similarity to the GT, achieving an MS-SSIM score of 0.98. (c) A 3D brain vasculature reconstruction in top view, created by stacking the elevational enhanced GT images along the scanning direction.(d) A 3D brain vasculature reconstruction in top view, created by stacking the DL-ULM images along the scanning direction, showing strong agreement with the GT. (e–h) An example comparison of 2D vascular images in top view: (e) Power Doppler, (f) Sparse ULM, (g) GT, and (h) DL-ULM, highlighting the DL-ULM's ability to recover vascular structures with high fidelity.
[0094] FIG.13 shows the performance of a deep learning model. (a) Reconstruction quality of ULM (similarity, PSNR and L1 norm) and DL-ULM images as a function ofacquisition time. (n=6). (b) Evolution of the deep learning loss with the number of epochs during the training. (c) Evolution of the l1 norm loss with the number of epochs during the training. (d) Evolution of the similarity loss with the number of epochs during the training. To quantitatively evaluate the reconstruction performance of the deep learning models, a series of metrics (i.e. MS-SSIM, PSNR and mean distance) may be used as the indicator. As shown in panel a of FIG.13, a function of MS-SSIM of the deep learning output and the complete ULM (200 sec data acquisition) may be plotted against different acquisition times. All three cases (i.e. sparse ULM input, power Doppler input, sparse ULM and power Doppler input) gain better similarity with the increasing acquisition time. However, not-surprisingly, it can be seen that DL-ULM with both inputs has faster speed to gain the same similarity as the other case. This is because the model learns overall vessel structure from doppler and fine vessel structure from sparse ULM, and as a result it depicts the super-resolution vessel structure faster than the other cases. One notable thing is that when only inputting LR, DL-ULM gets saturated much faster and only can achieve peak similarity at around 0.8. The comparison of PSNR shows that increasing acquisition time can improve PSNR of both DL-ULM and standard ULM. After 12 sec, PSNR of DL-ULM with two inputs and with A input is almost overlapped. While PNSR of DL-ULM with LR input slightly decreased after 15 sec. The curve of mean distance as a function of acquisition time shows that the difference between DL-ULM and GT decreases with increasing acquisition time. Similar to the results of MS-SSIM and PNSR, DL-ULM with two inputs has the lowest difference.
[0095] Thanks to the acceleration of multi-stream model, 3D tumor vasculature visualization may be enabled with fine scanning step size. The whole tumor area on mice may be scanned with the step size of 50 ^^m and total 240 slices were collected. To ensure each frame has similar image quality, a consistent low microbubble density is maintained by administering the microbubble solution via a syringe pump. After the DAQ, the images may be processed using multi-stream model to generate ULM images.
[0096] Panel a of FIG.14 shows the 3D power Doppler and 3D DL-ULM image coded with scanning distance. These two 3D images are generated at 1 second of DAQ per slice (or 240 seconds of total 3D acquisition). Compared with Doppler, which blurs the small blood vessels, 3D DL-ULM reveals more detailed vasculatures with much higher resolution. Panel d of FIG.14 shows the zoom-in 3D microvessel images of a selectedarea (the white box) in panel a at increasing acquisition time from 0.2 to 1.8 seconds per slice. Standard ULM generated with the same acquisition time is also shown as a comparison. How quickly the DL-ULM achieves complete vasculature reconstruction can be observed, within only 1 sec per slice. As a comparison, ULM at the same region still reveals sparse vessels even with 2 sec acquisition per slice. Furthermore, to quantitatively evaluate the 3D acceleration, the similarity curve of 2D DL-ULM and ULM is plotted over time, as presented in panel e of FIG.14. The results demonstrate that DL-ULM enhances the acquisition speed by up to 16 times per slice. Panel b of FIG.14 shows 3D Doppler image and 3D DL-ULM image encoded by imaging depth. The zoom-in micro-vessels reveal that 3D DL-ULM can capture vasculature as thin as 55 um along elevational direction and 25 um along lateral direction. The elevational resolution is limited by 50 um due to the step size, which may be further improved with even smaller step size. To evaluate the accuracy of DL-ULM, the tumor was scanned along y direction and x direction separately.3D-ULM and 3D DL-ULM with y direction scanning was shown in the yz plane in panel f and g of FIG.14. Due to poor elevational resolution in y direction, the vasculatures are blurred in panel f of FIG.14, while panel g of FIG.14 demonstrates detailed vasculatures due to the elevational enhancement with deep learning. Panel h of FIG.14 also demonstrates super-resolved vasculature in yz plane because yz plane is the imaging plane and elevational direction is x dimension. Compared to the panel h, the panel g shows similar vascular structures, validating the accuracy of deep learning model.
[0097] FIG.14 shows that deep learning enables 3D super-resolution imaging. (a) and (b): The comparison of 3D power Doppler and 3D DL-ULM with the same acquisition time (1 second per slice, total 240 slices). The scanning step is 50 um and the scanning direction is along y axis. The generated image size is 23 mm x 12 mm x 10 mm (xyz).3D DL-ULM is generated by stacking a series of 2D DL-ULM. Images in (a) are color coded by scanning distance. Images in (b) are color coded by imaging depth. (c) Zoom of vasculatures in a region depicted by a white-dashed box in (b). Vasculatures along elevational direction (y) and lateral direction (x) are selected, and their size are measured. (d) Gradual improvement of image quality of 3D DL-ULM and ULM in zoom-in region in (a) with increasing acquisition time from 0.2 seconds to 1.8 seconds. (e) Reconstructed quality of DL-ULM and ULM as the function of MS-SSIM verse the acquisition times. The error bars are standard deviations (N = 57). (f) A 3D ULM image is generated by stacking 2D ULM slices along the y dimension, with theresult visualized in the yz plane.(g) Similarly, a 3D DL-ULM image is generated by stacking 2D DL-ULM slices along the y dimension, and it is also visualized in the yz plane for comparison.(h) Another 3D ULM image is created by stacking 2D ULM slices along the x dimension and visualized in the yz plane. Since the yz plane corresponds to the imaging plane, it provides super-resolution and serves as the reference to evaluate the accuracy of the 3D ULM shown in (h). Dual modal PA / ULM imaging for hybrid molecular, anatomical and structural imaging in tumor
[0098] In some implementations, the feasibility of 3D hybrid imaging system may be demonstrated for cancer imaging. Detecting cancer at early stages and accurate cancer diagnosis are high desired to make more successful treatment. Although conventional ultrasound imaging can provide the anatomical information of tumor boundary, it usually requires the tumor grow to a certain size. Molecular imaging, on the other hand, is expected to contribute to the early detection of tumor due to the specific monitoring of key molecular information. In addition, molecular imaging can provide more accurate information about the type, location, and extent of the cancer, which can help guide treatment decisions. With the capability to perform molecular imaging using exogenous contrast agents, PA imaging plays an important role in molecular cancer imaging. In this study, epidermal growth factor receptor (EGFR), which has overexpression in many tumors, is used as a target for molecular imaging. It is well known that EGFR significantly contributes to cancer formation and angiogenesis. Using 3D dual PA / ULM, the expression level of EGFR and vascular information could be accessed simultaneously.
[0099] In some implementations, anti-EGFR antibody labeled with Alexa fluor 790 may be injected into A431 tumor bearing mice through tail veins. 3D ultrasound image overlays with photoacoustic Alexa fluor 790 signals before and after 24 h of injection shows in panel a of FIG.15 (Group 1). A stronger PA signal appears after injecting the particle. To validate the specific target of anti-EGFR antibody, IgG2a antibody labelled with the same dye may be injected as the control. Although PA signal also improves after injecting IgG2a (Group 2 in panel a of FIG.15), PA signal with anti-EGFR injection is much stronger (4.7 times) than PA signal with anti-IgG2a injection (2.3 times).
[0100] The PA signals depicted in panel a of FIG. 15 arise not only from the Alexa Fluor 790, but also from the endogenous contrast produced mainly by hemoglobin anddeoxyhemoglobin. In addition, ultrasound images can only reveal the tumor structure and surrounding tissues. To gain further insight into the distribution of antibodies and tumor vasculature, dual modal multi-wavelength PA / ULM imaging may be employed. By using spectrum unmixing, the distribution of antibodies was made more visible with high contrast. Furthermore, with the aid of ULM, the tumor vasculature may be visualized. Thanks to the faster DL-ULM technology, the microvascular structure of the entire tumor may be imaged with a speed of 1 sec per frame and a step size of 50 um. Panel d of FIG.15 shows the antibodies distribution and tumor vasculature of a mouse in Group 1 and Group 2 after 24 hours of injection. The signal of Alexa Fluor 790 in Group 1 is significantly higher than that in Group 2, which serves as evidence for the specific targeting of the anti-EGFR antibody.
[0101] The specific targeting using whole-body fluorescence imaging may be investigated. The biodistribution of antibodies (anti-EGFR and IgG2a) over time is shown in panel a of FIG.19. The antibodies are targeted at the tumor site. The higher fluorescence signal in Group 1 compared to Group 2 indicates the specificity of anti- EGFR antibody, which corroborates with the photoacoustic / ultrasound imaging results. After 36 h of antibodies injection, the mice were sacrificed and the key organs were excised to further evaluate the bio-distribution of antibodies. Panel b of FIG.15 shows the fluorescence signal of tumors in Group 1 are significantly higher than that in Group 2 (2.3 fold higher from panel c of FIG.15). Furthermore, the biodistribution of different organs (Panel b of FIG.19 and panel c of FIG.15) shows that other than tumors, antibodies mainly accumulated in the liver and partly accumulated in the kidney and spleen.
[0102] FIG.15 shows that Dual-modal 3D PA / ULM reveals molecular and structural information of tumor. (a) The comparison of PA signal before and after injecting either Alexa Fluro@790 (AF790) labeled anti-EGFR antibody (Group 1) or IgG2a (Group 2) via tail vein. Scale bar = 5 mm. (b) The fluorescence signal of harvested tumors from mice in Group 1 and Group 2. (c) Biodistribution of different organs after 36 h of injecting either Alexa Fluro@790 labeled anti-EGFR antibody or IgG2a. The error bars are standard deviations (N = 3). ** p<0.01 (p=0.0092). (d) Dual-modal PA / ULM shows tumor vasculature and the distribution of AF790 signal of mice in Group 1 and Group 2. Scale bar = 1 mm.
[0103] FIG.18 shows validation of EGFR antibody targeting. (a) Fluorescence signal of different volume of antibody incubated with A431 cells. (b) Fluorescence images of different antibody incubated with A431 cells.
[0104] FIG. 19 shows biodistribution study in tumor bearing mice. (a) In vivo fluorescence imaging of tumor-bearing mice after injection of either Alexa Fluro@790 (AF790) labeled anti-EGFR antibody (Group 1) or IgG2a (Group 2) via tail vein. (b) Fluorescence imaging of harvested organs after 36 h of particle injection in Group 1 and Group 2. Dual modal PA / ULM imaging monitor blood-brain barrier permeability
[0105] Non-invasive delivery of therapeutic agents to cerebral tissues by overcoming the blood-brain barrier (BBB) has posed a significant challenge for years. Focused ultrasound (FUS) in combination with circulating microbubbles has emerged as a potent method for achieving transient and reversible permeabilization of the vasculature, and it has been the subject of investigation in numerous pre-clinical studies. Nevertheless, the available imaging modalities for monitoring BBB disruption remain limited. In this context, an exploration of the feasibility of dual PA / ULM imaging for monitoring BBB disruption may be conducted.
[0106] Focused ultrasound with microbubbles was used to open BBB on the left-brain hemisphere of a mouse. Following FUS activation, an ICG solution was administered to the mouse via the tail vein. Panel a of FIG.16 displays the distribution of ICG before (0 min) and after (30 min, 110 min, and 170 min) BBB disruption. A more pronounced ICG signal is evident on the left hemisphere after BBB disruption, indicating the successful opening of the BBB through FUS. In panel b of FIG. 16, quantitative measurements of the ICG signal are presented as a function of BBB opening time, with the left and right regions for ICG signal measurement highlighted by orange circles in panel a of FIG.16. A stronger ICG signal in the left region was observed after BBB opening and remained elevated even 2 hours after BBB opening.
[0107] ICG distribution after BBB opening may be quantitatively evaluated. The brain region spanning from bregma -1.9 mm to -4.15 mm may be selected as the region of interest due to the accumulation of strong ICG in this area (panel a of FIG.16). Panel c of FIG.16 provides insights into the distribution of ICG from both coronal and sagittal views after 110 minutes of BBB opening. Additionally, the ICG signal in different major brain regions may be quantified, revealing that the ICG signal in the left hemisphere issignificantly higher than in the right hemisphere in the isocortex, hippocampal region, and thalamus (Panels d-e of FIG. 16). However, in deeper regions such as the midbrain, hypothalamus, and striatum, no significant difference was detected in the ICG signal between the left and right hemispheres. Furthermore, a quantitative comparison of microbubble density in the BBB opening regions (bregma -1.9 mm and -4.15 mm) may be conducted. It was observed that the right hemisphere exhibited a higher microbubble density than the left hemisphere in the isocortex, hippocampal region, thalamus, midbrain, and hypothalamus. However, no significant difference was detected in the ICG signal between the left and right hemispheres in the striatum region. This may be attributed to the relatively small area of the striatum within the field of interest (bregma -1.9 mm to -4.15 mm) and the short data acquisition time (1 second per frame).
[0108] Panel f of FIG.16 presents the velocity map spanning from bregma -1.9 mm to -4.15 mm. Similar to the microbubble density distribution, the right hemisphere exhibits higher velocity compared to the left hemisphere. Zooming in on the midbrain and hippocampus region (panel g of FIG. 16) reveals that the blood velocity in the right hemisphere is greater than in the left hemisphere. In panel h of FIG.16, a comparison of velocity across different brain regions is shown. The right hemisphere displays higher microbubble density than the left hemisphere in the isocortex, hippocampal region, thalamus, midbrain, and hypothalamus. However, no significant difference was detected in the ICG signal between the left and right hemispheres in the striatum region, like that in microbubble density quantification. This may be due to the relatively small area of the striatum within the field of interest (bregma -1.9 mm to -4.15 mm) and the short data acquisition time (1 second per frame).
[0109] FIG. 16 shows that dual-modal 3D PA / ULM detects blood-brain barrier disruption. (a) In vivo ICG distribution in a mouse brain before and after 30min, 120 min and 200 min of BBB disruption with focused ultrasound. The microvasculature image merged with ICG signals is collected after 120 min of BBB disruption. The scale bar is 1 mm. (b) ICG signal as function of time after BBB disruption. Left and right regions are indicated in (a). The error bars are standard deviations (N = 3). (c) Coronal views of 3D PA / ULM with the region from bregma -1.9 mm to -4.5 mm and sagittal views of left hemisphere reveals the ICG distribution and microvasculature of brain after 120 min of BBB disruption. The scale bar is 1 mm. (d) Average MB spatial density and ICG signal in the main cerebral regions of left hemisphere and right hemisphere.The average MB density in a specific region is calculated by the total MB counts over the areas of the specific region in the range of bregma from -1.9 mm to -4.5 mm. The average ICG signal in a specific region is calculated by the total MB counts over the areas of the specific region in the range of bregma from -1.9 mm to -4.5 mm. The scale bar is 1 mm. (e) The comparison of MB counts and ICG signal at left hemisphere and right hemisphere. The error bars are standard deviations (N = 5). (f) 3D blood speed map from bregma -1.9 mm to -4.15 mm after 120 min of BBB disruption. The scale bar is 1 mm. (g) The velocity difference in midbrain and hippocampus between left and right regions of brain. The scale bar is 1 mm. (h) The comparison of velocity at left hemisphere and right hemisphere. The error bars are standard deviations (N = 5). R1 to R6: isocortex, hippocampal formation, thalamus, midbrain, hypothalamus, striatum.
[0110] The present disclosure describes a fluorescent image, confirming BBB disruption after HIFU stimulation. Dual modal PA / ULM imaging for microvascular and lymph node imaging in hindlimb
[0111] Dual modal PA / ULM has the feasibility to simultaneously detect the blood vessel system and lymphatic system. Blood vessels are essential to support tissue growth by providing the oxygen, nutrient supply, and material change, while the lymphatic system, with the function to maintain the fluid homeostasis and immune function, is the complementary of blood vasculature. ICG is injected through the footpad of the mouse here to track the lymphatic drain. With the ability to detect the interaction of the lymphatic system and blood vessels, especially capillary bed, it would contribute to the understanding of different pathologies and assessment of therapeutic strategies.3D ultrasound image in panel a of FIG.17 shows the anatomy of the mouse hindlimb, where the lymph node position is barely to observe. Using photoacoustic imaging, the lymph node may be visualized by extracting ICG distribution using spectrum unmixing, as shown in panel b of FIG.17. The fluorescence image of the mouse further confirms the successful ICG accumulation in the popliteal lymph nodes, as shown in panel c of FIG.17. Using dual PA / ULM imaging, one can generate 3D DL-ULM with the acquisition of 1 sec per frame. When zooming in a smaller region (panel e of FIG.17), DL-ULM can reveal vessels as small as 10 um, while Doppler can only reveal large vessels. By co-registering the PA and ULM images, 3D lymph node combined with microvasculature image may be obtained (panel f of FIG.17).
[0112] FIG. 17 shows that dual-modal 3D PA / ULM reveals structural and anatomy information in hindlimb. (a) 3D dual PA / US image before injecting ICG dye through the footpad of a mouse. (b) 3D dual PA / US image after 15 min of injection. (c) Validation of ICG accumulation into the lymph node using fluorescence imaging. (d) 3D DL-ULM reveals blood vasculature of a hindlimb, but cannot show lymphatic vessel and lymph node. (e) Zoom in region in (d) shows the resolution improvement of DL-ULM compared to Doppler image. (f) Dual modality of PA / ULM shows blood vessel and lymph node. Discussion
[0113] Various embodiments in the present disclosure demonstrate the applicability and feasibility of 3D PA / ULM imaging. The main limitation to hamper the implementation of the hybrid imaging is the low acquisition time of ULM.
[0114] Various embodiments including neural network (e.g., GAN) may overcomes these limitations by predicting microbubble fluctuation. Even with low concentrated microbubbles injection, with the information of microbubble sparse locations and power doppler vessel image, GAN can depict the super-resolved microvascular image with high accuracy. In addition, elevation resolution of ULM may be improved with dynamic localization and deep learning.
[0115] In some implementations, considering dual-modal PA / ULM system, although it accelerates the acquisition of ULM, the speed of PAI, which may be restricted by the repetition rate of the laser source, hampers the real-time implementation of dual modal imaging. The repetition rate of the laser source is related to the laser power. Faster repetition rate cannot provide enough laser fluence for photoacoustic imaging. This limitation may be alleviated by combining several ultrafast laser sources.
[0116] Moreover, current deep learning only supports the morphologic prediction of super-resolution microvascular imaging and cannot detect the hemodynamics information like blood flow, which is a valuable and available functional parameter in conventional ULM. Some embodiments may improve the deep learning to extract the blood flow information.
[0117] Furthermore, for real-time implementation, instead of considering the acquisition time, the processing time is also essential. GAN has a short processing time (around 100 ms) after training. This fast processing provides more time for pre- processing. In GAN-ULM, the pre-processing time mainly includes the imagereconstruction, doppler process and the localization. The real-time clutter filtering for doppler process and GPU parallel programming for the localization may be implemented. Thanks to the low number of frames needed to be processed (200 to 300 frames of B-mode), the total pre-processing time can be controlled within 200 ms. With further advanced computation strategies, like GPU-based image reconstruction and deep-learning based localization, GAN-ULM has the feasibility to target the boundary of real-time imaging. Conclusion
[0118] The present disclosure shows the feasibility of achieving 3D dual modal ULM / PAI. ULM is accelerated by multi-stream model to match the photoacoustic acquisition speed. The deep learning results have been validated by different organs, including hindlimb, tumor and brain. In addition, dual modal ULM / PA are demonstrated in these organs, as well. In some implementations, limitations on laser repetition rate and processing time, including image reconstruction, microbubble localization, and spectral unmixing may limit the implementation of real-time imaging. In some implementations, such limitations can be improved by upgrading the laser source and using computational strategies like parallel programming to accelerate the process. Contrast agent characterization.
[0119] Commercial microbubbles (e.g., Vevo Micromarker) were used as ultrasound contrast agent for ULM imaging in this study. ICG dye (e.g., Indocyanine green, ATT Bioquest) and Alexa Fluor@790 (e.g., AF790) dye were used as photoacoustic contrast agents. The optical absorption of ICG dye was measured by Microplate reader (e.g., BioTek, Synergy) at room temperature.0.1 mL of 0.5 mg / mL ICG solution was used. Anti-EGFR Antibody (e.g., A-10) Alexa Fluor@790 (e.g., Santa Cruz Biotechnology) was used as the sample to measure the optical absorption of AF790. 0.1 mL samples were measured by the same Microplate reader at room temperature. Materials used in some exemplary embodiments
[0120] Cell culture: The human cancer cell lines A431 were obtained from were obtained from the American Type Culture Collection (ATCC). A431 tumor cells were maintained in Dulbecco's Modified Eagle's Medium (e.g., ATCC). The medium was supplemented with 10% fetal bovine serum. Cells were incubated at 37 °C in 5% CO2and passaged at 80%–90% confluence. All cell lines were regularly checked for mycoplasma contamination.
[0121] Cell culture uptake of EGFR-antibody: 1×104A431 cells were plated in a 96- well plate. Twenty-four hours after seeding, the cell media were removed, and the cells were washed twice with PBS. Then, the cells were incubated with either 10 µL Alexa Fluor 790‐labeled mouse antihuman EGFR monoclonal antibody (e.g., AF790‐EGFR‐ Ab, sc‐120 AF790; Santa Cruz Biotechnology) or 10 µL Alexa Fluor 790‐labeled normal IgG2a (e.g., sc‐24637 AF790; Santa Cruz Biotechnology) in the cell culture medium for 12 hours. After that, cells were imaged by a fluorescence microscope (e.g., BioTek Cytation 5, Agilent) with 10X objective lens, and Cy7 filter set.
[0122] Quantification of biodistribution of EGFR-antibody: When the tumor grew to around 1 cm3, 50 µg (250 µL) Alexa Fluor 790‐labeled mouse antihuman EGFR monoclonal antibody (e.g., AF790‐EGFR‐Ab, sc‐120 AF790; Santa Cruz Biotechnology) was retro-orbitally injected to mice (n=3). Alexa Fluor 790‐labeled normal IgG2a (e.g., sc‐24637 AF790; Santa Cruz Biotechnology) was used as a negative control (n=3). IVIS spectrum imaging system (e.g., PerkinElmer) was used with a 745 nm excitation filter and a 840 nm emission filter with 2 seconds of exposure time to record at 4 hours, 12 hours, 18 hours, 24 hours, and 36 hours the fluorescent images of the tumors after antibody injection. For the bio-distribution, mice were sacrificed 36 hours after injecting the antibody solution. Epi-fluorescence imaging of the excised organs was carried out using an IVIS spectrum imaging system (PerkinElmer) with an excitation filter centered at 745 nm and an emission filter centered at 840 nm with 2 seconds of exposure time. Quantitative analysis was performed using the Living Image 4.5 software.
[0123] Animal studies: All procedures performed on mice were approved by the Institutional Animal Care and Use Committee (IACUC). Mice were housed in an animal care facility approved by the American Association for Assessment and Accreditation of Laboratory Animal Care. A total 15 female BALB / cJ mice (e.g., Jackson Laboratory) and total 18 female Nu / J mice (e.g., Jackson Laboratory) at age 8-10 weeks were used. Human breast tumor xenografts were induced in 8 weeks old female immunodeficient nude mouse (e.g., Nu / J, Jackson Laboratory) (n=18). To develop breast cancer in mouse model, 50 µL of 3x107A431 tumor cells mixed with 1:1 volumeratio of Matrigel (e.g., Corning) were subcutaneously injected into the left flank of each mouse. Imaging system and imaging sequence used in some exemplary embodiments
[0124] Imaging system configuration: The dual modal 3D PAUL imaging system consists of a Verasonics ultrasound imaging research system (e.g., Vantage 256, Verasonics Inc., Kirkland, WA, USA), a wavelength tunable (690 nm to 950 nm) OPO laser source with 7-ns-pulse and 10 Hz pulse repletion rate (e.g., Phocus Essential, Opotek, Inc., Carlsbad, CA, USA), a 15 MHz linear array transducer (e.g., MS250, Visualsonics Inc., Toronto, ON, Canada) and a XYZ linear stage. A customized bifurcated fiber bundle is used to deliver light from the laser source. The function generator is used to synchronize the data acquisition and the laser firing. The fiber bundle is integrated to the transducer using a 3D-printed adaptor for the side- illumination. When performing imaging, the transducer is mounted on the linear stage and is connected to Verasonics hardware for data acquisition of ultrasound and photoacoustic signal.3D volume imaging is performed through mechanical scanning of the transducer.
[0125] Imaging sequences: Hybrid acquisition for the imaging sequence were designed. Each hybrid acquisition includes one photoacoustic data acquisition and multiple ultrafast ultrasound data acquisition. When the laser system generates a laser pulse, it simultaneously sends a trigger signal to Verasonics system, and the system starts to receive the generated PA signal. Then, the system transmits and receives multiple ultrasound signals for 100 milliseconds. For each ultrasound frame, the transducer transmits plane waves with seven-angle (-6° to 6°) at pulse repetition frequency of 500 Hz. The time interval between two adjacent frames (either PA / US or US / US) is 2 ms (or 500 Hz). In this case, the acquisition time of dual PAUL imaging is determined by the acquisition time of ULM. Imaging protocols used in some exemplary embodiments
[0126] In vivo PAUL imaging – Tumor: The mouse was placed on a heat pad and was anesthetized with 2% isoflurane at 2 L / min of oxygen flow. A 30-gauge home-made catheter was cannulated into the tail vein of a mouse before imaging. The transducer was placed on the top of the tumor and the acoustic gel was applied for ultrasound- coupling. Microbubbles with the concentration of 1x108bubbles / mL were injected viacatheter at the rate of 20 uL / min using a syringe pump. The hybrid imaging sequences with PA off were applied to collect ULM frames for deep learning dataset. To perform label-free ULM, the same imaging sequences were used but without microbubbles injection.
[0127] For molecular imaging, after two weeks of tumor growth, 250 µL Alexa Fluor@790 labelled anti-EGFR antibody was retro-orbitally injected into a tumor bearing mouse body. The same volume of Alexa Fluor@790 labelled IgG2a antibody was retro-orbitally injected into another tumor bearing mouse body. After 24 h of injection, the mouse was placed on a heat pad and was anesthetized with 2% isoflurane at 2 L / min of oxygen flow. A 30-gauge home-made catheter was cannulated into the tail vein of a mouse before imaging. The transducer was placed on the top of the tumor and the acoustic gel was applied for ultrasound-coupling. Microbubbles with the concentration of 1x108bubbles / mL were injected via catheter at the rate of 20 µL / min using a syringe pump. The transducer was then scanned laterally to collect 3D imaging data of the whole tumor. For each position, 0.8 sec of data acquisition was performed with 4 laser wavelengths of 700 nm, 780 nm, 800 nm and 850 nm. The total positions were based on the tumor size and the step size was 0.05 mm.
[0128] To longitudinal monitor tumor growth, seven days, eleven days, thirteen, fifteen and seventeen days after injecting tumor cells were selected as imaging date. During the experiment, the transducer was scanned laterally to collect 3D imaging data of the whole tumor. For each position, 0.9 sec of data acquisition was performed with 3 laser wavelengths of 750 nm, 800 nm, 850 nm. The total positions were based on the tumor size and the step size was 0.05 mm.
[0129] In vivo PAUL imaging – Brain: The animals were anesthetized with the same protocol as tumor imaging. A 30-gauge catheter was cannulated into the tail vein of a mouse. The animal was then placed in a stereotaxic frame (RWD, China) and the skin above the skull was removed but the skull was kept intact. The acoustic gel was applied on the skull and a home-made plastic water tank with a transparent bottom window was placed above the gel and head. The transducer was placed in the water tank and on the head of mouse. Microbubbles (1x108bubbles / mL) were injected at the rate of 20 µL / min using a syringe pump. Single ULM modality was collected using the hybrid imaging sequences with PA off for deep learning datasets.
[0130] To monitor blood brain barrier (BBB) disruption, focused ultrasound (FUS) mediated BBB disruption was applied before imaging. The animals were anesthetizedwith the same protocol as tumor imaging. A 30-gauge catheter was cannulated into the tail vein of a mouse. The animal was then placed in a stereotaxic frame (RWD, China) and the skin above the skull was removed but the skull was kept intact. Focused ultrasound transducer (central frequency 0.55 MHz, H104, Sonic Concept) was positioned with the mouse head coupled with acoustic gels.50 µL microbubble solutions (2x109bubbles / mL) was injected via the mouse tail vein. After 1 min of microbubble injection, FUS was activated with the following parameters: 0.2 MPa negative peak pressure, 2 ms of burst length, 1 Hz of pulse-repetition rate (PRF), 60 sec sonification on and 60 sec sonification off and this cycle repeated once. After sonification, 100 µL ICG solution (30 mg / kg) was injected via the mouse tail vein.
[0131] To longitudinal monitor BBB disruption, the hybrid imaging sequences were applied before and after 15 min, 30 min, 50 min, 70 min, 90 min, 120 min and 200 min of ICG solution injection. The transducer was placed in the water tank and on the head of mouse. Microbubbles with the concentration of 1x108bubbles / mL were injected via catheter at the rate of 20 µL / min only at 120 min of ICG solution injection. For each time point, the transducer was scanned laterally to collect 3D imaging data of the whole brain. For each position, 1.6 sec of data acquisition was performed with 4 laser wavelengths of 700 nm, 780 nm, 800 nm and 850 nm. The step size was 0.05 mm and a total of 260 positions were recorded.
[0132] To confirm BBB opening, after 4 hours of ICG solution injection, animals were sacrificed and perfused with 1X PBS. Then brains were excised and imaged using an IVIS spectrum imaging system with an excitation filter centered at 745 nm and an emission filter centered at 840 nm with 2 seconds of exposure time.
[0133] In vivo PAUL imaging – Hindlimb: The animals were anesthetized with the same protocol. A 30-gauge catheter was cannulated into the tail vein of a mouse. Microbubbles were injected at the rate of 20 µL / min using syringe pump. Single ULM modality was collected using the hybrid imaging sequences with PA off for deep learning datasets.
[0134] For the dual modality acquisition, 20 µL ICG solution (0.2 mg / mL) were injected through the footpad of the mouse before anesthetization. After 15 min of the injection, the mouse was placed on the heat pad and anesthetized. Microbubbles were injected via tail vein at the rate of 20 µL / min using a syringe pump. The transducer was then scanned laterally to perform 3D imaging. For each position, 1.6 sec of data acquisitionwas performed with 4 laser wavelengths of 700 nm, 780 nm, 800 nm and 850 nm. A total of 66 positions were collected with the step size 0.05 mm.
[0135] To validate that ICG was delivered into the lymph node, after 15 min of the injection, IVIS spectrum imaging system (PerkinElmer) was used to check the ICG biodistribution. An excitation filter centered at 745 nm and an emission filter centered at 840 nm with 2 seconds of exposure time was used. The images using Living Image 4.5 software were quantitatively analyzed. Imaging processing used in some exemplary embodiments
[0136] Power Doppler workflow: The acquired ultrasound data were beamformed using the delay and sum algorithm to reconstruct ultrasound images. By applying the singular value decomposition (SVD) based spatiotemporal filtering to reconstructed ultrasound cineloop, doppler signal over time can be obtained. Then a rigid motion model was used to correct the tissue motion. Briefly, a frame-to-frame rigid transformation was applied to ultrasound image frames and calculated the translation and rotation information using gradient descent optimization (imregtform function with regular step gradient descent optimizer in MATLAB). Doppler signal was corrected to compensate for tissue motion using translation and rotation information. The power doppler image was calculated as the power of the doppler signal.
[0137] ULM imaging workflow: Similar to power Doppler processing, the acquired ultrasound data were beamformed using the delay and sum algorithm to reconstruct ultrasound images. The singular value decomposition (SVD) based spatiotemporal filtering is applied to ultrasound images to extract the flowing microbubble signal. After filtering, there may still exist noise in each microbubble frame. Noise can be further suppressed using thresholding (to 40 dB) to reject the intensities lower than a defined threshold. Since the bubbles are much smaller than the wavelength and can be individually separated in space and time, they appear as the PSF of the ultrasound system. A system points spread function (PSF) is estimated using Gaussian fitting of one isolated microbubble signal. Then the cross-correlation between PSF and microbubble signals over frames is calculated to localize microbubbles. By rejecting cross-correlation coefficient index less than the threshold (0.7 experimentally), isolated blobs can be created. The locations of microbubbles are identified by the peak of these isolated blobs in each frame. The motion compensation model (same in power Doppler processing) was applied to correct tissue motion.
[0138] After correcting all microbubble locations over frames, the Hungarian tracking algorithm was implemented to track moving microbubbles. The algorithm calculates all distances between all microbubbles in the current frame and all microbubbles in the next frame. Then the algorithm minimizes the total distance and finds the optimal pairs of microbubbles in adjacent frames. Applying the algorithm to all frames can provide the collections of a series of microbubble tracks over time. The final super resolution ultrasound image was obtained by accumulating these positions over frames. The velocities of each microbubble may be calculated by differentiating each position in a trajectory according to the time vector to generate blood flow velocity map.
[0139] To obtain elevational-improved ULM image, dynamic localization was applied to each microbubble trajectory and discarded out-of-plane microbubbles. The final elevational-improved ULM image is obtained by accumulating microbubble positions after dynamic localization.
[0140] DL-ULM imaging workflow: DL-ULM is generated by feeding sparse ULM image and power Doppler image into the trained multi-stream GAN network. Here, sparse ULM image means the image is generated using short acquisition time (<< 1 min). Sparse ULM image is generated following ULM imaging workflow. Power Doppler image is generated following power Doppler imaging workflow.
[0141] Photoacoustic imaging workflow, Fluence compensation: After photoacoustic data acquisition, reconstruct photoacoustic images were constructed using the delay and sum algorithm.
[0142] MCXLab simulator may be used to model the light propagation. Different light propagation layers (i.e. air, skin, tissue and skull) may be manually segmented from ultrasound volumes using 3D slicer. Then the bi-side light illumination was simulated using two planer beams with an angle of 30 degrees. The optical properties of different components were assigned. The fluence map obtained from the Monte Carlo simulation was then used to compensate for the photoacoustic images.
[0143] Spectrum unmixing and oxygen saturation: Linear regression-based spectrum unmixing was performed to extract each contrast signal from compensated photoacoustic images. Briefly, the measured photoacoustic signal is a linear combination of the spectral signatures of different chromophores (i.e. hemoglobin, deoxyhemoglobin and exogenous contrast agent) with the relative concentration of each chromophores. If no exogenous contrast agents are used:Wherein, ^^^λ୍, ^^,^^^ is the photoacoustic signal recorded at wavelength λ୍, ϕ^λ୍^ is thelocal optical fluence at wavelength λ୍, ^^ு^ோ^λ^^and ^^ு^ைమ^λ^^are the molar extinctioncoefficients of HbR and HbO2 at the wavelength λ୍, respectively. ^^ு^ோ^^^, ^^^ andare the molar concentration of HbR and HbO2. If exogenous contrastagents are used, the contribution from the exogenous contrast agents need to beconsidered:Wherein ^^^௫^^λ^^ is the molar extinction coefficients of exogenous contrast agents and^^^௫^^^^,^^^ is the molar concentration of the exogenous contrast agents. In variousembodiments, ICG and Alexa-Fluro@790 may be used as exogenous contrast agents.
[0144] In some implementations, the absorption spectrum may be obtained either by measurement or database (i.e. the absorption of hemoglobin and deoxyhemoglobin is from https: / / omlc.org / spectra / hemoglobin / ). The following linear equation is used to estimate molar concentration of different contrast components:where ^^ᇱis photoacoustic signal after fluence compensation. The oxygen saturationis calculated by:
[0145] Brain Allen atlas registration: The registration of the PAUL image with the Allen Mouse Common Coordinate Framework (CCF) was performed with the registration software described in. Briefly, a power Doppler volume is first generated from PAUL data. Manually translating, scaling, and rotating the power Doppler volume is required to align Allen Mouse CCF. The software outputs an affine transformation matrix after manual registration. Then, this transformation could be applied to ULM or PA volumetric map and used to extract signals from any regions of interest for quantification.
[0146] Deep learning models, Network architecture: a neural network in various embodiments includes two main networks: a generator and a discriminator. First, generator G relies on U-net architecture to reconstruct super-resolved ultrasoundimages. A multi-stream encoder network may be used where each network takes as inputs 512×512 sized sparse ultrasound localization image (ULM stream) and a power Doppler ultrasound image (PD stream), respectively. The encoder is built on convolution layers with a stride of 2 for down-sampling, followed by batch normalization and leaky ReLUs for activation.
[0147] Similar to the U-Net architecture, skip layer connections are added to the encoder networks to concatenate feature maps located symmetrically between the encoder and the decoder. The difference between the two encoders lies in that the skip layer between the first layer to the layer is removed in the PD stream. This is to only pass down the high-level features of the power Doppler images to provide global guidance, allowing recovery of distinct structures.
[0148] The decoder network upsamples every input using convolutional layers, resulting in feature maps of a double on each dimension and half in the number of channels. The input is upsampled by a factor of 2 using nearest-neighbor interpolation followed by a convolutional layer. The last layer of the decoder network uses a Tanh activation function to output the reconstructed dense ULM image with the size of 512×512. Batch normalization is used after all convolutional layers with dropout units with a probability of p=0.5 applied to the second, third and fourth layers of the decoder. Skip-layer connections are added to the generator to symmetrically link convolution and deconvolution layers with the same size, passing the local and pixel-wise context from the encoder network to the decoder network.
[0149] Second, the discriminator D network comprises five convolutional layers and reduces the output to a size of 64×64. The discriminator takes 3-channel image as an input: sparse ultrasound image, the scaled-up power Doppler image, and either the reconstructed or real dense ULM image. The pixel values of the down-sampled output image indicate whether the third input image channel is a real dense ULM image, or an image produced by a generator.
[0150] Loss functions: GANs are trained to generate images that match the statistics of the real data through the joint optimization of a generator ^^ and a discriminator ^^. The generator learns to map a latent variable ^^ drawn for a probability density ^^^^^^^^to create synthetic image samples of a true data probability density ^^ௗ^௧^^^^^. Then, the discriminator attempts to distinguish the generated image and the real samples. The adversarial loss penalizes the generator such that it is updated to produce realisticsamples to fool the discriminator and penalizes the discriminator misclassifications. The minmax objective for GANs is formulated as following, arg m^ax ൫^^௫~^^ೌ^ೌ^௫^^^^^^^^^ − 1^ଶ^ + ^^௭~^^^௭^^^^^(^^(^^^)ଶ^൯^^⋆ = ^^^^^^ mீin ^^௭~^^(௭)^(^^(^^(^^)) − 1)ଶ^wherein the cross-entropy loss is replaced by least square loss to generate realistic images. The GAN network is extended into a conditional model where the network is conditioned on input images. Therefore, the discriminator loss ℒୈ(^^)and generatorloss ℒୋ(^^) is given as
[0151] In various embodiments, the input ^^ is the sparse ULM image ^^^^combined with the up-sampled power Doppler image ^^^^= B(P), where ^^ and ^^ indicate bilinear interpolation and power Doppler image, respectively. The objective function for theGAN network is rewritten as,
[0152] To improve the quality of the reconstructed image, additional objective may be incorporated for the training of the generator ^^. The first loss is the super-resolution reconstruction error that minimizes the difference between generator output ^^^^^and the real dense ULM image ^^^^. The loss is defined as the combination of multiscalestructural similarity index (MS-SSIM) and modified L1 norm expressed as,Here, * denotes the convolution and ^^^is a Gaussian window to make the L1 norm consistent with the MS-SSIM loss. Combining the two losses together, the objectiveof the generator network ^^ is given as,^^ீ(^^, ^^) = ^^^^ௌோ(^^) + ^^^^^ீ^ே(^^, ^^)
[0153] In summary, the final objective function can be defined as the following,
[0154] Image pre-processing and training / test: Training was performed on upsampled power Doppler (^^^^), sparse ULM (^^^^), and dense ULM (^^^^) images with the size of512 × 512 pixels. Before training, the images were rotated by a random angle between0° and 360° followed by elastic transformations. Then, the input ^^^^ may be normalizedby subtracting its mean and dividing by its standard deviation. The output image ^^^^is truncated to have the maximum value of 200 and then scaled to lie between a range of 0 and 1. ^^^^images are scaled to have a pixel value that lies between a minimum of 0 and a maximum of 1. If ^^^^is not present, the input is replaced by an image of zeros with the same shape.
[0155] In some implementations, the weights may be set as ^^ = 50 and ^^ = 1 and ^^ =0.84. The objective function was optimized with stochastic gradient descent (SGD) andAdam optimizer at the batch size of 1. The network may be trained for 20,000 iterations using single NVIDIA 2080 Ti graphics processing unit (GPU).
[0156] Dynamic Localization: After microbubble tracking, collections of a series of microbubble tracks over time may be obtained. During the localization process, backscattered signal of each microbubble (which is the pixel value of microbubble position) and associate with every microbubble track may be obtained. In each microbubble trajectory, backscattered signals of microbubbles as the function of time may be obtained. Gaussian curve fitting may be used with the backscattered curve. The width of the Gaussian function corresponds to the elevational resolution ofultrasound system. Gaussian fitting results in a Gaussian curve is ^^(^^)and its full width half maximum (FWHM) is 2√^^^^2 ^^. Based on the Gaussian curve, microbubble positions within a trajectory in 3D domain may be recovered. Thedistance of microbubble from the imaging plane can be estimated bywhere^^^^^௧ is the elevational resolution of ultrasound system. Therefore, the elevationallocation ^^ᇱ =ி^ுெ ^^^^^௧. From the geometry setup of the linear array, the reallocation and the real axial location ^^ᇱ^ = ^^^ . The final super-resolutionimage is obtained by accumulating these positions over frames. In some implementations, an exemplary dynamic localization process may besummaries as following. The input includes one microbubble trajectory: ^(^^^, ^^^, ^^^, ^^^),(^^ଶ, ^^ଶ, ^^ଶ,^^ଶ), … , (^^^, ^^^, ^^^, ^^^)^ , where (^^^ , ^^^) means localized microbubblepositions in lateral and axial direction and ^^^means time. ^^^means the backscatteredmicrobubble signal at position (^^^ , ^^^) and time ^^^. Elevational resolution of ultrasoundimaging system ^^^^^௧ um. The output includes one 3D microbubble trajectory:^(^^′^, ^^′^, ^^′^),… , (^^′^, ^^′^, ^^′^)^. The dynamic localization process mayinclude: Step 1: Gaussian fitting of the trajectory: ^^(^^)with the trajectorydata ^(^^^, ^^^), (^^ଶ,^^ଶ), … , (^^^,^^^)^. Step 2: Calculate FWHM of Gaussian curve:^^^^^^^^ = 2√^^^^2 ^^. Step 3: Recover 3D microbubble localization with the followingequation (^^ = 1,2, … , ^^):^^ᇱ^ = ^^^ ,ᇱ 2(t୧ − b)^^ ^ =^^^^^^^^ ^^^^^௧ ,
[0157] Simulation: To investigate the performance of dynamic localization, microbubble flow may be simulated within two parallel tubes in xy plane and with 45oto xz plane (panel c of FIG. 6). The tubes have a diameter of 10 um. MATLAB Ultrasound toolbox (https: / / www.biomecardio.com / MUST / ) is used to generate raw data cineloop. In the simulation, a 15 MHz linear array transducer was applied to transmit plane waves with seven angles (−6° to 6°) at a pulse repetition frequency of 500 Hz. Microbubbles are initialized in two tubes with the same reflectivity coefficient. Microbubbles are set to flow at 10 mm / s. Microbubbles in different tube flow in different directions. The transducer scans along y direction with the step size of 20 nm to get 3D image. At each position 1000 frames are recorded and a total 100 position are recorded. The received RF data was then reconstructed to a B-mode image using the delay-and-sum algorithm. The image data were then processed with the same procedure described in other section of the present disclosure.
[0158] Data and Statistical Analysis: MATLAB was used to do signal and image processing. For the image display, the ultrasound images are shown in logistic scale, photoacoustic, doppler and ULM are displayed in linear scale.3D volume with distance color encoded was rendered with 3D PHOVIS and the rest 3D volume was rendered with Amira 2022.1. Data was plotted using Origin 2020.
[0159] To segment brain vessels to different regions (FIG. 16), binary masks corresponding to six different regions of the brain (Isocortex, Hippocampus, Midbrain,Striatum, Thalamus and Hypothalamus) are computed and applied to the same MB density map, allowing to split and represent the micro vascularization for each region. In addition, HbO2 signal and MB counts of each region can be generated.
[0160] ICG signal overtime at left and right can be generated (panel b of FIG. 16). Standard deviation is calculated by randomly separating the regions indicated by orange circle. MB counts to 3 parts, ICG signal and velocity are calculated by separating volume (bregma -1.9 mm to -4.15 mm) to 5 parts along y direction.
[0161] The methods, devices, processing, and logic described above and below may be implemented in many different ways and in many different combinations of hardware and software. For example, all or parts of the implementations may be circuitry that includes an instruction processor, such as a Graphics Processing Unit (GPU), Central Processing Unit (CPU), microcontroller, or a microprocessor; an Application Specific Integrated Circuit (ASIC), Programmable Logic Device (PLD), or Field Programmable Gate Array (FPGA); or circuitry that includes discrete logic or other circuit components, including analog circuit components, digital circuit components or both; or any combination thereof. The circuitry may include discrete interconnected hardware components and / or may be combined on a single integrated circuit die, distributed among multiple integrated circuit dies, or implemented in a Multiple Chip Module (MCM) of multiple integrated circuit dies in a common package, as examples.
[0162] The circuitry may further include or access instructions for execution by the circuitry. The instructions may be embodied as a signal and / or data stream and / or may be stored in a tangible storage medium that is other than a transitory signal (i.e., non- transitory signal), such as a flash memory, a Random Access Memory (RAM), a Read Only Memory (ROM), an Erasable Programmable Read Only Memory (EPROM); or on a magnetic or optical disc, such as a Compact Disc Read Only Memory (CDROM), Hard Disk Drive (HDD), or other magnetic or optical disk; or in or on another machine- readable medium. A product, such as a computer program product, may particularly include a non-transitory storage medium and instructions stored in or on the medium, and the instructions when executed by the circuitry in a device may cause the device to implement any of the processing described above or illustrated in the drawings.
[0163] The implementations may be distributed as circuitry, e.g., hardware, and / or a combination of hardware and software among multiple system components, such as among multiple processors and memories, optionally including multiple distributedprocessing systems. Parameters, databases, and other data structures may be separately stored and managed, may be incorporated into a single memory or database, may be logically and physically organized in many different ways, and may be implemented in many different ways, including as data structures such as linked lists, hash tables, arrays, records, objects, or implicit storage mechanisms. Programs may be parts (e.g., subroutines) of a single program, separate programs, distributed across several memories and processors, or implemented in many different ways, such as in a library, such as a shared library (e.g., a Dynamic Link Library (DLL)). The DLL, for example, may store instructions that perform any of the processing described above or illustrated in the drawings, when executed by the circuitry.
[0164] While the particular disclosure has been described with reference to illustrative embodiments, this description is not meant to be limiting. Various modifications of the illustrative embodiments and additional embodiments of the disclosure will be apparent to one of ordinary skill in the art from this description. Those skilled in the art will readily recognize that these and various other modifications can be made to the exemplary embodiments, illustrated and described herein, without departing from the spirit and scope of the present disclosure. It is therefore contemplated that the appended claims will cover any such modifications and alternate embodiments. Certain proportions within the illustrations may be exaggerated, while other proportions may be minimized. Accordingly, the disclosure and the figures are to be regarded as illustrative rather than restrictive.
Claims
CLAIMS What is claimed is:
1. A method for obtaining a super-elevational resolution ultrasound localization microscopy (ULM) image, the method comprising: obtaining ultrasound signal acquired from a sample containing microbubbles by a transducer; reconstructing the ultrasound signal to form ultrasound images; filtering the ultrasound images to obtain microbubble signal with a filter; localizing and tracking the microbubbles based on the microbubble signal to obtain microbubble trajectories; performing dynamic localization by selecting in-plane microbubble signal and eliminating out-of-plane microbubble signal based on backscattered amplitudes of the microbubbles; and generating a super-elevational resolution ULM image based on microbubble positions corresponding to the in-plane microbubble signal.
2. The method according to claim 1, wherein: the transducer comprises a linear array transducer.
3. The method according to claim 1, wherein: the filter comprises a singular value decomposition (SVD) based spatiotemporal clutter filter.
4. The method according to claim 1, wherein the localizing and tracking the microbubbles based on the microbubble signal to obtain the microbubble trajectories comprises: applying a localization algorithm based on the microbubble signal to obtain positions of the microbubbles; and tracking the microbubbles with a Hungarian tracking algorithm based on the positions of the microbubbles to obtain the microbubble trajectories.
5. The method according to claim 4, wherein:the localization algorithm comprises a point spread function (PSF) cross correlation localization algorithm; and the tracking algorithm comprises a Hungarian tracking algorithm.
6. The method according to claim 1, wherein the performing the dynamic localization comprises: for each microbubble trajectory corresponding to a microbubble: retrieving backscattered amplitudes of the microbubble by identifying intensities associated with microbubble positions in the ultrasound image; determining an amplitude threshold for the microbubble based on the backscattered amplitudes of the microbubble; in response to a backscattered amplitude of the microbubble being below the amplitude threshold, determining microbubble signal corresponding to the backscattered amplitude of the microbubble as out-of-plan microbubble signal; and in response to a backscattered amplitude of the microbubble being above or equal to the amplitude threshold, determining microbubble signal corresponding to the backscattered amplitude of the microbubble as in-plan microbubble signal.
7. The method according to claim 6, wherein the performing the dynamic localization further comprises: fitting the microbubble trajectory with a Gaussian curve; and recovering the microbubble position for the microbubble within the microbubble trajectory in a three-dimensional (3D) domain based on the Gaussian curve.
8. The method according to claim 1, further comprising: generating a velocity image based on the microbubble positions corresponding to the in-plane microbubble signal; and generating a 3D super-elevational resolution ULM image based on a plurality of super-elevational resolution ULM images.
9. The method according to claim 1, further comprising:inputting the super-elevational resolution ULM image into a pre-trained neural network to generate a complete super-elevational resolution ULM image.
10. The method according to claim 9, wherein: the pre-trained neural network comprises a generative adversarial network (GAN), and the GAN comprises a generator neural network and a discriminator neural network.
11. The method according to claim 9, further comprising: generating a power doppler image based on the ultrasound signal; and inputting the power doppler image into the pre-trained neural network for generating the complete super-elevational resolution ULM image, wherein the pre-trained neural network comprises a multi-stream GAN.
12. An apparatus, comprising: a memory storing instructions; and a processor in communication with the memory, wherein, when the processor executes the instructions, the processor is configured to cause the apparatus to perform the method in any of claims 1 to 11.
13. A non-transitory computer readable storage medium storing computer readable instructions, wherein, the computer readable instructions, when executed by a processor, are configured to cause the processor to perform the method in any of claims 1 to 11.
Citation Information
Patent Citations
Rapid ultrasonic positioning microscopic imaging method based on generative adversarial network
CN114897689A
Systems and methods for ultrasound imaging and focusing
US20220011270A1
Probe having light delivery through combined optically diffusing and acoustically propagating element
US20220202296A1
Methods, systems, and computer readable media for generating images of microvasculature using ultrasound
US20220211350A1
Image processing method and apparatus based on contrast-enhanced ultrasound images
US20220296216A1
Cited By
Micro blood flow imaging method based on adaptive density filtering and related device
CN121176952A