Device diagnostic systems and methods for acquiring and analyzing images from ultrasound probes
By using machine learning algorithms and a pattern removal process, images acquired by different ultrasound probes are mapped to a common space, solving the problem of identifying and characterizing anatomical structures of interest in follow-up scans and improving the analytical accuracy and efficiency of ultrasound imaging systems.
Patent Information
- Application Number
- CN202380096941.7
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Priority Date
- 2023-02-13
- Filing Date
- 2023-02-16
- Publication Date
- 2025-11-07
AI Technical Summary
Existing ultrasound imaging systems struggle to automatically, quickly, and reliably identify and characterize anatomical structures of interest during follow-up scans, especially when image quality variations occur with different probes, leading to analytical difficulties.
By employing machine learning algorithms combined with domain knowledge, images acquired by different probes are mapped to a common space through a style removal process, reducing probe-specific effects, enabling automatic identification and characterization of anatomical structures of interest, and providing device-independent image evaluation guidance.
It improves the accuracy and efficiency of image analysis using ultrasound imaging systems under different probes, reduces diagnostic errors, and increases the speed and accuracy at which clinicians can acquire high-quality images.
Smart Images

Figure CN120916701A_ABST
Abstract
Description
TECHNICAL FIELD
[0001] The disclosed implementations generally relate to systems, methods, and apparatuses for utilizing an ultrasound probe. BACKGROUND
[0002] Ultrasound imaging is an imaging method that uses sound waves to produce images of structures or features within an area probed by the sound waves. In biological or medical applications, ultrasound images can be captured in real-time to show movement of internal organs and blood flowing through blood vessels. These images can provide valuable information for diagnosis and guidance of treatment of various diseases and conditions. SUMMARY
[0003] Medical ultrasound is an imaging modality based on reflections of propagated sound waves at interfaces between different tissues. Advantages of ultrasound imaging over other imaging modalities can include one or more of the following: (1) its non-invasiveness, (2) its reduced cost, (3) its portability, and (4) its ability to provide good temporal resolution (e.g., millisecond level or better). Point-of-care ultrasound (POCUS) can be used by healthcare providers at the bedside as a real-time tool for answering clinical questions (e.g., whether a patient has a developing hip dysplasia). Trained clinicians can perform both the acquisition task and the task of interpreting ultrasound images without the need for radiologists to analyze ultrasound images acquired by highly trained technicians. Depending on the details of the ultrasound examination, there can still be a highly specialized training to learn different protocols for acquiring medically relevant images as high-quality images.
[0004] In some embodiments, after positioning the patient in an appropriate manner, the clinician positions the ultrasound probe on the patient’s body and manually starts looking for an appropriate (e.g., optimal) image that allows the clinician to make an accurate diagnosis. Acquiring an appropriate image can be a time-consuming activity performed by trial and error and can require extensive knowledge of human anatomy. For diagnosis of some conditions, the clinician can need to perform manual measurements on the acquired images. Moreover, fatigue from repetitive tasks in radiology can lead to an increase in the number of diagnostic errors and a decrease in diagnostic accuracy.
[0005] Computer-aided diagnosis (CAD) systems can help clinicians acquire higher quality images and automatically analyze and measure relevant features in ultrasound images. The methods, systems, and apparatuses described herein can have one or more advantages, including: (1) the ability to incorporate medical expert domain knowledge in machine learning algorithms for efficient learning with a limited number of instances, (2) the ability to explain the output of machine learning algorithms, and (3) the ability to account for differences between probability distributions of images acquired with probes from different manufacturers. Systems that lack the ability to account for differences between probability distributions of images acquired with different probes can fail to ensure that an automated system that works on images acquired with a first probe (e.g., a probe from a first manufacturer) will also work on images acquired with a second probe (e.g., a probe from a second manufacturer different from the first manufacturer), thereby hindering the successful development of products that automatically analyze ultrasound images and provide guidance to clinicians to help them acquire better ultrasound images in less time.
[0006] In some embodiments, the methods, systems, and apparatuses described herein provide ultrasound image device-agnostic assessment and provide guidance to users to help them acquire images that meet the minimum requirements for making a diagnosis.
[0007] In some embodiments, the evolution of an anatomical structure of interest over time is automatically tracked using review scans, e.g., to monitor the presence or characteristics of a tumor in a patient, and then a sequence of ultrasound images acquired at different points in time is formed. In some embodiments, the tracking / review process includes automatically (e.g., without additional input from an operator of the ultrasound system): (1) identifying the same anatomical structure of interest in different images over time, (2) characterizing the anatomical structure of interest by extracting a set of relevant features in each different image, and (3) comparing the features extracted over time.
[0008] In some embodiments, to compare the anatomical structure of interest in a new image to the anatomical structure in images acquired in the past, the positioning of the patient and the probe should be as similar as possible to the positioning at which the original images were acquired. Since information about the positioning can not be stored, it can be very challenging to acquire images of the anatomical region of interest that are of sufficient quality to make a meaningful comparison. This problem can become even more challenging when images are acquired with different probes, since the quality of the images can vary.
[0009] Accordingly, there is a need to develop a method that allows a clinician to automatically, quickly, and reliably identify and characterize an anatomical structure of interest in review scans using ultrasound imaging, regardless of the ultrasound device or probe used to acquire such images.
[0010] Portable (e.g., handheld and / or battery-powered) ultrasound devices are capable of producing high-quality images because they contain many (e.g., hundreds or thousands) of transducers, each of which can produce sound waves and receive echoes to create an ultrasound image. As disclosed herein, an ultrasound probe (or computing device) guides an operator by providing guidance on how to position the ultrasound probe in order to obtain a high-quality frame containing an anatomical structure of interest.
[0011] The systems, methods, and devices of the disclosure each have several innovative aspects, one or more of which can be embodied solely or in combination with others. The various aspects of the disclosure can be implemented in any of a variety of ways.
[0012] According to some embodiments, a method of tracking a structure over time includes, at a computer system comprising one or more processors and a memory: obtaining a first frame of a probed region acquired at a first time using a first probe device, and a first set of control parameters used to acquire the first frame; processing the first frame to obtain a first set of properties of the first frame, wherein processing the first frame includes segmenting one or more features from the first frame after reducing one or more first probe-specific effects in the first frame based at least in part on the first set of control parameters. In accordance with a determination that the first frame contains a respective property that is present in a second frame acquired at a time earlier than the first time using a second set of control parameters, the method includes displaying, on a user interface, information related to a difference in the respective property based on the first frame and the second frame between the first time and the earlier time. The second frame is obtained by processing one or more features segmented from the second frame after reducing one or more second probe-specific effects in the second frame.
[0013] According to some embodiments, an ultrasound probe includes a plurality of transducers and a control unit. The control unit is configured to perform any of the methods disclosed herein.
[0014] According to some embodiments, a computer system includes one or more processors and a memory. The memory stores instructions that, when executed by the one or more processors, cause the computer system to perform any of the methods disclosed herein.
[0015] According to some embodiments of the disclosure, a non-transitory computer-readable storage medium stores computer-executable instructions. The computer-executable instructions, when executed by one or more processors of a computer system, cause the computer system to perform any of the methods disclosed herein.
[0016] Note that the various embodiments described above can be combined with any of the other embodiments described herein. The features and advantages described in the specification are not all inclusive and, in particular, many additional features and advantages will be apparent to one of ordinary skill in the art in view of the drawings, specification, and claims. Moreover, it should be noted that the language used in the specification has been principally selected for readability and instructional purposes and can not have been selected to delineate or circumscribe the subject application. BRIEF DESCRIPTION OF DRAWINGS
[0017] The disclosed aspects will be described with respect to the following figures, in which like numerals indicate like elements, and in which:
[0018] Figure 1 An ultrasound system for imaging a patient is shown in accordance with some embodiments.
[0019] Figure 2 A block diagram of an ultrasound device is shown in accordance with some embodiments.
[0020] Figure 3 A block diagram of a computing device is shown in accordance with some embodiments.
[0021] Figure 4 is a workflow in a device agnostic guidance system for acquiring ultrasound images in accordance with some embodiments.
[0022] Figure 5A and Figure 5B An example of acquiring medical ultrasound images for diagnosing hip dysplasia is provided in accordance with some embodiments.
[0023] Figure 6 An example of a multi-task approach for encoding domain specific knowledge into machine learning algorithms is shown in accordance with some embodiments.
[0024] Figure 7 An example of incorporating domain specific knowledge in a machine learning model is shown in accordance with some embodiments.
[0025] Figure 8 is a workflow for performing device agnostic ultrasound image follow-up acquisition in accordance with some embodiments.
[0026] Figure 9 An example input and example output of a system for automatically identifying relevant structures of interest is shown in accordance with some embodiments.
[0027] Figure 10 An example process for mapping anatomical structures of interest to a template is shown in accordance with some embodiments.
[0028] Figure 11 An example is shown for identifying an anatomical structure of interest according to some embodiments.
[0029] Figure 12 Identification of solid and cystic portions within a thyroid nodule is shown according to some embodiments.
[0030] Figures 13A to 13C A flowchart of a method for tracking a structure over time according to some embodiments is shown.
[0031] Figure 14 A flowchart of a method for tracking a structure over time according to some embodiments is shown.
[0032] Reference will now be made to implementations, examples of which are illustrated in the accompanying drawings. In the following description, numerous specific details are set forth in order to provide a thorough understanding of the present application. However, it will be apparent to one skilled in the art that the present application can be practiced without some or all of these specific details. DETAILED DESCRIPTION
[0033] Figure 1 An ultrasound system for imaging a patient is shown according to some embodiments.
[0034] In some embodiments, the ultrasound device 200 is a portable handheld device. In some embodiments, the ultrasound device 200 includes a probe portion that includes a transducer (e.g., transducer 220, Figure 2 ). In some embodiments, the transducer is arranged in an array. In some embodiments, the ultrasound device 200 includes an integrated control unit and user interface. In some embodiments, the ultrasound device 200 includes a probe that communicates with the control unit and user interface that is located outside of the housing of the probe itself. During operation, the ultrasound device 200 generates sound waves that are transmitted (e.g., via the transducer) toward an organ of a patient 110, such as a heart or a lung. The internal organ or other object(s) to be imaged can reflect a portion of the sound waves 120 back toward the probe portion of the ultrasound device 200, which is received by the transducer 220. In some embodiments, the ultrasound device 200 sends the received signals to a computing device 130, which uses the received signals to create an image 150, also referred to as an ultrasound scan. In some embodiments, the computing device 130 includes a display device 140 for displaying the ultrasound image, as well as other input and output devices (e.g., a keyboard, touch screen, joystick, touchpad, and / or speaker).
[0035] Figure 2 A block diagram of an example ultrasound device 200 is shown according to some embodiments.
[0036] In some embodiments, the ultrasound device 200 includes one or more processors 202, one or more communication interfaces 204 (e.g., network interface(s)), memory 206, and one or more communication buses 208 (sometimes called a chipset) for interconnecting these components.
[0037] In some embodiments, the ultrasound device 200 includes one or more input interfaces 210 that facilitate user input. For example, in some embodiments, the input interfaces 210 include port(s) 212 and button(s) 214. In some embodiments, the port(s) can be used to receive a cable for powering or charging the ultrasound device 200, or to facilitate communication between the ultrasound probe and other devices (e.g., the computing device 130, the computing device 300, the display device 140, a printing device, and / or other input-output devices and accessories).
[0038] In some embodiments, the ultrasound device 200 includes a power source 216. For example, in some embodiments, the ultrasound device 200 is battery-powered. In some embodiments, the ultrasound device is powered by a continuous AC power source.
[0039] In some embodiments, the ultrasound device 200 includes a probe portion that includes a transducer 220, which can also be referred to as a transceiver or imager. In some embodiments, the transducer 220 is based on photoacoustic or ultrasonic effects. For ultrasound imaging, the transducer 220 sends ultrasound waves to a target (e.g., a target organ, a blood vessel, etc.) to be imaged. The transducer 220 receives reflected sound waves (e.g., echoes) that are reflected from the body tissue. The reflected waves are then converted into electrical signals and / or ultrasound images. In some embodiments, the probe portion of the ultrasound device 200 is housed separately from the computing and control portion of the ultrasound device. In some embodiments, the probe portion of the ultrasound device 200 is integrated in the same housing as the computing and control portion of the ultrasound device 200. In some embodiments, part of the computing and control portion of the ultrasound device is integrated in the same housing as the probe portion, and part of the computing and control portion of the ultrasound device is implemented in a separate housing that is communicatively coupled with the part that integrates the probe portion of the ultrasound device. In some embodiments, the probe portion of the ultrasound device has a respective transducer array tailored for a respective scanner type (e.g., linear, convex, endocavity, phased array, transesophageal, 3D, and / or 4D). In this disclosure, an “ultrasound probe” can refer to the probe portion of the ultrasound device, or the ultrasound device that includes the probe portion.
[0040] In some embodiments, the ultrasound device 200 includes a radio 230. The radio 230 enables one or more communication networks, and allows the ultrasound device 200 to communicate with other devices (such as the computing device 130, the computing device 300, the display device 140, a printing device, and / or other input-output devices and accessories) over the one or more communication networks. Figure 1the computing device 130 in the Figure 1 the display device 140 in the Figure 3 the computing device 300 in the). In some embodiments, the radio 230 is capable of data communications using any of a variety of custom or standard wireless protocols (e.g., IEEE 802.15.4, Wi-Fi, ZigBee, 6L0WPAN, Thread, Z-Wave, Bluetooth Smart, ISA100.5a, WirelessHART, MiWi, Ultra-Wide Band (UWB), Software Defined Radio (SDR), etc.), custom or standard wired protocols (e.g., Ethernet, HomePlug, etc.), and / or any other suitable communication protocol, including communication protocols not yet developed as of the filing date of this document.
[0041] The memory 206 includes high-speed random access memory, such as DRAM, SRAM, DDR RAM or other random access solid state memory devices; and optionally includes non-volatile memory, such as one or more magnetic disk storage devices, one or more optical disk storage devices, one or more flash memory devices, or one or more other non-volatile solid state storage devices. The memory 206 optionally includes one or more storage devices remotely located from the one or more processors 202. The memory 206, or alternatively the non-volatile memory within the memory 206, comprises a non-transitory computer readable storage medium. In some embodiments, the memory 206 or the non-transitory computer readable storage medium of the memory 206 stores the following programs, modules, and data structures, or a subset thereof:
[0042] Operating logic 240 including processes for handling various basic system services and for performing hardware dependent tasks;
[0043] A communications module 242 (e.g., a radio communications module) for connecting to and communicating with other network devices (e.g., a local network, such as a router providing Internet connectivity, networked storage devices, network routing devices, server systems, computer device 130, computer device 300, and / or other connected devices, etc.) coupled to one or more communication networks via the communication interface(s) 204 (e.g., wired or wirelessly);
[0044] An application 250 for acquiring ultrasound data (e.g., imaging data) of a patient, and / or for controlling one or more components of the ultrasound device 200 and / or other connected devices (e.g., in accordance with a determination that the ultrasound data meets or does not meet certain conditions). In some embodiments, the application 250 includes:
[0045] An acquisition module 252 for acquiring ultrasound data. In some embodiments, the ultrasound data includes imaging data. In some embodiments, the acquisition module 252 activates transducers 220 (e.g., less than all transducers 220, a different subset(s) of transducers 220, all transducers 220, etc.) depending on whether one or more conditions associated with one or more quality requirements are satisfied by the ultrasound data;
[0046] A receiving module 254 for receiving ultrasound data;
[0047] A transmitting module 256 for transmitting ultrasound data to other device(s) (e.g., a server system, computer device 130, computer device 300, display device 140, and / or other connected devices, etc.);
[0048] An analysis module 258 for analyzing whether data (e.g., imaging data) acquired by the ultrasound device 200 satisfies one or more conditions associated with quality requirements of an ultrasound scan. For example, in some embodiments, the one or more conditions include one or more of the following conditions: the imaging data includes one or more newly acquired images that satisfy one or more threshold quality scores, the imaging data includes one or more newly acquired images that correspond to one or more anatomical planes that match a desired anatomical plane of a target anatomy, the imaging data includes one or more newly acquired images that have one or more landmarks / features (or combinations of landmarks / features), the imaging data includes one or more newly acquired images that have a feature of a particular size, the imaging data supports a prediction that images that satisfy one or more requirements will be acquired in one or more upcoming image frames, the imaging data supports a prediction that a first change (e.g., a percentage increase or a number) in a number of transducers used will support an improvement in a quality score of images acquired in one or more upcoming image frames, and / or other similar conditions; and
[0049] A transducer control module 260 for activating (e.g., adjusting) a plurality of transducers 220 during a portion of an ultrasound scan based on determining that the ultrasound data satisfies (or does not satisfy) one or more quality requirements; and
[0050] Device data 280 of the ultrasound device 200, including but not limited to:
[0051] Device settings 282 for the ultrasound device 200, such as default options and preferred user settings. In some embodiments, the device settings 282 include imaging control parameters. For example, in some embodiments, the imaging control parameters include one or more of the following: a number of transducers activated, a power consumption threshold for a probe, an imaging frame rate, a scan speed, a penetration depth, and other scan parameters that control power consumption, heat generation rate, and / or processing load of a probe;
[0052] User settings 284, such as preferred gain, depth, zoom, and / or focus settings;
[0053] Ultrasound scan data 286 (e.g., imaging data) acquired (e.g., detected, measured) by the ultrasound device 200 (e.g., via the transducer 220);
[0054] Image quality requirement data 288. In some embodiments, the image quality requirement data 288 includes clinical requirements for determining the quality of an ultrasound image; and
[0055] Atlas 290. In some embodiments, the atlas 290 includes an anatomical structure of interest. In some embodiments, the atlas 290 includes a three-dimensional representation of an anatomical structure of interest (e.g., a hip, a heart, a lung, and / or other anatomical structure).
[0056] Each of the above-described sets of executable modules, applications, or processes can be stored in one or more of the previously mentioned memory devices, and correspond to a set of instructions for performing the above-described functions. The above-described modules or programs (i.e., sets of instructions) need not be implemented as separate software programs, procedures or modules, and thus various subsets of these modules can be combined or otherwise rearranged in various embodiments. In some embodiments, the memory 206 stores a subset of the modules and data structures described above. Additionally, the memory 206 can store additional modules or data structures not described above. In some embodiments, a subset of the programs, modules, and / or data stored in the memory 206 are stored in and / or executed by a server system and / or external device (e.g., the computing device 130 or the computing device 300).
[0057] Figure 3 A block diagram of a computing device 300 is shown in accordance with some embodiments.
[0058] In some embodiments, the computing device 300 is a server or console that communicates with the ultrasound device 200. In some embodiments, the computing device 300 is integrated into the same housing as the ultrasound device 200. In some embodiments, the computing device is a smartphone, tablet device, game console, or other portable computing device.
[0059] According to some embodiments, the computing device 300 includes one or more processors 302 (e.g., processing units of CPU(s)), one or more network interfaces 304, memory 306, and one or more communication buses 308 for interconnecting these components (sometimes called a chipset).
[0060] In some embodiments, the computing device 300 includes one or more input devices 310, such as a keyboard, a mouse, a voice command input unit or microphone, a touch screen display, a touch-sensitive input pad, a gesture-capturing camera, or other input buttons or controls, to facilitate user input. In some embodiments, the computing device 300 uses a microphone and speech recognition or a camera and gesture recognition to supplement or replace a keyboard. In some embodiments, the computing device 300 includes one or more output devices 312, such as one or more speakers and / or one or more visual displays (e.g., the display device 140), to enable presentation of user interfaces and display content.
[0061] The memory 306 includes high-speed random access memory, such as DRAM, SRAM, DDR RAM or other random access solid state memory devices; and optionally includes non-volatile memory, such as one or more magnetic disk storage devices, one or more optical disk storage devices, one or more flash memory devices, or one or more other non-volatile solid state storage devices. The memory 306 optionally includes one or more storage devices remotely located from the one or more processors 302. The memory 306, or alternatively the non-volatile memory within the memory 306, comprises a non-transitory computer readable storage medium. In some implementations, the memory 306, or the non-transitory computer readable storage medium of the memory 306, stores the following programs, modules, and data structures, or a subset or superset thereof:
[0062] An operating system 322, including procedures for handling various basic system services and for performing hardware dependent tasks;
[0063] A communication module 323 (e.g., a radio communication module) for connecting to and communicating with other network devices (e.g., a local network, such as a router providing Internet connectivity, networked storage devices, network routing devices, server systems, computer devices 130, ultrasound devices 200, and / or other connected devices, etc.) coupled to one or more communication networks via the network interface 304 (e.g., wired or wireless);
[0064] A user interface module 324 for enabling presentation of information at the computing device 300 or another device (e.g., a graphical user interface for presenting applications, widgets, websites and webpages thereof, games, audio and / or video content, text, etc.);
[0065] Application 350, for acquiring ultrasound data (e.g., imaging data) from a patient. In some embodiments, application 350 is for receiving data (e.g., ultrasound data, imaging data, etc.) acquired via ultrasound device 200. In some embodiments, application 350 is for controlling one or more components of ultrasound device 200 (e.g., a probe portion and / or transducers) and / or other connected devices (e.g., according to determining that data meets or does not meet certain conditions). In some embodiments, application 350 includes:
[0066] Acquisition module 352 for acquiring ultrasound data. In some embodiments, the ultrasound data includes imaging data acquired by an ultrasound probe. In some embodiments, acquisition module 352 activates transducers 220 (e.g., less than all transducers 220, a different subset(s) of transducers 220, all transducers 220, etc.) according to whether ultrasound data meets one or more conditions associated with one or more quality requirements. In some embodiments, acquisition module 352 causes ultrasound device 200 to activate transducers 220 (e.g., less than all transducers 220, a different subset(s) of transducers 220, all transducers 220, etc.) according to whether ultrasound data meets one or more conditions associated with one or more quality requirements;
[0067] Receiving module 354 for receiving ultrasound data. In some embodiments, the ultrasound data includes imaging data acquired by an ultrasound probe;
[0068] Sending module 356 for sending ultrasound data (e.g., imaging data) to other device(s) (e.g., server system, computer device 130, display device 140, ultrasound device 200, and / or other connected devices, etc.);
[0069] an analysis module 358 to analyze data (e.g., imaging data, power consumption data, and other data related to the acquisition process) (e.g., received by the ultrasound probe) to determine whether the data satisfies one or more conditions associated with quality requirements for the ultrasound scan. For example, in some embodiments, the one or more conditions include one or more of the following conditions: the imaging data includes one or more newly acquired images that satisfy one or more threshold quality scores, the imaging data includes one or more newly acquired images that correspond to one or more anatomical planes that match a desired anatomical plane of a target anatomical structure, the imaging data includes one or more newly acquired images that have one or more landmarks / features (or combinations of landmarks / features), the imaging data includes one or more newly acquired images that have features of a particular size, the imaging data supports a prediction that images that satisfy one or more requirements will be acquired in one or more upcoming image frames, the imaging data supports a prediction that a first change (e.g., a percentage increase or a number) in the number of transducers used will support an improvement in quality scores of images acquired in one or more upcoming image frames, and / or other similar conditions; and
[0070] a transducer control module 360 to activate (e.g., adjust, control, and / or otherwise modify one or more operations of the transducers) or cause the ultrasound device 200 to activate (e.g., via the transducer control module 260) a plurality of transducers 220 during a portion of the ultrasound scan based on determining that the ultrasound data satisfies (or does not satisfy) one or more quality requirements. For example, in some embodiments, the transducer control module 360 activates a first subset of the transducers 220 during a first portion of the ultrasound scan. In some embodiments, the transducer control module 360 activates a second subset of the transducers 220 during a second portion of the scan that follows the first portion of the scan that is different from the first subset of the transducers when the imaging data corresponding to the first portion of the scan satisfies (or does not satisfy) one or more quality requirements. In some embodiments, the transducer control module 360 controls one or more operational modes of the ultrasound device 200. For example, in some embodiments, the ultrasound device 200 is configured to operate in a low power mode. In the low power mode, the transducer control module 360 activates only a subset (e.g., 10%, 15%, 20%, etc.) of all available transducers 220 in the ultrasound device 200. In some embodiments, the ultrasound device 200 is configured to operate in a full power mode. In the full power mode, the transducer control module 360 activates all available transducers 220 to acquire high quality images; and
[0071] a database 380 including:
[0072] ultrasound scan data 382 (e.g., imaging data) acquired (e.g., detected, measured) by one or more ultrasound devices 200;
[0073] image quality requirement data 384. In some embodiments, the image quality requirement data 384 includes clinical requirements for determining the quality of an ultrasound image;
[0074] atlas 386. In some embodiments, the atlas 386 includes an anatomical structure of interest. In some embodiments, the atlas 386 includes a three-dimensional representation of an anatomical structure of interest (e.g., a hip, a heart, or a lung);
[0075] imaging control parameters 388. For example, in some embodiments, the imaging control parameters include one or more of: a number of transducers that are activated, a power consumption threshold of a probe, an imaging frame rate, a scan speed, a penetration depth, and other scan parameters that control the power consumption, heat generation rate, and / or processing load of a probe;
[0076] ultrasound scan data processing models 390 for processing ultrasound data. For example, in some embodiments, the ultrasound scan data processing models 390 are trained neural network models trained to determine whether an ultrasound image satisfies a quality requirement corresponding to a scan type, or trained to output an anatomical plane corresponding to an anatomical structure of an ultrasound image, or trained to predict, based on an ultrasound image sequence and its quality scores, whether a subsequent frame to be acquired by an ultrasound probe will contain certain anatomical structures and / or landmarks of interest; and
[0077] labeled images 392 (e.g., a database of images) including images used to train models for processing new ultrasound data and / or new images that have been processed or need to be processed. In some embodiments, the labeled images 392 are images of anatomical structures that have been labeled with their respective identifiers and relative positions.
[0078] Each of the above-described elements can be stored in one or more of the memory devices described herein and correspond to sets of instructions for performing the above-described functions. The above- described modules or programs need not be implemented as separate software programs, procedures, modules, or data structures and thus various subsets of these modules can be combined or otherwise rearranged in various embodiments. In some embodiments, the memory 306 optionally stores a subset of the modules and data structures described above. Furthermore, the memory 306 optionally stores additional modules and data structures not described above. In some embodiments, a subset of the programs, modules, and / or data stored in the memory 306 are stored on and / or executed by the ultrasound device 200.
[0079] Figure 4is a workflow for acquiring an ultrasound image in a device-agnostic guidance system according to some embodiments. In some embodiments, the ultrasound image is a medical ultrasound image for diagnostic purposes. In some embodiments, the workflow 400 is executed by one or more processors (e.g., CPU(s) 302) of a computing device in communication connection with an ultrasound probe. For example, in some embodiments, the computing device is a server or console (e.g., server, standalone computer, workstation, smartphone, tablet device, medical system) in communication with an ultrasound probe. In some embodiments, the computing device is a control unit integrated into the ultrasound probe. In some embodiments, the ultrasound probe is a handheld ultrasound probe or an ultrasound scanning system.
[0080] In some embodiments, the workflow 400 includes acquiring (402) a medical image, such as an ultrasound image. In some embodiments, the ultrasound image is an ultrasound image frame that meets one or more quality requirements for making a diagnosis and / or other conditions. An ultrasound examination is typically performed by placing (e.g., pressing) a portion of an ultrasound device (e.g., an ultrasound probe or scanner) on the surface of or inside a cavity of a patient’s body, adjacent to the region under study. An operator (e.g., a clinician) moves the ultrasound device around the region of the patient’s body until the operator finds a position and pose of the probe that results in an image of the anatomical structure of interest with sufficiently high quality. According to some embodiments, the operator uses a user interface of the ultrasound device to control the acquisition of the image. For example, the operator can use the user interface to control the position and pose of the probe, the orientation of the probe, the depth of the image, the gain of the image, the focus of the image, the zoom of the image, and / or other parameters of the image. Figure 5A and Figure 5B The workflow 400 is further illustrated by an example of acquiring a medical ultrasound image for diagnosing hip dysplasia according to some embodiments of the present disclosure. The ultrasound probe can acquire the image using one or more of its acquisition modalities, including a single (e.g., static) 2D image, an automatic scan producing a series of (static) images for different 2D planes (referred to herein as a “2D scan”), a continuous dynamic image (cine clip) as a real-time video capture of the scanned anatomical structure, or a single scan acquiring a 3D volume of images corresponding to multiple different 2D planes (referred to herein as a “3D scan”). The raw images 552, 554, 556, and 558 acquired in step 402 using one or more ultrasound devices are processed in respective sequences 510, 512, 514, and 516 as shown in Figure 5A
[0081] In some embodiments, workflow 400 includes applying (404) a style-removal process to reduce (e.g., eliminate) probe-specific characteristics. As used herein, the term probe-specific effects refers to effects in the resulting images that are different for the probe than for the probe. Some probe-specific effects are based on the model or manufacturer of the probe, while other probe-specific effects can arise from differences between two instances of the same model. Probe-specific characteristics (e.g., caused by probe-specific effects) can reduce the performance of an artificial intelligence (AI) system used to analyze those images. For example, a classifier trained using images acquired with a first probe (e.g., a probe from a first manufacturer) can achieve 95% performance. But when the same classifier is used on images acquired with a second probe (e.g., a probe from a second manufacturer different from the first manufacturer), the classifier can have 70% performance. A probe-agnostic algorithm (e.g., including a style-removal step) can provide a predictor (e.g., a classifier or other machine learning model) that does not degrade in performance (or the degradation in performance is reduced) when provided with images acquired using different probes. Figure 5A Different aspects of the style-removal process are illustrated.
[0082] Style-removal process Style-removal process 502 (e.g., using a style-removal algorithm) is applied to the raw images (e.g., raw images 552, 554, 556, and 558) acquired during step 402 (e.g., the image acquisition step). Style-removal process 502 maps the raw images into a common space 560. In some embodiments, common space 560 is common to ultrasound images (e.g., all ultrasound images) of the same anatomical region in a particular view, even ultrasound images acquired by probes having different characteristics (e.g., power, scaling factor, depth, frequency, or size of field of view). Style-removal process 502 helps minimize the probability that a guidance system works well only for images acquired with a first probe and not for images acquired with a different probe. In some embodiments, style-removal process 502 removes from the images characteristics attributable to a particular probe (e.g., noise artifacts unique to a particular probe), leaving in the images characteristics related to the anatomical region interrogated by the probe but not unique to any particular probe.
[0083] In some embodiments, the style-removal algorithm includes a machine learning algorithm that minimizes (e.g., explicitly minimizes) divergence between multiple batches of images (or vector representations of those images) acquired with different probes. For example, as illustrated in FIG. 5, style-removal process 502 is applied to raw images 552, 554, 556, and 558 to generate style-removed images 562, 564, 566, and 568, respectively. In some embodiments, style-removed images 562, 564, 566, and 568 are generated by a style-removal algorithm that is trained using a plurality of batches of images acquired with different probes. For example, style-removal algorithm 504 can be trained using batches of images acquired with different probes. In some embodiments, style-removal algorithm 504 is trained using batches of images acquired with different probes from the same model. In some embodiments, style-removal algorithm 504 is trained using batches of images acquired with different probes from different models. Figure 5AAs shown, the original image 552 has a sector shaped field of view 561, while the other three original images 554, 556, and 558 have rectangular fields of view. By minimizing the differences between the batch of four images to remove style information (e.g., the geometry of the field of view), the original image 552 is transformed (e.g., via data transformations such as geometric projection) to map onto a common space 560 with a matching field of view (e.g., a rectangular field of view) as the other images.
[0084] Similarly, the image 556 can be acquired using a second probe different from the first probe used to acquire one or more of the original images 552, 554, and 558. The image 556 includes noise (e.g., a speckled pattern of random noise or other random noise) that can be unique to the second probe or otherwise not present in one or more of the original images 552, 554, and 558. After the style removal process 502 in step 404 of the processing sequence 514, the noise 562 is removed from the original image 556 to minimize the differences between the batch of four original images. Thus, the image 556 is mapped into the common space 560 without the noise 562 that was originally present in the image 556. In some embodiments, the common space is where discrepancies in pixel distribution (e.g., due to the use of different ultrasound probes) are minimized.
[0085] In some embodiments, the original images can include variations or artifacts associated with anatomical features captured in the images. For example, the original images 554 and 558 can have imaging artifacts 564 associated with anatomical features. Style removal reduces (e.g., eliminates) variations such as the imaging artifacts 564 produced by the ultrasound probes.
[0086] Some examples of difference measures are KL divergence, Wasserstein distance, maximum mean discrepancy. In some embodiments, the style removal step includes using an algorithm that indirectly minimizes the differences between the probability distributions of different probes. For example, by applying computer vision, signal processing, image processing, or other feature extraction techniques, the noise introduced by different probes is removed.
[0087] Invariant feature In some embodiments, the invariant feature is a feature related to the object being imaged rather than the image itself. For example, an organ (e.g., a liver) can be scanned with two different probe scanners, each probe having different noise levels, scaling, and other operating parameters. The invariant feature can include one or more of the following: the actual size of the organ, the volume of the organ, the percentage of fat in the organ, the shape of the organ, the geometric profile of the organ. Depending on the operating parameters of the ultrasound probe (e.g., the scaling factor), the scanned organ can appear larger or smaller in the acquired ultrasound image. As a result, if the two-dimensional area of the scanned organ is determined by summing the number of pixels located within the boundary of the liver, a different number of pixels will be obtained from the different ultrasound images depending on the operating parameters (e.g., the number of pixels will be higher for the magnified image than for the demagnified image). Thus, the number of pixels is not an invariant feature.
[0088] The pixel spacing corresponds to the physical dimension associated with a pixel. The number of pixels divided by the pixel spacing is an invariant feature. For a magnified image, the pixel spacing can be 10 μιη, while in a demagnified image the pixel spacing can be 1 mm.
[0089] In some embodiments, segmenting the anatomical structure of interest (e.g., an organ, such as a liver) can be based at least in part on locating "bright" pixels in the acquired image (e.g., having a pixel intensity greater than a predefined threshold). However, for images acquired with different operating settings, the predefined threshold can be different (e.g., some pixels will be brighter than others, different operating settings can include operating at different frequencies or amplitudes).
[0090] An invariant feature that can be extracted from an image can include the intensity gradient of the image after the image has been appropriately processed (e.g., subjected to a style removal process) (e.g., by taking the first derivative of the intensity of neighboring or other neighboring pixels to determine the gradient and / or local minima in the image). In some embodiments, the image is processed using histogram equalization prior to any determination of the intensity gradient. Histogram equalization is an image processing method of contrast adjustment using the histogram of the image. In some embodiments, the intensity gradient is invariant to a linear transformation of the pixel intensity.
[0091] In addition to extracting information based on the intensity gradient of the original image, other image parameters can also be changed in the style removal process, such as the intensity range recorded in each pixel of the original image. For example, the range of original pixel values in the original image can be from 0 to 255. By rescaling these values to a smaller range, such as between 0 and 20, noise in the original image can be reduced (e.g., eliminated). One example where such intensity range rescaling can be useful is for detecting B-lines in lung scans. In some embodiments, the rescaled pixel value range (e.g., to a range other than 0 to 20) can be determined automatically using a machine learning method to maximize performance (e.g., performance of a predictor). In some embodiments, the predictor is more robust when trained with images after pre-processing (e.g., after style removal) that have pixel intensity values with a narrower range.
[0092] Public space In some embodiments, the public space can also be referred to as a latent space. In some embodiments, the latent space is a common shared space rather than an observed space. For example, consider a classifier that predicts a health condition (sick vs. healthy) based on a person’s height. When height is measured in different units (e.g., feet and centimeters, and 6 feet is different from 6 centimeters), height data (e.g., physical height) in the observed space can not be directly comparable. One example of a latent space is a space in which variables are standardized. For example, for each observed variable value (e.g., height) in each of two datasets (e.g., one dataset corresponds to measurements recorded in feet, and the other dataset corresponds to measurements recorded in centimeters, respectively), subtract the statistical mean (e.g., of the set of heights measured in the same unit) and divide the result by the standard deviation (of the set of heights measured in the same unit). Thus, the two datasets (e.g., containing measurements in feet and centimeters, respectively) can be processed independently, and the output data can be combined and analyzed together in the latent space. The above standardization process produces a standard score that indicates how many standard deviations above or below the mean a particular observation (e.g., of height) falls. For example, a standardization value of 2 indicates that the observation falls 2 standard deviations above the mean. This allows height data to be compared regardless of the units in which the original data was collected. In some embodiments, the latent space is not observable but can be computed (e.g., by converting measurement units).
[0093] In some embodiments, the differences between original ultrasound images acquired using different probes are not caused by using different units, but rather by different intensity patterns, noise, and quality associated with each original image. And the style removal process 502 removes those variations in intensity patterns, transforming the images so that the processed images are in the same latent space.
[0094] ReturnFigure 4 Workflow 400 includes identifying (406) anatomical structures from the image to which style removal has been applied. Figure 5A Different aspects of feature identification and segmentation are illustrated.
[0095] Feature identification and segmentation Feature identification and segmentation process 504 is applied to the pre-processed image that has undergone the style removal process during step 404. Feature identification and segmentation process (e.g., anatomical structure identification process) 504 identifies relevant anatomical structures present in common space 560. In some embodiments, a convolutional neural network for segmentation is used to identify anatomical regions of interest from the acquired image. In some embodiments, a template matching method is used to identify anatomical regions of interest from the acquired image. In some embodiments, manually defined or selected features (e.g., based on shape and texture of the anatomical parts to be identified) are defined for identifying anatomical regions of interest from the acquired image.
[0096] As a first non-limiting example, for the diagnosis of developmental dysplasia of the hip, the structures of interest to be identified from segmentation process 504 include the ilium, the acetabulum, and the femoral head. For the diagnosis of fatty liver, the organs of interest include the liver itself and the kidneys. In Figure 5A In particular, segmentation masks 570, 572, 574, and 576 each show a circular femoral head. In each of segmentation masks 570, 572, and 574, the two linear structures corresponding to the ilium and the acetabulum form an obtuse angle with each other, while being oriented slightly differently and showing different lengths of the ilium and the acetabulum. Segmentation mask 576 shows only a single linear structure slightly angled from the horizontal position.
[0097] Workflow 400 includes automatically extracting (408) anatomical related landmarks using the segmentation masks obtained from segmentation process 504 as input, without user intervention.
[0098] Landmark identification The output of segmentation / identification process 504 is provided to landmark identification process 506 for identifying anatomical landmarks of interest. Figure 5A and Figure 6Different aspects of landmark identification are shown. In some embodiments, landmarks are obtained through geometric analysis of the segmentation masks obtained from the segmentation / identification process 504. For example, in embodiments where the diagnostic target is hip dysplasia, the segmentation mask can be circular (e.g., the femoral head), while the mask for identifying B-lines in a lung image can be substantially linear. For example, when the segmentation / identification process 504 provides a circular segmentation mask as output, the landmark identification process 506 can generate a landmark 522 around the circumference of the circular segmentation marker. In some embodiments, landmarks can be obtained through machine learning and image processing methods. As shown, the output 508 of the identification process 506 is provided to a computer system 532. The computer system 532 also receives minimum requirements 534 for making a clinical diagnosis based on the ultrasound image and data from the 3D model 530. The output of the computer system 532 is provided to the user via a display 536. In some embodiments, the output of the computer system 532 includes presenting metrics of interest to the clinician, in accordance with a determination that all elements for making a diagnosis are present in the image. In some embodiments, in accordance with a determination that the image does not contain all elements for making a diagnosis, the computer system 532 then proceeds to an analysis using a convolutional neural network (CNN) described below. Figure 5B
[0099] In some embodiments, the ultrasound image acquired in step 402 is provided as input to a convolutional neural network (CNN) trained to output an integer grade (e.g., from 1 to 5) indicating a range of proportion of requirements that the image satisfies (optionally, with more important requirements contributing more to the grade if satisfied). In some embodiments, the computer system 532 provides guidance on how to position the probe to acquire a better image containing all necessary elements for computing the metrics.
[0100] The workflow 400 includes generating (410) a statistical model of the relevant anatomical portion. In some embodiments, the statistical model is generated based on landmarks identified from the landmark identification process 506. For example, the landmarks are used as priors in a statistical shape model (SSM) to refine the boundaries of the anatomical region of interest. A statistical shape model is a geometric model that describes a collection of semantically similar objects and represents the average shape of many three-dimensional objects and their variations in shape. In some embodiments, semantically similar objects are objects that have similar visual (e.g., color, shape, or texture) features and have similarities in "higher-level" information such as possible relationships between objects. In some embodiments, each shape in the training set of the statistical shape model can be represented by a consistent set of landmark points from one shape to the next (e.g., for a statistical shape model involving a hand, the fifth landmark point can always correspond to the tip of the thumb).
[0101] For example, in diagnosing hip dysplasia, an initial circular segmentation mask for the femoral head can have a diameter of unit length. Based on the output of the landmark identification process 506, the size of the circular segmentation mask can be adjusted to track the position of the landmark and have a refined circular perimeter that is greater than the unit length. The refined segmentation mask (or other output) can be used to generate a statistical model of the relevant anatomical portion that provides an optimized boundary for the organ and / or anatomical structure present in the image. In some embodiments, the landmark is provided to the statistical shape model generated in step 410 to determine whether the image collected in step 402 contains sufficient anatomical information to provide a meaningful clinical decision.
[0102] Once the boundary of the organ or anatomical structure is optimized, the workflow 400 retrieves (412) field-specific requirements related to the metric of interest and uses the statistical model generated in step 410 to determine (414) whether the image obtained in step 402 meets the relevant requirements. For hip dysplasia, the field-specific requirements related to the metric of interest can include the angle between the acetabulum and the ilium, as well as the coverage of the femoral head relative to the ilium. For example, lower coverage of the femoral head relative to the ilium corresponds to a more visible femoral head, increasing the metric of its diagnostic relevance. Similarly, the angle between the acetabulum and the ilium provides an indication of the plane at which the image was acquired, and whether the acquired image meets the relevant requirements for making a diagnosis. The relevance of various image features is provided as the field-specific requirements related to the metric of interest.
[0103] Training a CNN In some embodiments, the ultrasound image acquired in step 402 is used as input to a trained neural network, such as a convolutional neural network (CNN), that has been trained to determine whether the image meets all of the clinical requirements. The output of this network can be one of n categories. When n is 2, the trained neural network can provide a binary output, such as “conforms” or “does not conform.” A conforming image is one that meets the image quality requirements, while a non-conforming image is one that does not meet at least one of the clinical requirements. In some embodiments, the image acquired in step 402 is used as input to a convolutional neural network (CNN) that is trained to output a real number (e.g., from 0 to 1, 0 to 100%, etc.) that indicates the proportion (e.g., percentage) or degree to which the image meets the requirements. In some embodiments, the neural network is configured (e.g., trained) to provide an indication as to which individual requirements are met and which individual requirements are not met.
[0104] In some embodiments, the neural network is trained by a training dataset comprising a set of p images that have been determined to be compliant, e.g., by an expert, and a set of q images that have been labeled as non-compliant, e.g., by an expert. Each image is input to a convolutional neural network, which comprises a set of convolutional layers, optionally followed by pooling layers, batch normalization layers, dropout layers, dense layers, or activation layers. The output of the selected architecture is a vector of length n, where n is the number of classes to be identified. Each entry in the output vector is interpreted as a computed probability of belonging to each of the n classes. The output vector is then compared to a ground truth vector, which contains the actual probabilities of belonging to each of the n classes. A loss function is then used to compute the distance between the output vector and the ground truth vector. A common loss function is cross-entropy and its regularized versions; however, there are many loss functions that can be used in this process. The loss function is then used to compute updates to the weights. A common optimization method to compute such updates is a gradient-based optimization method, such as gradient descent and its variants. The neural network can be configured to output a real number representing the percentage of requirements currently satisfied by the acquired image. The process of computing the loss and updating the weights is performed iteratively until a predetermined number of iterations is completed, or until a convergence criterion is met. One possible implementation of this method is to create a set of binary classifiers as described above. One binary classifier is trained for each clinical requirement, and then the percentage of classifiers with positive output is computed.
[0105] According to a determination that the image obtained in step 402 satisfies the relevant requirements, workflow 400 computes (416) and displays a metric of interest (e.g., a clinical metric of interest) associated with the image obtained in step 402.
[0106] Quality metrics The methods and systems described herein also automatically compute metrics related to assessing whether a current ultrasound image satisfies minimum quality requirements. For example, in the case of hip dysplasia, such requirements are that the femoral head is fully visible, and that the iliac bone appears as a horizontal line in the image. Figure 5B An example ultrasound image 538 of a hip satisfying the clinical requirements for determining the presence of hip dysplasia is shown. As a non-limiting example, the clinical requirements for an ultrasound image of a hip for determining the presence of hip dysplasia include the presence of the labrum 540, the ischium 542, the middle of the femoral head 544, a flat and horizontal iliac bone 546, and the absence of motion artifacts.
[0107] As another example, an ultrasound echocardiogram Figure 4 Apical view (e.g., as Figure 5BThe clinical requirements for the view shown) are: (i) a view of the four chambers of the heart (left ventricle, right ventricle, left atrium, and right atrium), (ii) the apex of the left ventricle is at the top and center of the region, while the right ventricle is triangular and smaller in area, (iii) the myocardium and mitral valve leaflets should be visible, and (iv) the walls and septum of each chamber should be visible.
[0108] According to a determination that the image obtained in step 402 does not satisfy the relevant requirements, workflow 400 maps 418 one or more anatomically relevant structures to a 3D anatomical model. In some embodiments, workflow 400 maps one or more anatomically relevant structures to a 3D anatomical model regardless of whether the relevant requirements are satisfied. In some embodiments, the 3D anatomical model comprises a 3D template. Mapping the anatomically relevant structures to the 3D anatomical model can be used to estimate the plane currently displayed by the image collected in step 402 (e.g., the image plane associated with the image collected in step 402). In some embodiments, the mapping of the anatomically relevant structures to the 3D anatomical model comprises combining a medical expert with a segmentation network (e.g., from step 406 and / or process 504) to map the image obtained in step 402 into the template. The relative position of the image with respect to the template allows workflow 400 to guide 420 the user to collect a new image of higher quality that satisfies the clinical requirements for the diagnosis. In some embodiments, the feedback is provided to the user on a display of the computer system, or the feedback can be displayed on a local device (e.g., a phone or tablet) of the clinician or patient collecting the image. Workflow 400 also includes informing the user how and why the image collected in step 402 does not satisfy the one or more requirements (e.g., why the image does not have the desired quality). In some embodiments, workflow 400 includes receiving or generating a compiled set of characteristics associated with a high-quality image, and workflow 400 also displays information to the user. Unlike other methods that use a black-box approach (e.g., other deep learning models) that receive an image as input and output a decision or a suggested action without providing any reason (e.g., output a message such as “the image has a 55% probability of poor quality”), the methods and systems described herein produce an output that can be interpreted by a clinical user (e.g., system output: “the iliac bone is not presented horizontally,” or “the left ventricle is not fully visible”). Workflow 400 then repeats through step 402.
[0109] In some embodiments, the guidance provided in step 420 includes retrieving an atlas of the anatomical structure of interest. In some embodiments, the atlas can include a 3D model of the anatomical structure of interest, for example Figure 5Bthe 3D model 530 in the 3D model database 520. In some embodiments, the atlas is used to determine image quality and / or provide guidance to improve image quality. In some embodiments, given a 3D model of an anatomical structure of interest (e.g., the 3D model 530) and an ultrasound image (e.g., the image 538 in the ultrasound image database 510), Figure 5B
[0110] In some embodiments, the computing device determines the anatomical plane corresponding to the anatomical structure whose image is currently acquired by the ultrasound probe using a trained neural network. In some embodiments, the trained neural network is a trained CNN. In some embodiments, the trained neural network is configured to output a point in a 6-dimensional space that indicates three positions and three rotation angles relative to a 3D model of the anatomical structure of interest (e.g., x, y, and z coordinates and pitch, roll, and yaw angles). In some embodiments, the trained neural network can output the angles of the imaging plane in the x-y, y-z, and x-z directions, as well as the distance of the plane to the origin of the 3D model.
[0111] In some embodiments, the trained neural network provides values associated with a 6-dimensional vector as output, rather than giving a vector representing probabilities as output. In some embodiments, a loss function that is a weighted sum of squared errors or any other loss function suitable for real-valued vectors that are not limited to probabilities can also be provided as output.
[0112] In some embodiments, the computing device determines the anatomical plane corresponding to the anatomical structure whose image is currently acquired by the ultrasound probe by dividing the angle-distance space into discrete classes, and then using a trained neural network that outputs the class of the input image. In some embodiments, the computing device includes (or is communicatively connected to) a library of images that have been labeled with their relative positions (e.g., a database of images, such as the labeled images 392 in the database 380). The computing device identifies the image in the library of images that is “closest” to the input image. Here, closest refers to the minimum distance function between the input image and each other image in the library.
[0113] In some embodiments, the computing device computes (e.g., measures) the respective distance between the image acquired in the probe position space (e.g., a six-dimensional space indicating (x, y, z) position and rotation on the x-axis, y-axis, and z-axis relative to a 3D model of the anatomical structure of interest) and a predicted plane in the probe position space that would provide a better image, and based on the computation, determines a sequence of steps that would guide the user to acquire said better image. In some embodiments, the computing device causes the sequence of steps or instructions to be displayed on a display device communicatively connected with the ultrasound probe.
[0114] In some embodiments, instead of computing the distance between the current image in the probe position space and the predicted plane, the computing device classifies the current image into one of n possible categories. In some embodiments, changing the tilt of the ultrasound probe can result in a different plane of an organ (e.g., a heart) being imaged. Because the views of anatomical planes are well known in medical literature, a classifier can be trained using a convolutional neural network that identifies what the current view captured by the image acquired in step 402 corresponds to.
[0115] In some embodiments, the computer system used to implement workflow 400 can be a local computer system or a cloud-based computer system.
[0116] The methods and systems described herein do not assume or require an optimal image, or that the “optimal image” would be the same for every patient scanned by the ultrasound probe. The use of an “optimal image” also does not account for differences in anatomical position of organs in different people, or provide any underlying rationale for why a large (or small) deviation is observed in the input image. Accordingly, the methods and systems described herein do not include recording a deviation between the input image and the optimal image. The methods and systems described herein also do not involve training a predictor to estimate any such deviation. Instead, the methods and systems segment one or more relevant organs from the obtained image and provide a set of segmentation masks as output of the predictor. The methods and systems described herein identify the respective position(s) of the organs in the image, and are therefore robust to differences in position of organs across different people.
[0117] Instead of training a neural network based on a set of rules to obtain an output quantifying the confidence that a particular feature is not or is present in the input image, the methods and systems described herein utilize medical information to identify relevant anatomical structures in the input image. This approach eliminates the need to collect a training set for training a neural network to quantify the confidence. Furthermore, the methods and systems described herein are based on the identification of anatomical structures, rather than providing a confidence about the presence of some landmark.
[0118] Instead of determining the quality of a particular ultrasound image acquired by the output of the neural network in step 402, which can imply a non-calibration issue present in the neural network, the methods and systems described herein additionally use a probabilistic approach. In some embodiments, the calibrated system provides an output that corresponds to a true probability. Typically, a neural network working with imaging data performs two tasks: (1) estimating the probability of an event by finding patterns in the training examples, and (2) finding a mapping between the training examples and the estimated probabilities. Estimating the probability based only on the analysis of the pixels, as a traditional approach can do, requires hundreds of thousands of images, which is often unobtainable. Instead, it is possible to directly provide such probabilities during the training process by using probability labels instead of classification labels. Thus, in some embodiments, to achieve the goal, the process can involve using a network that finds a mapping between the input images and the probabilities using much less data by several orders of magnitude. In the non-limiting example of diagnosing hip dysplasia, the angle between the acetabular angle and the coverage of the femoral head provides an indication of the likelihood of the condition. This approach can allow for using much less data at training time before the system generates calibrated outputs.
[0119] Figure 6 An example of a multi-task approach for encoding domain-specific knowledge into machine learning algorithms is shown, in accordance with some embodiments. Figure 6 A workflow 600 for segmenting an anatomical portion of interest in an acquired image is shown. In some embodiments, the acquired image has undergone style removal and has been mapped into a common space (e.g., common space 560, as shown). Figure 5A Figure 6 The example shown in the middle illustrates a schematic thyroid ultrasound image 602. In some embodiments, the output of workflow 600 (e.g., automatically or without user input) provides a mask 606 that identifies a property of interest (e.g., an abnormality, a tumor) in a body region (e.g., an organ, a thyroid) using a machine learning algorithm 604. In some embodiments, instead of a mask 606 that identifies a particular anatomical region of interest (e.g., a mask that identifies only a single anatomical region of interest), workflow 600 additionally or alternatively provides machine learning algorithm 604 as a multi-task learning algorithm. The multi-task learning algorithm uses medical expert knowledge 610 to define a new mask 612 that contains other relevant anatomical structures 608, 614, 616, and 618 that are not anatomical regions of interest (e.g., tumors). In some embodiments, the multi-task learning algorithm generates an output with multiple features (e.g., does not provide an output that is a single determination or feature). In some embodiments, additional anatomical structures in a set of training images in a common space are manually annotated by an expert, and the manually annotated training images are used to train machine learning algorithm 604 to output new mask 612, thereby providing medical expert knowledge 610. In some embodiments, medical expert knowledge 610 is provided by an expert manually annotating a training set that includes a set of new masks that include multiple relevant anatomical structures in addition to the anatomical feature of interest. In some embodiments, medical expert knowledge can be provided via an atlas of anatomical structures or other medical literature as information about spatial relationships between possible relevant anatomical structures. In some embodiments, the atlas includes a three-dimensional representation of an anatomical structure of interest (e.g., a hip, a heart, or a lung).
[0120] In some embodiments, machine learning algorithm 604 can include a classifier that identifies a current view associated with image 602 and determines relevant anatomical structures that can be associated with the current view. In some embodiments, this multi-task approach allows machine learning algorithm 604 to learn a better representation of image 602 that can improve its performance.
[0121] Figure 7 An example of incorporating domain-specific knowledge in a machine learning model is shown in accordance with some embodiments. Figure 7A workflow 700 including training a machine learning model is shown in FIG. 7. A computer system receives (702) as input a set of labeled images for training. In some embodiments, the set of labeled images includes training images in which one or more (e.g., all) relevant anatomical structures present in the image are labeled. In some embodiments, the set of labeled images includes training images depicting one or more healthy anatomical structures and training images depicting one or more diseased anatomical structures. In some embodiments, the respective anatomical structures are labeled with position information (e.g., relative position information, including an angle relative to one or more reference planes or relative to another anatomical structure) in addition to identification information. In some embodiments, the training images of the respective anatomical structures are labeled with quality information related to how well the respective image is suited for diagnostic purposes (e.g., a measure of the visibility of the head of the femur, a measure of how level a line corresponding to the ilium is). In some embodiments, the quality information includes assigning a quality score to an ultrasound image. In some embodiments, the assigning of the quality score occurs automatically in response to the ultrasound image being acquired during a scan. In some embodiments, the image quality is assessed one frame at a time. In some embodiments, the image quality of a newly acquired image is assessed based on the newly acquired image and a sequence of one or more images acquired immediately prior to the newly acquired image. The computer system also receives (704) as input encoded information containing medical expert knowledge.
[0122] A typical machine learning model can receive a large set of training data as input. In some embodiments, a smaller set of data (e.g., a much smaller set of data) is sufficient for training a machine learning model when domain expert knowledge (e.g., medical expert knowledge) is incorporated during training. In some embodiments, the domain expert knowledge is encoded by providing target labels to the training images. In some embodiments, using 200 images that incorporate domain expert knowledge increases the accuracy of the output by an additional 15% (e.g., from 70% to 85%) compared to a neural network model trained using 200 images that do not incorporate domain expert knowledge. For example, a medical expert manually annotates a set of training images for relevant anatomical structures (e.g., all relevant anatomical structures) present in the training images. In some embodiments, the domain expert knowledge is encoded using a multi-task approach. For example, when multiple anatomically relevant structures are annotated in each of multiple training images, the multi-task approach allows for determining a probability distribution of the respective anatomically relevant structures appearing together. In some embodiments, the multi-task approach allows the network to learn an additional set of patterns. For example, training a model using the label“tumor” and the label“non-tumor” can cause the system to treat all“non-tumor” areas or segments similarly without discrimination. This can not be the correct approach most of the time because there can be other elements (liver, kidney, septum, etc.) in the image that are very different from each other. In some embodiments, by providing additional labels in the image that are not related to the feature of interest (e.g., tumor) (e.g., kidney, liver, or bladder), the network is able to learn additional patterns that can improve its performance. By incorporating domain expert knowledge into the training data, a novel way of learning the parameters of a machine learning model uses a smaller set of training data to provide comparable performance compared to typical machine learning models.
[0123] Using the labeled set of images from step 702 and the encoded medical expert knowledge from step 704, the computer system trains (706) a machine learning model using the expert knowledge as a prior. The output from the training of step 706 produces a predictor that analyzes (708) an input image from an ultrasound probe that obtains (710) the input image. In some embodiments, the input image is processed to reduce or remove probe-specific effects described above with reference to style removal before the input image is provided as input to the predictor in step 708. The predictor automatically segments one or more anatomical features of interest and identifies (712) relevant anatomical landmarks from the input image of step 710. The computer system displays (714) the output from step 712 on a display of the computer system, or can display the output on a local device (e.g., a cell phone or tablet) of the medical provider and / or patient.
[0124] Generally, machine learning utilizes a stored dataset (e.g., a training dataset) extracted from a stored training image set to automatically learn patterns from the stored image set. This machine learning approach can be different from a pattern recognition approach in which a pattern to be recognized is predefined, and a pattern recognition module provides an output as to whether the predefined pattern is present in a particular input image (e.g., a test image or an image to be processed). Using a machine learning approach, a pattern can be extracted or learned without the need to predefine the pattern. For example, a probabilistic graphical model (PGM) encodes a probability distribution in a complex domain, where a joint (multivariate) distribution involves a large number of random variables that interact with each other.
[0125] In some embodiments, workflow 400 is based on anatomical identification of structures of interest combined with prior obtained from medical experts, rather than on tissue features. Traditional learning methods generally do not allow incorporation of domain expert knowledge. For example, workflow 400 does not exclusively extract information from training examples, in contrast to fully data-driven approaches. Fully data-driven approaches do not allow incorporation of expert domain knowledge (e.g., annotations) in the training dataset, as they are limited to analyzing patterns in the data used as input to the system. In contrast, workflow 400 combines medical expert knowledge with a segmentation network to map images into templates (e.g., Figure 4 Step 418 in FIG. 4B, and is further described in Figure 10 Step 420 in FIG. 4B). Figure 4
[0126] Workflow 400 provides a series of features of the image as output, leaving the final diagnosis to the medical expert. Workflow 400 does not require the position and orientation of the scanning device, thereby eliminating the need for a dedicated external sensor to measure these attributes.
[0127] Workflow 400 addresses domain adaptation to ensure successful implementation of the quality detection tool. Images acquired with different probes produce images with different probability distributions. To successfully use any computer algorithm to automatically evaluate the quality of the images, workflow 400 maps the acquired images into a common space that is independent of the device. Workflow 400 also includes semantic image segmentation, which adapts to different variations in the shape and location of the anatomical structure of interest. Workflow 400 maps the segmented images into a 3D template to suggest new positioning of the scanning device, without the need for the position and orientation of the scanning device. Workflow 400 maps the ultrasound images into a 3D template, thereby allowing direct computation of movements in real time, without the need to store pre-computed movements of the ultrasound probe in a database.
[0128] In some embodiments, workflow 400 focuses on the quality of ultrasound images in a single view, rather than providing guidance for acquiring images from different views. Workflow 400 does not require ultrasound images from different planes, or the use of images from one plane as a reference for the rest of the images acquired in other planes.
[0129] The methods, systems, and apparatus described herein relate to an automated, device-agnostic approach for analyzing an anatomical structure of interest in a follow-up exam. In a follow-up exam, an ultrasound image of an anatomical structure of interest (e.g., an abnormality in a patient’s body, an abnormality in an organ, or a nodule in a thyroid) is obtained and then compared to a different (e.g., second) image of the same anatomical structure of interest obtained at a different point in time (e.g., a month ago, three months ago, a year ago). To make a meaningful comparison, the images can be acquired as similar as possible in terms of probe position and orientation (e.g., probe angle). For example, in some embodiments, to compare the size of anatomical features, the second image is obtained from approximately the same plane in a comparable view.
[0130] Since information about the position and angle of an ultrasound probe is not typically available (e.g., as a separate stored information associated with a particular acquired ultrasound image, or as metadata not associated with an acquired ultrasound image), performing a follow-up scan to track an anatomical structure of interest can be a challenging task. For example, the follow-up scan can be performed by different clinicians using different ultrasound systems (e.g., made by different manufacturers, are different models than the ultrasound system used to collect the first image) at different medical facilities. The methods, systems, and apparatus described herein relate to finding, characterizing, and comparing an anatomical structure of interest in two or more ultrasound images acquired at different points in time. The images can be acquired by different experts using different ultrasound probes and different configurations of the probe. In some embodiments, the different configurations correspond to different sets of operating parameters of the ultrasound probe (e.g., frequency, phase, duration, direction of movement, and / or transducer array type) and / or requirements for the ultrasound image (e.g., presence, location, size, and / or spatial relationship of a marker selected based on scan type, clarity of the image, and / or other hidden requirements based on machine learning or other image quality scoring methods).
[0131] Figure 8A workflow 800 for automatically comparing two or more diagnostic images (e.g., ultrasound images) of an anatomical structure of interest acquired at different points in time is shown in accordance with some embodiments. In some embodiments, the workflow 800 is performed by one or more processors (e.g., CPU(s) 302) of a computing device that is communicatively coupled to an ultrasound probe. For example, in some embodiments, the computing device is a server or console (e.g., a server, a standalone computer, a workstation, a smartphone, a tablet device, a medical system) that is in communication with an ultrasound probe. In some embodiments, the computing device is a control unit that is integrated into the ultrasound probe. In some embodiments, the ultrasound probe is a handheld ultrasound probe or an ultrasound scanning system.
[0132] A first ultrasound image is acquired using one or more of the acquisition modalities of the ultrasound probe, including a single (e.g., static) 2D image, an automatic scan (e.g., 2D scan) that produces a series of (static) images for different 2D planes, a continuous dynamic image that is a real-time video capture of the scanned anatomical structure, or a single scan (e.g., 3D scan) that acquires a 3D volume of images corresponding to a plurality of different 2D planes.
[0133] In some embodiments, the workflow 800 maps (802) the first ultrasound image, which includes an anatomical region of interest, into a device-independent common space, as described above with reference to the workflow 400. Images acquired with different scanning devices or with the same scanning device with different image acquisition parameters can have different qualities. The mapping corrects for these differences so that the differences in the probability distributions of the images are reduced (e.g., the probability distributions of the images are substantially the same). The identification of the relevant anatomical structure can be done manually or automatically. In some embodiments, this process is done automatically via a digital image processing algorithm or a machine learning algorithm.
[0134] The workflow 800 segments (804) the relevant anatomical structure from the images that have been mapped into the common space. As described above with reference to the workflow 400, the segmented anatomical structure from step 804 is mapped to a 3D template. The segmentation in step 804 can also be done on the original ultrasound images or on the ultrasound images that are mapped to the common shared space. In some embodiments, when using ultrasound images, an automatic segmentation algorithm is trained using domain adaptation techniques. In some embodiments, medical expert knowledge or information about the anatomy of the human body is incorporated to improve the quality of the segmentation. The segmentation algorithm can identify several anatomical structures that appear on the ultrasound images, even if not all of the anatomical structures are analyzed in subsequent tasks (e.g., similar to the multi-task scenario described in Figure 6
[0135] As Figure 9 As shown, the input image 902 obtained using the ultrasound probe is analyzed by machine learning, computer vision, or image processing algorithms. An example of the output of the segmentation process outlined in step 804 is a segmentation mask 904, which automatically identifies several anatomical structures visible on the input image 902. For example, an anatomical region of interest 906 is the region to be characterized, while the rest of the anatomical structures can be used to map the region 906 to a template 1000 during step 806, as described below.
[0136] In some embodiments, an image segmentation algorithm with domain adaptation capabilities can provide good performance regardless of the ultrasound device used to acquire the images. Domain adaptation is the ability to apply an algorithm trained in one or more "source domains" to a different (but related) "target domain." Domain adaptation is a subcategory of transfer learning. In domain adaptation, the source and target domains have the same feature space (but different distributions); in contrast, transfer learning includes cases where the feature space of the target domain is different from one or more source feature spaces. In some embodiments, the feature space includes various characteristics of the anatomical structure of interest. The different distributions associated with the feature space can be due to probe-specific effects. In other words, the source domain can be a machine learning training using images obtained from a first ultrasound probe, and the target domain is for images obtained using a second ultrasound probe different from the first ultrasound probe.
[0137] The workflow 800 maps (806) the anatomical region of interest 906 segmented from step 804 into a template. As shown, the location of the anatomical region of interest 906 is mapped into a common template 1000. The template 1000 has markers 1002, 1004, 1006, 1008, 1010, 1012, and 1014 at the locations of different anatomical regions. In some embodiments, the location includes a set of coordinates, or any other identifier that indicates the relative location of the anatomical region of interest 906 appearing on the image 902. Figure 10
[0138] In some embodiments, after identifying the anatomical region of interest, the workflow 800 includes storing the location information of the structure by placing markers on a template representing the anatomical structure of the organ. In some embodiments, the template is a simplified representation of the organ (e.g., the thyroid), and the markers are a set of coordinates indicating the location of the anatomical region of interest (e.g., a thyroid nodule). In some embodiments, the location information of the anatomical region of interest is used for a subsequent review examination (e.g., in step 814, described below) to quickly identify the anatomical region of interest in later acquired images.
[0139] Workflow 800 uses machine learning methods to characterize (808) the anatomical feature. In some embodiments, a similar machine learning method as described in step 410 in FIG. 4 is used. In some embodiments, as shown in FIG. 8, the anatomical region of interest 906 is characterized by a set of clinically relevant features. For example, a clinician can first select the anatomical region of interest 906 from the image to be analyzed 902. In some embodiments, the selection includes creating a first bounding box 1102 and a second bounding box 1104 around the structure to be analyzed. Then, step 808 includes providing a more refined segmentation mask 1106 and 1108 of the anatomical structure of interest within the bounding boxes 1102 and 1104. Figure 4 Figure 11 Step 808 also includes extracting a set of features that are clinically relevant to the anatomical region of interest 906 and representative of the anatomical region of interest 906.
[0140] Step 808 also includes extracting a set of features that are clinically relevant to the anatomical region of interest 906 and representative of the anatomical region of interest 906. Figure 12 One possible feature that can be extracted is shown. Step 808 includes selecting a first bounding box 1202 and a second bounding box 1204 for identifying the state of the anatomical region of interest 906 within the two bounding boxes. The refined segmentation masks corresponding to each of the first bounding box 1202 and the second bounding box 1204 identify the solid portion and the cystic (liquid) portion within the anatomical region of interest 906. Further statistical data can also be extracted, such as the percentage of solid tissue within the nodule, and / or the shape and size of the nodule.
[0141] After the location of the anatomical structure of interest is stored, a set of features is extracted from this anatomical structure of interest using a machine learning algorithm, which are clinically relevant to provide a diagnosis. For example, in the analysis of thyroid nodules, workflow 800 can include calculating a TIRADS score for the nodule. This score is calculated by characterizing the composition, echogenicity, margin, size, and echolucent lesions of the nodule. This characterization can be obtained by a combination of digital image processing, machine learning, and / or computer vision algorithms. All features related to the characterization of the anatomical structure of interest can also be stored on a local computer system or a cloud-based database for future use in review exams.
[0142] Steps 802, 804, 806, and 808 can be performed by a first clinician using a first ultrasound system at a first medical facility at a first time range. In some embodiments, a second ultrasound image is acquired at a second time range separate from the first time range (e.g., the second time range is one month later than the first time range, the second time range is three months later than the first time range, the second time range is six months later than the first time range, or the second time range is one year or more later than the first time range).
[0143] During the review session, new images of the anatomical region of interest are acquired. Workflow 800 maps (810) the second images into the common shared space. Then, workflow 800 segments (812) the relevant anatomical structures from the second images. Thereafter, workflow 800 performs a local search (814) around the locations of the anatomical structures in the template as determined in step 806. In some embodiments, workflow 800 segments the anatomical structures of interest by performing a local search from a window centered on the coordinates of the anatomical structures of interest derived from the first ultrasound images of the first exam. This segmentation process can be performed manually or automatically by a segmentation algorithm. For example, using the markers 1002, 1004, 1006, 1008, 1010, 1012, and 1014 on the template 1000, Figure 10 Step 814 includes performing a local search for the anatomical region of interest 906 on the second images (e.g., new images acquired at a second time range). In some embodiments, the clinician selects a bounding box around the anatomical structure of interest in a window centered on one or more of the markers 1002, 1004, 1006, 1008, 1010, 1012, and 1014.
[0144] Based on the search results from step 814, workflow 800 characterizes (816) the anatomical features using a machine learning method. In some embodiments, once the anatomical structure of interest is segmented in the second images, the same set of features used to characterize the anatomical structure of interest from the first images is used to characterize the anatomical structure of interest.
[0145] Steps 810, 812, 814, and 816 can be performed at different medical facilities using different ultrasound systems by different clinicians at a second time range different from the first time point.
[0146] Then, workflow 800 compares (818) the changes between the two images 818 based on the information characterizing the anatomical features from step 808 and from step 816. In some embodiments, based on the comparison results from step 818, workflow 800 also highlights the differences between the features extracted from the first and second images and displays the differences to the clinician.
[0147] Workflow 800 updates (820) a report and displays (822) guidance related to the medical condition involved in both acquired images. For example, the guidance can be the most up-to-date official clinical guidance that determines the likely next step. If the next step includes another review scan, workflow 800 also stores information related to future performance of data analysis (e.g., information about operational parameters (intensity, amplitude, frequency) of the ultrasound probe and / or information about positioning of the ultrasound probe). In some embodiments, workflow 800 includes generating or updating a report that indicates a clinically relevant difference between characteristics of the object of interest identified on the original scan and on the review scan. For example, the comparison can include a change in size, class, or any other relevant clinical feature of the anatomical structure of interest.
[0148] In some embodiments, workflow 800 does not base on convolutional neural networks, but rather uses probabilistic graphical models, image processing, and machine learning methods to predict TIRADS classification. In some embodiments, workflow 800 includes using a U-Net that produces a coarse estimate of the boundary of the anatomical structure of interest (e.g., a nodule) rather than using an ellipse fitting based segmentation scheme (e.g., nodule segmentation).
[0149] In contrast to tracking the anatomical object of interest in real-time or near real-time during a medical procedure (e.g., radiation therapy), workflow 800 includes tracking changes (or lack thereof) in the anatomical region of interest over a longer period of time. For the latter, real-time changes in the position, orientation, and characteristics of the anatomical region of interest are relatively small and are typically known at least to the clinician performing the ultrasound scan. To compare changes over a longer period of time, many factors can change significantly, such as the hardware used to track the anatomical region of interest, the position of the patient, the orientation and position of the probe. Much of the information about the image acquisition protocol of the first image can be unknown when the subsequent image is acquired at a later time, which can make the acquisition task more difficult.
[0150] Workflow 800 can be used to detect bone as well as soft tissue. Workflow 800 can automatically perform a difference analysis (e.g., between images initially acquired at an earlier time point and images acquired in a later review scan) for diagnostic purposes without direct user input or intervention (e.g., without manual analysis by a clinician). In some embodiments, workflow 800 is not intended to visualize images acquired at different time points. In some embodiments, workflow 800 does not include a registration step for images acquired at two different time points. Registration steps can be very difficult for ultrasound images because ultrasound images can be acquired from different angles and positions such that there can be no correspondence between the two images. Workflow 800 circumvents this problem by not including a registration step and instead maps the anatomical features of interest to a common space (e.g., step 806 in FIG. 8), which indicates the anatomical region in which the object or structure of interest lies. In other words, mapping two ultrasound images to a common coordinate space without intentionally registering the images (e.g., without comparing one image to another and determining a coordinate transformation based on the comparison) allows for an approximate location of the anatomical structure of interest to be determined even if the ultrasound images are not perfectly registered. Figure 8
[0151] In some embodiments, workflow 800 extracts features from the acquired ultrasound images and then compares the images in feature space, avoiding the need to store, render, and overlay the images. For example, workflow 800 does not include performing a 3D rendering of the anatomical structure of interest (e.g., a tumor) and then overlaying the two rendered lesions or regions of interest for comparison. Workflow 800 includes mapping the ultrasound images to a feature space, computing a set of metrics in this new space before comparing the extracted features in the feature space. Thus, in some embodiments, workflow 800 does not directly measure the volumetric difference in the anatomical structure of the features in image space. Instead, workflow 800 characterizes the object of interest (or anatomical structure) by a set of measurements (e.g., categories of TIRADS scores) and then displays the change in this measurement between two time points or time ranges. Workflow 800 does not need to access both ultrasound images to compare because this process can be done independently for each image and workflow 800 compares the extracted features.
[0152] In some embodiments, workflow 800 independently identifies the anatomical region of interest in the image set acquired at the first time range and the anatomical region of interest in the image set acquired at the second time range, rather than segmenting the anatomical region of interest in the second image set (acquired during a later time range) based on the segmentation of the first image set (obtained during an earlier time range).
[0153] Workflow 800 addresses the issue of field adaptation, which makes images acquired under different protocols or with different hardware potentially incompatible. In some embodiments, workflow 800 also eliminates the need to detect landmarks and reference points to compare the anatomical object of interest at two different points in time.
[0154] Figures 13A to 13C A flowchart illustrating a method 1300 of guiding an ultrasound probe (e.g., to achieve a diagnostic goal of tracking a structure, an anatomical structure, an imaging artifact associated with an anatomical structure using ultrasound imaging with the ultrasound probe over time) is shown, in accordance with some embodiments. In some embodiments, the ultrasound probe (e.g., ultrasound device 200) is a handheld ultrasound probe or an ultrasound scanner with an automated probe. In some embodiments, method 1300 is performed at a computing device (e.g., computing device 130 or computing device 300) that includes one or more processors (e.g., CPU(s) 302) and memory (e.g., memory 306). For example, in some embodiments, the computing device is a server or a console (e.g., a server, a standalone computer, a workstation, a smartphone, a tablet device, a medical system) that is in communication with the handheld ultrasound probe or the ultrasound scanning system. In some embodiments, the computing device is a control unit that is integrated into the handheld ultrasound probe or the ultrasound scanning system.
[0155] At a computer system comprising one or more processors and memory, and optionally an ultrasound device, the computer system obtains (1302) a first frame of a probed region acquired at a first time using a first probe device, and a first set of control parameters used to acquire the first frame. The computer system processes (1304) the first frame (e.g., the first frame is an image frame selected from a first plurality of image frames corresponding to different planes within a measurement volume; the measurement volume is obtained in a single scan; in some embodiments, machine learning (ML) is used to select the first frame from the plurality of image frames) to obtain a first set of attributes. In some embodiments, the term “attribute” as used herein broadly encompasses characteristics related to “features” as well as “structures,” e.g., brightness, image quality, and other parameters related to image acquisition. In some embodiments, the attributes of the first image refer to features or signals present in the first image, or can refer to attributes of anatomical structures. The features can refer to aspects of a content image (e.g., a content image that can be segmented from other content of the first frame). In some embodiments, processing the first frame includes segmenting one or more features (e.g., the one or more attributes include size of the features, characteristics of the features; whether the features are solid or liquid; geometric arrangement around the features) from the first frame after reducing one or more first probe-specific effects in the first frame based at least in part on the first set of control parameters. Processing the first frame can include running feature extraction; adjusting pixel size, eliminating noise profiles specific to the probe device; and rescaling the magnitude of values associated with the pixels (e.g., including determining the rescaled values by machine learning).
[0156] In accordance with a determination that the first frame contains a corresponding attribute present in a second frame, wherein the second frame is acquired at a time earlier than the first time using a second set of control parameters: the computer system displays (1318) information related to differences in the corresponding attribute on a user interface based on the first frame and the second frame. In some embodiments, the corresponding attribute corresponds to an anatomical structure or an artifact. In some embodiments, a plurality of anatomical structures are used in a multi-task machine learning process. For example, the information relates to changes in the structure (e.g., changes in size of the structure, location of the structure, and / or number of structures) between the first time and the earlier time. The information can also include displaying additional medical information based on the changes. In some embodiments, the medical information includes clinical guidance with suggested courses of action to take based on the changes. In some embodiments, the second set of control parameters is the same as the corresponding set of control parameters used to acquire the first frame. In some embodiments, the second set of control parameters is different from the corresponding set of control parameters used to acquire the first frame.
[0157] After reducing the one or more second-probe-specific effects in the second frame, the second frame is obtained by processing the one or more features segmented from the second frame. For example, reducing the second-probe-specific effects includes eliminating the second-probe-specific effects. The probe-specific effects are reduced by: mapping the first frame into a common shared space; running feature extraction; adjusting pixel size; and / or eliminating noise profiles specific to the second-probe device. In some embodiments, eliminating the noise profiles includes rescaling the magnitude of the pixel values, and the range of the rescaling is determined by a machine learning algorithm.
[0158] In some embodiments, reducing the one or more first-probe-specific effects in the first frame includes one or more elements selected from the group consisting of: pre-processing (1306) the first frame to account for the first set of control parameters and the second set of control parameters, and removing noise from the first frame.
[0159] In some embodiments, pre-processing the first frame to account for the first set of control parameters includes: adjusting the size of the pixels in the first frame (1308), and removing noise from the first frame includes: subtracting noise from the first frame using noise profiles specific to the first-probe device, and rescaling the magnitude of the values associated with respective pixels in the first frame.
[0160] In some embodiments, rescaling the magnitude of the values associated with respective pixels to a rescaled range includes: using machine learning to dynamically determine (1310) the rescaled range (e.g., after each additional frame is acquired). In some embodiments, the machine learning model is trained using frames that are of lower resolution and / or are otherwise device-agnostic denoised.
[0161] In some embodiments, segmenting the one or more features from the first frame includes: using the one or more spatial locations derived from the second frame to segment (1312) the one or more features from the first frame. In some embodiments, segmenting the first frame into the one or more features based on the spatial locations derived from the second frame includes: obtaining (1314) the spatial locations from a template. The method further includes: performing a local search in a window centered at coordinates corresponding to the spatial locations. For example, the coordinates of the one or more features are coordinates in the common shared space, and the subset of the segmentation results of the second modified frame includes the one or more features.
[0162] In some embodiments, displaying information related to the differences in the structure includes: displaying (1320) information about changes (e.g., movement, deformation, changes in size that exceed a predefined change) in one or more segmented features that are clinically relevant.
[0163] In some embodiments, processing the one or more features segmented from the second frame includes mapping (1322) the one or more features onto a template. In some embodiments, mapping the one or more features into the template includes mapping the first / second frame into a common shared space. In some embodiments, the template includes a 3-dimensional anatomical model. In some embodiments, the template is an array of feature representations, each feature representation having a respective value, and the template is not a 3D anatomical model.
[0164] In some embodiments, the template includes one or more markers identifying (1324) a location of one or more anatomical regions of interest, including the first anatomical region of interest in which the respective attribute is located. The location can include a set of coordinates, or any other identifier indicating a relative location of the anatomical region of interest within the acquired image.
[0165] In some embodiments, the second frame is acquired (1316) using a second probe device different from the first probe device. In some embodiments, the second probe device is different from the first probe device when the second probe device has a different model or is manufactured by a different manufacturer, has a different hardware configuration, or has a different set of control parameters. In some embodiments, images acquired with different probes or with different settings have different appearances. In some embodiments, images generated by different probes or different configurations of a probe are samples from different probability distributions over intensities of pixels / voxels.
[0166] In some embodiments, the computer system further includes providing (1326) guidance for positioning the first probe device at a different location from the location at which the first frame was acquired (e.g., displaying on a user interface, or providing audio guidance feedback) in accordance with a determination that the first frame fails to satisfy a first criterion related to the anatomical region of interest (e.g., the first criterion is a domain-specific requirement provided via a set of annotated training images, where the set of annotated training images includes a boundary of the anatomical region of interest). For example, the different location has the same x, y, z coordinates but different angles of rotation to obtain a third frame of the structure within the anatomical region of interest, where the third frame has a higher image quality metric compared to the first frame. In some embodiments, the image quality metric indicates a closeness of correspondence of the obtained frame to a desired image plane of the anatomical region of interest of the structure.
[0167] In some embodiments, the computer system further generates (1328) a statistical model based on the one or more features to determine a presence of the anatomical region of interest in the first frame. For example, the structure is in the anatomical region of interest; the one or more features include anatomically-related landmarks. In some embodiments, the landmarks are used as a prior in a statistical shape model to refine a boundary of the anatomical region of interest.
[0168] In some embodiments, the computer system further includes calculating (1330) a metric of interest associated with the structure in accordance with a determination that a first criterion related to the anatomical region of interest has been met. In some embodiments, the metric of interest relates to a size of the structure, the metric of interest relates to an angle associated with the structure (or between the structure and another element). In some embodiments, the method further includes displaying the calculated metric of interest.
[0169] In some embodiments, the computer system further processes (1332) a third frame by segmenting one or more features from the third frame after one or more probe-specific effects are reduced in the third frame, and displays information related to differences in the structure on the user interface based on the third frame and the second frame. In some embodiments, the third frame is obtained from approximately the same plane as the second frame and / or in a field of view comparable to the second frame (corresponding to one or more anatomical planes matching a desired anatomical plane of the target anatomical structure).
[0170] In some embodiments, the one or more features segmented from the first image include (1336) features that are invariant to the first probe device.
[0171] In some embodiments, the first frame is obtained by selecting (1338) the frame from a plurality of frames acquired in a scan. In some embodiments, the computer system further obtains (1340) a second frame of the probed region acquired using a second probe device, and a second set of control parameters used to acquire the second frame; and processes the second frame to obtain a second set of attributes of the second frame. In some embodiments, processing the second frame includes segmenting one or more features from the second frame after one or more second probe-specific effects are reduced in the second frame based at least in part on the second set of control parameters.
[0172] Figure 14A flowchart of a method 1400 of guiding an ultrasound probe (e.g., to achieve a diagnostic goal of tracking a structure, an anatomical structure, an imaging artifact associated with an anatomical structure using ultrasound imaging with an ultrasound probe over time) is shown, in accordance with some embodiments. In some embodiments, the ultrasound probe (e.g., ultrasound device 200) is a handheld ultrasound probe or an ultrasound scanner with an automated probe. In some embodiments, the method 1400 is performed at a computing device (e.g., computing device 130 or computing device 300) that includes one or more processors (e.g., CPU(s) 302) and memory (e.g., memory 306). For example, in some embodiments, the computing device is a server or a console (e.g., a server, a standalone computer, a workstation, a smartphone, a tablet device, a medical system) that is in communication with a handheld ultrasound probe or an ultrasound scanning system. In some embodiments, the computing device is a control unit that is integrated into a handheld ultrasound probe or an ultrasound scanning system.
[0173] At a computer system that includes one or more processors and memory and optionally an ultrasound device, the computer system displays (1402) a user interface for presenting an analysis of a second frame of a structure obtained using a second probe device. For example, the second frame is an image frame selected from a second plurality of image frames corresponding to different planes within a measurement volume. In some embodiments, the measurement volume is obtained by a single scan. In some embodiments, machine learning is used to select the second frame from the plurality of image frames, the second frame having a highest quality metric associated with the structure. In some embodiments, the second probe device is the same as the first probe device. In some embodiments, the second probe device is different from the first probe device. In some embodiments, the second frame of the structure is obtained at a second time after the first frame of the respective attribute. For example, the first frame is an image frame selected from a first plurality of image frames corresponding to different planes within a measurement volume. In some embodiments, the measurement volume is obtained by a single scan. In some embodiments, a machine learning algorithm is used to select the first frame from the plurality of image frames, the first frame having a highest quality metric associated with the structure. In some embodiments, the first frame is obtained at a first time using a first probe device. In some embodiments, the first probe device is the same as the second probe device. In some embodiments, the first probe device is different from the second probe device.
[0174] The computer system displays (1404), on the user interface, information about a difference in the structure based on the first frame and the second frame, wherein the difference is characterized by processing one or more features segmented from the second frame after reducing one or more probe-specific effects in the second frame, and using one or more spatial locations derived from the first frame.
[0175] In one aspect, an electronic device includes one or more processors; and a memory storing instructions that, when executed by the one or more processors, cause the electronic device to perform the functions described above. Figures 13A to 13C and Figure 14 The method described.
[0176] On the other hand, a non-transitory computer-readable storage medium stores program code instructions thereon, which, when executed by a processor, cause the processor to perform the above-mentioned [reference]... Figures 13A to 13C and Figure 14 The method described.
[0177] In another aspect, an electronic device includes an input unit configured to receive a first frame of a probed region acquired using a first probe device having a plurality of transducers and a first set of control parameters for acquiring the first frame. In some embodiments, the first probe device is a handheld ultrasound probe or an ultrasound scanner, the ultrasound probe being communicatively connected to the electronic device. The electronic device includes a memory unit configured to store data associated with a second frame acquired using a second set of control parameters. The electronic device includes a processing unit (e.g., the processing unit includes one or more processors and memory, a control circuitry system, an ASIC and / or other electrical and semiconductor controllers) configured to: acquire one or more attributes of the first frame; segment one or more features from the first frame, at least in part based on the first set of control parameters, after reducing one or more probe-specific effects in the first frame; the processing unit is further configured to process the one or more features segmented from the second frame after reducing one or more probe-specific effects in the second frame to obtain information relating to differences in structures recorded in the first and second frames; and a user interface configured to display information relating to differences in structures.
[0178] Although some of the accompanying figures illustrate certain logical levels in a specific order, the levels can be reordered regardless of the order, and other levels can be combined or separated. While some reorderings or other groupings are specifically mentioned, other reorderings will be obvious to those skilled in the art, and therefore the orderings and groupings presented herein are not an exhaustive list of alternatives. Furthermore, it should be recognized that these levels can be implemented in hardware, firmware, software, or any combination thereof.
[0179] It will also be understood that although the terms first, second, etc., are used in some cases herein to describe various elements, these elements should not be limited by these terms. These terms are used only to distinguish one element from another. For example, without departing from the scope of the various described embodiments, a first transducer may be referred to as a second transducer, and similarly, a second transducer may be referred to as a first transducer. Both the first sensor and the second sensor are sensors, but they are not the same type of sensor.
[0180] The terminology used in the description of the various implementations described herein is for the purpose of describing particular implementations only and is not intended to be limiting. As used in various descriptions of implementations and the appended claims, the singular forms "a," "an," and "the" are intended to include the plural forms as well, unless the context clearly indicates otherwise. It will also be understood that the term "and / or" as used herein refers to and encompasses any and all possible combinations of one or more of the associated listed items. It will be further understood that the terms "includes," "including," "comprises," and / or "comprising," when used in this specification, specify the presence of stated features, integers, steps, operations, elements, and / or components, but do not preclude the presence or addition of one or more other features, integers, steps, operations, elements, components, and / or groups thereof.
[0181] As used herein, the term "if' can be construed to mean "when" or "if," depending on the context. Similarly, the phrase "if it is determined" or "if [a stated condition or event] is detected" can be construed to mean "upon determining," or "in response to determining," or "upon detecting [the stated condition or event]," depending on the context.
[0182] The foregoing description of implementations has been presented for the purposes of illustration and description. It is not intended to be exhaustive or to limit the claims to the precise form disclosed. Many modifications and variations are possible in light of the above teaching. It is intended that the disclosed implementations be used in conjunction with all alternative embodiments and that the claimed implementations not be limited to only those implementations illustrated and described herein. Nothing in the foregoing description, and no element, component, or method step in the claims, is intended to be relied upon to convey an appreciation of the patent to the public.
Claims
1. A method comprising: at a computer system comprising one or more processors and memory: obtaining a first frame of a probed region acquired at a first time using a first probe device, and a first set of control parameters used to acquire the first frame; processing the first frame to obtain a first set of attributes of the first frame, wherein processing the first frame comprises segmenting one or more features from the first frame after reducing one or more first probe-specific effects in the first frame based at least in part on the first set of control parameters; in accordance with a determination that the first frame contains a respective attribute that is present in a second frame, wherein the second frame was acquired at a time earlier than the first time using a second set of control parameters: based on the first frame and the second frame, displaying information on a user interface relating to a difference in the respective attribute, wherein the second frame was obtained by processing the one or more features segmented from the second frame after reducing one or more second probe-specific effects in the second frame.
2. The method of claim 1, wherein, reducing the one or more first probe-specific effects in the first frame comprises one or more elements selected from the group consisting of: pre-processing the first frame to account for the first set of control parameters and the second set of control parameters, and removing noise from the first frame.
3. The method of claim 2, wherein, pre-processing the first frame to account for the first set of control parameters comprises adjusting a size of a pixel in the first frame, and removing noise from the first frame comprises subtracting noise from the first frame using a noise profile specific to the first probe device, and rescaling a magnitude of a value associated with the respective pixel in the first frame.
4. The method of claim 3, wherein, rescaling the magnitude of the value associated with the respective pixel to a rescaled range comprises dynamically determining the rescaled range using machine learning.
5. The method of claim 1, wherein, segmenting the one or more features from the first frame comprises segmenting the one or more features from the first frame using one or more spatial locations derived from the second frame.
6. The method of claim 5, wherein, segmenting the first frame into the one or more features based on the spatial locations derived from the second frame comprises obtaining the spatial locations from a template, and the method further comprises performing a local search in a window centered at coordinates corresponding to the spatial locations.
7. The method of claim 1, wherein, displaying information relating to a difference in the respective attribute comprises displaying information about a change in the one or more segmented features that is clinically relevant.
8. The method of claim 1, wherein, processing the one or more features segmented from the second frame comprises mapping the one or more features onto a template.
9. The method of claim 8, wherein, the template comprises one or more markers identifying locations of one or more anatomical regions of interest, the one or more anatomical regions of interest including a first anatomical region of interest in which the respective attribute is located.
10. The method of claim 1, wherein, the second frame was acquired using a second probe device different from the first probe device.
11. The method of claim 1, further comprising: in accordance with a determination that the first frame fails to satisfy a first criterion related to an anatomical region of interest: providing guidance to position the first probe device at a different location than the location at which the first frame was acquired to obtain a third frame of structures within the anatomical region of interest, the third frame having a higher image quality metric than the first frame.
12. The method of claim 11, further comprising: generating a statistical model based on the one or more features to determine a presence of an anatomical region of interest in the first frame.
13. The method of claim 11, further comprising: in accordance with a determination that a first criterion related to the anatomical region of interest has been met, computing a metric of interest associated with the structures.
14. The method of claim 11, further comprising: processing the third frame by segmenting one or more features from the third frame after reducing one or more probe-specific effects in the third frame; and displaying information related to differences in the structures on the user interface based on the third frame and the second frame.
15. The method of claim 14, wherein, the third frame is obtained from approximately the same plane and / or in a comparable field of view as the second frame.
16. The method of claim 1, wherein, the one or more features segmented from the first frame include features that are invariant to the first probe device.
17. The method of claim 1, wherein, the first frame is obtained by selecting a frame from a plurality of frames acquired in a scan.
18. The method of claim 1, further comprising: obtaining the second frame of the probed region acquired using a second probe device and the second set of control parameters used to acquire the second frame; and processing the second frame to obtain a second set of attributes of the second frame, wherein processing the second frame includes segmenting one or more features from the second frame after reducing one or more second probe-specific effects in the second frame based at least in part on the second set of control parameters.
19. A method comprising: at a computer system comprising one or more processors and memory: displaying a user interface for presenting an analysis of a second frame of a structure obtained using a second probe device, wherein the second frame of the structure is obtained at a second time after a first frame of a corresponding attribute is obtained at a first time using a first probe device; and displaying information related to differences in the structure on the user interface based on the first frame and the second frame, wherein the differences are characterized by processing one or more features segmented from the second frame after reducing one or more probe-specific effects in the second frame and using one or more spatial locations derived from the first frame.
20. An electronic device, comprising: one or more processors; and memory storing instructions that, when executed by the one or more processors, cause the electronic device to perform the method of any of claims 1-19.
21. A non-transitory computer-readable storage medium having program code instructions stored therein, the program code instructions, when executed by a processor, causing the processor to perform the method of any of claims 1-19.
22. An ultrasound probe, comprising: a plurality of transducers; one or more processors; and memory storing one or more programs, the one or more programs comprising instructions, which when executed by the one or more processors, cause the one or more processors to perform operations comprising: obtaining a first frame of a probed region acquired at a first time using a first probe device, and a first set of control parameters used to acquire the first frame; processing the first frame to obtain a first set of attributes of the first frame, wherein processing the first frame comprises segmenting one or more features from the first frame after reducing one or more first probe-specific effects in the first frame based at least in part on the first set of control parameters; in accordance with a determination that the first frame contains a respective attribute that exists in a second frame, wherein the second frame was acquired at a time earlier than the first time using a second set of control parameters: based on the first frame and the second frame, displaying information on a user interface relating to a difference in the respective attributes, wherein the second frame was obtained by processing the one or more features segmented from the second frame after reducing one or more second probe-specific effects in the second frame.
23. The ultrasound probe of claim 22, wherein, the memory stores instructions for performing the method of any one of claims 2-19.