Assisting users in performing medical ultrasound examinations

By using machine learning models to assist ultrasound examination systems, key anatomical features can be analyzed and highlighted in real time, solving the problem of ultrasound physicians having difficulty capturing high-quality images and improving examination efficiency and accuracy.

CN115666400BActive Publication Date: 2025-12-02KONINKLIJKE PHILIPS NV
View PDF 5 Cites 0 Cited by

Patent Information

Application Number
CN202180036923.0
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Priority Date
2020-04-16
Filing Date
2021-04-13
Publication Date
2025-12-02
Estimated Expiration
2041-04-13

AI Technical Summary

Technical Problem

In existing technologies, ultrasound physicians may not be able to capture sufficiently high-quality images, leading to diagnostic difficulties and wasted resources, especially in cases of repeated examinations.

Method used

The ultrasound examination system employs a machine learning model to analyze ultrasound images in real time and highlight anatomical features relevant to the examination, providing real-time guidance to the user through the processor and display.

Benefits of technology

It improves the imaging quality of ultrasound examinations, reduces the risk of diagnostic errors and duplicate examinations, especially for untrained users, and ensures that key features are not overlooked.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN115666400B_ABST
    Figure CN115666400B_ABST
Patent Text Reader

Abstract

A system for assisting a user in performing a medical ultrasound examination includes a memory including instruction data representing a set of instructions; a processor; and a display. The processor is configured to communicate with the memory and execute the set of instructions. When executed by the processor, the set of instructions causes the processor to perform the following operations: i) receive a real-time sequence of ultrasound images captured by an ultrasound probe during the medical ultrasound examination; ii) using a model trained using a machine learning process, taking image frames from the real-time sequence of ultrasound images as input and outputting the correlation between one or more image components in the image frames and a prediction of the medical ultrasound examination being performed; and iii) highlighting the image components predicted by the model and associated with the medical ultrasound examination to the user in real time on the display for further consideration by the user.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The disclosure herein relates to ultrasound imaging. Specifically, but not exclusively, the embodiments herein relate to systems and methods for recording ultrasound images. Background Technology

[0002] Ultrasound imaging (US) is used in a range of medical applications, such as fetal monitoring. Medical ultrasound imaging involves moving a probe containing an ultrasound transducer that generates high-frequency sound waves on the skin. These high-frequency sound waves travel through the tissue and are reflected from the inner surface (e.g., tissue boundaries). The reflected waves are detected and used to construct an image of the internal structures of interest.

[0003] Ultrasound imaging can be used to create two-dimensional or three-dimensional images. In a typical workflow, users (e.g., sonographers, radiologists, clinicians, or other medical professionals) can use two-dimensional imaging to locate anatomical features of interest. Once the features are located in two dimensions, the user can activate three-dimensional mode to capture three-dimensional images.

[0004] One objective of the embodiments described herein is to improve these methods. Summary of the Invention

[0005] Sonographers are trained to acquire image frames capturing normal features as well as those containing pathological features. These images are then used by radiologists for diagnosis. The fact that image capture and analysis can be performed by different people may result in the sonographer not capturing the image views needed by the radiologist. For example, an inexperienced user (sonographer) may not be able to capture sufficiently high-quality images (in terms of depth, focus, number of views, etc.) with relevant diagnostic content because they are not well-trained or do not have a grasp of the anatomical features and abnormalities that the radiologist considers most important. This can lead to a waste of time and resources, especially when repeat ultrasound examinations are necessary. Some of the embodiments described herein aim to improve this situation.

[0006] Therefore, according to a first aspect, there exists a system for assisting a user in performing a medical ultrasound examination, the system comprising a memory, a processor, and a display, the memory including instruction data representing a set of instructions. The processor is configured to communicate with the memory and execute the set of instructions. The set of instructions, when executed by the processor, causes the processor to perform the following operations: i) receive a real-time sequence of ultrasound images captured by an ultrasound probe during a medical ultrasound examination; ii) using a model trained using a machine learning process, taking image frames from the real-time sequence of ultrasound images as input, and outputting the correlation between one or more image components in the image frames and a prediction of the medical ultrasound examination being performed; and highlighting the image components predicted by the model and associated with the medical ultrasound examination on the display in real time for the user to further consider.

[0007] Therefore, this system can guide users in real time to image the anatomical features most relevant to the medical ultrasound examination being performed, as predicted by a model trained using a machine learning process. This helps ensure that users do not miss relevant features important to the diagnostic process.

[0008] According to the second aspect, there is a method for assisting a user in performing a medical ultrasound examination. The method includes: receiving a real-time sequence of ultrasound images captured by an ultrasound probe during the medical ultrasound examination; using a model trained using a machine learning process, taking image frames from the real-time sequence of ultrasound images as input, and outputting the correlation between one or more image components in the image frames and a prediction of the medical ultrasound examination being performed; and highlighting the image components predicted by the model and related to the medical ultrasound examination to the user in real-time on a display for further consideration by the user.

[0009] According to a third aspect, there is a method for training a model to assist a user in performing a medical ultrasound examination. The method includes: obtaining training data, the training data including: example ultrasound images; and ground truth annotations for each example ultrasound image, the ground truth annotations indicating the relevance of one or more image components in the corresponding example ultrasound image to the medical ultrasound examination; and training a model based on the training data to predict the relevance of one or more image components in the ultrasound images to the medical ultrasound examination.

[0010] According to a fourth aspect, a computer program product including a computer-readable medium is provided, the computer-readable medium having computer-readable code contained therein, the computer-readable code being configured such that, when executed by a suitable computer or processor, it causes the computer or processor to perform the methods described in the second aspect. Attached Figure Description

[0011] To better understand the embodiments and to more clearly illustrate how they are implemented and effective, reference will now be made to the accompanying drawings by way of example only, wherein:

[0012] Figure 1 Example systems according to some embodiments described herein are shown;

[0013] Figure 2 The illustration shows gaze tracking as used in some of the embodiments described herein;

[0014] Figure 3 The illustrations depict example methods for training and using neural network models according to some embodiments of this document;

[0015] Figure 4 The illustration shows example neural network architectures according to some embodiments of this document;

[0016] Figure 5 Example systems according to some embodiments herein are illustrated;

[0017] Figure 6 Example methods according to some embodiments herein are illustrated; and

[0018] Figure 7 An example system according to some embodiments described herein is illustrated. Detailed Implementation

[0019] As mentioned above, there is generally little feedback communication between radiologists and sonographers. Sonographers may follow standard imaging protocols, hoping that these protocols are broad enough to cover the diagnostic imaging needs of each patient. Radiologists rarely have a direct influence on how diagnostic imaging is performed on a particular patient.

[0020] In addition, quality assurance and ultrasound physician performance evaluation are often limited and are primarily achieved through certification / training and direct supervision.

[0021] In addition, the increasing use of portable ultrasound may lead to fewer trained users (such as first responders) performing ultrasound examinations on-site.

[0022] The purpose of these embodiments is to provide users of ultrasound imaging equipment with intelligent real-time image interpretation and guidance assistance to encourage improved image quality and facilitate the use of ultrasound imagers by inexperienced users.

[0023] Figure 1A system (e.g., apparatus) 100 for recording ultrasound images according to some embodiments thereof is illustrated. System 100 is used to record (e.g., acquire or capture) ultrasound images. System 100 may include a medical device or a part of a medical device, such as an ultrasound system.

[0024] refer to Figure 1 System 100 includes a processor 102 that controls the operation of system 100 and can implement the methods described herein. Processor 102 may include one or more processors, processing units, multi-core processors, or modules configured or programmed to control system 100 in the manner described herein. In a particular implementation, processor 102 may include multiple software and / or hardware modules, each configured to perform or be used to perform one or more steps of the methods described herein.

[0025] In short, the processor 102 of system 100 is configured to: i) receive a real-time sequence of ultrasound images captured by an ultrasound probe during a medical ultrasound examination; ii) use a model trained using a machine learning process, taking image frames from the real-time sequence of ultrasound images as input, and outputting the correlation between one or more image components in the image frames and a prediction of the medical ultrasound examination being performed; and iii) highlight the image components predicted by the model and associated with the medical ultrasound examination to the user in real time on a display for further consideration by the user.

[0026] In this way, communication between sonographers and radiologists / physicians can be minimized by incorporating a general knowledge base of features, regions, or image components that they believe are relevant to a given image and a given examination type from numerous radiologists within the model. As will be described in more detail below, this information can be embedded, for example, in a large deep learning network and projected / highlighted onto the current US view to assist the sonographer in real time during a medical ultrasound examination. Technically, this provides an improved way of obtaining ultrasound images, ensuring that all views relevant to the medical ultrasound examination (e.g., significant) are fully acquired. Therefore, this reduces the risk of misdiagnosis or the need for repeat examinations due to insufficient data. The system can also be used in remote imaging settings where no sonographer may be available to perform the medical examination (e.g., on-site or in emergency locations). Typically, the system can be used to guide untrained users to perform ultrasound examinations of acceptable quality.

[0027] In some embodiments, such as in Figure 1As illustrated, system 100 may further include memory 104, and memory 106 is configured to store program code that can be executed by processor 102 to perform the methods described herein. Alternatively or additionally, one or more memories 104 may be external to system 100 (e.g., separate or remote). For example, one or more memories 104 may be part of another device. Memory 106 may be used to store images, information, data, signals, and measurements acquired or generated by processor 102 of system 100 or by any interface, memory, or storage device external to system 100.

[0028] In some embodiments, such as Figure 1 As shown, system 100 may also include a transducer 108 for capturing ultrasound images. Alternatively or additionally, system 100 may receive (e.g., via a wired or wireless connection) a data stream of two-dimensional images captured using ultrasound transducer 108 external to system 100.

[0029] Transducer 108 may be formed from multiple transducer elements. Such transducer elements may be arranged to form an array of transducer elements. Transducer 108 may be included in a probe, such as a handheld probe, which may be held and moved across a patient's skin by a user (e.g., an sonographer, radiologist, or other clinician). Those skilled in the art will be familiar with the principles of ultrasound imaging, but in brief, an ultrasound transducer includes a piezoelectric crystal that can be used to both generate and detect / receive sound waves. Ultrasound generated by the ultrasound transducer enters the patient's body and is reflected from underlying tissue structures. The reflected waves (e.g., echoes) are detected by the transducer and processed by a computer to produce an ultrasound image of the underlying anatomical structures, also known as a sonograph.

[0030] In some embodiments, transducer 108 may include a matrix transducer capable of probing volume space.

[0031] In some embodiments, such as Figure 1 As shown, system 100 may also include at least one user interface, such as user display 106. Processor 102 may be configured to control user display 106 to display or present, for example, a real-time sequence of ultrasound images captured by an ultrasound probe. User display 106 may also be used to highlight image components predicted by a model to be relevant to a medical ultrasound examination to the user in real time. This may be in the form of an overlay (e.g., markers, colors, or other shadows displayed on a real-time sequence of images in a fully or partially transparent manner). User display 106 may include a touchscreen or application (e.g., on a tablet or smartphone), a display screen, a graphical user interface (GUI), or other visual presentation components.

[0032] Alternatively or additionally, at least one user display 106 may be external to the system 100 (i.e., separate or remote). For example, at least one user display 106 may be part of another device. In such an embodiment, the processor 102 may be configured to send instructions (e.g., via a wireless or wired connection) to the user display 106 external to the system 100 to trigger (e.g., induce or initiate) the external user display to display a real-time sequence of ultrasound images to the user and / or to highlight image components predicted by the model to be relevant to the medical ultrasound examination in real time.

[0033] It should be understood that Figure 1 Only the components necessary to illustrate this aspect of the disclosure are shown, and in actual implementation, system 100 may include additional components beyond those shown. For example, system 100 may include a battery or other devices for connecting system 100 to a mains power supply. In some embodiments, such as Figure 1 As shown, system 100 may also include a communication interface (or circuitry) for enabling system 100 to communicate with any interface, memory, and device inside or outside system 100, for example, via a wired or wireless network.

[0034] More specifically, users can include operators of the ultrasound probe, such as the person performing the ultrasound examination. Typically, this could be an ultrasound physician, radiologist, or other physician. Users can also be individuals not trained in medical imaging, such as clinicians or other users operating away from the medical environment, such as remotely or on-site. In such an example, the user could be guided by system 100 to consider or image portions of anatomical structures predicted to be important to the radiologist.

[0035] Medical ultrasound examinations can include any type of ultrasound examination. For example, a model can be trained to determine the relevant image components associated with any type of (pre-specified) medical ultrasound examination. Examples of medical ultrasound examinations to which the teachings herein can be applied include, but are not limited to, oncological examinations of injuries, neonatal examinations of fetuses, examinations to assess fractures, or any other type of ultrasound examination.

[0036] As described above, the instruction set enables the processor to, i) receive a real-time sequence of ultrasound images captured by an ultrasound probe during a medical ultrasound examination. In this sense, the processor can receive a sequence of ultrasound images from an ongoing ultrasound examination while the examination is in progress. Therefore, the sequence of ultrasound images can be considered as a real-time stream or feed of ultrasound images, as captured by the ultrasound probe.

[0037] Ultrasound image sequences can include two-dimensional (2D), three-dimensional, or any other dimensional ultrasound image sequences. The ultrasound image frames can include image components. In a 2D image frame, the image components are pixels; in a 3D image frame, the image components are voxels. Ultrasound image sequences can be any type of ultrasound image, such as B-mode images, Doppler ultrasound images, elastography images, or any other type or mode of ultrasound images.

[0038] In block ii), the processor uses a model trained using a machine learning process to take image frames from a real-time sequence of ultrasound images as input and outputs the correlation between one or more image components in the image frames and a prediction of the medical ultrasound examination being performed.

[0039] Those skilled in the art will be familiar with machine learning processes and models. However, in brief, the model can include any type of model that can be trained or has been trained, using machine learning processes to take an image (e.g., a medical image) as input and output the predicted relevance of one or more image components in the image frame to the medical ultrasound examination being performed. In some embodiments, the model can be trained according to method 700 as described below.

[0040] In some embodiments, the model may include a trained neural network, such as a trained F-net or a trained U-net. Those skilled in the art will be familiar with neural networks, but in short, a neural network is a supervised machine learning model that can be trained to predict the desired output for a given input data. The neural network is trained using training data containing example input data and the desired corresponding "correct" or ground-factual result. The neural network comprises multiple layers of neurons, each neuron representing a mathematical operation applied to the input data. The output of each layer in the neural network is fed into the next layer to produce an output. For each training data point, the weights associated with the neurons are adjusted until optimal weights are found to produce a prediction that reflects the corresponding real-world situation for the training example.

[0041] While this document describes examples including neural networks, it should be understood that the teachings herein are more generally applicable to any type of model that can be used or trained to output the predicted relevance of one or more image components in an image frame to an examination being performed by medical ultrasound. For example, in some embodiments, the model includes a supervised machine learning model. In some embodiments, the model includes a random forest model or a decision tree. The model may include a classification model or a regression model. Examples of both types are provided below. In other possible embodiments, the model may be trained using support vector regression or random forest regression or other nonlinear regressors. Those skilled in the art will be familiar with these other types of supervised machine learning models that can be trained to predict the desired output for a given input data.

[0042] In some embodiments, the trained model may have been trained using training data including: example ultrasound images; and ground truth annotations for each example ultrasound image, the ground truth annotations indicating the relevance of one or more image components in the corresponding example ultrasound image to a medical ultrasound examination. In this sense, the ground truth annotations represent examples of which pixels in the example ultrasound image are "correctly" predicted to be relevant to a medical ultrasound examination.

[0043] Those skilled in the art will be familiar with methods for training machine learning models using training data. Examples include gradient descent, backpropagation, and loss functions.

[0044] Typically, model training can be performed incrementally by training the model in the field, such as at the location where a radiologist examines ultrasound images. Once trained, the trained model can then be installed on another system, such as an ultrasound machine. In other examples, the model can be located on a remote server and accessed and updated dynamically. In still other examples, the model can be trained based on historical data.

[0045] In short, ground truth annotations can be obtained from one or more radiologists. In some examples, ground truth annotations can be specific to the type of ultrasound examination being performed. For example, a model can be trained for a specific type of medical ultrasound examination. In such embodiments, ground truth annotations can indicate image components or regions of the image associated with that type of medical ultrasound examination. In other examples, a model can be trained for more than one type of medical ultrasound examination. In such embodiments, ground truth annotations can indicate image components or regions of the image that may be more generally associated with many types of medical ultrasound examinations.

[0046] In some embodiments, the annotation may include image component level (e.g., pixel or voxel annotation for 2D and 3D images, respectively) annotations for the corresponding example ultrasound image frames, indicating the relevance or relative relevance of each image component (e.g., one or more pixels / one or more voxels) in the image frame. In some embodiments, this may be referred to as an annotation map or annotation heatmap.

[0047] As used herein, the term "relevance" can relate to the level of importance a radiologist would attribute to a different region or group of image components or image frames within the context of performing a medical ultrasound examination. For example, an image component or region of an image component can be labeled as relevant if the radiologist will be looking at (e.g., considering or examining) them as part of an ultrasound examination or wishes to investigate them further.

[0048] In some embodiments, the real-world annotation can be based on gaze tracking information obtained from the observing radiologist. Gaze tracking is a method of tracking a person's gaze focus on a 2D screen. Gaze techniques have been significantly improved by data-driven models (e.g., see the 2016 paper by Krafka et al. entitled "Eye Tracking for Everyone") and are accurate and cost-effective because they can be implemented using simple cameras and basic portable devices to implement computational nodes. Gaze tracking allows for the collection and annotation of relevant input features without requiring user-provided input. For example, annotations can be collected as part of a radiologist's routine examination of images.

[0049] This is Figure 2 As shown in the figure, Figure 2 The illustration shows different points observed by a radiologist in two ultrasound images of the chest cavity. In image 202, eye fixation data is represented as points 204 on the image, which the radiologist views while analyzing the image. In image 206, fixation data is represented as a set of circular regions 208 centered on the points observed by the radiologist. A model can be trained based on training data to predict any type of annotation for new (e.g., unseen) images. It should be understood that the model can also be trained to predict other outputs as described below.

[0050] Gaze information can be obtained by considering the location of the gaze and the duration of déjà vu on the image while the radiologist / physician is examining it. This "attention heatmap" is then used as input, with déjà vu serving as the relevance score. In other words, a heatmap can be generated based on gaze information, where the level of the heatmap is proportional to the amount of time the radiologist (or annotator) observes each specific area. The relevance score can indicate, for example, pathological lesions or hard-to-see areas, both of which are important for the sonographer to perform a correct scan.

[0051] In some embodiments, the model is trained to take image frames (e.g., image frames only) from a real-time sequence of ultrasound images as input. In other embodiments, the model may include additional input channels (e.g., employing additional inputs). The model may take, for example, an indication of the type of medical ultrasound examination being performed as input. In other words, the model can be trained to predict the relevance of pixels in an ultrasound image to different types of ultrasound examinations, depending on the indicated examination type.

[0052] In some embodiments, the model can be further trained to take an indication of the likelihood of a user overlooking a feature as input. For example, annotations can be graded based on the relevance and skill level of the sonographer required before the user might image a feature without prompting. This allows the system to provide highlighting relevant to the user's experience level (e.g., in block iii) and / or reduce the number of highlightings provided by offering only the most relevant highlighting that is most likely to be ignored by the user.

[0053] Other examples of inputs can include radiologist annotations, sonographer annotations, and ultrasound imaging settings, which can further improve the model's accuracy. Other possible input parameters include elastography or contrast images.

[0054] Turning now to the model's output, in some embodiments, the model's output (e.g., the correlation between one or more image components in an image frame and a prediction of the medical ultrasound examination being performed) may include a correlation value or score for each image component in the image frame. In such an example, the predicted correlation of one or more image components in an image frame may include a mapping of the correlation values ​​for each image component in the image frame.

[0055] In other examples, the predicted relevance of one or more image components in an image frame may include the relevance values ​​or scores of a subset of the image components in the image frame. For example, a subset of image components may have a relevance value higher than a predetermined threshold.

[0056] In some embodiments, one or more correlation thresholds may be used to group regions of an ultrasound image frame together. In such embodiments, the predicted correlation of one or more image components in an image frame may include one or more bounding boxes surrounding the image components, or regions of image components in the image frame that have a correlation value higher than a threshold (or between two thresholds). In some embodiments, the maximum correlation of each (or average) within each bounding box may be provided as the output of the model.

[0057] Using thresholds in this way allows, for example, highlighting only the most relevant areas to the user, such as only the top 10% of relevant image components. This enables sonographers to select specific thresholds to display only the most relevant annotations.

[0058] In some embodiments, the model can be trained to provide further outputs (e.g., with additional output channels). For example, the model can be further trained to output an indication of confidence associated with the relevance of predictions of one or more image components in an image frame.

[0059] In some examples, confidence levels can include (or reflect) the estimated accuracy of the relevance of predictions for one or more image components as model outputs. Alternatively, it can be a rating of how relevant each pixel / voxel is in an image frame, as determined by the model.

[0060] In other examples, the confidence level for the one or more image components may include a prediction of the priority a radiologist would place in studying the region containing that image component compared to other regions when performing the medical ultrasound examination. For example, the confidence level may include an estimate of the relative importance of different areas or regions of an image component (pixel / voxel) within an image frame.

[0061] In other examples, the model's output may include a combination of the options described above. For example, confidence may include a measure of both the predicted relevance and the estimated accuracy of the predicted relevance. In some embodiments, for each image component, the model may output (or the system may calculate from the model's output) the predicted relevance multiplied by the estimated accuracy for the predicted relevance.

[0062] In block iii), the processor then highlights the image components related to the medical ultrasound examination predicted by the model to the user in real time on the display for further consideration.

[0063] For example, the processor can send instructions to the display to provide markings, annotations, or overlays on the ultrasound frame to indicate relevant areas of the image frame to the user. This can guide the user to consider further or perform further imaging on anatomical areas that have already been highlighted to the user.

[0064] In some embodiments, block iii) includes enabling the processor to display the model's output to the user in the form of a heatmap superimposed on an ultrasound image frame. For example, the level of the heatmap can be based on the predicted correlation values ​​of the image components in the image frame. The level of the heatmap can be colored or highlighted according to the correlation values. In this way, the most relevant regions of the image frame can be effectively superimposed with "center"-style annotations so that the user can focus their imaging on them.

[0065] In other embodiments, the level of the heatmap can be based on, for example, the output confidence of image components in an image frame. The level of the heatmap can be colored according to the confidence. In this way, the most relevant regions of the image frame can be effectively overlaid with "center" style annotations so that the user can focus their imaging on them.

[0066] In other embodiments, the level of the heatmap can be based on the predicted relevance multiplied by the estimated accuracy of the predicted relevance, as described above.

[0067] In the implementation of output confidence, regions or areas of image frames predicted to include highly correlated (e.g., higher than threshold correlation) image components with high confidence (e.g., above threshold confidence) can be more significantly annotated compared to other regions, such as those including areas most likely to be associated with medical ultrasound examinations.

[0068] In other embodiments of the output confidence, regions or areas of image frames predicted to include image components with high correlation (e.g., above the threshold correlation) and low confidence (e.g., below the threshold confidence) can be more significantly annotated compared to other regions. In other words, displaying highly significant regions with low detection confidence can be useful because these regions may represent, for example, small lesions or other features that a radiologist might want to analyze in more detail. Ultrasound physicians can use this information to improve the imaging quality of these regions.

[0069] In other embodiments, in block iii), the processor may highlight image components predicted by the model to be relevant to a medical ultrasound examination to the user in any of the following ways (alone or in combination):

[0070] • Bounding boxes, circles, or polygons, where the color of the bounding box indicates the correlation of image components within the defined region.

[0071] • Bounding boxes, circles, or polygons, whose scale / size represents the correlation of image components within the defined region.

[0072] • A bounding box, circle, or polygon, whose thickness represents the correlation of image components within the defined region.

[0073] • A bounding box, circle, or polygon centered on the "centroid" of the image components in the region.

[0074] Values ​​on or near the bounding box

[0075] • Use a color map to color each image component according to the confidence level of the correlation.

[0076] • Transparency (alpha blending of each image component and its corresponding correlation as weights and color mappings).

[0077] • Highlighting can be dynamic in nature, thus highlighting relevant areas of image components to the user within a certain proximity of the mouse cursor (e.g., within a 2cm circle).

[0078] It should be understood that the processor can also repeat blocks ii) and iii) for multiple image frames in a real-time sequence of ultrasound images. For example, the processor can repeat blocks ii and iii in a continuous (e.g., real-time) manner. The processor can repeat blocks ii and iii for all images in a real-time sequence of ultrasound images. Thus, in some embodiments, images from an ultrasound examination can be overlaid with an annotation map as described above, which changes in real time as the user moves the ultrasound probe. In this way, real-time guidance is provided to the user for the most relevant area of ​​the imaged anatomical features during a medical ultrasound examination.

[0079] In some embodiments, a pixel-wise (or 3D voxel-wise) flow model can be used to link the predicted correlation of image components in an image frame to the predicted correlation of image components in another image frame in a real-time sequence of ultrasound images. This can provide a smooth overlay of highlighted correlated image components during ultrasound examination. The pixel-wise flow model can utilize temporal information from US imaging when available. In embodiments where the model outputs a graph describing the correlation of each image component in an ultrasound image frame, the model can correlate the predicted graph of subsequent ultrasound image frames in the ultrasound image sequence with regularization terms, such as smoothness and temporal consistency. The correlation graph can be extended from pixels to voxels, predicting the saliency level of each voxel in 3D space (see Girdhar et al.'s 2018 paper, "Detect-and-Track: Efficient Pose Estimation in Videos" (arXiv: 1712.09184v2)). The model architecture can be the same as a pixel correlation detection model (e.g., with the following). Figure 4 (The same as shown).

[0080] Figure 3 A system according to some embodiments herein is illustrated. In this embodiment, a model is trained to predict regions of an input ultrasound image associated with a specific type of medical ultrasound examination. Multiple radiologists 302 provide annotations for example ultrasound image frames 306, which are used as training data to train a neural network. In this embodiment, the annotations are in the form of an annotation graph 304, including bounding boxes indicating regions in each image frame that the annotating radiologists believe are associated with the type of medical examination being performed. The annotation graph is then used to train a neural network 308 to predict the true annotated graph 304 based on the input example ultrasound frame 306. In some versions of this embodiment, the neural network 308 may include the following regarding… Figure 4The neural network described is 400.

[0081] Once trained, the neural network 308 can be used for inference to take image frames 310 from a real-time sequence of ultrasound images as input (e.g., unseen images) and output a predicted annotation map 312 for that ultrasound image frame, indicating the relevance of each image component in the image frame to the medical ultrasound examination being performed. In this embodiment, the relevance is graded according to confidence, as described above. Therefore, the annotation map has the appearance of a heatmap, or multiple “center” style targets indicating the most relevant regions of the ultrasound frame.

[0082] The processor can then highlight image components predicted by the model to be relevant to the medical ultrasound examination to the user in real time on the display by overlaying the predicted annotation map onto the ultrasound image 314 on the display to create a thermal annotation on image 314 (as highlighted by the white circle 316 in image 314). This can be performed in real time, for example, by having the annotation map overlaid on the ultrasound image frame as the user views the ultrasound image. The user can thus use the predicted annotation map as a guide to the image areas that should be considered, for example, for further imaging.

[0083] Now turning to other embodiments, in one embodiment, such as Figure 4 As shown, the model includes a fully convolutional neural network FCN 400. FCN can be used to capture salient features in ultrasound images that are manually defined by the sonographer from many ultrasound images.

[0084] The FCN 400 takes ultrasound frames (e.g., 512x512 US images) as input through input layer 402. The network first stacks one or more layers of convolutions, batch normalization, and pooling (max pooling in this figure) 404, 406. Each such layer can have a different number of convolutional kernels, stride, normalization operations, and pooling kernel size. After each pooling, the size of the input image is scaled down proportionally to the size of the pooling kernel. Above these layers, one or more unpooling layers and deconvolution layers 408, 410 are added to upsample the intermediate-sized feature maps to the size of the original input image. Unpooling uses pixel interpolation to generate a larger image compared to pooling. The final output layer 412 outputs a correlation map of the original size, including correlation values ​​or scores for each image component in the original image.

[0085] The entire architecture can be trained end-to-end using backpropagation. The loss function of the last layer consists of a regression loss, which is the sum of all regression losses that regress the feature map of the last deconvolutional layer of each pixel to its relevance (or significance) score.

[0086] As described above, the training data used to train the FCN 400 can be performed on diagnostic images and can involve direct annotation by annotating radiologists / physicians (e.g., via mouse on the screen). Any standard annotation tool is acceptable, such as bounding boxes; circle centers and radii, polygons, clicks, etc. Furthermore, a saliency score (1-10, 10 being the most significant) can be assigned to each annotated region.

[0087] To reduce the amount of data required for model training, the image set may be limited to a specific medical ultrasound examination (e.g., a specific protocol or the anatomical region being imaged). One approach is to add an input to the network that specifies the type of medical ultrasound examination being performed (or a step in the protocol being performed). In other embodiments, a separate deep learning model 400 may be trained for each step in the medical protocol.

[0088] In some embodiments, such as Figure 4 As shown, the model can have a last fully connected layer the same size as the input image. In this way, the global per-pixel context can be determined (e.g., a map can be built based on image components, such as...). Figure 3 (As shown). Transposed convolutional layers (also known as deconvolutions) can achieve this upsampling. The loss function in such embodiments may include per-pixel regression. Standard data augmentation methods can be used to reduce bias in the dataset and correct for size.

[0089] Check the reasoning: Once trained, such as Figure 4 The FCN shown can be used during a routine ultrasound examination (the type of FCN trained on), where each image frame in the ultrasound image sequence can be evaluated by a trained model (e.g., provided as input to it) to output a corresponding annotation map, where each pixel has a relevant score. As described above, the relevant scores can then be visually presented on the original US images used for inference, such as... Figure 3 As shown in the image.

[0090] Turning to other embodiments, in some examples, a slice-by-slice approach can be employed, thereby dividing the ultrasound frame into smaller subframes. Subframes can be input into any embodiment of the model described herein. Once all slices have been processed, the outputs of the subframes can be combined to reconstruct the output of the entire image frame (e.g., a mapping of predicted relevance values ​​for each image component in the image frame). For example, the annotation map can be reconstructed in the same configuration as the original image subdivision. This can potentially reduce the amount of data required for training, as the relative positions of the annotations are not taken into account. A similar alternative would use a bounding box region detector, such as YOLO (see Redmon et al.'s 2018 paper entitled "YOLOv3: An Incremental Improvement"), which performs region detection at different scales.

[0091] In some embodiments, a classification model can be used. For example, in embodiments where the model outputs relevance values ​​or scores, the scores can be discretized. For instance, a softmax layer in a neural network can be used to output relevance values ​​as predefined levels (e.g., 0.1, 0.2, 0.3, ..., 1.0) that can be used as classification labels. This can reduce the computational effort required to implement the methods described herein.

[0092] Whether image components are correlated can be determined by the relative positions of different anatomical features in an image. For example, the correlation of annotations / features may be related to the presence of multiple features in a single view and / or their spatial context (relative position / or orientation). Such information can be used to measure annotation importance and can be embedded using methods trained with methods such as convolutional pose machines (e.g., see the paper "Convolutional Pose Machines" by Wei et al. (2016 IEEE Conference on Computer Vision and Pattern Recognition (CVPR), Las Vegas, NV, 2016, pp. 4724-4732)). In other words, in some embodiments, block ii) may include instructing the processor to consider the relative spatial context of different anatomical features to predict the correlation of one or more image components in an image frame. In some embodiments, the processor may use a convolutional pose machine trained to take into account the relative spatial context of different anatomical features to predict the correlation of one or more image components in an image frame.

[0093] For quality assurance, gaze tracking can be used during medical ultrasound examinations. For example, system 100 may also include a camera, and the processor may also monitor the user's gaze relative to display 106 during the ultrasound examination. In some embodiments, the user may be asked to look at locations in image frames that have been predicted by the model to be relevant (or above a certain relevant threshold). If, based on gaze information, it is determined that the user has not viewed an area of ​​the ultrasound frame predicted as relevant by the model, the processor may also be configured to provide visual assistance to the user to prompt them to look at the relevant area. In other words, the instruction set, when executed by the processor, may also cause the processor to: determine the user's gaze information. In block iii), the processor may then also be configured to: highlight one or more portions of the image frames that the gaze information indicates the user has not yet viewed on the display in real time.

[0094] In some embodiments, the underlying image frame may need to be visible to the user. To make the underlying image frame more visible to the user and to highlight it, in some embodiments, block iii) may also include instructing the processor to display a marker highlighting image components predicted by the model to be relevant to the medical ultrasound examination, and to remove or fade the marker after a predetermined time interval. In other words, a temporarily visible marker that disappears after a period of time can be used.

[0095] In another example, in block iii), the processor may display markers highlighting image components predicted by the model to be relevant to a medical ultrasound examination, wherein the markers are added or increased after predetermined time intervals. For example, if the user does not image the area (e.g., moves the transducer towards it), the markers may become brighter over time.

[0096] In another example, block iii may include a processor configured to use a heads-up display (HUD) or augmented reality to highlight components predicted by the model as relevant to a medical ultrasound examination.

[0097] In some embodiments, system 100 can be used to train users or ultrasound physicians. For example, the system can be used without an ultrasound machine. A type of medical imaging procedure and one or more images from an examination database can be presented to the ultrasound physician. For each image, the ultrasound physician may be asked to select a clinically important region (using a mouse or gaze as input), which can then be compared with the output of the model described herein.

[0098] System 100 can also be configured to identify typical areas missed by the user. This can be general or user-specific. For example, over time, a new (ultrasound physician-specific) model can be trained to highlight relevant image components missed by the user. This can encode anatomical features in the model and display them when they are detected. This is used to proactively guide the ultrasound physician while reducing clutter on the screen.

[0099] Turn now Figure 5 , Figure 5 An example embodiment of an ultrasound system 500 constructed according to the principles described herein is shown. Figure 5 One or more of the components shown may be included in a system configured to: i) receive a real-time sequence of ultrasound images captured by an ultrasound probe during a medical ultrasound examination; ii) use a model trained using a machine learning process, taking image frames from the real-time sequence of ultrasound images as input, and outputting the correlation between one or more image components in the image frames and a prediction of the medical ultrasound examination being performed; and iii) highlight the image components predicted by the model and associated with the medical ultrasound examination to the user in real time on a display for further consideration by the user.

[0100] For example, any of the aforementioned functions of processor 102 can be programmed into the processor of system 500, for example, via computer-executable instructions. In some examples, the functions of processor 102 can be provided by... Figure 5 One or more processing components are shown that implement and / or control, including, for example, an image processor 536.

[0101] exist Figure 5 In the ultrasound imaging system, the ultrasound probe 512 includes a transducer array 614 for transmitting ultrasound waves into a body region and receiving echo information in response to the transmitted waves. The transducer array 514 may be a matrix array comprising multiple transducer elements configured to be individually activated. In other embodiments, the transducer array 514 may comprise a one-dimensional linear array. The transducer array 514 is coupled to a microwave beamformer 516 in the probe 412, where the probe 512 can control the transmission and reception of signals by the transducer elements in the array. In the illustrated example, the microwave beamformer 516 is connected by a probe cable to a transmit / receive (T / R) switch 518, which switches between transmission and reception and protects the main beamformer 522 from high-energy transmitted signals. In some embodiments, the T / R switch 518 and other components of the system may be included in the transducer probe rather than in a separate ultrasound system base.

[0102] The transmission of an ultrasonic beam from transducer array 514, under the control of microwave beamformer 516, is indicated by a transmit controller 520 coupled to T / R switch 518 and beamformer 522, which receives input from, for example, a user-to-user interface or control panel 524. One of the functions controlled by transmit controller 520 is the direction in which the beam is steered. The beam can be steered vertically forward from the transducer array (perpendicular to the transducer array) or at different angles for a wider field of view. The partially beamformed signal generated by microwave beamformer 516 is coupled to beamformer 522, where the partially beamformed signals from individual facets of the transducer elements are combined into a fully beamformed signal.

[0103] The beamforming signal is coupled to signal processor 526. Signal processor 526 can process the received echo signal in various ways, such as bandpass filtering, decimation, I and Q component separation, and harmonic signal separation. Data generated by the different processing techniques employed by signal processor 526 can be used by a data processor to identify the internal structure and its parameters.

[0104] Processor 526 can also perform signal enhancement, such as ripple reduction, signal recombination, and noise cancellation. The processed signal can be coupled to B-mode processor 528, which can employ amplitude detection to image structures and tissues in the body. The signal generated by the B-mode processor is coupled to scan converter 530 and multiplane reformer 532. Scan converter 530 arranges the echo signals in a desired image format according to the spatial relationships in which the echo signals are received. For example, scan converter 530 can arrange the echo signals in a two-dimensional (2D) fan-shaped format. Multiplane reformer 532 is capable of converting echoes received from points in a common plane of a volumetric region of the body into an ultrasound image of that plane, as described in U.S. Patent 6,663,896 (Detmer). Volume plotter 534 converts the echo signals of a 3D dataset into a 3D image as a projection seen from a given reference point, for example, as described in U.S. Patent 6,530,885 (Entrekin et al.).

[0105] 2D or 3D images are coupled from the scan converter 530, the multi-plane reformer 532, and the volume plotter 534 to the image processor 536 for further enhancement, caching, and temporary storage for display on the image display 538.

[0106] The graphics processor 540 can generate graphic overlays for display alongside ultrasound images. These graphic overlays may contain, for example, mappings of correlation values ​​or fractions, as output by the model described herein.

[0107] The overlay may also contain other information, such as standard identification information, such as patient name, image date and time, imaging parameters, etc. The graphics processor may receive input from the user interface 524, such as a typed patient name. The user interface 524 may also receive input prompts to adjust settings and / or parameters used by the system 500. The user interface may also be coupled to a multiplane reformer 532 for selecting and controlling the display of multiple multiplane reformulated (MPR) images.

[0108] Those skilled in the art will understand that Figure 5 The embodiments shown are merely examples, and the ultrasound system 500 may also include features for... Figure 5 The additional components shown include, for example, a power supply or battery.

[0109] Turn now Figure 6 In some embodiments, there is a method 600 that assists a user in performing a medical ultrasound examination. This method may be performed, for example, by system 100 or system 700.

[0110] The method includes, in block 602, receiving a real-time sequence of ultrasound images captured by an ultrasound probe during a medical ultrasound examination. In block 604, the method includes taking image frames from the real-time sequence of ultrasound images as input, using a model trained using a machine learning process, and outputting the correlation between one or more image components in the image frames and a prediction of the medical ultrasound examination being performed. In block 606, the method includes highlighting, in real-time, the image components predicted by the model and related to the medical ultrasound examination to the user on a display for further consideration.

[0111] The above detailed description of the functionality of system 100 describes receiving a real-time sequence of ultrasound images captured by an ultrasound probe during a medical ultrasound examination, and the details therein are understood to also apply to block 602 of method 600. The above detailed description of the functionality of system 100 describes using a model trained using a machine learning process to take image frames from the real-time sequence of ultrasound images as input and output the correlation between one or more image components in the image frames and a prediction of the medical ultrasound examination being performed, and the details therein are understood to also apply to block 604 of method 600. The image components predicted by the model to be relevant to the medical ultrasound examination are highlighted to the user in real time on a display for further consideration. The above detailed description of the functionality of system 100, and the details therein are understood to also apply to block 606 of method 600.

[0112] Turning Figure 7In some embodiments, a method 700 for training a model to assist a user in performing a medical ultrasound examination also exists. In a first block 702, method 700 includes obtaining training data, which includes: example ultrasound images; and ground truth annotations for each example ultrasound image, the ground truth annotations indicating the relevance of one or more image components in the corresponding example ultrasound image to the medical ultrasound examination. In a second block 704, the method includes training a model to predict the relevance of the medical ultrasound examination to the image components in the ultrasound images based on the training data. The model described above with respect to system 100 discusses in detail how the model is trained in this manner, and the details therein will be understood to apply equally to model 700.

[0113] In another embodiment, a computer program product including a computer-readable medium having computer-readable code contained therein is provided, the computer-readable code being configured such that, when executed by a suitable computer or processor, it causes the computer or processor to perform one or more methods described herein.

[0114] Therefore, it should be recognized that this disclosure also applies to computer programs suitable for putting the embodiments into practice, particularly computer programs on or in a carrier. The program may be in the form of source code, object code, code between source code and object code (e.g., in a partially compiled form), or any other form suitable for use in implementing methods according to the embodiments described herein.

[0115] It should also be understood that such a program can have many different architectural designs. For example, the program code implementing the functionality of a method or system can be subdivided into one or more subroutines. The various ways in which functionality is distributed among these subroutines will be apparent to those skilled in the art. Subroutines can be stored together in an executable file to form a self-contained program. Such an executable file can include computer-executable instructions, such as processor instructions and / or interpreter instructions (e.g., Java interpreter instructions). Alternatively, one or more of the subroutines can be stored in at least one external library file and linked statically or dynamically (e.g., at runtime) with the main program. The main program includes at least one call to at least one of the subroutines. Subroutines may also include function calls to each other.

[0116] The carrier of a computer program can be any entity or device capable of carrying the program. For example, the carrier may include a data storage device, such as a ROM (e.g., a CD-ROM or semiconductor ROM), or a magnetic recording medium (e.g., a hard disk). Furthermore, the carrier can be a transmissible carrier, such as an electrical or optical signal, which can be transmitted via cable or optical fiber, or by radio or other means. When the program is implemented in such a signal, the carrier can consist of such a cable or other device or unit. Alternatively, the carrier can be an integrated circuit with an embedded program, the integrated circuit being adapted to perform the relevant method or be used in the implementation of the relevant method.

[0117] Those skilled in the art, through studying the accompanying drawings, the disclosure, and the claims, will be able to understand and implement variations of the disclosed embodiments. In the claims, the word "comprising" does not exclude other elements or steps, and the words "a" or "an" do not exclude multiple. A single processor or other unit can perform the functions of several items recited in the claims. Although specific measures are recited in dissimilar dependent claims, this does not imply that combinations of these measures cannot be advantageously used. Computer programs can be stored / distributed on suitable media such as optical storage media or solid-state media provided with or as part of other hardware, but can also be distributed in other forms such as via the Internet or other wired or wireless telecommunications systems. Any reference numerals in the claims should not be construed as limiting the scope.

Claims

1. A system for assisting a user in performing a medical ultrasound examination, the system comprising: A memory that includes instruction data, the instruction data representing a set of instructions; processor; as well as monitor; The processor is configured to communicate with the memory and to execute the set of instructions, wherein the set of instructions, when executed by the processor, causes the processor to perform the following operations: i) Receive a real-time sequence of ultrasound images captured by an ultrasound probe during the medical ultrasound examination. ii) Using a model trained via a machine learning process, image frames from the real-time sequence of ultrasound images are taken as input, and the model outputs the correlation between one or more image components in the image frames and a predicted medical ultrasound examination being performed, wherein the predicted correlation includes the importance level that a radiologist would attribute to the image component or different regions or groups of image components in the image frame within the context of performing the medical ultrasound examination; and iii) Based on the predicted relevance, the image components predicted by the model to be relevant to the medical ultrasound examination are highlighted to the user in real time on the display for the user to consider further.

2. The system according to claim 1, wherein, The processor also causes the processor to repeat blocks of multiple image frames in the real-time sequence of ultrasound images (ii) and (iii).

3. The system according to claim 1 or 2, wherein, The model is trained on training data using a machine learning process, the training data including: example ultrasound images; and ground truth annotations for each example ultrasound image, the ground truth annotations indicating the relevance of one or more image components in the corresponding example ultrasound image to the medical ultrasound examination.

4. The system according to claim 3, wherein, The real-world annotations are based on gaze tracking information obtained from observing a radiologist who analyzes the corresponding example ultrasound images for the purpose of the medical ultrasound examination.

5. The system according to claim 1 or 2, wherein, The model is also trained to output an indication of confidence associated with the relevance of the predictions to the one or more image components in the image frame.

6. The system according to claim 5, wherein, The confidence level reflects the accuracy of the estimate of the relevance of the predictions for one or more image components output by the model.

7. The system according to claim 5, wherein, The confidence level for the one or more image components includes a prediction of the priority a radiologist would place on studying the region containing that image component compared to other regions when performing the medical ultrasound examination.

8. The system according to claim 5, wherein, Block iii) includes enabling the processor to: The output of the model is displayed to the user in the form of a heatmap superimposed on the image frame, wherein the level of the heatmap is based on the output confidence of the image components in the image frame.

9. The system according to claim 1 or 2, wherein, Block ii) includes enabling the processor to consider the relative spatial background of different anatomical features to predict the correlation of the one or more image components in the image frame.

10. The system according to claim 1 or 2, wherein, The set of instructions, when executed by the processor, also causes the processor to: Determine the user's gaze information; and Block iii) further includes enabling the processor to: The gaze information is highlighted on the display in real time to indicate one or more portions of the image frame that the user has not yet viewed.

11. The system according to claim 1 or 2, wherein, Block iii) also includes enabling the processor to: Display markers highlight the image components predicted by the model to be relevant to the medical ultrasound examination, and wherein the markers are removed or faded out after a predetermined time interval; Display markers highlight image components predicted by the model to be relevant to the medical ultrasound examination, wherein the markers are added or their salience is increased after predetermined time intervals; and / or Augmented reality is used to highlight the image components that the model predicts are relevant to the medical ultrasound examination.

12. The system according to claim 1 or 2, wherein, The set of instructions, when executed by the processor, also causes the processor to: A pixel-by-pixel flow model is used to link the predicted correlation of image components in the image frame to the predicted correlation of image components in another image frame in the real-time sequence of ultrasound images.

13. A method for assisting a user in performing a medical ultrasound examination, the method comprising: During the medical ultrasound examination, a real-time sequence of ultrasound images captured by an ultrasound probe is received. The model, trained using a machine learning process, takes image frames from the real-time sequence of ultrasound images as input and outputs the correlation between one or more image components in the image frames and a predicted outcome of the medical ultrasound examination being performed. This predicted correlation includes the level of importance a radiologist would attribute to the image component or different regions or groups of image components within the image frame in the context of performing the medical ultrasound examination. Based on the predicted relevance, image components predicted by the model to be relevant to the medical ultrasound examination are highlighted to the user in real time on the display for further consideration.

14. A method for training a model to assist a user in performing a medical ultrasound examination, the method comprising: Training data is obtained, comprising: example ultrasound images; and ground truth annotations for each example ultrasound image, the ground truth annotations indicating the relevance of one or more image components in the corresponding example ultrasound image to the medical ultrasound examination; and The model is trained based on the training data to predict the correlation between one or more image components in an ultrasound image and the medical ultrasound examination, wherein the predicted correlation includes the level of importance that a radiologist would attribute to the image component or different regions or groups of image components in the ultrasound image in the context of performing the medical ultrasound examination.

15. A computer program product comprising a computer-readable medium having computer-readable code embodied therein, the computer-readable code being configured such that, when executed by a suitable computer or processor, it causes the computer or processor to perform the method according to claim 13 or 14.

Citation Information

Patent Citations

  • Spatially compounded three dimensional ultrasonic images

    US6530885B1

  • Delayed release aspirin for vascular obstruction prophylaxis

    US6663896B1

  • Ultrasound image recognition systems and methods utilizing an artificial intelligence network

    US20170262982A1

  • Training a neural network model

    US20190156204A1

  • Adaptive ultrasound scanning

    WO2019201726A1