Automatic analysis of pathology by ultrasound imaging
Automatic classification and pathological evaluation of ultrasound images through machine learning models, solving the professional dependence and high cost problems of aortic valve stenosis detection in the prior art, achieving early recognition and improving detection accuracy, and is suitable for bedside ultrasound equipment.
Patent Information
- Application Number
- CN202510130309.0
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Priority Date
- 2024-02-05
- Filing Date
- 2025-02-05
- Publication Date
- 2025-08-05
AI Technical Summary
The existing echocardiography methods depend on professional skills, are expensive and difficult to identify asymptomatic AS in early stages when detecting aortic valve stenosis (AS), resulting in many patients being diagnosed only in the late stage of the disease development. Conventional testing relies on high-quality image acquisition, and sub-critical images are difficult to be effectively utilized.
Machine learning models are used to automatically classify and evaluate multiple ultrasound images, including automatic labeling of training data and pathological severity estimation, and image analysis is performed in multiple cardiac cycles using sliding window technology, combining standard and non-standard views to improve the accuracy of pathological detection and early recognition capabilities.
It realizes automated, fast and low-cost detection of pathology such as aortic valve stenosis, can identify AS in the early stage, reduce dependence on professional skills, improves the accuracy of detection and the ability to cover sub-critical images, and is suitable for bedside ultrasound equipment.
Smart Images

Figure CN120420006A_ABST
Abstract
Description
Background Art
[0001] Aortic stenosis (AS) and other pathologies diagnosable by ultrasound imaging are a significant public health problem, affecting over 12.6 million adults and resulting in an estimated 102,700 deaths annually worldwide. Recently, there has been a surge of interest in the early identification of AS, and evidence suggests that many patients may not receive appropriate treatment. This interest has fueled research into novel approaches to identifying AS and other pathologies, yet little is known about how to improve the identification and treatment of AS. Comprehensive, population-based transthoracic echocardiographic (TTE) screening approaches would be prohibitively expensive. Automated interpretation of limited echocardiographic datasets is an attractive alternative for disease detection, particularly with the advent of point-of-care ultrasound equipment. Barriers to automated detection of AS (and other pathologies) include the complex nature of these diagnoses, the need to integrate information from multiple images within a single study, and the challenges posed by datasets often being unlabeled in routine clinical care. Furthermore, even routine echocardiographic assessment of AS is highly complex, often requiring the use of all standard imaging modalities and multiple scan views and windows. Therefore, the results of conventional detection of AS are highly dependent on the skills of the ultrasound operator, especially in making measurements of two-dimensional structures and generating optimized pulsed and continuous-wave Doppler signals.
[0002] Another major difficulty in detecting AS is its long asymptomatic period. During this period, the disease progresses, often unbeknownst to the patient. Given the current level of expertise required to diagnose AS, most patients first learn of their condition after undergoing a comprehensive echocardiographic study, often following referral from their primary care physician, as symptoms have begun to develop. The dramatic result is that many patients go undiagnosed simply because they exhibit no symptoms. Furthermore, the diagnosis is often made late in the disease process. Symptomatic, severe aortic stenosis is associated with a high mortality rate, as high as 50% within 1 year, and the incidence is likely to increase with the aging population. Therefore, improved methods and systems for the early detection of AS and other pathologies are needed. Summary of the Invention
[0003] In one aspect, a method for ultrasound imaging is described herein. In some aspects, the method includes acquiring a plurality of ultrasound images of at least a portion of an organ of a subject using an ultrasound imaging system. In some aspects, the plurality of ultrasound images includes images captured across at least a portion of at least one complete cardiac cycle.
[0004] In some aspects, the method includes processing the acquired plurality of ultrasound images to automatically classify a pathology of the subject, including providing the acquired plurality of ultrasound images as input to a trained machine learning model. In some aspects, the method includes outputting an indication of the condition of the subject.
[0005] In some aspects, the output is based at least in part on the output of a trained machine learning model. In some aspects, the output includes: (i) an indication of the presence or absence of a pathology in the subject; and / or (ii) a confidence estimate for the indication of (i).
[0006] In some aspects, the machine learning model is trained by a training method that does not include computer vision analysis of tissue landmarks and / or manual labeling of training data with cardiac phase information. In some aspects, the indication of (i) further includes an estimate of the severity of the pathology. In some aspects, the pathology is aortic stenosis and the organ is the subject's heart.
[0007] In some aspects, the method does not include visually detecting heart valve closure within the acquired plurality of images. In some aspects, automatic classification is performed on at least a subset of the plurality of ultrasound images as secondary images. In some aspects, the training method includes individually evaluating the prediction accuracy of each of a plurality of discrete views included in at least a subset of the training data.
[0008] In some aspects, the plurality of discrete views includes standard views and non-standard views. In some aspects, the one or more video clips include a plurality of acquired ultrasound images. In some aspects, the processing includes sliding a window over a plurality of frames of the one or more video clips to select frames that provide an improved confidence level for the indication of (i) compared to an indication based on an indication generated using a stationary window that includes a single cardiac cycle. In some aspects, the frames of the one or more video clips include frames acquired from more than one cardiac cycle. In some aspects, the sliding window slides over frames from more than one cardiac cycle.
[0009] In some aspects, each of the one or more video clips is associated with one or more discrete views. In some aspects, for each of the one or more video clips, estimating the severity of the pathology includes: determining a cardiac cycle of a subject for the video clip; sliding a window of a size based on the determined cardiac cycle over the clip; and / or calculating a confidence distribution for the pathology and / or an associated likelihood of successful severity classification.
[0010] In some aspects, estimating the severity of the pathology includes extracting windows associated with one or more video clips that meet a threshold for a likelihood of successful severity classification. In some aspects, estimating the severity of the pathology includes combining the extracted windows from two or more of the one or more discrete views to obtain an increased likelihood of successful severity classification compared to a classification based on a single one of the discrete views.
[0011] In some aspects, any method steps described herein may be repeated for all views supported by the ultrasound imaging system.In some aspects, the ultrasound imaging system supports 30 or more views.
[0012] In some aspects, estimating the severity of the pathology further comprises: sorting windows of all clips associated with a particular view based on a likelihood of successful severity classification to obtain a subset containing fewer than all windows associated with the particular view, the subset comprising three or more windows associated with the particular view having a highest likelihood of successful severity classification; and after combining the extracted windows, re-sorting windows of all combined clips based on a likelihood of successful severity classification to obtain a combination of windows having a highest global likelihood of successful severity classification.
[0013] In some aspects, the output includes an instruction to a user of the ultrasound system to acquire more images of the subject. In some aspects, the instruction to acquire more images is provided based at least in part on detecting at least a threshold likelihood of the presence of pathology in one or more of the plurality of acquired ultrasound images. In some aspects, the instruction to acquire more images is provided based at least in part on detecting less than a threshold confidence level in the presence or absence of pathology.
[0014] In some aspects, the acquired plurality of ultrasound images are two-dimensional ultrasound images. In another aspect, described herein is a non-transitory computer-readable medium storing instructions that, when executed by a processor of a computer, cause the computer to perform any of the methods described herein.
[0015] In another aspect, an ultrasound imaging system is described herein. In some aspects, the ultrasound imaging system includes an ultrasound imaging probe and a computing system. In some aspects, any of the ultrasound systems described herein can be configured to perform any of the methods described herein. In some aspects, the ultrasound imaging system can also include any of the non-transitory computer-readable storage media described herein.
[0016] Another aspect of the present disclosure provides a system comprising one or more computer processors and a computer memory coupled thereto, wherein the computer memory comprises machine executable code that, when executed by the one or more computer processors, implements any of the methods described above or elsewhere herein.
[0017] Further aspects and advantages of the present disclosure will become readily apparent to those skilled in the art through the detailed description that follows, wherein only illustrative aspects of the present disclosure are shown and described. As will be appreciated, the present disclosure is capable of other and different aspects, and its several details are capable of modification in various obvious respects, all without departing from the present disclosure. Therefore, the drawings and description are to be regarded as illustrative in nature, and not restrictive.
[0018] Incorporated by reference
[0019] All publications, patents, and patent applications mentioned in this specification are herein incorporated by reference to the same extent as if each individual publication, patent, or patent application was specifically and individually indicated to be incorporated by reference. To the extent that publications and patents or patent applications incorporated by reference conflict with the disclosure contained in this specification, this specification is intended to supersede and / or take precedence over any such conflicting material. BRIEF DESCRIPTION OF THE DRAWINGS
[0020] The novel features of the present invention are set forth with particularity in the appended claims. A better understanding of the features and advantages of the present invention will be obtained by reference to the following detailed description which sets forth illustrative aspects utilizing the principles of the invention and the accompanying drawings (also referred to herein as "Figures" and "FIG."), in which:
[0021] Figure 1A An example workflow for automatic classification of pathologies according to the methods and / or using the systems described herein is shown.
[0022] Figure 1B An alternative example workflow for automatic classification of pathologies, such as aortic stenosis, according to the methods and / or systems described herein is shown.
[0023] Figure 2 An example model architecture is shown that can be used to implement aspects of the methods and systems described herein.
[0024] Figure 3 A detailed schematic workflow of an example implementation of the methods and systems described herein is shown.
[0025] Figure 4An example ultrasound image of a parasternal long-axis view of the heart of a healthy subject is shown. This view shows a normal aortic valve, left ventricle, and left atrium, mitral valve, etc., with thin valve leaflets and no calcification.
[0026] Figure 5 An example ultrasound image focused on a stenotic aortic valve is shown. It has high diagnostic value for aortic stenosis but does not contain all the anatomical targets of a standard canonical parasternal long axis view. A clip selector that is not trained to select images missing those features or that is not trained to select images with pathology indications may reject such an image and not process it. The methods and systems described herein can process such clips, for example, using the methods described herein for classifying pathologies such as aortic stenosis.
[0027] Figure 6 A scatter plot showing the empirical density distribution of aortic stenosis severity in measurement space according to example implementations of the methods and systems described herein is shown. Figure 3 This is achieved by the example workflow shown in .
[0028] Figure 7 Shown (left) is an example frame (frame #5) from an example video clip of multiple ultrasound images including a parasternal long axis (PLAX) view of the heart and (right) shows the use of Figure 3 An example workflow implementation is shown in Figure 1. A diagram of the predicted cardiac cycle coordinates determined by the model.
[0029] Figure 8 Shown (left) is an example frame (frame #5) from an example video clip of multiple ultrasound images including a parasternal short-axis aortic valve (PSAX-AoV) view of the heart and (right) shows the use of Figure 3 An example workflow implementation is shown in Figure 1. A diagram of the predicted cardiac cycle coordinates determined by the model.
[0030] Figure 9 Shown by Figure 3 Scatter plot of the observed errors in aortic stenosis classification for the example model implemented with the workflow shown in Figure 4.
[0031] Figure 10 Shows the use of Figure 3Figure 1. Data density versus quality score for an example binary prediction task for an example model implemented by the workflow shown in Figure 2. Data density helps model the expected performance of the example model on the binary prediction task based on the quality score (QS). Using maximum likelihood, a model logistic regression classifier is fit to the binary problem (classifying whether we made a successful binary prediction of none / mild vs. moderate / severe) using a single feature as input: the quality score. Assume that when the quality score is low, the model reduces to 50% of random chance, while achieving perfect accuracy (100%) when the quality score is high. Furthermore, the classifier has only 2 free parameters, the scaling factor and the intercept.
[0032] Figure 11 Shows the Figure 10 The results of fitting the data.
[0033] Figure 12 A computer system is shown that is programmed or otherwise configured to implement the methods provided herein. DETAILED DESCRIPTION
[0034] Although various aspects of the present invention have been shown and described herein, it will be apparent to those skilled in the art that these aspects are provided by way of example only. Many variations, changes, and substitutions may occur to those skilled in the art without departing from the present invention. It will be understood that various alternatives to the various aspects of the present invention described herein may be employed.
[0035] Whenever the term "at least," "greater than," or "greater than or equal to" precedes the first value in a series of two or more values, the term "at least," "greater than," or "greater than or equal to" applies to every value in the series. For example, greater than or equal to 1, 2, or 3 is equivalent to greater than or equal to 1, greater than or equal to 2, or greater than or equal to 3.
[0036] When the term "not greater than," "less than," or "less than or equal to" precedes the first value in a series of two or more values, the term "not greater than," "less than," or "less than or equal to" applies to every value in the series. For example, less than or equal to 3, 2, or 1 is equivalent to less than or equal to 3, less than or equal to 2, or less than or equal to 1.
[0037] Certain inventive aspects herein contemplate numerical ranges. When a range is present, the range includes the range endpoints. In addition, each subrange and the values within the range are present as if explicitly written out. The terms "about" or "approximately" can mean within an acceptable error range for a particular value, which will depend in part on how the value is measured or determined, such as limitations of the measurement system. For example, according to practice in the art, "about" can mean within 1 or more than 1 standard deviation. Alternatively, "about" can mean a range of up to 20%, up to 10%, up to 5%, or up to 1% of a given value. Where a specific value is described in the application and claims, unless otherwise indicated, it can be assumed that the term "about" means within an acceptable error range for the specific value.
[0038] Diagnostic image quality
[0039] A particular challenge in ultrasound medical imaging is accurately determining what probe posture or movement will result in a clinical or diagnostic quality image. As used herein, image quality (e.g., diagnostic quality or clinical quality) can be used to refer to one or more aspects of image quality. In some aspects, image quality refers to an image that can be viewed by a trained expert or machine learning tool in a manner that identifies anatomical structures and can make a diagnostic interpretation. In some aspects, image quality refers to an image in which the target is displayed in a clear and well-defined manner, for example, with minimal extraneous noise or clutter, the grayscale display shows subtle changes in tissue type and texture, the blood flow signal is clear and distinct, the frame rate is high, thereby providing an accurate depiction of tissue or blood flow motion, the boundaries between tissue type or blood flow and blood vessels or other structures are well resolved, ultrasonic artifacts such as rastering and sidelobes are minimized, acoustic noise is absent, the place where measurements are taken in the image is obvious and clear, or any combination thereof, depending on the nature of the ultrasound examination. In some aspects, image quality refers to an image that contains the necessary anatomical targets to represent a standard diagnostic view. For example, an apical four-chamber view of the heart should show the apex of the heart, the left and right ventricles, the myocardium, the mitral and tricuspid valves, the left and right atria, and the atrial septum. As another example, a long-axis view of the carotid artery at the bifurcation should show the common, external, and carotid arteries, as well as the carotid bulb. In some aspects, image quality refers to an image in which a disease condition, abnormality, or pathology is well visualized. For example, medical images can be labeled by a cardiologist, radiologist, or other healthcare professional based on whether they are considered to have a disease condition, abnormality, or pathology that is well visualized, and then used to train a machine learning algorithm to distinguish between images based on image quality.
[0040] In some respects, image quality means the presence of some combination of these aforementioned characteristics. Effective navigation guidance will need to be provided to ensure that captured ultrasound images meet the combination of these image quality characteristics necessary to produce an overall clinical or diagnostic quality image, as patient presentation can pose challenges in ultrasound imaging, making it difficult to obtain high-resolution, low-noise images. For example, when attempting to assess blood flow in the kidneys of an obese individual, obtaining a sufficiently strong Doppler signal can be particularly challenging, as the kidneys are located so deeply beneath the adipose tissue. In chronic smokers, lung disease can make it difficult to obtain high-quality cardiac images. These conditions are very common, and in such cases, image quality can mean that the image may be suboptimal in terms of noise and resolution, yet still provide sufficient information for diagnosis. Similarly, patient presentation and pathology can make it impossible to obtain a view that shows all anatomical components of a standard, normative image. For example, a patient with technically challenging cardiac disease may make it impossible to obtain an apical four-chamber view with all four chambers well defined. However, if an image, for example, clearly shows the left ventricle, it can be considered a quality image, as many key diagnostic conclusions can only be drawn from this image.
[0041] In some aspects, the anatomical views used in the present disclosure include one or more of a probe position or window, an imaging plane, and a visualized area or structure. Examples of probe positions or windows include parasternal, apical, subcostal, and suprasternal notches. Examples of imaging planes include long axis (LAX), short axis (SAX), and four chambers (4C). Examples of visualized areas or structures include two chambers, aortic valve, mitral valve, etc. For example, the anatomical views may include parasternal long axis (LV inflow / outflow), RV inflow + / - RV outflow, parasternal short axis (aortic valve level, mitral valve level, papillary muscle level, apical LV level), apical four chambers, apical five chambers, apical two chambers, apical three chambers, subcostal four chamber views, subcostal short axis and long axis, suprasternal long axis (aortic arch), and suprasternal short axis (aortic arch).
[0042] Substandard image
[0043] As used herein, a "substandard" image may be an off-axis view, an incomplete anatomical view shown, and / or a non-canonical view. In some cases, a substandard image may be substandard due to low image quality. In some aspects, the quality may be low due to inexperienced users, incorrect ultrasound device setup, and / or patient size (such as narrow rib spaces or obesity). In some cases, a substandard image may be of substandard quality, where the quality is considered low compared to an ideal high quality image because of the presence of some pathology that makes the acquired image difficult to interpret due to noise or artifacts. In some cases of substandard images, conventional automated image analysis or image quality analysis methods may reject the substandard image and not include it in the automated processing of pathology (e.g., valvular stenosis) estimation.
[0044] Standard and non-standard views
[0045] Some examples of standard and nonstandard views that may be present with ultrasound imaging of the aortic valve: A parasternal long-axis view of the aortic valve in a standard, high-quality example will display many features of the heart, not just the aortic valve. Specifically, it will show the left ventricular cavity, left ventricular myocardium, mitral valve, left atrium, portion of the right ventricle, and the aortic valve. In a standard view, the aortic valve will appear at the right edge of the image. A nonstandard PLAX view that is useful for evaluating the aortic valve can exclude anatomical structures that are not the aortic valve and only show the aortic valve. In this case, the valve may be in the middle of the image or even at the opposite edge compared to a standard orientation. Traditional image evaluation algorithms may reject this type of image because it is very different from a textbook image.
[0046] Basic echo views that display the aortic valve, such as the parasternal long axis, parasternal short axis, and apical five-chamber views, ideally clearly display the aortic valve leaflet structure and its motion. Visualization of the valve leaflets can inform the assessment of stenosis and the severity of stenosis. If the valve leaflets are not thickened by calcification or other pathology and if they move freely and open fully, the possibility of stenosis can be ruled out. However, if the valve leaflets are stenotic and pathological, their motion may be abnormal and the opening may decrease during systole. In these cases, speckle noise and shadowing caused by the calcified valve can produce low-quality images that do not have the standard appearance or clarity of optimal images.
[0047] As used herein, a “standard” view generally refers to a textbook image (e.g., for a normative view), while a “non-standard” view generally refers to an image that deviates from the textbook image while still showing one or more features of interest relevant to the classification of the pathology assessed by the methods described herein.
[0048] Pathological classification
[0049] Described herein are methods and systems that provide accurate, rapid, simple, and / or low-cost methods for assessing pathologies (e.g., heart valve stenosis) by ultrasound imaging. In some cases, such assessment can be performed using two-dimensional imaging. In some cases, the pathology classified is aortic stenosis and / or stenosis of other valves of the subject's heart.
[0050] There are several possible workflows described herein that can be used for classification, e.g. Figure 1A As shown. A sonographer, cardiologist, or other user may acquire an ultrasound image 101 of a subject or patient, for example, by retrieving the ultrasound image from a data storage medium, from a data storage medium of an ultrasound imaging system, or by directly acquiring the ultrasound image from a patient being imaged in real time using an ultrasound imaging system. The image 103 is then processed, for example, by a machine learning model operating directly on the ultrasound imaging system or by another system that retrieves the image in 101. One or more pathologies are then classified 105 by the machine learning model or algorithm, and an output 107 is provided to the user. For example, the output may include a severity and / or a confidence value for the estimate of severity. In some cases, the output may include the identity of the detected pathology.
[0051] Alternative workflows may include any combination of the features or method steps described herein. Figure 1BAs shown in FIG, a sonographer, cardiologist, or other user may acquire ultrasound image clips of a subject or patient by directly acquiring them from the patient being imaged in real time using an ultrasound imaging system 102. The image clips may then be processed using an annotator model to classify individual views 104, for example, using a 2D convolutional neural network. One or more pathologies are classified by the annotator and the subviews are classified on a subview basis 106, and a separate confidence distribution 108 is generated for each pathology and / or each subview classified. Combinations and / or subsets of the individual confidence distributions are evaluated, and the top N (e.g., top 3, top 10, top 15, etc.) are selected 110 and aggregated 112 to produce a combined confidence distribution that optimizes the confidence of the classification of the one or more pathologies. The optimized confidence distribution is then used to evaluate the images collected from the subject in 102, and an output of optimized calculated severity and estimated confidence values is provided to a user 114. In some cases, this output may include the identification of multiple pathologies and / or their severity (where more than one pathology is detected in the acquired image). A rich and diverse training data set can be used for training. In some cases, a 2D neural network in the spatial domain is used together with a 1D neural network in the time domain. In some cases, a third neural network is utilized as an annotator for preprocessing using transfer learning, where the pre-trained weights from the annotator can be used to initialize the spatial portion of the model. In some cases, the model is trained as a multi-task using an auxiliary cardiac cycle regression function to teach the model to predict the cardiac cycle moment of each individual frame. In some aspects, a Gaussian mixture is fit to an empirical data distribution of aortic valve stenosis severity in a measurement space (a three-dimensional space of parameters such as aortic valve area, mean pressure flow gradient, and peak jet velocity). Severity grade probabilities are returned. A probability mass function is used to represent the likelihood of each grade.
[0052] The example model architecture utilizes a cascade of two CNNs in two domains, which operate on 24 consecutive frames that can be placed in a sliding window. MobileNetV3. Each frame is individually processed by a 2D network into a 1D embedding vector. This is processed in 1D along the time axis. An example, the automatic aortic stenosis service 1) Image preprocessing: Pixel normalization and rescaling. Frames are processed along the time axis to add auxiliary dynamic frames for temporal features. 2) Example annotator: The first step in the method. Identifies view, scan mode, and cardiac cycle. Filter for minimum clip length. Image quality rating and other attributes can be assigned. Clips are run through the example annotator to select clips that meet the AS severity assessment criteria. 3) Prediction of cardiac cycle of images: The initial algorithm implementation uses a trained neural network with EKG traces as input, other image-based methods can be used or combined with it. Determine the R-wave to R-wave period. 4) Aortic stenosis prediction: The model generates a confidence distribution for the AS severity level. It can also predict other relevant parameters from 2D (B-mode) images, such as valve area, mean pressure gradient, peak jet velocity, etc. The confidence distribution can be averaged over multiple cardiac cycles and multiple views can be used and combined. 5) Quantified probability of successful classification: Two features are used: the maximum value of the confidence distribution and the concentration of the distribution around this maximum value. These are plotted to produce a scalar that serves as a quality score for the probability of success. A high score is associated with a high probability of no error.
[0053] C) Interoperability with Standard Machine Learning Image Quality Scoring and Clip Selection When implemented to run integrated into an AI real-time acquisition device (such as our Guidance product), the AS severity score output can be run in real-time and weighted with a standard image quality assessment score to allow images with a high probability of pathology detection to not be rejected by an image quality clip selection function that may be looking for common canonical high-quality images and not trained to accept non-canonical images that include pathology. This can be done without specific training of a model for such distinctions. The function can also be run on images that have already been acquired, such as on a PACS viewer.
[0054] Automatic and accurate classification of aortic stenosis severity can be performed from 2D images alone. In some cases, the standard image quality clip selector function is not required to identify clips for processing. In such cases, abnormal images that do not meet the standard image quality definition are not rejected for processing. In such cases, the algorithm can assess the prediction accuracy of images independent of the standard image quality, without the need to train the clip selector separately for the pathology. This aspect can also be important and unique when integrated into an ultrasound imaging system that includes scan acquisition guidance, but it is also valuable when run on images that have already been acquired.
[0055] In some cases, the models described herein can be trained without the need for landmark computer vision analysis of tissue and without manual labeling of cardiac phase information for the training data. In such cases, cardiac phase information can be determined directly from analysis of one or more ultrasound images. Such approaches can provide significant advantages over earlier non-EKG approaches. For example, approaches based on valve closure detection can suffer from poor image quality, particularly when the patient has valvular stenosis. In some cases, the sliding window can use frames from more than one cardiac cycle, including time intervals within one cardiac cycle where the frames comprising the interval are from portions of two separate cardiac cycles.
[0056] In some cases, the methods described herein include estimating the severity of a pathology such as aortic stenosis, comprising: for each distinct clip associated with a given view, the method is as follows: 1) using a machine learning model to determine the patient's cardiac cycle for that clip; 2) sliding a cardiac cycle-sized window over the clip. For each window, calculate an AS confidence distribution and its associated likelihood of successful severity classification; 3) extract the top N such windows (based on likelihood); 4) once this has been done for all clips associated with a given view, rank the top N windows across all clips; 5) repeat the above process for all supported views; 6) combine the top N windows "across" all views and recalculate likelihood values based on the view combination. Rerank all window combinations and return the combination with the highest likelihood of successful severity classification.
[0057] Described herein are methods for early detection of pathologies such as aortic stenosis. In some cases, the methods can be used to provide a means of conveniently helping to identify cases of aortic stenosis in adults before the aortic stenosis becomes symptomatic (e.g., detecting AS during its asymptomatic period). In some cases, the methods and systems described herein can allow patients who exhibit the presence of multiple risk factors (e.g., advanced age, the presence of heart murmurs, pathological findings such as a bicuspid aortic valve) to be screened and / or evaluated by a non-professional (such as a family physician). If the initial assessment suggests the possibility of moderate to severe AS, the patient will be referred to an echocardiography laboratory because the risk of adverse outcomes with moderate to severe AS is significantly increased compared to mild or no AS.
[0058] Example implementation:
[0059] The example model described here was built with this scenario in mind. The patient data selected for evaluation of the model came from two institutions: Northwestern University Hospital (NM), which has at least 4 different collection sites, and the Minneapolis Heart Institute (MHI), which has at least 39 different collection sites, ranging from rural clinics to academic centers.
[0060] Considering the relatively low prevalence of mild, moderate, and severe AS in the echo database, echocardiographic studies were selectively sampled to enrich the dataset with as many AS cases as possible. For each study, the following information was collected when possible:
[0061] AS severity, determined by the cardiologist who initially read the images and established according to the ASE / ACC guidelines for quantification of AS
[0062] Aortic valve area (AVA) (usually by the continuity equation method)
[0063] Peak aortic valve velocity (also called jet velocity - JV)
[0064] Mean pressure gradient (MPG) across the aortic valve
[0065] Any information indicating the presence of a prosthetic aortic valve (TAVR, mechanical valve)
[0066] The patient's sex, height, weight, and age
[0067] Ultrasound machine model
[0068] The morphology of the aortic valve leaflets (tricuspid and mitral)
[0069] Initial screening yielded a set of just over 200,000 studies to choose from. The training data was then restricted to the following requirements:
[0070] Patients should retain their native valve (so no patients with prosthetic aortic valves). This requirement is imposed to explain the fact that prosthetic aortic valves generally increase hemodynamic parameters (especially the values of JV and MPG) when compared with native valves,
[0071] This could potentially confound the model and lead to systematic biases in severity predictions.
[0072] Patients should have at least one of AVA, MPG, or JV available or an AS severity label available
[0073] Patient sex, age, weight, and height should be available in most cases
[0074] Ultrasound machine models should always be available
[0075] These constraints limited the number of studies allowed to approximately 80,000. All data points from studies with aortic stenosis severity greater than "none" were then used to train the model. To supplement these, a subset of studies with "none" AS was randomly sampled to maximize diversity in the following features: sex, BMI, ultrasound machine model, and number of unique patients.
[0076] The final development dataset obtained contains:
[0077] Evaluation dataset characteristics: 1771 unique studies from 1427 unique patients, 1158 unique studies from 999 unique patients from MHI; 613 unique studies from 428 unique patients from NM; 852 studies from female patients, 919 studies from male patients; no patients shared with the following training datasets
[0078] BMI categories (height, weight, or both were not available in patient records for 15 studies): normal BMI (BMI < 25), overweight (25 ≤ BMI < 30), obese (30 ≤ BMI): studies 543 | 546 | 667.
[0079] Manufacturer Classification:
[0080] Manufacturers General Electric Company Philips Siemens Research
[0081] 202 Vivid E95
[0082] 118 Vivid E9
[0083] 115 Vivid i
[0084] 20 Vivid7
[0085] 440 iE33
[0086] 113 CX50
[0087] 44 EPIQ 7C
[0088] 510 SEQUOIA
[0089] 209 ACUSON SC2000
[0090] Year of collection:
[0091] Year: 2008|2009|2010|2011|2012|2013|2014|2015|2016|2017|2018|2019|2020
[0092] Research: 29|65|69|34|83|174|227|278|199|39|264|277|33
[0093] The evaluation dataset is constructed as a representative set of AS severity. Specifically, the AS severity classification is:
[0094] AS severity: None | Mild | Moderate | Severe
[0095] Research: 1,409|248|57|57
[0096] Training dataset characteristics:
[0097] 29,527 unique studies from 26,981 unique patients
[0098] 12,418 unique studies from 11,245 unique patients originating from the MHI
[0099] 17,109 unique studies from 15,736 unique patients originating from NM
[0100] A study of 14,856 female patients and 14,671 male patients
[0101] BMI classification:
[0102] Normal BMI (BMI < 25) | Overweight (25 ≤ BMI < 30) | Obese (30 ≤ BMI)
[0103] Research 8,846 | 9,761 | 10,711
[0104] Manufacturer Classification:
[0105] Manufacturer General Electric Philips Siemens
[0106] Research
[0107] 7,749 Vivid E95
[0108] 3,698Vivid E9
[0109] 2,725Vivid i
[0110] 197Vivid7
[0111] 7,117iE33
[0112] 1,347CX50
[0113] 641EPIQ 7C
[0114] 4,290SEQUOIA
[0115] 1,763ACUSON SC2000
[0116] Year of collection:
[0117] Year 2008|2009|2010|2011|2013|2014|2015|2016|2017|2018|2019|2020
[0118] Research 132|347|481|737|1,389|2,059|2,016|5,257|1,638|6,527|7,838|1,106
[0119] In approximately 97.5% of these studies (28,787 in total), the original AS severity diagnosis was available and was categorized as follows:
[0120] AS severity None | Mild | Mild to moderate | Moderate | Moderate to severe | Severe
[0121] Research 20,622|3,015|718|1,667|1,035|1,730
[0122] In the remaining 2.5% of these studies (740 in total), Doppler-derived measures (AVA, MPG, and JV) were available.The fact that the raw AS severity for all patients was not present in the training dataset does not pose a problem, as discussed in detail below.
[0123] Training methods
[0124] Collecting a large and diverse training dataset is only the first step. The way you train your model is also important, especially to reduce bias in the model’s predictions.
[0125] Outlined below are techniques used to reduce the chance of overfitting, improve generalization, and combat model bias in the example model:
[0126] The model capacity is purposefully limited. The example model is a sequence of a lightweight 2D convolutional neural network core operating in the spatial domain, followed by a 1D convolutional neural network operating in the temporal domain. In particular, 3D convolutional architectures are avoided to reduce the computational burden associated with high parameter counts and generally increased computational costs, and also to reduce overfitting. To further reduce overfitting and improve generalization, randomness is injected at different layers.
[0127] Specifically, heavy Dropout regularization is used while training the model with a small batch size (16 or 32). In addition, a fine-grained oversampling strategy is used to train the model.
[0128] 24 study groups were created corresponding to the Cartesian product of 2 patient gender groups, 3 ultrasound device manufacturer groups, and 4 aortic stenosis severity groups (a total of 2×3×4=24 groups). During training, examples were randomly and uniformly sampled from these 24 groups. Therefore, the model was presented with data from male patients 50% of the time and data from female patients the other 50% of the time. Similarly, the model was presented with data from General Electric medical devices 1 / 3 of the time, data from Philips medical devices 1 / 3 of the time, and data from Siemens medical devices another 1 / 3 of the time. The same applies to AS severity, with each of the four severities being presented 25% of the time.
[0129] The example model training strategy successfully minimizes the bias in the model predictions for several features of interest (here, patient sex, ultrasound machine type, and aortic stenosis severity level).
[0130] To further improve model performance, transfer learning and pre-trained weights were used to initialize the spatial portion of the example model (while randomly initializing the temporal portion). These pre-trained weights came from the annotator model as described herein. Another performance improvement occurred after training the example model as a multi-task. An auxiliary task of cardiac cycle regression was added to teach the model to predict the precise moment of the cardiac cycle on each frame. In some cases, knowledge distillation was used to exclude the use of one-hot encoded labels as supervisory signals to train the example model. Instead, a Gaussian mixture was fitted to the empirical data distribution of each AS severity in the measurement space (the three-dimensional space of AVA, MPG, and JV). As in Figure 6 As can be seen in Figure 2, AS severity density performs well in this space. For each study in the training set, the associated triplet (AVA, MPG, JV) is used and input into each Gaussian mixture model to obtain four positive values (each Gaussian mixture model returns its probability density value for the input triplet). Bayes' rule is then used to derive a probability mass function (PMF), which represents the likelihood that the input triplet is from any of the four AS severities. This PMF is in turn used as a supervisory signal to train the example model using a straightforward cross-entropy loss.
[0131] In some cases, the use of these "smoothed" labels improves overall performance and reduces the variability of predictions. The example model uses a cascade of two convolutional neural networks operating in different domains (see Figure 2). An input sequence of 24 consecutive video clip frames is input to a 2D convolutional neural network (CNN). The network uses the MobileNetV3 Large architecture. Each frame in the sequence is independently processed by this 2D network into a 1D embedding vector. The resulting 1D frame embedding sequence is then processed along the time axis by another 1D convolutional neural network. The latter network consists of an initial dimensionality reduction operation (reducing the size of the embedded feature space from 1280 to 256) followed by a sequence of 2 residual blocks, each block consisting of 2 1D convolutions in series followed by a residual connection. Finally, a final dimensionality reduction operation is inserted (from 128 dimensions to 64 dimensions). All 1D convolutions used in the example have a kernel size of 5. Figure 2 The architecture of the example model's temporary CNN is described in detail.
[0132] Figure 2 The example model shown in comprises a cascade of two convolutional neural networks operating in different domains. An input sequence of 24 consecutive video clip frames is input to a 2D convolutional neural network (CNN). The network adopts the MobileNetV3 Large architecture. Each frame in the sequence is independently processed by the 2D network into a 1D embedding vector. The resulting 1D frame embedding sequence is then processed along the time axis by a 1D convolutional neural network. The latter network consists of an initial dimensionality reduction operation (reducing the size of the embedded feature space from 1280 to 256) and a subsequent sequence of 2 residual blocks, each block consisting of two 1D convolutions in series and a subsequent residual connection. Finally, a final dimensionality reduction operation is inserted (from 128 dimensions to 64 dimensions). All 1D convolutions have a kernel size of 5.
[0133] Operational details of the sample model service for assessing aortic stenosis
[0134] The example model is implemented as part of a service that provides an end-to-end software system that independently processes an entire echocardiographic study and generates an estimate of aortic stenosis severity as output. The service consists of two main models: an annotator and the example model itself.
[0135] Image preprocessing
[0136] Both the annotator and the example model are implemented to operate directly on ultrasound image files. These images can be in a variety of formats. For training and development, video clip images are used. These images are typically obtained in DICOM format. Other image formats can be used for algorithm training and development. For example, in addition to or as an alternative to DICOM format images, the example model can also use scan-converted video images that have not yet been converted to a fully DICOM-compliant format.
[0137] In some cases, pre-scanned polar coordinate images can be used. To account for this variation, the terms image, clip, or video are used interchangeably. Several transformations may be required before these files are input to the annotator or example model. For example, the image content can first be extracted from the image file and decompressed into a 4-dimensional tensor representing the RGB frames of the video clip. The dimensions of this tensor can be [N, H, W, 3], where N represents the number of frames in the video clip, H and W are the height and width in pixels of each frame, respectively, and 3 represents the number of color channels. The data type can be implemented as uint8 (unsigned byte). Given an input tensor, it can be checked whether the number of frames N is greater than or equal to a threshold window size (for the example model, this is implemented as 24). The example model operates on a temporal window of 24 consecutive video frames and is rejected if the clip is not long enough.
[0138] Next, it can be determined whether the image file contains the necessary information to identify the location of the ultrasound cone region. The vast majority of image files that display a valid ultrasound region of interest store this information in a metadata field called "Sequence of Ultrasound Regions" (SOUR). If this information is not available, the clip is rejected, otherwise the cone region is cropped from the full frame.
[0139] The cone-shaped region clipped color frame can then be transformed into a grayscale frame using a dot product operation along the channel dimension for the red, green, and blue channels respectively. Once grayscaled, the frame can be quantized back to uint8. The grayscale frame can then be scaled to a height of 256 pixels while maintaining the aspect ratio (resizing of the example model is performed in uint8). A 256 pixel wide portion can then be cropped along the width of the frame centered on the middle vertical line (so that the "tip" of the cone-shaped region remains in the middle). Finally, the central 224×224 portion is extracted from the resulting 256×256 square.
[0140] The pixels can then be normalized in a variety of ways. For the example model, the first normalization is to simply rescale the pixel values from the original uint8 range (i.e., [0,255]) to [0,1) by multiplying all frames by 1 / 255. The resulting normalized frames are referred to as appearance frames in this article. The second normalization exploits the periodicity of the heartbeat. First, an average time image is extracted by averaging all frames in the clip (averaging across the frame index or "time" dimension). Second, the average time image is subtracted from each frame in the clip. Finally, the resulting (centered) frame is multiplied by a pre-calculated scaling factor so that the expected standard deviation of the pixel values is 1. The resulting normalized frames are referred to as dynamic frames or time frames in this article.
[0141] For each frame in the original clip, two different normalized frames are obtained: an appearance frame and a dynamic frame. As a final step, these normalized frames are concatenated along the channel dimension.
[0142] Example Annotator Model
[0143] Before an image can trigger a response from the example model that classifies aortic stenosis, it is first processed by an annotator model. The annotator model's role is to identify the clip's mode (B-mode ultrasound or color Doppler flow imaging) and view (echocardiogram view).
[0144] The example annotator is a 2D CNN that is architecturally identical to the spatial part of the example model (i.e., MobileNetV3Large) used to classify aortic stenosis. It is trained on a very large dataset using a multi-task learning approach. Specifically, the annotator is trained to jointly predict the following set of attributes on each frame:
[0145] Patient's gender
[0146] Patient age
[0147] Patient's BMI
[0148] Echocardiographic modality (B-mode ultrasound or color Doppler flow imaging)
[0149] Echocardiographic view (one of 48 possible views)
[0150] A set of physiological features present in the image (from a set of 26 possible features, such as aortic valve, mitral valve, left ventricle, right atrium, left ventricular outflow tract, hepatic vein, etc.)
[0151] • “Echo distance” (the ground truth echo distance is calculated using the AI guidance model). As used herein, echo distance generally refers to the deviation between the contemporaneous ultrasound probe pose and the ideal pose determined by the AI probe guidance model.
[0152] Heart cycle coordinates
[0153] The training dataset for the example annotator consisted of 452,778 image files (each file contributed 20 randomly sampled frames, for a total of over 9M individual frames) extracted from 45,783 unique studies originating from 35,262 unique patients. Similar to the approach taken by the example model for classifying aortic stenosis, the example annotator model was trained using a fine-grained oversampling approach with a total of 144 groups (2 patient sex groups x 3 patient BMI groups x 3 ultrasound device manufacturer groups x 8 patient age groups).
[0154] The example annotator training dataset overlaps with the training dataset for the example aortic stenosis model, with the following classifications:
[0155] 22,188 patients are unique to the training dataset for the example aortic stenosis model
[0156] 30,469 patients were unique to the example annotator’s training dataset
[0157] 4793 patients were common to both training datasets
[0158] Crucially, none of the patients used to validate the example aortic stenosis model had any overlap with the training dataset of the example aortic stenosis model, the training dataset of the example annotator, or the evaluation dataset of the example aortic stenosis model.
[0159] Editing Acceptance Criteria
[0160] For the example aortic stenosis model to return a prediction of aortic stenosis severity, the input image file must meet several conditions. These conditions are:
[0161] The image clip must be successfully pre-processed as outlined previously
[0162] Example: The annotator needs to identify the clip as a B-mode ultrasound. The color Doppler blood flow imaging clip is rejected.
[0163] • The example annotator needs to identify one supported cardiac view (e.g., PLAX, PSAX-AoV, or AP5 in the example implementation), other views will be rejected.
[0164] The example aortic stenosis model needs to be able to estimate the patient's cardiac cycle based on a video clip. Additionally, the clip length must be greater than or equal to the cardiac cycle to ensure that at least one complete cardiac cycle is present.
[0165] Predicting a patient's heart cycle
[0166] The example aortic stenosis model is able to predict a patient's cardiac cycle by looking only at a video clip. To achieve this, the model is trained to predict, for each frame, the coordinates of a point on a unit circle corresponding to the precise moment in the cardiac cycle at that frame. This circle is defined by arbitrarily associating the R-wave of the EKG with the coordinate point (1,0). For each intermediate time point between two consecutive R-waves, a simple temporal interpolation method is used to assign a timestamp and (cos(), sin()) coordinates to that point.
[0167] Given an input video clip, the example aortic stenosis model returns two time series corresponding to the estimated heart cycle coordinates at each frame. Figures 7 and 8 Two examples of predicted time series of cardiac cycle coordinates, predicted cardiac cycle (red horizontal line), and one random frame extracted from the clip are shown in Figure 2. To go from estimated cardiac cycle coordinates to cardiac cycle, the following post-processing operations are applied to each time series separately:
[0168] 1. Apply the Fast Fourier Transform (FFT) of the time series to obtain the Fourier coefficients.
[0169] 2. Extract the Fourier frequencies associated with these coefficients.
[0170] 3. Take the modulus of each Fourier coefficient (to make them positive).
[0171] 4. Apply a softmax operation to the vector of Fourier coefficients to obtain a normalized probability-like weight distribution.
[0172] 5. Calculate the weighted average of the Fourier frequencies using the normalized vector of weights from 4. Then obtain the cardiac cycle by taking the inverse of the average frequency.
[0173] This produces two estimates of the cardiac period, one from post-processing the cosine time series and another from post-processing the sine time series. The final cardiac period estimate, as implemented in the example annotator, is then obtained by averaging these two quantities.
[0174] Converting from image clips to confidence distributions on AS severity
[0175] The example aortic stenosis model operates on a window of 24 consecutive frames. Given an input sequence 24 frames long, the example aortic stenosis model returns 4 positive numbers that sum to 1, which is called a confidence distribution. These numbers represent the guarantee that the model has in attributing the input sequence to any of the 4 aortic stenosis severity levels implemented in the example (i.e., "none," "mild," "moderate," and "severe"). Given an image clip containing N frames (assuming N ≥ 24), a 24-frame long window can be slid along the time axis, and a total of (N-24+1) sequences can be extracted. For each sequence, the example aortic stenosis model returns a confidence distribution.
[0176] In order to generate medically meaningful confidence distributions, it is typically necessary to average these distributions over one or more cardiac cycles. To do this, the cardiac cycle T can be determined individually for each clip using the process described in the previous section. For the example model, this cycle spans a 24-frame window of Δ = (T - 24 + 1) (assuming T ≥ 24). Therefore, the confidence distributions for the example aortic stenosis model are averaged over the Δ window to produce a total of N - T + 1 predictions.
[0177] To obtain a confidence distribution for a view combination (i.e., a combination of cardiac cycles from two or more different cardiac views), the confidence distribution for each cardiac cycle is extracted as outlined herein and then considered based on their Cartesian product. Each element of the product consists of a tuple of two or more confidence distributions. To obtain a combined confidence distribution, the individual distributions of the tuples are averaged together. To return an aortic valve stenosis severity classification for a given confidence distribution, the severity with the highest confidence can be returned.
[0178] Quantify the probability of successful classification
[0179] In addition to producing an estimate of the severity of aortic stenosis (or other pathology), the example aortic stenosis model can also quantify the probability of producing a successful classification. Any such "confidence" metric can be defined statistically to observe how well the model performs on a sufficiently large dataset of held-out patients, and then hopefully find an appropriate feature space that can separate the performance of the example aortic stenosis model into different states.
[0180] In the example aortic stenosis model, two features are used for this purpose. The first feature is direct and corresponds to the maximum value of the confidence distribution of the example aortic stenosis model (see Figure 2 ). This feature is called the highest confidence score. The second feature quantifies how concentrated the distribution is around the maximum value and penalizes distributions with sub-peaks. This second feature is called the unimodal score.
[0181] For any given prediction, the example aortic stenosis model might be completely correct (error 0), off by 1 severity level (error 1), off by 2 severity levels (error 2), or off by 3 severity levels (error 3). The evaluation set described in this article (minus a subset of 40 patients held aside for a pilot study) was used to visualize how the model's error was distributed in feature space. Figure 9 400 random representative samples of each error type (error 0, 1, 2, or 3) are shown (note that the features were normalized to have zero mean and unit standard deviation before plotting). The black arrows on the plot show the direction of "optimal separability" of the errors. Figure 9 A scatter plot of the error for an example aortic stenosis model in feature space is shown. If we project each data point in our evaluation set along this optimal direction, we obtain a scalar quantity for each of these examples called a quality score (QS). Figure 10 The normalized data density is shown as a function of QS for the binary prediction tasks “none / mild” and “moderate / severe” in the evaluation dataset.
[0182] Can be obtained from Figure 10 Several conclusions can be inferred. First, the example model produces correct binary predictions most of the time. Second, high QS values are associated with a high probability of not making an error, while low QS values are associated with a high probability of making an error.
[0183] Figure 10 The density of data plotted against the quality score for a binary prediction task is shown. This data can be used to model the expected performance of an example aortic stenosis model on this binary prediction task based on QS. Using maximum likelihood, a model logistic regression classifier can be fit to the binary problem (classifying whether we made a successful binary prediction of none / mild vs. moderate / severe) using a single feature as input: the quality score. Assume that when the quality score is low, the model reduces to 50% random chance, while achieving perfect accuracy (100%) when the quality score is high. Furthermore, the classifier has only 2 free parameters, the scaling factor and the intercept. The results of the fit are shown in Figure 11 , which shows the logistic regression classifier probability curve for quality score.
[0184] In some cases, this "probability of success" can be used as a definition of confidence in the predictions of the example aortic stenosis model. Figure 3 An example model workflow is summarized. For example, an echocardiographic study may be performed on a large group of patients to accumulate DICOM files comprising images of multiple views of each patient's heart. Preprocessing may then be performed to annotate the clips included in one or more DICOM files to generate processed clips. The processed clips may then be submitted to a confidence algorithm and / or confidence machine learning model that estimates cardiac cycle and / or individual view quality scores for one or more (e.g., each) clips. The annotated clips may then be sorted by one or more of these parameters to produce a distribution of clips for each view that yields the highest confidence in classifying a pathology (e.g., aortic stenosis) of the patient's organ. The algorithm or machine learning model may then combine the insights gained from the clips with the highest predicted confidence to determine a combined confidence distribution based at least in part on a subset of each individual view and / or by selecting an optimized subset of views that were calculated or predicted to provide the highest confidence for the classification of the pathology to be measured.
[0185] Subsequent images can then be submitted to the example model workflow for evaluation of pathology in the newly imaged subject (e.g., for diagnosing aortic stenosis in a patient in a clinical setting). Output parameters from the model workflow can include the presence or absence of pathology, the severity of the pathology (e.g., aortic stenosis) in the submitted image, the model prediction confidence in the presence / absence of pathology and / or the severity of the pathology, and / or combinations thereof.
[0186] For example, when evaluating for aortic stenosis, such as Figure 4 An ultrasound image of a healthy heart in the parasternal long axis view, as shown in , would yield a high confidence value for a finding indicating the absence (or no / low severity) of pathology because the view shows a normal aortic valve with thin valve leaflets and no calcification, a left ventricle, and a left atrium, and a mitral valve, all of which appear normal. In contrast, a Figure 5 The ultrasound image of aortic stenosis shown in will produce a high confidence level for the prediction or classification of the presence of severe aortic stenosis (e.g., the output may include an indication that severe aortic stenosis is present with a confidence level of 95% or greater), even though the image shown does not contain the full anatomical targets of a standard canonical parasternal long axis view (the absence of which may often confuse quality assessment algorithms and models).
[0187] Compared to methods that require visual detection of valve opening and / or valve closure within ultrasound image clips, using a large number of views on the training data (including subviews containing non-canonical views) can provide the example model with significantly improved robustness in classifying aortic valve stenosis. In some cases, the example model is able to classify aortic valve stenosis without visually detecting valve opening and / or closure in the classified clips.
[0188] Machine Learning Algorithms
[0189] Disclosed herein are platforms, systems, and methods for providing ultrasound image classification using machine learning algorithms. In particular, in some aspects, the machine learning algorithm comprises a deep learning neural network configured to evaluate ultrasound images. The algorithm may comprise one or more of a positioning algorithm, a scoring algorithm, a probe guidance algorithm, and an intrinsic image quality algorithm. The positioning algorithm may comprise one or more neural networks that estimate the probe positioning relative to an ideal anatomical view or perspective and / or the distance or deviation of the current probe position from the ideal probe position. The intrinsic image quality algorithm may determine that the intrinsic image quality is below a threshold based in part on a determination by the positioning algorithm that one or more images have been acquired at a probe position where clinical quality images are expected to be obtained.
[0190] The development of each machine learning algorithm spans three phases: (1) dataset creation and management, (2) algorithm training, and (3) design elements necessary to tune product performance and usability. The dataset used to train the algorithm can be generated by acquiring ultrasound images, which are then collated and labeled by expert radiologists, for example, based on localization, scoring, and other metrics. Each algorithm is then trained using the training dataset, which can include one or more different target organs and / or one or more different views of a given target organ. The training dataset for the localization algorithm can be labeled based on known probe pose deviations from the optimal probe pose. A non-limiting description of the training and application of a localization algorithm or estimator can be found in U.S. patent application Ser. No. 15 / 831,375, which is incorporated herein by reference in its entirety. Another non-limiting description of a localization algorithm and a probe guidance algorithm can be found in U.S. patent application Ser. No. 16 / 264,310, which is incorporated herein by reference in its entirety. Design elements can include a user interface that includes an omnidirectional guidance feature.
[0191] The machine learning model may include supervised, semi-supervised, unsupervised or self-supervised machine learning models. In some cases, one or more ML methods perform classification or clustering of MS data. In some examples, the machine learning method includes a classical machine learning method, such as but not limited to a support vector machine (SVM) (e.g., a class SVM, a linear or radial kernel, etc.), a K-nearest neighbor (KNN), an isolation forest, a random forest, a logistic regression, an AdaBoost classifier, an extra tree classifier, an extreme gradient boost, a Gaussian process classifier, a gradient boost classifier, a light gradient boost, a linear discriminant analysis, a naive Bayesian, a quadratic discriminant analysis, a ridge classifier, or any combination thereof. In some examples, the machine learning method includes a deep learning method (e.g., a deep neural network (DNN)), such as but not limited to a fully connected network, a convolutional neural network (CNN) (e.g., a class CNN), a recursive neural network (RNN), a transformer, a graph neural network (GNN), a convolutional graph neural network (CGNN), a multi-level perceptron (MLP), or any combination thereof.
[0192] In some aspects, the classical ML method includes one or more algorithms that learn from existing observations (i.e., known features) to predict outputs. In some aspects, one or more algorithms perform clustering of data. In some examples, the classical ML algorithms for clustering include K-means clustering, mean shift clustering, density-based spatial clustering of applications with noise (DBSCAN), expectation maximization (EM) clustering (e.g., using Gaussian mixture models (GMM)), agglomerative hierarchical clustering, or any combination thereof. In some aspects, one or more algorithms perform classification of data. In some examples, the classical ML algorithms for classification include logistic regression, naive Bayes, KNN, random forest, isolation forest, decision tree, gradient boosting, support vector machine (SVM), or any combination thereof. In some examples, SVM includes one-class SMV or multi-class SVM.
[0193] In some aspects, the deep learning method includes one or more algorithms that learn to predict output by extracting new features. In some aspects, the deep learning method includes one or more layers. In some aspects, the deep learning method includes a neural network (e.g., a DNN including more than one layer). A neural network typically includes a connection node in the network that can perform functions such as transforming or translating input data. In some aspects, the output from a given node is passed as input to another node. The nodes in the network typically include input units in the input layer, hidden units in one or more hidden layers, output units in the output layer, or a combination thereof. In some aspects, the input node is connected to one or more hidden units. In some aspects, one or more hidden units are connected to the output unit. A node can typically receive input through an input unit and generate output from an output unit using an activation function. In some aspects, the input or output includes a tensor, a matrix, a vector, an array, or a scalar. In some aspects, the activation function is a rectified linear unit (ReLU) activation function, a sigmoid activation function, a hyperbolic tangent activation function, or a softmax activation function.
[0194] The connection between the nodes also includes a weight for adjusting the input data (i.e., activating input data or deactivating input data) to a given node. In some aspects, the weight is learned by a neural network. In some aspects, the neural network is trained to learn the weight using gradient-based optimization. In some aspects, gradient-based optimization includes one or more loss functions. In some aspects, gradient-based optimization is gradient descent, conjugate gradient descent, stochastic gradient descent, or any variant thereof (e.g., adaptive moment estimation (Adam)). In some further aspects, back propagation is used to calculate the gradient in gradient-based optimization. In some aspects, the nodes are organized into graphs to generate a network (e.g., a graph neural network). In some aspects, the nodes are organized into one or more layers to generate a network (e.g., a feedforward neural network, a convolutional neural network (CNN), a recurrent neural network (RNN), etc.). In some aspects, CNN includes a class of CNN or multiple classes of CNN.
[0195] In some aspects, the neural network includes one or more recursive layers. In some aspects, the one or more recursive layers are one or more long short-term memory (LSTM) layers or gated recursive units (GRUs). In some aspects, the one or more recursive layers perform sequential data classification and clustering, taking into account data ordering (e.g., time series data). In these aspects, future predictions are made by the one or more recursive layers based on the sequence of past events. In some aspects, the recursive layers retain or "remember" important information while selectively "forgetting" information that is not necessary for classification.
[0196] In some aspects, the neural network includes one or more convolutional layers. In some aspects, the input and output are tensors representing variables or attributes (e.g., features) in the data set, which can be referred to as feature maps (or activation maps). In these aspects, the one or more convolutional layers are referred to as feature extraction stages. In some aspects, the convolution is a one-dimensional (1D) convolution, a two-dimensional (2D) convolution, a three-dimensional (3D) convolution, or any combination thereof. In other aspects, the convolution is a 1D transposed convolution, a 2D transposed convolution, a 3D transposed convolution, or any combination thereof.
[0197] The layer in the neural network can also include one or more pooling layers before or after the convolutional layer. In some aspects, one or more pooling layers use filters that summarize the area of the matrix to reduce the dimension of the feature map. In some aspects, this downsamples the number of outputs and therefore reduces the parameters and computing resources required for the neural network. In some aspects, one or more pooling layers include maximum pooling, minimum pooling, average pooling, global pooling, standard pooling or its combination. In some aspects, maximum pooling reduces the dimension of data by only taking the maximum value in the matrix area. In some aspects, this helps to capture the most important one or more features. In some aspects, one or more pooling layers are one-dimensional (1D), two-dimensional (2D), three-dimensional (3D) or its any combination.
[0198] The neural network may also include one or more flattening layers that can flatten the input to be passed to the next layer. In some aspects, the input (e.g., feature map) is flattened by reducing the input to a one-dimensional array. In some aspects, the flattened input can be used to output the classification of the object. In some aspects, the classification includes binary classification or multi-class classification of visual data (e.g., images, videos, etc.) or non-visual data (e.g., measurements, audio, text, etc.). In some aspects, the classification includes binary classification of images (e.g., need for contrast agent or not). In some aspects, the classification includes multi-class classification of text (e.g., recognition of handwritten digits). In some aspects, the classification includes binary classification of measurements. In some examples, the binary classification of measurements includes classification of system performance using the physical measurements described herein (e.g., normal or abnormal, normal or abnormal).
[0199] The neural network may also include one or more dropout layers. In some aspects, dropout layers are used during training of the neural network (e.g., to perform binary or multi-class classification). In some aspects, the one or more dropout layers randomly set some weights to 0 (e.g., approximately 10%, 20%, 30%, 40%, 50%, 60%, 70%, 80% of the weights). In some aspects, setting some weights to 0 also sets the corresponding elements in the feature map to 0. In some aspects, the one or more dropout layers can be used to prevent overfitting of the neural network.
[0200] The neural network may also include one or more dense layers, which include a fully connected network. In some aspects, information is passed through the fully connected network to generate a predicted classification of the object. In some aspects, the error associated with the predicted classification of the object is also calculated. In some aspects, the error is backpropagated to improve the prediction. In some aspects, one or more dense layers include a Softmax activation function. In some aspects, the Softmax activation function converts a vector of numbers into a vector of probabilities. In some aspects, these probabilities are then used for classification, such as pathology classification according to any method or system described herein.
[0201] Computer system
[0202] The present disclosure provides a computer system programmed to implement the method of the present disclosure. Figure 12 A computer system 1201 is shown that is programmed or otherwise configured to assess one or more pathologies based on ultrasound images using any of the methods or systems described herein. The computer system 1201 can facilitate various aspects of the present disclosure, such as, for example, calculating the severity and confidence of one or more pathologies based on ultrasound images and / or alerting a user to the severity and corresponding confidence. In some cases, the computer system 1201 can be the user's electronic device or a computer system remotely located relative to the electronic device. The electronic device can be a mobile electronic device.
[0203] Computer system 1201 includes a central processing unit (CPU, also referred to herein as a "processor" and "computer processor") 1205, which can be a single-core or multi-core processor, or multiple processors for parallel processing. Computer system 1201 also includes memory or memory locations 1210 (e.g., random access memory, read-only memory, flash memory), electronic storage 1215 (e.g., a hard disk), a communication interface 1220 (e.g., a network adapter) for communicating with one or more other systems, and peripherals 1225, such as cache, other memory, data storage, and / or electronic display adapters. Memory 1210, storage 1215, interface 1220, and peripherals 1225 communicate with CPU 1205 via a communication bus (solid lines), such as a motherboard. Storage 1215 can be a data storage unit (or data repository) for storing data. Computer system 1201 can be operatively coupled to a computer network ("network") 1230 with the aid of communication interface 1220. Network 1230 may be the Internet, the Internet and / or an extranet, or an intranet and / or an extranet in communication with the Internet. In some cases, network 1230 is a telecommunications and / or data network. Network 1230 may include one or more computer servers that may implement distributed computing, such as cloud computing. In some cases, with the assistance of computer system 1201, network 1230 may implement a peer-to-peer network, which may enable devices coupled to computer system 1201 to act as either clients or servers.
[0204] The CPU 1205 can execute a series of machine-readable instructions, which can be embodied in a program or software. The instructions can be stored in a memory location, such as the memory 1210. The instructions can be sent to the CPU 1205, which can then be programmed or otherwise configured to implement the methods of the present disclosure. Examples of operations performed by the CPU 1205 can include fetching, decoding, executing, and writing back.
[0205] CPU 1205 may be part of a circuit, such as an integrated circuit. One or more other components of system 1201 may be included in the circuit. In some cases, the circuit is an application-specific integrated circuit (ASIC).
[0206] Storage unit 1215 can store files such as drivers, libraries, and saved programs. Storage unit 1215 can store user data, such as user preferences and user programs. In some cases, computer system 1201 may include one or more additional data storage units external to computer system 1201, such as on a remote server that communicates with computer system 1201 via an intranet or the Internet.
[0207] Computer system 1201 can communicate with one or more remote computer systems via network 1230. For example, computer system 1201 can communicate with a remote computer system of a user (e.g., a family doctor, an untrained technician, a patient, and / or a cardiologist or other specialist). Examples of remote computer systems include personal computers (e.g., portable PCs), tablets or tablet PCs (e.g., iPad, Galaxy Tab), phones, smartphones (e.g. iPhone, Android-enabled devices, ) or personal digital assistant. Users can access computer system 1201 via network 1230.
[0208] The methods described herein may be implemented by machine (e.g., computer processor) executable code stored on an electronic storage location (such as, for example, memory 1210 or electronic storage unit 1215) of computer system 1201. Machine executable code or machine readable code may be provided in the form of software. During use, the code may be executed by processor 1205. In some cases, the code may be retrieved from storage unit 1215 and stored on memory 1210 for ready access by processor 1205. In some cases, electronic storage unit 1215 may be eliminated and machine executable instructions may be stored on memory 1210.
[0209] The code may be precompiled and configured for use with a machine having a processor suitable for executing the code, or may be compiled during runtime. The code may be provided in a programming language that may be selected so that the code can be executed in a precompiled or compiled manner.
[0210] Various aspects of the systems and methods provided herein (such as computer system 1201) can be embodied in programming. Various aspects of the present technology can be considered to be "products" or "articles" in the form of associated data carried or embodied on a type of machine-readable medium, usually in the form of machine (or processor) executable code and / or on a type of machine-readable medium. Machine executable code can be stored on an electronic storage unit such as a memory (e.g., read-only memory, random access memory, flash memory) or a hard disk. "Storage" type media can include any or all tangible memories or their associated modules of a computer, processor, etc., such as various semiconductor memories, tape drives, disk drives, etc., which can provide non-transient storage for software programming at any time. All or part of the software can sometimes be communicated via the Internet or various other telecommunications networks. Such communication, for example, can enable software to be loaded from one computer or processor to another, such as from a management server or a host computer to a computer platform of an application server. Therefore, another type of medium that can carry software elements includes light, electricity, and electromagnetic waves used in physical interfaces between local devices across wired and optical landline networks and on various air links. The physical elements that carry such waves, such as wired or wireless links, optical links, etc., can also be considered to be the medium that carries the software. As used herein, unless restricted to non-transitory, tangible "storage" media, terms such as computer or machine "readable medium" refer to any medium that participates in providing instructions to a processor for execution.
[0211] Thus, machine-readable media such as computer executable code can take many forms, including but not limited to tangible storage media, carrier media, or physical transmission media. Non-volatile storage media include, for example, optical or magnetic disks, such as any storage device in any computer, such as can be used to implement the databases shown in the accompanying drawings. Volatile storage media include dynamic memory, such as the main memory of such a computer platform. Tangible transmission media include coaxial cables; copper wire and optical fiber, including the wires that make up the bus within a computer system. Carrier transmission media can take the form of electrical or electromagnetic signals, or sound or light waves, such as those generated during radio frequency (RF) and infrared (IR) data communications. Thus, common forms of computer-readable media include, for example: floppy disks, diskettes, hard disks, magnetic tape, any other magnetic medium, CD-ROMs, DVDs or DVD-ROMs, any other optical medium, punched card tape, any other physical storage medium with a pattern of holes, RAM, ROMs, PROMs and EPROMs, FLASH-EPROMs, any other memory chips or cartridges, carrier waves that transmit data or instructions, cables or links that transmit such carrier waves, or any other medium from which a computer can read programming code and / or data. Many of these forms of computer readable media may be involved in carrying one or more sequences of one or more instructions to a processor for execution.
[0212] The computer system 1201 may include or be in communication with an electronic display 1235 that includes a user interface (UI) 1240 for providing, for example, a readout of the severity and confidence level of the presence of one or more pathologies in real time. Examples of UIs include, but are not limited to, graphical user interfaces (GUIs) and web-based user interfaces.
[0213] The methods and systems of the present disclosure may be implemented by one or more algorithms. The algorithms may be implemented by software when executed by the central processing unit 1205. The algorithms may, for example, implement any method or facilitate the operation of any system described herein.
[0214] Example
[0215] The following illustrative examples are representative of aspects of the software applications, systems, and methods described herein and are not meant to be limiting in any way.
[0216] Example 1 Automatic Classification of Pathologies During Ultrasound Imaging Procedures
[0217] A sonographer acquires ultrasound images of a patient according to the methods and / or using the systems described herein. While a stenographer collects the images, an ultrasound imaging system implementing the methods described herein processes the images and assesses the severity of one or more pathologies (such as aortic stenosis). The system then provides the stenographer with an estimate of the severity of the one or more pathologies affecting the subject, along with an estimate of the confidence level in the estimate.
[0218] Although preferred aspects of the present invention have been shown and described herein, it will be apparent to those skilled in the art that these aspects are provided by way of example only. This does not mean that the present invention is limited by the specific examples provided in the specification. Although the present invention has been described with reference to the foregoing description, the description and illustration of the various aspects herein are not intended to be interpreted in a limiting sense. Without departing from the present invention, those skilled in the art will appreciate many variations, changes, and substitutions. In addition, it will be understood that all aspects of the present invention are not limited to the specific description, configuration, or relative proportions set forth herein, which depend on a variety of conditions and variables. It will be understood that various alternatives to the various aspects of the present invention described herein can be adopted in practicing the present invention. Therefore, it is contemplated that the present invention also encompasses any such substitutions, modifications, variations, or equivalents. The following claims are intended to define the scope of the invention, and methods and structures within the scope of these claims and their equivalents are thereby encompassed.
Claims
1. A method for ultrasound imaging, comprising: processing multiple ultrasound images to automatically classify a subject's pathology; as well as outputting an indication of the subject's condition, the output comprising: (i) an indication of the severity of a pathology in the subject; and (ii) a confidence estimate for the indication of (i).
2. The method of claim 1 , wherein the plurality of ultrasound images comprises images captured across at least a portion of at least one complete cardiac cycle of the subject, or wherein processing comprises providing the acquired plurality of ultrasound images as input to a trained machine learning model.
3. The method of claim 2, wherein the machine learning model is trained by a training method that does not include computer vision analysis of tissue landmarks and / or manual labeling of training data using cardiac phase information.
4. The method of claim 2, wherein the pathology is aortic stenosis and the organ is the subject's heart.
5. The method of claim 3, wherein the method does not include visually detecting heart valve closure within the acquired plurality of images. The method of claim 2 , wherein automatic classification is performed on at least a subset of the plurality of ultrasound images as secondary images. 7 . The method of claim 2 , wherein the training method comprises individually evaluating the prediction accuracy of each of a plurality of discrete views included in at least a subset of the training data. The method of claim 7 , wherein the plurality of discrete views comprises standard views and non-standard views.
9. The method of claim 3, wherein the one or more video clips include a plurality of acquired ultrasound images.
10. A method according to claim 9, wherein the processing includes sliding a window over multiple frames of the one or more video clips to select frames that provide improved confidence in the indication of (i) compared to an indication generated using a stationary window including a single cardiac cycle, wherein the frames of the one or more video clips include frames obtained from more than one cardiac cycle and the sliding window slides over frames from more than one cardiac cycle.
11. The method of claim 9, wherein each of the one or more video clips is associated with one or more discrete views, and estimating the severity of the pathology comprises: (I) for each of the one or more video clips: determining a cardiac cycle of the subject of the video clip; sliding a window having a size based on the determined cardiac cycle over the clip; as well as calculating a confidence distribution and / or an associated likelihood of successful severity classification for the pathology; (II) extracting windows associated with the one or more video clips that meet a threshold for likelihood of successful severity classification; as well as (III) combining the extracted windows from two or more of the one or more discrete views to obtain an increased likelihood of successful severity classification compared to classification based on a single one of the discrete views.
12. The method of claim 11, wherein (I)-(III) are repeated for all views supported by the ultrasound imaging system. The method of claim 12 , wherein the ultrasound imaging system supports 30 or more views.
14. The method of claim 11, wherein estimating the severity of the pathology further comprises: sorting the windows of all clips associated with a particular view based on likelihood of successful severity classification to obtain a subset comprising fewer than all of the windows associated with the particular view, the subset comprising three or more windows associated with the particular view having the highest likelihood of successful severity classification; and After the combining of (III), the windows of all combined clips are reordered based on the likelihood of successful severity classification to obtain the combination of windows with the highest global likelihood of successful severity classification.
15. The method according to claim 1, wherein The output includes instructions to a user of the ultrasound system to acquire more images of the subject.
16. The method of claim 15, wherein the instruction to acquire more images is provided based at least in part on detecting at least a threshold likelihood of the presence of the pathology in one or more of the acquired plurality of ultrasound images.
17. The method of claim 15, wherein the instruction to acquire more images is provided based at least in part on detecting less than a threshold confidence level in the presence or absence of the pathology. The method according to claim 17 , wherein the acquired plurality of ultrasound images are two-dimensional ultrasound images.
19. An ultrasound imaging system, comprising: an ultrasound imaging probe for acquiring a plurality of ultrasound images of a subject; computing systems; and A non-transitory computer-readable storage medium storing instructions that, when executed by a processor of the computing system, cause the ultrasound imaging system to perform the method of any one of claims 1 to 18.
20. A non-transitory computer-readable medium storing instructions that, when executed by a processor of a computer, cause the computer to perform the method of any one of claims 1 to 18.
Citation Information
Patent Citations
Guided navigation of an ultrasound probe
US20180153505A1
Prescriptive guidance for ultrasound diagnostics
US20200245970A1