Echocardiographic Estimation of Right Atrial Pressure Using a Lightweight and Open-World Artificial Intelligence System
Patent Information
- Authority / Receiving Office
- JP · JP
- Patent Type
- Applications
- Current Assignee / Owner
- THE GOVERNMENT OF THE UNITED STATES OF AMERICA AS REPRESENTED BY THE SECRETARY DEPARTMENT OF HEALTH & HUMAN SERVICES
- Filing Date
- 2023-06-09
- Publication Date
- 2026-06-02
AI Technical Summary
Manual analysis of echocardiograms for estimating inferior vena cava (IVC) collapse rate and right atrial pressure (RAP) is time-consuming, prone to errors, subjective, and lacks consistency, particularly in resource-constrained regions, affecting patient management and treatment outcomes.
An AI system with an image retrieval network for quality assessment, a region segmentation network for IVC localization, and an open-world active learning system for real-time IVC collapse rate and RAP estimation, utilizing a lightweight and efficient Tri-polar Attention Network (TaNet) for segmentation and distance calculations.
Enables accurate, efficient, and reproducible estimation of IVC collapse rate and RAP, improving clinical decision-making and reducing the need for expert interpretation, suitable for resource-limited settings.
Smart Images

Figure 00000000_0000_ABST
Abstract
Description
Technical Field
[0001] Cross - Reference to Related Applications This patent application claims the benefit of U.S. Provisional Patent Application No. 63 / 458,054, filed Apr. 7, 2023, and U.S. Provisional Patent Application No. 63 / 350,720, filed Jun. 9, 2022, which are hereby incorporated by reference.
[0002] Statement Regarding Federally Sponsored Research or Development This invention was made with government support under grant numbers LM010018, Z99CL999999, and HL006199 awarded by each of the National Institutes of Health. The government has certain rights in this invention.
Background Art
[0003] Background of the Invention Echocardiography (``echo'') is performed by capturing ultrasonic images of the heart and its related anatomical structures using a dedicated bedside or portable imaging system. Using ultrasound to capture the structure and function of the heart has several advantages over other imaging modalities, including high temporal resolution, non-invasiveness, low cost, and portability. Images / videos captured during an echo exam are typically analyzed manually by a sonographer or cardiologist, who interprets the results to guide the diagnosis, treatment planning, and prognosis of various cardiovascular diseases. However, this manual analysis is time-consuming and requires advanced training, which can be costly or hampered by a lack of expertise in resource-constrained regions. Furthermore, human interpretation is subjective and may lack consistency due to specific factors, resulting in low reproducibility and potentially affecting patient management and treatment outcomes. Therefore, accurate and automated analysis of echocardiograms has the potential to mitigate these issues: 1) it can lead to increased efficiency, allowing physicians to secure more analysis time; 2) it can provide consistent and reproducible results; and 3) it will improve clinical interpretation and decision-making. SUMMARY OF THE INVENTION
[0004] In certain aspects of the present disclosure, an artificial intelligence (AI) system is provided for real-time estimation of inferior vena cava (IVC) collapse rate and right atrial pressure (RAP) from echocardiogram. The system has: an image retrieval network configured to receive images from an echocardiogram and perform image quality assessment to generate selected images; a region segmentation network configured to localize and segment the selected images to obtain the IVC region in the selected images; an IVC quantification and RAP estimation network configured to perform distance calculations to determine the diameters of the IVC region at different spatial and temporal points; and an open-world active learning system having a classification engine and a clustering engine. By performing IVC thickness calculation and collapse rate analysis at different spatial positions and different temporal points, the AI system can measure the IVC and RAP values with higher reliability.
[0005] In another aspect of the present disclosure, a method for estimating inferior vena cava (IVC) collapse rate and right atrial pressure (RAP) from an echocardiogram executed in real time by an artificial intelligence (AI) system is provided. The method has: receiving images from an echocardiogram and performing image quality assessment to generate selected images; localizing and segmenting the selected images to obtain the IVC region in the selected images; performing distance calculations to determine the diameters of the IVC region at different spatial and temporal points; classifying known and unknown views in the selected images; and grouping the unknown views into one or more clusters used to update the classification ability in the AI system. By performing IVC thickness calculation and collapse rate analysis at different spatial positions and temporal points, the method can measure the IVC and RAP values with higher reliability.
[0006] In yet another aspect of the present disclosure, a non-transitory computer-readable medium implemented in an artificial intelligence (AI) system for real-time estimation of inferior vena cava (IVC) collapse rate and right atrial pressure (RAP) from echocardiogram examinations is provided. The non-transitory computer-readable medium has instructions that, when executed by a computer, configure the computer to perform the following steps, which include: receiving an image from an echocardiogram examination and performing an image quality assessment to generate a selected image; localizing and segmenting the selected image to obtain an IVC region in the selected image; performing distance calculations to determine the diameters of the IVC region at different spatial and temporal points; classifying known and unknown views in the selected image; and grouping the unknown views into one or more clusters used to update the classification ability in the AI system. By performing IVC thickness calculations and collapse rate analysis at different spatial positions and temporal points, the AI system can more reliably measure IVC and RAP values.
Brief Description of the Drawings
[0007]
Figure 1
Figure 2
Figure 3
Figure 4
Figure 5
Figure 6
Figure 7
Figure 8
Figure 9
Figure 10
[0008] Detailed description of the invention Accurate estimation of right atrial pressure (RAP) by echocardiography is important for assessing hemodynamics and vascular volume status and guiding clinical decision-making to provide optimal patient care. Embodiments of the present disclosure provide an echocardiographic artificial intelligence (AI) system that can estimate inferior vena cava (IVC) collapse rate and RAP in real time. This system employs a lightweight and open-world architecture to quickly generate reproducible and interpretable results. The lightweight nature facilitates the integration of this system into portable devices, thereby improving accessibility, and the open-world nature makes the system more robust when detecting new / unknown / unpredictable cases or scenarios in real-world clinical settings.
[0009] Embodiments of the present disclosure provide the following: 1) automatic echocardiographic evaluation of IVC diameter and collapse rate and RAP estimation; 2) a lightweight system for real-time quantitative echocardiographic evaluation; and 3) an open-world system for active echocardiographic interpretation and learning.
[0010] The inferior vena cava (IVC) is a blood vessel responsible for the circulation that returns deoxygenated blood from the lower extremities and abdomen to the right atrium. In several studies [3, 4, 7], the IVC diameter and its change associated with inspiration (also known as the IVC collapsibility rate) can be easily captured using ultrasonic imaging devices and have been reported to be useful for determining the patient's fluid status in various settings such as acute heart failure (HF) and critically ill patients. The IVC collapsibility rate is measured based on the difference between the maximum and minimum diameters. This collapsibility rate can be used to easily estimate the right atrial pressure (RAP) as follows [7]: An IVC diameter of 2.1 cm with a collapse of 50% or more supports an estimated RA pressure of 3 mmHg (range, 0 - 5 mmHg), while an IVC diameter of 2.1 cm with a collapse of less than 50% supports an estimated RA pressure of 15 mmHg (range, 10 - 20 mmHg). RAP is an important indicator of right ventricular diastolic function, volume status, and right heart compliance and has been used as a predictor of mortality in patients with heart failure and cardiogenic shock.
[0011] Clinically, as shown in FIG. 1, the IVC diameter is measured perpendicular to the long axis 102 of the IVC within 1.0 to 2.0 cm from the junction of the large vein and the right atrium. Then, the collapse rate is visually estimated based on the change in the IVC diameter associated with inspiration. Specifically, there are several steps involved in the current practice of estimating the IVC and RAP. First, a cardiologist needs to manually select the subcostal long-axis view of the IVC and visually evaluate the quality of the view. Next, she / he needs to measure the IVC diameter at the correct position and determine the collapse rate. Finally, the determined IVC collapse rate is integrated into a predetermined formula to manually estimate the RAP. This practice is troubled by the fact that 1) it is time-consuming and error-prone, 2) it is a manual operation by a cardiologist and is inaccurate because it depends on visual evaluation, and 3) it is subjective and inconsistent, which may lead to variations in decision-making. Considering that the IVC collapse rate has a strong potential to predict the fluid status of patients in heart and non-heart diseases, it is extremely important to establish and develop a standardized and fully automated approach for real-time IVC collapse rate analysis and RAP estimation.
[0012] FIG. 2 shows an end-to-end and automated AI system 200 for real-time estimation of the IVC collapse rate and RAP. As shown in FIG. 2, an echocardiogram 202 is performed, where images are captured and high-quality images 204 from the examination are input into the system 200 for analysis. Using the high-quality images 204, the system 200 performs real-time subcostal IVC view selection and IVC image quality evaluation 206, IVC region segmentation 208, and IVC diameter quantification, collapse rate, and RAP estimation 210. Hereinafter, these algorithms for performing the IVC view search and quality evaluation 206, IVC region segmentation 208, and IVC quantification and collapse rate and RAP estimation 210 will be described.
[0013] IVC View Selection and Quality Evaluation The image search network 206 of the system 200 searches for specific views with acceptable (medium to good) quality. This search is performed by a lightweight model with a shared encoder and two heads (see Figure 7 and related considerations below for further details), where the first head is for view classification to identify IVC views, and the second head is for quality assessment to evaluate the quality of the identified views.
[0014] The image search network 206 includes five inverted residual bottleneck (see the residual bottleneck of MobileNetV2-s [8]) blocks and a final pooling layer. The two heads are configured to perform view classification and quality assessment and operate together with the encoder. Each head has the following layers: a global average pooling (GAP) layer, a dropout layer, a fully connected (FC) layer, and an OpenMax layer. Each head, together with the encoder, is adjusted by: (1) initializing the encoder with echo-specific weights; and (2) fine-tuning the encoder and its view classification layer. The view classification head detects IVC views from a given echo examination, while the quality assessment head labels a given IVC view as good or bad in terms of quality. Thus, the model includes three task-specific layers, and both heads share the same encoder. Also, both heads can be configured to operate simultaneously or in parallel, thereby further improving the network efficiency of the image search component 206.
[0015] In clinical practice, echocardiography technicians visually identify echocardiogram views and manually exclude low-quality echos as they can lead to inaccurate measurements. The image search component 206 enables automated real-time IVC view classification and quality assessment in clinical practice. The performance of view classification and quality assessment using the image search component 206 is shown in Table 1 and Table 2 respectively. From the tables, it can be seen that the image search component 206 achieves equivalent performance, if not better, compared to state-of-the-art echocardiography examination systems while having a much lower computational complexity. Furthermore, the image search component 206 has 55,620 parameters in certain embodiments, which is smaller than most prior art networks in the literature; for example, VGG16 has 138MM parameters (528MB) and ResNet18 has 11MM (44MB).
[0016]
Table 1
[0017]
Table 2
[0018] These results shown in Table 1 and Table 2 demonstrate the superiority of the image search component 206 based on a shared encoder with two heads for view classification and quality assessment. After using the image search component 206 to search for high-quality images, poor-quality IVC images are automatically excluded from further analysis while good-quality echos are sent to the IVC region segmentation network 208.
[0019] IVC Region Segmentation The IVC region segmentation network 208 is a novel, lightweight, and technically efficient network called the Tri-polar Attention Network (TaNet) that detects and segments the IVC region (see Figures 8 - 9 and related discussions below for further details). The TaNet of the IVC region segmentation network 208 achieves excellent performance in segmenting the IVC in each frame from the image retrieval component 206. In certain embodiments, the TaNet network achieves a high accuracy of 98% and an intersection over union of 96%, and does so at a high speed of 85 frames per second (FPS). This speed is much higher than that of the most widely used segmentation networks (e.g., Unet and FCN). Figure 3 shows the TaNet of the IVC region segmentation network 208 and IVC segmentation using ground truth annotations.
[0020] Real-time IVC quantification The IVC quantification and RAP estimation network 210 performs a temporal analysis of cardiac metrics over all frames, instead of extracting cardiac biomarkers at specific frames (e.g., end-diastole or end-systole) as done manually, thereby improving reproducibility in clinical cardiology practice and research. The analysis of all frames provides information regarding the temporal changes during respiration over multiple cardiac cycles. This temporal IVC analysis enables measurement of the IVC collapse rate, which is an important component for estimating RAP.
[0021] Prior to defining the IVC boundary and calculating the diameter, the IVC quantification and RAP estimation network 210 performs morphological cleaning to remove isolated unwanted pixels and leave only the closed region of interest. Next, the IVC quantification and RAP estimation network 210 uses a Moore-neighborhood trace algorithm modified by Jacob's stopping criterion to calculate the contour of the clean region (see Pradhan, Ratika, et al. “Contour line tracing algorithm for digital topographic maps.” International Journal of Image Processing (IJIP) 4.2 (2010):156-163, which is incorporated herein by reference for the description of the algorithm). Next, the IVC quantification and RAP estimation network 210 divides the defined region into equal segments (or sectors). To calculate the IVC diameter (IVCD), the IVC quantification and RAP estimation network 210 determines the major axis of the sub-segment located approximately 2 cm proximal to the small hole of the RA. Next, the IVC quantification and RAP estimation network 210 calculates the Euclidean distance between the endpoints of the major axis. Finally, the IVC quantification and RAP estimation network 210 converts the calculated pixel distance to millimeters (mm).
[0022] After calculating the IVCD, the IVC quantification and RAP estimation network 210 constructs an IVCD curve by plotting the IVCD values across the frames. Next, the IVC quantification and RAP estimation network 210 uses a Savitzky-Golay filter to obtain a smoothed IVCD curve. The smoothed curve is then used to calculate the RAP.
[0023] RAP is calculated as follows: (1) The IVC quantification and RAP estimation network 210 calculates the IVC collapse rate based on the difference between the absolute maximum peak and the minimum valley in the IVC D curve; and (2) The IVC quantification and RAP estimation network 210 calculates the RAP value by substituting the values of the IVC diameter and the collapse rate into Equation (1).
[0024]
Equation
[0025] Equation (1) is for calculating RAP using the standard equation set by the American Society of Echocardiography (ASE). Equation (1) is an exemplary equation for calculating RAP and can be adapted for various site-specific applications. For example, the system 200 (see Figure 2) can be adjusted to utilize any site-specific version of the above RAP equation to determine RAP within the context of the relevant site.
[0026] Figure 4 shows the calculated IVC diameter in each frame plotted as a time signal. From this curve, the IVC quantification and RAP estimation network 210 can estimate the absolute maximum value (highest peak) and the absolute minimum value (lowest valley) of this measurement. Also, the IVC quantification and RAP estimation network 210 can estimate the average maximum value and the average minimum value by averaging the peaks and valleys of the curve.
[0027] To evaluate the degree of agreement between the automated values and the values estimated by experts, and thereby determine the effectiveness of the IVC quantification and RAP estimation network 210, Pearson correlation coefficient and Bland-Altman analysis are performed. Figure 5 shows the correlation and Bland-Altman plots for the automated IVC of the IVC quantification and RAP estimation network 210 compared to the standard manual IVC quantification. From Figure 5, it is observed that the IVC quantification extracted by the IVC quantification and RAP estimation network 210 has a high correlation with the values calculated by human experts.
[0028] To estimate RAP, the IVC quantification and RAP estimation network 210 measures the collapse rate of the IVC. Next, the IVC quantification and RAP estimation network 210 substitutes the absolute value and the ratio of the collapse rate of the IVC into Equation (1) to estimate RAP. To evaluate the effectiveness of the RAP estimation performed by the IVC quantification and RAP estimation network 210, the automated RAP values are compared with the RAP values estimated by experts. Figure 6 shows the confusion matrix indicating this comparison. From the confusion matrix in Figure 6, it is concluded that the IVC quantification and RAP estimation network 210 can accurately estimate the RAP values based on the segmented IVC regions.
[0029] The automatically calculated IVC can then be combined with other clinical and laboratory biomarkers for diagnostic purposes. Such automation can significantly enhance current practices, especially in low- and middle-resource settings where there is a lack of expertise and limited availability of advanced testing and imaging diagnostic resources. An additional advantage of this automation technology is that it enables standardized and individualized diagnosis and prognosis based on the IVC and other testing and clinical biomarkers. Other advantages include the potential to save physicians' time by accelerating and enabling the streamlining of mundane tasks (such as visual IVC observation) in clinical practice, allowing them to focus on innovation and discovery. Furthermore, the ability to simultaneously monitor and analyze multiple data sources by artificial intelligence (AI)-based technologies plays an important role in preventive medicine and may lead to even better patient outcomes. Further details regarding the components of the present invention, including IVC search, segmentation, diameter quantification, and RAP estimation, can be found in Ghada Zamzmi, Sivarama Krishanan Rajaraman, Li-Yueh Hsu, Vandana Sachdev, and Sameer Antani, Real time Echocardiography Image Analysis and Quantification of Cardiac Indices, Medical Image Analysis 80(2022)102438, which is incorporated herein by reference.
[0030] Lightweight System for Real-Time Quantitative Echocardiographic Assessment The above-described system 200 from FIG. 2 is designed to be efficient in terms of space and speed. This efficiency results from the efficient design of the classification and segmentation algorithms. The structural components and related functions of system 200 will be described below.
[0031] Lightweight Search Network The IVC view classification and quality assessment 206 is performed by the IVC search network 700 shown in FIG. 7. As described above, the IVC search network 700 uses a single shared encoder 702 as well as two task-specific heads 704 and 706. In the illustrated embodiment, each head has three task-specific layers. In certain embodiments, the two heads 704 and 706 can operate simultaneously to further enhance the efficiency of the IVC search network 700. Table 3 below shows the size and inference time of the IVC search network 700 according to certain embodiments.
[0032] [Table 3]
[0033] Lightweight segmentation network Regarding the IVC region segmentation network 208 (see FIG. 2) implemented in the form of TaNet, from the above experimental results, it has been shown that the proposed network achieves very fast performance in segmenting the IVC (i.e., TaNet achieved an IVC segmentation rate of 85 frames per second (FPS) in certain embodiments). Therefore, the IVC region segmentation network 208 is a lightweight network.
[0034] The IVC region segmentation network 208 (hereinafter referred to as TaNet208) is used to localize the IVC region. TaNet208 uses a localization component and three pathways for learning rich texture, low-level, and context features. TaNet208 is fine-tuned end-to-end to learn localization and segmentation. The co-learning of localization and segmentation within TaNet208 prevents unnecessary repetition of training the individual models separately and enables TaNet208 to focus on a specific IVC region. By using a single network for both region localization and segmentation, the efficiency of the system 200 is improved. A diagram of the efficient TaNet208 is shown in FIG. 8.
[0035] TaNet for Localization and Segmentation of the IVC Region Convolutional neural networks (CNNs) often operate on the entire image and are limited by the spatial invariance of the input data. Traditional approaches to address these issues involve using separate models for spatial transformation and localization. Jaderberg et al. [6] proposed a more efficient spatial transformation network called STN that applies these transformations (e.g., scaling, translation, attention) to the input image or feature map without additional training supervision. STN is a plug-and-play module that can be inserted into existing CNNs. It is also differentiable in the sense that it calculates the derivatives of the transformations within the module, enabling the learning of the loss gradients with respect to the module parameters. In IVC segmentation, the target region occupies a relatively small portion of the entire image (see FIG. 1). Therefore, considering the entire image for segmentation adds noise due to irrelevant regions. Thus, TaNet208 uses STN802 to focus the segmentation attention on specific regions while suppressing irrelevant regions.
[0036] IVC Segmentation Segmentation is performed using three paths: the spatial or detailed path 804, the handcrafted path 806, and the context path 808. Each of these paths extracts a unique set of features. The spatial path 804 (which in certain embodiments is shallow having only three convolutional layers with high channel capacity) extracts rich low-level details at low computational cost. Specifically, this path has three blocks, each block containing a 3×3 convolutional layer with a stride of 2, followed by batch normalization and ReLU activation. In certain embodiments, the number of filters in the first, second, and third blocks are 64, 64, and 128 respectively. In these embodiments, this path outputs a feature map that is 1 / 8 the size of the input image size.
[0037] The handcrafted path 806 is a shallow path that has only three local binary pattern (LBP)-encoded convolutional layers and extracts rich texture features from the echo images. The mathematical formulation of these LBP-encoded convolutional kernels is shown in [1]. Each LBP block has a layer with fixed anchor weights (m), followed by a second layer with a learnable convolutional filter of size 1×1. The anchor weights are probabilistically generated with various ranges of sparsity. Similar to propagating gradients through layers with learnable and non-learnable parameters (e.g., ReLU, max pooling), the entire path can be trained by backpropagating gradients through both the learnable weights and the anchor weights. In other words, the anchor weights remain unaffected and only the weights of the learnable filters are updated. Compared with the spatial path 804, the handcrafted path 806 is designed to extract specific features (e.g., texture, geometric features) different from those extracted by the spatial path 804. For example, texture features (e.g., LBP) have a strong ability to distinguish small differences in texture and topography, especially at the boundaries between complex regions where separation is particularly difficult. Therefore, the handcrafted path 806 is used to extract rich texture features in different orientations while enhancing the spatial path 804 without increasing the computational burden, while the spatial path 804 is used to extract general low-level details from the image.
[0038] Finally, the context path 808 is used for fast downsampling of the feature map to obtain a sufficient receptive field. In the context path, a lightweight model (e.g., SqueezeNet) is used for fast downsampling of the feature map of the input image, a sufficient receptive field is obtained, and high-level context information is encoded. Then, global average pooling is added at the end of the lightweight model to provide the largest receptive field with global context information. In segmentation, the network analyzes the feature map of the input image in various receptive fields. The receptive field provides an indicator of the extent of the range of input data that neurons or units within a layer can expose to, and is determined by the filter size of the layer within the convolutional neural network. The context path enables the analysis of images in receptive fields of various sizes, and it starts from the image and continues with downsampling. Specifically, since downsampling can reduce the feature representation and expand the receptive field, a lightweight mode is used for downsampling the feature representation of the image. Global average pooling provides the largest receptive field with global context information. Finally, the output of global average pooling is upsampled, combined with the output of other paths, and after extracting features from the three paths, they are combined.
[0039] The process described above can be performed by simply summing the feature representations; however, since the embeddings differ in nature and length, it can degrade performance via the summation approach and complicate network optimization. Thus, in other embodiments, a route fusion 810 is performed to efficiently combine these routes. FIG. 9 shows the steps performed by the route fusion 810. The route fusion 810 first fuses the features from the three routes by concatenating the outputs of the routes at step 902, and then uses batch normalization to balance the various scales of the features. Next, at step 904, the concatenated features are combined into a single feature vector. This feature vector is sent to global pooling at step 906, followed by a convolutional layer (1×1) at step 908, where a rectified linear unit (ReLU) activation is performed at step 910, a convolutional layer (1×1) is applied at step 912, and finally a sigmoid function is used at step 914 to generate a weight vector. This weight vector is used to re-weight the concatenated feature vector.
[0040] Training of TaNet for Joint IVC Localization and Segmentation TaNet208 (see FIG. 8) is trained in two stages: pre-training and fine-tuning. In the pre-training stage, there are two steps: 1) pre-training of the Coarse Segmentation Model (CSM), and 2) pre-training of the Localization Network (L). The Coarse Segmentation Model (CSM) is trained to obtain a rough prediction of the IVC Region of Interest (ROI). Next, the Localization Network (L) is trained to estimate the approximate position of the IVC ROI. In certain embodiments, the Coarse Segmentation Model (CSM) is trained with a batch size of 32 to generate a rough ROI. This can be performed using the Adam optimizer to minimize the loss between the GT mask and the predicted coarse segmentation mask. Next, the output (z) of the coarse segmentation is used as the input to the Localization Network (L). In certain embodiments, the Localization Network (L) is trained with a batch size of 32 to optimize the Smooth L1 loss. The Smooth L1 loss is commonly used in box regression and is less sensitive to outliers. The Localization Network aims to minimize the Smooth L1 loss between the prediction and the ground truth. In the end-to-end fine-tuning stage, the entire network is fine-tuned end-to-end using the Adam optimizer with the pre-trained parameters (stage 1) loaded. The goal of optimization is to minimize the loss between the IVC prediction and the IVC ground truth label.
[0041] Table 3 above shows the size and inference time of the IVC segmentation network. The small size and high inference speed are due to the three shallow paths 804, 806, and 808, each of which has only three layers. Also, the fact that these paths 804, 806, and 808 extract features simultaneously also contributes to the further improvement of the efficiency of the IVC Region Segmentation Network 200 (see FIG. 2). This efficient design of the IVC Region Segmentation Network 208 enables the use in devices with limited resources and for other echo applications, not just for IVC segmentation.
[0042] IVC Quantification The IVC quantification and RAP estimation network 210 performs distance calculations to determine the diameter of each frame, adding low computational complexity to system 200. As shown in Table 3, system 200 is shallow, has small classification, segmentation, and quantification components, and is lightweight. By being lightweight, the system can perform real-time IVC diameter and RAP estimation in each frame, as well as real-time quantification in other echo applications.
[0043] Open-World System for Active Cardiac Echo Interpretation and Learning Traditional machine learning models require that samples of a specific set of classes be available during training. Therefore, when faced with new data not considered during training, these models fail or their performance degrades. To overcome this challenge, system 200 is designed to operate efficiently in an open-world clinical setting. Existing systems for automatic cardiac echo analysis are designed under the assumption that examples in the test or deployment phases must belong to the same limited number of classes that appeared in the training phase. This assumption, known as a closed-world setting, is too restrictive for real-world environments that are open and often have unseen examples, which can dramatically weaken the robustness of machine learning classifiers. Here, system 200 integrates an open-world active learning approach into the IVC cardiac echo system by providing a feedback network. The open-world approach is described below for IVC, but this approach can be integrated and used for other cardiac echo applications.
[0044] The open-world active learning approach disclosed in this specification has two main engines: a classification engine and a clustering engine. The classification engine in this system consists of a view classification model, a quality assessment model, and a disease diagnosis model. Each of these classification engines has a corresponding clustering engine that clusters new or unknown instances based on similarity or other clustering metrics and then sends these clusters to human experts to obtain feedback. For example, the view classification engine contains a classifier trained to recognize various echocardiogram views, including IVC views, but it is also trained using OpenMax (described later) to recognize unknown views. In each operation, the view classifier either detects a specific view or labels it as an unknown view. The clustering engine then groups the unknown views into clusters (based on their similarity), and after being labeled by a human expert, the newly labeled clusters / classes are passed to the classification engine to update the model. In the case of quality assessment, if the classification engine cannot reliably provide a quality label for a specific image / view, that image is labeled as unknown, similar unknown images are grouped together, and then these images are sent to a human expert to obtain feedback. At the final stage of the system, the open-world active learning approach is integrated into a disease classification model that uses automated IVC diameter and collapse rate. This disease classification model classifies known heart diseases into their respective classes and identifies new (unseen) diseases as unknown. Instances with unknown labels are grouped into different clusters, and each cluster is sent back to a human expert for feedback. Based on the provided feedback, the disease classification model is then updated.
[0045] The importance of open-world learning can be demonstrated in several IVC applications. One possible application is to group rare cases of IVC morphology that do not exist in the training data. Examples of such rare cases include a very dilated IVC with a poor collapse rate due to heart failure, or a very small IVC with a complete collapse rate due to dehydration. Another "unknown" cluster may be IVC images that appear to be collapsed but actually represent artifacts where the image is out of the plane of the ultrasound beam due to respiration. Grouping patient populations with similar IVC collapse rate patterns is a useful application of open-world learning.
[0046] Classification engine, open-world active learning classifier In certain embodiments, the method for identifying unknown classes is thresholding of the Softmax output, i.e., a given input image is labeled as unknown if none of the classes reach a predetermined threshold. The performance of this approach is sensitive to the threshold used. OpenMax predicts unknown classes using the concepts of probability of failure and meta-recognition. Scores from the second-to-last layer (i.e., the fully connected layer) of the CNN are used to estimate whether the input is unknown or "far" from known classes. Inputs that are (distributively) far from known classes are then rejected. Note that the OWE function can be replaced by one of the three methods above (Softmax thresholding, one-vs.-rest layer, and OpenMax), or any new function that can distinguish new classes with different distributions.
[0047] Clustering engine, cluster-based active learning In certain embodiments, a clustering algorithm is used to group similar images of unknown classes into clusters and a human expert performs the labeling. Several methods can be used to cluster images of unknown classes. Examples of these methods include K-medoids, K-Means, and K-Centers. The selected clustering algorithm is used to group the unknown images into cluster representatives. Specifically, the embedded features of the unknown images are used to generate k clusters of unknown classes. The optimal number of clusters Κ can be determined empirically (e.g., the elbow method) or specified by a human expert. After grouping the unknown images into different clusters, instead of labeling all the unknown images, having a cardiologist label each cluster group of the unknown images led to a significant reduction in the time required and the number of human interactions. Finally, the newly labeled (previously unknown) images are sent back to the classification stage and used for enhancing training and updating the open-world classifier.
[0048] Figure 10 shows an iterative process 1000 for training and updating an open-world classifier. At step 1002, process 1000 trains an initial open-world classifier to classify known classes while identifying new (unseen) classes as unknown. At step 1004, process 1000 uses the feature embeddings of unknown images (extracted by encoder 702 (see FIG. 7)) and a clustering algorithm to group similar unknown samples into K clusters. At step 1006, process 1000 presents the K clusters of unknown samples to a human expert for labeling. At step 1008, process 1000 adds the newly labeled samples to the initially labeled data, and at step 1010, process 1000 retrains the open-world classifier using the newly labeled data (returns to step 1002). Process 1000 is repeated to label clusters and use them to update the open-world classifier each time a new cluster of unknown samples is created.
[0049] To evaluate the effectiveness of the open-world approach, the open-world active learning approach is integrated with a view classification model. In a particular embodiment, integrating the OpenMax function with the clustering algorithm improved the performance (F-score) of the view classification model from 0.698 to 0.858 (where the F-score of 0.698 is obtained using a closed-world classifier).
[0050] Automating the assessment of the cardiac echo IVC diameter and collapse rate using a lightweight open-world AI system such as system 200 (see Fig. 2) is a novel and efficient approach that standardizes the estimation of RAP, improves reproducibility, and has a positive impact on clinical management. System 200 is lightweight and can be easily implemented in any echocardiography imaging system for real-time automatic RAP assessment. Furthermore, the open-world active learning capabilities of system 200 enable it to be extended to even broader echocardiogram interpretations and applications.
[0051] Throughout this disclosure, the following references are cited: [1] Rao Muhammad Anwer, Fahad Shahbaz Khan, Joost Van De Weijer, Matthieu Molinier, and Jorma Laaksonen. Binary patterns encoded convolutional neural networks for texture recognition and remote sensing scene classification. ISPRS journal of photogrammetry and remote sensing, 138:74-85, 2018. [2] Abhijit Bendale and Terrance E Boult. Towards open set deep networks. In Proceedings of the IEEE conference on computer vision and pattern recognition, pages 1563-1572, 2016. [3] Morgan Caplan, Arthur Durand, Perrine Bortolotti, Delphine Colling, Julien Goutay, Thibault Duburcq, Elodie Drumez, Anahita Rouze, Saad Nseir, Michael Howsam, et al. Measurement site of inferior vena cava diameter affects the accuracy with which fluid responsiveness can be predicted in spontaneously breathing patients: a post hoc analysis of two prospective cohorts. Annals of intensive care, 10(1):1-10, 2020. [4] Salvatore Di Somma, Silvia Navarin, Stefania Giordano, Francesco Spadini, Giuseppe Lippi, Gian- franco Cervellin, Bryan V Dieffenbach, and Alan S Maisel. The emerging role of biomarkers and bio-impedance in evaluating hydration status in patients with acute heart failure. Clinical Chemistry and Laboratory Medicine (CCLM), 50(12):2093-2105, 2012. [5] Chuanxing Geng, Sheng-jun Huang, and Songcan Chen. Recent advances in open set recognition: A survey. IEEE transactions on pattern analysis and machine intelligence, 43(10):3614-3631, 2020. [6] Max Jaderberg, Karen Simonyan, Andrew Zisserman, et al. Spatial transformer networks. Advances in neural information processing systems, 28, 2015. [7] Roberto M Lang, Luigi P Badano, Victor Mor-Avi, Jonathan Afilalo, Anderson Armstrong, Laura Ernande, Frank A Flachskampf, Elyse Foster, Steven A Goldstein, Tatiana Kuznetsova, et al. Rec- ommendations for cardiac chamber quantification by echocardiography in adults: an update from the american society of echocardiography and the european association of cardiovascular imaging. European Heart Journal-Cardiovascular Imaging, 16(3):233-271, 2015. [8] Mark Sandler, Andrew Howard, Menglong Zhu, Andrey Zhmoginov, and Liang-Chieh Chen. Mo- bilenetv2: Inverted residuals and linear bottlenecks. In Proceedings of the IEEE conference on com- puter vision and pattern recognition, pages 4510-4520, 2018. [9] Changqian Yu, Jingbo Wang, Chao Peng, Changxin Gao, Gang Yu, and Nong Sang. Bisenet: Bilateral segmentation network for real-time semantic segmentation. In Proceedings of the European conference on computer vision (ECCV), pages 325-341, 2018.
[0052] All references cited herein, including publications, patent applications, and patents, are hereby incorporated by reference in their entirety to the same extent as if each reference had been individually and specifically indicated to be incorporated by reference and had been set forth in its entirety herein.
[0053] The use of the terms "a", "an", "the", "at least one", and similar reference terms in the context of describing the present invention (in particular in the context of the following claims) should be construed to cover both the singular and the plural forms, unless otherwise indicated herein or unless clearly contradicted by the context. The use of the term "at least one" followed by a list of one or more items (e.g., "at least one of A and B") should be construed to mean one item selected from the listed items (A or B), or any combination of two or more of the listed items (A and B), unless otherwise indicated herein or unless clearly contradicted by the context. The terms "comprising", "having", "including", and "containing" should be construed as open-ended terms (i.e., having the meaning of "including, but not limited to") unless specifically stated otherwise. The recitation of a range of values herein is intended to serve merely as a shorthand way of referring individually to each separate value falling within the range, and each separate value is incorporated herein as if it were individually recited herein. All methods described herein can be performed in any suitable order, unless otherwise indicated herein or unless clearly contradicted by the context. The use of any examples, or exemplary language (e.g., "such as") provided herein is merely intended to better illustrate the invention and does not limit the scope of the invention unless otherwise claimed. No language in this specification should be construed as indicating any non-claimed element as essential to the practice of the invention.
[0054] Preferred embodiments of the present invention are described herein, including the best mode known to the inventors of practicing the present invention. Variations of those preferred embodiments will be apparent to those skilled in the art upon reading the foregoing description. The inventors expect those skilled in the art to adopt such variations as appropriate, and the inventors intend for the present invention to be practiced otherwise than as specifically described herein. Accordingly, the present invention includes all modifications and equivalents of the subject matter recited in the claims appended hereto as permitted by applicable law. Further, unless otherwise specifically indicated herein or clearly inconsistent with context, combinations in any conceivable variation of the above-described elements are included in the present invention.
Claims
1. An artificial intelligence (AI) system for real-time estimation of inferior vena cava (IVC) collapse rate and right atrial pressure (RAP) from echocardiography, wherein the system: The system has an image retrieval network, which is configured to receive images from echocardiography and perform quality evaluation of the images to generate selected images; The system has a region segmentation network, which is configured to localize and segment the selected image to obtain the IVC region in the selected image; It has an IVC quantification and RAP estimation network, which is configured to perform distance calculations to determine the diameter of the IVC region at different spatial (different locations) and temporal (over time) points; and, It has an open-world active learning system, and the open-world active learning system has a classification engine and a clustering engine. The aforementioned system.
2. The system according to claim 1, wherein the classification engine is configured to detect known views and unknown views in the selected images.
3. The system according to claim 2, wherein the clustering engine is configured to group the unknown views into one or more clusters for use in updating the classification engine.
4. The aforementioned open-world active learning system: The system is configured to classify known views into known classes and unknown views into unknown classes; The system is configured to group the unknown views of the unknown class based on the feature embedding of the selected images determined by the image search network; It is configured to obtain expert labeling for the aforementioned unknown view of the aforementioned unknown class; The system is configured to add the aforementioned expert labeling to the aforementioned unknown view of the aforementioned unknown class to determine the updated class; and, The classification engine of the open-world active learning system is configured to retrain in order to know the updated class. The system according to claim 3.
5. The system according to claim 1, wherein the region segmentation network has a spatial transformation network for focusing on a specific region of the selected image.
6. The aforementioned domain segmentation network further: It has a spatial path for extracting low-level details of the aforementioned specific region of the selected image; It has a handcrafted path for extracting the rich texture features of the aforementioned specific region of the selected image; and, The system according to claim 5, further comprising a contextual pathway for obtaining a receptive field by downsampling a feature map.
7. The system according to claim 6, wherein the outputs of the spatial path, the handcraft path, and the context path are combined using path fusion to obtain a weighted feature vector.
8. The aforementioned pathway fusion is: The process involves concatenating the outputs of the spatial path, the handcraft path, and the context path to obtain a concatenated feature; It has the ability to combine the aforementioned linked features into a feature vector; The feature vector is sent to the global pooling; The process involves performing a 1x1 convolution on the feature vector after the global pooling; The process involves performing normalized linear unit (ReLU) activation on the feature vector after the 1x1 convolution; The process involves performing a second 1x1 convolution on the feature vector after the ReLU activation; It has the ability to generate a weight vector by applying a sigmoid function; and, The process involves re-weighting the feature vector based on the weight vector to obtain the weighted feature vector. The system according to claim 7.
9. The aforementioned image search network: It has a shared encoder; It has a first task-specific head; and, It has a second task-specific head; Here, the first task-specific head and the second task-specific head each have at least three task-specific layers. The system according to claim 1.
10. The system according to claim 1, wherein the domain segmentation network is a lightweight domain segmentation network configured to achieve an IVC segmentation rate of 85 frames per second (FPS).
11. A method for estimating inferior vena cava (IVC) collapse rate and right atrial pressure (RAP) from echocardiographic examinations performed in real time by an artificial intelligence (AI) system, the method being: The system includes receiving images from the echocardiography examination, performing quality evaluation of the images, and generating selected images; The process involves localizing and segmenting the selected image to obtain the IVC region within the selected image; This involves performing distance calculations to determine the diameter of the IVC region at different spatial and temporal points; The system includes classifying known and unknown views in the selected image; and, The method involves grouping the unknown views into one or more clusters used to update the classification capabilities of the AI system. The aforementioned method.
12. moreover: The system includes classifying known views into known classes and unknown views into unknown classes; The method includes grouping the unknown views of the unknown class based on the feature embedding of the selected images; It has the ability to obtain expert labeling for the aforementioned unknown view of the aforementioned unknown class; The process involves adding the aforementioned expert labeling to the aforementioned unknown view of the aforementioned unknown class to determine the updated class; and, This involves retraining the AI system to know the updated class. The method according to claim 11.
13. The selected image can be localized and segmented to obtain the IVC region in the selected image: It has the ability to extract low-level details from a specific region of the selected image; It comprises extracting rich texture features from the aforementioned specific region of the selected image; and, This involves downsampling the feature map to obtain the receptive field. The method according to claim 11.
14. The method according to claim 13, further comprising combining the low-level details of the particular region of the selected image, the rich texture features of the particular region of the selected image, and the receptive field using path fusion to obtain a weighted feature vector.
15. The low-level details of the aforementioned specific region of the selected image, the rich texture features of the aforementioned specific region of the selected image, and the receptive field are combined using path fusion: The method involves concatenating the low-level details of the specific region of the selected image, the rich texture features of the specific region of the selected image, and the receptive field to obtain a concatenated feature; It has the ability to combine the aforementioned linked features into a feature vector; The feature vector is sent to the global pooling; The process involves performing a 1x1 convolution on the feature vector after the global pooling; The process involves performing normalized linear unit (ReLU) activation on the feature vector after the 1x1 convolution; The process involves performing a second 1x1 convolution on the feature vector after the ReLU activation; It has the ability to generate a weight vector by applying a sigmoid function; and, The process involves re-weighting the feature vector based on the weight vector to obtain the weighted feature vector. The method according to claim 14.
16. A non-transient computer-readable medium implemented in an artificial intelligence (AI) system for real-time estimation of inferior vena cava (IVC) collapse rate and right atrial pressure (RAP) from echocardiography, wherein the non-transient computer-readable medium has instructions, and when executed by the computer, the instructions set the computer to perform the following steps, the steps being: The process includes receiving images from the echocardiography and performing quality evaluation of the images to generate selected images; The step of localizing and segmenting the selected image to obtain an IVC region in the selected image; The process includes the step of performing distance calculations to determine the diameter of the IVC region at different spatial and temporal points; The step includes classifying known views and unknown views in the selected image; and, The process includes the step of grouping the unknown views into one or more clusters used to update the classification capabilities of the AI system. The aforementioned non-transient computer-readable medium.
17. The instruction further has, when executed, set the computer to perform the following further steps, which are: The process includes the steps of classifying the known views into known classes and the unknown views into unknown classes; The process includes the step of grouping the unknown views of the unknown class based on the feature embedding of the selected images; The process includes the step of obtaining expert labeling for the aforementioned unknown view of the aforementioned unknown class; The step of determining the updated class is to add the aforementioned expert labeling to the aforementioned unknown view of the aforementioned unknown class; and, The step includes retraining the AI system to know the updated class, The non-transient computer-readable medium according to claim 16.
18. The steps of localizing and segmenting the selected image to obtain the IVC region in the selected image are: It has the ability to extract low-level details from a specific region of the selected image; It comprises extracting rich texture features from the aforementioned specific region of the selected image; and, This involves downsampling the feature map to obtain the receptive field. The non-transient computer-readable medium according to claim 16.
19. A non-transient computer-readable medium according to claim 18, further comprising instructions, which, when executed, configure the computer to perform the following further steps, the further steps comprising combining the low-level details of the particular region of the selected image, the rich texture features of the particular region of the selected image and the receptive field using path fusion to obtain a weighted feature vector.
20. The steps include combining the low-level details of the specific region of the selected image, the rich texture features of the specific region of the selected image, and the receptive field using path fusion: The method involves concatenating the low-level details of the specific region of the selected image, the rich texture features of the specific region of the selected image, and the receptive field to obtain a concatenated feature; It has the ability to combine the aforementioned linked features into a feature vector; The feature vector is sent to the global pooling; The process involves performing a 1x1 convolution on the feature vector after the global pooling; The process involves performing normalized linear unit (ReLU) activation on the feature vector after the 1x1 convolution; The process involves performing a second 1x1 convolution on the feature vector after the ReLU activation; It has the ability to generate a weight vector by applying a sigmoid function; and, The process involves re-weighting the feature vector based on the weight vector to obtain the weighted feature vector. The non-transient computer-readable medium according to claim 19.