Methods and systems for mitral valve detection using generative modeling
By classifying cardiac ultrasound videos using a generative model based on variational autoencoders and automatically detecting mitral valves using reconstruction error differences, the problem of early detection of mitral aortic valves has been solved, achieving efficient mitral valve identification.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- GE PRECISION HEALTHCARE LLC
- Filing Date
- 2022-03-16
- Publication Date
- 2026-07-21
AI Technical Summary
Existing technologies make it difficult to detect the relatively rare mitral aortic valve in routine screening procedures, especially since it is difficult to visually distinguish from the tricuspid valve leaflets, making early detection challenging.
A generative model based on variational autoencoder is used to classify cardiac ultrasound videos. The model is trained to identify cardiac structures in ultrasound image frames, and the mitral and tricuspid valves are automatically detected by utilizing reconstruction error differences. The classification results of multiple frames are aggregated to improve accuracy.
It enables automatic detection of the mitral aortic valve during routine cardiac ultrasound screening, improving the accuracy and reliability of mitral valve detection and overcoming the problem of data imbalance.
Smart Images

Figure CN115192074B_ABST
Abstract
Description
Technical Field
[0001] The embodiments of the subject matter disclosed herein relate to ultrasound imaging, and more specifically to the detection of the mitral aortic valve using ultrasound imaging. Background Technology
[0002] Aortic stenosis, or aortic narrowing, occurs when the aortic valve in the heart narrows, resulting in an increased blood flow pressure gradient at the aortic valve orifice. The aortic valve does not open completely, thus reducing blood flow from the heart to the aorta. The likelihood of aortic stenosis is significantly increased if the aortic valve is mitral rather than tricuspid, where the aortic valve consists of only two leaflets instead of the typical three. This type of mitral aortic valve is relatively rare and occurs in less than 2% of the normal population.
[0003] Ultrasound is the primary imaging modality used for screening aortic stenosis. Early detection of mitral aortic valve disease can enable proactive intervention and treatment to potentially reduce the likelihood of aortic stenosis. However, due to its low incidence and the difficulty in visually distinguishing mitral leaflets from tricuspid leaflets, the presence of mitral aortic valve disease is challenging to detect during routine screening procedures. Summary of the Invention
[0004] In one implementation, a method includes acquiring ultrasound video of the heart during at least one cardiac cycle, identifying frames in the ultrasound video corresponding to at least one cardiac phase, and classifying cardiac structures in the identified frames as mitral or tricuspid valves. Generative models, such as variational autoencoders trained on ultrasound image frames at at least one cardiac phase, can be used to classify cardiac structures. In this way, the relatively rare occurrence of mitral aortic valve can be automatically detected during routine cardiac ultrasound screening.
[0005] It should be understood that the above brief description is provided to introduce selected concepts further described in the detailed embodiments in a simplified form. This is not intended to identify key or essential features of the claimed subject matter, the scope of which is uniquely defined by the claims following the detailed embodiments. Furthermore, the claimed subject matter is not limited to embodiments that address any shortcomings mentioned above or in any part of this disclosure. Attached Figure Description
[0006] The invention will be better understood by referring to the following description of non-limiting embodiments, in which:
[0007] Figure 1 A block diagram illustrating an exemplary ultrasound imaging system according to one embodiment is shown;
[0008] Figure 2A block diagram of an exemplary image processing system according to one embodiment is shown, the image processing system being configured to classify anatomical structures in ultrasound image frames using a generative model;
[0009] Figure 3 A block diagram illustrating an exemplary architecture for generating a model according to one embodiment is shown, the model being configured to classify anatomical structures in ultrasound image frames;
[0010] Figure 4 A high-level flowchart illustrating an exemplary method for classifying anatomical structures in ultrasound video according to one embodiment is shown;
[0011] Figure 5 A high-level flowchart illustrating an exemplary method for identifying keyframes including anatomical structures in a given cardiac phase of ultrasound video, according to one embodiment, is shown.
[0012] Figure 6 A set of exemplary ultrasound images, according to one embodiment, illustrates exemplary identification of keyframes including anatomical structures at a given cardiac phase;
[0013] Figure 7 A high-level flowchart illustrating an exemplary method for classifying anatomical structures as mitral or tricuspid valves, according to one embodiment, is shown.
[0014] Figure 8 The illustration shows a set of images illustrating keyframes in an ultrasound video, generated images, and a set of images showing how reconstruction errors lead to tricuspid valve classification, according to one embodiment; and
[0015] Figure 9 The diagram illustrates a set of images showing the generated images of keyframes in an ultrasound video and the classification of the mitral valve caused by reconstruction errors, according to one embodiment. Detailed Implementation
[0016] The following describes various implementation schemes involving ultrasound imaging. Specifically, a method and system for automated mitral valve detection using an ultrasound imaging system are provided. Figure 1 An example of an ultrasound imaging system that can be used to acquire images according to the technology of the present invention is shown. The ultrasound imaging system can be communicatively coupled to an image processing system, such as... Figure 2 An image processing system. An image processing system may include one or more deep learning models, such as one or more neural networks and one or more generative models, stored in non-transitory memory. For example... Figure 3 As shown, the generative model configured as a variational autoencoder can be trained based on ultrasound images of the tricuspid valve at a selected cardiac phase. Ultrasound videos comprising frames or sequences of ultrasound images can then be processed, for example... Figure 4 As shown, this type of trained generative model classifies cardiac structures in ultrasound videos into mitral or tricuspid valves. Cardiac structures in ultrasound videos can be extracted, such as... Figure 5 and Figure 6 As shown, the frame corresponding to the selected cardiac phase can then be identified. (See figure) Figure 7 As shown, a method for classifying cardiac structures using a generative model can then include using extracted ultrasound images of cardiac structures at selected cardiac phases as input to the corresponding generative model. Since the generative model is trained based on images of the tricuspid valve, if the input image depicts the mitral valve, the image output by the generative model can have a relatively high reconstruction error. By processing multiple ultrasound images acquired over multiple cardiac cycles and obtaining multiple classifications of cardiac structures as either the mitral or tricuspid valve, individual classifications can be aggregated to determine the final classification of the cardiac structure as either the mitral or tricuspid valve, as shown. Figure 8 and Figure 9 As shown. Therefore, despite the data imbalance problem, where less training data is available for the mitral valve compared to the tricuspid valve, which would normally raise significant questions about training deep learning models to classify heart structures as mitral valves, it is still possible to use deep learning models to achieve automatic detection of the mitral valve.
[0017] Now turn to the diagrams, Figure 1 A schematic diagram of an ultrasound imaging system 100 according to one embodiment of the present disclosure is shown. The ultrasound imaging system 100 includes a transmit beamformer 101 and a transmitter 102 that drives elements (e.g., transducer elements) 104 within a transducer array (referred to herein as probe 106) to transmit pulsed ultrasound signals (referred herein as transmit pulses) into a body (not shown). According to one embodiment, probe 106 may be a one-dimensional transducer array probe. However, in some embodiments, probe 106 may be a two-dimensional matrix transducer array probe. As further explained below, transducer element 104 may be made of a piezoelectric material. When a voltage is applied to a piezoelectric crystal, the crystal physically expands and contracts, thereby emitting ultrasonic spherical waves. In this way, transducer element 104 can convert an electronic emission signal into an acoustic emission beam.
[0018] After the element 104 of probe 106 transmits a pulsed ultrasound signal into the patient's body, the pulsed ultrasound signal is backscattered from internal structures (such as blood cells or muscle tissue) to generate an echo returning to element 104. The echo is converted into an electrical signal or ultrasound data by element 104, and the electrical signal is received by receiver 108. The electrical signal representing the received echo passes through receiver beamformer 110, which outputs ultrasound data. Additionally, transducer element 104 can generate one or more ultrasound pulses based on the received echo to form one or more transmit beams.
[0019] According to some embodiments, probe 106 may include all or part of electronic circuitry to perform transmit beamforming and / or receive beamforming. For example, all or part of transmit beamformer 101, transmitter 102, receiver 108, and receive beamformer 110 may be located within probe 106. In this disclosure, the terms “scanning” or “under scanning” may also be used to refer to the process of acquiring data by transmitting and receiving ultrasound signals. In this disclosure, the term “data” may be used to refer to one or more datasets acquired using an ultrasound imaging system. In one embodiment, data acquired via ultrasound system 100 may be used to train a machine learning model. User interface 115 may be used to control the operation of ultrasound imaging system 100, including for controlling the input of patient data (e.g., patient history), for changing scan or display parameters, for initiating probe repolarization sequences, etc. User interface 115 may include one or more of the following: a rotary element, a mouse, a keyboard, a trackball, hard keys linked to specific actions, soft keys configurable to control different functions, and a graphical user interface displayed on display device 118.
[0020] The ultrasound imaging system 100 also includes a processor 116 for controlling the transmitting beamformer 101, the transmitter 102, the receiver 108, and the receiving beamformer 110. The processor 116 communicates electronically (e.g., is communicatively connected) with the probe 106. For the purposes of this disclosure, the term "electronic communication" may be defined to include both wired and wireless communication. The processor 116 can control the probe 106 to acquire data according to instructions stored in the processor 116's memory and / or memory 120. The processor 116 controls which of the elements 104 are active and the shape of the beam emitted from the probe 106. The processor 116 also communicates electronically with a display device 118, and the processor 116 can process data (e.g., ultrasound data) into images for display on the display device 118. The processor 116 may include a central processing unit (CPU) according to one embodiment. According to other embodiments, the processor 116 may include other electronic components capable of performing processing functions, such as a digital signal processor, a field-programmable gate array (FPGA), or a graphics board. According to other embodiments, processor 116 may include multiple electronic components capable of performing processing functions. For example, processor 116 may include two or more electronic components selected from a list of electronic components, including: a central processing unit, a digital signal processor, a field-programmable gate array, and a graphics board. According to another embodiment, processor 116 may also include a composite demodulator (not shown) that demodulates RF data and generates raw data. In another embodiment, demodulation may be performed earlier in the processing chain. Processor 116 is adapted to perform one or more processing operations based on multiple selectable ultrasound modalities on the data. In one example, data may be processed in real time during a scanning session because the echo signal is received by receiver 108 and transmitted to processor 116. For the purposes of this disclosure, the term "real time" is defined as including a procedure executed without any intentional delay. For example, embodiments may acquire images at a real-time rate of 7 frames / second to 20 frames / second. Ultrasonic imaging system 100 is capable of acquiring 2D data of one or more planes at a significantly faster rate. However, it should be understood that the real-time frame rate may depend on the length of time spent acquiring each frame of data for display. Therefore, the real-time frame rate may be slow when acquiring relatively large amounts of data. Therefore, some embodiments may have a real-time frame rate significantly faster than 20 frames per second, while other embodiments may have a real-time frame rate less than 7 frames per second. Data may be temporarily stored in a buffer (not shown) during a scanning session and processed in a less real-time manner in real-time or offline operation. Some embodiments of the invention may include multiple processors (not shown) to handle the processing tasks processed by processor 116 according to the exemplary embodiments described above.For example, before displaying an image, a first processor can be used to demodulate and extract the RF signal, while a second processor can be used to further process the data (e.g., by augmenting the data as further described herein). It should be understood that other embodiments may use different processor arrangements.
[0021] The ultrasound imaging system 100 can continuously acquire data at frame rates, for example, from 10 Hz to 30 Hz (e.g., 10 to 30 frames per second). Images generated from the data can be refreshed on a display device 118 at a similar frame rate. Other embodiments are capable of acquiring and displaying data at different rates. For example, depending on the frame size and the intended application, some embodiments may acquire data at frame rates less than 10 Hz or greater than 30 Hz. A memory 120 is included for storing frames of processed acquired data. In an exemplary embodiment, the memory 120 has sufficient capacity to store at least several seconds of ultrasound data frames. The data frames are stored in a manner that facilitates retrieval based on their acquisition order or time. The memory 120 may include any known data storage medium.
[0022] In various embodiments of the invention, processor 116 can process data through different mode-related modules (e.g., B-mode, color Doppler, M-mode, color M-mode, spectral Doppler, elastography, TVI, strain, strain rate, etc.) to form 2D or 3D data. For example, one or more modules can generate B-mode, color Doppler, M-mode, color M-mode, spectral Doppler, elastography, TVI, strain, strain rate, and combinations thereof. As an example, one or more modules can process color Doppler data, which may include conventional color flow Doppler, power Doppler, HD flow, etc. Image lines and / or frames are stored in memory and may include timing information indicating the time when image lines and / or frames are stored in memory. These modules may include, for example, a scan conversion module for performing scan conversion operations to convert the acquired images from beam space coordinates to display space coordinates. A video processor module may be provided that reads the acquired images from memory and displays the images in real time while performing procedures (e.g., ultrasound imaging) on a patient. The video processor module may include a separate image memory, and ultrasound images may be written to the image memory for reading and display by the display device 118.
[0023] In various embodiments of this disclosure, one or more components of the ultrasound imaging system 100 may be included in a portable handheld ultrasound imaging device. For example, a display device 118 and a user interface 115 may be integrated into the external surface of the handheld ultrasound imaging device, which may also include a processor 116 and a memory 120. A probe 106 may include a handheld probe that electronically communicates with the handheld ultrasound imaging device to collect raw ultrasound data. Transmit beamformer 101, transmitter 102, receiver 108, and receive beamformer 110 may be included in the same or different parts of the ultrasound imaging system 100. For example, transmit beamformer 101, transmitter 102, receiver 108, and receive beamformer 110 may be included in the handheld ultrasound imaging device, the probe, and combinations thereof.
[0024] After performing a two-dimensional ultrasound scan, a data block containing scan lines and their samples is generated. Following the application of a back-end filter, a process called scan transformation is performed to convert the two-dimensional data block into a displayable bitmap image with additional scan information, such as depth, the angle of each scan line, etc. During scan transformation, interpolation techniques are applied to fill in missing holes (i.e., pixels) in the resulting image. These missing pixels occur because each element of the two-dimensional block should typically cover many pixels in the resulting image. For example, in current ultrasound imaging systems, bicubic interpolation is used, which utilizes the adjacent elements of the two-dimensional block. Therefore, if the two-dimensional block is relatively small compared to the size of the bitmap image, the scan-transformed image will include areas of poor or low resolution, especially for deeper regions.
[0025] Figure 2A block diagram illustrating an exemplary system 200 according to one embodiment is shown. This exemplary system includes an image processing system 202 configured to classify anatomical structures in ultrasound image frames. In some embodiments, the image processing system 202 may be disposed within an ultrasound imaging system 100 as a processor 116 and a memory 120. In some embodiments, at least a portion of the image processing system 202 is disposed at a device (e.g., an edge device, server, etc.) communicatively coupled to the ultrasound imaging system 100 via a wired and / or wireless connection. In some embodiments, at least a portion of the image processing system 202 is disposed at a separate device (e.g., a workstation) that may receive images from the ultrasound imaging system 100 or from a storage device (not shown) storing image data generated by the ultrasound imaging system 100. The image processing system 202 may be communicatively coupled to a user input device 232 and a display device 234. In some examples, the user input device 232 may include a user interface 115 of the ultrasound imaging system 100, while the display device 234 may include a display device 118 of the ultrasound imaging system 100. For example, the image processing system 202 may be further communicatively coupled to the ultrasound probe 236, which may include the ultrasound probe 106 described above.
[0026] Image processing system 202 includes processor 204 configured to execute machine-readable instructions stored in non-transitory memory 206 of image processing system 202. Processor 204 may be single-core or multi-core, and programs executing thereon may be configured for parallel or distributed processing. In some embodiments, processor 204 may optionally include individual components distributed across two or more devices, which may be remotely located and / or configured for coordinated processing. In some embodiments, one or more aspects of processor 204 may be virtualized and executed by a remotely accessible networked computing device configured in a cloud computing configuration.
[0027] Non-transitory memory 206 may store keyframe extraction module 210, classification module 220, and ultrasound image data 226. Keyframe extraction module 210 is configured to extract keyframes including regions of interest (ROIs) from ultrasound image data 226. Keyframe extraction module 210 may therefore include region of interest (ROI) model module 212 and cardiac phase recognition module 214. ROI model module 212 may include a deep learning model (e.g., a deep learning neural network) and instructions for implementing the deep learning model to identify desired ROIs within ultrasound images. ROI model module 212 may include trained and / or untrained neural networks, and may also include various data or metadata related to one or more neural networks stored therein. As an illustrative and non-limiting example, the deep learning model of ROI model module 212 may include a region-based convolutional neural network (R-CNN), such as masked R-CNN. The target ROI may include the aortic valve region in the heart, and therefore ROI model module 212 may extract the ROI corresponding to the aortic valve region in an ultrasound image of the heart. The ROI model module 212 further aligns the ROI temporally within the ultrasound image sequence (e.g., throughout an ultrasound video containing the ultrasound image sequence) to remove motion outside the motion within the ROI from the probe 236 or the imaged patient.
[0028] The cardiac phase recognition module 214 is configured to identify key cardiac phases in the extracted ROIs from the ROI model module 212. For example, the cardiac phase recognition module 214 processes each extracted ROI from the ROI model module 212 and selects frames corresponding to the key cardiac phases. Key cardiac phases may include one or more of the following: end-diastolic (ED), end-systolic (ES), diastole, and systole. To this end, the cardiac phase recognition module 214 can process the extracted ROIs by modeling the extracted ROI frame sequence as a linear combination of ED and ES frames and applying nonnegative matrix factorization (NMF) to identify keyframes. Therefore, the keyframe extraction module 210 takes ultrasound video (i.e., an ordered sequence of ultrasound image frames) as input and outputs keyframes of the extracted ROIs that are at the key cardiac phases.
[0029] The classification module 220 includes at least two generative models configured to generatively model the Region of Interest (ROI) during key cardiac phases. For example, a first generative model 222 is configured to generatively model a first cardiac phase, while a second generative model 224 is configured to generatively model a second cardiac phase. As an illustrative and non-limiting example, the first generative model 222 can be trained to generatively model the aortic valve during the end-systolic phase, while the second generative model 222 can be trained to generatively model the aortic valve during the end-diastolic phase. In some examples, the classification module 220 may include more than two generative models, where each generative model is trained for a corresponding cardiac phase, while in other examples, the classification module 220 may include a single generative model trained based on a single cardiac phase. Generative models 222 and 224 may include variational autoencoders (VAEs), generative adversarial networks (GANs), or another type of deep generative model. As illustrative examples, generative models 222 and 224 may include a VAE configured to generatively model the tricuspid aortic valve during the end-systolic and end-diastolic phases, respectively. This article is about... Figure 3 An exemplary generative model is further described.
[0030] Figure 3 A block diagram illustrating an exemplary architecture for a generative model 300, according to one embodiment, is shown. This generative model is configured to classify anatomical structures in ultrasound image frames. Specifically, the generative model 300 includes a variational autoencoder configured to process an input ultrasound image 305 and output a generated image 345. For this purpose, the generative model 300 includes an encoder 310 configured to encode the input ultrasound image 305 into a low-dimensional representation 330, and a decoder 340 configured to decode the low-dimensional representation 330 into the generated image 345.
[0031] Encoder 310 includes a neural network configured to accept an input ultrasound image 305 as input and output a representation of the input ultrasound image 305 in a latent space. Therefore, encoder 310 includes at least two layers comprising multiple neurons or nodes, with the last layer of encoder 310 comprising a minimum number of neurons, such that encoder 310 reduces the dimension of the input ultrasound image 305 to the latent space corresponding to the last layer of encoder 310. Instead of outputting the latent variables of the last layer of encoder 310, encoder 310 outputs a mean (μ) tensor 312 and a standard deviation or variance (σ) tensor 314, comprising the mean and standard deviation of the encoded latent distribution, respectively.
[0032] To randomize the latent values from the input ultrasound image 305 encoded as a mean tensor 312 and a variance tensor 314, the generative model 300 randomly samples a distribution 316, which may include a unit Gaussian distribution as an illustrative example, to form a randomly sampled distribution tensor 318. The variance tensor 314 is multiplied by the randomly sampled distribution tensor 318, and then summed 320 with the mean tensor 312 to form a low-dimensional representation 330 of the input ultrasound image 305. This reparameterization of the distribution tensor 318 enables backpropagation of the error, and thus enables the training of the encoder 310.
[0033] The low-dimensional representation 330 of the input ultrasound image 305 is then fed into a decoder 340, which decodes the low-dimensional representation 330 into a generated image 345. The decoder 340 includes a neural network configured to accept the low-dimensional representation 330 as input and output the generated image 345. Therefore, the decoder 340 includes at least one layer comprising a plurality of neurons or nodes.
[0034] During the training of generative model 300, a loss function that includes reconstruction error and Kulback-Leibler (KL) divergence error is minimized. Minimizing the reconstruction error, which includes the difference between the input ultrasound image 305 and the generated image 345, improves the overall performance of the encoder-decoder architecture. The reconstruction error is computed at error computation 350 of generative model 300. Minimizing the KL divergence error in the latent space regularizes the distribution of the encoder 310's output (i.e., the mean tensor 312 and variance tensor 314) so that tensors 312 and 314 approximate a standard normal distribution, such as a unit Gaussian distribution 316. As an example, backpropagation with the loss function and gradient descent can be used to update the weights and biases of encoder 310 and decoder 340 during the training of generative model 300.
[0035] Because the mitral valve is rarer than the tricuspid valve, the generative model 300 is trained based on the input ultrasound image of the tricuspid valve. In this way, when the input ultrasound image 305 includes an image of the tricuspid valve, the trained generative model 300 outputs a generated image 345 of the tricuspid valve, where the reconstruction error, or the difference, between the input ultrasound image 305 and the generated image 345 is minimal. In contrast, when the input ultrasound image 305 includes an image of the mitral valve, the trained generative model 300 outputs a generated image 345 that depicts an anatomical structure closer to the tricuspid valve than the mitral valve. Therefore, the reconstruction error between the input ultrasound image 305 and the generated image 345 can be relatively high and exceed an error threshold because the generative model 300 is trained based on an image of the tricuspid valve, not the mitral valve. As further discussed herein, the mitral valve can be detected by continuously measuring reconstruction errors above such an error threshold using the trained generative model 300. Therefore, error calculation 350 can measure the reconstruction error and compare the measured reconstruction error with an error threshold. Furthermore, the generative model 300 can be trained based on ultrasound images of the tricuspid valve at critical cardiac phases (such as the end-diastolic phase or the end-systolic phase). As discussed above, the generative model 300 can be implemented as a first generative model 222 and a second generative model 224, wherein the first generative model 222 is trained based on ultrasound images of the tricuspid valve at the end-diastolic phase, and the second generative model 224 is trained based on ultrasound images of the tricuspid valve at the end-systolic phase.
[0036] Figure 4 A high-level flowchart illustrating an exemplary method 400 for classifying anatomical structures in an ultrasound video, according to one embodiment, is shown. Specifically, method 400 involves analyzing an ultrasound video to classify cardiac structures depicted in the ultrasound video as mitral or tricuspid valves. (Refer to...) Figures 1 to 3 The system and components described herein constitute method 400, but it should be understood that method 400 may be implemented with other systems and components without departing from the scope of this disclosure. Method 400 may be implemented as executable instructions in non-transitory memory (such as non-transitory memory 206) that, when executed by a processor (such as processor 204), cause the processor to perform the actions described herein.
[0037] Method 400 begins at 405. At 405, method 400 acquires ultrasound video of the heart over at least one cardiac cycle. The ultrasound video may include a sequence of ultrasound image frames of the heart acquired over multiple cardiac cycles. The ultrasound image frames of the ultrasound video may include short-axis images across one or more cardiac cycles, such that the ultrasound image frames depict a cross-sectional view of the heart, including the ventricles and valve annulus. Specifically, the ultrasound image frames may include a cross-sectional view of at least the aortic valve. In some examples, method 400 may acquire ultrasound video by controlling an ultrasound probe (such as ultrasound probe 236), and therefore method 400 may acquire ultrasound video in real time. In other examples, method 400 may acquire ultrasound video by loading ultrasound image data 226 from non-transitory memory 206. In other examples, method 400 may acquire ultrasound video from a remote storage system, such as a Picture Archiving and Communication System (PACS).
[0038] At 410, method 400 identifies frames in the ultrasound video corresponding to the selected cardiac phase. For example, method 400 may identify frames in the ultrasound video corresponding to the systolic and diastolic phases. As an illustrative and non-limiting example, method 400 may identify frames in the ultrasound video corresponding to the end of the systolic and diastolic phases. In one example, identifying frames corresponding to the selected cardiac phase includes: identifying ROIs (such as the aortic valve) within each ultrasound image frame of the ultrasound video; segmenting or extracting the ROI from each ultrasound image frame; and identifying the segmented or extracted ROI corresponding to the key cardiac phase. Figure 5 An exemplary method for identifying frames in an ultrasound video that correspond to a selected cardiac phase is further described.
[0039] At 415, method 400 classifies the cardiac structures in the identified frames as either mitral or tricuspid valves. Method 400 can classify the aortic valve as mitral or tricuspid valves, for example, by inputting each identified frame into a generative model and measuring the reconstruction error based on the output of the generative model. For example, method 400 can input extracted image frames corresponding to a first cardiac phase into a first generative model trained on images for the first cardiac phase, and method 400 can further input extracted image frames corresponding to a second cardiac phase into a second generative model trained on images for the second cardiac phase, such as the first generative model 222 and the second generative model 224 discussed above. The generative model can be trained based on images of the tricuspid valve at the corresponding cardiac phase, such that the reconstruction error may be low when the cardiac structures depicted in the identified frames include the tricuspid valve, but relatively high when the cardiac structures depicted in the identified frames include the mitral valve. Therefore, method 400 can classify the cardiac structures in each identified frame as mitral or tricuspid valves based on whether the reconstruction error is above or below an error threshold.
[0040] As an illustrative and non-limiting example, an image quality metric such as mean squared error (MSE) or mean absolute error (MAE) may be used, and an error threshold may include a value of the image quality metric selected such that an input image with a reconstruction error that causes the image quality metric to exceed the error threshold depicts a mitral valve image.
[0041] Furthermore, to avoid false positives or false negatives, method 400 can further aggregate the number of tricuspid and mitral valve classifications in the ultrasound video to classify the cardiac structure as either tricuspid or mitral. For example, if the ultrasound video includes multiple ultrasound image frames corresponding to a selected cardiac phase, method 400 can classify the cardiac structure as mitral if the number of cardiac structures classified as mitral is greater than or equal to an aggregation error threshold. The aggregation error threshold can include the percentage of multiple ultrasound image frames classified as mitral. For example, if the multiple ultrasound image frames corresponding to a selected cardiac phase include ten ultrasound image frames, but only one frame is classified as mitral, method 400 can classify the cardiac structure as tricuspid because the number of aggregated mitral valve classifications is less than the aggregation error threshold. For example, the aggregation error threshold can be 50%, or in some examples it can be less than 50% or even greater than 50%. For example, the aggregation error threshold can include 100%, such that method 400 can classify the cardiac structure as mitral when all identified frames are classified as mitral. Figure 7 An exemplary method for classifying cardiac structures in the identified frames as either mitral or tricuspid valves is further described.
[0042] At 420, method 400 outputs a classification of the cardiac structure. For example, method 400 outputs a aggregate classification of the cardiac structure identified as a mitral or tricuspid valve at 415. For example, method 400 may output the classification to display devices 118 or 234. Alternatively, method 400 may output the classification to, for example, non-transitory memory 120 or 206, and / or to a PACS for reporting. Method 400 then returns.
[0043] Figure 5 A high-level flowchart illustrating an exemplary method 500 for identifying keyframes including anatomical structures in a given cardiac phase of an ultrasound video is shown. Specifically, method 500 involves identifying frames in the ultrasound video corresponding to a selected cardiac phase. Method 500 may therefore include subroutines of method 400. (See reference...) Figures 1 to 3The system and components described herein constitute method 500; however, it should be understood that method 500 can be implemented with other systems and components without departing from the scope of this disclosure. Method 500 can be implemented as executable instructions in non-transitory memory (such as non-transitory memory 206) that, when executed by a processor (such as processor 204), cause the processor to perform the actions described herein. As an example, method 500 can be implemented as a keyframe extraction module 210.
[0044] Method 500 begins at 505. At 505, method 500 identifies a region of interest (ROI) in each frame of the ultrasound video acquired at 405. As discussed above, the ROI in each frame may include the region depicting the aortic valve in each ultrasound image frame. Therefore, method 500 can identify the aortic valve in each frame of the ultrasound video. Furthermore, at 510, method 500 extracts the identified ROI in each frame. Method 500 extracts the identified ROI in each frame by removing the identified ROI from each frame. Continuing at 515, method 500 registers the extracted ROI. For example, method 500 may temporally align the extracted ROI to prevent any motion beyond the leaflet motion (such as probe motion and / or patient motion). To identify, extract, and align or register the ROI throughout the ultrasound video, the ultrasound video can be processed by a deep neural network (such as Mask R-CNN of ROI model module 212) configured to perform aortic valve localization.
[0045] Continuing at 520, method 500 identifies frames within the registered and extracted region of interest that include the selected cardiac phase. Method 500 can use nonnegative matrix factorization (NMF) to identify the selected cardiac phase within the registered and extracted region of interest. For example, method 500 can model the video sequence as a linear combination of end-diastolic and end-systolic frames and apply NMF to identify keyframes within the selected cardiac phase. Method 500 can use the cardiac phase recognition module 214 of the keyframe extraction module 210 to identify frames within the registered and extracted region of interest that include the selected cardiac phase. Once method 500 identifies the keyframes, method 500 then returns.
[0046] As an illustrative example, Figure 6A set of exemplary ultrasound images 600 is shown illustrating exemplary identification of keyframes including anatomical structures at a given cardiac phase. Specifically, ultrasound image 600 includes a first ultrasound image 610 with a region of interest 615 including the aortic valve at the end-systolic phase, and a second ultrasound image 620 with a region of interest 625 including the aortic valve at the end-diastolic phase. Regions of interest 615 and 625 can be extracted and aligned as discussed above, and the resulting extracted frames corresponding to regions of interest 615 and 625 can be input into a generative model to classify the cardiac structures depicted in regions of interest 615 and 625 as mitral or tricuspid valves.
[0047] Figure 7 A high-level flowchart illustrating an exemplary method 700 for classifying anatomical structures as mitral or tricuspid valves according to one embodiment is shown. Specifically, method 700 involves classifying cardiac structures depicted in identified frames of ultrasound video corresponding to selected cardiac phases as mitral or tricuspid valves. Method 700 may therefore include subroutines of method 400. (See reference...) Figures 1 to 3 The system and components described herein constitute method 700, but it should be understood that method 700 may be implemented with other systems and components without departing from the scope of this disclosure. Method 700 may be implemented as executable instructions in non-transitory memory (such as non-transitory memory 206) that, when executed by a processor (such as processor 204), cause the processor to perform the actions described herein. In some examples, method 700 may be implemented as classification module 220.
[0048] Method 700 begins at 705. At 705, method 700 inputs image frames, including extracted regions of interest corresponding to selected cardiac phases, into appropriate generative models to generate images. For example, method 700 may input extracted frames corresponding to diastolic or end-diastolic phases into a first generative model 222, and extracted frames corresponding to systolic or end-systolic phases into a second generative model 224. The generative models may include variational autoencoders that encode each extracted frame into a latent space and then decode the low-dimensional representation of the extracted frames into an image space to generate images. The generative models may be trained based on images of the tricuspid valve, such that if the input frames depict the mitral valve, the resulting reconstruction error may be relatively high.
[0049] At 710, method 700 measures the reconstruction error between each generated image and its corresponding input image. The reconstruction error may include the difference between the generated image and the input image. Furthermore, method 700 may compute an error metric (as an illustrative and non-limiting example, such as mean squared error or mean absolute error) to quantify the reconstruction error as a single value rather than a two-dimensional array of differences.
[0050] Continuing at 715, method 700 classifies each input image with a reconstruction error greater than or equal to an error threshold as mitral lobe. For example, since the generative model is trained based on tricuspid lobe images, the reconstruction error of input images depicting the tricuspid lobe may be relatively low, while the reconstruction error of input images depicting the mitral lobe may be relatively high. Therefore, method 700 can classify each input image as mitral lobe when the reconstruction error, or the error metric based on the reconstruction error as described above, is greater than or equal to the error threshold. Furthermore, at 720, method 700 classifies each input image with a reconstruction error less than the error threshold as tricuspid lobe.
[0051] At 725, method 700 determines whether the number of mitral valve classifications exceeds a threshold. This threshold may include an aggregation error threshold as discussed above, which can be selected to avoid classifying the tricuspid valve as mitral valve overall due to a small number of correct or incorrect classifications of the tricuspid valve as mitral valve. If the number of mitral valve classifications is less than the aggregation error threshold (“No”), method 700 proceeds to 735. At 735, method 700 classifies the cardiac structure as tricuspid valve. Method 700 then returns.
[0052] However, if method 700 determines at 725 that the number of mitral valve classifications is greater than or equal to a threshold (“Yes”), then method 700 alternatively continues to 735. At 735, method 700 classifies the heart structure as the mitral valve. Then, method 700 returns.
[0053] Therefore, Method 700 uses a generative model trained on tricuspid valve images at a given cardiac phase to classify ultrasound image frames as depicting the tricuspid or mitral valve.
[0054] As an illustrative example of the individual and aggregate classification of cardiac structures in ultrasound video, Figure 8A set of images 800 is shown, illustrating generated images 820 and reconstruction errors 830 related to keyframes 810 in an ultrasound video, thereby producing tricuspid valve classification. Multiple keyframes 810 comprise extracted regions of interest in key cardiac phases within the ultrasound video and include a first image 811, a second image 812, a third image 813, a fourth image 814, a fifth image 815, and a sixth image 816. Each image frame in the multiple keyframes 810 is input to a generative model corresponding to the cardiac phase of the respective keyframe, thereby generating multiple generated images 820. The multiple generated images 820 include a first generated image 821 corresponding to the first image 811, a second generated image 822 corresponding to the second image 812, a third generated image 823 corresponding to the third image 813, a fourth generated image 824 corresponding to the fourth image 814, a fifth generated image 825 corresponding to the fifth image 815, and a sixth generated image 826 corresponding to the sixth image 816.
[0055] Quantitatively, the generated image 820 visually resembles keyframe 810, as depicted. Quantitatively, the reconstruction error 830 is measured by subtracting the generated image 820 from the corresponding keyframe 810. Therefore, the reconstruction error 830 includes a first error 831 for the first image 811, a second error 832 for the second image 812, a third error 833 for the third image 813, a fourth error 834 for the fourth image 814, a fifth error 835 for the fifth image 815, and a sixth error 836 for the sixth image 816. As discussed above, the anatomical structure depicted in keyframe 810 can be classified as a mitral valve if the reconstruction error is greater than or equal to an error threshold, and as a tricuspid valve if the reconstruction error is less than the error threshold. Error metrics such as MSE or MAE can be calculated from the reconstruction error 830 to more easily compare error thresholds. In the depicted example, the first error 831, second error 832, fourth error 834, and fifth error 835 are below the exemplary error threshold and are therefore classified as tricuspid valve. However, the third error 833 and the sixth error 836 are higher than the exemplary error threshold and are therefore classified as mitral valves. However, the anatomical structure is classified as a tricuspid valve because the number of mitral valve classifications is not higher than the exemplary aggregation error threshold.
[0056] As another example, Figure 9A set of images 900 is shown, illustrating generated images 920 and reconstruction errors 930 related to keyframes 910 in an ultrasound video, thus yielding mitral valve classification. Multiple keyframes 910 comprise extracted regions of interest in the ultrasound video at key cardiac phases and include a first image 911, a second image 912, a third image 913, a fourth image 914, a fifth image 915, and a sixth image 916. Each image frame in the multiple keyframes 910 is input to a generative model corresponding to the cardiac phase of the respective keyframe, thereby generating multiple generated images 920. The generative model used to generate the generated images 920 includes the same generative model used to generate the generated images 920. The multiple generated images 920 include a first generated image 921 for a first image 911, a second generated image 922 for a second image 912, a third generated image 923 for a third image 913, a fourth generated image 924 for a fourth image 914, a fifth generated image 925 for a fifth image 915, and a sixth generated image 926 for a sixth image 916.
[0057] Quantitatively, the generated image 920 appears significantly different from the keyframe 910 because the generative model is trained on images of the tricuspid valve while the keyframe depicts the mitral valve. Quantitatively, the reconstruction error 930 is measured by subtracting the generated image 920 from the corresponding keyframe 910. Therefore, the reconstruction error 930 includes a first error 931 for the first image 911, a second error 932 for the second image 912, a third error 933 for the third image 913, a fourth error 934 for the fourth image 914, a fifth error 935 for the fifth image 915, and a sixth error 936 for the sixth image 916. Using the same exemplary error threshold applied to the reconstruction error 930, each error of the reconstruction error 930 is above the error threshold, and therefore each image of the keyframe 910 is classified as the mitral valve. Therefore, the number of mitral valve classifications is greater than or equal to the aggregated error threshold, and therefore the heart structure is classified as the mitral valve.
[0058] The technical advantages of this disclosure include the automatic detection of the mitral valve in cardiac ultrasound images. Another technical advantage of this disclosure is the acquisition of ultrasound images. Yet another technical advantage of this disclosure is the classification and display of cardiac structures as either the mitral or tricuspid valve.
[0059] In one embodiment, a method includes acquiring an ultrasound video of the heart during at least one cardiac cycle, identifying frames in the ultrasound video corresponding to at least one cardiac phase, and classifying cardiac structures in the identified frames as mitral or tricuspid valves.
[0060] In a first example of the method, the method further includes: inputting the identified frames into at least one generative model to obtain a generated image corresponding to the identified frames; measuring the error between the identified frames and the corresponding generated image; classifying cardiac structures in each identified frame with an error higher than an error threshold as mitral valves; and classifying cardiac structures in each identified frame with an error lower than an error threshold as tricuspid valves. In a second example of the method, optionally including the first example, classifying cardiac structures in the identified frames as mitral or tricuspid valves includes: classifying the cardiac structure as a mitral valve if the number of identified frames is equal to or greater than the aggregated error threshold; and otherwise classifying the cardiac structure as a tricuspid valve. In a third example of the method, optionally including one or more of the first and second examples, inputting the identified frames into at least one generative model includes: inputting a first subgroup of identified frames corresponding to a first cardiac phase into a first generative model; and inputting a second subgroup of identified frames corresponding to a second cardiac phase into a second generative model. In a fourth example, optionally including one or more of the methods from the first to third examples, a first generative model is trained based on ultrasound images of the tricuspid valve at a first cardiac phase, and a second generative model is trained based on ultrasound images of the tricuspid valve at a second cardiac phase. In a fifth example, optionally including one or more of the methods from the first to fourth examples, the first and second generative models include variational autoencoders. In a sixth example, optionally including one or more of the methods from the first to fifth examples, identifying frames in an ultrasound video corresponding to at least one cardiac phase includes: extracting regions of interest (ROIs) from each frame of the ultrasound video to obtain a set of extracted ROIs; modeling a sequence of extracted ROIs as a linear combination of a first cardiac phase and a second cardiac phase; and applying nonnegative matrix factorization to the modeled sequence to identify the extracted ROIs corresponding to the first and second cardiac phases, wherein the identified frames corresponding to at least one cardiac phase include the identified extracted ROIs corresponding to both the first and second cardiac phases. In a seventh example, which optionally includes one or more of the methods from the first to the sixth examples, the method further includes temporally aligning the extracted region of interest before modeling the sequence of the extracted region of interest as a linear combination of a first cardiac phase and a second cardiac phase.
[0061] In another embodiment, a method includes: acquiring multiple images of a cardiac structure in a cardiac phase using an ultrasound probe; generating multiple output images corresponding to the multiple images using a generative model; measuring the error between each of the multiple output images and each corresponding image in the multiple images; classifying the cardiac structure as abnormal if the number of images with measured errors above an error threshold is higher than an aggregate error threshold, and otherwise classifying the cardiac structure as normal.
[0062] In a first example of the method, the cardiac structure includes the aortic valve, and classifying the cardiac structure as abnormal includes classifying the aortic valve as a mitral aortic valve, and classifying the cardiac structure as normal includes classifying the aortic valve as a tricuspid aortic valve. In a second example of the method optionally including the first example, acquiring multiple images of the cardiac structure at a cardiac phase includes acquiring a first plurality of images of the cardiac structure at a first cardiac phase and a second plurality of images at a second cardiac phase. In a third example of the method optionally including one or more of the first and second examples, generating multiple output images corresponding to the multiple images using a generative model includes generating a first plurality of output images corresponding to the first plurality of images using a first generative model trained based on cardiac images at a first cardiac phase; and generating a second plurality of output images corresponding to the second plurality of images using a second generative model trained based on cardiac images at a second cardiac phase. In a fourth example of the method optionally including one or more of the first to third examples, the first and second generative models include a variational autoencoder configured to encode an input image into a latent spatial representation and decode the latent spatial representation into an output image.
[0063] In another embodiment, the system includes an ultrasound probe and a processor, the processor having executable instructions configured in a non-transitory memory that, when executed, cause the processor to: acquire a sequence of ultrasound images of cardiac structures over at least one cardiac cycle via the ultrasound probe; identify frames in the ultrasound image sequence corresponding to at least one cardiac phase; classify the identified frames as normal or abnormal; classify the cardiac structure as a mitral valve if the number of identified frames classified as abnormal exceeds a threshold; and otherwise classify the cardiac structure as a tricuspid valve.
[0064] In a first example of the system, the processor also has executable instructions configured in non-transitory memory that, when executed, cause the processor to input identified frames into at least one generative model to obtain an output image corresponding to the identified frames, and to measure the error between the identified frames and the corresponding output images. In a second example of the system optionally including the first example, to classify identified frames as normal or abnormal, the processor also has executable instructions configured in non-transitory memory that, when executed, cause the processor to classify each identified frame with an error exceeding an error threshold as abnormal, and to classify the remaining identified frames as normal. In a third example of the system optionally including one or more of the first and second examples, at least one generative model is trained based on ultrasound images of the tricuspid valve at at least one cardiac phase. In a fourth example of the system optionally including one or more of the first to third examples, the at least one generative model includes a variational autoencoder configured to encode input frames into low-dimensional representations and decode the low-dimensional representations into output images. In a fifth example of a system optionally including one or more of the first to fourth examples, in order to identify frames in an ultrasound image sequence corresponding to at least one cardiac phase, the processor further comprises executable instructions configured in non-transitory memory that, when executed, cause the processor to extract regions of interest (ROIs) in each frame of the ultrasound image sequence to obtain a set of extracted ROIs. In a sixth example of a system optionally including one or more of the first to fifth examples, the processor further comprises executable instructions configured in non-transitory memory that, when executed, cause the processor to model a set of extracted ROIs as a linear combination of a first cardiac phase and a second cardiac phase; and to apply non-negative matrix factorization to the modeled set of extracted ROIs to identify a first set of extracted ROIs corresponding to the first cardiac phase and a second set of extracted ROIs corresponding to the second cardiac phase, wherein the identified frames corresponding to at least one cardiac phase include the identified extracted ROIs corresponding to the first cardiac phase and the identified extracted ROIs corresponding to the second cardiac phase.
[0065] As used herein, elements or steps listed in the singular and beginning with the word "a" or "an" should be understood to not exclude multiple said elements or steps unless such exclusion is explicitly stated. Furthermore, references to "an embodiment" of the invention are not intended to be construed as excluding the existence of additional embodiments that also include the referenced features. Moreover, unless explicitly stated to the contrary, embodiments that "comprise," "include," or "have" elements or multiple elements having a particular characteristic may include additional such elements that do not have that characteristic. The terms "comprise" and "in..." are used as concise linguistic equivalents to the corresponding terms "comprising" and "wherein". Furthermore, the terms "first," "second," and "third," etc., are used merely as notations and are not intended to impose numerical requirements or a particular order of position on their objects.
[0066] This written description uses examples to disclose the invention, including the best mode, and also enables those skilled in the art to practice the invention, including making and using any device or system and performing any included methods. The scope of patentability of the invention is defined by the claims and may include other examples that would occur to those skilled in the art. Such other examples are intended to fall within the scope of the claims if they have structural elements that are not indistinguishable from the literal language of the claims, or if they include equivalent structural elements that differ only slightly from the literal language of the claims.
Claims
1. A method for classifying anatomical structures in ultrasound video, the method comprising: Acquire ultrasound video of the heart during at least one cardiac cycle; Identify frames in the ultrasound video that correspond to at least one cardiac phase; as well as The identified cardiac structures in the frames are classified as either mitral or tricuspid valves, including: The identified frames are input into at least one generative model to obtain a generated image corresponding to the identified frames; Measure the error between the identified frame and the corresponding generated image; The cardiac structure in each identified frame with an error exceeding the error threshold is classified as the mitral valve; and The cardiac structure in each identified frame with an error below the stated error threshold is classified as the tricuspid valve.
2. The method of claim 1, wherein classifying the cardiac structure in the identified frame as the mitral valve or the tricuspid valve comprises: If the heart structure in the identified frames is classified as the mitral valve when the number of identified frames is equal to or greater than the aggregation error threshold, the heart structure is classified as the mitral valve. as well as Otherwise, the cardiac structure in the identified frame is classified as the tricuspid valve.
3. The method of claim 1, wherein inputting the identified frame into the at least one generative model comprises: The first subgroup of the identified frames corresponding to the first cardiac phase is input into the first generative model; as well as The second subgroup of the identified frames corresponding to the second cardiac phase is input into the second generative model.
4. The method of claim 3, wherein the first generative model is trained based on an ultrasound image of the tricuspid valve at the first cardiac phase, and wherein the second generative model is trained based on an ultrasound image of the tricuspid valve at the second cardiac phase.
5. The method according to claim 3, wherein the first generative model and the second generative model comprise a variational autoencoder.
6. The method of claim 1, wherein identifying the frame in the ultrasound video corresponding to the at least one cardiac phase comprises: Extract the region of interest from each frame of the ultrasound video to obtain a set of extracted regions of interest; The extracted region of interest sequence is modeled as a linear combination of the first and second cardiac phases; as well as Nonnegative matrix factorization is applied to the modeled sequence to identify extracted regions of interest corresponding to the first cardiac phase and the second cardiac phase, wherein the identified frames corresponding to the at least one cardiac phase include the identified extracted regions of interest corresponding to the first cardiac phase and the second cardiac phase.
7. The method according to claim 6, further comprising: The extracted region of interest is temporally aligned before the sequence is modeled as a linear combination of the first cardiac phase and the second cardiac phase.
8. A method for classifying anatomical structures in an ultrasound image frame, the method comprising: Multiple images of cardiac structures in the cardiac phase are acquired using an ultrasound probe; A generative model is used to generate multiple output images corresponding to the multiple images; Measure the error between each of the plurality of output images and each corresponding image in the plurality of images; If the number of images with measured errors exceeding the error threshold is higher than the aggregate error threshold, the cardiac structure is classified as abnormal. as well as Otherwise, the heart structure is classified as normal.
9. The method of claim 8, wherein the cardiac structure includes the aortic valve, wherein classifying the cardiac structure as abnormal includes: Classifying the aortic valve as a mitral aortic valve, and classifying the heart structure as normal, includes classifying the aortic valve as a tricuspid aortic valve.
10. The method of claim 8, wherein acquiring the plurality of images of the cardiac structures at the cardiac phase comprises: Acquire a first plurality of images of the heart structure in a first cardiac phase and a second plurality of images in a second cardiac phase.
11. The method of claim 10, wherein generating the plurality of output images corresponding to the plurality of images using the generative model comprises: A first generative model trained based on a heart image at the first heart phase is used to generate a first plurality of output images corresponding to the first plurality of images; as well as A second generative model trained based on a heart image in the second heart phase is used to generate a second plurality of output images corresponding to the second plurality of images.
12. The method of claim 11, wherein the first generative model and the second generative model comprise a variational autoencoder configured to encode an input image into a latent spatial representation and decode the latent spatial representation into an output image.
13. A system for classifying anatomical structures in ultrasound image frames, the system comprising: Ultrasonic probe; and A processor having executable instructions configured in non-transitory memory, the instructions causing the processor, when executed, to: Acquired via the ultrasound probe a sequence of ultrasound images of cardiac structures during at least one cardiac cycle; Identify frames in the ultrasound image sequence that correspond to at least one cardiac phase; The identified frames are classified as normal or abnormal. If the number of identified frames classified as abnormal exceeds a threshold, the heart structure is classified as the mitral valve. as well as Otherwise, the heart structure is classified as the tricuspid valve.
14. The system of claim 13, wherein the processor further comprises executable instructions configured in the non-transitory memory, the executable instructions causing the processor, when executed, to: The identified frames are input into at least one generative model to obtain an output image corresponding to the identified frames; and Measure the error between the identified frame and the corresponding output image.
15. The system of claim 14, wherein, in order to classify the identified frames as normal or abnormal, the processor further comprises executable instructions configured in the non-transitory memory, the executable instructions causing the processor, when executed, to: Each identified frame with an error exceeding the error threshold is classified as an anomaly; and The remaining identified frames are classified as normal.
16. The system of claim 14, wherein the at least one generative model is trained based on ultrasound images of the tricuspid valve at the at least one cardiac phase.
17. The system of claim 14, wherein the at least one generative model comprises a variational autoencoder configured to encode an input frame into a low-dimensional representation and decode the low-dimensional representation into an output image.
18. The system of claim 13, wherein, in order to identify the frame in the ultrasound image sequence corresponding to the at least one cardiac phase, the processor further comprises executable instructions configured in the non-transitory memory, the executable instructions causing the processor, when executed, to: Regions of interest are extracted from each frame of the ultrasound image sequence to obtain a set of extracted regions of interest.
19. The system of claim 18, wherein the processor further comprises executable instructions configured in the non-transitory memory, the executable instructions causing the processor, when executed, to: The extracted regions of interest are modeled as a linear combination of the first and second cardiac phases; and Nonnegative matrix factorization is applied to the modeled set of extracted regions of interest to identify a first set of extracted regions of interest corresponding to the first cardiac phase and a second set of extracted regions of interest corresponding to the second cardiac phase, wherein the identified frames corresponding to the at least one cardiac phase include the identified extracted regions of interest corresponding to the first cardiac phase and the identified extracted regions of interest corresponding to the second cardiac phase.