Evaluating cardiac parameters

An automated method using view and quality classification models improves cardiac parameter evaluation by accurately selecting optimal ultrasound image clips, addressing inconsistencies and inefficiencies in manual echocardiography.

WO2025242532A1PCT designated stage Publication Date: 2025-11-27KONINKLIJKE PHILIPS NV
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
PCT/EP2025/063374
Authority / Receiving Office
WO · WO
Patent Type
Applications
Current Assignee / Owner
Priority Date
2024-05-22
Filing Date
2025-05-15
Publication Date
2025-11-27

AI Technical Summary

Technical Problem

Existing echocardiography methods for determining cardiac parameters suffer from inconsistencies due to manual image classification, time-consuming processes, and susceptibility to human error, leading to variability and potential inaccuracies in cardiac parameter evaluation.

Method used

An automated method using view and quality classification models to identify and rank ultrasound image clips, generating composite scores for accurate selection of optimal clips to determine cardiac parameters, leveraging machine learning techniques to improve efficiency and reliability.

Benefits of technology

The method enhances the accuracy and reliability of cardiac parameter measurements by automating the selection of high-quality ultrasound image clips, reducing variability and human error, and enabling efficient cardiac assessment.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure EP2025063374_27112025_PF_FP_ABST
    Figure EP2025063374_27112025_PF_FP_ABST
Patent Text Reader

Abstract

A method is provided for selecting optimal ultrasound image clips to determine cardiac parameters. A view identification model is used to classify a view represented by each of a plurality of image clips and assign view confidence scores. A quality classification model is used to assess clip quality and generate quality scores indicative of predicted usefulness of each clip for computing a cardiac parameter of interest. Composite clip scores are computed by combining view and quality scores. Clip sets corresponding to predetermined combinations of image views are compiled and ranked based on the composite scores of their component clips. The highest-ranking clip sets are selected first for cardiac parameter determination.
Need to check novelty before this filing date? Find Prior Art

Description

[0001] EVALUATING CARDIAC PARAMETERS

[0002] FIELD OF THE INVENTION

[0003] This invention relates to the field of cardiovascular systems, and in particular to determining the value of cardiac parameters of cardiovascular systems.

[0004] BACKGROUND OF THE INVENTION

[0005] Echocardiography is an imaging technique clinical assessment of the heart. It can be used to determine the values of a variety of clinically relevant parameters of the cardiovascular system. Global left ventricular (LV) function evaluation is often performed by clinicians during an echocardiogram. Common parameters evaluated during echocardiography include left-ventricular ejection fraction (LVEF), global longitudinal strain (GLS), longitudinal strain per view (meaning the average longitudinal strain calculated from one standard apical view: A4C, A2C or A3C), segmental strain (meaning longitudinal strain measured in individual myocardial segments of the left ventricle), and ejection fraction per view (meaning an estimation of the ejection fraction (EF) from one individual apical view: A4C, A2C or A3C).

[0006] Traditionally, these parameters are evaluated by manual assessment of ultrasound images by a health care professional or by using semi-automated or automated software developed by different manufacturers. This may often lead to inconsistencies between practitioners and between the systems of different manufacturers.

[0007] Further, it is usual for many different ultrasound images to be taken during one examination of a patient, corresponding to different views of the heart. These may vary in quality. This may lead to further inconsistencies in the evaluation of cardiac parameters as the use of different images may result in different values for a given cardiac parameter.

[0008] For example, during an echocardiogram, the sonographer captures ultrasound images from several standard views, such as Apical 2-Chamber (A2C), Apical 3 -Chamber (A3C), Apical 4-Chamber (A4C), Parasternal Long Axis (PLAX), and RV-focused 4-chamber view (RV), among others. For each view, multiple images (or image clips) may be acquired to achieve optimal image quality, observe dynamic changes over time, reduce artifacts and noise, and ensure reproducibility and consistency. Following image acquisition, clinicians manually classify the view for each image and select the best quality images for patient evaluation. This process is time-consuming, subject to intra and inter-observer variability, and prone to errors. To address these challenges, there is a growing need for automated systems to classify the view and quality of each image quickly and accurately.

[0009] SUMMARY OF THE INVENTION

[0010] The invention is defined by the claims.

[0011] An aspect of the invention is a computer-implemented method for determining a value of a cardiac parameter of a cardiovascular system based on a study dataset of ultrasound image clips of the cardiovascular system. Each clip comprises a set of one or more ultrasound image frames corresponding to a same view of the cardiovascular system. The method comprises: processing the study dataset with a view identification model to predict a view represented by each ultrasound image clip and an associated view score indicative of a likelihood that the predicted view of the ultrasound image clip is correct; determining for each of the clips, using a quality classification model, a clip quality score indicative of a predicted usefulness of the clip for determining a value of the cardiac parameter; determining for each of the clips a composite clip score, wherein the composite clip score is computed as a function of the view score and clip quality score; compiling from the clips of the study dataset, a plurality of clip sets, each clip set comprising a set of clips corresponding to a pre-determined combination of views required for computing the cardiac parameter; determining, for each of the plurality of clip sets, a clip set ranking, wherein the clip set ranking is computed for each clip set at least in part as a function of the individual composite clip scores of the set of clips comprised by the clip set; selecting one or more of the clip sets based on the clip set rankings; and determining, using the clips comprised by the selected one or more of the clip sets, a value of the cardiac parameter of the cardiovascular system.

[0012] Embodiments of the invention are based on the concept of generating for each clip a separate view score (reflecting confidence in the view classification) and quality score (reflecting predicted usefulness of the clip for computing the relevant cardiac parameter) and using these in combination, via a composite clip score which is a function of both, to rank and select the best clips to use for computing the cardiac parameter. This improves accuracy and reliability of cardiac parameter measurements.

[0013] Optionally, the view classification and the view score may be generated by a view classification model and the quality classification may be performed by one or more further models which are different to the view classification model. Optionally, a different respective quality classification model may be employed depending upon the clip view class.

[0014] It is to be understood that the term 'clip' as used herein may comprise just a single acquired image or may comprise a sequence or two or more acquired image frames, e.g. representing a cine-loop. The term ‘clip’ may be understood synonymously with the term ‘acquisition’. During an ultrasound examination, a sonographer positions the ultrasound probe at a particular position, for acquiring image(s) of a certain view, and triggers acquisition. The acquisition may be comprised of multiple image frames or may comprise just a single image frame for the relevant view. Each clip accordingly corresponds to a specific view obtained during a particular study. The term 'clip' thus encompasses both multi-frame and single-frame acquisitions. Single-frame clips may be treated analogously to multi-frame clips within the context of the invention.

[0015] In some embodiments, the method further comprises, after determining the view score for each clip, identifying and discarding any clips having a view score which fails to meet a pre-determined criterion, for example which has a value outside of a pre-determined range, e.g. below a pre-determined threshold.

[0016] This may help eliminate clips which do not correctly capture the relevant view early in the process.

[0017] Regarding the composite clip score, in some embodiments this is computed as a function of the product of the view score and clip quality score. It may be computed as equal to the product of the view score and clip quality score, or a function thereof.

[0018] In some embodiments, the method further comprises, for at least a subset of the clips: identifying depth information for each of the at least subset of clips, relating to an imaging depth of ultrasound image frames contained in each clip. The method may further comprise, after applying the view identification model to the study dataset, and prior to determining the composite clip scores, performing depth-adjustment of the view scores for each of the at least subset of clips by applying a depth-adjustment to each view score which is a function of an imaging depth associated with the image frames contained in the respective clip.

[0019] This depth-adjustment step accounts for the impact of imaging depth on view classification confidence, potentially improving the accuracy of the composite clip scores and subsequent clip selection.

[0020] In some embodiments, the method further comprises: after determining the composite clip scores, ranking the clips according to the composite clip scores; and selecting clips for inclusion in the clip sets based on the rankings of the clips. This allows for the highest-quality clips to be prioritized for inclusion in the clip sets, potentially improving the accuracy of the cardiac parameter determination.

[0021] In some embodiments, the method further comprises, after determining the composite clip scores and before compiling the clip sets: ranking the clips according to composite clip scores; identifying a top ranking subset of clips; for the top-ranking subset of clips, identifying depth information for each of the clips relating to an imaging depth of ultrasound image frames contained in each clip; and modifying the composite clip scores of the top ranking subset of the clips based on the depth information to prioritize clips comprising lower imaging depth image frames (i.e. shallower imaging depth). Lower imaging depth corresponds to lesser image penetration depth. Clips with shallow image depth images may provide higher image quality for computing cardiac parameters. This additional depth-based modification of composite clip scores for top-ranking clips adds granularity to the rankings, potentially allowing for more informed selection among high-quality clips and favoring clips with shallower imaging depths that may provide clearer images.

[0022] In some embodiments, the steps of selecting one or more of the clip sets based on the rankings, and determining the cardiac parameter using the selected clip sets comprises: determining an ordered list of the clip sets, ordered from higher ranking clip sets to lower ranking clip sets; and indexing through the clips sets of the ordered list one by one, applying a cardiac parameter determination model to each clip set, evaluating an output of a cardiac parameter determination model against a pre-defined success criterion and, continuing to the next clip set in the list if the success criterion is not met until the success criterion is met.

[0023] In some embodiments, determining the clip quality score comprises processing the image frames of the respective clip with a view-specific quality classifier selected from among a bank of different view-specific quality classifiers based on the view classification of the clip.

[0024] In some embodiments, the method further comprises a preliminary step of processing the study dataset with a contrast classification model configured to classify each clip as to whether it contains any contrast images, and discarding any clips classified as containing contrast images.

[0025] This preliminary filtering step ensures that only non-contrast images are processed by the subsequent steps of the method.

[0026] In some embodiments, the pre-determined combination of views to which the clips of at least a subset of the clip sets correspond comprises a triplet of image views consisting of: a two-chamber apical (A2C) view; a three-chamber apical (A3C) view; and a four-chamber apical (A4C) view. Optionally, the cardiac parameter is the left ventricular ejection fraction or the average global longitudinal strain.

[0027] In some embodiments, the pre-determined combination of views to which at least a subset of the clips of each clip set corresponds comprises a pair of image views consisting of: a two-chamber apical (A2C) view; and a four-chamber apical (A4C) view.

[0028] Depending upon the target cardiac parameter, the clip sets may include both triplets and pairs. By way of brief explanation: some cardiac parameters require all three of A2C, A3C and A4C for determination of the relevant parameter. These parameters include, by way of non-limiting example, global longitudinal strain (GLS). For such parameters, it is necessary that the clip sets include only the triplets mentioned above. Some cardiac parameters require only a pair of images A2C and A4C (known as a Simpson biplane). These parameters include, by way of non-limiting example, left ventricular ejection fraction (LVEF). For these parameters, either the above-mentioned triplet or the above-mentioned pair would be acceptable as both contain A2C and A4C views. For some parameters, only one of A2C and A4C is needed. For example, for computing longitudinal strain per view or ejection fraction per view, only one of A4C, A2C or A3C is required in each case. For these cardiac parameters, either a triplet, a pair or a single image view of A2C or A4C would be acceptable. Therefore, depending on the cardiac parameter, the combination of views, and the number of views, included in the clip sets can be varied.

[0029] In some embodiments, the cardiac parameter value comprises at least one of: a left ventricular ejection fraction value; global longitudinal strain value; a segmental strain value; a longitudinal strain per view value; and an ejection fraction per view value.

[0030] Longitudinal strain per view corresponds to the average longitudinal strain calculated from a single standard apical view, such as the apical 4-chamber (A4C), apical 2-chamber (A2C), or apical 3 -chamber (A3C) view. The ejection fraction per view corresponds to an estimation of the ejection fraction (EF) calculated from a single apical view, such as the A4C, A2C, or A3C view. These "per view" measurements provide localized assessments of cardiac function based on individual imaging perspectives, in contrast to global measurements that incorporate data from multiple views.

[0031] A further aspect of the invention is a computer program product comprising computer program code configured, when run on a processor, to implement a method in accordance with any embodiment of the invention or in accordance with any claim.

[0032] A further aspect of the invention is a processing device comprising one or more processors configured to perform a method in accordance with any embodiment described in this disclosure or in accordance with any claim. These and other aspects of the invention will be apparent from and elucidated with reference to the embodiment s) described hereinafter.

[0033] BRIEF DESCRIPTION OF THE DRAWINGS

[0034] For a better understanding of the invention, and to show more clearly how it may be carried into effect, reference will now be made, by way of example only, to the accompanying drawings, in which:

[0035] Fig. 1 outlines steps of an example method in accordance with one or more embodiments of the invention;

[0036] Fig. 2 illustrates a block diagram of an example processing device and system in accordance with one or more embodiments of the invention;

[0037] Fig. 3 outlines steps of an example method in accordance with a particular set of embodiments of the invention;

[0038] Fig. 4 illustrates a flowchart of a first subset of the steps performed in accordance with the method of Fig. 3;

[0039] Fig. 5 illustrates a flowchart of a second subset of steps performed in accordance with the method of Fig. 3; and

[0040] Fig. 6 outlines steps of an example process for ranking and iteratively selecting image clips to use for determining the cardiac parameter in accordance with one or more embodiments of the invention.

[0041] DETAILED DESCRIPTION OF THE EMBODIMENTS

[0042] The invention will be described with reference to the Figures.

[0043] It should be understood that the detailed description and specific examples, while indicating exemplary embodiments of the apparatus, systems and methods, are intended for purposes of illustration only and are not intended to limit the scope of the invention. These and other features, aspects, and advantages of the apparatus, systems and methods of the present invention will become better understood from the following description, appended claims, and accompanying drawings. It should be understood that the Figures are merely schematic and are not drawn to scale. It should also be understood that the same reference numerals are used throughout the Figures to indicate the same or similar parts.

[0044] Embodiments of the invention provide a method for selecting optimal ultrasound image clips to determine cardiac parameters. A view identification model is used to classify a view represented by each of a plurality of image clips and assign view confidence scores. A quality classification model is used to assess clip quality and generate clip quality scores indicative of predicted usefulness of each clip for computing a cardiac parameter of interest. Composite clip scores are computed by combining view and quality scores. A plurality of candidate clip sets are compiled, each comprising a set of clips corresponding to a predetermined combination of image views required for computing a target cardiac parameter. The candidate clip sets are then ranked based on the composite scores of the clips comprised by each. The highest-ranking clip sets are selected for cardiac parameter determination. This approach enables efficient and accurate cardiac assessment by leveraging machine learning techniques to identify the most suitable ultrasound image clips to use for computing cardiac parameters.

[0045] In the current state of the art, it is usual for many different ultrasound images to be taken during one examination of a patient, including several images for each of a plurality of different views. Following image acquisition, clinicians manually classify the view for each image and select the best quality image for each view for patient evaluation.

[0046] The current manual classification methodology for echocardiogram images is associated with several limitations. One drawback is the substantial time investment required for this process. Clinicians must dedicate considerable effort to evaluate and select high-quality images from various cardiac views. This time-intensive task can be particularly challenging in high-volume clinical environments where efficiency is important.

[0047] Another notable concern is the inherent inconsistency in manual classifications. Both intra-ob server and inter-observer variability are prevalent in this approach. These inconsistencies can lead to discrepancies in image selection processes. Such variations may potentially impact the accuracy and reliability of subsequent diagnoses derived from these selected images.

[0048] Furthermore, manual classification methods are inherently susceptible to human error. This vulnerability can manifest in various forms, including misclassification of images or failure to identify subtle diagnostic indicators. These errors have the potential to compromise the quality of patient care.

[0049] Embodiments of the present invention provide an automated system and method for classifying a selecting the best combination of images to use for computing one or more target cardiac parameters.

[0050] Fig. 1 outlines in block diagram form steps of an example computer-implemented method according to one or more embodiments. The steps will be recited in summary, before being explained further in the form of example embodiments. The method 10 is for determining a value of a cardiac parameter of a cardiovascular system based on a study dataset 12 of ultrasound image clips of the cardiovascular system. Each clip comprises a set of one or more ultrasound image frames. Optionally, the ultrasound image frames of any one clip may all represent a same view of the cardiovascular system.

[0051] The method 10 comprises processing the study dataset 12 with a view identification model to predict 14 a view represented by each ultrasound image clip and an associated view score indicative of a likelihood that the predicted view of the ultrasound image clip is correct. In some embodiments, the view identification model may be configured to classify the view into one of a pre-defined set of view classifications. By way of example, the set of view classifications may include: apical 2-chamber view (A2C), apical 3-chamber view (A3C), and apical 4-chamber view (A4C). Preferably, the set of view classifications also includes parasternal long-axis view (PLAX), and RV-focused 4-chamber view (RV). The set of classifications may also include an ‘other’ class (OTHR) for views that do not fall into one of the other classifications.

[0052] The method 10 further comprises determining 16 for each of the clips, using a quality classification model, a clip quality score indicative of a predicted usefulness of the clip for determining a value of the cardiac parameter.

[0053] The method 10 further comprises determining 18 for each of the clips a composite clip score, wherein the composite clip score is computed as a function of the view score and clip quality score. In some embodiments, the composite clip score may be computed based on the product of the view score and clip quality score. For example, the composite clip score may be computed as equal to the product of the view score and clip quality score or a function thereof.

[0054] The method 10 further comprises compiling 20 from the clips of the study dataset, a plurality of clip sets, each clip set comprising a set of clips corresponding to a pre-determined combination of views required for computing the cardiac parameter. The clip sets might otherwise be referred to as a clip combinations. The plurality of clip sets compiled at this stage in the method might be understood as candidate clip sets, in the sense that they include multiple instances of the same combination of image views, and wherein only one or a subset of the compiled clip sets will ultimately be used to compute the cardiac parameter. In other words, the compilation process 20 generates multiple candidate clip sets that represent the same combination of image views. For example, if a particular cardiac parameter requires views A, B, and C, the method may compile multiple clip sets, each containing different clips representing views A, B, and C from the study dataset. From these compiled clip sets, only one or a subset of them will ultimately be selected for use in computing the cardiac parameter. The compilation of multiple candidate sets allows for subsequent evaluation and selection of the optimal set or sets based on quality and other criteria.

[0055] The method 10 further comprises determining 22, for each of the plurality of clip sets, a clip set ranking, wherein the clip set ranking is computed for each clip set at least in part as a function of the individual composite clip scores of the set of clips comprised by the clip set. For example, in a simple case, the clip sets may be ranked in order of clip score from highest to lowest. However, in other examples, the ranking may be determined based on a more complex calculation, for example based on modifying or modulating the clip scores based on imaging depth of the clips.

[0056] The method 10 further comprises selecting 24 one or more of the clip sets based on the clip set rankings.

[0057] The method 10 further comprises determining 26, using the clips comprised by the selected one or more of the clip sets, a value of the cardiac parameter of the cardiovascular system. By way of example, in some embodiments, the cardiac parameter may include any one or more of: a left-ventricular ejection fraction (LVEF), global longitudinal strain (GLS), longitudinal strain per view (meaning the average longitudinal strain calculated from one standard apical view: A4C, A2C or A3C), segmental strain (meaning longitudinal strain measured in individual myocardial segments of the left ventricle), and ejection fraction per view (meaning an estimation of the ejection fraction EF from one individual apical view: A4C, A2C or A3C).

[0058] It is to be understood that the term 'clip' as used herein may comprise just a single acquired image or may comprise a sequence or two or more acquired image frames, e.g. representing a cine-loop. The term ‘clip’ may be understood synonymously with the term ‘acquisition’. During an ultrasound examination, a sonographer positions the ultrasound probe at a particular position, for acquiring image(s) of a certain view, and triggers acquisition. The acquisition may be comprised of multiple image frames or may comprise just a single image frame for the relevant view. Each clip accordingly corresponds to a specific view obtained during a particular study. The term 'clip' thus encompasses both multi-frame and single-frame acquisitions. Single-frame clips may be treated analogously to multi-frame clips within the context of the invention.

[0059] The method may be understood as being composed of two main functionalities: (i) view and quality classification of image clips, and (ii) selection and ranking of candidate combinations (sets) of image clips for use in computing the one or more cardiac image parameters.

[0060] Functionality (i) corresponds to steps 12, 14, 16 and 18 of Fig. 1 (and steps 12, 52, 14, 16, 54, 56, and 18 of Fig. 3, to be described below).

[0061] Functionality (ii) corresponds to steps 20, 22, 24 and 26 of Fig. 1 (and steps 58, 60, 20, 22, 24 and 26 of Fig. 3 to be described below).

[0062] Regarding the view and quality classification functionality (i), this functionality involves recognizing which of a pre-defined set of image views is depicted in image frames using a view classifier model, e.g. the heart apical view (e.g. a detected one of A2C, A3C, A4C, RV, PL AX, and OTHER), and assessing the quality of the image frames. The quality assessment may be achieved with an image quality classification model which may output a numerical quality score, e.g. between 0 and 100, or between 0 and 1. Optionally a quality class (e.g. BAD, MID, GOOD) may further be determined from the quality score, with each class corresponding to a defined sub-range of the quality score range.

[0063] The view and quality classification may be performed using respective machine learning models utilizing one or a plurality of convolutional neural networks (CNNs). The CNNs may be deep learning neural networks.

[0064] Regarding the functionality (ii) of selection and ranking of candidate combinations (sets) of image clip, this functionality comprises generating and ranking candidates for computation of one or more cardiac parameters (termed ‘clinical evaluation’ herein), for example ejection fraction (EF) and strain analysis, prioritizing images from highest quality to lowest quality for clinical evaluation. Image clips suitable for inclusion in candidate image clip sets are identified based on their view classifications. The candidate clips sets may then be ranked based on combined clip scores of the clips comprised by each candidate clip set.

[0065] In some embodiments, the view identification model is configured to output a view class and a view score. The view identification model may be applied individually to each of the one or more image frames of each clip, and the overall view class for a clip determined based on amalgamation of the individual view classifications of the one or more frames of the clip, for example by averaging or highest frequency selection. The view score may be a prediction confidence score output by the model concurrently with output of the view class.

[0066] In some embodiments, determining the clip quality scores comprises processing each of the clips with a quality classification model. In some embodiments, the quality classification model may be configured to output a quality score ranging from 0 to 1. The clip quality score serves as an indicator of the clip's suitability and reliability for use in subsequent measurements or analyses.

[0067] In some embodiments, the quality classification model is configured to receive individual image frames as input and output a quality score for the frame. To determine the clip quality score, quality scores for individual frames may be amalgamated, for example by averaging. In other words, the quality classification model operates at the individual frame level, with post-processing applied to combine the frame-level predictions into an overall clip-level prediction. This aggregation may be accomplished, for example, by averaging the scores of all predicted frames within the clip.

[0068] Additionally, as part of the post-processing, optionally the aggregated scores may be mapped to quality classes based on predefined thresholds. For instance, scores in the range of 0 to 0.6 may be classified as "bad" quality, scores from 0.6 to 0.8 as "mid" quality, and scores from 0.8 to 1.0 as "good" quality. The specific thresholds used for this classification may be adjustable based on the model's performance characteristics.

[0069] In some embodiments, the composite clip score may be computed 18 as a function of the product of the clip view score and clip quality score. It may be computed as equal to the product of the two scores, or a correlate thereof for example. Both the clip view score and clip quality score may take values from 0 to 1, with 1 representing highest confidence (in the view) or highest quality. In this case, a product of the two scores results in a composite clip score also ranging from 0 to 1, where 1 represents the highest possible combined confidence in view classification and image quality. A score closer to 1 indicates both high confidence in the view classification and high image quality, while scores closer to 0 suggest lower confidence in view classification, lower image quality, or both.

[0070] In some embodiments, the method may be performed in real time with image acquisition, for example during an imaging study itself. Alternatively the method may be applied to study datasets offline but still with near real-time processing. The automated system may permit classification of images almost instantaneously, allowing for rapid assessment and minimal delays in the diagnostic process.

[0071] As noted above, the method can also be embodied in hardware form, for example in the form of a processing device which is configured to carry out a method in accordance with any example or embodiment described in this document, or in accordance with any claim.

[0072] To further aid understanding, Fig. 2 presents a schematic representation of an example processing device 32 configured to execute a method in accordance with one or more embodiments of the invention. The processing device 32 is shown in the context of a system 30 which comprises the processing device. The processing device 32 alone represents an aspect of the invention. The system 30 is another aspect of the invention. The provided system need not comprise all of the illustrated hardware components, it may comprise only subset of them.

[0073] The processing device 32 comprises one or more processors 36 configured to perform a method 10 in accordance with that outlined above, or in accordance with any embodiment described in this document or any claim of this application. In the illustrated example, the processing device further comprises an input / output 34 or communication interface.

[0074] In the illustrated example of Fig. 2, the system 30 further comprises an ultrasound acquisition or imaging apparatus 40 for acquiring a study dataset 12 of ultrasound image clips.

[0075] The system 30 may further comprise a memory 38 for storing computer program code (i.e. computer-executable code) which is configured for causing the one or more processors 36 of the processing device 32 to perform the method as outlined above, or in accordance with any embodiment described in this disclosure, or in accordance with any claim.

[0076] As mentioned previously, the invention can also be embodied in software form. Thus another aspect of the invention is a computer program product comprising computer program code configured, when run on a processor, to cause the processor to perform a method in accordance with any example or embodiment of the invention described in this document, or in accordance with any claim.

[0077] Regarding the view identification model, this may be implemented using a machine learning model, for example comprising one or more convolutional neural networks (CNNs).

[0078] Methods of training a machine-learning algorithm are well known. Typically, such methods comprise obtaining a training dataset, comprising training input data entries and corresponding training output data entries. An initialized machine-learning algorithm is applied to each input data entry to generate predicted output data entries. An error between the predicted output data entries and corresponding training output data entries is used to modify the machinelearning algorithm. This process can be repeated until the error converges, and the predicted output data entries are sufficiently similar (e.g. ±1%) to the training output data entries. This is commonly known as a supervised learning technique. For example, weightings of the mathematical operation of each neuron may be modified until the error converges. Known methods of modifying a neural network include gradient descent, backpropagation algorithms and so on.

[0079] The view identification model may be trained in a supervised training procedure in which the training input data entries correspond to real or simulated ultrasound images of cardiovascular systems. The training output data entries may correspond to the view classes of these ultrasound images as manually annotated by a health care professional. In other words, the view identification model may be trained using a deep learning algorithm configured to receive an array of training inputs and respective known outputs, wherein a training input comprises an ultrasound image of a cardiovascular system and a respective known output comprises a known view class of the ultrasound image. In this way, the view identification model may be trained to automatically determine the view of an ultrasound image from an echocardiogram examination. The machine learning model is configured to also output a confidence level in the determination of the output and this is used as the view score, indicative of the likelihood that the predicted view of the ultrasound image clip is correct.

[0080] Regarding the quality classification model, this may comprise one or more trained Convolutional Neural Networks (CNNs). In some embodiments, the model may comprise a plurality of view-specific CNNs, each trained to determine a quality score indicative of predicted usefulness of an input image frame depicting a particular view. A bank of quality classification models may therefore be trained corresponding to the different image views, and wherein the quality classification model to be applied to a given clip is determined based on the view class assigned to the clip using the view identification model.

[0081] Each of the one or more quality classification models may be trained for application to individual images, and wherein determination of a clip quality score is achieved by amalgamating the individual clip quality scores of the individual image frames comprised by the clip (e.g. by averaging).

[0082] Each respective view-specific quality classification model may be trained in a supervised training procedure in which the training input data entries correspond to real or simulated ultrasound images of the relevant view of the cardiovascular system. The training output data entries may correspond to a quality score (e.g. from 0 to 1 or from 1 to 100) of these ultrasound images indicative of a predicted usefulness of the clip for determining a value of the cardiac parameter. The quality score may be manually annotated by a health care professional. Alternatively, each ultrasound image in the training data may be manually classified by a human with a quality class (e.g. BAD, MID, GOOD), and wherein these manual classifications are then converted into numerical scores (e.g. using the mapping: BAD=0, MID=0.5, GOOD=1), and wherein the images in the training data are then annotated with the converted scores. This allows the model to be trained to predict a numerical quality score from 0 to 1, based on training data which was initially annotated by a human with a discrete set of quality classes. Accordingly, the quality classification model may be trained using a deep learning algorithm configured to receive an array of training inputs and respective known outputs, wherein a training input comprises an ultrasound image of a cardiovascular system and a respective known output comprises a known quality score of the ultrasound image.

[0083] The determining of the value of the cardiac parameter from the selected one or more clip sets may comprise use of a cardiac parameter determination model.

[0084] The cardiac parameter determination model may comprise a machine learning model, for example comprising one or more CNNs. As discussed above, there are many known ways of training machine learning models to generate a required output from input data. By way of example, the cardiac parameter determination model may be trained using a deep learning algorithm configured to receive an array of training inputs and respective known outputs, wherein each respective training input comprises a set of ultrasound images of a cardiovascular system corresponding to a pre-defined combination of views and each respective known output comprises a known value of a cardiac parameter of the cardiovascular system. In this way, the cardiac parameter determination model may be trained to automatically determine the value of a cardiac parameter from a selected set of ultrasound images, and wherein the set of ultrasound images may correspond to a particular combination of image views. The set of images which is actually input to the model may comprise one image per required view. The one image for each view may be extracted from an acquired clip corresponding to the given view.

[0085] The model may be trained to compute any of a variety of different cardiac parameters. By way of non-limiting example, these may include any one or more of left- ventricular ejection fraction (LVEF), global longitudinal strain (GLS), longitudinal strain per view (meaning the average longitudinal strain calculated from one standard apical view: A4C, A2C or A3C), segmental strain (meaning longitudinal strain measured in individual myocardial segments of the left ventricle), and ejection fraction per view (meaning an estimation of the ejection fraction EF from one individual apical view: A4C, A2C or A3C).

[0086] As discussed above, embodiments of the invention involve compiling 20 from the clips of the study dataset, a plurality of clip sets, each clip set comprising a set of clips corresponding to a pre-determined combination of views required for computing the cardiac parameter. The views included in each clip set may be determined in dependence upon the cardiac parameter which is to be computed.

[0087] In some embodiments, the pre-determined combination of views to which the clips of at least a subset of the clip sets correspond comprises a triplet of image views consisting of: a two-chamber apical (A2C) view, a three-chamber apical (A3C) view; and a four-chamber apical (A4C) view. Additionally or alternatively, in some embodiments, the pre-determined combination of views to which at least a subset of the clips of each clip set corresponds comprises a pair of image views consisting of: a two-chamber apical (A2C) view; and a four- chamber apical (A4C) view. In some embodiments, the clip sets may include both triplets and pairs.

[0088] To explain further, dependent upon the cardiac parameter being computed, different combinations of views are required or recommended. For example, for computing global longitudinal strain (GLS), it is recommended to use a triplet of image views comprising A2C, A3C and A4C. For computing left-ventricular ejection fraction (LVEF), the modified Simpson biplane may be used, which comprises a pair of views consisting of A2C and A4C. For some cardiac parameters, only a single image view is required. For example, for computing longitudinal strain per view or ejection fraction per view, only one of A4C, A2C or A3C is required in each case.

[0089] In some embodiments, the method is operable in different modes corresponding to different target cardiac parameters. If the target cardiac parameter requires all three of A2C, A3C and A4C, then the clip sets may include only triplets (comprising A2C, A3C, A4C). If the target cardiac parameter requires only A2C and A4C (e.g. Simpson biplane), then the clip sets may include just triplets (A2C, A3C, A4C), just pairs (A2C, A4C), or a combination of triplets and pairs. If the target cardiac parameter requires only a single one of A2C or A4C, then the clips sets may include any combination of: triplets, pairs or single views.

[0090] Fig. 3 outlines steps of an example method in accordance with a particular set of embodiments of the invention. The method is an example embodiment of the method outlined in Fig. 1. The method of Fig. 3 includes all of the steps comprised by the method of claim 1, with some additional steps. Each one of these additional steps is entirely optional and, unless explicitly described otherwise below, independent from any other additional step and, hence, can be omitted. Thus, other exemplary embodiments of a method in accordance with the present invention may comprise fewer, or even just one, of the additional steps described in the context of Fig. 3. Steps of the method of Fig. 3 which correspond to steps of the method of Fig. 1 are labelled with the same reference numerals. These steps may be implemented in the same way as for the method of Fig. 1 for example. The method of Fig. 3 is notionally divided into two subportions, labelled A and B. Portion A comprises steps which can be understood as being performed at a clip-level, i.e. comprising steps which are applied to each individual clip. Portion B comprises steps which can be understood as being performed at a study level, i.e. comprising steps which are applied to the overall dataset. Further to the steps described with reference to Fig. 1, the method of Fig. 3 comprises a preliminary step 52, applied after receiving or obtaining the study dataset 12, comprising processing the study dataset with a contrast classification model configured to classify each clip as to whether it contains any contrast images, and discarding any clips classified as containing contrast images.

[0091] For example, the various classification models may be trained specifically for non-contrast image, and so filtering out contrast images avoids false results. A contrast image means an image acquired with contrast agent.

[0092] This may be performed based on use of a machine learning algorithm utilizing one or a plurality of convolutional neural networks (CNNs). This may be trained to identify and exclude contrast images.

[0093] Optionally, the preliminary step 52 may additionally or alternatively comprise determining whether each image clip comprises non B-mode images and discarding image clips which comprise non-B-mode images. This may also be achieved using a trained machine learning algorithm. Optionally a same single machine learning algorithm may be used which is trained to identify both non B-mode images and contrast images. This provides a filtering operation, wherein the system filters out non-valid B-mode images and contrast images, ensuring that only relevant non-contrast images are processed.

[0094] The method of Fig. 3 further comprises, after determining 14 the view score for each clip, identifying and discarding 54 any clips having a view score which fails to meet a predetermined criterion. For example, this step may comprise discarding any clips having view scores which fail to exceed a pre-determined threshold.

[0095] In some embodiments, and as illustrated in Fig. 3, the method may optionally comprise performing depth adjustment 56 of view scores. This may comprise, for at least a subset of the clips: identifying depth information for each of the at least subset of clips, relating to an imaging depth of ultrasound image frames contained in each clip. Following this, after applying the view identification model to the study dataset, and prior to determining the composite clip scores, depth-adjustment 56 may be performed of the view scores for each of the at least subset of clips by applying a depth-adjustment to each view score which is a function of an imaging depth associated with the image frames contained in the respective clip. Generally, lower (i.e. more shallow) imaging depths are preferable for optimal cardiac parameter determination. This means that image clips comprising frames with less penetration into the body, i.e., more shallow images, are preferred. The depth adjustment may comprise applying a depth-adjustment function. The depth adjustment function may comprise any function that modifies the view scores based on the imaging depth information. For example, it may increase scores for clips containing frames with imaging depths below a certain threshold. Conversely, it may reduce scores for clips comprising image frames associated with an imaging depth above a certain threshold, indicating that these frames are relatively shallow. Alternatively, a linear adjustment of the scores may be applied as a function of depth.

[0096] By way of example, in one advantageous implementation, the depth adjustment function may be a linear function of both the original view score and a depth parameter which is correlated with imaging depth.

[0097] By way of example, in one advantageous implementation, the depth adjustment function may take the following form:

[0098] Depth- Adjusted View Score = (a * Original View Score) - ( * Depth) + y where Original View Score is the view score for the relevant image prior to depthadjustment, a and P are tunable coefficients, y is a tunable offset, Depth is a parameter indicative of imaging depth of the relevant image, and Depth -Adjusted View Score is the view score for the relevant image after adjustment. This may be referred to as linear interpolation.

[0099] The optimal values for a, P and y may be determined using optimization techniques, for example linear regression.

[0100] By way of non-limiting example, in one example implementation, a depthadjustment function was used: Depth- Adjusted View Score = (0.077 * Original View Score) - (0.004 * Depth) + 0.943, and wherein the adjusted view score is then clipped between 0 to 1. It is noted that in this example implementation the Original View Score and Depth were not normalized (i.e. Depth » Original View Score).

[0101] In some embodiments, and as illustrated in Fig. 3, the method may comprise, after determining the composite clip scores, ranking 58 the clips according to the composite clip scores and selecting clips for inclusion in the clip sets based on the rankings of the clips. For example this may comprise prioritizing higher ranking clips.

[0102] In some embodiments, and as illustrated in Fig. 3, after ranking 58 the clips according to composite clip scores, the method may optionally comprise modifying 60 the clip scores of a top-ranking subset of the clips according to imaging depth. This may comprise: identifying a top ranking subset of clips and, for the top-ranking subset of clips: identifying depth information for each of the clips relating to an imaging depth of ultrasound image frames contained in each clip, and subsequently modifying 60 the composite clip scores of the top ranking subset of the clips based on the depth information to prioritize clips comprising image frames corresponding to lower imaging depth (i.e. shallower depth). The top ranking clips may all be very close together with respect to their composite clip scores. Therefore this steps helps to add additional granularity to the rankings near the top. In some embodiments, the selecting one or more of the clip sets based on the rankings comprises selecting a top-ranking subset of one or more of the ranked clip sets. The modifying the clip scores may comprise increasing the clip scores of clips having imaging depths falling within a preferred depth range, for example exceeding a pre-determined depth threshold.

[0103] By way of one example implementation, the optional step of modifying 60 the clip scores procedure may comprise, for each recognized anatomical view, identifying the clip with the highest composite clip score. Once this top-scoring clip is determined, a subset of clips having the same view is selected within a defined score proximity, such as those clips having scores within a range of [S max - A, S max], where S max denotes the highest composite clip score for that view and A is a predetermined proximity margin, for instance, 0.05.

[0104] Subsequently, the depth information associated with each clip in the subset is extracted. This depth information corresponds to the imaging depth of the ultrasound image frames constituting the respective clips. If it is determined that the maximum difference in imaging depth among the selected clips exceeds a predefined threshold, e.g. 2.1 cm, the scores of the subset clips are further modified based on depth.

[0105] The modification comprises reordering the subset of clips in ascending order of imaging depth, such that clips with shallower imaging depths are prioritized. The original composite clip scores of the subset are disregarded for the purpose of this reordering. A new set of scores is then assigned by linearly scaling the clips across the same numeric interval [S max - A, S max], based on their relative position in the reordered list.

[0106] Through this optional procedure, clips within the high-ranking group are differentiated based on depth-related image content, thereby enabling finer prioritization in situations where composite scores alone do not provide sufficient separation. Optionally, this scoring refinement may be applied only when the range of imaging depths among the top clips exceeds a minimum variation threshold, so that the depth-based prioritization is performed only when a minimum variability is present.

[0107] An example implementation of the method is outlined in more detail in Fig. 4 and

[0108] Fig. 5. Fig. 4 shows steps associated with functionality (i) mentioned above: view and quality classification of image clips. Fig. 4 covers steps 12, 14 and 16 of the method of Fig. 1 and steps 12, 52, 14 and 16 of the method of Fig. 3.

[0109] Fig. 5 shows steps associated with functionality (ii) mentioned above: selection and ranking of candidate combinations (sets) of image clips for use in computing the one or more cardiac image parameters.

[0110] Fig. 5 covers steps 18, 20 and 22 to 22 of the method of Fig. 1 and steps 54, 56, 18, 58, 60, 20 and 22 of the method of Fig. 3

[0111] Referring first to Fig. 4, the method begins with an input study dataset 12 comprising a plurality of ultrasound image clips of the cardiovascular system, each clip comprising a set of one or more ultrasound image frames corresponding to a same view of the cardiovascular system.

[0112] Optionally, the method may comprise a pre-processing step 72 which may comprise extracting or separating the individual one or more image frames of each image clip.

[0113] The pre-processing step 72 may comprise applying a preprocessing module which may comprise sampling a subset of the frames in each clip, e.g. 20% of frames from each clip, either in absolute terms or at fps (frame-per-second) intervals. The pre-processing module may be configured to perform sector mask extraction and cropping, comprising extracting for each image frame a sector mask and cropping the sector using this mask. The preprocessing module may be configured to perform normalization and resizing of each image frame comprising normalizing the image frames and resizing them to fit the input size required by the view identification model and the quality classification model.

[0114] Optionally, the method further may comprise a step 52 of processing the study dataset with a contrast classification model configured to classify each clip as to whether it contains any contrast images, and disregarding (for purposes of the remainder of the method 10) as invalid 53 (referred to for shorthand as ‘discarding’) any clips classified as containing contrast images. The remaining steps of the method are applied only to image clips which have been classified as not containing any image frames which are contrast images.

[0115] The model may for example be applicable only for non-contrast images. A contrast image means an image acquired with contrast agent.

[0116] Additionally or alternatively, step 52 may comprise detecting and discarding from the remainder of the method any image clips containing image frames which are not B-mode images, for example based on detecting DICOM tags. In some embodiments, step 52 may comprise applying a filtering module which is configured to perform both non-B-mode filtering and contrast filtering as described above. The filtering module may be configured to detect and filter out (i.e. discard or disregard from the remainder of the method) non- B-mode images, for example using DICOM tags or classical image processing techniques. This can be achieved using a trained CNN model comprising MobileNetV2 architecture, achieving 99% accuracy.

[0117] Step 52 may be omitted if preferred, or any other validity filtering criteria could be applied instead, as appropriate to the model input requirements, to ensure that the images comprised by the image clips meet input requirements for the view identification module and quality classification module.

[0118] The method further comprises processing the study dataset with a view identification model (view classifier) trained to predict 14 a view represented by each ultrasound image clip and an associated view score indicative of a likelihood that the predicted view of the ultrasound image clip is correct. The view classification module may for example comprise a trained CNN comprising MobileNetV2 architecture for view classification. An output of the view classification module is one of a set of possible view classes 74, and a view score representing a probability that the view class is correct, i.e. a model confidence score. Many machine learning models, including convolutional neural networks like MobileNetV2, inherently generate confidence scores as part of their classification process. These scores typically represent the model's estimated probability for each possible class. The highest probability is usually associated with the predicted class, and this probability value serves as the confidence score. In the context of the view classification, this score indicates the model's certainty in the generated view prediction. Higher scores suggest greater confidence in the classification, while lower scores may indicate less certainty or potential ambiguity in the image view.

[0119] In the illustrated example, the view identification model is trained to classify each input image into one of a set of view classes which comprise: 2-chamber apical view (A2C), 3- chamber apical view (A3C), 4-chamber apical view (A4C), RV-focused 4-chamber apical view (RV), and parasternal long axis view (PLAX). The set of classifications may also include an ‘other’ class (OTHR) for views that do not fall into one of the other classifications.

[0120] The method further comprises a quality classification process 16 comprising determining, for each of the clips, using a quality classification model, a clip quality score 78 indicative of a predicted usefulness of the clip for determining a value of a cardiac parameter of interest. In the illustrated example in Fig. 4, the process of determining 16 the clip quality score 78 comprises processing the image frames of each respective clip with a view-specific quality classifier selected from among a bank 76 of different view-specific quality classifiers based on the view classification 74 of the clip. In other words, a different quality classifier is trained for application to images corresponding to different views of the cardiovascular system.

[0121] The output from each view-specific quality classifier includes a quality score (e.g. from 0 to 1 or 1 to 100) indicative of a predicted usefulness of an image frame for use in determining a cardiac parameter of interest, e.g. a probability that the image will result in accurate prediction of a cardiac parameter using a cardiac parameter determination model.

[0122] Each view-specific quality classifier may be trained to analyze a single frame input, and wherein an overall clip quality score may be generated for each clip in a post processing step 78 by amalgamation of the individual quality scores / classes generated for the individual image frames of the clip. For example, a clip quality score may be generated by taking an average of the quality scores generated by the relevant quality classification model for the set of image frames comprised by the clip. Instead of averaging, a most common accumulation method may be employed to determine the clip-level quality score from the individual image frame quality scores.

[0123] The same principles apply also the view classification 14, wherein the view identification module may be configured to process individual image frames and generate a view classification for each image frame, and wherein an overall clip view class can be determined by taking a most commonly / frequently occurring view class from among the image frames comprised by the clip and an overall clip view score might be obtained by taking an average of the clip scores of the individual frames comprised by the clip.

[0124] The output 80 from the view and quality classification process represented in Fig. 4 is a clip view class, a clip view score and a clip quality score for each clip (for each clip that was not rejected by the filter module 52).

[0125] Referring now to Fig. 5, this illustrates steps in accordance with one example set of embodiments for identifying candidate clips and clip sets to use for determination of at least one cardiac parameter based on the computed clip view scores and clip quality scores (and optionally also image depth).

[0126] The input to the functionality represented in Fig. 5 is the output 80 from the functionality represented in Fig. 4.

[0127] In the example of Fig. 5, the method includes a view score filter step 54, comprising identifying and ignoring 92 (or discarding or disregarding) for the remainder of the method any clips having a view score which fails to meet a pre-determined criterion. By way of example, the criterion may be that the score is above a threshold. The threshold is labelled View thresh in Fig. 5. For example, the method may comprise identifying and disregarding 92 from the remainder of the method any clips having a clip view score which is below the threshold View thresh, e.g. a threshold of 0.5 (although of the course the threshold can be set as desired).

[0128] For clips having a view score which meets the pre-determined criterion, e.g. is above a threshold, a step 56 is performed of adjusting the view scores based on image depth. An example implementation of this step was described above with reference to Fig. 3 and the same implementation details are equally applicable here.

[0129] Following the depth-adjustment step 56, the method comprises determining 18 a composite clip score for each clip. The composite clip score is computed as a function of the view score and quality score for each clip, for example based on the product of the (depth- adjusted) view score and the quality score for each clip.

[0130] Optionally, the method further comprises a step 94 of updating the composite clip score for clips corresponding to the 4-chamber apical view (A4C) or the RV-focused 4-chamber view (RV). For example, in some embodiments, the step 94 of updating the composite clips scores may comprise updating the clip scores to prioritize diagnostically valuable views of the left ventricle (LV) chamber. For example, the updating may be configured to favor high-quality four-chamber (4CH) views while also incorporating high-quality right ventricle RV-focused 4- chamber views, provided they demonstrate sufficient LV chamber visibility.

[0131] In some embodiments, an updated composite clip score may be computed for each A4C and RV clip based on the view class and the original clip score in accordance with the following scheme. This scheme assumes that the original composite clip scores are in a range [0, 1], When a clip is identified as a 4CH view and is assigned an original score greater than 0.6, it is considered to have high diagnostic utility and is therefore scaled into the range [0.6, 1.0] in the updated scoring schema. This range for example corresponds to clips that are likely to provide clear and reliable representations of cardiac chamber morphology. In cases where the clip is identified as an RV-focused 4-chamber view and the original clip score exceeds a threshold of 0.65, the clip is deemed to contain diagnostically useful information regarding the LV chamber despite its RV classification. Accordingly, such RV clips are mapped to an intermediate score range of [0.4, 0.6], thereby assigning them a preferential status above low-quality 4CH clips, yet below high-quality 4CH clips. When a clip is identified as a 4CH clip with an original clip score of 0.6 or less, this may be scaled into a lower score range of [0, 0.4] in the updated scoring representation. This approach ensures that diagnostically inferior 4CH clips are deprioritized relative to higher- quality content from either view type.

[0132] Through this optional score adjustment step 94, high-quality 4CH clips are favored, while high-quality RV clips are conditionally prioritized to ensure prioritization of LV chamber visibility.

[0133] The method may further comprise ranking 58 the clips according to clip score and subsequently adjusting the top-scoring scores according to imaging depth of the images comprised by the clips, in the manner already described previously. Clips comprising images with shallower image depth may be prioritized over clips comprising images with greater penetration depth. An example implementation of this step was described above with reference to Fig. 3 and the same implementation details are equally applicable here.

[0134] The method further comprises compiling 20 from the clips of the study dataset, a plurality of clip sets, each clip set comprising a set of clips corresponding to a pre-determined combination of views required for computing the cardiac parameter.

[0135] The method further comprises determining 22 for each of the plurality of clip sets, a clip set ranking, wherein the clip set ranking is computed for each clip set at least in part as a function of the individual composite clip scores of the set of clips comprised by the clip set.

[0136] Regarding the step of compiling 20 the clip sets, this may comprise compiling a plurality of clips sets comprising clips corresponding to the same combination of views, composed from different selections of clips representing those views. An overall clip set ranking can then be determined 22 for each clip set based on a combination of the clip scores (e.g. based on a product of the clip scores) of the clips comprised by each clip set. The method may then comprise selecting from the plurality of clip sets based on the rankings, e.g. prioritizing higher ranking clip sets, i.e. those whose component clips have higher clip scores.

[0137] An example process for selecting the clips for use in computing the cardiac parameter is outlined in more detail in Fig. 6.

[0138] Fig. 6 shows an example implementation of steps 20, 22, 24 and 26 of the method of Fig. 1 and Fig. 3, including a process for selecting clip sets to utilize for determining the cardiac parameter. The process comprises compiling 20 the plurality of clip sets as discussed previously, each clip set comprising a set of clips corresponding to a pre-determined combination of views required for computing the cardiac parameter. The clip sets may comprise various combinations, depending on the specific cardiac parameter being computed. For instance, some parameters may require triplets (e.g., two-chamber, three-chamber, and four- chamber views), while others may be determinable using pairs (e.g., two-chamber and four- chamber views), or from triplets or pairs, or even from a single view.

[0139] The method further comprises determining 22, for each of the plurality of clip sets, a clip set ranking, wherein the clip set ranking is computed for each clip set at least in part as a function of the individual composite clip scores of the set of clips comprised by the clip set.

[0140] Once the clip sets are ranked, the method proceeds to an iterative selection and evaluation process.

[0141] The method comprises determining an ordered list of the clip sets, ordered from higher ranking clip sets to lower ranking clip sets. The method further comprises indexing 24 through the clip sets of the ordered list one by one, applying 26 a cardiac parameter determination model to each clip set, evaluating 27 an output of a cardiac parameter determination model against a pre-defined success criterion and, continuing 24 to the next clip set in the list if the success criterion is not met until the success criterion is met, at which point the computed cardiac parameter is output 102. With reference to Fig. 6 step 24 comprises selecting the first (or next) clip set in rank order. This selection strategy ensures that the highest- ranked sets, which are likely to yield the most accurate results, are evaluated first. Step 26 comprises applying the cardiac parameter determination model to the selected clip set to attempt to compute the cardiac parameter using the selected clip set.

[0142] The success of the computation is evaluated at decision node 27. If the computation is unsuccessful (following the "No" path from node 27), the process loops back to step 24, where the next clip set in the ordered list is selected. This iterative approach continues until a successful computation is achieved or all clip sets have been exhausted.

[0143] Upon successful computation (following the "Yes" path from node 27), the process concludes with step 102, where the calculated cardiac parameter is output.

[0144] With regards to the step 27 of evaluating success of the cardiac parameter determination, this may comprise evaluating a pre-defined success criterion. There are different options for the success criterion. In some embodiments, the cardiac parameter determination model may be configured to output a confidence score associated with each output cardiac parameter value. The pre-determined success criterion may relate to the confidence score, for example it may comprise a threshold value for the confidence score.

[0145] As discussed above, embodiments of the invention employ use of a view identification (or classification) model and a quality classification model. An example process of training of the models will now be briefly discussed. For each of the view classification and quality classification models, a training image dataset is obtained comprising a large number of ultrasound image clips representing the anatomy of interest, e.g. the cardiovascular system in this case, e.g. the heart. A training dataset for the view classification model is then obtained by annotating the image frames of each clip in accordance with the view represented by the image frames. A training dataset for the quality classification model is obtained by annotating the image frames of each clip in accordance with a quality classification and / or quality score of the image frames.

[0146] By way of example, in one example implementation, the image dataset for the view classification model included 6103 clips from 1136 echo examinations and the image dataset for the quality classification model included 1338 clips from 712 echo exams. The patient population primarily consisted of adults with a mean age of 62.8 years, ranging from 18 to 100 years old. This large and diverse dataset, sourced from multiple systems and sites, ensures robustness and provides a well-rounded representation of real-world clinical validation data. This diversity supports the generalizability of models across different imaging systems and patient populations.

[0147] The image data in each dataset comprised DICOM clips from retrospectively collected echocardiographic examinations. These examinations were acquired using a variety of ultrasound systems from different manufacturers. The clips were captured at frame rates ranging from 25 to 80 frames per second.

[0148] A preprocessing module may be implemented to prepare the data for model training. This module performs three main functions. First, it samples 20% of frames from each ultrasound clip, either in absolute terms or at regular frame-per-second intervals. Second, it extracts the sector mask from each sampled frame and applies tight cropping based on this mask. Finally, the module normalizes and resizes the images to match the input requirements of the Convolutional Neural Networks (CNNs) used in the system.

[0149] The data annotation process may be performed by expert sonographers. For example, for the view classification model, annotation of the view represented in each image was performed independently by two sonographers for each image. In cases of disagreement, a third independent sonographer reviews and tags the view to ensure accuracy. Annotation of the quality classification is performed by one of the sonographers, with categories assessed based on endocardium visibility and the feasibility of visually estimating heart function. View labels in the dataset include A2C, A3C, A4C, RV, PL AX, OTHER, CONT A2C, CONT A4C, and CONT OTHER. Quality labels are categorized as BAD, MID, and GOOD. The quality label annotations were subsequently converted into numerical scores (e.g. using the mapping: BAD=0, MID=0.5, GOOD=1), and wherein the final images in the training data were annotated with the converted numerical quality scores.

[0150] The model training process involves 200 epochs. The validation set is evaluated at each epoch to generate a corresponding validation confusion matrix. The optimal checkpoint is selected based on a combination of model stability and superior performance metrics. This selection considers both parameter optimization and architectural suitability. Specific success criteria are established for each classification task. The training process is considered complete when these predefined criteria are met.

[0151] The tuning data is kept separate from the training data. It is partitioned to mirror the features and demographics of the training data. This process is conducted independently of the algorithm development team to ensure proper data representation and robustness. After each training session, the model is evaluated on the tuning data.

[0152] The testing data is also kept separate from both the training and tuning data. It is partitioned to mirror the training data's features and demographics. This process is also conducted independently of the algorithm development team. The final models are evaluated on this testing data.

[0153] Embodiments of the invention outlined above provide a variety of advantages. One advantage is improved diagnosis speed. By automating the view and quality classification process, the proposed solution significantly increases the speed of diagnosis, and frees up clinician time. A further advantage is reducing the cognitive load and time pressure on clinicians, which can help to mitigate burnout and improve overall job satisfaction. A further advantage is enhancing consistency and accuracy of cardiac parameter determination. Embodiments of the invention provide more consistent and accurate classification of echocardiogram images compared with human view and quality classification. A further advantage is the potential to support more junior sonographers or less skilled sonographers in performing echocardiography examinations which previously could only be performed by, or under the direct supervision of, highly skilled sonographers.

[0154] The proposed structure allows for the overall model to be compact, requiring minimal computational resources while still maintaining high performance.

[0155] Embodiments of the invention described above employ a processing device 32.

[0156] The processing device may in general comprise a single processor or a plurality of processors. It may be located in a single containing device, structure or unit, or it may be distributed between a plurality of different devices, structures or units. Reference therefore to the processing device being adapted or configured to perform a particular step or task may correspond to that step or task being performed by any one or more of a plurality of processing components, either alone or in combination. The skilled person will understand how such a distributed processing device can be implemented. The processing device includes a communication module or input / output for receiving data and outputting data to further components.

[0157] The one or more processors of the processing device can be implemented in numerous ways, with software and / or hardware, to perform the various functions required. A processor typically employs one or more microprocessors that may be programmed using software (e.g., microcode) to perform the required functions. The processor may be implemented as a combination of dedicated hardware to perform some functions and one or more programmed microprocessors and associated circuitry to perform other functions.

[0158] Examples of circuitry that may be employed in various embodiments of the present disclosure include, but are not limited to, conventional microprocessors, application specific integrated circuits (ASICs), and field-programmable gate arrays (FPGAs).

[0159] In various implementations, the processor may be associated with one or more storage media such as volatile and non-volatile computer memory such as RAM, PROM, EPROM, and EEPROM. The storage media may be encoded with one or more programs that, when executed on one or more processors and / or controllers, perform the required functions. Various storage media may be fixed within a processor or controller or may be transportable, such that the one or more programs stored thereon can be loaded into a processor.

[0160] Variations to the disclosed embodiments can be understood and effected by those skilled in the art in practicing the claimed invention, from a study of the drawings, the disclosure and the appended claims. In the claims, the word "comprising" does not exclude other elements or steps, and the indefinite article "a" or "an" does not exclude a plurality.

[0161] A single processor or other unit may fulfill the functions of several items recited in the claims.

[0162] The mere fact that certain measures are recited in mutually different dependent claims does not indicate that a combination of these measures cannot be used to advantage.

[0163] A computer program may be stored / distributed on a suitable medium, such as an optical storage medium or a solid-state medium supplied together with or as part of other hardware, but may also be distributed in other forms, such as via the Internet or other wired or wireless telecommunication systems.

[0164] If the term "adapted to" is used in the claims or description, it is noted the term "adapted to" is intended to be equivalent to the term "configured to".

[0165] Any reference signs in the claims should not be construed as limiting the scope.

Claims

CLAIMS:

1. A computer-implemented method (10) for determining a value of a cardiac parameter of a cardiovascular system based on a study dataset (12) of ultrasound image clips of the cardiovascular system, each clip comprising a set of one or more ultrasound image frames corresponding to a same view of the cardiovascular system, the method comprising: processing the study dataset with a view identification model to predict (14) a view represented by each ultrasound image clip and an associated view score indicative of a likelihood that the predicted view of the ultrasound image clip is correct; determining (16) for each of the clips, using a quality classification model, a clip quality score indicative of a predicted usefulness of the clip for determining a value of the cardiac parameter; determining (18) for each of the clips a composite clip score, wherein the composite clip score is computed as a function of the view score and clip quality score; compiling (20) from the clips of the study dataset, a plurality of clip sets, each clip set comprising a set of clips corresponding to a pre-determined combination of views required for computing the cardiac parameter; determining (22), for each of the plurality of clip sets, a clip set ranking, wherein the clip set ranking is computed for each clip set at least in part as a function of the individual composite clip scores of the set of clips comprised by the clip set; selecting (24) one or more of the clip sets based on the clip set rankings; and determining (26), using the clips comprised by the selected one or more of the clip sets, a value of the cardiac parameter of the cardiovascular system.

2. The method (10) of claim 1, wherein the method further comprises, after determining the view score for each clip, identifying and discarding (54) any clips having a view score which fails to meet a pre-determined criterion.

3. The method of claim 1 or 2, wherein the composite clip score is computed (18) as a function of the product of the view score and clip quality score.

4. The method of any preceding claim, wherein the method further comprises, for at least a subset of the clips: identifying depth information for each of the at least subset of clips, relating to an imaging depth of ultrasound image frames contained in each clip; and after applying the view identification model to the study dataset, and prior to determining the composite clip scores, performing depth-adjustment (56) of the view scores for each of the at least subset of clips by applying a depth-adjustment to each view score which is a function of an imaging depth associated with the image frames contained in the respective clip.

5. The method of any preceding claim, wherein the method further comprises: after determining the composite clip scores, ranking (58) the clips according to the composite clip scores; and selecting clips for inclusion in the clip sets based on the rankings of the clips.

6. The method of any preceding claim, wherein the method further comprises, after determining the composite clip scores and before compiling the clip sets: ranking (58) the clips according to composite clip scores; identifying a top ranking subset of clips; for the top-ranking subset of clips, identifying depth information for each of the clips relating to an imaging depth of ultrasound image frames contained in each clip; and modifying (60) the composite clip scores of the top ranking subset of the clips based on the depth information to prioritize clips comprising shallower imaging depth image frames.

7. The method of any preceding claim, wherein the steps of selecting (24) one or more of the clip sets based on the rankings, and determining the cardiac parameter using the selected clip sets comprises: determining an ordered list of the clip sets, ordered from higher ranking clip sets to lower ranking clip sets; and indexing through the clips sets of the ordered list one by one, applying a cardiac parameter determination model to each clip set, evaluating an output of a cardiac parameter determination model against a pre-defined success criterion and, continuing to the next clip set in the list if the success criterion is not met until the success criterion is met.

8. The method of any preceding claim, wherein the determining the clip quality score comprises processing the image frames of the respective clip with a view-specific quality classifier selected from among a bank of different view-specific quality classifiers based on the view classification of the clip.

9. The method of any preceding claim, wherein the method further comprises a preliminary step (52) of processing the study dataset with a contrast classification model configured to classify each clip as to whether it contains any contrast images, and discarding any clips classified as containing contrast images.

10. The method of any preceding claim, wherein the pre-determined combination of views to which the clips of at least a subset of the clip sets correspond comprises a triplet of image views consisting of: a two-chamber apical (A2C) view; a three-chamber apical (A3C) view; and a four-chamber apical (A4C) view.

11. The method of any of claims 1-9, wherein the pre-determined combination of views to which at least a subset of the clips of each clip set corresponds comprises a pair of image views consisting of: a two-chamber apical (A2C) view; and a four-chamber apical (A4C) view.

12. The computer-implemented method of any preceding claim, wherein the cardiac parameter value is at least one of: a left ventricular ejection fraction value; global longitudinal strain value; a segmental strain value; a longitudinal strain per view value; and an ejection fraction per view value.

13. A computer program product comprising computer program code configured, when run on a processor, to implement the method of any of claims 1-12.

14. A processing device comprising one or more processors configured to perform a method in accordance with any of claims 1-12.

Citation Information

Patent Citations

  • Automated cardiac function assessment by echocardiography

    US20210000449A1

  • Systems and methods for automated physiological parameter estimation from ultrasound image sequences

    US20210330285A1