Angiographic image processing method and system
The angiographic image processing method uses a convolutional transformer network to analyze spatiotemporal features from angiographic videos, addressing the limitations of invasive FFR/iFR methods by providing accurate, non-invasive stenosis assessment.
Patent Information
- Application Number
- PCT/IB2025/056252
- Authority / Receiving Office
- WO · WO
- Patent Type
- Applications
- Current Assignee / Owner
- Priority Date
- 2024-06-21
- Filing Date
- 2025-06-19
- Publication Date
- 2025-12-26
AI Technical Summary
Conventional methods for determining the severity of coronary stenoses, such as FFR and iFR, are invasive, time-consuming, costly, and subject to variability due to human expertise, with existing image analysis techniques struggling to accurately assess stenosis in complex medical images.
An angiographic image processing method using a convolutional transformer network to analyze spatiotemporal features from angiographic videos, providing direct and indirect estimates of FFR/iFR values without laborious operations, leveraging artificial neural networks to classify stenosis severity and regress physiological indices.
This approach reduces invasiveness and variability by autonomously focusing on relevant frames, facilitating accurate assessment of stenosis severity through non-invasive medical imaging, thereby supporting clinical decision-making.
Smart Images

Figure IB2025056252_26122025_PF_FP_ABST
Abstract
Description
[0001] "Angiographic image processing method and system"
[0002] TEXT OF THE DESCRIPTION
[0003] Technical field
[0004] The disclosure relates to methods and systems for processing images, in particular digital images.
[0005] One or more embodiments may be applied in the medical field, for example for the cardiological characterization of coronary lesions or multi-vessel stenoses.
[0006] One or more embodiments may be used to support diagnostic activity in the medical field.
[0007] Background
[0008] The ability to determine the stage of decompensation of coronary stenoses or multi-vessel lesions can play a significant role in the health sector, for example in determining whether or not to perform a surgical intervention to apply stents to prevent the occlusion of one or more blood vessels.
[0009] Conventional techniques for making this determination are known as: fractional flow reserve (FFR) measurement and instantaneous fractional flow ratio (iFR). These techniques involve the use of catheters and therefore represent invasive tests for patients. The invasive nature of these measurements also poses some limitations in obtaining them: time needed to perform the measurements, cost of the medical procedure, risk of unwanted complications, and so on.
[0010] Although physicians can perform a careful examination of the FFR or iFR results, the heuristic experience of the physician may play a role in the localization of major stenoses. Consequently, these medical procedures may be subject to significant variations based on the skill and experience of the analyzer.
[0011] The possibility of reducing the impact of the "human factor" by providing some kind of technical assistance to the medical diagnosis has been investigated to some extent.
[0012] The measurement of FFR / iFR is being spread, despite the limitations illustrated, as a guide to revascularization strategies in multi-vessel disease. There are studies that analyze how to use them as indicators of the opportunity to subject a patient to revascularization interventions or not.
[0013] Especially in high-volume medical centers, there is the problem of reducing the time to obtain estimates of the hemodynamic significance of stenoses, in order to provide support to physicians in making correct decisions, as well as reduce the procedures to which patients must undergo.
[0014] A possible alternative approach to the invasive examination consists in quantifying the degree of stenosis and through image analysis techniques obtained through medical imaging.
[0015] In the literature, there are proposals for the application of artificial neural network processing procedures (also known as machine learning or deep learning, in short called ANN - artificial neural network) in the analysis of cardiovascular images.
[0016] Examples of existing approaches are discussed in the following documents:
[0017] WO 2021 / 117043 Al discusses an automated image- based blood vessel analysis solution, wherein an analysis system receives longitudinal 2D images of patient vessels, obtained during a X-ray angiography and applies a pre-trained classifier on those images to provide an indication of the presence of stenosis in the vessels and a location of the localized stenosis, which can be displayed on a display.
[0018] US 2022 / 215542 Al discusses a computer-based method for determining the presence of a coronary stenosis for a patient, comprising a step of receiving at least one multiplanar medical tomographic CT (X-scanner) image of the patient comprising a coronary stenosis, the method further comprising a step of detecting the coronary stenosis on the image or a part of the image using a first deep neural network.
[0019] US 2015 / 112182 Al discusses a method and system for determining the fractional flow reserve (FFR) for a coronary artery stenosis of a patient from a medical image of the patient.
[0020] US 2022 / 215541 Al discusses a computer-based method for determining the presence of a coronary lesion for a patient, the method comprising: receiving at least one multiplanar medical CT image of a coronary artery of the patient and determining a Coronary Artery Disease- Reporting and Data System (CAD-RADS) classification value of a coronary lesion on the image or a portion of the image using a previously trained deep neural network.
[0021] US 2020 / 113449 Al discusses a system comprising a computer-readable storage medium with computer- executable instructions, including a biophysical simulator configured to determine a fractional flow reserve value.
[0022] Existing solutions may be limited in recognizing the variety of stenosis, for example, due to the presence of overlapping and occlusive vessels and the relatively small size of the involved image region.
[0023] Another limitation is the use of complex medical machines whose use involves specialists who may be absent in many facilities.
[0024] Additional existing limitations include, for example: the use of a plurality of tests to be performed on patients, and the impact on the activities of the healthcare technician who is in any case involved in identifying frames for the different tests.
[0025] The presence of an external intervention by the physician, in addition to requiring time, introduces a criterion of subjectivity that limits the homogeneity of the results provided on a large scale by existing techniques .
[0026] Despite the extensive activity in this area, as evidenced by the existing literature, further improved solutions are therefore desirable.
[0027] Object and summary
[0028] An object of one or more embodiments is to contribute to providing such an improved solution.
[0029] According to one or more embodiments, this object can be achieved by means of a method having the features set forth in the claims that follow. An image processing method can be an example of such a method.
[0030] One or more embodiments can relate to a corresponding system. An image processing system can be an example of such a system.
[0031] One or more embodiments can relate to a corresponding apparatus.
[0032] One or more embodiments can include a computer program product loadable into the memory of at least one processing circuit (e.g., a computer) and including portions of software code for executing the method steps when the computer program product is executed on at least one processing circuit. A reference to such a computer program is intended to be equivalent to a reference to a computer-readable medium containing instructions for controlling the processing system to coordinate an implementation of the method according to one or more embodiments. A reference to "at least one computer" is intended to highlight the possibility that one or more embodiments may be implemented in a modular and / or distributed form.
[0033] The claims are an integral part of the technical teaching provided herein with respect to the embodiments .
[0034] Embodiments provide an approach to assessing the severity of stenosis from angiographic videos. For example, one or more embodiments provide a direct and indirect estimate of FFR / iFR values from digital images.
[0035] One or more embodiments have the advantage of doing without laborious operations such as collecting multiple views. One or more embodiments may benefit from external certification of the model from data already made available.
[0036] One or more embodiments advantageously use artificial neural network processing (e.g., a convolutional transformer) to facilitate the extraction of spatiotemporal features from angiographic examination videos. For example, this may facilitate the assessment of the significance of coronary stenosis from a single medical image view.
[0037] One or more embodiments exploit a branched architecture to support different assessment modalities. For example, both a classification of hemodynamic significance of the stenosis and a regression of the FFR / iFR values can be provided.
[0038] One or more embodiments facilitate the learning of robust features even under training conditions of the artificial neural network on heterogeneously labeled data sets.
[0039] One or more embodiments use calibration operations and facilitate interpretability.
[0040] One or more embodiments may include verification and validation operations of the artificial neural network processing system starting from data already made available.
[0041] One or more embodiments address the problem of how to extrapolate deformable and dynamic information embedded in angiographic sequences to allow a non- invasive and medical imaging-based estimation of physiological index values.
[0042] One or more embodiments facilitate doing without manual or semi-automatic keyframe selection systems. For example, the process receives a video as input and applies an analysis processing that, through explainability mechanisms, is observed to tend to autonomously focus on relevant frames (which may coincide with keyframes). Advantageously, this result is due to an emergent behavior of the processing model.
[0043] Brief description of various views of the drawings
[0044] One or more embodiments will now be described, purely by way of example, with reference to the attached figures, in which:
[0045] Figure 1 is an exemplary diagram of a digital image processing system according to this disclosure;
[0046] Figures 2 and 3 are exemplary input signals for processing and / or training of artificial neural network stages according to this disclosure;
[0047] Figure 4 is an exemplary diagram of operations in a portion of the processing chain of the system exemplified in Figure 1;
[0048] Figure 5 is an exemplary diagram of an alternative image processing chain according to this disclosure;
[0049] Figure 6 is an exemplary diagram of a training data set for artificial neural network processing according to this disclosure;
[0050] Figure 7 is a further exemplary diagram of a distribution of features in the training data;
[0051] Figure 8 is exemplary of performance of the method according to this disclosure;
[0052] Figure 9 is exemplary of performance parameters of the method according to this disclosure;
[0053] Figure 10 is exemplary of performance parameters of operations of the method according to this disclosure;
[0054] Figure 11 is exemplary of further performance parameters of operations of the method according to this disclosure;
[0055] Figure 12 is exemplary of performance benchmarks of operations of the method according to this disclosure; and
[0056] Figure 13 is exemplary of further performance benchmarks of operations of the method according to this disclosure.
[0057] Corresponding numbers and symbols in the various figures generally refer to corresponding parts, unless otherwise indicated.
[0058] The figures are drawn to clearly illustrate the relevant aspects of the embodiments and are not necessarily drawn to scale.
[0059] The edges of features drawn in the figures do not necessarily indicate the end of the feature extent.
[0060] Detailed description of examples of embodiments
[0061] In the following description, one or more specific details are illustrated in order to provide an in-depth understanding of examples of embodiments of this specification. Embodiments may be achieved without one or more of the specific details or with other methods, components, materials, etc. In other cases, known operations, materials, or structures are not illustrated or described in detail so that certain aspects of the embodiments will not be obscured.
[0062] A reference to "an / one embodiment" in this specification is intended to indicate that a particular configuration, structure, or feature described with respect to the embodiment is included in at least one embodiment. Accordingly, phrases such as "in an / one embodiment" that may appear in one or more places in this specification do not necessarily refer to the same embodiment .
[0063] Furthermore, particular configurations, structures, or features may be combined in any suitable way in one or more embodiments.
[0064] The references used herein are provided merely for convenience and therefore do not define the extent of protection or the scope of embodiments.
[0065] Figure 1 is an exemplary diagram of an artificial network neural processing method, ANN 10, which can be applied to digital videos (i.e., representable by three- dimensional matrices of numerical values) to perform various types of analysis, in particular for the computation of FFR and iFR starting from temporal sequences of images collected, in a way known per se, via angiography. Preferably, the system 10 is configured to extract spatiotemporal features from at least one video comprising a temporal sequence of successive frames (or "still image") of a recording relating to the third temporal dimension observed by detectors of a medical imaging system for coronary angiography (known per se).
[0066] For simplicity, in the following reference is mainly made to the case in which the analyzed images are collected via coronary radiography, i.e., coronary angiography, it being understood that this case is purely exemplary and not limiting. Again for simplicity, reference is mainly made to the case in which the input video Min includes greyscale frames, i.e. representable by associating each pixel with intensity values (for example, values in a range between 0 and 255). This description is purely exemplary and not limiting, it being understood that one or more embodiments could also be applicable to color videos or videos in other color scales. As appreciated by those skilled in the art, the result of a coronary angiography is a coronary angiogram that includes a series of successive X-rays (for example, taken at an interval of 7.5 / 10 / 12.5 / 15 / 30 / 60 Hz) or another type of video diagnostics (medical imaging) known in itself that shows the coronary circulation. This is achieved by injecting a contrast agent (a substance, often iodine-based, that appears opaque to irradiation) into the coronary arteries of the patient, injected intravenously through a catheter inserted into a blood vessel, such as the femoral artery.
[0067] As people in the field may appreciate, a coronary angiography is related to the coronary arteries located on the surface of the heart, which are subject to continuous contractions and dilations due to cardiac systole and diastole. This means that quantitatively estimating blood flow and pressure loss in the coronary arteries may involve dynamic analysis of hemodynamics and arterial deformation. For example, the non-rigid deformation of the coronary arteries due to systolic- diastolic cycles and the variation of the anatomical structure due to cardiac motion make conventional methods unsuitable for identifying a temporal correspondence and alignment between the frames useful for estimating these parameters.
[0068] For example, the input signal matrix Min comprises a sequence comprising a number (integer) of frames T, i.e. two-dimensional images of width W and height H. Figure 1 exemplifies the case in which width T, height H and number of frames T are equal, it being understood that this case is purely exemplary and not limiting. For example, the method is suitable for processing a number of time frames T greater than or equal to two and of any value of width W and height H. One or more embodiments may comprise the operation of collecting the time series of frames of the video Min via an angiography apparatus A comprising, for example, an X-ray source and a detector R (known per se) configured to detect signals indicative of the distribution of X-rays in time and space. The apparatus A also comprises a processing circuitry RC coupled to the detector R to receive the collected signals from it and to perform a medical image formation processing of the video Min starting from the signals themselves, in a way known per se.
[0069] In one or more embodiments, the processing circuitry RC may be configured to include additional artificial neural network processing circuits configured to perform the operations of the processing method ANN 10 in an way integrated with the apparatus A.
[0070] One or more embodiments may include a computer product loadable into the memory of at least one processing circuit and including portions of software code to perform the steps of one or more embodiments of the method 10 when the computer product is executed on at least one processing circuit RC.
[0071] For example, one or more embodiments use graphics processing circuitry (such as a graphics processing unit (GPU), a tensor processing unit (TPU), or a neural processing unit (NPU).
[0072] As exemplified in Figure 1, the processing system 10 comprises a plurality of signal processing blocks, for example: an input processing stage 100 configured to receive the input video and to apply thereto a first artificial neural network processing, in particular a convolutional neural network (CNN) processing; as a result, the input processing stage 100 provides a set of features Sf whose size depends on the number of frames T and the height f_h, width f_w and number of channels f_c parameters of the features to be extracted; a set of network processing (or transformer) layers 102, 104 comprising a first spatial transformation network processing stage 102 and a second temporal transformation network processing stage 104 with both stages 102, 104 comprising self-attention processing layers, as discussed below; for example, both stages 102 and 104 maintain the size of the input data unchanged, processing or modifying the content through self- attention (or importance) matrices; an output CNN neural network stage 106, and a multi-branch analysis stage 110 coupled to the output CNN network stage 106 and configured to provide a plurality of signals 01, 02, 03 indicative of the hemodynamic significance of a stenosis photographed over time through the medical imaging system with which the input signal matrix Min was obtained.
[0073] For example, the multi-branch stage 110 comprises: a first classification branch 114 configured to provide a first signal 01 indicating the significance classification of the stenosis; a second regression branch 116 configured to provide a second signal 02, for example an estimate of an FFR value, and a third regression branch 114 configured to provide a third signal 03, for example an estimate of an iFR value.
[0074] It is noted that the order in which the branches 112, 114, 116 of the multi-branch stage are discussed is purely exemplary and not limiting.
[0075] For example, the multi-branch stage 110 comprises: a first classification branch 112 configured to provide a first signal 01 indicating the significance classification of the stenosis; a second regression branch 114 configured to provide a second signal 02, for example an estimate of an FFR value, and a third regression branch 116 configured to provide a third signal 03, for example an estimate of an iFR value.
[0076] For example, the classification stage may take into account the FFR and iFR values, in particular the comparison of the estimated values with threshold values (known in themselves). For example, the threshold values may be 0.8 for the FFR value and 0.89 for the iFR value, as discussed in the following documents:
[0077] F.-J. Neumann, M. Sousa-Uva, A. Ahlsson, F. Alfonso, and A. P. e. a. Banning, "2018 ESC / EACTS Guidelines on myocardial revascularization," EHJ, vol. 40, 2018;
[0078] P. A. Tonino, B. De Bruyne, N. H. Pijls, U. Siebert, and F. e. a. Ikeno, "Fractional flow reserve versus angiography for guiding percutaneous coronary intervention," New England Journal of Medicine, vol. 360, 2009;
[0079] B. De Bruyne, N. H. Pijls, B. Kalesan, E. Barbato, and P. A. e. to. Tonino, "Fractional flow reserve-guided PCI versus medical therapy in stable coronary disease, " New England Journal of Medicine, vol. 367, 2012;
[0080] S. Baumann, L. Chandra, E. Skarga, M. Renker, M. Borggrefe, I. Akin, and D. Lossnitzer, "Instantaneous wave-free ratio (ifr®) to determine hemodynamically significant coronary stenosis: a comprehensive review," World Journal of Cardiology, vol. 10, 2018. Without restrictions on any specific theoretical model, the processing network topology 10 as exemplified in Figure 1 was developed based on the following observations by the Inventors: the CNN processing stages exhibit a strong inductive bias to learn local spatiotemporal features via 3D convolutions and are suitable for extracting local visual features, and the transformer stages are specialized for finding global and / or local correlations between regions of the input data, particularly those where a self-attention mechanism is applied.
[0081] Without restrictions on any specific theoretical model, the Inventors observed that decoupling the spatial and temporal axes reflects the phenomenon that, in addition to the anatomical aspect in a static scenario, blood flow over time also facilitates the evaluation of the state of a pathology by reducing the computational load.
[0082] For example, the processing exploits the temporal dynamics of blood flow (e.g., a rapid variation of contrast intensity) to identify where the stenosis is located, as a function of the fluid dynamic evolution.
[0083] In one or more embodiments, a transformer stage is configured to rearrange input tokens (i.e., representation vectors extracted from features), accessing ways to analyze the dynamic blood flow.
[0084] One or more embodiments exploit the observation over time of the propagation of the contrast agent through the coronary arteries, as visible in dynamic angiographic sequences, in order to estimate - for example, through pattern recognition mechanisms - the presence and / or severity of a stenosis, associated with a drop in blood pressure.
[0085] In one or more embodiments, the use of spatiotemporal attention layers facilitates: capturing a deformation induced by the motion of the vessel walls; monitoring intra-sequential changes related to blood flow and propagation of the contrast agent; coding physiological pressure gradients in a learned representation.
[0086] As exemplified in Figures 2 and 3, the input data matrix Min comprises a set of frames Min (TO) having width WO and height HO collected at a first time kO and a second image frame Min(Tl) having width W1 and height Hl collected at a second time T1 subsequent to TO, e.g. WO=W1, HO=H1, and Min= (Min (kO); Min(kl)).
[0087] In each Min(TO), Min(Tl) image, a respective stenosis LOO, Lil can be identified. Such identification is initially provided by expert physicians in order to create a set of training frames for the neural network CNN 100, as discussed below. For example, in order to create a set of training frames, the region of interest comprising the stenosis LOO, Lil can be analyzed based on predefined criteria or metadata associated with the angiographic sequences. Such metadata may include, for example, medical information or classifications assigned by the medical staff to the sequences, in a way known per se.
[0088] One or more embodiments may use an explainability procedure to detect which parts of the input signal / datum activate the model the most (corresponding to where the model focus on the most attention), so as to facilitate the recognition of the type of stenosis LOO, Lil. For example, a explainability process per se known as Grad- CAM discussed in the paper Selvaraju, Ramprasaath R., et al.: "Grad-CAM: Visual Explanations from Deep Networks via Gradient-Based Localization", arXiv [Cs.CV], 2016, abs / 1610.02391 can be used in one or more embodiments, preferably for convolutional networks. This type of explainability process is purely exemplary and not limiting, as embodiments could use different explainability processes, for example for transformer networks. For example, an explainability process per se known as attention rollout discussed in the paper Abnar, Samira and Willem Zuidema: "Quantifying Attention Flow in Transformers", Annual Meeting of the Association for Computational Linguistics (2020), arXiv, abs / 2005.00928 can be used in one or more embodiments.
[0089] For example, during the inference phase the explainability process is configured to provide a heat map that emerges a region of interest from a set of images Min(TO), Min(Tl); Min'(Tl) in a temporal sequence of images (Min; Min'), where at least one image in the set of images includes a region of interest LOO, Lil indicative of a stenosis, preferably a coronary stenosis.
[0090] Figures 2 and 3 are an example of angiographic projections related to the left coronary artery (consisting of the common trunk, anterior interventricular artery and circumflex artery) of a single patient. For example, Figures 2 and 3 represent a right caudal (left) and left cranial (right) projection. As identified by the areas labeled LOO, Lil in Figures 2 and 3, the patient has an intermediate- grade (in terms of angiographic definition) stenosis (occlusion or narrowing) of the anterior interventricular coronary artery, in its proximal segment (highlighted with a dotted rectangle in the figures), while the remaining coronary segments are free of significant stenosis.
[0091] According to the recommendations of international guidelines, the hemodynamic impact of the stenosis found in the patient whose angiographic views are exemplified in Figures 2 and 3 should be assessed by assessing FFR or iFR via an invasive procedure. Using the method as exemplified here, it is possible to estimate the hemodynamic impact of the stenosis exemplified in Figures 2 and 3 by processing the collected angiographic images, thereby reducing the invasiveness of the measurement .
[0092] The procedure was tested with images acquired via different angiography systems.
[0093] Table I provided at the end of the description represents by way of example the features of some of the X-ray image acquisition systems used.
[0094] One or more embodiments operate on sequences of angiographic images that can be indicated as (2D + t) type, i.e. two-dimensional images acquired over time and containing dynamic and temporally evolving representations of the same region of a blood vessel, preferably coronary. For example, the temporal dimension of the images provides an indication of the variations in flow induced by the presence of stenosis, allowing the model to capture the haemodynamic patterns characteristic of the narrowing.
[0095] As exemplified in Figure 1, the input video Min comprising a set of frame having size TxHxW is provided to the 3D convolutional network processing stage 100 which subsequently provides a set of features Sf having sizes f_h x f_w x f_c smaller than the size of the input video, due to the dimensionality reduction effect of 3D convolutional networks. In one or more embodiments, a 3D convolutional network known as ResNet3D may be suitable for use in the input CNN stage 100.
[0096] Alternatively, networks known as ResNet2+lD or ResNet-MC (Mixed Convolution) may be suitable for use in the input CNN stage 100 in one or more variant embodiments . The operation of the ResNet3D network (and other ResNet networks) is known, for example, from the paper D. Tran, et al.: "A Closer Look at Spatiotemporal Convolutions for Action Recognition", IEEE Conference on Computer Vision and Pattern Recognition, CVPR, 2018, abs / 1711.11248, arXiV.
[0097] In one or more embodiments, the first layer of convolutional kernels of a ResNet(3D, 2+1D, MC) network is "flattened" by computing the average value of the values of all the input channels f_c in order to adapt the network to the processing of greyscale images.
[0098] In one or more embodiments, the input CNN processing stage 100 may be (pre)trained on a training data set known as Kinetics-400 and discussed in the paper J. Carreira and A. Zisserman: "Quo vadis, action recognition? a new model and the kinetics data set", in CVPR, 2017, abs / 1705.07750, arXiV.
[0099] As exemplified in Figure 1, the features in the set of 3D features Sf extracted by the 3D convolutional network stage 100 are fed to the transformer processing stages 102, 104.
[0100] The Inventors have observed that the use of transformer processing facilitates overcoming limitations due to the spatial deformation of the region being imaged and the periodicity of the cardiac cycle that would otherwise make the extraction of perturbative dynamics complex.
[0101] In order to improve computational efficiency, the Inventors have observed that it may be convenient to create a chain in which the transformer processing is distributed between two modules: a first spatial transformer processing module 102 followed by a second temporal transformer processing module 104. Preferably, both modules 102, 103 comprise layers in which (multi head) self-attention processing 103a,103b, 103n, 10-5 comprising scaled scalar products is applied. A multi- head self-attention method suitable for use in one or more embodiments is discussed for example in the paper Vaswani, Ashish et al.: "Attention is All you Need." Neural Information Processing Systems (2017), arXiV, abs / 1706.03762 .
[0102] Given the set of features Sf including for example a plurality of features Fl, F2, these can be distributed on rows of matrices having size f_h*f_w x F so that the spatial transformer layers can produce a new set of transformed features Sff that can be expressed as follows :
[0103] Sff = LN (MLP (Sf')+Sf') where
[0104] LN is a layer normalization operation known per se from the paper Ba, Jimmy et al.: "Layer Normalization", ArXiv, abs / 1607.06450 (2016);
[0105] MLP is an operator representing two layers of multilayer perceptron processing with input and output stages having size equal to F and hidden layers equipped with a number d of neurons configured to apply a rectification linear unit (ReLU), in a per se known way;
[0106] Sf'= LN (MA (Sf)+Sf) with MA representing the multi- head self-attention operator, defined as follows;
[0107] In each layer 103a,103b, 103n, the self-attention function A computes query values, keys, and vectors for each feature in the set of features Sf by computing respective projections Q, K, and V into a vector space having size Nxd, with d equal to the size of the features. For example, as a result of the projection operation, the projected matrices WQ, WK, and WV are obtained which are learnable by layers 103a, 103b, 103c. The multi-head self-attention operator MA is configured to provide, based on a set of h projected matrices fWiQ, W1K, WiV} for 1=1...h, outputs where such projected matrices are concatenated and linearly projected so as to reduce their dimensionality from d to F / h.
[0108] For example, the multi-head self-attention operator MA receives as input a set of "heads" (e.g., a set of a number h of heads), each defined by the projection matrices {WiQ, WiK, WiV} for i=l...h, and then each head computes attention (with dimension d = F / h), producing an output (of dimension d). Again in the example considered, the outputs (e.g., in a number equal to h) are then concatenated into a vector (of dimension F where F = h -d) and then linearly projected through a matrix (e.g., WO e IRFxF).
[0109] As exemplified in Figure 1, the multi-head self- attention stages 103a,103b,103c in the spatial transformer processing stage 102 receive a set of convolutional features Fl, F2 and process each frame autonomously, so as to obtain new transformed features Sf' which are then concatenated as a new space of features Sff having size txhxwxF.
[0110] As exemplified in Figure 1, the temporal transformer processing stage 104 receives the set of transformed features Sff from the spatial transformer processing stage 102 and applies further multi-head self-attention processing 105 to it. As a result of such processing, a set of ordered temporal features can be obtained, which capture the dynamic dependencies between different instants of the sequence.
[0111] For example, the further multi-head self-attention processing 105 is applied between frame-level concatenated feature vectors whereby the processing comprises : concatenating all features within each frame of the transformed features Sff, resulting in a tensor H' E R^timex / wf (where Ntime = t); processing with self-attention the obtained tensor, resulting in an output tensor Ht of size txhxwxf.
[0112] Preferably, both the spatial and temporal transformer stages 102 each comprise two layers and a number h=4 of attention heads per layer (for a total of 8 heads).
[0113] As exemplified herein, a multi-head spatiotemporal self-attention stage 102, 104 may be configured to perform operations comprising:
[0114] - given a number h of heads of the stage, computing for each head the value of a self-attention operator SA which may be expressed as: hi=SA=softmax where
[0115] Q, K and G are queries, keys and input values, respectively; dk is the size of the keys;
[0116] - concatenating the outputs of each head, obtaining a set of concatenated self-attentions that can be expressed as: l|hi,h2, hh|| computing a product between the concatenated outputs of the different heads with a matrix of transformed output weights V, product that can be expressed as:
[0117] MHA(Q, K, G)=||h1,h2,•••,hh||xV
[0118] As exemplified in Figure 1, the temporal transformer stage 104 is coupled to the output convolutional stage 106, for example comprising two layers configured to apply CNN processing with stride having value 2, providing an output vector V.
[0119] The Inventors have observed that the use of the output convolutional network 106 facilitates a further reduction of dimensionality by a scalar factor (e.g., equal to four, so that the tensor Ht is reduced to a vector V having size fthw / 64).
[0120] The multi-branch processing stage 110 receives the vector V and provides it as input to a set (e.g., three) of processing branches 112, 114, 116.
[0121] For example, each of the processing branches in the set of processing branches 112, 114, 116 may comprise a same topology, exemplified in Figure 4.
[0122] As exemplified in Figure 4, an i-th branch (e.g., the regression branch 114) in the set of branches 110 is configured to receive the vector V and to provide an i- th output indicator Oi (e.g., the indicator 02 indicative of the FFR), wherein the i-th branch comprises a topology comprising a chain of artificial neural network processing stages 40 (e.g., multi-layer perceptron, MLP with two hidden layers), wherein each processing stage 40 comprises: block 402: a fully connected (FC) layer configured to learn correlations with the augmented clinical data, as discussed below; block 404: a batch normalization (BC) stage; block 406: a dropout stage configured to exclude some neurons during training to introduce noise into the processing, as discussed in the paper Nitish Srivastava, et al.: "Dropout: a simple way to prevent neural networks from overfitting," 2014, J. Mach. Learn. Res. 15, 1 (January 2014), 1929-1958; block 408: an activation function stage, preferably of the rectification linear unit type.
[0123] As exemplified in Figure 4, the hidden layers of each layer of the MLP topology halve the number of input neurons or perceptrons (e.g., from 512 to 256 perceptrons) so that the output stage provides a single scalar value (e.g., from 256 values to 1).
[0124] As exemplified in Figure 1, each branch 112, 114, 116 provides a scalar value as output: a first indicator 01 with values constrained to the interval 0 to 1 for the first branch 112, indicative of a binary class of significance of the stenosis; a second indicator 02 indicative of an estimate of the FFR value, and a third indicator 03 indicative of an estimate of the iFR value.
[0125] To differentiate the branches 112, 114, 116 from each other in the case they share the same topology exemplified in Figure 4, different learning or "loss" functions are defined.
[0126] In particular, given the "ground-truth" values yc, yFFR and yiFR, the following training objectives are defined:
[0127] L = LiBCE+LFFR+LiFR =
[0128] = -yclog (01)- (l-yc)*log (1-01)+A*|yFFR-02|+X*|y±FR- 03| where
[0129] LBCE is the loss function of the binary classification,
[0130] LEER and LIFR are the loss functions of the linear regression;
[0131] X is a weight value (e.g., equal to 10) between the regression and classification terms.
[0132] As exemplified here, the convolutional networks and the transformer are trained to minimize a cross entropy loss and a LI loss function also called mean-absolute- error (MAE).
[0133] As exemplified in Figures 1 and 5, each branch of the multi-branch stage 110 may be trained using a different set of training data and with a different objective function to be achieved. For example: the first branch 112 is trained with a first training data set TDBa comprising a set of angiographic images for which the FFR and / or iFR values are above or below the operability threshold and therefore belong to two classes (e.g., clinically significant stenoses) and is configured to minimize an objective or loss function called "binary cross entropy loss" which can be expressed as: [tjlogh’^pj +(l-tj)logh’i(l-pj)] so as to provide the output indicator 01 with a value corresponding to one of the two classes; the second branch 114 is trained with a first training data set TDBa comprising a set of angiographic images belonging to patients with different FFR values and is configured to minimize the value of an objective or loss function LI which can be the mean value or the reduction sum in the case of batch (to overcome the limitation otherwise present with regard to the weak- labelling mechanism, i.e. in the absence of one of the values of ifr or ffr, as discussed below), which can be expressed as
[0134] LILossFunction in order to provide the second output indicator 02 with a real value of FFR estimate; the second branch 114 is trained with a first training data set TDBa comprising a set of angiographic images belonging to patients with different iFR values and is configured to minimize the value of an objective or loss function LI which can be the value average or the reduction sum in the case of batches (to overcome the limitation otherwise present with regard to the weak- labelling mechanism, as discussed below), so as to provide the second output indicator 02 with a real value of iFR estimate.
[0135] In one or more embodiments it is used an objective function of technique of aggregation of the loss for batch according to the sum of the individual values that facilitates the weak labelling, i.e. the possibility of using partial, noisy or imprecise labels / annotations that can still provide useful information for training models. For example, in the training phase, for 80% of the total training data set only one of the FFR and iFR values can be present.
[0136] One or more embodiments facilitate the contrast of the phenomenon whereby a missing (therefore null) value in a batch of tests is averaged with the other acceptable values existing in the same batch, disturbing the pattern recognition learning procedure of the corresponding processing branch 112, 114.
[0137] As exemplified in Figure 5, in an alternative embodiment of the method 50 it may comprise the following further operations: providing a plurality of image sequences Min, Min' collected by means of the medical imaging apparatus A, preferably an X-ray apparatus for coronary angiography; and / or providing a plurality of artificial neural network processing chains 10a, 10b, e.g. a first processing chain 10a and a second processing chain 10b, wherein each processing chain 10 comprises a subset of the stages illustrated with reference to the method 10 exemplified in Figure 1.
[0138] As exemplified in Figure 5, the method may further comprise a pre-processing operation 502a, 502b of the images in the plurality of image sequences Min, Min'. Optionally, the pre-processing stage 502 may be configured to provide data amplification processing, preferably comprising the inclusion of additional patient clinical data among the input data. For example, if you opted for amplification pre-processing, such additional clinical data would include: patient age, left ventricular ejection fraction (LVEF), AC index, ST segment elevation myocardial infarction (STEMI), non-ST- segment elevation myocardial infarction (NSTEMI), stable or unstable angina (UA), stress test results. In the optional scenario you are considering, for example, such additional clinical information could be divided into intervals and incorporated as a multi-dimensional vector comprising binary values (e.g., 1 for presence, -1 for absence).
[0139] In a further or complementary scenario, pre- processing stages 502a, 502b (which may also be present in the embodiment exemplified in Figure 1) may comprise, preferably during a training phase, the operations of: applying a normalization, for example in uniform scale, to the sequence of angiographic images Min, Min' (for example, normalizing the values so that they are between 0 and 1 with a mean value of 0 and a standard deviation of 1); applying a sequence of rotation and shift operations with angles randomly extracted from a random distribution, so as to improve the generalization capacity of the artificial neural network processing, applying a temporal shift operation, for example by identifying 60 frames in the vicinity of a key frame that is considered the most representative within the sequence.
[0140] As exemplified in Figure 5, the first processing chain 10a and the second processing chain 10b comprise a first stage 100a and a second stage 100b (e.g., each comprising a ResNetl8 type network) configured to extract respective features (e.g., 2D+T)) from each frame Min(TO), Min'(Tl) of each sequence in the plurality of input image sequences Min, Min' so as to apply one or more parallel self-attention processing operations 102a,104a;102b,104b followed by one or more further output convolutional network operations 106a, 106b, finally comprising concatenating the respective outputs Ht, Ht' in a further concatenation stage 508 before providing them to the multi-branch stage 110.
[0141] Table II provided at the end of the description represents by way of example a summary of the possible combinations of signals in input Min; Min' and processing chains 10; 10a, 10b in terms of processing performance. As exemplified in Table II, the performance of the procedure in the different variants can be represented in terms of accuracy, area under the region of interest ("Area under the ROC curve", AUG), sensitivity and specificity.
[0142] As mentioned, the output signals 01, 02, 03 processed by the plurality of output processing branches 112, 114, 116 can be provided to user circuits, such as a video screen equipped on board the apparatus A or external to the apparatus itself.
[0143] In an exemplary scenario, the method 10, 50 can be implemented on a processing device installed locally in the hospital, on a centralized computer system with a graphic processing unit or remotely on a "cloud" processing infrastructure. For example, via a web interface with an Internet browser, the physicians can upload a video of a coronary angiography or an angiogram and display the values of the output indicators 01, 02, 03 on a video screen.
[0144] In one or more embodiments where there may be limits to the available computational power, the multi-branch stage 110 may also comprise only the binary classification stage 112 or only one of the FFR / iFR point estimation stages or even subsets of two stages in any possible combination of those illustrated.
[0145] As appreciated by the person skilled in the art, the different processing stages 100, 102, 104, 106; 10a, 10b, 40 and the artificial neural network of the method 10, 50 are trained to perform pattern recognition in a dedicated training phase, for example supervised training .
[0146] As exemplified in Figures 1 and 5, training databases TrDB, TDBa, TDBb, TDBc are provided to train the networks, including angiographic images containing stenoses labeled by medical personnel and for which the measured values (via conventional technique) of FFR and iFR are available.
[0147] For collecting the set of training images TrDB, TDBa, TDBb, TDBc a coronary pressure sensing wire (itself known by the trade name Philips Volcano) equipped with a pressure sensor coupled to the catheter tip can be used to measure FFR and iFR once advanced and equalized to the aortic pressure (for example, the catheter is advanced at least 3 diameters of the vessel distal to the stenosis to provide a record of the distal coronary and aortic pressure).
[0148] For example, the iFR of each patient can be computed using an automated process applied to time-aligned pressure traces in the wave-free diastole period over a minimum of five beats prior to adenosine administration. Maximal coronary hyperemia can then be induced with continuous intravenous infusion of adenosine (standard rate, 140 pg / kg / min for 2 minutes) or intracoronary bolus injections (recommended dose, 100-200 micrograms).
[0149] For example, the FFR of each patient can be computed as the ratio of mean distal coronary pressure to mean aortic pressure during stable hyperemia (the lowest mean of 3 consecutive beats).
[0150] For example, as a training database TrDB, TDBa, TDBb, TDBc, the same database can be used, including coronary angiographies collected between 2020 and 2022 from 495 patients hospitalized in 5 hospitals, with the diagnosis of Chronic Coronary Syndrome (CCS) and acute coronary syndrome (ACS). In the total of 495 diseased patients, the presence of 376 men and 119 women was found, with the computed mean age equal to 67.719.23 years.
[0151] As mentioned, each angiography in the training database TrDB, TDBa, TDBb, TDBc has associated at least one medical opinion (preferably two) from the staff of each hospital, where the major stenosis in each image is labeled as operable for a value of FFR and / or iFR measured (via conventional invasive procedure) below the clinically significant threshold (e.g., FFR<0.8; iFR<0.89).
[0152] In the sample database, 129 patients (26.1% of the total) are classified as candidates for surgery (indicated as "positive"), while the remaining patients are classified as not candidates for surgery (indicated as "negative").
[0153] For example, the applicability of FFR / iFR assessment during coronary angiography facilitates the authorization of a percutaneous coronary intervention in case of a positive result of pattern recognition tasks.
[0154] In the at least one training database TrDB, angiographies collected with different types of apparatus A can be converged, which can vary from hospital to hospital. Consequently, the size of the images WO, HO, Wl, Hl in the image sequences Min, Min' can vary, for example from W0=H0=512 to W0=H0=1024 as well as the duration of the videos, that can include from T=7.5 fps (frames per second) to T'=60 fps, can vary.
[0155] As mentioned, these differences can be made homogeneous through pre-processing 502a, 502b: for example, it is possible to rescale the image samples in order to provide a homogeneous size (for example, H0=W0=Hl=Wl= 256 and T=T'=15 fps through sub-sampling). For example, a duration T=T' of 60 frames can be defined by selecting (i.e., analyzing or processing) 4 seconds (the conventional average duration of an angiogram) from each collected sequence Min, Min'. For example, the procedure includes sub-sampling or padding operations of the first / last images. In this way each sequence Min, Min' is analyzed in its entirety, with processing of a homogeneous time segment for the entire spatiotemporal pipeline. Preferably, the selected (and therefore analyzed) frames are centered around the region of interest (i.e. a main frame or "key-frame", for example identified through explainability mechanisms such as the heat map, as discussed above) that presents the highest level of contrast agent perfusion. Alternatively, a single view of the patient can be used, both for training and inference, setting aside the remaining collected images.
[0156] In one embodiment, the training phase was performed by minimizing a combination of objective functions (e.g., binary cross-entropy loss L_e reduction sum LI) by applying a mini-batch gradient descent algorithm. For example, one can use an optimization procedure known as AdamW and discussed in Loshchilov, Ilya and Frank Hutter: "Decoupled Weight Decay Regularization", arXiv, abs / 1711 .05101. For example, the learning rate is le-5 with a batch size of 8 and a duration of 300 epochs.
[0157] The Inventors noted that such a sizing of the phase counteracts the overfitting on the training data set, which could otherwise cause a loss of generalization for the validation data set.
[0158] To validate the accuracy of the procedure 10, 50, a 5-fold balanced nested cross-validation strategy can be used (known in itself) that uses 60% data of the database TrDB during the training phase, 20% during the validation phase, i.e. choice of the model considered best and the remaining 20% as inference data (i.e. test data, to have a verification on real data and not correlated to the training data). This subdivision, for example, takes into account the original proportion of labels applied to the data in the TrDB data set.
[0159] As exemplified in Figure 6, the training phase can include a data filtering operation in the TrDB training data set, for example for an ablation study and / or for use of the neural network. For example, among 763 cases reported as having intermediate coronary stenosis and initially included in the cohort, a filtered subset El (e.g., 495 cases) passed a filtering while others were excluded based on the following parameters (exemplified in portion b) of Figure 6):
[0160] - some cases E2 (e.g., 98 in the case under study) presenting iFR and FFR assessments with accuracy below a filtering threshold; or
[0161] - some cases E3 (e.g., 170 in the exemplary scenario considered) presenting angiograms in a format not compatible with digital processing.
[0162] For example, a training database TrDB is provided (used to train all stages, so TrDB=TDBa=TDBb=TDBc) containing: measured values of FFR for a subset of 205 cases (corresponding to 41.4% of the total 495 used), measured values of iFR for a subset of 176 cases (corresponding to 35.5% of the total 495 used); measured values of both FFR and iFR for a subset of 114 cases (corresponding to 23.1% of the total 495 used).
[0163] Figure 7 is an exemplary histogram of a distribution of stenosis diameters in the cohort of enrolled patients.
[0164] Table ITT provided at the end of the description represents by way of example the reference and procedural characteristics of the cases for which images were collected from an exemplary set of TrDB training images for the ANN processing stages according to this disclosure. As exemplified in Table III, STEMI is an acronym that indicates ST-segment elevation myocardial infarction while NSTEMI indicates non ST-segment elevation myocardial infarction.
[0165] As exemplified in Table TV, the performance of different configurations of the processing chain of the procedure according to this disclosure can be represented by PPV (which indicates positive estimated values) and NPV (which indicates negative estimated values indicated at 95% confidence intervals).
[0166] Table TV provided at the end of the description represents by way of example the performance of the processing procedure in some exemplary scenarios. In particular: in a first scenario, indicated in Table IV as "STARFLOW", the analysis stage 110 includes the classification branch 112; in a second scenario, indicated in Table IV as "STARFLOW-FFR", the analysis stage 110 includes the classification branch 112 and a regression branch 114; in a third scenario, indicated in the table as "STARFLOW-iFR", the analysis stage 110 includes the classification branch 112 and the two regression branches 114, 116; in a fourth scenario, referred to in Table IV as "STARFLOW-Clinical", the analysis stage 110 comprises at least one pre-processing stage configured to perform data-augmentation processing in addition to the classification branch 112, the classification branch 112 and the two regression branches 114, 116.
[0167] Table V provided at the end of the description represents by way of example the performance of the regression branches 114, 116 (e.g., in the absence of the classification stage 112) in estimating the FFR and iFR values during the validation phase of the model trained with the TrDB training data. As exemplified in Table V, the performance of the regression branch procedure according to this disclosure can be expressed in terms of mean-square error (MSE) and mean absolute error (MAE).
[0168] In one or more embodiments, processing with the method 10, 50 discussed in this disclosure may take a total time of approximately 62.56 milliseconds (which however depends entirely on the characteristics of the hardware used) to process a input video Min comprising 60 image frames.
[0169] Table VI provided at the end of the description represents by way of example performance improvements in the use of the method according to this disclosure compared to existing methods. As exemplified in Table VI, the performances of the method according to this disclosure, briefly called STARFLOW, are compared in benchmarks .
[0170] The finding indicated in Table VI as Direct quantification is discussed, for example, in the paper D. Zhang, et al.: "Direct Quantification of Coronary Artery Stenosis Through Hierarchical Attentive Multi- View Learning", IEEE Transactions on Medical Imaging, vol. 39, no. 12, pp. 4322-4334, Dec. 2020.
[0171] The finding indicated in Table VI as Direct quantification with attention is discussed, for example, in the paper W. Xue, et al.: "Full left ventricle quantification via deep multitask relationships learning," Med. Image Anal., vol. 43, pp. 54-65, Jan. 2018.
[0172] The finding indicated in Table VI as full left- ventricle quantification is discussed, for example, in the paper Morris PD, Ryan D., et al.: "Virtual fractional flow reserve from coronary angiography: modeling the significance of coronary lesions: results from the VIRTU-1 (VIRTUal Fractional Flow Reserve from Coronary Angiography) study," JACC Cardiovasc Interv. 2013 Feb;6(2):149-57 .
[0173] Figure 8 shows a plot where, for an AUG of 0.93 for all cases in the TrDB training data set, the false positive rate is plotted on the x-axis and the valid positive rate is plotted on the y-axis for the classification of the AUG of the training images.
[0174] Figure 9 shows a plot where, for an AUG of 0.81 for all cases in the TrDB training data set, the recall rate R is plotted on the x-axis and the precision rate P is plotted on the y-axis.
[0175] The precision rate P can be defined as the number of valid positives TP divided by the sum of valid positives TP and false positives FP.
[0176] The recall rate can be defined as the number of valid positives TP divided by the sum of valid positives TP and false negatives FN.
[0177] Performance analyses were also performed using subsets of the training data.
[0178] For example: portion a) of Figure 10 includes an exemplary plot where, for an AUG of 0.79 for the 205 cases with measured values of the FFR alone in the TrDB training data set, the rate of false positives is represented on the abscissa and the rate of valid positives on the ordinate with regard to the classification of the AUG of the training images; portion b) of Figure 10 includes an exemplary plot where, for an AUG of 0.83 for the 205 cases with measured values of FFR only in the TrDB training data set, the recall rate R is represented on the abscissas while the precision rate P is represented on the ordinates.
[0179] For example: portion a) of Figure 11 includes an exemplary plot where, for an AUG of 0.98 for the 176 cases with measured values of iFR only in the TrDB training data set, the false positive rate is represented on the abscissas and the valid positive rate is represented on the ordinates with respect to the AUG classification of the training images; portion b) of Figure 11 includes an exemplary plot in which, for an AUG of 0.77 for the 176 cases with measured values of the FFR only in the TrDB training data set, the recall rate R is represented on the abscissas while the precision rate P is represented on the ordinates.
[0180] Figures 8, 10 and 12 also show the diagonal of the quadrant that acts as a reference line for the trend of the graph in the case of assigning labels according to a "random choice" procedure (in which therefore TPR = FPR).
[0181] As exemplified in Figures 1 and 5, a method includes: receiving at least one temporal sequence of digital images RC acquired through a medical imaging device A; selecting a set of images in the temporal sequence of images (e.g., by analyzing or processing a segment of a few seconds duration within a minute of sequence), wherein at least one image in the selected set of images includes a region of interest indicative of a stenosis, preferably a coronary stenosis; processing the set of images via an artificial neural network, ANN, processing chain.
[0182] For example, the ANN processing chain comprises: an input ANN processing stage configured to apply convolutional neural network, CNN, processing to the image set, resulting in a set of features; a set of transformer ANN processing layers coupled to the input ANN processing stage to receive the set of features therefrom and configured to apply at least one self-attention ANN processing layer to the features in the set of features, resulting in a set of transformed features; an output ANN processing stage coupled to the set of transformer ANN processing layers to receive the set of transformed features and configured to apply further CNN processing to the transformed features in the set of transformed features, resulting in an output vector; an analysis stage coupled to the output ANN processing stage for receiving the output vector, the analysis stage configured to apply at least one multilayer perceptron, MLP, processing to the output vector, resulting in a set of signals indicative of a classification and / or regression of the functionality of the stenosis.
[0183] As exemplified in figures 1 and 5, the analysis stage comprises a classification stage configured to produce a classification indicator of the stenosis, and at least one regression stage configured to provide an estimate of an FFR and / or iFR value.
[0184] As exemplified in figures 1 and 5, the classification stage is configured to perform a comparison between the estimate of the FFR and / or iFR value and at least one reference value, and produce a classification signal indicative of whether the estimate of the FFR and / or iFR value meets or fails to meet said reference value.
[0185] As exemplified in figures 1 and 5, the digital image temporal sequence (Min; Min') comprises images acquired via an X-ray coronary angiography apparatus, and / or the digital image temporal sequence (Min; Min') comprises greyscale images.
[0186] As exemplified in figures 1 and 5, the set of transformer ANN processing layers comprises: a first spatial transformer ANN processing layer configured to apply spatial transformer processing, and a second temporal transformer ANN processing layer configured to apply temporal transformer processing.
[0187] For example, each transformer ANN processing layer in the set of transformer ANN processing layers is configured to apply at least one (factorized) multi-head self-attention ANN operator to features in the set of features, resulting in a respective set of transformed features.
[0188] For example, each transformer ANN processing layer in the set of transformer ANN processing layers comprises two layers and four heads for each layer.
[0189] As exemplified in figures 1 and 5, the output CNN processing comprises: applying a set of rectification linear unit, ReLU, activation functions to the transformed features in the set of transformed features (Ht, Ht'), and concatenating the signals produced as a result of applying said set of ReLu activation functions, resulting in the output vector (V).
[0190] As exemplified in figures 1 and 5, the method further comprises: providing a set of training images comprising digital training images acquired by at least one medical imaging apparatus, the digital images comprising respective regions of interest indicative of at least one stenosis, preferably a coronary stenosis, wherein each digital image in the training image set has associated at least one measured FFR and / or iFR value; training the analysis stage to produce the set of signals as a function of the at least one measured FFR and / or iFR value associated with each training image in the training image set.
[0191] As exemplified herein, a system for processing at least one temporal sequence of digital images acquired by an imaging apparatus comprises processing circuits configured to operate in accordance with the method exemplified in figures 1 and 5.
[0192] For example, at least one image in the at least one temporal sequence of digital images acquired in the system comprises a region of interest indicative of a stenosis, preferably a coronary stenosis.
[0193] As exemplified herein, an imaging apparatus comprises a system according to this disclosure, preferably in combination with an X-ray source and an array of X-ray detectors to provide digital images of stenosis to be processed / analyzed.
[0194] As exemplified herein, a computer product loadable into the memory of at least one computer comprises portions of software code for executing the steps of the method exemplified in figures 1 and 5 when the product is executed on at least one computer (or even microprocessor or embedded system or edge computing unit).
[0195] Notwithstanding the underlying principles, the details and embodiments may vary, even appreciably, from what has been described, purely by way of example, without departing from the extent of protection. The extent of protection is defined by the attached claims.
[0196]
Claims
CLAIMS1. A method (10; 50), comprising: receiving at least one temporal sequence of digital images (Min; Min') acquired (RC) by means of an imaging apparatus (A); analyzing a set of images (Min(TO), Min(Tl); Min' (Tl)) in the temporal sequence of images (Min; Min'), wherein at least one image in the analyzed set of images (Min (TO), Min(Tl); Min'(Tl)) comprises a region of interest (LOO, Lil) indicative of a stenosis, preferably a coronary stenosis; processing the set of images (Min(TO), Min(Tl); Min' (Tl)) via an artificial neural network, ANN, processing chain (10; 10a, 10b; 110), wherein the ANN processing chain (10; 10a, 10b; 110) comprises : an input ANN processing stage (100) configured to apply a convolutional neural network, CNN, processing to the images in the set of images (Min(TO), Min(Tl); Min' (Tl)), producing a set of features (Sf) as a result; a set of transformer ANN processing layers (102, 104) coupled to the input ANN processing stage (100) to receive the set of features (Sf) therefrom and configured to apply at least one self-attention ANN processing layer (102a, 104a; 102b, 104b) to the features in the set of features (Sf), producing a set of transformed features (Sff, Ht; Ht') as a result; an output ANN processing stage (106) coupled to the set of transformer ANN processing layers (102, 104) to receive the set of transformed features (Sff, Ht; Ht') and configured to apply further CNN processing to the transformed features in the set of transformed features (Sff, Ht; Ht'), producing an output vector (V) as a result;an analysis stage (110) coupled to the output ANN processing stage (106) for receiving the output vector (V), the analysis stage (110) being configured to apply at least one multilayer perceptron, MLP, processing to the output vector (V), producing as a result a set of signals (01; 02, 03) indicative of a classification (112) and / or regression (114; 116) of the functionality of the stenosis.
2. The method (10; 50) according to claim 1, wherein the analysis stage (110) comprises: a classification stage (112) configured to produce a classification indicator (01) of the stenosis, and / or at least one regression stage (114; 116) configured to provide the estimate value of fractional flow reserve, FFR and / or instantaneous flow reserve, iFR.
3. The method (10; 50) according to claim 2, wherein the classification stage (112) is configured to: perform a comparison between the FFR and / or iFR value estimate and at least one reference value, and provide (40) a classification signal (01) indicative of whether the FFR and / or iFR value estimate reaches or fails to reach said at least one reference value.
4. The method (10; 50) according to any one of the previous claims, wherein: the digital image temporal sequence (Min; Min') comprises images acquired via an X-ray angiocardiography apparatus, and / or the temporal sequence of digital images (Min; Min') comprises greyscale images.
5. The method (10; 50) according to any one of theprevious claims, wherein the set of transformer ANN processing layers (102, 104) comprises: a first spatial transformer ANN processing layer (102) configured to apply spatial transformer processing, and a second temporal transformer ANN processing layer (104) configured to apply temporal transformer processing, wherein each transformer ANN processing layer in the set of transformer ANN processing layers (102, 104) is configured to apply at least one ANN multi-head self- attention operator (102a, 104a; 102b, 104b) to the features in the set of features (Sf), producing the respective set of transformed features (Sff, Ht; Ht'), preferably wherein each transformer ANN processing layer in the set of transformer ANN processing layers (102, 104) comprises two layers and an equal number of four heads per layer.
6. The method (10; 50) according to any one of the previous claims, wherein the output ANN processing stage (106a, 106b) comprises: applying a set of rectification linear unit, ReLU, activation functions, to the transformed features in the set of transformed features (Ht, Ht'), and concatenating the signals produced as a result of applying said set of ReLu activation functions, producing in the output vector (V) as a result.
7. The method (10; 50) according to any one of the previous claims, comprising: providing a set of training images (TrDB; TDBa, TDBb, TDBc) comprising digital training images acquired by at least one medical imaging apparatus (A), the digital training images comprising respective regions ofinterest (LOO, Lil) indicative of at least one stenosis, preferably a coronary stenosis, wherein each digital image in the set of training images (TrDB; TDBa, TDBb, TDBc) has associated therewith at least one measured value of FFR and / or iFR; training said analysis stage (110) to produce the set of signals (01; 02, 03) as a function of the at least one measured value of FFR and / or iFR associated with each training image in the set of training images (TrDB; TDBa, TDBb, TDBc).
8. A processing system configured to process at least one temporal sequence of digital images (Min; Min') acquired (RC) by an imaging apparatus (A), wherein at least one image in the at least one temporal sequence of digital images (Min; Min') comprises a region of interest (LOO, Lil) indicative of a stenosis, preferably a coronary stenosis, wherein the processing system comprises processing circuits configured to operate according to the method of any one of claims 1 to 7.
9. Medical imaging apparatus (A) comprising a processing system according to claim 8 coupled to an angiocardiography apparatus equipped with an X-ray source (R) and an array of X-ray detectors, the angiocardiography apparatus being configured to provide digital images (Min, Min') of stenosis (LOO, Lil) to be processed.
10. A computer product loadable into the memory of at least one processing device, preferably a microprocessor, and comprising portions of software code for performing the operations of the method of any one of claims 1 to 7 when the computer product is executedon at least one processing device.
Citation Information
Patent Citations
Method and System for Machine Learning Based Assessment of Fractional Flow Reserve
US20150112182A1
Machine learning spectral FFR-ct
US20200113449A1
Method, device and computer-readable medium for automatically classifying coronary lesion according to CAD-RADS classification by a deep neural network
US20220215541A1
Method, device and computer-readable medium for automatically detecting hemodynamically significant coronary stenosis
US20220215542A1
Automatic stenosis detection
WO2021117043A1