Device for carrying out an automated assessment of images obtained from mammography examinations by using transformer networks

The device addresses the challenges of automated mammography diagnosis by employing a novel Transformer neural network architecture for multi-view analysis, achieving enhanced classification accuracy and efficiency in detecting cancerous lesions.

WO2025120475A1PCT designated stage expired Publication Date: 2025-06-12HEALTH TRIAGE SRL +1
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
PCT/IB2024/062085
Authority / Receiving Office
WO · WO
Patent Type
Applications
Current Assignee / Owner
Priority Date
2023-12-05
Filing Date
2024-12-02
Publication Date
2025-06-12

AI Technical Summary

Technical Problem

Current systems for automated mammography diagnosis face challenges such as low lesion prevalence, high image resolution, and the need to combine multiple views, which complicates the use of deep neural networks for accurate and efficient analysis.

Method used

A device utilizing a novel Transformer neural network architecture, specifically designed for multi-view analysis of mammography images, which operates in conjunction with other neural network architectures to enhance classification accuracy and efficiency.

Benefits of technology

The device achieves improved performance in classifying mammography images by effectively combining local and global features, outperforming conventional convolutional neural networks and enabling accurate detection of cancerous lesions with reduced computational and memory requirements.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure IB2024062085_12062025_PF_FP_ABST
    Figure IB2024062085_12062025_PF_FP_ABST
Patent Text Reader

Abstract

A device for carrying out an automated assessment of images obtained from an equipment for carrying out mammography examinations, said images being analysed by using a transformer neural network, possibly combined with at least one further neural network.
Need to check novelty before this filing date? Find Prior Art

Description

[0001] DEVICE FOR CARRYING OUT AN AUTOMATED ASSESSMENT OF IMAGES

[0002] OBTAINED FROM MAMMOGRAPHY EXAMINATIONS BY USING TRANSFORMER

[0003] NETWORKS

[0004] The present invention relates to a device for assessing and classifying, in an automated fashion and without the need of human intervention, an examination consisting of one or more mammogram images according to one or more predetermined diagnostic categories .

[0005] Mammography is the main imaging method in screening for breast cancer and therefore one of the most important instruments available to reduce mortality caused by such disease . Given the high number of women who annually undergo this type of examination, combined with a well-defined diagnostic task and a sufficiently standardised data acquisition process , mammography screening is an ideal candidate for the use of systems which, using a suitable software, allow to automate, even partially, the mammography diagnosis proces s , currently carried out by one or more expert radiologists depending on the reference guidelines . At the same time, various recent studies show that deep neural networks can be trained to identify the presence of potential cancerous lesions in a mammography, similarly to the classification provided by a radiologist or human reader .

[0006] However, providing a system having a device that is capable of examining images obtained from a common mammography equipment and which operates using a deep neural network appropriately designed and trained (and which allows to identify the lesions mentioned above with high accuracy and with a high degree of safety) is technically complex . As a matter of fact , such system must operate considering several concomitant factors such as : low prevalence of lesions and as a result , the high prevalence of negative examinations ; the high resolution of the mammogram images , especially compared with the photograph images ; and lastly the presence, in a mammography examination, of a plurality of images or views , which must be suitably combined to identify any abnormalities .

[0007] Therefore, it is advantageous to use image analysis devices which use deep neural networks , so-called multiview, which receive as input a plurality of images which form the examination, process them simultaneously, rather than process each view individually, and which assign to the image a classification score according to predetermined categories indicate the probability of a cancerous lesion being detected therein . This score may for example indicate the probability that the images show the presence of a cancer, or the presence of abnormalities that require the need for the physician to recall the woman for further examinations , or even the probability that the examination shows the absence of lesions or abnormalities and therefore can be considered surely negative . The publication on behalf of ZHONG YUTONG ET AL : "Multi-view fusion-based local-global dynamic pyramid convolutional cross-transformer network for density classification in mammography" , 20231115 , vol . 68 , no . 22 , 15 November 2023 , XP020483762 , DOI : 10 . 1088 / 1361- 6560 / AD02D7 , describes a device for carrying out an assessment on a plurality of images according to the preamble of the independent claims of a device and of a method of the present document .

[0008] I f trained on suitably annotated data, the neural network of the analysis device may be used for carrying out other types of classification, such as the risk of developing the disease at five years , or the breast density category .

[0009] Specifically, the images which form a mammography examination, also referred to as projections or views in the jargon of the industry, are generally two per breast and they can be of the cranio-caudal (CC) type or medio-lateral oblique (MLO) type . Combining all the views that form a mammography, multi-view deep neural networks allow to simultaneously carry out both a unilateral and contralateral analysis . As known, unilateral analysis combine the CC and MLO views of the same breast to compensate for the effects of the overlapping of tissues , due to the projective nature of the examination, which could simulate the presence of lesions ( false positives ) . Contralateral analysis allows to compare the structure of the two breasts and can, for example, detect asymmetries or other diagnostic signs which would not emerge were the individual views to be considered separately .

[0010] However, in many multi-view architectures , the image is downsampled significantly to reduce the computational weight and the memory footprint of the network ; this can lead to various drawbacks : as a matter of fact , this downsampling is not desired given that some diagnostic signs , such as microcalcifications , can generally be distinguished only with higher resolution .

[0011] In addition, it is also known that the multi-view architectures which generate a classification score further reveal the technical advantage of being trainable starting from labels associated with the entire examination, without requiring to label the individual pixels which form the images , and therefore reducing the cost required for the implementation thereof .

[0012] In the light of the above, there arises the need to highlight that the performance of a deep neural network significantly depends not only on the quality and quantity of data on which the network was trained, but also on the specific architecture used, where architecture refers to the type and the sequence of mathematical operations calculated by the network, with the relevant weights whose value is determined through a training procedure known to the person skilled in the art . Scientific literature regarding conventional RGB image analysis clearly highlights this fact . In mammography, the state-of-the-art solutions are based on deep Convolutional Neural Networks (CNN) , and above all residual networks .

[0013] It should be observed that Convolutional Neural Networks are characterised by a sequence of layers each comprising at least one convolution operation followed by a non-linear operator ( for example the Rectified Linear Unit or ReLU) . The convolution operators extract from the images local features , whose calculation is based on identical parameters for the entire image, which are then aggregated to carry out the classification . This characteristic allows to reduce the number of examples required for the training, but simultaneously reduces the variety of features that the network can learn and therefore it creates limit s to the maximum performance that it can obtain .

[0014] However, more recently Transformer-type architectures are replacing convolutional neural networks in various medical and non-medical fields . Compared to convolutional neural networks , Transformer networks do not assume that the features are calculated in the same manner on the entire image, but they base the classification on the ability to compare different parts of the image through a particular mathematical operation known as self-attention .

[0015] The self-attention mechanism allows to extract from each image features that not only depend on the combination of adjacent pixels (as implicitly assumed by the use of convolution operators ) , but also on the combination of localised pixels in any point of the image, and therefore compare and correlate parts of the organ examined both inside the same view and on different views .

[0016] Being more flexible that the convolution operation, the self-attention mechanism allows Transformer networks to learn classifications that would not be possible with convolutional neural networks ; at the same time it requires a greater amount of training data to achieve desirable performance .

[0017] In addition, this type of neural networks requires a larger memory for their use and, therefore, their application in assessing mammography examinations was only limited to analysing individual views or analysing low-resolution images .

[0018] However, current research has not yet demonstrated that Transformer networks outperform convolutional neural networks in all scenarios , just like they did not reveal the advantages that can be obtained by combining these two types of architectures .

[0019] An object of the present invention is to provide a device for automatically analysing a mammography examination .

[0020] Another object of the invention is to provide a device that is capable of examining images obtained by a mammographic equipment and which is also capable of generating a result from the examination mentioned above comprising a classification score of the image highlighting, for example a probability of presence or absence of cancer .

[0021] A further object is to provide a device of the type mentioned above which uses deep neural networks for analysing images obtained from the mammography equipment where said neural networks can be efficiently trained on a small training set without requiring annotations on the individual pixels of the images .

[0022] Another object is to provide a device of the type mentioned above which comprises a novel Transformer neural network architecture, designed for the multi-view analysis of images obtained from a mammography examination, which operates also combined with other neural network architectures .

[0023] These and other objects which shall be more apparent to the person skilled in the art are attained by a device according to the attached claims .

[0024] For a better understanding of the present invention, the following drawings are attached hereto, purely by way of non-limiting example, wherein : figure 1 shows a block flow diagram of a first Transformer architecture on which the method according to the invention operates ; figure 2 shows a representation of a mammogram image processed and divided into patches for carrying out a step of the method according to the invention; figure 3 shows a representation of images deriving from that of figure 2 used so that the Transformer architecture of figure 1 can be trained to carry out also the classification of the single patches shown in figure 2 ; figure 4 shows a block flow diagram of a part of the architecture according to figure 1 ; figure 5 shows the block flow diagram of a second architecture with which there can be combined the Transformer architecture of figure 1 , according to the invention; figure 6 shows the block flow diagram of a third architecture with which there can be combined the Transformer architecture of figure 1 according to the invention; figure 7 shows a representation of pseudo-landmark points extracted from a mammogram image and the corresponding division into patches or tessellation of the image ; figure 8 shows a diagram or blocks of a combination function used according to the invention; and figure 9 shows a chart for comparing the performance (Area under the ROC curve ) of devices operating according to the invention in predicting the probability that a carcinogenic lesion be detected on an examined breast with respect to devices operating according to the prior art methods .

[0025] In the present document there is described a device for examining images obtained from a mammography equipment which can operate according to multi-view architectures therefore characterised by different inductive biases . In particular :

[0026] • there is described a first embodiment of the invention mainly based on the use by said analysis device (a microprocessor unit ) of only one transformer-type architecture and subsequently the combination of such architecture with convolutional neural network architectures for analysing mammogram images or views obtained in the mammography examination . In particular, there is introduced a novel network-based architecture with unilateral and cross-view contralateral attention;

[0027] • the various architectures are assessed not only in terms of performance, but also with respect to how they supplement local and global features . The results suggest that various architectures are of complementary nature, in the sense that they are preferentially sensitive to specific signs , and the detection of breast cancer can draw advantage from their supplementation although the Transformers outperform convolutional neural network architectures .

[0028] The present invention relates to the use, by said device for analysing mammogram images , of a deep neural network, also combined with other deep neural networks , which allows classify mammography examinations consisting of two or more images or views autonomously and without human intervention . Although by way of principle any classification can be learnt by the neural network which will be described, hereinafter reference will be made by way of example to a neural network which simultaneously carries out two classifications (multitask configuration) : presence or absence of cancer (that is in the detection of the presence or absence of cancer) , and recall vs . normal examination (that is in the prediction of whether a radiologist would report the mammography as suspicious requiring additional investigation) . It is clear that the neural network can be trained in any case to carry out only one of these two classifications .

[0029] For every woman subjected to examination there are obtained in a per se known manner, the four standard mammogram views or images , indicated below with R-CC, L-CC, L-MLO and L-MLO and shown in the figures , obtained in a per se known manner during mammography . The images are acquired through an equipment or conventional mammography or tomosynthesis ; the present invention can be also used for synthetic mammogram images obtained using prior art techniques starting from three-dimensional tomosynthesis images . Each breast ( left and right ) is associated with binary labels which indicate the presence or absence of cancer (yR, m and yL, m) and if the image was recalled by radiologists for further investigations (yR, r e yL, r) .

[0030] Therefore, the training is formulated by providing the two labels for each side (AyR, m,AyL, m,AyR, r andAyL, r) .

[0031] The device for analysing mammogram images according to the invention mainly uses a neural network, possibly combined with different neural network architectures , each operating on the four views obtained during mammography examination .

[0032] When the present invention uses multiple neural networks characterised by different architectures , at least one is always based on the Transformer-type architecture ; each architecture, operating according to different principles , is particularly apt at recognising different diagnostic signs , and therefore the combination of different neural networks simulates the combined work of different human "experts" .

[0033] As mentioned, the present invention uses a ( first ) Transformer-based multi-view architecture, hereinafter referred to as Mammography Multi-view Transformer (MaMVT) for the sake of brevity .

[0034] The MaMVT Transformer architecture is shown in figure 1 . The four views in input (L-CC, R-CC, L-MLO and R-MLO) are processed simultaneously by a Transformer network which acts as a backbone ( indicated with 20 in figure 1 ) .

[0035] First and foremost , as known from literature, the views are broken down into several patches ( see figure 2 ) , through the block indicated with 10 in figure 1 . In the standard embodiment each patch T corresponds to a rectangular portion of the image, but alternative embodiments can divide the image according to anatomical criteria, as described hereinafter .

[0036] Each patch is associated with its position in the image, and it is converted into an embedding vector by multiplying with a matrix of weights , which forms the first element of the backbone network . The Transformer network which operates as a backbone is characterised by several layers ( indicated as 201 , 202 , 203 , 204 and 205 in figure 1 ) , each containing at least one self-attention block; such network is characterised by weights whose values are obtained through a known training process .

[0037] In a preferred embodiment , the weights used to calculate self-attention are the same for each view, although other possible alternatives are possible ( for example, MLO views could use weights identical to each other, but different from the CC ones ) .

[0038] The self-attention operation supplements the information present in each view by taking in input an embedding vector for each patch and returning in output a new embedding vector for each patch . Possibly sets of adjacent patches may be aggregated to reduce the required number of operations .

[0039] In a preferred embodiment , the sequence and composition of the layers is structured according to the known Swin-vl or Swin-v2 architecture . However, the invention is also compatible with other Transformer networks known in the prior art , such as for example DeiT .

[0040] One of the peculiarities of the MaMVT network architecture is the cross-view attention (CVA) module 40 which suitably combines the representations of the different views . Such module is suitably inserted at an intermediate point of the backbone network, which was experimentally detected enabling to obtain the best performance . The CVA block may be inserted for example between the block 203 and 204 of the backbone . The pairs of anatomically relevant views are corresponding views associated with different breasts ( for example, L-CC and R-CC) or different views associated with the same breast (R-CC and R-MLO, L-CC and L-MLO) .

[0041] Lastly, the embedding vectors suitably processed by the Transformer backbone are aggregated through pooling operations (block 50 of figure 1 ) and data, each, in input at a corresponding classification or classifier level (block 60 of figure 1 ) , known to the person skilled in the art , which calculates the required classification score . In a preferred embodiment , the classifier 60 calculates two classification scores 70 (one per breast ) using weights shared between the two views , but other configurations can be easily derived . During training, such scores are compared with a standard reference to calculate a loss function and train the network according to known optimisation techniques based on the gradient descent .

[0042] Figure 1 is only a specific non-limiting example of the MaMVT architecture used in this work : the four views or images (processed in block 10 ) are passed through a shared Swinv2 backbone (block 20 ) and processed by the CVA block 40 described above which is inserted into the backbone to carry out the cross-referenced attention between each pair of views .

[0043] The CVA block 40 is the critical element for the network to effectively combine the visual information coming from different views . It calculates the attention between the embeddings of the patches T belonging to anatomically correlated pairs of views and then aggregates them, for example through a sum operation, in order to re-obtain in output a sequence of patches for each view with the same size as the input one, as exemplified in figure 4 . This layer implements the attention between two views deriving the query matrix Q of the attention mechanism from the embedding of the patches of a target view a ( for example R-MLO) , while the known and common key matrix K and the known value matrix V are derived from the embedding of the patches of a source view p ( for example R-CC) . This operation is carried out for different copies of target and source views , and the respective results are then combined, for example summed, with the embeddings of the target view . In this manner, both the unilateral attention and the contralateral attention are implemented .

[0044] The CVA module may be implemented in various ways . In a first preferred embodiment , the combinations belonging to the same side are carried out first (L-CC and L-MLO, R-CC and R-MLO) , followed by the combinations of the same type of view (L-CC and L-MLO, R-CC and R-MLO) , therefore calculating the attention four times . It should be observed that the self-attention operation does not have the commutative property (the result changes upon inverting the source and target view) , therefore the result calculated by this implementation depends on the order of the views .

[0045] In a second preferred embodiment ( see figure 4 ) a single attention module is applied and the attention is computed bi-directionally by calculating, for each pair of target view a and source view p , both the attention from view a to view p (e . g . , L-CC and L-MLO) , and from view p to view a (e . g . , L-MLO and L-CC) . The bidirectional attention operations are firstly carried out on each pair of views (blocks 400 ) and then the respective sum operations (nodes 401 ) are carried out simultaneously for each view, as shown in figure 4 , in which, for the sake of clarity, only the sum operations for the views L-CC and R-MLO are shown .

[0046] This second embodiment does not change with respect to the order in which the views are analysed : in other words , the numerical result does not change if the right and left breast are inverted .

[0047] This second embodiment is more effective, in the presence of a limited training data set , so as to prevent the network from learning spurious correlations between the laterality of the lesions and their level of suspicion .

[0048] In a possible embodiment , the MaMVT architecture provides in output a unique classification score per breast or for the entire examination . However, exploiting the patchbased nature of the Transformer architecture, there can be introduced variants , such as for example the prediction of a specific score for each patch into which the image is divided, whose value is linked to the overall classification .

[0049] For example, if the MaMVT network is trained to predict the presence of a lesion, the specific score for each patch will show the presence or absence of the lesion or part of the lesion in that specific area (patch T) of the view or image .

[0050] The classification may be carried out for each patch or, if a Swin backbone is used, for aggregations of adjacent patches patch . Figure 3 shows in its two different parts 3A and 3B, a simplified example on a reduced number of patches T . The first part 3A of figure 3 shows that the mask corresponding to the lesion is converted into a binary label for each patch; the second part 3B of figure 3 shows an example of classification carried out by the neural network . This classification may be obtained from the MaMVT architecture simply by modifying each classification block 60 ( figure 1 ) . The local classification may be used by the radiologist , but it is also useful when training the neural network to calculate an additional loss function for each patch .

[0051] This weak supervision has a minimum computational and memory overload, introducing only one additional classification layer which carries out a few operations for each patch (blocks 60 of figure 1 ) .

[0052] As mentioned, the MaMVT Transformer network may operate alone or combined with one or more neural networks , where this set is characterised by at least two neural networks having different architectures , only one of which is based on the Transformer architecture . The way these architectures are combined is described below .

[0053] A first example of known architecture that can be used to this end is based on the convolutional network ( or CNN) . In this architecture, each view is processed by a convolutional backbone ( see figure 5 ) which maps each view in a multidimensional matrix of fixed size features (blocks 500 ) . The convolutional backbone may have an architecture per se known as a residual network with 18 or 22 layers , also respectively known as ResNetl 8 and ResNet22 . The multidimensional matrices of different views are combined through mathematical operations , for example known pooling operations (blocks 501 ) and concatenation (blocks 502 ) , and data input into one or more classification layers (blocks 503 ) whose output represents the classification score . Of this architecture there are known in literature various variants , characterised by the number of layers of the backbone, by the size of the feature vector, in that the backbones share the same weights or that each view or type of view is processed by a backbone with dedicated weights , and the number and type of layers which form the classification block .

[0054] A second example of architecture that is known and which can be used in the device according to the invention combined with the transformer network is the "Anatomy-aware Graph Convolutional Network" (AGN) ( figure 6 ) , which was designed to more closely mimic how radiologists integrate information on contralateral and lateral views . Generally, these architectures try to explicitly combine the corresponding regions of the various views depending on their geometric and visual properties to emphasise both the abnormalities that appear coherently in the CC and MLO visualisations , and the asymmetries between the right and left breast .

[0055] In order to avoid registering the usual four views obtained during a mammography, given that this might not be feasible due to the effect of the compression on the soft tissues , an attempt was made to introduce additional modules in the standard Convolutional Neural Networks architecture such as relational networks or Graph Convolutional Networks (GCNs ) . In the latter case, a weighted graph models the relation between the local regions from different points of view .

[0056] The architecture (AGN) in particular introduces two graph-based modules , each dedicated to modelling a unilateral analysis ( convolutional network operating on a bipartite graph network or BGN - blocks 605 ) or of a contralateral network ( Inception Graph convolution Network or IGN - block 606 ) . Of this architecture there are known in literature versions that classify a mammogram view at a time, while the architecture described in the present invention was extended to classify four views at the same time .

[0057] The architecture of figure 6 mainly consists in a convolutional backbone (blocks 600 ) , which processes each view similarly to the CNN architecture . The output of the backbone, for each view, is a multidimensional matrix of features , similarly to the Convolutional Neural Networks . Starting from these matrices and from specific points hereinafter referred to as pseudo-landmark points , there are constructed three graphs with a specific mapping function . The pseudo-landmark points are extracted from the block 601 ( figure 6 ) as detailed hereinafter . The mapping function 602 encodes the relationship between each pseudo-landmark point and the features of the corresponding region : each node of the graph associated with a pseudo-landmark point is in turn associated with a (possibly irregular) region of the breast (divided into patches ) .

[0058] The mapping function 602 associates a feature vector with each node of the graph by aggregating, for example summing, the values of the features extracted from the backbone for the pixels belonging to the corresponding patch . An example of pseudo-landmark points , with the corresponding division into patches T (numbered progressively) may be seen in Figure 7 .

[0059] The nodes of the graphs are connected to weighted arcs whose weights , still calculated by the block 602 , respectively model the geometric relations of the unilateral views (module BGN - blocks 605 ) and the geometric and visual relations between the right and left breast ( IGN module - block 606 ) .

[0060] In each BGN module 605 , the weight of the connection is determined as a function of the frequency with which two pseudo-landmark points correspond to the same anatomical region, estimated based on the anatomical structures with known position identified in a set of training images, and as a function of the distance between the feature vectors associated with each node, that is their visual resemblance .

[0061] In this architecture there is the duplication of the BGN module 605 with shared weights , so as to classify both breasts at the same time . These graphs are processed by one or more layers of neural network adapted to operate on graphs , like the graph convolutional networks (block 603 ) . The output of these layers is lastly re-mapped on the image by the block 604 which calculates the reverse function with respect to the block 602 , and combined with the original output of the backbone, for example through a multiplication operation (block 607 ) . The latter operation serves to highlight the features which, based on the existing relations between the images , are most relevant for the classification . Lastly, the result is given in input at one or more layers which calculate the final classification score (block 608 ) .

[0062] The operation of the IGN block 606 is similar to that of the BGN block 605 , but the input images change .

[0063] To extract the pseudo-landmark points , the rules below are followed, whose application is exemplified in figure 7 : • each pseudo-landmark point represents a region 0 , 1 , 2 , ...20 with positions relatively similar to the same projections of different breasts ; • distinct pseudo-landmark points represent distinct regions 0 , > 20 of the breast ;

[0064] • the combination of all pseudo-landmark points covers the entire breast .

[0065] To extract pseudo-landmark points from the CC views , in a known manner, the inventors opted to start from a single anatomical landmark point available, that is the nipple . To extract further pseudo-landmark points , both the longitudinal position of the landmark points just calculated and the breast contour were used, so that the final pseudolandmark points were equally spaced apart along both axes . In the MLO views , the position of the pectoral muscle, together with the nipple, was used to extract the landmark points .

[0066] Starting from the position of the landmark point of the nipple, a line perpendicular to the pectoral muscle is drawn to find the intersection between these two, and subsequently the pseudo-landmark points located on the pectoral muscle are arranged at regular intervals . Lastly, as done for the CC projection using the breast contour and, in this case, two lines parallel to the pectoral muscle, the remaining pseudo-landmark points are positioned .

[0067] The architectures used generate actual results relating to the classification or prediction of the presence of cancer or other clinically relevant variables . Such result s are combined through an appropriate mathematical function (combination function) to obtain a single prediction, whose accuracy exceeds those of the single architectures .

[0068] The mathematical function may be fixed and identical for all mammogram images , such as for example an arithmetic mean or a weighted mean; in the latter case, the weights can be assigned by the designer or they can be obtained by optimising a measurement of the performance of the overall classifier, like the area under the ROC curve, on a set of reference data .

[0069] In an alternative embodiment , the function may be more complex and depend on the features of the individual examination . For example, the combination function may be a weighted mean in which the weights are predicted by a further neural network which, based on the images in question and the classification scores assigned to each network, allocates a different weight to each architecture, according to the technique known to the person skilled in the art as Mixture of Expert s . This is shown in figure 8 which for the sake of simplicity shows the process applied to one of the predictionsAyL, m . In figure 8 the combination function is implemented by a block 703 which carries out a weighted sum of the predictions output by N neural networks 702 , in which the prediction of the i-th neural networkAyiL, m is multiplied by a weight w1and where each weight w1is in turn predicted by a neural network 704 . The neural network 704 could be a neural network with, in output , a known softmax activation function so that the sum of the weights w1is equal to 1 . The neural network is trained in a per se known manner starting from a labelled dataset .

[0070] In another variant , the combination may be carried out by a decision tree which takes as input the scores of the individual architectures and generates in output the overall classification .

[0071] Such combination functions are in turn trained, using known optimisation techniques , on a suitably identified labelled dataset .

[0072] TEST

[0073] Figure 9 shows the ROC (Receiver Operating Characteristics ) curves of the three architectures and of their mean in identifying cancers in the breasts examined using their mammogram images . Such curves were generated by the Inventors after training and assessing the three architectures on a set of mammogram images . As observable, the combination of the results obtained results in having a greater accuracy, measured by the area under the ROC curve .

[0074] In such figure, there are indicated the curves obtained with two transformer networks (MaMVT-vl and MaMVT-v2 ) operating according to the invention and having as backbone the known Swin-vl or Swin-v2 architectures , already mentioned above in the present document .

[0075] From the comparative analysis of three different multiview architectures for classifying breast cancer, that is a convolutional neural network, an AGN architecture and the MaMVT architecture ( in two embodiments ) based on Trasformer, it is observed that not only do these architectures achieve different performance, but they also tend to concentrate on different areas of the breast . Although the Transformerbased architecture individually obtains a more accurate classification, the results indicate that an ensemble of the models may generally improve the performance increasing the area subtended by the ROC curve and identifying a larger number of cancers .

[0076] The presence of such Transformer neural network is however necessary, according to the invention, but also sufficient to allow an automated, independent and reliable assessment of views obtained from a mammography examination so as to identify the presence of cancer in the breasts .

[0077] Therefore, the present invention relates to a device for analysing images obtained from a mammography examination using a known mammography equipment . The device comprises a microprocessor unit which operates using at least one Transformer neural network, possibly combined with another different neural network . The invention also relates to a method for carrying out an automated assessment of images obtained from a mammography examination and which allows to classify such images according to a classification score which allows to detect the presence or absence of a cancerous lesion in a breast .

[0078] The characteristics of such device and of said method are indicated in the claims below .

Claims

CLAIMS1 . Device for carrying out an automated assessment of a plurality of images obtained from a mammography examination, said device being adapted to examine said images by using a neural network architecture, said neural network receiving in input at least two images , the neural network being based on the use of at least one transformer-type module, each image being divided into areas or patches (T) which are analysed by the Transformer neural network, the Transformer network consisting of self-attention layers so as to highlight features relevant for the automated assessment , said Transformer neural network further comprising at least one cross-view attention block ( 40 ) operating on the output of at least one intermediate layer calculated on all examined images , the Transformer neural network therefore carrying out an overall and simultaneous analysis of all images so as to combine the information contained in such images , characterised in that the images are acquired on both breasts , the device comprising mapping blocks ( 605 , 606 ) of each image so as to model geometric relations of unilateral images and the structural resemblance between the left breast and the right breast , in each image of the breast divided into areas or patches (T) there being defined corresponding pseudo-landmark points , corresponding pseudo-landmark points of various breasts representing positions that are relatively similar in the two breasts , the pseudo-landmark points combined covering each entire breast .2 . Device according to claim 1 , characterised in that the pseudo-landmark points are defined considering axes passing through the nipple of the breast and the pectoral muscle, said pseudo-landmark points being equally spaced from both axes .3 . Device according to claim 1 , characterised in thatsaid transformer neural network is adapted to generate a predetermined classification score relating to the examined images .

4. Device according to claim 1, characterised in that the classification score represents a prediction relating to the presence or absence of a cancer in the organ subjected to mammography examination.

5. Device according to claim 1, characterised in that the Transformer neural network consists of several blocks (10, 20, 50, 60) , each having one or more layers (201, 202, 203, 204, 205) each comprising a self-attention block, said blocks progressively aggregating the features calculated on the adjacent areas or patches (T) in the examined images.

6. Device according to claim 1, characterised in that the transformer neural network operates on areas or patches (T) defined by a grid overlapped on each image.

7. Device according to claim 1, characterised in that the transformer neural network operates on four images separately inserted into a backbone (20) of the neural network, there being included classification blocks (60) for classifying corresponding images of the right and left breasts and the cranial-caudal and medio-lateral-oblique image of each breast, said images being coupled through a shared architecture and with a cross-view attention block (40) inserted into such backbone (20) .

8. Device according to claim 7, characterised in that the images of the left breast and of the right breast are concatenated and therefore used to individually classify each patch or area (T) .

9. Device according to claim 7, characterised in that the backbone architecture (20) is of the Swin type.

10. Device according to claim 7, characterised in that the outlet of the cross-view attention block is independentfrom the order of permutation of the input images .11 . Device according to claim 3 , characterised in that the classification is carried out on the global regrouping of the images .12 . Device according to claim 1 , characterised in that it is arranged to combine the use of the transformer neural network with at least one second neural network which autonomously as sesses the images obtained from the mammography examinations , each neural network generating an output related to the assessment thereof of the image, the output data of all the neural networks being combined so as to generate a single result defining the final assessment of the image .13 . Device according to claim 12 , characterised in that the at least one second neural network is a convolutional neural network .14 . Device according to claim 13 , characterised in that the convolutional neural network is alternatively a CNN network or an Anatomy-aware Convolutional Network or AGN operating on a plurality of images .15 . Method for carrying out an automated assessment of a plurality of images obtained from a mammography examination, said method being carried out using the device according to claim 1 , the images being examined by a neural network, said neural network receiving in input at least two images which are examined by said neural network, such neural network being based on the use of at least one Transformertype module, each image being divided into areas or patches (T) which are analysed by the transformer neural network, the transformer network consisting of self-attention layers so as to highlight features relevant for the automated assessment , said Transformer neural network further comprising at least one cross-view attention block operatingon the outputs of at least one intermediate layer calculated on all examined images , the Transformer neural network therefore carrying out an overall and simultaneous analysis of all images so as to combine the information contained in such images , said Transformer neural network therefore generating a predetermined classification score according to one or more diagnostic categories , characterised in that the images are acquired on both breasts , each image being mapped so as to model the geometric relations of unilateral images and the structural resemblance between the left breast and the right breast , in each image of the breast divided into areas or patches (T) there being defined corresponding pseudo-landmark points , corresponding pseudo-landmark points of different breasts representing positions relatively similar in the two breasts , the combined pseudo-landmark points covering each entire breast .16 . Method according to claim 15 , characterised in that the pseudo-landmark points are defined considering axes passing through the nipple of the breast and the pectoral muscle, said pseudo-landmark points being equally spaced from both axes .17 . Method according to claim 15 , characterised in that the classification score represents a prediction relating to the presence or absence of a cancer in the organ subjected to mammography examination .18 . Method according to claim 15 , characterised in that the transformer neural network operates on four images separately inserted into a backbone network ( 20 ) , corresponding images of the right and left breasts and the cranial-caudal and medio-lateral-oblique image of each breast being classified and coupled through a shared architecture and with a cross-view attention block inserted into the backbone .19 . Method according to claim 15 , characterised in that it is provided for to combine the use of the transformer neural network with at least one second neural network which autonomously as sesses the images obtained from the mammography examinations , each neural network generating an output data relating to the assessment thereof of the image, the output data of all the neural networks being combined so as to generate a single result defining the final assessment of the image .