device and method for annotating medical images
The method decomposes medical images into tiles, compares vector representations, and transfers annotations to enhance image interpretation accuracy and efficiency, addressing the imprecision and variability in medical image analysis.
Patent Information
- Application Number
- FR2024001946
- Authority / Receiving Office
- FR · FR
- Patent Type
- Applications
- Current Assignee / Owner
- Filing Date
- 2024-02-28
- Publication Date
- 2025-08-29
- Estimated Expiration
- 2044-02-28
AI Technical Summary
The interpretation of medical images, such as CT scans and MRIs, is imprecise and time-consuming, and can lead to different interpretations by different practitioners or at different times, posing a risk of missed pathologies.
A method for annotating medical images using a foundation model to decompose images into tiles, comparing vector representations of target and reference images, and transferring annotations based on similarity measurements to enhance accuracy and efficiency.
Enables precise and efficient annotation of medical images without re-training neural networks, improving pathology detection and reducing variability in image interpretation.
Smart Images

Figure 00000000_0000_ABST
Abstract
Description
Title of the invention: device and method for annotating medical images Technical field
[0001] The present invention relates to the field of annotation of radiological images. Prior art
[0002] In the field of medical imaging, it is important to note that the interpretation of images from radiology devices, such as CT scans and MRIs, requires considerable expertise and experience on the part of the radiologist. They may also refer to previous images of the patient or request additional studies, such as thinner sections or contrast examinations, to obtain better visualization or a more precise evaluation.
[0003] The radiologist methodically examines the different parts of the body or specific regions that require special attention based on an indication. He evaluates the bone structures, internal organs, blood vessels, soft tissues, etc. He examines the images in search of normal anatomical structures and evaluates their appearance, size, shape and position. This involves recognizing the organs, bones, blood vessels and surrounding tissues. These examinations are often triggered when looking for pathologies, anomalies, such as fractures, lesions (e.g. tumors or coronary artery disease), infections, thrombi, congenital anomalies, signs of disease, etc. He compares the suspicious structures to those that are considered normal.Finally, once the abnormalities have been detected or not, the radiologist synthesizes his observations in a detailed written report, which is then sent to the requesting physician. The report generally includes a description of the normal structures, the abnormalities detected, their location, size, appearance and other relevant characteristics.
[0004] All these steps can make the disease identification process imprecise and time-consuming for a practitioner, and there is a risk of not identifying a pathology. In addition, the interpretation of the same image by two or more different practitioners, or even the interpretation of an image at different time intervals by the same radiologist, can lead to different interpretations of the image. This is, for example, illustrated in the article by Daniela Muenzel et al entitled: “Intra- and interobserver variability in measurement of target lesions: implication on response evaluation according to RECIST 1.1” published on January 2, 2012. There is therefore a need to improve the detection of pathologies during the interpretation of images. medical. Statement of the invention
[0005] The present invention aims to overcome at least one of the drawbacks of the prior art.
[0006] To this end, the present invention proposes a method for annotating at least one target image obtained by medical imaging representing at least one structure of interest: - an application, on a tile decomposition of a plurality of unannotated reference images, of a foundation model previously trained to obtain a vector representation of a medical image decomposed into tiles, - an annotation of a plurality of medical data on said reference images, at least one medical data being annotated on the tiles of said at least one structure of interest present in said reference images, - an application, on a decomposition into tiles of said target image, of said foundation model, to obtain a vector representation of the tiles of said target image - a comparison of the vector representations of the tiles of the target image and the vector representations of the tiles of at least one of said reference images, - a selection, for each tile of a structure of interest of said target image, of at least one tile of at least one reference image, the distance between their vector representation of which is less than a threshold or the distance of which is minimal, - obtaining an annotation of at least one medical data item for the tiles of said structure of interest of said target image, by assigning to each tile of said structure of interest of the target image, the annotation of a medical data item associated with the selected tile of the reference image.
[0007] Thus, advantageously, the present invention can make it possible to annotate medical images without re-training neural networks on new images when it is desired to benefit from the content of new images during the analysis of target medical images. The vectors of the reference images obtained are associated, or even paired with one or more annotations. The annotations are transferred to the target image based on a comparison of the reference vectors with the vectors of the target image.
[0008] According to certain embodiments, said annotation of a plurality of medical data comprises - an association of one or more labels with a tile of a reference image, said labels belonging to one or more atlases according to a determined classification.
[0009] According to certain embodiments, said classification of said labels is chosen among one or more of: - a hierarchical classification of an organ of the human body, the hierarchical classification corresponding to a level of detail of the components of said organ - a pathological classification, - a classification relating to a structure of the human body.
[0010] According to certain embodiments, the method comprises, prior to comparing the vector representations of the tiles of the target image and the vector representations of the tiles of at least one of said reference images, - a selection from said plurality of reference images, reference images representing at least said structure of interest, - said comparison being carried out between the tiles of the target image and the tiles of the selected reference images.
[0011] According to certain embodiments, the method comprises, - a segmentation of at least a portion of said reference images from among said plurality of reference images, to identify one or more structures of interest in said images - a segmentation of the target image, to identify one or more structures of interest in the target image.
[0012] According to certain embodiments, the method comprises, - a segmentation of at least a portion of said reference images from among said plurality of reference images, to identify one or more structures of interest in said images - an E2' segmentation of the target image, to identify one or more structures of interest in the target image.
[0013] According to certain embodiments, the method comprises, - obtaining a vector representation for said segmented structures of interest, from said vector representations of the tiles composing said segmented structures of interest, for the reference images and for the target image - a comparison of the vector representations of the tiles of the target image and the vector representations of the tiles of at least one of said reference images being obtained by a comparison of the vector representations of the segmented structures of interest.
[0014] According to certain embodiments, the method comprises - a selection, for said structures of interest of the target image, of at least one structure of interest of a reference image whose distance between their vector representation is less than a threshold or whose distance is minimal - obtaining an annotation of at least one medical data for said structure of interest of said target image, by assigning to said structure of interest the target image, the annotation of medical data associated with the structure of interest of the selected reference image.
[0015] According to certain embodiments, the method comprises, following the assignment of said annotation to each tile of said structure of interest of the target image, - a grouping of tiles associated with the same annotation, - a determination of the outline of said structure of interest, said outline encompassing said tiles having the same annotation, - an association of said annotation of said tiles of said structure of interest to said structure of interest.
[0016] According to certain embodiments, the method further comprises - smoothing of the contours of said annotated structure of interest.
[0017] According to certain embodiments, the method further comprises, prior to said annotation of the target image, - a selection of an atlas based on one or more criteria chosen from: - a hierarchical level of detail relating to a pathology, or a malformation, or a suspected anomaly, or a structure, -one or more types of structures chosen from vessels, and / or arteries, and / or muscles, and / or bones, - an area of the body, associated with one or more hierarchical levels of detail of this area, said annotation containing one or more labels of the atlas selected for the at least one structure of interest.
[0018] According to certain embodiments, the method further comprises, - a display of said target image with said annotation.
[0019] According to certain embodiments, the method comprises, following the display of said target image with said annotation, - the selection of a new atlas different from a previously selected atlas, the new selected atlas including more precise annotations than the previously selected atlas, - the annotation of at least one medical data item for the tiles of said structure of interest of said target image, by assigning to each tile of said structure of interest of the target image, the annotation of a medical data item associated with the selected tile of the reference image from the new selected atlas.
[0020] According to certain embodiments, the method comprises - a correction of the annotations of at least one medical data for the tiles of said structure of interest of said target image, when said annotation is incorrect, - an enrichment of the set of reference images with said target image on which the annotations have been corrected and an addition of said at least one an- corrected rating for said labels.
[0021] According to certain embodiments, following the obtaining of an annotation of at least one medical data item for said structure of interest of said target image, said target image is added to the set of reference images, said at least one annotation of at least one medical data item for said structure of interest of said target image being added to said labels.
[0022] According to certain embodiments, said at least one target image and / or said at least one reference image are images obtained by computed tomography or by magnetic resonance imaging.
[0023] According to certain embodiments, said foundation model is trained by self-supervised learning.
[0024] According to certain embodiments, said structure may be one or more of - a part of the human body, - a pathology of a part of the human body.
[0025] According to certain embodiments, said part of the human body may be one or more of - bone structures, - internal organs, - blood vessels, - soft tissues.
[0026] According to certain embodiments, said medical data relates to one or more of: - a naming of an organ, and / or a bone structure, and / or a blood vessel, and / or soft tissue, and / or - a naming of a pathology.
[0027] The present invention also relates to a computer program comprising instructions for executing the steps of the method according to any one of its embodiments when said program is executed by a computer.
[0028] The present invention also relates to a computer-readable recording medium on which is recorded a computer program comprising instructions for carrying out the steps of the method according to any of its embodiments.
[0029] The invention also relates to a device for annotating at least one target image obtained by medical imaging representing at least one structure of interest, the device comprising one or more processors configured to - apply, on a tile decomposition of a plurality of unannotated reference images, a previously trained foundation model to obtain a representation vector representation of a medical image broken down into tiles, - annotating a plurality of medical data on said reference images, at least one medical data being annotated on the tiles of said at least one structure of interest present in said reference images, - applying, on a decomposition into tiles of said target image, said foundation model, to obtain a vector representation of the tiles of said target image - comparing the vector representations of the tiles of the target image and the vector representations of the tiles of at least one of said reference images, - selecting, for each tile of a structure of interest of said target image, at least one tile of at least one reference image, the distance between their vector representation of which is less than a threshold or the distance of which is minimal, - obtaining an annotation of at least one medical data item for the tiles of said structure of interest of said target image, by assigning to each tile of said structure of interest of the target image, the annotation of a medical data item associated with the selected tile of the reference image. Brief description of the drawings
[0030] [Fig.l] [Fig.l] represents a method according to certain embodiments of the invention.
[0031] [Fig.2] [Fig.2] represents a schematic view of a medical image comprising several structures on which a grid of patches is positioned,
[0032] [Fig.3] [Fig.3] represents an example of segmentation of structures on a medical image,
[0033] [Fig.4a] [Fig.4a] represents an example of an annotation according to a first classification,
[0034] [Fig.4b] [Fig.4b] represents an example of an annotation according to a second classification,
[0035] [Fig.5] [Fig.5] represents an example of a patch annotation of a medical image,
[0036] [Fig.6] [Fig.6] represents an example of the use of segmentation during annotation,
[0037] [Fig.7] [Fig.7] represents an example of a device according to an embodiment of the invention. Description of the embodiments
[0038] The present invention relates mainly to the field of medical imaging. By medical imaging, we can understand, in a non-exhaustive manner, images of the type X-ray, CT-scan or tomodensitometry (computer tomography), MRI (magnetic resonance imaging), ultrasound, mammography, angiography, en- doscopy...
[0039] The present invention can help to improve the process of diagnosing pathologies from the observation of medical images. In particular, it can help to automatically identify anomalies observed on medical images.
[0040] It can in particular help to annotate medical images from reference images and in particular from structures identified in reference images.
[0041] A reference image may for example be a representative image of an individual identified as a representative individual for a population category. For example, one may take a reference image by gender, by age or age group, by origin. One may thus determine a plurality of reference images representing a human liver according to different ages, genders, etc. One may also determine a plurality of images by pathology. One may thus obtain an image representative of a particular pathology. One may also obtain, for a particular pathology, of a determined organ, a plurality of images representative of subgroups in this particular pathology.
[0042] Finally, it is also possible to acquire reference images for a plurality of different acquisition methods; for example, the reference image database may contain different MRI sequences or even different images associated with different injection times in CT scan.
[0043] By structure, we mean in the remainder of the description, bone structures, internal organs, blood vessels, soft tissues, pathologies, anomalies, such as fractures, lesions (e.g. tumors or coronary diseases), infections, thrombi, congenital anomalies, signs of diseases, etc. thus a structure can be defined as a part of the human body or a pathology relating to a part of the human body or a pathological zone relating to a part of the human body.
[0044] [Fig. 1] shows an embodiment of a method according to the present disclosure.
[0045] A foundation model (known by the English acronym "foundation model") or also called basic model, is trained on a set of medical images, step EL Preferably, to obtain remarkable performances, this set of medical images comprises at least 100 million images. However, this set of medical images can be reduced to a lower number. By image is meant here either a 2D radiograph, or a slice of a 3D acquisition (CT scan or MRI), which typically contains from 100 to 1000 of these images. These images are representative of the diversity of radiological practice and preferably: - cover many different organs / indications, - represent a plurality of varied examples for the same pathology, - are obtained with a plurality of acquisition techniques: CT-Scanner (with or without injection) / MRI (via the different available sequences, which are very numerous), - are obtained using a plurality of imaging systems (CT-scanner, MRI, etc.) supplied by a plurality of different manufacturers, - are representative of individuals of different genders, ages, origins, - can come from a plurality of acquisition centers.
[0046] This step E1 is optional in the sense that the method as disclosed uses the foundation model as obtained during this step E1, which can be obtained well upstream of the method steps described later.
[0047] This step involves adjusting the weights and parameters of the foundation model using the training data. This is typically done by optimizing a loss function using optimization algorithms such as stochastic gradient descent (SGD) or variants thereof. The objective is to minimize the error between the model's predictions and the actual values. The foundation model used may include from a few hundred million parameters to at least one billion parameters to be relevant to the application of the present disclosure. In some embodiments, the number of parameters may be as high as 20 billion.
[0048] Several techniques for training a foundation model are described in the literature and are not the subject of this disclosure. The techniques themselves, typically Self-Supervised Learning, as described in the document entitled "DINOv2: https: / / ai.meta.com / blog / dino-v2-computer-vision-self-supervised-learning / )" are public but the radiological data are not, which means that radiology-specific foundation models are not easily accessible.
[0049] The foundation model used is preferably based on visual transformers, as proposed for example by the ViT (https: / / arxiv.org / abs / 2010.11929) or SWIN (https: / / arxiv.org / abs / 2103.14030) models, which define the patches used for the similarity measurement in the method of the present disclosure.
[0050] The foundation model thus trained can be used in the method steps as described below.
[0051] The foundation model used may be a self-supervised learning (SSL) type model, which is a machine learning approach that allows the model to learn from unlabeled, unannotated data by creating artificial supervised tasks from that data. The foundation model outputs a vector or latent representation, or a latent vector, (known as "embedding" in English), of the data entered, here in this case the medical image received as input or a part of the medical image received as input such as one or more structures of interest.
[0052] More precisely, the image is cut into patches, also called tiles, of intermediate size (typically but not exclusively 16x16 pixels) and the foundation model associates with each of the patches a latent vector or representative vector.
[0053] [Fig.2] represents a schematic view of a medical image and its patch grid or on which a grid of patches is placed.
[0054] The patches may constitute a grid defined by the foundation model before training, and statically placed on the image. Advantageously, the size of a patch may be defined as a function of the size of the smallest structure. In other words, the size of a patch may be such that it is not larger than the size of the smallest medical structure to be studied. Thus, a structure as defined previously may be covered by one or more patches, or parts of patches. Similarly, a patch may be positioned entirely on a structure or overlap one or more structures. The size of the patches, when composed of a square of 16*16 pixels, remains relatively small compared to the majority of the structures and thus, the edge phenomenon induced by the patches included only partially in a structure, may be considered negligible.According to certain embodiments, the method according to the invention may comprise a step of zooming or enlarging the image or a step of increasing the resolution of the image, prior to its coding by the foundation model, so as to adapt the size of the structures according to the size of the patches. This may be advantageous when the size of the patches used by the foundation model is too large compared to the size of the structures to be studied (structures of interest coded by the foundation model).
[0055] According to some embodiments, the size of the vector representation is fixed to the training of the foundation model. For example, the vector representation may be a vector of size 2048.
[0056] Different foundation model training architectures and algorithms can be used to create a foundation model dedicated to medical images. For example, the “CLIP” and “DINO” algorithms are known, which can be used in the present disclosure.
[0057] When it is desired to characterize a structure on a target medical image, for example to detect a pathology on a target medical image (by detecting a pathological structure), for example during a new medical examination, at least one piece of data from this target image or this target image is entered into the foundation model.
[0058] As shown in [Fig.l], the target image data used by the model of foundation can be one or more target images.
[0059] Optionally, the method may comprise a step E2 and a step E2'. Steps E2 and E2' are segmentation steps for segmenting the images into structures of interest present in the images, or part of the structures of interest, to identify one or more structures of interest. This segmentation may be carried out using the method as described in the publication "segment anything" by Alexander Kirillov et al published on April 5, 2023. The segmentation makes it possible to group together parts of the image having visual similarities. The segmentation further groups together patches that are located in the same structure.
[0060] Such segmentation methods may, for example, be based on contour detection methods. Parts of the image having visual similarities are generally parts of the image forming part of the same structure. This can then advantageously make it possible to automatically obtain the contours of the annotated structures.
[0061] At the end of steps E2 and E2', reference images are thus obtained, respectively target images, in which one or more structures are identified or segmented. [Fig. 3] represents an example of segmentation of a medical image, on which two different structures, structure 1 and structure 2, are identified. It can be noted that the segmentation steps consist of isolating or even detecting or determining areas of the medical image according to their characteristics, but do not aim to recognize them anatomically.
[0062] It can of course be noted that steps E2, E3, E4 and E5 (E5 applied to the data relating to the reference images), are carried out upstream of steps E2' and E5 (E5 applied to the data relating to the target images), so as to have vector representations of the patches of the reference images, or even of the segmented images, during step E6. Of course, steps E2, E3, E4 and E5 (E5 applied to the data relating to the reference images) can be carried out as the database of reference images is enriched with other images.
[0063] During a step E3, annotations are determined that can make it possible to obtain one or more atlases. An atlas can be described as being a database of annotated reference patches. In other words, an atlas represents a visual map.
[0064] During step E3, an annotation of at least one medical data item is obtained on said reference images, at least one medical data item being annotated on the patches present in the structures of interest of the reference images.
[0065] The medical data relates to one or more of: - a naming of an organ, and / or a bone structure, and / or a vessel blood, and / or soft tissue, and / or other anatomical structure or substructure and / or - a naming of a pathology
[0066] These annotations are performed on at least one reference image and make it possible to define classes associated with patches, which are called reference patches. An annotation can therefore be performed at the patch level, each patch being able to be annotated, or in other words labeled. A reference patch can correspond to at least one structure of interest, or a part of a structure of interest, which can be identified, and is annotated with the label(s) associated with the structure of interest. Thus, a reference patch can be associated with several labels, for example in a hierarchical manner. A patch can be associated with one or more anatomical data, such as a label of a structure (liver, radius, aorta, etc.) and in addition with one or more pathological data (hepatic steatosis). A patch can for example be associated with the liver label and the hepatic steatosis label.
[0067] In other words, the annotation of medical data comprises the association of one or more labels with a patch of a reference image, the labels belonging to one or more atlases according to a determined classification.
[0068] The classification of labels can be chosen from one or more of: - a hierarchical classification of an organ of the human body, the hierarchical classification corresponding to a level of detail of the components of said organ - a pathological classification, - a classification relating to a structure of the human body,
[0069] According to certain embodiments, the annotations of the reference images are carried out manually by an operator such as a radiologist, for example using a graphical interface on a computer, a tablet, etc.
[0070] According to certain embodiments, the annotations can be carried out more automatically, by software. For example, the “IMAIOS” software can be used, which annotates a reference image and more particularly the structures present in this reference image.
[0071] According to certain embodiments, the annotations comprise at least one anatomical class. By anatomical class, we can mean an organ name, such as liver, heart, lung, etc.
[0072] According to certain embodiments, the annotations comprise at least one pathology relating to an anatomical class. For example, an annotation may consist of a cancerous lesion (for example a lung tumor), metabolic (hepatic steatosis), infectious...
[0073] According to certain embodiments, the annotations made on the reference images are made by structure, in order to facilitate the association of an annotation with a patch. The annotations of the structure are then automatically associated with the patches composing this structure.
[0074] When step E2 is not present, structures of interest are further determined during the annotation step E3, for example by tracing the contours of the structures of interest on the image to determine them geographically. The annotations are then applied to these determined geographical areas (corresponding to anatomical structures).
[0075] For example, when an image contains a representation of a fatty liver, the contours thereof are traced and annotations indicating at least that this structure is a fatty liver are associated with this structure.
[0076] When step E2 is present, the annotations made in step E3 can be made on the reference images segmented during step E2. The segmentation or even the visual identification of the reference structures in the reference images allows for simpler annotation of the structures than when this does not exist. Indeed, according to certain methods, by selecting a segmented structure of interest, for example via a graphical interface, for example by a mouse click on the structure, a window can open in which the radiologist can enter the name of the structure.
[0077] The method may also comprise an optional step E4 comprising both a personalization, E41 and a selection, E42 of the atlas. Indeed, if the target image is obtained following a suspicion of a certain pathology or malformation, a personalization of the type of annotations can help to refine the detection of the anomaly or pathology. The personalization may consist of - determine an atlas class based on a hierarchical level of detail relating to a pathology, a malformation, a suspected anomaly, a structure, - determine a plurality of atlas classes for one or more structures of interest, each linked to a hierarchical level, - determine a specialized atlas class by type of structure, for example a class for vessels, a class for arteries, for muscles, for bones... - determine an atlas class for areas of the body, for example an atlas class for the upper back, for the lower back, for the leg, the arm... this can also be combined at a plurality of hierarchical levels.
[0078] These personalized atlas classes are recorded, like the reference images and the vectors of the associated patches and can be used and enriched collaboratively by a group of practitioners.
[0079] Thus, personalization can have the technical effect of helping with diagnosis by providing a level of detail adapted to a pathology, malformation, anomaly suspected, structure. In addition, depending on their medical specialty, radiologists may wish to obtain annotations differently from their colleagues.
[0080] Thus, without the need to re-train the foundation model, practitioners can then, while benefiting from the trained and shared foundation model, use their own annotation (atlas) for the similarity measurements and annotations of the target image and customize the atlases.
[0081] This customization can make it possible to define one or more versions of the atlas. Figures 4a and 4b illustrate, for example, different levels of granularity of the brain analyzed as a structure. A version of the atlas can correspond to a different naming of the structures, to a particular level of detail for a given structure. More generally, a version of the atlas can correspond to a hierarchical level of detail associated with a structure as illustrated in Figures 4a and 4b. Thus, customization has the advantage of improving the detection of pathologies or anomalies of a structure by providing a significant level of detail on this structure.
[0082] For example, an oncologist may focus on pathologies, sometimes even on pathologies linked to a specific organ. A version of an atlas personalized for this practitioner may consist of an atlas containing a plurality of very precise annotations of this organ. By increasing the level of detail, one can obtain more precision and / or robustness in the analysis of a medical image.
[0083] A rheumatologist will prefer to obtain a version of the atlas containing details on bone structures and such an atlas can allow better detection of diseases such as rheumatism and osteoarthritis.
[0084] The customization of the atlas can be done or redone several times and is not static over time. It can also evolve according to the discovery of new pathologies. This can therefore allow for continually better annotating of medical images, without the need for training of the foundation model.
[0085] A selection of an atlas from among the personalized atlases can be made, step E42, when several atlases are available. The selection can be made according to a pathology sought or a structure of the target image or according to the preferences of a practitioner or even his specialty. The selection can be made according to: - a hierarchical level of detail relating to a pathology, a malformation, a suspected anomaly, a structure, - a type of structure, for example a class for vessels, a class for arteries, for muscles, for bones... - of a body area, for example an atlas class for the upper back, for the lower back, for the leg, the arm...this can also be combined at a plurality of hierarchical levels.
[0086] It can be noted that the selection step E42 can be iterated several times (following step E8 described later) if the previously selected atlas does not allow a sufficiently precise annotation of said structure of interest. By sufficiently precise analysis we can mean that the labels used are too generic or too high level, for example at the level of a structure corresponding to an organ (see [Fig.4a]) and not at the detail of this organ (see [Fig.4b]). We can also mean that the initially selected atlas was linked to a classification by structure while a pathological classification is more appropriate.
[0087] Thus, with a single trained foundation model and a single similarity measurement step, several types or levels of annotation can be obtained.
[0088] During a step E5, the foundation model receives as input, on the one hand, reference data and, on the other hand, target data. As mentioned previously, step E5 is carried out in a decorrelated manner on the reference images and on the target images. The application of the foundation model to the reference images can be carried out in the background from a database of reference images, for example, and when new reference images are added to this database. The method requires at least one reference image to have been used by the foundation model in order to carry out, during a step E6, the similarity measurement. The greater and more diverse or adapted the number of reference images is to the target image, the more relevant the result of the similarity measurement is.By adapted we can mean that the method allows better precision when we have a large quantity of annotated reference images that have been coded by the foundation model which contain the structures of interest present in the target image.
[0089] As indicated in [Fig.l], the data of at least one reference image used by the foundation model can be so-called arbitrary reference images, i.e. not selected (or in other words not pre-sorted) by type of structure.
[0090] As mentioned previously, the foundation model associates a vector representation with a patch of an image. The target images, the reference images are therefore decomposed or even gridded into patches (as can be seen in [Fig.2]) of determined size, and for example of size 16*16 pixels prior to their use by the foundation model. The size of the patches is determined by the parameterized patch size of the foundation model. In other words, the foundation model works on patches of determined size of the reference image and the target image. In other words, the foundation model is applied to a decomposition into patches of the data of a plurality of reference images and the data of the target image. The vector representation obtained for each patch is recorded so as to be used in subsequent steps.
[0091] The method therefore comprises an application, on a decomposition into tiles or patches, of said plurality of reference images, of the trained foundation model to obtain a vector representation of an image decomposed into patches.
[0092] The method therefore comprises an application, on a decomposition into tiles or patches, of the target image, of the trained foundation model to obtain a vector representation of an image decomposed into patches.
[0093] According to certain embodiments, the foundation model receives as input one or more previously selected reference images focusing for example on a structure of interest. This embodiment may be of interest when there is no access to a database of reference images.
[0094] In such an embodiment, it can be considered that the vector representations of the patches of the reference images have not been generated by the foundation model beforehand nor recorded in a database beforehand but are generated at the same time as the generation for the target image.
[0095] As described later, new reference patches can enrich the method when new pathologies are detected, or when new patches are considered particularly relevant to serve as reference patches.
[0096] At the end of step E5, it is therefore possible to have a plurality of reference patches with which a reference vector is associated.
[0097] It is also possible, according to certain embodiments, to have reference vectors per reference structure and target vectors per structure of the target image. Indeed, when the patches are associated with a structure, for example following the segmentation carried out during step E2 (or E2'), a vector can be determined per structure from the vectors of patches associated with the structure obtained by the foundation model. The determination of a vector per structure can be done according to different methods, following the obtaining of a vector per patch. According to one embodiment, the vector associated with the structure can be obtained by taking an arithmetic mean of the vectors of the patches identified as patches of the structure. According to another embodiment, all of the vectors of the structure are all compared to the reference vectors in step E6.
[0098] The application of the foundation model to the reference images therefore generates a plurality of reference vectors, associated with patches, these vectors associated with patches being able to be grouped to determine a vector associated with a structure common to the patches. The reference images, as well as the vectors, associated with the patches or the structures are recorded, for example in a database so as to be used during the similarity measurement step E6.
[0099] The vector description of the target image is compared to the vector description of the reference image or images, step E6. More specifically, the vector representation of a target patch of a target image is compared to the vector representation of the reference patches. According to certain embodiments, when vectors per structure are obtained from the vectors per patch, the similarity measurement can be made on the vectors obtained per structure.
[0100] Step E6 comprises: - the comparison between the vector representation of the patches of the target image and the vector representations obtained for the patches of at least one reference image, - the selection, for each tile of the target image, of at least one tile of at least one reference image, the distance between their vector representation of which is less than a threshold or the distance of which is minimal.
[0101] As mentioned previously, to shorten the calculation times, it may be advantageous to select the reference images used during the similarity measurement according to structures of interest sought in the target image. Thus, prior to the similarity measurement, the method may comprise a step (not shown) of selecting a plurality of reference images comprising the structure(s) of interest present in the target image or assumed to be present in the target image, in order to carry out the similarity measurement with the vectors of the patches of the selected images. Advantageously, the selection consists of selecting the reference images comprising the anatomical structure(s) present in the target image. This selection may advantageously be based solely on anatomical structures and not on pathologies.This selection can be carried out based on received medical data, the medical data being able to relate to a pathology or an anomaly. This data can be considered as statistically potentially present in at least one structure of the target image, for example by a practitioner having carried out a clinical assessment on the patient to whom the target image belongs prior to the creation of the medical image. This selection can consist of selecting reference images containing structures present in the target image, healthy and pathological. The selection can be even more precise by selecting reference images containing the structure(s) present in the target image showing a suspected pathology in the target image.
[0102] Thus, for example, a target liver image can be compared to all or at least a plurality of reference liver images comprising both pathological livers and healthy livers. By selecting only the liver images, the similarity step is faster since the number of images is reduced and in particular reduced to the relevant images.
[0103] More specifically, the structures of interest may be linked to a pathology or a anatomical abnormality. For example, when the radiologist suspects that a patient has NASH disease, the radiologist can select reference images that include a structure corresponding to a liver annotated as a liver with NASH disease and use the reference vectors from these images as reference vectors to be used for the similarity measurement. This advantageously avoids comparing the vectors of the target image with all the vectors of the reference images, most of which are very far apart, thus shortening the computation time.
[0104] The comparison can be carried out by a similarity measurement between the vector representation of the target patch (respectively of the target structure) and each of the vector representations of the reference patches (respectively of the reference structures). The similarity measurement can for example be carried out by measuring the Euclidean distance between vector representations.
[0105] One or more reference vector representations may be selected whose distance is the smallest with the vector representation of the target patch. When several vector representations are selected, the vector representations whose distance is less than a predetermined or parameterized distance may be selected. This can improve robustness because if several reference patches representative of a pathology, for example Nash disease, are selected, this can confirm the diagnosis of this pathology on the target liver.
[0106] We therefore obtain the reference patch associated with the selected vector representation, the reference patch and the associated vector representation being for example linked by an index, a pointer, or linked in a conventional manner in a database.
[0107] The foundation model producing a vector per patch, a grouping of patches forming a structure of interest can be carried out and the similarity measurement can be carried out at the scale of the structure and not of the patch. As mentioned above, when the target image and the reference image are segmented (steps E2' and E2) one vector per structure can be determined. Thus, the similarity measurement can be carried out by comparing a vector of a segmented structure of the target image with the vectors of the segmented structures of the reference images. This can advantageously be faster than comparing the patches one by one.
[0108] Once the reference patch or the reference structure has been selected during step E6, the annotations of this selected reference patch or of this selected reference structure are transferred to the target patch, respectively to the structure of interest of the target image, step E7. [Fig.5] represents an example of annotated patches.
[0109] Following the obtaining of an annotation of at least one medical data for the structure of interest of the target image, the target image is added to the set of images reference, at least one annotation of at least one medical data for the structure of interest of the target image being added to the labels.
[0110] Advantageously, the method can comprise, following the assignment of the annotation to each patch of the structure of interest, - obtaining the outline of the structure of interest, the outline encompassing the patches having the same annotation, - an association of the annotation of the patches of the structure of interest to the structure of interest.
[0111] This allows for clearer reading of the annotations. Indeed, an annotation by patch can be less readable than an annotation by structure. It can be noted, as in [Fig.5] that the annotations can be indicated in the form of color per patch, a color being associated with an annotation. Here, for example, the liver patches are colored light gray and light gray is associated with the liver.
[0112] The structures are generally of a geometric shape different from the shape of the patches. Indeed, the shape of the structures is closer to an ovoid shape while the patches as they are defined are square (or possibly rectangular) in shape. Thus, it can be easily understood that when an image or a structure is decomposed into patches, at least some of the patches may be located on the edge of a structure and include both pixels belonging to the structure and pixels not belonging to the structure creating edge effects. As indicated in [Fig.5], the representation of the patches without them being grouped into a structure offers less good visibility than when the patches are grouped into a structure, [Fig.6].
[0113] Indeed, the patches included on the edges of the structures of interest have representative vectors that are probably far from the representative vectors of the reference patches of this structure of interest since they include both pixels of the structure and pixels outside the structure. A vote can be carried out to determine whether a patch comprising both pixels of the structure and pixels outside the structure is associated with the structure. This vote can be carried out manually by an operator or automatically. This vote can for example be based on a proportion of pixels forming part of the structure, and if this proportion is greater than a first threshold, the patch is selected as belonging to the structure. This determination can be carried out using the segmentation information obtained during step E2'.
[0114] [Fig.6] illustrates an embodiment of obtaining the annotation of the structure of interest from the annotation of the patches and the segmentation. According to this embodiment, the contours of structures (bottom right image) are obtained by the segmentation obtained during step E2'. The classification of the vectors by tile obtained during the similarity measurement of step E6 is illustrated on the top right image. The left image is obtained on which a liver and spleen annotation is obtained by combining the contours of the bottom right image and the classified tiles of the top left image. Thus the tiles of the edges of the structure are associated with the structure according to the contours of the structure.
[0115] According to other embodiments, when the segmentation into the structure of interest and therefore the bottom right image are not available (absence of steps E2 and E2'), the annotated image (left image of [Fig.6]) is obtained from the classification of the patches (top right image of [Fig.6]) on which a vote can be carried out to determine whether a patch comprising both pixels of the structure and pixels outside the structure is associated with the structure. The association of a patch comprising only a part of these pixels in a structure with this structure can be established from anatomical considerations and thus the contours of the structures can be determined.
[0116] Thus, we group the tiles associated with the same annotation, as mentioned above, - we determine the outline of the structure of interest, the outline encompassing the tiles having the same annotation, - we associate, with the structure of interest, the annotation of the tiles of the structure of interest.
[0117] An example is given in [Fig.5] where we can observe the problems of the structure edge effects. Since the reference annotations are moved to each of the patches independently, this is a source of noise: neighboring patches can be associated with different annotations. It can therefore be decided that some of these patches are liver patches and others are not. Smoothing methods can be used to adjust the edges or inconsistencies. The smoothing method can consist of using convolution kernels to attenuate this noise.
[0118] In some embodiments, it may be necessary to adjust the annotations, to adapt to certain morphological constraints of the morphology of the patient whose image is the target image. For example, it may be necessary to adjust contours of the organ of interest.
[0119] In order to be used by the radiologist, the annotated target image can be displayed, for example on a screen, step E8. The radiologist can thus, simply and directly, view the analyzed image of the patient. The method therefore advantageously allows him to benefit from a report created automatically from the reference images. In addition, the level of detail of the annotations may have been personalized during step E41, the annotations are therefore adapted to the personalized level of detail to allow anomalies or pathologies to be detected. If the displayed annotations are not sufficient for the interpretation of a pathology or an anomaly, it may be considered to select a new atlas, step E42, whose annotations are more precise or more suitable. For example, an atlas corresponding to the arm may have been selected. When analyzing the displayed image including the annotations present in the arm atlas, the radiologist may wish to display details about the wrist if he suspects a problem with the wrist. The method may include a new selection of a wrist atlas. Steps E7 and E8 are repeated following this new selection.
[0120] Following the selection of a new atlas, the method comprises steps E7 to E9. During step E7, at least one medical data item is annotated for the tiles of the structure of interest of the target image, by assigning to each tile of the structure of interest of the target image, the annotation of a medical data item associated with the selected tile of the reference image from the new selected atlas.
[0121] In certain embodiments, the method comprises a correction step, E9. This step can advantageously be implemented, for example when the annotations viewed during step E8 are incorrect, in particular at the scale of the structure. This correction can consist of a modification of the annotation or annotations made during step E7, for example by receiving, through a user interface, corrective annotation information. For example, it may happen that new target images are analyzed, containing structures different from those present in the reference images.The similarity measurement between the representative vectors of these new images and the representative vectors of the reference images can lead to selecting reference vectors far from the reference vectors of the target images, even if they are the closest, which can lead to an annotation that is not very precise, or even very far away.
[0122] For example, when there is ascites around the liver, in certain liver pathologies, the appearance of the liver may be altered and a target image containing such a liver may be difficult to annotate as a liver image. Step E9 may correct this annotation error.
[0123] Thus, for example, when a new target image showing a diseased liver bathed in ascites is analyzed, its reference vectors are close to the reference vectors of the previously analyzed and corrected reference image containing such a liver. The annotations of this reference image containing the liver surrounded by ascites are therefore used to be transferred during step E7 to this target image.
[0124] Thus, advantageously, step E9 can allow the method to improve when the corrected images are added to, or even enriched, the database of reference images. These new reference images adding to the reference image database are injected into the foundation model to obtain their reference vector. Thus, the model has a new reference image, its reference vector and the associated annotations. The method as described allows therefore advantageously, without re-training the foundation model, to benefit from the medical information of new reference images.
[0125] According to certain embodiments, during step E9, the radiologist, in addition to correcting, can also adjust the annotations obtained from the annotations of the selected reference image. Thus, it may be possible to customize the annotations to the target image, for example if the radiologist detects new pathologies. Thus, following step E9, the method can move on to step E3 in which the existing atlases can be modified or supplemented from the information identified during the display, step E8. This can make the annotations more robust. In certain cases, new atlases can also be created.
[0126] This customization can also be linked to the radiologist's preferences.
[0127] According to some embodiments, it may be possible to modify the contours by smoothing methods in step E9. This is particularly relevant when the images have not been segmented as described in steps E2 and E2'. The smoothing method may consist of using convolution kernels to attenuate this noise.
[0128] According to certain embodiments, a collaborative platform can be made available allowing practitioners to use such a method by enriching it with their own medical images and / or their own reference images. Thus, when certain practitioners or care centers are specialized in a pathology, an organ, etc., they can annotate reference images targeted on this pathology and / or this organ of interest, these new annotations being able to constitute new semantic classes of the atlas. The present disclosure can thus advantageously make it possible to detect pathologies which are not present in the training images of the model but which are discovered during examinations of the target images. These new target images become reference images and provide new annotations.
[0129] For this purpose, new reference images can be added by practitioners so as to enrich the method and make the similarity measurement more relevant. The higher the number of reference images, the more efficient the method is in terms of annotation and the less adjustment it requires by the practitioner in the formatting of a report.
[0130] The present disclosure may be used for the identification of characteristic “signatures” of pathologies. For example, there are several subgroups of hepatocellular carcinomas (HCC), the most common liver cancer: typically there are the subgroups of pseudo-progressors, progressors, and hyper-progressors. Representative images of a liver affected by each of the pathology subgroups may be annotated by a practitioner and the foundation model may be applied to these new images constituting reference images. By comparing these reference images with those of a new patient with HCC, we can very simply identify the subgroup associated with this HCC. We can thus use the foundation model to define a signature of HCC subgroups.
[0131] The present disclosure can in particular be applied to the detection of Nash disease (an acronym for “non-alcoholic steatosis-hepatitis”). This disease potentially affects a significant number of people in the general population, between 1.5 and 6% of the population, and currently remains underdiagnosed. There are therefore potentially many patients suffering from this disease who are unaware of it in the cohorts of images already collected and which have, for example, been used to train the foundation model. Thanks to the present disclosure, practitioners can annotate images comprising livers presenting this pathology. These images are reference images which are entered into the foundation model and each produce a vector representation (per patch) as output.These vector representations are then also used, with all the vector representations of the patches of the reference images already collected, for the similarity measurements carried out during step E6. Thus, the present disclosure does not require retraining the foundation model with images representative of NASH pathology. A similarity measurement of the vector representations of NASH pathological livers from the foundation model with the vector representations of the reference images presenting NASH livers from the foundation model and annotated can make it possible to quickly and automatically annotate images of NASH livers. The foundation model can thus make it possible to identify a signature of livers having NASH disease.
[0132] The present invention also relates to an annotation device implementing the method as described previously in at least one of these embodiments. To this end, the present disclosure relates to a device for annotating at least one target image obtained by medical imaging representing at least one structure of interest, the device comprising one or more processors configured to - applying, on a tile decomposition of a plurality of unannotated reference images, a previously trained foundation model to obtain a vector representation of a medical image decomposed into tiles, - annotating a plurality of medical data on the reference images, at least one medical data being annotated on the tiles of the at least one structure of interest present in the reference images, - apply, on a tile decomposition of the target image, the foundation model, to obtain a vector representation of the tiles of the target image - compare the vector representations of the tiles of the target image and the vector representations of the tiles of at least one of the reference images, - select, for each tile of a structure of interest of the target image, at least one tile of at least one reference image, the distance between their vector representation of which is less than a threshold or the distance of which is minimal, - obtaining an annotation of at least one medical data item for the tiles of the structure of interest of the target image, by assigning to each tile of the structure of interest of the target image, the annotation of a medical data item associated with the selected tile of the reference image.
[0133] [Fig.7] represents an example of hardware architecture 10 of a device annotation allowing the implementation of a method according to the present disclosure and as represented for example in [Fig. 1]. This hardware architecture is that of a computer. Other hardware architecture elements are present in the device 10 and not represented here.
[0134] The hardware architecture 10 comprises one or more processors 11 (only one is shown in [Fig.2]) implementing a method according to the present disclosure, a read-only memory 12 (of the “ROM” type), a rewritable non-volatile memory 13 (of the “EEPROM” or “NAND Flash” type for example), a rewritable volatile memory 14 (of the “RAM” type), a communication interface 15 with an external network (for example a cellular network for accessing external databases that can record the reference data, the databases being able to be shared and fed collaboratively) and a communication interface 16 (or several communication interfaces) with radiology equipment for example.The read-only memory 12 constitutes a recording medium in accordance with an exemplary embodiment of the invention, readable by the processor or processors 11 and on which is recorded a computer program Prog in accordance with an exemplary embodiment of the invention comprising instructions for executing steps of the method according to one or more of the embodiments of the invention. Alternatively, the computer program Prog is stored in the rewritable non-volatile memory 13. In certain embodiments, the hardware architecture 10 may also comprise a graphics card of the GPU type (for "graphics processor unit" in English), not shown.
[0135] The computer program Prog may enable the device 10 to implement at least part of the method in accordance with the present disclosure and as illustrated for example in [Fig.l].
[0136] This computer program Prog can thus define functional and software modules, configured to implement the steps of an annotation method in accordance with an exemplary embodiment of the invention, or at least part thereof. of these steps. These functional modules rely on or control the hardware elements 11, 12, 13, 14, 15 mentioned above.
Claims
Claims
1. Method for annotating at least one target image obtained by medical imaging representing at least one structure of interest: - an application (E5), on a decomposition into tiles of a plurality of unannotated reference images, of a foundation model previously trained to obtain a vector representation of a medical image decomposed into tiles, - an annotation (E3) of a plurality of medical data on said reference images, at least one medical data being annotated on the tiles of said at least one structure of interest present in said reference images, - an application (E5), on a decomposition into tiles of said target image, of said foundation model, to obtain a vector representation of the tiles of said target image - a comparison (E6) of the vector representations of the tiles of the target image and the vector representations of the tiles of at least one of said reference images,- a selection (E6), for each tile of a structure of interest of said target image, of at least one tile of at least one reference image, the distance between their vector representation of which is less than a threshold or the distance of which is minimal, - an obtaining (E7) of an annotation of at least one medical data item for the tiles of said structure of interest of said target image, by assigning to each tile of said structure of interest of the target image, the annotation of a medical data item associated with the selected tile of the reference image.,
2. Method according to claim 1 wherein said annotation (E3) of a plurality of medical data comprises - an association of one or more labels with a tile of a reference image, said labels belonging to one or more atlases according to a determined classification.
3. Method according to claim 2 in which said classification of said labels is chosen from one or the other or several of: - a hierarchical classification of an organ of the human body, the hierarchical classification corresponding to a level of detail of the components of said organ - a pathological classification, - a classification relating to a structure of the human body,
4. Method according to one of the preceding claims comprising, prior to the comparison of the vector representations of the tiles of the target image and the vector representations of the tiles of at least one of said reference images, - a selection from said plurality of reference images, of the reference images representing at least said structure of interest, - said comparison being carried out between the tiles of the target image and the tiles of the selected reference images.
5. Method according to one of the preceding claims comprising, - a segmentation (E2) of at least a part of said reference images among said plurality of reference images, to identify one or more structures of interest in said images - a segmentation (E2') of the target image, to identify one or more structures of interest in the target image.
6. Method according to claim 5 comprising, - obtaining a vector representation for said segmented structures of interest, from said vector representations of the tiles composing said segmented structures of interest, for the reference images and for the target image - a comparison (E6) of the vector representations of the tiles of the target image and the vector representations of the tiles of at least one of said reference images being obtained by a comparison of the vector representations of the segmented structures of interest, - a selection, for said structures of interest of the target image, of at least one structure of interest of a reference image whose distance between their vector representation is less than a threshold or whose distance is minimal - obtaining an annotation of at least one medical data item for said structure of interest of said target image,by assigning to said structure of interest of the target image, the annotation of medical data associated with the structure of interest of the selected reference image.,
7. Method according to one of claims 1 to 5 comprising, following the assignment of said annotation to each tile of said structure of interest of the target image, - a grouping of the tiles associated with the same annotation, - a determination of the contour of said structure of interest, said contour encompassing said tiles having the same annotation, - an association of said annotation of said tiles of said structure of interest to said structure of interest.
8. The method of claim 7 further comprising - smoothing the contours of said annotated structure of interest.
9. Method according to one of claims 2 to 8 comprising, prior to said annotation of the target image, - a selection (E42) of an atlas according to one or more criteria chosen from: - a hierarchical level of detail relating to a pathology, or a malformation, or a suspected anomaly, or a structure, - one or more types of structures chosen from vessels, and / or arteries, and / or muscles, and / or bones, - an area of the body, associated with one or more hierarchical levels of detail of this area, said annotation containing one or more labels of the atlas selected for the at least one structure of interest.
10. Method according to one of the preceding claims further comprising - a display (E8) of said target image with said annotation.
11. Method according to claims 9 and 10 comprising, following the display of said target image with said annotation, - the selection (E42) of a new atlas different from a previously selected atlas, the new selected atlas comprising more precise annotations than the previously selected atlas, - the annotation of at least one medical data for the tiles of said structure of interest of said target image, by assigning to each tile of said structure of interest of the target image, the annotation of a medical data associated with the selected tile of the reference image from the new selected atlas.
12. Method according to one of claims 10 or 11 comprising - a correction (E9) of the annotations of at least one medical data for the tiles of said structure of interest of said target image, when said annotation is incorrect, - an enrichment of the set of reference images with said target image on which the annotations have been corrected and an addition of said at least one corrected annotation to said labels.
13. Method according to one of the preceding claims in which, following the obtaining of an annotation of at least one medical data for said structure of interest of said target image, said target image is added to the set of reference images, said at least one annotation of at least one medical data for said structure of interest of said target image being added to said labels.
14. A method according to any preceding claim wherein said structure may be one or more of - a part of the human body, - a pathology of a part of the human body.
15. A method according to any preceding claim wherein said human body part may be one or more of - bone structures, - internal organs, - blood vessels, - soft tissues.
16. Method according to one of claims 2 or 3 in which said medical data relates to one or more of: - a naming of an organ, and / or a bone structure, and / or a blood vessel, and / or soft tissues, and / or - a naming of a pathology
17. Computer program comprising instructions for executing the steps of the method according to one of claims 1 to 16 when said program is executed by a computer.
18. A computer-readable recording medium on which is recorded a computer program comprising instructions for carrying out the steps of the method according to one of claims 1 to 16.
19. Device for annotating at least one target image obtained by medical imaging representing at least one structure of interest, the device comprising one or more processors configured to - apply, on a decomposition into tiles of a plurality of unannotated reference images, a foundation model previously trained to obtain a vector representation of a medical image decomposed into tiles, - annotate a plurality of medical data on said reference images, at least one medical data being annotated on the tiles of said at least one structure of interest present in said reference images, - apply, on a decomposition into tiles of said target image, said foundation model, to obtain a vector representation of the tiles of said target image -comparing the vector representations of the tiles of the target image and the vector representations of the tiles of at least one of said reference images, - select, for each tile of a structure of interest of said target image, at least one tile of at least one reference image, the distance between their vector representation of which is less than a threshold or the distance of which is minimal, - obtaining an annotation of at least one medical data item for the tiles of said structure of interest of said target image, by assigning to each tile of said structure of interest of the target image, the annotation of a medical data item associated with the selected tile of the reference image.