device and method for annotating medical images

The method decomposes medical images into tiles and transfers annotations based on vector similarity, addressing imprecision and time inefficiencies in medical image interpretation, enhancing pathology detection.

FR3159699B1Active Publication Date: 2026-03-27RAIDIUM
View PDF 0 Cites 0 Cited by

Patent Information

Authority / Receiving Office
FR · FR
Patent Type
Patents
Current Assignee / Owner
Filing Date
2024-02-28
Publication Date
2026-03-27

AI Technical Summary

Technical Problem

The interpretation of medical images, such as CT scans and MRIs, is imprecise and time-consuming, and can lead to different interpretations by different practitioners or at different times, posing a risk of missed pathologies.

Method used

A method for annotating medical images using a foundation model to decompose images into tiles, comparing vector representations of target and reference images, and transferring annotations based on similarity measurements, allowing for efficient annotation without retraining neural networks.

Benefits of technology

Enables accurate and efficient annotation of medical images, improving pathology detection and reducing variability in image interpretation.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 00000030_0000
    Figure 00000030_0000
  • Figure 00000031_0000
    Figure 00000031_0000
  • Figure 00000031_0001
    Figure 00000031_0001
Patent Text Reader

Abstract

Device and method for annotating medical images. A method for annotating a target image obtained by medical imaging is described, comprising: - applying a foundation model trained to obtain a vector representation of a medical image decomposed into tiles to the reference image tiles, - annotating medical data on said reference images, - applying the foundation model to obtain a vector representation of the tiles of said target image, - comparing the vector representations of the tiles of the target image and the reference images, - selecting, for each tile of said target image, at least one tile of a reference image, based on the distance between their vector representations, - obtaining a medical annotation by assigning to each tile of the target image the annotation of medical data associated with the selected tile of the reference image.Figure for the abridged version: Fig. 1.
Need to check novelty before this filing date? Find Prior Art

Description

Title of the invention: Device and method for annotating medical images technical field

[0001] The present invention relates to the field of radiological image annotation. Previous technique

[0002] In the field of medical imaging, it is important to note that the interpretation of images from radiology equipment, such as CT scans and MRIs, requires considerable expertise and experience on the part of the radiologist. Radiologists may also refer to previous images of the patient or request additional studies, such as thinner slices or contrast-enhanced examinations, to obtain better visualization or a more precise assessment.

[0003] The radiologist methodically examines the different parts of the body or specific regions that require particular attention based on an indication. They evaluate bone structures, internal organs, blood vessels, soft tissues, etc. They examine the images for normal anatomical structures and assess their appearance, size, shape, and position. This involves recognizing organs, bones, blood vessels, and surrounding tissues. These examinations are often triggered when investigating pathologies or abnormalities, such as fractures, lesions (e.g., tumors or coronary artery disease), infections, thrombi, congenital anomalies, signs of disease, etc. They compare suspicious structures to those considered normal.Finally, once abnormalities have been detected or not, the radiologist summarizes their observations in a detailed written report, which is then sent to the referring physician. The report generally includes a description of normal structures, detected abnormalities, their location, size, appearance, and other relevant characteristics.

[0004] All these steps can make the disease identification process imprecise and time-consuming for a practitioner, and there is a risk of failing to identify a pathology. Furthermore, the interpretation of the same image by two or more different practitioners, or even the interpretation of an image at spaced intervals by the same radiologist, can lead to different interpretations of the image. This is illustrated, for example, in the article by Daniela Muenzel et al. entitled: “Intra- and interobserver variability in measurement of target lesions: implication on response evaluation according to RECIST 1.1,” published on January 2, 2012. Therefore, there is a need to improve the detection of pathologies during the interpretation of images. medical. Description of the invention

[0005] The present invention aims to overcome at least one of the drawbacks of the prior art.

[0006] To this end, the present invention proposes a method for annotating at least one target image obtained by medical imaging representing at least one structure of interest: - an application, on a tile decomposition of a plurality of unannotated reference images, of a foundation model previously trained to obtain a vector representation of a medical image decomposed into tiles, - an annotation of a plurality of medical data on said reference images, at least one medical data point being annotated on the tiles of said at least one structure of interest present in said reference images, - an application, on a tile decomposition of said target image, of said foundation model, to obtain a vector representation of the tiles of said target image -a comparison of the vector representations of the tiles in the target image and the vector representations of the tiles in at least one of said reference images, - a selection, for each tile of a structure of interest in said target image, of at least one tile from at least one reference image, the distance between their vector representations of which is less than a threshold or whose distance is minimal, - obtaining an annotation of at least one medical data for the tiles of said structure of interest of said target image, by assigning to each tile of said structure of interest of the target image, the annotation of a medical data associated with the selected tile of the reference image.

[0007] Thus, advantageously, the present invention can enable the annotation of medical images without retraining neural networks on new images when it is desired to benefit from the content of new images during the analysis of target medical images. The vectors of the reference images obtained are associated, or paired, with one or more annotations. The annotations are transferred to the target image based on a comparison of the reference vectors with the vectors of the target image.

[0008] According to certain embodiments, said annotation of a plurality of medical data includes - an association of one or more labels with a tile of a reference image, said labels belonging to one or more atlases according to a determined classification.

[0009] According to certain embodiments, the said classification of said labels is chosen from one or more of the following: - a hierarchical classification of a human body organ, the hierarchical classification corresponding to a level of detail of the components of said organ - a pathological classification, - a classification relating to a structure of the human body.

[0010] According to certain embodiments, the method comprises, prior to comparing the vector representations of the tiles of the target image and the vector representations of the tiles of at least one of said reference images, - a selection from said plurality of reference images, reference images representing at least said structure of interest, - said comparison being carried out between the tiles of the target image and the tiles of the selected reference images.

[0011] According to certain embodiments, the process comprises, - a segmentation of at least a portion of said reference images from said plurality of reference images, to identify one or more structures of interest in said images - a segmentation of the target image, to identify one or more structures of interest in the target image.

[0012] According to certain embodiments, the process comprises, - a segmentation of at least a portion of said reference images from said plurality of reference images, to identify one or more structures of interest in said images - an E2' segmentation of the target image, to identify one or more structures of interest in the target image.

[0013] According to certain embodiments, the process comprises, - obtaining a vector representation for said segmented structures of interest, from said vector representations of the tiles composing said segmented structures of interest, for the reference images and for the target image - a comparison of the vector representations of the tiles of the target image and the vector representations of the tiles of at least one of said reference images being obtained by a comparison of the vector representations of the segmented structures of interest.

[0014] According to certain embodiments, the process includes - a selection, for said structures of interest in the target image, of at least one structure of interest from a reference image whose vector representation is less than a threshold or whose distance is minimal - obtaining an annotation of at least one medical data point for said structure of interest of said target image, by assigning to said structure of interest the target image, the annotation of medical data associated with the structure of interest of the selected reference image.

[0015] According to certain embodiments, the method comprises, following the assignment of said annotation to each tile of said structure of interest of the target image, - a grouping of tiles associated with the same annotation, - a determination of the outline of said structure of interest, said outline encompassing said tiles having the same annotation, - an association of said annotation of said tiles of said structure of interest to said structure of interest.

[0016] According to certain embodiments, the process further comprises - a smoothing of the contours of said annotated structure of interest.

[0017] According to certain embodiments, the method further comprises, prior to said annotation of the target image, - a selection of an atlas based on one or more criteria chosen from: - a hierarchical level of detail relating to a pathology, malformation, suspected anomaly, or structure, -one or more types of structures chosen from among blood vessels, and / or arteries, and / or muscles, and / or bones, - a body area, associated with one or more hierarchical levels of detail for that area, said annotation containing one or more labels from the atlas selected for at least one structure of interest.

[0018] According to certain embodiments, the method further comprises, - a display of said target image with said annotation.

[0019] According to certain embodiments, the method comprises, following the display of said target image with said annotation, - the selection of a new atlas different from a previously selected atlas, the new selected atlas including more precise annotations than the previously selected atlas, - the annotation of at least one medical data for the tiles of said structure of interest of said target image, by assigning to each tile of said structure of interest of the target image, the annotation of a medical data associated with the selected tile of the reference image from the new selected atlas.

[0020] According to certain embodiments, the process includes - a correction of the annotations of at least one medical data point for the tiles of said structure of interest of said target image, when said annotation is incorrect, - an enrichment of the set of reference images with said target image on which the annotations have been corrected and an addition of said at least one an- rating corrected to said labels.

[0021] According to certain embodiments, following the obtaining of an annotation of at least one medical data for said structure of interest of said target image, said target image is added to the set of reference images, said at least one annotation of at least one medical data for said structure of interest of said target image being added to said labels.

[0022] According to certain embodiments, said at least one target image and / or said at least one reference image are images obtained by computed tomography or by magnetic resonance imaging.

[0023] According to some embodiments, said foundation model is trained by self-supervised learning.

[0024] According to certain embodiments, said structure may be one or more of the following: - a part of the human body, - a pathology of a part of the human body.

[0025] According to certain embodiments, said part of the human body may be one or more of - bone structures, - internal organs, - blood vessels, - soft tissues.

[0026] According to certain embodiments, said medical data relates to one or more of the following: - the naming of an organ, and / or a bone structure, and / or a blood vessel, and / or soft tissue, and / or - a name for a pathology.

[0027] The present invention also relates to a computer program comprising instructions for executing the steps of the process according to any one of its embodiments when said program is executed by a computer.

[0028] The present invention also relates to a computer-readable recording medium on which is recorded a computer program comprising instructions for carrying out the steps of the process according to any one of its embodiments.

[0029] The invention also relates to a device for annotating at least one target image obtained by medical imaging representing at least one structure of interest, the device comprising one or more processors configured to - apply, to a tile decomposition of a plurality of unannotated reference images, a foundation model previously trained to obtain a representation Vector representation of a medical image broken down into tiles, - annotate a plurality of medical data on said reference images, at least one medical data point being annotated on the tiles of said at least one structure of interest present in said reference images, - apply, on a tile decomposition of said target image, said foundation model, to obtain a vector representation of the tiles of said target image - compare the vector representations of the tiles of the target image and the vector representations of the tiles of at least one of said reference images, - select, for each tile of a structure of interest of said target image, at least one tile of at least one reference image, whose distance between their vector representation is less than a threshold or whose distance is minimal, - obtain an annotation of at least one medical data for the tiles of said structure of interest of said target image, by assigning to each tile of said structure of interest of the target image, the annotation of a medical data associated with the selected tile of the reference image. Brief description of the drawings

[0030] [Fig.1] Fig.1 represents a method according to certain embodiments of the invention.

[0031] [Fig.2] Fig.2 represents a schematic view of a medical image comprising several structures on which a grid of patches is positioned,

[0032] [Fig.3] Fig.3 represents an example of structure segmentation on a medical image,

[0033] [Fig.4a] Figure [Fig.4a] represents an example of an annotation according to a first classification,

[0034] [Fig.4b] Figure [Fig.4b] represents an example of an annotation according to a second classification,

[0035] [Fig.5] Figure [Fig.5] represents an example of patch annotation of a medical image,

[0036] [Fig.6] Figure 6 represents an example of the use of segmentation during annotation,

[0037] [Fig.7] Fig.7 represents an example of a device according to one embodiment of the invention. Description of the implementation methods

[0038] The present invention relates primarily to the field of medical imaging. Medical imaging may be understood, without limitation, to include images such as radiographs, CT scans (computed tomography), MRI (magnetic resonance imaging), ultrasound, mammography, angiography, etc. doscopy...

[0039] The present invention can help improve the process of diagnosing pathologies from the observation of medical images. In particular, it can help to automatically identify anomalies observed on medical images.

[0040] In particular, it can help to annotate medical images from reference images and especially from structures identified in reference images.

[0041] A reference image can, for example, be a representative image of an individual identified as representative for a population category. For example, a reference image can be chosen for each gender, age or age group, or origin. Thus, a plurality of reference images representing a human liver can be determined according to different ages, genders, etc. A plurality of images can also be determined for each pathology. In this way, a representative image of a particular pathology can be obtained. Similarly, for a particular pathology of a specific organ, a plurality of representative images of subgroups within that particular pathology can be obtained.

[0042] Finally, it is also possible to acquire reference images for a plurality of different acquisition methods; For example, the reference image database can contain different MRI sequences or different images associated with different injection times in CT scan.

[0043] By structure, we mean in the following description, bony structures, internal organs, blood vessels, soft tissues, pathologies, anomalies, such as fractures, lesions (e.g. tumors or coronary diseases), infections, thrombi, congenital anomalies, signs of diseases, etc. Thus a structure can be defined as a part of the human body or a pathology relating to a part of the human body or a pathological area relating to a part of the human body.

[0044] Figure 1 represents an embodiment of a process according to this disclosure.

[0045] A foundation model (also called a basic model) is trained on a set of medical images, step EL. Preferably, to obtain outstanding performance, this set of medical images comprises at least 100 million images. However, this set of medical images can be reduced to a smaller number. An image is understood here to be either a 2D radiograph or a slice of a 3D acquisition (CT scan or MRI), which typically contains 100 to 1000 such images. These images are representative of the diversity of radiological practice and preferably: - cover many different organs / indications, - represent a plurality of varied examples for the same pathology, - are obtained with a plurality of acquisition techniques: CT scan (with or without injection) / MRI (via the various available sequences, which are very numerous), - are obtained using a plurality of imaging systems (CT scanner, MRI, etc.) supplied by a plurality of different manufacturers, - are representative of individuals with different genders, ages, and origins, - can come from a plurality of acquisition centers.

[0046] This step El is optional in the sense that the process as disclosed uses the foundation model as obtained during this step El, which can be obtained well before the steps of the process described later.

[0047] This step involves adjusting the weights and parameters of the foundation model using the training data. This is generally done by optimizing a loss function using optimization algorithms such as stochastic gradient descent (SGD) or its variants. The objective is to minimize the error between the model predictions and the actual values. The foundation model used may include from a few hundred million parameters to at least one billion parameters to be relevant to the application in this disclosure. In some embodiments, the number of parameters can reach 20 billion.

[0048] Several techniques for training a foundation model are described in the literature and are not the subject of this disclosure. The techniques themselves, typically Self-Supervised Learning, as described in the document entitled "DINOv2: https: / / ai.meta.com / blog / dino-v2-computer-vision-self-supervised-learning / ", are publicly available, but the radiological data are not, which means that radiology-specific foundation models are not readily accessible.

[0049] The foundation model used is preferably based on visual transformers, such as those proposed for example by the ViT (https: / / arxiv.org / abs / 2010.11929) or SWIN (https: / / arxiv.org / abs / 2103.14030) models, which define the patches used for the similarity measurement in the process of the present disclosure.

[0050] The foundation model thus trained can be used in the steps of the process as described below.

[0051] The foundation model used may be a self-supervised learning (SSL) model, which is a machine learning approach that allows the model to learn from unlabeled, unannotated data by creating artificial supervised tasks from that data. The foundation model outputs a vector or latent representation, or a latent vector, (known as "embedding" in English), of the input data, here in this case the medical image received as input or a part of the medical image received as input such as one or more structures of interest.

[0052] More precisely, the image is cut into patches, also called tiles, of intermediate size (typically but not exclusively 16x16 pixels) and the foundation model associates each of the patches with a latent vector or representative vector.

[0053] Fig. 2 represents a schematic view of a medical image and its patch grid or on which a grid of patches is placed.

[0054] The patches can form a grid defined by the foundation model before training and statically placed on the image. Advantageously, the size of a patch can be defined based on the size of the smallest structure. In other words, the size of a patch can be such that it is no larger than the size of the smallest medical structure to be studied. Thus, a structure as defined above can be covered by one or more patches, or parts of patches. Similarly, a patch can be positioned entirely on a structure or overlap one or more structures. The size of the patches, when composed of a 16x16 pixel square, remains relatively small compared to the largest part of the structures, and thus, the edge phenomenon induced by patches only partially contained within a structure can be considered negligible.According to certain embodiments, the method according to the invention may include a step of zooming or enlarging the image or a step of increasing the image resolution, prior to its encoding by the foundation model, so as to adapt the size of the structures according to the size of the patches. This can be advantageous when the patch size used by the foundation model is too large compared to the size of the structures to be studied (structures of interest encoded by the foundation model).

[0055] According to some embodiments, the size of the vector representation is fixed during the training of the foundation model. For example, the vector representation may be a vector of size 2048.

[0056] Various architectures and algorithms for training foundation models can be used to create a foundation model dedicated to medical images. For example, the "CLIP" and "DINO" algorithms are known and can be used in this disclosure.

[0057] When it is desired to characterize a structure on a target medical image, for example to detect a pathology on a target medical image (by detecting a pathological structure), for example on the occasion of a new medical examination, at least one data point from this target image or this target image, has entered into the foundation model.

[0058] As shown in [Fig. 1], the target image data used by the model of foundations can be one or more target images.

[0059] Optionally, the method may include an E2 step and an E2' step. E2 and E2' are segmentation steps that segment the images into structures of interest present in the images, or a portion thereof, to identify one or more structures of interest. This segmentation can be performed using the method described in the publication "segment anything" by Alexander Kirillov et al., published on April 5, 2023. The segmentation groups together parts of the image that have visual similarities. Furthermore, the segmentation groups patches that are located within the same structure.

[0060] Such segmentation methods can, for example, be based on edge detection methods. Parts of the image with visual similarities are generally parts of the image that belong to the same structure. This can then advantageously allow the contours of the annotated structures to be obtained automatically.

[0061] Following steps E2 and E2', reference images, or target images respectively, are obtained, in which one or more structures are identified or segmented. Figure 3 shows an example of the segmentation of a medical image, in which two different structures, structure 1 and structure 2, are identified. It should be noted that the segmentation steps consist of isolating, detecting, or determining areas of the medical image based on their characteristics, but do not aim to identify them anatomically.

[0062] It can of course be noted that steps E2, E3, E4 and E5 (E5 applied to the data relating to the reference images) are carried out upstream of steps E2' and E5 (E5 applied to the data relating to the target images), so as to have vector representations of the patches of the reference images, or even of the segmented images, during step E6. Of course, steps E2, E3, E4 and E5 (E5 applied to the data relating to the reference images) can be carried out as the database of reference images is enriched with other images.

[0063] During step E3, annotations are defined that can be used to obtain one or more atlases. An atlas can be described as a database of annotated reference patches. In other words, an atlas represents a visual map.

[0064] During step E3, an annotation of at least one medical data is obtained on said reference images, at least one medical data being annotated on the patches present in the structures of interest of the reference images.

[0065] The medical data relates to one or more of the following: - the naming of an organ, and / or a bone structure, and / or a vessel blood, and / or soft tissue, and / or another anatomical structure or substructure, and / or - the naming of a pathology

[0066] These annotations are performed on at least one reference image and allow the definition of classes associated with patches, which are called reference patches. An annotation can therefore be performed at the patch level, with each patch being annotated, or in other words, labeled. A reference patch can correspond to at least one structure of interest, or a part of a structure of interest, that can be identified, and is annotated with the label(s) associated with the structure of interest. Thus, a reference patch can be associated with several labels, for example, hierarchically. A patch can be associated with one or more anatomical features, such as a label for a structure (liver, radius, aorta, etc.), and also with one or more pathological features (hepatic steatosis). A patch can, for example, be associated with the label "liver" and the label "hepatic steatosis."

[0067] In other words, the annotation of medical data includes the association of one or more labels to a patch of a reference image, the labels belonging to one or more atlases according to a determined classification.

[0068] The classification of the labels can be chosen from one or more of the following: - a hierarchical classification of a human body organ, the hierarchical classification corresponding to a level of detail of the components of said organ - a pathological classification, - a classification relating to a structure of the human body,

[0069] According to some embodiments, the annotations of the reference images are carried out manually by an operator such as a radiologist, for example using a graphical interface on a computer, a tablet...

[0070] According to some embodiments, the annotations can be performed more automatically, by software. For example, the "IMAIOS" software can be used, which annotates a reference image and, more specifically, the structures present in that reference image.

[0071] According to some embodiments, the annotations include at least one anatomical class. By anatomical class, one can understand an organ name, such as liver, heart, lung...

[0072] According to some embodiments, the annotations include at least one pathology related to an anatomical class. For example, an annotation may consist of a cancerous lesion (e.g., a lung tumor), a metabolic lesion (hepatic steatosis), an infectious lesion...

[0073] According to some embodiments, the annotations made on the reference images Annotations are made by structure to facilitate the association of an annotation with a patch. The structure's annotations are then automatically associated with the patches that make up that structure.

[0074] When step E2 is not present, structures of interest are further determined during annotation step E3, for example by tracing the outlines of the structures of interest on the image to define them geographically. The annotations are then applied to these defined geographical areas (corresponding to anatomical structures).

[0075] For example, when an image contains a representation of hepatic steatosis, the contours of the latter are drawn and annotations indicating at least that this structure is hepatic steatosis are associated with this structure.

[0076] When step E2 is present, the annotations made in step E3 can be applied to the reference images segmented during step E2. Segmentation, or the visual identification of reference structures in the reference images, allows for simpler annotation of the structures than when this segmentation is absent. Indeed, according to certain methods, by selecting a segmented structure of interest, for example via a graphical interface (e.g., by clicking on the structure), a window can open in which the radiologist can enter the name of the structure.

[0077] The method may also include an optional step E4 comprising both a personalization, E41, and a selection, E42, of the atlas. Indeed, if the target image is obtained following a suspicion of a certain pathology or malformation, a personalization of the annotation type can help refine the detection of the anomaly or pathology. The personalization may consist of - to determine an atlas class based on a hierarchical level of detail related to a pathology, malformation, suspected anomaly, or structure, - determine a plurality of atlas classes for one or more structures of interest, each linked to a hierarchical level, - determine a specialized atlas class by type of structure, for example a class for vessels, a class for arteries, for muscles, for bones... - determine an atlas class for body areas, for example an atlas class for the upper back, for the lower back, for the leg, the arm... this can also be combined with a plurality of hierarchical levels.

[0078] These custom atlas classes are recorded, like the reference images and vectors of the associated patches, and can be used and enriched collaboratively by a group of practitioners.

[0079] Thus, personalization can have the technical effect of aiding diagnosis by providing a level of detail adapted to a pathology, malformation, anomaly suspected structure. Furthermore, depending on their medical specialty, radiologists may wish to obtain annotations in a different way than their colleagues.

[0080] Thus, without needing to retrain the foundation model, practitioners can then, while benefiting from the trained and shared foundation model, use their own annotation (atlas) for similarity measures and annotations of the target image and customize the atlases.

[0081] This customization allows for the definition of one or more versions of the atlas. Figures 4a and 4b, for example, illustrate different levels of granularity of the brain analyzed as a structure. A version of the atlas may correspond to a different naming convention for structures, or to a particular level of detail for a given structure. More generally, a version of the atlas may correspond to a hierarchical level of detail associated with a structure, as illustrated in Figures 4a and 4b. Thus, the advantage of customization is that it improves the detection of pathologies or anomalies of a structure by providing a significant level of detail about that structure.

[0082] For example, an oncologist may focus on pathologies, sometimes even pathologies related to a specific organ. A customized atlas version for this practitioner might consist of an atlas containing a plurality of very precise annotations of that organ. By increasing the level of detail, greater accuracy and / or robustness can be obtained in the analysis of a medical image.

[0083] A rheumatologist would prefer to obtain a version of the atlas containing details on bone structures and such an atlas can allow better detection of diseases such as rheumatism, osteoarthritis.

[0084] The atlas can be customized multiple times and is not static over time. It can also evolve based on the discovery of new pathologies. This allows for continuously improved annotation of medical images, without the need to train the foundational model.

[0085] A selection of an atlas from among the customized atlases can be made, step E42, when several atlases are available. The selection can be made based on a pathology being sought, a structure of the target image, a practitioner's preferences, or their specialty. The selection can be made based on: - a hierarchical level of detail relating to a pathology, malformation, suspected anomaly, structure, - of a type of structure, for example a class for vessels, a class for arteries, for muscles, for bones... - of a body area, for example an atlas class for the upper back, for the lower back, for the leg, the arm...this can also be combined with a plurality of hierarchical levels.

[0086] It can be noted that the selection step E42 can be iterated several times (following step E8 described later) if the previously selected atlas does not allow for sufficiently precise annotation of the structure of interest. By sufficiently precise analysis, we can mean that the labels used are too generic or too high-level, for example, at the level of a structure corresponding to an organ (see [Fig. 4a]) and not at the level of the detail of that organ (see [Fig. 4b]). We can also mean that the initially selected atlas was linked to a structural classification, whereas a pathological classification is more appropriate.

[0087] Thus, with a single trained foundation model and a single similarity measurement step, several types or levels of annotation can be obtained.

[0088] During step E5, the foundation model receives as input reference data and target data. As mentioned previously, step E5 is performed independently on the reference and target images. The foundation model can be applied to the reference images in the background, for example, using a database of reference images, and as new reference images are added to this database. The process requires that at least one reference image has been used by the foundation model to perform the similarity measurement during step E6. The larger and more diverse the number of reference images, or the more closely they are adapted to the target image, the more relevant the result of the similarity measurement.By "adapted" we can mean that the process allows for better accuracy when a large quantity of annotated reference images are available, having been coded by the foundation model and containing the structures of interest present in the target image.

[0089] As shown in [Fig.1], the data from at least one reference image used by the foundation model can be so-called arbitrary reference images, i.e., not selected (or in other words not pre-sorted) by structure type.

[0090] As mentioned previously, the foundation model associates a vector representation with a patch of an image. The target images, the reference images, are therefore decomposed or gridded into patches (as can be seen in [Fig. 2]) of a predetermined size, for example, 16 x 16 pixels, prior to their use by the foundation model. The size of the patches is determined by the parameterized patch size of the foundation model. In other words, the foundation model works on patches of a predetermined size of the reference image and the target image. The foundation model is applied to a patch decomposition of the data from a plurality of reference images and the data from the target image. The vector representation obtained for each patch is saved for use in subsequent steps.

[0091] The method therefore includes an application, on a decomposition into tiles or patches, of said plurality of reference images, of the trained foundation model to obtain a vector representation of an image decomposed into patches.

[0092] The method therefore includes an application, on a decomposition into tiles or patches, of the target image, of the trained foundation model to obtain a vector representation of an image decomposed into patches.

[0093] According to certain embodiments, the foundation model receives as input one or more pre-selected reference images focusing, for example, on a structure of interest. This embodiment may be useful when access to a database of reference images is unavailable.

[0094] In such an embodiment, it can be considered that the vector representations of the patches of the reference images have not been generated by the foundation model beforehand nor recorded in a database beforehand but are generated at the same time as the generation for the target image.

[0095] As described later, new reference patches can enrich the process when new pathologies are detected, or when new patches are considered particularly relevant to serve as reference patches.

[0096] At the end of step E5, we can therefore have a plurality of reference patches to which a reference vector is associated.

[0097] In certain embodiments, reference vectors can also be provided for each reference structure and target vectors for each structure of the target image. Indeed, when patches are associated with a structure, for example following the segmentation performed in step E2 (or E2'), a vector can be determined for each structure from the patch vectors associated with the structure obtained from the foundation model. The determination of a vector for each structure can be done using different methods, following the acquisition of a vector for each patch. In one embodiment, the vector associated with the structure can be obtained by taking an arithmetic mean of the vectors of the patches identified as patches of the structure. In another embodiment, all the vectors of the structure are compared to the reference vectors in step E6.

[0098] Applying the foundation model to the reference images thus generates a plurality of reference vectors, associated with patches. These vectors associated with patches can be grouped to determine a vector associated with a structure common to the patches. The reference images, as well as the vectors associated with the patches or structures, are stored, for example, in a database so as to be used during the similarity measurement step E6.

[0099] The vector description of the target image is compared to the vector description of the reference image(s), step E6. More specifically, the vector representation of a target patch of a target image is compared to the vector representation of the reference patches. According to some embodiments, when vectors by structure are obtained from vectors by patch, the similarity measure can be performed on the vectors obtained by structure.

[0100] Step E6 includes: - the comparison between the vector representation of the patches of the target image and the vector representations obtained for the patches of at least one reference image, - the selection, for each tile of the target image, of at least one tile from at least one reference image, whose vector representation distance is less than a threshold or whose distance is minimal.

[0101] As mentioned previously, to shorten computation time, it can be advantageous to select the reference images used in the similarity measurement based on structures of interest sought in the target image. Thus, prior to the similarity measurement, the method may include a step (not shown) of selecting a plurality of reference images containing the structure(s) of interest present in the target image or presumed to be present in the target image, in order to perform the similarity measurement with the patch vectors of the selected images. Advantageously, the selection consists of selecting the reference images containing the anatomical structure(s) present in the target image. This selection may advantageously be based solely on anatomical structures and not on pathologies.This selection can be made based on received medical data, which may relate to a pathology or anomaly. This data may be considered statistically potentially present in at least one structure of the target image, for example, by a practitioner who has performed a clinical assessment on the patient to whom the target image belongs prior to acquiring the medical image. This selection may consist of selecting reference images containing structures present in the target image, both healthy and pathological. The selection can be even more precise by selecting reference images containing the structure(s) present in the target image that show a pathology suspected in the target image.

[0102] Thus, for example, a target liver image can be compared to all or at least a plurality of reference liver images, including both pathological and healthy livers. By selecting only liver images, the similarity step is faster since the number of images is reduced, and in particular, reduced to the relevant images.

[0103] More specifically, the structures of interest may be linked to a pathology or a Anatomical anomaly. For example, when a radiologist suspects a patient has NASH, they can select reference images containing a structure corresponding to a liver annotated as having NASH and use the reference vectors from these images as reference vectors for the similarity measurement. This advantageously avoids comparing the vectors of the target image with all the vectors of the reference images, most of which are very different, thus shortening computation time.

[0104] The comparison can be performed by measuring the similarity between the vector representation of the target patch (respectively, of the target structure) and each of the vector representations of the reference patches (respectively, of the reference structures). The similarity measurement can, for example, be performed by measuring the Euclidean distance between vector representations.

[0105] One or more reference vector representations can be selected whose distance is the smallest from the vector representation of the target patch. When selecting multiple vector representations, those whose distance is less than a predetermined or parameterized distance can be chosen. This can improve robustness because if several reference patches representative of a pathology, for example NASH disease, are selected, this can confirm the diagnosis of that pathology in the target liver.

[0106] We thus obtain the reference patch associated with the selected vector representation, the reference patch and the associated vector representation being linked for example by an index, a pointer, or linked in a conventional way in a database.

[0107] Since the foundation model produces one vector per patch, a grouping of patches forming a structure of interest can be performed, and the similarity measure can be carried out at the structure level rather than at the patch level. As mentioned above, when the target image and the reference image are segmented (steps E2' and E2), one vector per structure can be determined. Thus, the similarity measure can be performed by comparing a vector from a segmented structure of the target image to the vectors of the segmented structures of the reference images. This can advantageously be faster than comparing the patches one by one.

[0108] Once the reference patch or reference structure has been selected in step E6, the annotations of this selected reference patch or this selected reference structure are transferred to the target patch, respectively to the structure of interest of the target image, step E7. [Fig.5] shows an example of annotated patches.

[0109] Following the acquisition of an annotation of at least one medical data point for the structure of interest of the target image, the target image is added to the set of images of reference, at least one annotation of at least one medical data for the structure of interest of the target image being added to the labels.

[0110] Advantageously, the method may include, following the assignment of the annotation to each patch of the structure of interest, - obtaining the outline of the structure of interest, the outline encompassing patches with the same annotation, - an association of the annotation of the patches of the structure of interest to the structure of interest.

[0111] This allows for a clearer reading of the annotations. Indeed, an annotation per patch can be less readable than an annotation per structure. As shown in [Fig. 5], the annotations can be indicated by color per patch, with each color associated with a specific annotation. Here, for example, the liver patches are colored light gray, and light gray is associated with the liver.

[0112] Structures are generally of a different geometric shape than patches. Indeed, the shape of structures is closer to an ovoid shape, while patches, as defined, are square (or possibly rectangular). Thus, it is easy to understand that when an image or structure is decomposed into patches, at least some of the patches may be located on the edge of a structure and include both pixels belonging to the structure and pixels not belonging to the structure, creating edge effects. As shown in [Fig. 5], representing patches without grouping them into structures offers less visibility than when patches are grouped into structures, [Fig. 6].

[0113] Indeed, patches located on the edges of structures of interest have representative vectors that are likely quite different from the representative vectors of the reference patches of that structure of interest, since they include both pixels within and outside the structure. A vote can be performed to determine whether a patch containing both pixels within and outside the structure is associated with the structure. This vote can be performed manually by an operator or automatically. For example, this vote can be based on a proportion of pixels belonging to the structure, and if this proportion exceeds a first threshold, the patch is selected as belonging to the structure. This determination can be made using the segmentation information obtained during step E2'.

[0114] Figure 6 illustrates an embodiment of obtaining the annotation of the structure of interest from the annotation of the patches and the segmentation. According to this embodiment, the contours of the structures (bottom right image) are obtained by the segmentation obtained during step E2'. The tile-based vector classification obtained during the similarity measurement of step E6 is shown in the top right image. The image on the left shows an annotation of the liver and spleen. This was obtained by combining the outlines of the bottom right image and the sorted tiles of the top left image. Thus, the tiles at the edges of the structure are associated with the structure based on its outlines.

[0115] According to other embodiments, when the segmentation into the structure of interest is not available, and therefore the image in the lower right is not available (absence of steps E2 and E2'), the annotated image (left image of [Fig. 6]) is obtained from the classification of the patches (upper right image of [Fig. 6]). A vote can then be performed on this classification to determine whether a patch containing both pixels of the structure and pixels outside the structure is associated with that structure. The association of a patch containing only some of these pixels within a structure with that structure can be established based on anatomical considerations, and thus the contours of the structures can be determined.

[0116] Thus, we group together the tiles associated with the same annotation, as mentioned above, - we define the outline of the structure of interest, the outline encompassing the tiles having the same annotation, - we associate, with the structure of interest, the annotation of the tiles of the structure of interest.

[0117] An example is given in [Fig. 5] which illustrates the problems of structure edge effects. Since the reference annotations are offset independently on each patch, this is a source of noise: neighboring patches may be associated with different annotations. It may therefore be decided that some of these patches are liver patches and others are not. Smoothing methods can be used to adjust the edges or inconsistencies. The smoothing method may involve using convolution kernels to attenuate this noise.

[0118] In certain embodiments, it may be necessary to adjust the annotations to adapt to certain morphological constraints of the patient's morphology, whose image is the target image. For example, it may be necessary to adjust the contours of the organ of interest.

[0119] For use by the radiologist, the annotated target image can be displayed, for example on a screen, step E8. The radiologist can thus easily and directly view the analyzed image of the patient. This process advantageously allows them to benefit from a report automatically generated from the reference images. Furthermore, the level of detail of the annotations may have been customized during step E41, so the annotations are adapted to the customized level of detail to enable the detection of anomalies or pathologies. If the displayed annotations are insufficient for the interpretation of a pathology or anomaly, a new atlas with more precise or more appropriate annotations can be selected, step E42. For example, an atlas corresponding to the arm may have been Selected. When analyzing the displayed image, which includes annotations from the arm atlas, the radiologist may wish to display details about the wrist if a problem is suspected. This procedure may involve reselecting a wrist atlas. Steps E7 and E8 are repeated after this reselecting.

[0120] Following the selection of a new atlas, the process includes steps E7 to E9. During step E7, at least one medical data is annotated for the tiles of the structure of interest of the target image, by assigning to each tile of the structure of interest of the target image, the annotation of a medical data associated with the selected tile of the reference image from the new selected atlas.

[0121] In some embodiments, the method includes a correction step, E9. This step can advantageously be implemented, for example, when the annotations displayed during step E8 are incorrect, particularly at the structural level. This correction can consist of modifying the annotation(s) made during step E7, for example, by receiving corrective annotation information through a user interface. For example, it may happen that new target images are analyzed, containing structures different from those present in the reference images.The similarity measurement between the representative vectors of these new images and the representative vectors of the reference images can lead to the selection of reference vectors that are far removed from the reference vectors of the target images, even if they are the closest, which can lead to an inaccurate, or even very far removed, annotation.

[0122] For example, when there is ascites around the liver, in certain liver pathologies, the appearance of the liver may be altered, and a target image containing such a liver may be difficult to annotate as a liver image. Step E9 can correct this annotation error.

[0123] Thus, for example, when a new target image showing a diseased liver surrounded by ascites is analyzed, its reference vectors are close to the reference vectors of the previously analyzed and corrected reference image containing such a liver. The annotations of this reference image containing the liver surrounded by ascites are therefore used to be transferred to this target image in step E7.

[0124] Advantageously, step E9 allows the process to improve when the corrected images are added to, or enriched by, the reference image database. These new reference images added to the reference image database are injected into the foundation model to obtain their reference vector. Thus, the model has a new reference image, its reference vector, and associated annotations. The process as described allows therefore advantageously, without retraining the foundation model, to benefit from the medical information of new reference images.

[0125] According to some embodiments, during step E9, the radiologist, in addition to correcting, can also adjust the annotations obtained from the annotations of the selected reference image. Thus, it may be possible to personalize the annotations to the target image, for example, if the radiologist detects new pathologies. Following step E9, the process can proceed to step E3, in which existing atlases can be modified or supplemented using the information identified during display in step E8. This can make the annotations more robust. In some cases, new atlases can also be created.

[0126] This personalization can also be linked to the radiologist's preferences.

[0127] According to some embodiments, it may be possible to modify the contours by Smoothing methods are used in step E9. This is particularly relevant when the images have not undergone segmentation as described in steps E2 and E2'. The smoothing method may involve using convolution kernels to reduce this noise.

[0128] In certain embodiments, a collaborative platform can be made available, enabling practitioners to use such a process by enriching it with their own medical images and / or reference images. Thus, when certain practitioners or healthcare centers specialize in a pathology, an organ, etc., they can annotate reference images focused on that pathology and / or organ of interest. These new annotations can constitute new semantic classes in the atlas. This disclosure can therefore advantageously enable the detection of pathologies that are not present in the model's training images but are discovered during examinations of the target images. These new target images become reference images and provide new annotations.

[0129] To this end, new reference images can be added by practitioners to enrich the process and make the similarity measurement more relevant. The greater the number of reference images, the more efficient the process is in terms of annotation and the less adjustment is required by the practitioner in formatting a report.

[0130] This disclosure can be used for the identification of characteristic “signatures” of pathologies. For example, there are several subgroups of hepatocellular carcinoma (HCC), the most common liver cancer: typically, there are the subgroups of pseudoprogressors, progressors, and hyperprogressors. Representative images of a liver affected by each of the pathology subgroups can be annotated by a practitioner, and the foundation model can be applied to These new images serve as reference images. By comparing these reference images with those of a new patient with HCC, the subgroup associated with that HCC can be easily identified. The foundation model can then be used to define a signature for HCC subgroups.

[0131] This disclosure can be applied, in particular, to the detection of NASH (non-alcoholic steatohepatitis, also known as NASH). This disease potentially affects a significant number of people in the general population, between 1.5% and 6%, and remains currently underdiagnosed. Therefore, there are potentially many patients with this disease who are unaware of it in the image cohorts already collected, which have, for example, been used to train the foundation model. Thanks to this disclosure, practitioners can annotate images containing livers exhibiting this pathology. These images are reference images that have been fed into the foundation model and each produces a vector representation (per patch) as output.These vector representations are then also used, along with all the vector representations of the patches of the reference images already collected, for the similarity measurements performed in step E6. Thus, the present disclosure does not require retraining the foundation model with images representative of NASH pathology. A similarity measurement of the vector representations of pathological NASH livers from the foundation model with the vector representations of reference images of NASH livers from the foundation model and annotated can allow for the rapid and automatic annotation of NASH liver images. The foundation model can therefore be used to identify a signature of livers with NASH disease.

[0132] The present invention also relates to an annotation device implementing the method as described above in at least one of these embodiments. To this end, the present disclosure relates to an annotation device for at least one target image obtained by medical imaging representing at least one structure of interest, the device comprising one or more processors configured to - to apply, to a tile decomposition of a plurality of unannotated reference images, a foundation model previously trained to obtain a vector representation of a medical image decomposed into tiles, - annotate multiple medical data points on the reference images, with at least one medical data point annotated on the tiles of at least one structure of interest present in the reference images, - Apply the foundation model to a tile decomposition of the target image to obtain a vector representation of the target image tiles. - compare the vector representations of the tiles in the target image and the vector representations of the tiles in at least one of the reference images, - select, for each tile of a structure of interest in the target image, at least one tile from at least one reference image, whose distance between their vector representations is less than a threshold or whose distance is minimal, - obtain an annotation of at least one medical data for the tiles of the structure of interest of the target image, by assigning to each tile of the structure of interest of the target image, the annotation of a medical data associated with the selected tile of the reference image.

[0133] Figure 7 represents an example of the hardware architecture 10 of a device Annotation enabling the implementation of a process according to this disclosure and as represented, for example, in [Fig. 1]. This hardware architecture is that of a computer. Other hardware architecture elements are present in device 10 and are not shown here.

[0134] The hardware architecture 10 includes one or more processors 11 (only one is shown in [Fig.2]) implementing a method according to this disclosure, a read-only memory 12 (of the "ROM" type), a rewritable non-volatile memory 13 (of the "EEPROM" or "NAND Flash" type, for example), a rewritable volatile memory 14 (of the "RAM" type), a communication interface 15 with an external network (for example, a cellular network to access external databases capable of storing reference data, the databases being able to be shared and fed collaboratively) and a communication interface 16 (or several communication interfaces) with radiology equipment, for example.The read-only memory 12 constitutes a storage medium according to an exemplary embodiment of the invention, readable by the processor or processors 11, and on which is stored a computer program Prog according to an exemplary embodiment of the invention comprising instructions for executing steps of the process according to one or more embodiments of the invention. Alternatively, the computer program Prog is stored in the rewritable non-volatile memory 13. In some embodiments, the hardware architecture 10 may also include a GPU (graphics processing unit) type card, not shown.

[0135] The computer program Prog can enable the device 10 to implement at least a part of the process in accordance with this disclosure and as illustrated for example in [Fig.1].

[0136] This computer program Prog can thus define functional and software modules, configured to implement the steps of an annotation process conforming to an example embodiment of the invention, or at least a part thereof. of these steps. These functional modules rely on or control the hardware elements 11, 12, 13, 14, 15 mentioned previously.

Claims

Demands

1. A method for annotating at least one target image obtained by medical imaging representing at least one structure of interest: - an application (E5), on a tile decomposition of a plurality of unannotated reference images, of a foundation model previously trained to obtain a vector representation of a tiled medical image, - an annotation (E3) of a plurality of medical data on said reference images, at least one medical data being annotated on the tiles of said at least one structure of interest present in said reference images, - an application (E5), on a tile decomposition of said target image, of said foundation model, to obtain a vector representation of the tiles of said target image - a comparison (E6) of the vector representations of the tiles of the target image and the vector representations of the tiles of at least one of said reference images,- a selection (E6), for each tile of a structure of interest in said target image, of at least one tile from at least one reference image, the distance between their vector representations of which is less than a threshold or whose distance is minimal, - obtaining (E7) an annotation of at least one medical data point for the tiles of said structure of interest in said target image, by assigning to each tile of said structure of interest in the target image, the annotation of a medical data point associated with the selected tile of the reference image.

2. Method according to claim 1 wherein said annotation (E3) of a plurality of medical data comprises - an association of one or more labels to a tile of a reference image, said labels belonging to one or more atlases according to a determined classification.

3. A method according to claim 2 wherein said classification of said labels is chosen from one or more of the following: - a hierarchical classification of a human body organ, the hierarchical classification corresponding to a level of detail of the components of said organ - a pathological classification, - a classification relating to a structure of the human body,

4. A method according to any one of the preceding claims comprising, prior to comparing the vector representations of the tiles of the target image and the vector representations of the tiles of at least one of said reference images, - a selection from said plurality of reference images, of the reference images representing at least said structure of interest, - said comparison being carried out between the tiles of the target image and the tiles of the selected reference images.

5. A method according to any one of the preceding claims comprising, - a segmentation (E2) of at least a part of said reference images among said plurality of reference images, to identify one or more structures of interest in said images - a segmentation (E2') of the target image, to identify one or more structures of interest in the target image.

6. A method according to claim 5 comprising: - obtaining a vector representation for said segmented structures of interest, from said vector representations of the tiles composing said segmented structures of interest, for the reference images and for the target image; - a comparison (E6) of the vector representations of the tiles of the target image and the vector representations of the tiles of at least one of said reference images being obtained by a comparison of the vector representations of the segmented structures of interest; - a selection, for said structures of interest of the target image, of at least one structure of interest from a reference image whose vector representations are separated by a distance less than a threshold or whose distance is minimal; - obtaining an annotation of at least one medical data point for said structure of interest of said target image.by assigning to said structure of interest in the target image, the annotation of medical data associated with the structure of interest in the selected reference image.

7. A method according to any one of claims 1 to 5 comprising, following the assignment of said annotation to each tile of said structure of interest in the target image, - a grouping of the tiles associated with the same annotation, - a determination of the contour of said structure of interest, said contour encompassing said tiles having the same annotation, - an association of said annotation of said tiles of said structure of interest to said structure of interest.

8. Method according to claim 7 further comprising - smoothing of the contours of said annotated structure of interest.

9. A method according to any one of claims 2 to 8 comprising prior to said annotation of the target image, - a selection (E42) of an atlas based on one or more criteria chosen from: - a hierarchical level of detail relating to a pathology, or a malformation, or a suspected anomaly, or a structure, - one or more types of structures chosen from vessels, and / or arteries, and / or muscles, and / or bones, - a body area, associated with one or more hierarchical levels of detail of this area, said annotation containing one or more labels of the atlas selected for at least one structure of interest.

10. A method according to any one of the preceding claims further comprising - a display (E8) of said target image with said annotation.

11. A method according to claims 9 and 10 comprising, following the display of said target image with said annotation, - the selection (E42) of a new atlas different from a previously selected atlas, the new selected atlas comprising more precise annotations than the previously selected atlas, - the annotation of at least one medical data for the tiles of said structure of interest of said target image, by assigning to each tile of said structure of interest of the target image, the annotation of a medical data associated with the selected tile of the reference image from the new selected atlas.

12. A method according to any one of claims 10 or 11 comprising - a correction (E9) of the annotations of at least one medical data for the tiles of said structure of interest of said target image, when said annotation is incorrect, - an enrichment of the set of reference images with said target image on which the annotations have been corrected and an addition of said at least one corrected annotation to said labels.

13. A method according to any one of the preceding claims, wherein, following the annotation of at least one medical data point for said structure of interest of said target image, said target image is added to the set of reference images, said at least one annotation of at least one medical data for said structure of interest of said target image being added to said labels.

14. A method according to any one of the preceding claims wherein said structure may be one or more of - a part of the human body, - a pathology of a part of the human body.

15. A method according to any one of the preceding claims in which said part of the human body may be one or more of - bone structures, - internal organs, - blood vessels, - soft tissues.

16. A method according to any one of claims 2 or 3, wherein said medical data relates to one or more of the following: - the naming of an organ, and / or a bone structure, and / or a blood vessel, and / or soft tissue, and / or - the naming of a pathology

17. Computer program comprising instructions for carrying out the steps of the process according to any one of claims 1 to 16 when said program is executed by a computer.

18. Computer-readable recording medium on which is recorded a computer program comprising instructions for carrying out the steps of the process according to any one of claims 1 to 16.

19. A device for annotating at least one target image obtained by medical imaging representing at least one structure of interest, the device comprising one or more processors configured to: - apply, to a tile decomposition of a plurality of unannotated reference images, a foundation model previously trained to obtain a vector representation of a tiled medical image; - annotate a plurality of medical data on said reference images, at least one medical data being annotated on the tiles of said at least one structure of interest present in said reference images; - apply, to a tile decomposition of said target image, said foundation model, to obtain a vector representation of the tiles of said target image -compare the vector representations of the tiles in the target image and the vector representations of the tiles in at least one of said reference images, - select, for each tile of a structure of interest in said target image, at least one tile from at least one reference image, the distance between their vector representations of which is less than a threshold or whose distance is minimal, - obtain an annotation of at least one medical data for the tiles of said structure of interest of said target image, by assigning to each tile of said structure of interest of the target image, the annotation of a medical data associated with the selected tile of the reference image.