Training method for a semantic segmentation model of medical images
The method generates synthetic images from segmentation masks and uses semi-supervised training with weak annotations and a teaching network to improve medical image segmentation, addressing overtraining and rare class biases, achieving high-performance segmentation.
Patent Information
- Application Number
- FR2023008686
- Authority / Receiving Office
- FR · FR
- Patent Type
- Patents
- Current Assignee / Owner
- Filing Date
- 2023-08-11
- Publication Date
- 2025-07-25
- Estimated Expiration
- 2043-08-11
AI Technical Summary
Existing methods for training semantic segmentation neural networks in medical imaging face challenges such as the need for large amounts of manually segmented data, particularly for rare anatomical abnormalities, leading to over-learning, insufficient segmentation performance, and learning biases against rare classes, with insufficient medical expertise involvement.
A method involving the generation of synthetic images using a generator neural network from segmentation masks, combined with real and unsegmented medical images with weak annotations, and using an exponential moving average to update a teaching network for semi-supervised training, incorporating pseudo-segmentation masks to refine predictions.
This approach enhances segmentation performance by limiting overtraining and biases, ensuring effective segmentation of rare classes with reduced expert input, achieving high-quality medical image segmentation.
Smart Images

Figure 00000025_0000 
Figure 00000025_0001 
Figure 00000026_0000
Abstract
Description
Title of the invention: Method for training a semantic segmentation model of medical images Field of invention
[0001] The present application belongs to the field of semantic segmentation of medical images. In particular, a method for training a segmentation neural network is proposed, as well as a computer-readable storage medium comprising a segmentation neural network trained by such a method. State of the art
[0002] Machine learning algorithms for semantic segmentation of medical images are generally implemented by artificial neural networks. In particular, it is known that training a convolutional neural network based on a U-net type model makes it possible to perform semantic image segmentation. Semantic segmentation involves labeling each pixel of an image with a class corresponding to what is represented in the image. In the field of medical images, a class can represent a particular anatomical structure, such as an anatomy of interest (this can be an organ such as the liver, a lung or a kidney, or another anatomical structure such as a bone or a blood vessel), an anomaly within this anatomy of interest (for example a cyst, a tumor or an ablation zone), or a surgical artifact (for example a part of a medical instrument).
[0003] Training a segmentation neural network requires a large amount of training data to achieve good prediction quality. Also, underrepresentation of a class in the training data can produce a significant prediction bias against that class.
[0004] This problem is of particular importance in the medical field. Indeed, to train the neural network in a supervised manner, the training data must contain a large number of segmented medical images. Each medical image must then be segmented manually, or possibly semi-automatically, by a medical expert. This segmentation of medical images is particularly costly in terms of time and expertise.
[0005] It is even more difficult to collect a large number of medical images showing rare anatomical abnormalities (due to the rarity of diseases, patient confidentiality, the effort and expense required to conduct medical imaging operations, etc.).
[0006] The document “Deep semi-supervised segmentation with weight-averaged consistency targets”, Christian S. PERONE et al., describes a method for semi-supervised training of a segmentation neural network using the “Mean Teacher” approach.
[0007] Several prior art solutions are also based on a relatively similar approach in which the segmentation neural network that one seeks to train is reproduced to form a teaching network. In particular, patent applications WO 2021 / 140426 Al, US 2022 / 0292689 Al, US2021 / 0407656 Al and US2022 / 0358658 Al describe different methods for training a medical image segmentation neural network.
[0008] These solutions, however, have several drawbacks, including over-learning of the segmentation algorithm on the supervised training data, sometimes insufficient segmentation performance, learning biases against the segmentation of rare classes, and / or an insufficient level of involvement of medical expertise in the learning of the segmentation neural network.
[0009] Patent application WO 2022 / 238640 A1 describes a method for generating synthetic images presenting rare anatomical anomalies using generative antagonistic networks. Statement of the invention
[0010] The present application aims to propose a solution to all or part of the drawbacks of the prior art, in particular those set out above.
[0011] For this purpose, and according to a first aspect, a method is proposed for training a segmentation neural network to segment anatomical structures on a medical image representing an anatomy of interest of a patient. The method comprises the following steps: - obtaining real segmented medical images, each associated with a segmentation mask, including: • majority real medical images associated with majority segmentation masks and representing the anatomy of interest of a patient without anomaly, • minority real medical images associated with minority segmentation masks and representing the anatomy of interest of a patient with an anomaly, - training a generator neural network to generate a synthetic image from a segmentation mask, - obtaining artificial segmentation masks, each artificial segmentation mask being obtained by combining a segmentation mask majority with a segmentation of an anomaly of a minority segmentation mask, - obtaining synthetic images, from artificial segmentation masks, using the trained generator neural network, - obtaining real, non-segmented medical images, each representing the anatomy of interest of a patient and each associated with a weak annotation, - an identical reproduction of the segmentation neural network to form a teaching network.
[0012] The method further comprises several iterations of the following steps: - supervised training of the segmentation neural network from supervised training data comprising the segmented real medical images and their associated segmentation masks, as well as the synthetic images and their associated artificial segmentation masks, - updating the weights of the segmentation neural network by gradient descent, - updating the weights wt of the teaching network by an exponential moving average calculated as a function of the previous weights wTprec of the teaching network and the updated weights of the segmentation neural network, - generating pseudo-segmentation masks corresponding to predictions of the teaching network obtained from semi-supervised training data comprising the unsegmented real medical images and their associated weak annotations, - an addition to the supervised training data of the generated pseudo-segmentation masks and the real unsegmented medical images which allowed their generation.
[0013] The construction of artificial segmentation masks and the generation of associated synthetic images makes it possible to limit overtraining of the segmentation neural network. Indeed, it is possible to construct an infinity of different artificial segmentation masks by a random combination of the majority segmentation masks with the minority segmentation masks, and possibly by a transformation of the segmentation of an anomaly of a minority segmentation mask before its combination with a majority segmentation mask. Each synthetic image thus generated can then be used only once in the training of the segmentation algorithm.
[0014] Real medical images can be segmented by medical experts. Weak annotations can also be provided by medical experts. This ensures a sufficient level of involvement of medical expertise in learning the segmentation neural network.
[0015] A weak annotation is significantly less expensive to produce (in time and expertise) than a segmentation mask. A weak annotation corresponds to the marking of one or more pixels of the image. It can correspond to an approximate segmentation (of low precision compared to the segmentations of the majority or minority segmentation masks) or to a simple localization of a structure to be segmented. The use of weak annotations makes it possible to refine the predictions of the pseudo-segmentation masks made by the teaching network.
[0016] Updating the weights of the teacher network using an exponential moving average (EMA) allows the accumulation of different versions of the weights of the segmentation neural network (which plays the role of a student network) over time. The "conservative" predictions of the teacher network thus remain slightly different, which has a regularization effect for the training of the student network (the neural network that we are trying to train) and limits overfitting (in other words, we are trying to avoid the predictions of the student network being too close to certain segmentation masks that are too specific to the supervised training data).
[0017] The combination in the supervised training data of real segmented medical images, synthetic images comprising rare anomalies, and real non-segmented medical images for which a pseudo-segmentation has been generated, makes it possible to obtain very good segmentation performances at the end of the training.
[0018] In particular embodiments, the training method may further comprise one or more of the following characteristics, taken in isolation or in all technically possible combinations.
[0019] In particular modes of implementation, the method comprises, at each iteration: - a calculation of a supervised cost function Cs representative of a dissimilarity between a prediction of the segmentation neural network, made from an image belonging to the supervised training data, and the segmentation mask associated with this image, - a calculation of a semi-supervised cost function representative of a dissimilarity between a prediction of the teaching network and a prediction of the segmentation neural network made for an image belonging to the semi-supervised training data, - a calculation of a global cost function C as a function of the supervised cost function Cs and the semi-supervised cost function CS5, - updating the weights of the segmentation neural network by gradient descent being performed as a function of the global cost function.
[0020] The combination of the supervised cost function with the semi-supervised cost function contributes to the good segmentation performance obtained at the end of the training.
[0021] In particular modes of implementation, a prediction of the teaching network carried out on an image belonging to the semi-supervised training data is enriched with the weak annotation associated with said image. The enriched prediction, denoted $ , can be written in the form = + , expression in which y? is a raw prediction of the teacher network, Vp corresponds to the weak annotation, and fi is a contribution ratio to the weak annotation.
[0022] In particular modes of implementation, the contribution ratio fi of the weak annotation is between 0.45 and 0.55.
[0023] This combination of weak annotations with the predictions of the teaching network makes it possible to refine the generation of pseudo-segmentation masks.
[0024] In particular embodiments, each weak annotation comprises an approximate location of an anatomical structure to be segmented. This approximate location may be represented by one or more points of the anatomical structure to be segmented, a geometric shape covering a region within the anatomical structure to be segmented, or a geometric shape surrounding or covering the anatomical structure to be segmented.
[0025] In particular embodiments, the global cost function C is calculated in the form of a weighted sum of the supervised cost function Cs and the semi-supervised cost function Css.
[0026] In particular modes of implementation, the overall cost function is written in the form C = (ly)Cç + yCss, with between 0.05 and 0.15.
[0027] In particular embodiments, the supervised cost function Cs is calculated in the form of a linear combination of a real cost function CRe calculated from the segmented real medical images and their associated segmentation masks and a synthetic cost function C$yn calculated from the synthetic images and their associated artificial segmentation masks.
[0028] In particular modes of implementation, the global cost function C is written in the form (J — C Re + apC expression in which î is a contribution ratio of the semi-supervised cost function CSs, a is a contribution ratio of the synthetic cost function C$vn, and P is a realism score of the synthetic images calculated using a discriminator neural network forming with the generator neural network a pair of generative adversarial networks.
[0029] In particular embodiments, the contribution ratio a is understood between 0.35 and 0.45 and the contribution ratio T is between 0.05 and 0.15.
[0030] In particular embodiments, obtaining artificial segmentation masks involves a transformation of the segmentation of the anomaly of the minority segmentation mask.
[0031] In particular embodiments, the transformation of the segmentation of the anomaly corresponds to a rotation, an enlargement, a reduction, a deformation and / or a displacement of the segmentation of the anomaly.
[0032] According to a second aspect, a method for automatically segmenting a medical image is proposed. This automatic segmentation method comprises a step of training a segmentation neural network according to any one of the preceding implementation modes, then using the trained segmentation neural network to segment the medical image.
[0033] According to a third aspect, there is provided a computer-readable storage medium comprising a segmentation neural network trained according to any one of the preceding implementation modes. Presentation of figures
[0034] The invention will be better understood on reading the following description, given by way of non-limiting example, and made with reference to Figures 1 to 12 which represent:
[0035] [Fig.l] a schematic representation of the main steps of an implementation mode of a method for training a segmentation neural network according to the invention,
[0036] [Fig.2] a schematic representation of the training of a neural network generator for generating synthetic images from a segmentation mask,
[0037] [Fig.3] a schematic representation of the method for training a network of segmentation neurons according to the invention,
[0038] [Fig.4] a schematic representation of the calculation of a supervised cost function representative of a dissimilarity between a prediction of the segmentation neural network and the segmentation mask associated with the image from which the prediction is made,
[0039] [Fig.5] a schematic representation of the calculation of a semi- cost function supervised representative of a dissimilarity between a prediction of the teaching network and a prediction of the segmentation neural network,
[0040] [Fig.6] a schematic representation of the updating of the weights of the network of segmentation neurons and weights of the teaching network from a global cost function calculated based on the supervised cost function and the semi-supervised cost function,
[0041] [Fig.7] a schematic representation of a step of obtaining real medical images and their associated segmentation masks,
[0042] [Fig.8] a schematic representation of a step of obtaining artificial segmentation masks from majority segmentation masks and minority segmentation masks,
[0043] [Fig.9] a schematic representation of a step of obtaining artificial segmentation masks with transformation of the segmentation of the anomaly,
[0044] [Fig. 10] an illustration of obtaining an artificial segmentation mask from a majority segmentation mask and a minority segmentation mask,
[0045] [Fig. 11] an illustration of a real medical image, its associated segmentation mask, and a synthetic medical image generated by the generator neural network from the segmentation mask, for a majority case (at the top of the figure) and for a minority case (at the bottom of the figure),
[0046] [Fig. 12] an illustration, for three different majority real medical images, of the majority real medical image, its associated majority segmentation mask, an artificial segmentation mask generated from the majority segmentation mask, and a synthetic medical image generated by the generator neural network from the artificial segmentation mask.
[0047] In these figures, identical references from one figure to another designate identical or similar elements. For reasons of clarity, the elements represented are not necessarily on the same scale, unless otherwise stated. Detailed description of the invention
[0048] Figures 1 and 3 schematically represent the main steps of an implementation mode of a method 100 for training a segmentation neural network 30 to segment anatomical structures on a medical image representing an anatomy of interest of a patient.
[0049] In the present application, a “neural network” corresponds to a model, implemented by a computer, whose operation is inspired by that of the neurons of the human brain. It is a variety of deep learning technology, which itself is part of machine learning algorithms. Machine learning algorithms form a category in the field of artificial intelligence.
[0050] The anatomy of interest may correspond to an organ (e.g., the liver, pancreas, gallbladder, lung, or kidney) or another anatomical structure (e.g., a bone or blood vessel). An abnormality within the anatomy of interest generally corresponds to a lesion, such as a tumor, cyst, ablation zone, aneurysm, etc. An ablation zone corresponds to a lesion that has undergone ablation treatment using a known method (microwave, laser, radiofrequency, etc.), it is a necrotic region. The anomaly may also correspond to a surgical artifact such as part of a medical instrument.
[0051] The segmentation neural network 30 is trained to take as input a real medical image and to provide as output a segmentation mask associated with this image. A segmentation mask is an image in which each voxel provides particular information on an element represented at the position of the voxel on the real medical image. A voxel can thus take a particular numerical value associated with the element represented at the position of the voxel on the real medical image. Different specific numerical values are for example defined respectively for a healthy part of the anatomy of interest, for an anomaly within the anatomy of interest, for other anatomical structures (for example bones or blood vessels), for the background of the image, etc. Different numerical values thus represent different classes of elements visible on the real medical image.It should be noted that the term "voxel" is used generically to define a particular area of an image (a voxel identifies a position of said area on the image and takes a value representative of what is represented in said area on the image). It can be a two-dimensional or three-dimensional image. If it is a two-dimensional image, the term "voxel" then takes on the same meaning as the term "pixel".
[0052] The term "real medical image" means a medical image of a patient acquired using a medical imaging device, for example by computed tomography (CT), by positron emission tomography (PET), by magnetic resonance imaging (MRI), by ultrasound, or by X-rays.
[0053] The segmentation neural network 30 takes for example the form of a U-Net type convolutional neural network. The segmentation neural network 30 can also be based on a model derived from the U-Net architecture, such as for example the RITnet model which combines the U-Net and DenseNet models (see “RITnet: Real-time Semantic Segmentation of the Eye for Gaze Tracking”, AK Chaudhary et al.). The segmentation neural network 30 could however also be based on other models. It could for example be a transformer neural network such as the “Swin Transformer” network (see for example “Swin UNETR: Swin Transformers for Semantic Segmentation of Brain Tumors in MRI Images” A. Ha-tamizadeh et al.).
[0054] As illustrated in Figures 1 and 3, the method 100 comprises a step of obtaining 101 of segmented real medical images 11, 12 each associated with a segmentation mask 21, 22. These real medical images comprise: • majority real medical images 11 associated with majority segmentation masks 21 and representing the anatomy of interest of a patient without anomaly, • minority real medical images 12 associated with minority segmentation masks 22 and representing the anatomy of interest 41 of a patient with an anomaly.
[0055] The terms "majority" and "minority" are used because there are generally a significantly greater number of medical images depicting anatomy of interest without abnormality than medical images depicting anatomy of interest with abnormality.
[0056] [Fig.7] schematically represents the step 101 of obtaining real medical images 11, 12 and their associated segmentation masks 21, 22.
[0057] On the left of [Fig.7] is illustrated the generation of majority segmentation masks 21 from majority real medical images 11. The anatomy of interest 41 is visible on a majority real medical image 11. On a so-called “majority” real medical image, there is no anomaly within the anatomy of interest. A majority segmentation mask 21 comprises a segmentation of the anatomy of interest.
[0058] On the left of [Fig.7] is illustrated the generation of minority segmentation masks 22 from minority real medical images 12. The anatomy of interest 41 and an anomaly 42 present within the anatomy of interest 41 are visible on a minority real medical image 12. A minority segmentation mask 22 comprises a segmentation of the anatomy of interest and a segmentation 44 of the anomaly.
[0059] Segmentation can be implemented manually by a medical expert. In this case, the medical expert himself defines the contours of the different anatomical regions on the medical image using a graphical interface (mouse, stylus, touch screen, etc.) of an electronic device (computer, tablet, etc.) on which the image is displayed. Segmentation can also be implemented semi-automatically, using a conventional image segmentation algorithm. For example, bones visible on medical images can be segmented by the voxel intensity thresholding method with an intensity window of width 1800 HU and center 400 HU (HU is the English acronym for "Hounsfield Unit", it is a quantitative scale describing radio-density, i.e. a unit of measurement representative of the opacity of a material to X-rays).
[0060] As illustrated in Figures 1 and 3, the method 100 comprises: - a step 102 of training a generator neural network 31 to generate a synthetic image 13 from a segmentation mask, - a step 103 of obtaining artificial segmentation masks 23, - a step of obtaining 104 synthetic images 13, from the masks of artificial segmentations 23, using the trained generator neural network 31.
[0061] [Fig.2] schematically represents an example of implementation of the training step 102 of the generator neural network 31. The generator neural network 31 is trained to generate a synthetic image from a segmentation mask associated with a real medical image. The majority 11 and minority 12 real medical images and their associated segmentation masks 21, 22 can be used to train the generator neural network 31. It should however be noted that nothing would prevent the generator neural network 31 from being trained with segmentation masks other than the majority 21 and minority 22 segmentation masks.
[0062] Different types of neural networks can be considered to generate a synthetic image from a segmentation mask. For example, it is possible to use a neural network of the auto-encoder type ("Variational Au-toencoder" or VAE in the English literature). However, it is preferable, as illustrated in [Fig.2], to use a generator of a pair of generative adversarial neural networks (GAN). These types of neural networks make it possible to generate images with a high degree of realism. A GAN is a generative model where two neural networks are placed in competition in a zero-sum game scenario. The first network, the generator 31, generates an image, and its adversary, the discriminator 32, tries to detect whether the generated image is real or whether it is a synthetic image generated by the generator 31.
[0063] In a first step, the generator 31 is trained to generate synthetic images from the segmentation masks. In a second step, the synthetic images generated by the generator are analyzed by the discriminator 32 which has been previously trained to recognize, by taking as input an image and its associated segmentation mask, whether the pair formed by the image and the segmentation mask is real. There is therefore at the output of the discriminator 32 a “True” or “False” decision depending on whether the pair formed by the image and the segmentation mask is considered real or not. By a backpropagation loop, depending on the veracity of the decision taken by the discriminator 32, the parameters of the generator 31 are modified until the generated synthetic images are considered real by the discriminator 32.
[0064] Generator 31 is a neural network for translating one image into another image. For example, this is a convolutional neural network of the “pix2pix” type as described in the document “Image-to-Image translation with conditional adversial networks” by Isola, P et al. In the example considered, the neural network comprises a first part of the encoder type composed of “Batch Normalization Leaky ReLU” convolution layers with 4x4 convolution filters of sizes 64, 128, 256, 512, 512, 512, 512 and a second part of the decoder type composed of “Batch Normalization Dropout ReLU” convolution layers with 4x4 convolution filters of sizes 512, 512, 512 and then of “Batch Normalization ReLU” convolution layers with 4x4 convolution filters of sizes 256, 128, 64.The image size reduction in the encoder part is produced by strides of two pixels and the image size increase in the decoder part is produced by a 2D upsampling layer (Nearest Neighbors method) of size 2x2. The output is produced by a hyperbolic tangent (Tanh) activation layer.
[0065] The discriminator 32 is a neural network. For example, it is a convolutional neural network of the “PatchGAN” type as described in the document “Image-to-Image translation with conditional adversial networks” by Isola, P et al, modified to accept two input images which are concatenated into a single image. The rest of the neural network consists of a “Leaky ReLU” convolution layer with 4x4 convolution filters of size 64 then four “Batch Normalization Leaky ReLU” convolution layers with 4x4 convolution filters of sizes 128, 256, 512, 512. The output is produced using a sigmoid activation layer.
[0066] The neural network 31 for translating an image into another image (“pix2pix”) is combined with the PatchGAN type discriminator 32 in such a way that the output predictions of the generator 31 (the synthetic medical images) constitute the second input of the discriminator 32. The first input of the discriminator corresponds to the medical image associated with the segmentation mask which is provided as input to the generator 31. The output is a 70x70 probability matrix. The weights of the neurons of the discriminator 32 cannot be modified during the training of the generator 31. The weights of the generator 31 can be updated during training. The cost calculation function is composed of the cross entropy and the norm 1 in a ratio of 1 to 100.
[0067] The discriminator 32 and the generator 31 are trained alternately, in turn, on a training set. The output of the discriminator 32 is optimized using the Adam-type stochastic gradient algorithm (Adaptive Moment Estimation, beta_l: 0.9, beta_02: 0.999, epsilon: le-08) against a 70x70 matrix of 1 values when its input is a “true” pair (i.e. a pair including a mask segmentation mask and its associated real medical image), and of value 0 when its input is a "false" pair (i.e. a pair comprising a segmentation mask and a synthetic medical image produced by the generator 31 from the segmentation mask). The output of the generator 31 is optimized using the Adam-type stochastic gradient algorithm against a 70x70 matrix of values 1 so that the weights of the neurons of the generator 31 are updated, but not those of the discriminator 32, when the discriminator 32 detects that the input pair is not sufficiently close to the "real" pairs already encountered. Updating the weights of the discriminator 32 alternately allows it to be ahead of the generator 31 and force it to update.
[0068] The method 100 also comprises a step 103 of obtaining artificial segmentation masks 23 from the majority segmentation masks 21 and the minority segmentation masks 22. As illustrated in [Fig.8], this step 103 may comprise: - a selection of a majority segmentation mask 21, - a selection of a minority segmentation mask 22, - an identification, on the selected minority segmentation mask 22, of a set of voxels whose numerical value encodes the anomaly, - a replacement, on the selected majority segmentation mask 21, of the numerical value of the identified voxels by the numerical value encoding the anomaly.
[0069] The majority segmentation mask 21 and the minority segmentation mask 22 can be selected randomly. The number of combinations of artificial segmentation masks can thus be equal to the product of the number of majority segmentation masks 21 by the number of minority segmentation masks 22.
[0070] Several minority segmentation masks 21 with different semantic values (i.e. with different types of anomaly: tumor, cyst, ablation region, artifact, etc.) can be combined with the majority segmentation masks 22 and thus increase the number of potential combinations. Manipulating the proportions of the different minority masks 21 that will be used in the combinations of artificial segmentation masks 23 makes it possible to control the characteristics of the synthetic images produced. Similarly, several majority masks 22 with different semantic values (liver, lung, pancreas, gallbladder, etc.) can be combined with the minority segmentation masks 21. These different majority masks 22 make it possible to control in which organ or structure the minority segmentation masks 21 can appear using Boolean logic rules (for example rules of the “AND”, “OR” type, “NOT”) between the different majority masks 22 and minority masks 21 for each pixel of the artificial segmentation masks 23. This characteristic makes it possible to control, for example, in which organ the random distribution of a minority segmentation mask 21 of a tumor can appear (“AND”) and in which structure or organ this minority segmentation mask 21 cannot appear (“NOT”).
[0071] The combinations of majority segmentation mask 21 and minority segmentation mask 22 of greatest interest can also be selected according to an estimated distance between a segmentation of the anomaly 25 on the minority segmentation mask 22 and a segmentation of interest of the majority segmentation mask 21, such as for example the segmentation of the gallbladder, the vessels, the hilum or even the liver capsule.
[0072] As illustrated in [Fig.9], the step 103 of obtaining artificial segmentation masks 23 may also comprise a transformation of the segmentation 44 of the anomaly of the selected minority segmentation mask 22. Such arrangements make it possible to generate a greater number of different artificial segmentation masks.The transformation of the segmentation 44 of the anomaly corresponds for example to a rotation (as illustrated for the artificial segmentation mask 23-1), a displacement (as illustrated for the artificial segmentation mask 23-2), an enlargement (as illustrated for the artificial segmentation mask 23-3), a reduction (as illustrated for the artificial segmentation mask 23-4), a deformation (as illustrated for the artificial segmentation mask 23-5) or a combination of these different possible transformations (as illustrated for the artificial segmentation mask 23-6 for which the segmentation 44 of the anomaly has been simultaneously displaced, reduced, and rotated).
[0073] [Fig. 10] illustrates the generation 103 of an artificial segmentation mask 23 from a majority segmentation mask 21 and a minority segmentation mask 22. On the majority segmentation mask 21, the segmentation 43 of the anatomy of interest can be seen. On the minority segmentation mask 22, both the segmentation 43 of the anatomy of interest and the segmentation 44 of the anomaly can be seen. On the generated artificial segmentation mask 23, the segmentation 43 of the anatomy of interest of the majority segmentation mask 21 can be seen combined with the segmentation 44 of the anomaly of the minority segmentation mask 22 (in the example considered, the segmentation 44 of the anomaly has also been moved).
[0074] This step of generating 103 the artificial segmentation masks 23 is implemented by a computer. The computer comprises a memory with a set of program code instructions which, when the program is executed, configure one or more processors of the computer to generate, as described above with reference to Figures 8 to 10, artificial segmentation masks 23 from the majority segmentation masks 21 and the minority segmentation masks 22.
[0075] As illustrated in Figures 1 and 3, the method 100 comprises a step 104 of obtaining synthetic images 13, from the artificial segmentation masks 23, using the trained generator neural network 31.
[0076] [Fig. 11] is an illustration of a real medical image 11 (respectively 12), of its associated segmentation mask 21 (respectively 22), and of a synthetic medical image 13 generated by the trained neural network 31, from the segmentation mask 21 (respectively 22), for a majority case without anomaly (respectively for a minority case with anomaly).
[0077] [Fig. 12] illustrates by way of example, for three different majority real medical images 11: the majority real medical image 11, its associated majority segmentation mask 21, an artificial segmentation mask 23 generated from the majority segmentation mask 21, and a synthetic medical image 13 generated from the artificial segmentation mask 23 by the trained generator neural network 31.
[0078] It is thus possible to generate a very large diversity of artificial segmentation masks 23. This then makes it possible to generate a very large diversity of synthetic medical images 13 presenting anatomical anomalies.
[0079] It is possible to train the segmentation neural network 30 from the supervised training data 25 comprising on the one hand the real segmented medical images 11, 12 and their associated segmentation masks 21, 22, and on the other hand the synthetic images 13 and their associated artificial segmentation masks 23. The supervised training of the segmentation neural network 30 corresponds to step 107 in FIGS. 1 and 3.
[0080] The solution proposed in the present application goes further and proposes to further enrich the supervised training data 25.
[0081] For this, the method 100 comprises a step 105 of obtaining real non-segmented medical images 14 each representing the anatomy of interest 41 of a patient and each being associated with a weak annotation. The non-segmented medical images 14 may comprise images with anomaly and / or images without anomaly.
[0082] A weak annotation may correspond to an approximate segmentation (of low precision compared to the aforementioned majority or minority segmentations) or to a simple localization of a structure to be segmented. This approximate localization may be materialized by one or more points of the anatomical structure to be segmented, by a geometric shape covering a region to be segmented, or by a geometric shape covering a region to be segmented. inside the anatomical structure to be segmented, or by a geometric shape surrounding or covering the anatomical structure to be segmented (bounding box). The case where a weak annotation is materialized by a geometric shape covering a region inside the anatomical structure to be segmented is particularly advantageous. Weak annotations can advantageously be generated or at least validated by a medical expert.
[0083] The real unsegmented medical images 14 and their associated weak annotations form semi-supervised training data 26.
[0084] As illustrated in Figures 1 and 3, the method 100 comprises an identical reproduction 106 (a cloning) of the segmentation neural network 30 to form a teaching network 30'. Thus, the teaching network 30' is initialized with weights equal to those of the segmentation neural network 30 at the time of the reproduction 106.
[0085] As will be seen later, thanks to the use of weak annotations, it is not necessary to train the segmentation neural network 30 before cloning it to form the teaching network 30'. The weak annotations in fact make it possible to initiate the segmentation even from a naive teaching network (little or not trained).
[0086] The weights of the segmentation neural network 30 (which plays the role of a student network) and the weights of the teacher network 30' are updated iteratively. Each iteration corresponds to a training phase based on the use of a batch of images among the training data.
[0087] At each iteration, the weights of the segmentation neural network 30 are updated by gradient descent. The weights of the teaching network 30' are updated by moving exponential average as a function of the weights of the teaching network 30' calculated at the previous iteration and the weights of the segmentation neural network 30 calculated at the current iteration.
[0088] The teaching network 30' can then be used to generate pseudo-segmentation masks 24 which correspond to predictions of the teaching network 30' obtained from semi-supervised training data 26. The pseudo-segmentation masks 24 thus generated can then be used as supervised training data to continue the training 107 of the segmentation neural network during a following iteration.
[0089] Each iteration therefore includes:
[0090] - a supervised training step 107 of the segmentation neural network 30 from supervised training data 25,
[0091] - a step 108 of updating the weights of the segmentation neural network 30 by gradient descent,
[0092] - a step of updating 109 of the weights of the teaching network 30' by average expo mobile nential,
[0093] - a step 110 of generation of pseudo-segmentation masks 24 by the network teacher 30',
[0094] - a step 111 of adding to the supervised training data 25 of the pseudo 24 generated segmentation masks and 14 unsegmented real medical images which allowed their generation 110.
[0095] It should be noted that the above steps could be performed in a different order (i.e., it is not essential to perform the steps listed above in the proposed order). Nothing would prevent, for example, for a current iteration, updating the weights of the teaching network 30' by moving exponential average before updating the weights of the segmentation neural network 30 by gradient descent (in this case, the updating of the weights of the teaching network 30' by moving exponential average is done according to the weights of the segmentation neural network 30 updated in the previous iteration).
[0096] As illustrated in FIG. 4, the supervised training 107 comprises a calculation of a supervised cost function Cs representative of a dissimilarity between a prediction y of the segmentation neural network 30, carried out from an image belonging to the supervised training data 25, and the segmentation mask Ty associated with this image.
[0097] The supervised cost function Cs is for example calculated according to the Sorensen-Dice index:
[0098] _ s~
[0099] At the same time, and as illustrated in FIG. 5, the method 100 may also comprise a calculation of a semi-supervised cost function Css representative of a dissimilarity between a prediction ÿ of the teaching network 30' and a prediction y^ of the segmentation neural network 30 carried out for an image belonging to the semi-supervised training data 26.
[0100] The semi-supervised cost function Css is for example calculated by binary cross entropy:
[0101] Css = 4(1 -yT)log(1 --yTlog(^)]
[0102] The predictions of the teaching network 30' are a source of truth for the calculation of the semi-supervised cost function (just as the segmentation masks 21, 22 are for the calculation of the supervised cost function).
[0103] As illustrated in Figure 6, a global cost function C can then be calculated based on the supervised cost function Cs and the semi-supervised cost function Qs.
[0104] The global cost function C can in particular be calculated in the form of a weighted sum of the supervised cost function Cs and the semi-supervised cost function Css. For example, the global cost function can be written in the form:
[0105] C = (l <r)Cs + yC^
[0106] The combination of the supervised cost function Cs with the semi-supervised cost function Css contributes to the good segmentation performance obtained at the end of the training.
[0107] It has been empirically observed that an optimal value for the coefficient V is between 0.05 and 0.15.
[0108] The update 108 of the weights of the segmentation neural network 30 by gradient descent can then be carried out according to a result of the global cost function C,
[0109] The weights wt of the teaching network 30' are updated in the form of an exponential moving average calculated as a function of the previous weights wTprec of the teaching network 30' and the updated weights of the segmentation neural network 30.
[0110] As illustrated in Figure 6, the weights wt of the teaching network 30' are for example calculated in the form: [YES] wT = ÔWTpree + ( 1 - Ô)
[0112] In this expression, the term <5 is a parameter whose value is predetermined. It has been observed, empirically, that an optimal value for parameter 5 is between 0.85 and 0.95.
[0113] Updating 109 the weights wt of the teacher network 30' using an exponential moving average makes it possible to accumulate the different versions of the weights of the student network over time. The conservative predictions of the teacher network thus remain slightly different; this has a regularization effect for the training of the student network (i.e. for the segmentation neural network 30 that we are seeking to train), and this makes it possible to limit overfitting.
[0114] It is also important to note that, still with the aim of limiting overtraining, the supervised training data 25 are not consumed by the teaching network 30'. Thus, the weights of the teaching network 30' are updated from a training data set different from the data that the teaching network 30' consumes to generate the pseudo-annotations 24. This limits overtraining.
[0115] Advantageously, and as illustrated in [Fig.5], a prediction of the teaching network 30' carried out on an image belonging to the semi-supervised training data 26 is enriched with the weak annotation associated with said image (the prediction obtained is merged with the weak annotation).
[0116] The enriched prediction, denoted A, can be written in the form: ■'T 101171 Ï'T= (l-0)yT+llyp
[0118] In this expression, y is a raw prediction of the teacher network 30', ^'f corresponds to the weak annotation, and is a contribution ratio of the weak annotation. It has been observed that an optimal value for the contribution ratio P of the weak annotation is between 0.45 and 0.55. The enriched prediction is used instead -T of the prediction ÿ in the calculation of the semi-supervised cost function Css.
[0119] As indicated previously, the weak annotations correspond to location information of the structures to be segmented, this information being produced by medical experts. Combining the predictions of the teaching network 30' with the weak annotations thus makes it possible to guide the training of the segmentation on valid structures. This makes it possible to refine the generation of the pseudo-segmentation masks by the teaching network 30'.
[0120] In the example considered, the probability maps formed by the combination C of "T the prediction ÿ of the teaching network 30' and the weak annotation are binarized at a threshold of value 0.5.
[0121] The supervised training 107 of the student network 30 then makes it possible to update the teacher network 30' which in turn can produce more precise pseudo-segmentation masks 24 thanks to the weak annotations.
[0122] The supervised cost function Cs can be calculated as a linear combination of the following: - a real cost function CRe calculated from the real segmented medical images 11, 12 and their associated segmentation masks 21, 22, - a synthetic cost function C$yn calculated from the synthetic images 13 and their associated artificial segmentation masks 23.
[0123] The overall cost function C can then be written in the form:
[0124] cs = (1 - y) (CRe + apCSyfl) + yCss
[0125] In this expression, f is a contribution ratio of the semi-supervised cost function Css, a is a contribution ratio of the synthetic cost function Cgm, and P is a realism score of the synthetic images 13 calculated using the discriminator neural network 32 forming with the generator neural network 31 a pair of generative adversarial networks. The more confident the discriminator 32 is that the synthetic images produced by the generator 31 are real, the more these synthetic images contribute to the total cost function. It has been observed that a optimal value for contribution ratio a is between 0.35 and 0.45. It has also been observed that an optimal value for contribution ratio T is between 0.05 and 0.15.
[0126] The training of the segmentation neural network 30 can then continue until the convergence of the model by cross-validation. The predictions of the segmentation neural network 30 are binarized using a threshold to produce the segmentation masks of the images to be segmented.
[0127] Each of the different artificial intelligence algorithms (the segmentation neural network 30, the teaching network 30', the generator neural network 31, and the discriminator neural network 32) is implemented by a computer. It is conceivable to use a single computer to implement all of these algorithms. Alternatively, some of these algorithms may be implemented on separate computers.
[0128] Once the segmentation neural network 30 is sufficiently trained, it can be used in a method for automatically segmenting a medical image. The trained segmentation neural network 30 can be stored on a computer-readable storage medium. One or more processors of the computer can then be configured to take as input a medical image to be segmented, and to produce as output a segmentation mask associated with the medical image using the trained segmentation neural network 30.
[0129] The above description clearly illustrates that, through its various characteristics and their advantages, the present invention achieves the set objectives.
[0130] In particular, the construction of the artificial segmentation masks 23 and the generation of the associated synthetic images 13 makes it possible to limit the over-learning of the segmentation neural network 30. Indeed, it is possible to construct an infinity of different artificial segmentation masks 23 by a random combination of the majority segmentation masks 21 with the minority segmentation masks 22, and by a transformation of the segmentation 44 of an anomaly of a minority segmentation mask 22 before its combination with a majority segmentation mask 21. Each synthetic image 13 thus generated can then be used only once in the training of the segmentation algorithm.
[0131] The majority 21 and minority 22 segmentation masks, as well as the weak annotations, are preferably obtained by medical experts. This makes it possible to guarantee a sufficient level of involvement of medical expertise in the training of the segmentation neural network.
[0132] Weak annotations, while being significantly less expensive to produce than a segmentation mask, make it possible to refine the predictions of pseudo-masks. segmentation by the teacher network 30'.
[0133] The pseudo-segmentation masks generated by the teaching network 30' make it possible to enrich the supervised training data set 25. The supervised training data 25 can thus contain a very large number of different images.
[0134] The fact that the teaching network 30' does not consume the supervised training data 25 also makes it possible to limit overlearning.
[0135] The combination in the supervised training data of real segmented medical images 11, 12, of synthetic images 13 comprising rare anomalies, and of real non-segmented medical images 14 for which a pseudo-segmentation has been generated, makes it possible to obtain very good segmentation performances at the end of the training.
Claims
Claims
1. A method (100) of training a segmentation neural network (30) to segment anatomical structures on a medical image representing an anatomy of interest (41) of a patient, the method (100) comprising the following steps: - obtaining (101) real segmented medical images (11, 12), each associated with a segmentation mask (21, 22), among which: • majority real medical images (11) associated with majority segmentation masks (21) and representing the anatomy of interest (41) of a patient without anomaly, • minority real medical images (12) associated with minority segmentation masks (22) and representing the anatomy of interest (41) of a patient with an anomaly (42), - training (102) a generator neural network (31) to generate a synthetic image from a segmentation mask, - obtaining (103) artificial segmentation masks (23), each artificial segmentation mask (23) being obtained by combining a majority segmentation mask (11) with a segmentation (44) of an anomaly of a minority segmentation mask (12), - obtaining (104) synthetic images (13), from the artificial segmentation masks (23), using the trained generator neural network (31), - obtaining (105) real non-segmented medical images (14) each representing the anatomy of interest (41) of a patient and each being associated with a weak annotation, - an identical reproduction (106) of the segmentation neural network (30) to form a teaching network (30'), the method further comprising several iterations of the steps following: - supervised training (107) of the segmentation neural network (30) from supervised training data (25) comprising the segmented real medical images (11, 12) and their associated segmentation masks (21, 22), as well as the synthetic images (13) and their associated artificial segmentation masks (23), - an update (108) of the weights M s of the segmentation neural network (30) by gradient descent, - an update (109) of the weights wt of the teaching network (30') by an exponential moving average calculated according to the previous weights wTprec of the teaching network (30') and the updated weights M7> of the segmentation neural network (30), - a generation (110) of pseudo-segmentation masks (24) corresponding to predictions of the teaching network (30') obtained from semi-supervised training data (26) comprising the real non-segmented medical images (14) and their associated weak annotations, - an addition (111) to the supervised training data (25) of the pseudo-segmentation masks (24) generated and of the real non-segmented medical images (14) which allowed their generation (110).
2. Method (100) according to claim 1, comprising, at each iteration: - a calculation of a supervised cost function Cs representative of a dissimilarity between a prediction of the segmentation neural network (30), made from an image belonging to the supervised training data (25), and the segmentation mask associated with this image, - a calculation of a semi-supervised cost function Css representative of a dissimilarity between a prediction of the teaching network (30') and a prediction of the segmentation neural network (30) carried out for an image belonging to the semi-supervised training data (26), - a calculation of a global cost function C as a function of the supervised cost function Cs and the semi- supervised Css, - the update (108) of the weights of the segmentation neural network (30) by gradient descent being carried out according to the global cost function.
3. Method (100) according to claim 2 in which a prediction of the teaching network (30') carried out on an image belonging to the semi-supervised training data (26) is enriched with the weak annotation associated with said image, the enriched prediction, noted , being able to be written in the form ( j _ / J ) + flv ■> expression in which y is a raw prediction of the teaching network (30'), corresponds to the weak annotation, and P is a contribution ratio to the weak annotation.
4. The method (100) of claim 3 wherein the contribution ratio of the weak annotation is between 0.45 and 0.
55.
5. Method (100) according to any one of claims 2 to 4 in which each weak annotation comprises an approximate location of an anatomical structure to be segmented, said approximate location being materialized by one or more points of the anatomical structure to be segmented, a geometric shape covering a region inside the anatomical structure to be segmented, or a geometric shape surrounding or covering the anatomical structure to be segmented.
6. Method (100) according to any one of claims 2 to 5 wherein the global cost function C is calculated as a weighted sum of the supervised cost function Cs and the semi-supervised cost function Css.
7. Method (100) according to claim 6 in which the overall cost function is written in the form C = (l -y)C+ yCss, with f between 0.05 and 0.
15.
8. Method (100) according to any one of claims 2 to 7 in which the supervised cost function Cs is calculated in the form of a linear combination of a real cost function CRe calculated from the segmented real medical images (11, 12) and their associated segmentation masks (21, 22) and a synthetic cost function CSyn calculated from the synthetic images (13) and their associated artificial segmentation masks (23).
9. Method (100) according to claim 8 in which the global cost function C is written in the form C = ( 1 - y ) ( CRe + apCSyn ) + yCs# expression in which Y is a contribution ratio of the semi-supervised cost function a is a contribution ratio of the synthetic cost function Csyn, and P is a realism score of the synthetic images (13) calculated using a discriminator neural network (32) forming with the generator neural network (31) a pair of generative adversarial networks.
10. The method (100) of claim 9 wherein the contribution ratio a is between 0.35 and 0.45 and the contribution ratio Y is between 0.05 and 0.
15.
11. Method (100) according to any one of claims 1 to 10 in which the obtaining (103) of artificial segmentation masks (23) comprises a transformation of the segmentation (44) of the anomaly of the minority segmentation mask.
12. Method (100) according to claim 11 wherein the transformation of the segmentation (44) of the anomaly corresponds to a rotation, an enlargement, a reduction, a deformation and / or a displacement of the segmentation (44) of the anomaly.
13. Method for automatic segmentation of a medical image comprising: - training a segmentation neural network (30) according to any one of claims 1 to 12, - using the trained segmentation neural network (30) to segment the medical image.
14. A computer-readable storage medium comprising a segmentation neural network (30) trained according to any one of claims 1 to 12.