DEVICE AND METHOD FOR GENERATING n VIRTUAL IMMUNOHISTOCHEMICAL (IHC) STAIN IMAGES FROM ONE HEMATOXYLIN AND EOSIN (H&E) STAIN IMAGE

The device generates virtual IHC stain images using a machine learning system trained on unpaired H&E and IHC datasets, addressing the limitations of H&E and IHC staining by providing efficient and cost-effective protein visualization and quantification.

WO2025168731A1PCT designated stage Publication Date: 2025-08-14INST DU CERVEAU & DE LA MOELLE EPINIERE ICM +5
View PDF 1 Cites 0 Cited by

Patent Information

Application Number
PCT/EP2025/053153
Authority / Receiving Office
WO · WO
Patent Type
Applications
Current Assignee / Owner
Priority Date
2024-02-09
Filing Date
2025-02-06
Publication Date
2025-08-14

AI Technical Summary

Technical Problem

H&E staining provides limited information on protein expression within cells, requiring additional and resource-intensive IHC staining, which is costly and damaging to tissues.

Method used

A device and method for generating virtual IHC stain images using a machine learning system trained on unpaired datasets of H&E and IHC images, allowing visualization and quantification of proteins without physical staining, reducing time and cost.

Benefits of technology

Enables efficient, non-destructive visualization and quantification of proteins in tissues, overcoming the limitations of H&E and IHC staining by generating virtual IHC images quickly and cost-effectively.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure EP2025053153_14082025_PF_FP_ABST
    Figure EP2025053153_14082025_PF_FP_ABST
Patent Text Reader

Abstract

The invention also related to a device and a computer-implemented method for generating n virtual Immunohistochemical (IHC) stain images from one Hematoxylin and Eosin (H&E) stain image, using the IHC machine learning system obtained thanks to the device for training.
Need to check novelty before this filing date? Find Prior Art

Description

DEVICE AND METHOD FOR GENERATING n VIRTUAL IMMUNOHISTOCHEMICAL (IHC) STAIN IMAGES FROM ONE HEMATOXYLIN AND EOSIN (H&E) STAIN IMAGEFIELD OF INVENTION

[0001] The present invention relates to the technical field of image processing and more particularly to a device for generating a IHC machine learning system for generating n virtual Immunohistochemical (IHC) stain images from one Hematoxylin and Eosin (H&E) stain image. The invention also relates to a device and method for generating n virtual Immunohistochemical (IHC) stain images from one Hematoxylin and Eosin (H&E) stain image, using a IHC machine learning system trained using the training parameters obtained thanks to said device for training.BACKGROUND OF INVENTION

[0002] Hematoxylin and Eosin (H&E) staining is a widely used histological staining technique in the field of pathology and microscopy. It is commonly employed to visualize the cellular and tissue structures of biological samples under a microscope. More precisely, Hematoxylin is a blue-purple dye that stains cell nuclei and other acidic structures in tissues, and Eosin is a pink-red dye that stains the cytoplasm and extracellular matrix of cells.

[0003] While H&E staining is a valuable staining technique in histology and pathology, it does have its limitations. Indeed, H&E staining provides information about tissue and cell morphology but it does not offer insights into the expression of specific proteins within cells. To study proteins, additional staining methods are often required.

[0004] Immunohistochemistry (IHC) staining is particularly adapted for visualizing and quantifying the presence, localization, and distribution of specific proteins in tissues. Notably, IHC staining is particularly adapted for the classification of various tumor types. It is also key in pinpointing the origin of metastatic tumors. In addition, it can revealminute tumor cells that might escape detection through standard staining procedures. In particular, this technique proves particularly beneficial in the diagnosis of diseases that traditional biopsy cultures and serological diagnoses struggle to detect.

[0005] IHC staining involves the use of antibodies that selectively bind to the target protein of interest. Antibodies become detectable by using secondary antibodies or antibody-conjugated markers, labeled or tagged with a reporter molecule (e.g., enzyme, fluorescent dye, or radioactive isotope).

[0006] Despite these strengths, IHC procedures are not either without limitations. Indeed, IHC staining requires considerable resources in terms of time and money for sample preparation, and expert oversight, raising the probability of errors and delays that could potentially influence disease diagnosis and treatment. Moreover, IHC staining is toxic and may damage the tissues to the point that further analysis of the same tissue becomes impossible.

[0007] The shortcomings of both H&E and IHC staining techniques underscore the pressing need for an improved way of analyzing tissues, that is less expensive than IHC staining and that does not result in the destruction of the analyzed tissues.SUMMARY

[0008] This invention thus relates to a device for generating a IHC machine learning system for generating n virtual Immunohistochemical (IHC) stain images from one Hematoxylin and Eosin (H&E) stain image, said device comprising: at least one input configured to receive a training dataset of unpaired images comprising: o a H&E stain images ensemble comprising a plurality of H&E stain images, and o n IHC stain images ensembles, wherein each ithIHC stain images ensemble comprises a plurality of IHC stain images obtained using an ithIHC stain, and wherein the ithIHC stain used to obtain the images of the ithIHC stain images ensemble is different from the IHC stain used to obtain the images of the others IHC stain images ensembles;at least one processor configured to train a trainable machine learning system to obtain training parameters, wherein: o said trainable machine learning system comprises:■ n IHC discriminative models and a IHC generative model comprising one H&E encoder and n IHC decoders, and■ one H&E discriminative model and a H&E generative model comprising n IHC encoders and one H&E decoder, and o during the training:■ in a first step, the H&E encoder from said IHC generative model is fed with the images from the H&E stain images ensemble from said training dataset and said IHC generative model outputs a first set of n virtual IHC stain images, wherein each ithIHC decoder of the IHC generative model provides as an output a ithvirtual IHC stain image of the n virtual IHC stain images, the ithIHC discriminative model receiving as an input said ithvirtual IHC stain image and the IHC stain images from the ithIHC stain images ensemble from the training dataset; each ithIHC encoder from the H&E generative model also receiving as input the ithvirtual IHC stain image outputted by the ithIHC decoder; the H&E decoder from said H&E generative model providing as an output a first set of n virtual H&E stain images, the H&E discriminative model receiving as an input the first set of n virtual H&E stain images outputted by the H&E generative model and the H&E stain images ensemble from said training dataset;■ in a second step, each ithIHC encoder from the H&E generative model receives as an input the IHC stain images from the ithIHC stain images ensemble from the training dataset and the ithvirtual IHC stain image outputted by the ithIHC decoder, and the H&E decoder from said H&E generative model provides as an output a second set of n virtual H&E stain images, the H&E discriminative model receiving as an input the second set of nvirtual H&E stain images outputted by the H&E generative model and the H&E stain images ensemble from said training dataset; the H&E encoder from said IHC generative model also receiving as an input the second set of n virtual H&E stain images, said IHC generative model providing as an output a second set of n virtual IHC stain images, wherein each ithIHC decoder of the n IHC decoders provides as an output a ithvirtual IHC stain image of the second set of n virtual IHC stain images, the ithIHC discriminative model receiving as an input the ithvirtual IHC stain image outputted by the ithIHC decoder and the IHC stain images from the ithIHC stain images ensemble from the training dataset; at least one output configured to provide the obtained training parameters for said IHC machine learning system, wherein said IHC machine learning system comprises the H&E encoder, the n IHC decoders and the n IHC discriminative models.

[0009] In other words, the present invention relates to a device for generating a IHC machine learning system for generating n virtual Immunohistochemical stain images from one Hematoxylin and Eosin stain image, said device comprising: at least one input configured to receive a training dataset of unpaired images comprising: o a H&E stain images ensemble comprising a plurality of H&E stain images, and o n IHC stain images ensembles, wherein each ithIHC stain images ensemble comprises a plurality of IHC stain images obtained using an ithIHC stain, and wherein the ithIHC stain used to obtain the images of the ithIHC stain images ensemble is different from the IHC stain used to obtain the images of the others IHC stain images ensembles; at least one processor configured to train a trainable machine learning system to obtain training parameters, wherein: o said trainable machine learning system comprises:■ n IHC discriminative models and a IHC generative model comprising one H&E encoder and n IHC decoders, and■ one H&E discriminative model and a H&E generative model comprising n IHC encoders and one H&E decoder, andng the training:■ the H&E encoder from said IHC generative model is fed with each image of the H&E stain images ensemble from said training dataset (e.g., sequentially) and said IHC generative model outputs, for each of said image of the H&E stain images ensemble, a first set of n virtual IHC stain images, wherein each ithIHC decoder of the IHC generative model provides as an output a ithvirtual IHC stain image of the n virtual IHC stain images, each of the ithIHC discriminative model receiving as an input said ithvirtual IHC stain image and the IHC stain images from the ithIHC stain images ensemble from the training dataset; each ithIHC encoder from the H&E generative model receiving as input the ithvirtual IHC stain image outputted by the ithIHC decoder; the H&E decoder from said H&E generative model providing as an output a first set of n virtual H&E stain images, the H&E discriminative model receiving as an input said first set of n virtual H&E stain images outputted by the H&E generative model and the H&E stain images ensemble from said training dataset; wherein each ithvirtual H&E stain image is generated using a latent representation generated by the ithIHC encoder of the H&E generative model;■ each ithIHC encoder from the H&E generative model receives as an input the IHC stain images from the ithIHC stain images ensemble from the training dataset and the ithvirtual IHC stain image outputted by the ithIHC decoder, and the H&E decoder from said H&E generative model provides as an output a second set of n virtual H&E stain images, the H&E discriminative model receiving as an input each image of the second set of n virtualH&E stain images outputted by the H&E generative model, one by one, and the H&E stain images ensemble from said training dataset; the H&E encoder from said IHC generative model further receiving as an input the second set of n virtual H&E stain images, said IHC generative model providing as an output a second set of n virtual IHC stain images, wherein each ithIHC decoder of the n IHC decoders provides as an output a ithvirtual IHC stain image of the second set of n virtual IHC stain images, the ithIHC discriminative model receiving as an input the ithvirtual IHC stain image outputted by the ithIHC decoder and the IHC stain images from the ithIHC stain images ensemble from the training dataset; at least one output configured to provide the obtained training parameters for said IHC machine learning system, wherein said IHC machine learning system comprises the H&E encoder, the n IHC decoders and the n IHC discriminative models.

[0010] According to the invention, a virtual IHC stain image is an image generated by the IHC generative model of the invention. The virtual IHC stain image simulates the appearance of a real IHC stain image by creating staining patterns and visual characteristics observed in actual IHC-stained tissue samples. On the contrary, a IHC stain image is an image taken from a microscope slide of a tissue stained with a real IHC stain.

[0011] According to the invention, a virtual H&E stain image is an image generated by the H&E generative model of the invention. The virtual H&E stain image simulates the appearance of a real H&E stain image by creating staining patterns and visual characteristics observed in actual H&E-stained tissue samples. On the contrary, a H&E stain image is an image taken from a microscope slide of a tissue stained with a real H&E stain.

[0012] In other words, the device of the present invention allows to obtain a trained IHC machine learning system capable of generating n virtual IHC images based on a singleH&E image. Therefore, the device of the invention advantageously allows to overcome the shortcomings of H&E staining by generating images wherein proteins within the cells are visualizable and quantifiable. Moreover, the device of the invention also allows to overcome the shortcomings of IHC staining, because the generation of virtual IHC images is almost instantaneous, contrary to real IHC images, that requires a long time to prepare the sample for visualization. Additionally, the same tissue sample can be virtually processed with different IHC stains, which is almost impossible on real tissues because they are damaged by the staining process.

[0013] In the present invention, a training dataset of unpaired images refer to a type of dataset used for training machine learning models, wherein there is no one-to-one correspondence or pairing between images in the different IHC stain images ensembles. In other words, unlike paired image datasets, where each image in one domain has a corresponding image in the other domain, there is no explicit mapping between images from one IHC stain images ensemble to another. More precisely, in the case of the invention, there is no connection between the tissue samples used to make the H&E stain images ensemble and IHC stain images ensembles. In this context, the at least one processor is configured to train said trainable machine learning system in an unsupervised manner.

[0014] “Unsupervised training” refers to a training wherein the algorithm is trained on a dataset without explicit supervision or labeled output. In other words, the algorithm explores the inherent structure and patterns within the input data without guidance on the correct output. The main goal of unsupervised learning is to identify hidden structures or relationships within the data. This type of learning is particularly useful when the task involves discovering patterns, grouping similar data points, reducing the dimensionality of the data, or generating new samples.

[0015] According to other advantageous aspects of the invention, the device comprises one or more of the features described in the following embodiments, taken alone or in any possible combination.

[0016] According to one embodiment, said at least one input is configured to receive a training dataset of paired images, said at least one processor being further configured, during training, to implement an adaptation phase wherein said trainable machine learning system is re-trained using said training dataset of paired images.

[0017] Advantageously, the trainable machine learning system is first trained on a huge volume of unpaired data (e.g. less exigent and less expensive and time-consuming to collect). Once the trainable machine learning system has converged, it may be re-trained using a smaller dataset of paired images (e.g. more expensive and time-consuming to collect and label) in order to refine the model training and make it more performant at generating virtual IHC stain images that look like real IHC stain images.

[0018] The adaptation phase may be a fine-tuning phase, a distillation phase or a transfer learning phase.

[0019] According to the invention, the training dataset of paired images comprises a paired H&E stain images ensemble comprising a plurality of H&E stain images, and n paired IHC stain images ensembles, wherein each ithIHC paired images ensemble comprises a plurality of IHC stain images obtained using an ithIHC stain, and wherein the ithIHC stain used to obtain the images of the ithIHC paired images ensemble is different from the IHC stain used to obtain the images of the others IHC paired images ensembles.

[0020] In the present invention, a training dataset of paired images corresponds to a training dataset wherein each jthH&E stain image from the H&E paired stain images ensemble is paired with the jthIHC stain image from each ithIHC paired stain images ensemble.

[0021] More precisely, the pairing of the H&E and IHC stain images typically involves ensuring that each H&E stain image is associated with its corresponding IHC stain image from the same tissue specimen. In other words, each tissue specimen is stained with H&E stain, then a first image is captured. Afterwards, the stain is removed and a first IHC stain is used to color the tissue specimen and a second image is captured. The process is repeated for each IHC stain of the n IHC stains used to construct the n ensembles of IHCstain images. This pairing is crucial for supervised learning tasks and enables the trainable machine learning system to learn the relationships between the structural and molecular information within the same tissue sample. Alternatively, the staining process may involve repeatedly using new raw tissue samples to obtain the H&E stain image and its corresponding IHC stain images. In that case, there are differences in tissue morphology and staining characteristics. Therefore, pairing may be performed by establishing a correspondence between the images. For instance, a manual or automatic annotation of corresponding regions or structures in both H&E and IHC images may be performed. Alternatively, a landmark-based registration (e.g. based on distinctive features present in all the images), a cross-correlation registration or deep learning registration may be performed. The images pairing may be performed beforehand or during the generation of the training dataset of paired images.

[0022] According to the invention, training on the training dataset of paired images may involve a first step of feeding the H&E encoder from said IHC generative model with the H&E stain images from the H&E stain images ensemble from said training dataset and the IHC generative model provides as an output n virtual IHC stain images, each ithIHC decoder of the n IHC decoders providing as an output one ithvirtual IHC stain image from the n virtual IHC stain images, the ithIHC discriminative model receiving as an input said ithvirtual IHC stain image outputted by the ithIHC decoder and the IHC stain images of the ithIHC stain images ensemble from the training dataset.

[0023] According to the invention, training on the training dataset of paired images may involve a second step wherein each ithIHC encoder from said H&E generative model is fed with the IHC stain images from the ithIHC stain images ensemble from said training dataset and with the ithvirtual IHC stain image outputted by the ithIHC decoder, and the H&E decoder from said H&E generative model provides as an output n virtual H&E stain images, the H&E discriminative model receiving as an input the n virtual IHC stain images outputted by the H&E generative model and the H&E stain images ensemble from said training dataset.

[0024] According to one embodiment, the structure of each discriminative model comprises at least one of the following: CNN, classifier, transformer.

[0025] In other words, the structure of each discriminative model may be a CNN, or a classifier or a transformer or a CNN followed by a classifier or even a transformer followed by a classifier.

[0026] According to one embodiment, during training, the at least one processor is further configured to minimize a global generator loss functions comprising at least one consistency component, at least one adversarial component, at least one regularization component and at least one dynamic weighting factor.

[0027] According to one embodiment, said dynamic weighting factor is calculated based on the total number of pixels in IHC-activated regions and in IHC-non-activated regions of the IHC stain images from the n IHC stain images ensembles.

[0028] According to the invention, IHC-activated regions typically refer to specific areas within the image wherein the target protein of interest is detected and visualized thanks to the IHC staining process. These regions are highlighted by a specific staining signal, such as a specific color, for instance due to the presence of the antigen- antibody complex, that is labeled with a detectable marker (e.g., a colored enzyme reaction product or a fluorescent signal) during the IHC staining process. On the contrary, IHC-non-activated regions refer to areas within the image where the target protein of interest is not detectable or visualizable. These regions are devoid of the staining signal typically associated with the presence of the target protein. For instance, IHC-non-activated regions are areas where the primary antibody did not bind to its target antigen, resulting in a lack of specific staining.

[0029] According to one embodiment, during training, each IHC discriminative model of the n IHC discriminative models is configured to output a prediction, the at least one processor being further configured to minimize n IHC discriminator loss functions associated with the n IHC discriminative models, wherein each ithIHC discriminator loss function comprises calculating, for each image fed to the corresponding ithIHC discriminative model, a difference between the prediction of said ithIHC discriminative model and the image fed to said ithIHC discriminative model.

[0030] In the context of the invention, the image fed to the ithIHC discriminative model is either the ithvirtual IHC stain image or one image from the ithIHC stain image ensemble.

[0031] Moreover, the prediction of the ithIHC discriminative model refers to its assessment of whether the fed image is real (e.g. a IHC stain image coming from the IHC stain images ensembles from the dataset) or fake (e.g. a virtual IHC stain image generated by a IHC decoder from the IHC generative model). The prediction may be a continuous value, for instance comprised between a first value and a second value (e.g. for example 0 and 1), where the first value (e.g. for example 0) indicates a prediction of a “fake image” and the second value (e.g. for example 1) indicates a prediction of a “real image”. Accordingly, IHC stain images coming from the IHC stain images ensembles from the dataset are associated with the first value indicative of a “real image” (e.g. for example 1), while virtual IHC stain images generated by IHC decoders from the IHC generative model are associated with the second value indicative of a “fake image” (e.g. for example 1). Hence, calculating, for each image fed to the corresponding ithIHC discriminative model, a difference between the prediction of said ithIHC discriminative model and the image fed to said ithIHC discriminative model comprises comparing the value comprised between the first value and the second value predicted by the ithIHC discriminative model and the value associated with the image fed to the corresponding ithIHC discriminative model, indicative of a “fake image” or a “real image”.

[0032] The learning process typically involves iteratively training the generative model and the discriminative models until a convergence criterion is met. The setting of the convergence criterion depends on the specific goals of the training and the desired quality of generated virtual IHC images.

[0033] According to one embodiment, the calculated difference is a pixel-wise difference.

[0034] In the context of the invention, a pixel-wise difference refers to the comparison of the input image and the prediction of the IHC discriminative model at a pixel level. Typically, the input image may be associated with a binary matrix containing only thefirst value (e.g. for example 0) indicative of a “fake image” or only the second value (e.g. for example 1) indicative of a “real image”. The IHC discriminative model produces a matrix of values comprised between the first value and the second value (e.g. for instance between 0 and 1) as its prediction for the input image. The corresponding values of the two matrices may be compared by means a distance calculation such as the LI Distance (e.g. sum of absolute differences between corresponding values, L2 Distance (e.g. square root of the sum of squared differences between corresponding values, Mean Squared Error (average of squared differences between corresponding values), Binary CrossEntropy.

[0035] According to one embodiment, during training, the H&E discriminative model is configured to output a prediction, the at least one processor being further configured to minimize a H&E discriminator loss functions associated with the H&E discriminative model, wherein said H&E discriminator loss function comprises calculating, for each image (e.g. either a virtual H&E stain image generated by the H&E generative model or a H&E stain image from the H&E stain image ensemble) fed to the H&E discriminative model, a difference between the prediction of said H&E discriminative model and the image fed to said H&E discriminative model.

[0036] In other words, calculating the H&E discriminator loss function comprises comparing the value comprised between the first value and the second value predicted by the H&E discriminative model and the value associated with the image fed to the H&E discriminative model, indicative of a “fake image” or a “real image”.

[0037] According to one embodiment, the calculated difference is a pixel-wise difference.

[0038] According to one embodiment, the H&E stain images from the H&E stain images ensemble and / or the IHC stain images from the n IHC stain images ensembles are obtained by capturing at least one microscope slide with a magnification of xlO. Indeed, experiments have shown that this magnification advantageously gives the best results in the generation of the virtual IHC stain images. In practice, the mean square error (MSE), the LI error (absolute error), the SSIM (structural similarity index measure) or the PSNR(sigle de Peak Signal to Noise Ratio) may be employed to compare the generated virtual IHC stain images with their real counterpart and a magnification of xlO showed reduced MSE compared to other magnifications.

[0039] The present invention further relates to a device for generating n virtual Immunohistochemical (IHC) stain images from one Hematoxylin and Eosin (H&E) stain image, using an IHC machine learning system trained using the training parameters obtained thanks to the device for training described above, wherein said device comprises: at least one input configured to receive said H&E stain image; at least one processor configured to feed said H&E stain image to said trained IHC machine learning system and output n virtual IHC stain images; and at least one output configured to provide said n virtual IHC stain images.

[0040] Advantageously, the obtained n virtual IHC stain images may be used for research purpose or diagnostic investigation. Indeed, the virtual IHC stain images allow researchers to visualize the presence, localization, and relative abundance of specific proteins within tissues at a less expensive cost than traditional IHC staining techniques.

[0041] The IHC machine learning system may be a sub-part of the trainable machine learning system.

[0042] According to a first embodiment, the IHC machine learning system may comprise only the H&E encoder, the n IHC decoders and the n IHC discriminative models.

[0043] According to another embodiment, the IHC machine learning system further comprises the H&E decoder and the H&E discriminative model. In that case, the H&E encoder and the H&E decoder may form a verification generative model, said verification generative model being configured to receive said H&E stain image and to output a virtual H&E stain image, said virtual H&E stain image being fed to the H&E discriminative model so as to obtain a H&E reliability score.

[0044] According to an embodiment, when the obtained H&E reliability score does not satisfy a predefined first reliability criterion, the at least one processor is further configured to generate a heatmap to provide visual Al explicability for a user.

[0045] According to an embodiment, the H&E reliability score may be a pixel-wise reliability score.

[0046] According to one embodiment, the at least one processor is further configured to segment IHC-activated regions of at least two virtual IHC stain images from the n virtual IHC stain images and overlay said segmented regions so as to obtain at least one virtual multiplex IHC stain image.

[0047] Advantageously, multiplex IHC staining allows to simultaneously visualize and analyze multiple target proteins within a single tissue sample. This approach enables the examination of the relative locations and interactions of different types of proteins, thus allowing researchers and pathologists to gain insights into the complex interplay of proteins and cellular interactions.

[0048] According to one embodiment, each of the n virtual IHC stain images outputted by the device described above may be provided as input to the corresponding IHC discriminative model of the IHC machine learning system in order to obtain a pixel- wise reliability score.

[0049] In other words, the ithIHC discriminative model may receive as an input the virtual IHC stain image outputted by the ithIHC decoder and the ithIHC discriminative model is configured to output a pixel-wise reliability score. Thus, at the end of the process, n pixel-wise reliability scores are obtained.

[0050] According to an embodiment, when the obtained pixel- wise reliability scores do not satisfy a predefined second reliability criterion, the at least one processor is further configured to generate a heatmap to provide visual Al explicability for a user.

[0051] According to the invention, the heatmap for visual Al explicability refers to a graphical representation that provides to the user an evaluation of the quality of the H&Estain image by highlighting the regions or features within an image that have lower quality.

[0052] Lower quality in the H&E stain images may be explained by the quality of the images comprised in the training dataset. For instance, variations in H&E chemical concentrations (higher or lower), deviations in tissue preparation protocols, such as difference in thickness of the samples, differences in scanning techniques or in the hardware used, the presence of impurities or contaminants on the tissue surface, issues during scanning, like folded tissue, abnormal stretching or external mechanical pressure on the tissue, aggressive compression on the digital slide.

[0053] According to one embodiment, the at least one processor is further configured to: divide the H&E stain image into a set of tiles, each tile being fed to said trained IHC machine learning system so as to output n corresponding virtual IHC stain tiles; for each tile of the set, select thehoutputted virtual IHC stain tile so as to reconstruct the k,hvirtual IHC stain image; and output the reconstructed ensemble of n virtual IHC stain images.

[0054] Advantageously, dividing the H&E stain image into a set of tiles allows for parallel processing, which can significantly speed up image analysis. Each tile can be processed independently, utilizing multiple CPU or GPU cores for improved computational efficiency.

[0055] The present invention further relates to a computer-implemented method for generating n virtual Immunohistochemical (IHC) stain images from one Hematoxylin and Eosin (H&E) stain image, using a IHC machine learning system trained using the training parameters outputted by the device for training described above, wherein said method comprises: receiving said H&E stain image; feeding said H&E stain image to said IHC machine learning system so as to generate said n virtual IHC stain images, and outputting said n virtual IHC stain images.

[0056] The present disclosure further pertains to a computer-readable medium comprising instructions which, when executed by a computer, cause the computer to carry out the method for generating n virtual IHC stain images from one H&E stain image compliant with any of the above execution modes.

[0057] In addition, the disclosure relates to a computer program product comprising instructions which, when the program is executed by a computer, cause the computer to carry out the method for generating n virtual IHC stain images from one H&E stain image described above.

[0058] The present disclosure further pertains to a computer-readable storage medium comprising instructions which, when executed by a computer, cause the computer to carry out the method for generating n virtual IHC stain images from one H&E stain image described above.

[0059] The present disclosure further pertains to a non-transitory program storage device, readable by a computer, tangibly embodying a program of instructions executable by the computer to perform a method for generating n virtual IHC stain images from one H&E stain image, compliant with the present disclosure.

[0060] Such a non-transitory program storage device can be, without limitation, an electronic, magnetic, optical, electromagnetic, infrared, or semiconductor device, or any suitable combination of the foregoing. It is to be appreciated that the following, while providing more specific examples, is merely an illustrative and not exhaustive listing as readily appreciated by one of ordinary skill in the art: a portable computer diskette, a hard disk, a ROM, an EPROM (Erasable Programmable ROM) or a Flash memory, a portable CD-ROM (Compact-Disc ROM).DEFINITIONS

[0061] In the present invention, the following terms have the following meanings:

[0062] The terms “adapted” and “configured” are used in the present disclosure as broadly encompassing initial configuration, later adaptation or complementation of thepresent device, or any combination thereof alike, whether effected through material or software means (including firmware).

[0063] The term “processor” should not be construed to be restricted to hardware capable of executing software, and refers in a general way to a processing device, which can for example include a computer, a microprocessor, an integrated circuit, or a programmable logic device (PLD). The processor may also encompass one or more Graphics Processing Units (GPU), whether exploited for computer graphics and image processing or other functions. Additionally, the instructions and / or data enabling to perform associated and / or resulting functionalities may be stored on any processor- readable medium such as, e.g., an integrated circuit, a hard disk, a CD (Compact Disc), an optical disc such as a DVD (Digital Versatile Disc), a RAM (Random- Access Memory) or a ROM (Read-Only Memory). Instructions may be notably stored in hardware, software, firmware or in any combination thereof.

[0064] “Machine learning (ML)” designates in a traditional way computer algorithms improving automatically through experience, on the ground of training data enabling to adjust parameters of computer models through gap reductions between expected outputs extracted from the training data and evaluated outputs computed by the computer models.

[0065] A “hyper-parameter” presently means a parameter used to carry out an upstream control of a model construction, such as a remembering-forgetting balance in sample selection or a width of a time window, by contrast with a parameter of a model itself, which depends on specific situations. In ML applications, hyper-parameters are used to control the learning process.

[0066] “Datasets” are collections of data used to build an ML mathematical model, so as to make data-driven predictions or decisions. In “supervised learning” (i.e. inferring functions from known input-output examples in the form of labelled training data), three types of ML datasets (also designated as ML sets) are typically dedicated to three respective kinds of tasks: “training”, i.e. fitting the parameters, “validation”, i.e. tuning ML hyperparameters (which are parameters used to control the learning process), and “testing”, i.e. checking independently of a training dataset exploited for building a mathematical model that the latter model provides satisfying results.

[0067] A “neural network (NN)” designates a category of ML comprising nodes (called “neurons”), and connections between neurons modeled by “weights”. For each neuron, an output is given in function of an input or a set of inputs by an “activation function”. Neurons are generally organized into multiple “layers”, so that neurons of one layer connect only to neurons of the immediately preceding and immediately following layers.

[0068] The above ML definitions are compliant with their usual meaning, and can be completed with numerous associated features and properties, and definitions of related numerical objects, well known to a person skilled in the ML field. Additional terms will be defined, specified or commented wherever useful throughout the following description.BRIEF DESCRIPTION OF THE DRAWINGS

[0069] The present disclosure will be better understood, and other specific features and advantages will emerge upon reading the following description of particular and non-restrictive illustrative embodiments, the description making reference to the annexed drawings wherein:

[0070] Figure 1 is a block diagram representing schematically a particular mode of a device for training a IHC machine learning system for generating n virtual Immunohistochemical (IHC) stain images from one Hematoxylin and Eosin (H&E) stain image compliant with the present disclosure;

[0071] Figure 2 is a flow chart showing successive steps executed with the device for training a IHC machine learning system for generating n virtual Immunohistochemical (IHC) stain images from one Hematoxylin and Eosin (H&E) stain image figure 1;

[0072] Figure 3 is a block diagram representing schematically a particular mode of a device for generating n virtual Immunohistochemical (IHC) stain images from one Hematoxylin and Eosin (H&E) stain image, using a IHC machine learning system trained using the training parameters obtained thanks to the device represented in Figure 1, compliant with the present disclosure;

[0073] Figure 4 is a flow chart showing successive steps executed with the device for generating n virtual Immunohistochemical (IHC) stain images from one Hematoxylin and Eosin (H&E) stain image of Figure 3;

[0074] Figure 5 is a flow chart showing the architecture and successive steps performed by the trainable machine learning system during training;

[0075] Figure 6 is a flow chart showing further elements of the architecture and further steps that may be performed by the trainable machine learning system during training;

[0076] Figure 7 is a flow chart showing the architecture and successive steps performed by the IHC machine learning system trained using training parameters obtained with the device for training according to the invention; and

[0077] Figure 8 shows an exemplary embodiment of a visual Al heatmap generated for a H&E stain image given as an input to the device for generating n virtual Immunohistochemical (IHC) stain images from one Hematoxylin and Eosin (H&E) stain image.ILLUSTRATIVE EMBODIMENTS

[0078] The present description illustrates the principles of the present disclosure. It will thus be appreciated that those skilled in the art will be able to devise various arrangements that, although not explicitly described or shown herein, embody the principles of the disclosure and are included within its scope.

[0079] All examples and conditional language recited herein are intended for educational purposes to aid the reader in understanding the principles of the disclosure and the concepts contributed by the inventor to furthering the art, and are to be construed as being without limitation to such specifically recited examples and conditions.

[0080] Moreover, all statements herein reciting principles, aspects, and embodiments of the disclosure, as well as specific examples thereof, are intended to encompass both structural and functional equivalents thereof. Additionally, it is intended that such equivalents include both currently known equivalents as well as equivalents developed inthe future, i.e., any elements developed that perform the same function, regardless of structure.

[0081] Thus, for example, it will be appreciated by those skilled in the art that the block diagrams presented herein may represent conceptual views of illustrative circuitry embodying the principles of the disclosure. Similarly, it will be appreciated that any flow charts, flow diagrams, and the like represent various processes which may be substantially represented in computer readable media and so executed by a computer or processor, whether or not such computer or processor is explicitly shown.

[0082] The functions of the various elements shown in the figures may be provided through the use of dedicated hardware as well as hardware capable of executing software in association with appropriate software. When provided by a processor, the functions may be provided by a single dedicated processor, a single shared processor, or a plurality of individual processors, some of which may be shared.

[0083] It should be understood that the elements shown in the figures may be implemented in various forms of hardware, software or combinations thereof. Preferably, these elements are implemented in a combination of hardware and software on one or more appropriately programmed general-purpose devices, which may include a processor, memory and input / output interfaces.

[0084] The present disclosure will be described in reference to a particular functional embodiment of a device 1 for training a trainable machine learning system 40, as illustrated on Figure 1. The device 1 is adapted to output training parameters 42 and a (corresponding trained) IHC machine learning system 41 for generating n virtual IHC stain images 71 based on a single H&E stain image 25, as illustrated on Figure 3.

[0085] To that end, the device 1 is configured to train the trainable machine learning system 40 using a training dataset (21, 22, 23, 24) so as to obtain the training parameters 42 for said IHC machine learning system 41. With trainable machine learning system 40 it as to be understood the unparametrized architecture of a machine learning system. The device 1 may be configured to receive as input said training dataset (i.e.; from a database 10) or generate the training data set in a module 13 for constructing the training dataset.

[0086] The training dataset may comprise a H&E stain images ensemble 21 comprising a plurality of H&E stain images.

[0087] According to the present invention, a H&E stain image is an image of a type of histological or pathological microscope slide that is commonly used in the field of medicine and biology. In a H&E stain, two different dyes, hematoxylin and eosin, are used to stain different cellular components, making it easier to visualize and study tissues under a microscope. Hematoxylin stains cell nuclei and other structures that are rich in DNA. Nuclei appear blue-purple when stained with hematoxylin, making it possible to identify the number, size, and shape of cell nuclei within a tissue sample. Eosin is an acidic dye that stains the cytoplasm and extracellular matrix of cells. It highlights the pink to red coloration in the H&E stain image, making it possible to differentiate between various cellular and tissue components based on their staining intensity and color. The H&E stain images may be obtained thanks to a digital imaging system mounted on a microscope or a dedicated microscope slide scanner.

[0088] The training dataset may also comprise n IHC stain images ensembles 22, 23, 24.

[0089] IHC stain images are image of a type of histological slides used in pathology and research to visualize the presence, location, and abundance of specific proteins within tissue samples. IHC staining relies on the biding of markers on the proteins of interest. The IHC stain images may also be obtained thanks to a digital imaging system mounted on a microscope or a dedicated microscope slide scanner.

[0090] Theoretically, there are as many stain types as there are proteins. For instance, the target protein may be stained using a direct method wherein a primary antibody is labeled and applied to the tissue in a one-step process. Alternatively, the target protein may be stained using an indirect method wherein a secondary antibody is labeled, allowing for signal amplification and use with many different primary antibodies. The stains may be chosen from a wide variety, including enzyme conjugates such as horseradish peroxidase (HRP) or alkaline phosphatase (AP), and fluorescently labeled secondary antibodies or streptavidin. Options also extend to the Avidin-Biotin Complex (ABC) reagents, Diaminobenzidine (DAB), and Alkaline phosphatase substrates like BCIP / NBT. Classic histological stains include Hematoxylin and Eosin (H&E), Ziehl-Neelsen (ZN), Whartin-Starry (WS), von Kossa (vK), and Masson's trichrome (MT). Specialized stains like thioflavine, rouge Sirius (RS), rhodanine, congo red (CR), mucicarmine, Gram stain (GS), oil red O (ORO), and black Sudan (BS) can also be selected. For specific tissue structures or organisms, one might choose silver stain, orcein, reticulin, Peris' Prussian blue (PPB), Periodic acid-Schiff (PAS), Amylase PAS (APAS), Hale's colloidal iron (HCI), or Gomori methenamine silver (GMS). For cytological studies, the Papanicolaou stain (Pap) is commonly utilized. Other notable stains include May Grunwald-Giemsa (MGG), Fontana-Masson (FM), esterases, toluidine blue (TB), and cresyl violet (CV).

[0091] According to the invention, the number n of IHC stain images ensembles may be equal to 1 or 2 or 3 but more preferably, n is superior or equal to 4. In one alternative embodiment, n is superior or equal to 2.

[0092] Among the n IHC stain images ensembles 22, 23, 24, each IHC stain images ensemble preferably comprises images of tissue samples stained using the same stain. Each IHC stain images ensemble may be obtained using a different stain. In other words, among the n IHC stain images ensembles 22, 23, 24, n different stains may be used.

[0093] The stain used in a given IHC stain images ensemble preferably stains at least one same protein. Alternatively, the IHC stain may stain a group of different proteins with a common chemical structure. A different protein or the same protein may be stained in another IHC stain images ensemble.

[0094] Moreover, the biological tissues used to obtain the IHC images and H&E images may come from a same subject or from different subjects.

[0095] The device 1 for training the trainable machine learning system 40 is associated with a device 2, represented on Figure 3, for generating n virtual IHC stain images 71, using the IHC machine learning system 41 trained using the training parameters 42 obtained from the device 1, which will be subsequently described.

[0096] Though the presently described devices 1 and 2 are versatile and provided with several functions that can be carried out alternatively or in any cumulative way, other implementations within the scope of the present disclosure include devices having only parts of the present functionalities.

[0097] Each of the devices 1 and 2 is advantageously an apparatus, or a physical part of an apparatus, designed, configured and / or adapted for performing the mentioned functions and produce the mentioned effects or results. In alternative implementations, any of the device 1 and the device 2 is embodied as a set of apparatus or physical parts of apparatus, whether grouped in a same machine or in different, possibly remote, machines. The device 1 and / or the device 2 may have functions distributed over a cloud infrastructure and be available to users as a cloud-based service, or have remote functions accessible through an API.

[0098] The device 1 for generating the IHC machine learning system 41 and the device 2 for obtaining the n virtual IHC stain images 71 may be integrated in a same apparatus or set of apparatus, and intended to same users. In other implementations, the structure of the device 2 may be completely independent from the structure of the device 1, and may be provided for other users. For example, the device 2 may have a trained IHC machine learning system 41 available to operators for generation of n virtual IHC stain images, wholly set from previous training effected upstream by other players with the device 1.

[0099] In what follows, the modules are to be understood as functional entities rather than material, physically distinct, components. They can consequently be embodied either as grouped together in a same tangible and concrete component, or distributed into several such components. Also, each of those modules is possibly itself shared between at least two physical components. In addition, the modules may be implemented in hardware, software, firmware, or any mixed form thereof as well. They are preferably embodied within at least one processor of the device 1 or of the device 2.

[0100] The device 1 may comprise a module 11 for receiving, the H&E stain images ensemble 21, the n IHC stain images ensembles 22, 23, 24 and the trainable machine learning system 40, stored in one or more local or remote database(s) 10. The latter can take the form of storage resources available from any kind of appropriate storage means, which can be notably a RAM or an EEPROM (Electrically-Erasable Programmable Read- Only Memory) such as a Flash memory, possibly within an SSD (Solid-State Disk).

[0101] The trainable machine learning system 40, as illustrated on Figure 5 and Figure 6, may comprise n IHC discriminative models 27 and a IHC generative model 30comprising one H&E encoder 31 and n IHC decoders 32. Additionally, the trainable machine learning system 40 may further comprise a H&E discriminative model 37 and a H&E generative model 35 comprising n IHC encoders 33 and one H&E decoder 36.

[0102] In one embodiment, the one H&E encoder 31 and each ithIHC decoder 39 of the IHC generative model 30 form an encoder-decoder pair. In the same manner, each ithIHC encoder 34 and the H&E decoder 36 may form an encoder-decoder pair. In the present disclosure, ithis used as an increment to designate an element among n elements and is not restricted to any specific element of the n elements. In other words, ithmay refer to the 1st, 2nd, 3rd, ..., or nlhelement of the n elements.

[0103] The images from the received ensemble of H&E stain images 21 and n ensembles of IHC stain images 22, 23, 24 may be unpaired. When the images in a dataset are unpaired, it means that the H&E stain images and IHC stain images do not have a direct, one-to-one correspondence between them. In such cases, each H&E image is not explicitly associated with a specific corresponding IHC image from the same tissue specimen. When using unpaired images, it is possible to implement an unsupervised training.

[0104] Supervised training and image pairing is generally the most commonly used training technique. However, in the case of H&E stain images and IHC stain images, it is very difficult to obtain identical images of tissues because the process of removing a stain and putting another stain on a tissue sample is very expensive and time-consuming and tends to damage the tissue sample. Therefore, the use of unpaired H&E stain images and IHC stain images and unsupervised training in the present invention becomes interesting as it is less expensive and time-consuming and does not require identical images.

[0105] The device 1 further comprises optionally a module 12 for preprocessing the H&E stain images and / or the IHC stain images. For instance, module 12 may be adapted to preprocess the H&E stain images by performing automatic segmentation of the tissue regions, followed by normalization to reduce batch effects and variability.

[0106] Additionally, the H&E encoder 31 is configured to be fed with a H&E stain image (e.g., the ensemble of H&E stain images 21) and used to obtain a latent representation of said H&E image (e.g., during the training the latent space of H&Eencoder 31 is learned using the H&E stain images ensemble 21). In other words, during the training, the H&E encoder 31 may receive one after the other each image of the ensemble of H&E stain images 21. Each latent representation may then be passed on to the H&E decoder 36 to generate a denoised H&E image (e.g., a virtual H&E stain image). The denoised H&E images may then be used for training.

[0107] This process can address challenges such as variations in H&E chemical concentration and can also reduce background noise to a certain extent. However, for extreme cases, this method might be ineffective. In such instances, a heatmap may be provided to the user, highlighting the specific spatial origin of the issue and the user may choose to exclude the problematic H&E stain image from the H&E stain images ensemble 21.

[0108] The device may comprise a module 13 for the construction of the training dataset based on the received ensemble of H&E stain images 21 and n ensembles of IHC stain images 22, 23, 24.

[0109] The device 1 may further comprise a module 14 configured to train the trainable machine learning system 40 using the training dataset received (or constructed by module 13) to obtain training parameters 42.

[0110] As previously mentioned, the trainable machine learning system 40 relies on a IHC generative model 30 configured to produce virtual IHC stain images (71, 77) and to feed images to IHC discriminative models 27. The IHC discriminative models 27 are then configured to evaluate the data received and discriminate between the data generated by the IHC generative model 30 (e.g. virtual IHC stain images) and real data (e.g. real IHC stain images coming from the n ensembles of IHC stain images 22, 23, 24). In other words, each ilhIHC discriminative model 28 is used to evaluate the ability of its corresponding IHC generator (i.e., pair H&E encoder 31 and ilhIHC decoder 39) to produce virtual IHC images that are indistinguishable from real IHC images associated with the same IHC stain.

[0111] The trainable machine learning system 40 also relies on a H&E generative model 35 configured to produce virtual H&E stain images 73, 74 and to feed images to a H&E discriminative model 21. The H&E discriminative model 21 is then configured toevaluate the data received and discriminate between the data generated by the H&E generative model 35 (e.g. virtual H&E stain images) and real data (e.g. real H&E stain images coming from the H&E stain image ensemble 21).

[0112] More precisely, as illustrated on Figure 5, training may involve a first step of feeding the H&E encoder 31 from the IHC generative model 30 with the images from the H&E stain images ensemble 21.

[0113] The IHC generative model 30 is based on an architecture comprising the H&E encoder 31 (contracting path) and the n IHC decoders 32 (expansive path). The group of n IHC decoders 32 comprises a first,..., i*,..., and nlhIHC decoder 39. The H&E encoder 31 is configured to map the H&E stain images into a shared latent space. In other words, the H&E stain images are mapped from their original high-dimensional representation (e.g., pixel values) to a lower-dimensional space called the “latent space”. The latent space is a compressed and more abstract representation of the input images.

[0114] Afterward, each ithIHC decoder 39 aims to recover spatial information from the shared latent space and produces a segmentation map. It utilizes a series of upsampling and convolutional layers to gradually increase spatial resolution and generate an ithvirtual IHC stain image 72 from points in this shared latent space. In the end, a first set of n virtual IHC stain images 71 is outputted by the IHC generative model 30.

[0115] In other words, during the first step of the training, each H&E image from the H&E stain images ensemble 21 is provided as input, one by one (i.e., subsequently, in series) to the H&E encoder 31 from said IHC generative model 30 and each of the n IHC decoders 39 of the IHC generative model 30 provides as an output the corresponding a ithvirtual IHC stain image 72. The ensemble of the obtained n virtual IHC stain images 72 obtained for one H&E image is called in the present description first set of n virtual IHC stain images 71.

[0116] Each ithIHC discriminative model 28 from the n IHC discriminative models 27 then receives the corresponding ithvirtual IHC stain image 72 outputted by the ithIHC decoder 39 and images from the ithIHC stain images ensemble 23, and each ithIHC discriminative model 28 is configured to output predictions about these images.T1

[0117] Each ithvirtual IHC stain image 72 outputted by the ithIHC decoder 39 is also fed to the ithIHC encoder 34 from the H&E generative model 35. The n IHC encoders 33 share a common latent space with the H&E decoder 36. Each ithIHC encoder 34 is therefore configured to obtain a latent representation of the received ithvirtual IHC stain image 72. The latent representations may then be passed on to the H&E decoder 36 (e.g., sequentially or in parallel) so as to output a first set of n virtual H&E stain images 73, wherein each ithvirtual H&E stain image 74 is generated using the latent representation generated by the ithIHC encoder 34.

[0118] Afterwards, the first set of n virtual H&E stain images 73 may be fed (e.g., one by one, a virtual H&E stain image after the other) to the H&E discriminative model 37, together with the images from the H&E stain images ensemble 21. The H&E discriminative model 37 is then configured to output a prediction on the image that is fed to it (e.g., whether it is a “real” image coming from the H&E stain images ensemble 21 or a “fake’7”generated” image generated by the IHC encoders 33).

[0119] Additionally, as illustrated on Figure 6, training may involve a second step of feeding the n IHC stain images ensembles 22, 23, 24 and the virtual IHC stain images 71 outputted by the IHC decoders 32, to the H&E generative model 35. More precisely, each ithIHC stain images ensemble 23 and each ithvirtual IHC stain image 72 outputted by the ithIHC decoder 39 is fed to the ithIHC encoder 34 of the H&E generative model 35. The n IHC encoders 33 share a common latent space with the H&E decoder 36. Each ithIHC encoder 34 is therefore configured to obtain a latent representation of the IHC stain images from the ithIHC stain images ensemble 23. The latent representations may then be passed on to the H&E decoder 36 so as to output a second set of n virtual H&E stain images 76, wherein each ithH&E stain image 75 is generated using the latent representation generated by the ithIHC encoder 34. The H&E decoder 36 may receive one by one the latent representations each corresponding to an ithIHC stain images ensemble 23 (i = 1, ..., n) of the n IHC stain images ensembles 23.

[0120] Afterwards, the second set of n virtual H&E stain images 76 may be fed to the H&E discriminative model 37, in addition with the images from the H&E stain images ensemble 21. The H&E discriminative model 37 is then configured to output a prediction on the image that is fed to it (e.g. whether it is a “real” image coming from the H&E stainimages ensemble 21 or a “fake” image generated by the IHC encoders 33 and H&E decoder 36).

[0121] The second set of n virtual H&E stain images 76 is also fed to the H&E encoder 31 from the IHC generative model 30. The H&E encoder 31 is configured to map the second set of n virtual H&E stain images 76 into the shared latent space. Afterward, each ithIHC decoder 39 aims to recover spatial information from the shared latent space and generate an ithvirtual IHC stain image 78 from points in this shared latent space. In the end, a second set of n virtual IHC stain images 77 is outputted by the IHC generative model 30.

[0122] Each ithIHC discriminative model 28 from the n IHC discriminative models 27 then receives the ithvirtual IHC stain image 78 outputted by the ithIHC decoder 39 and images (e.g., one by one) from the ithIHC stain images ensemble 23 and is configured to output predictions about these images.

[0123] During training, a global generator loss function is calculated for the trainable machine learning system 40. The global generator loss function may be a sum or a weighted sum or a dynamically weighted sum of several components.

[0124] The dynamic weighting factors are configured to control the influence of the consistency components, adversarial components and regularization components. Choice of the weights may be based on the total number of pixels in IHC-activated regions and in IHC-non-activated regions of the IHC stain image.

[0125] The first component (consistency component) is derived from the calculation of n IHC generator loss functions computed for the IHC generative model 30. In other words, a IHC generator loss function is computed for each encoder-decoder pair formed by the H&E encoder 31 and one of the IHC decoders 32.

[0126] Each IHC generator loss function may be configured to measure how well the virtual H&E stain images 73 generated by the H&E generative model 35 match with true H&E images from the training dataset (e.g. give a feedback on the performances of the H&E encoder 31 and the n IHC decoders 32). For instance, a distance between the true H&E images and the virtual H&E stain images 73 may be computed. Computing the distance between the true H&E image and the virtual H&E stain image involvesquantifying the dissimilarity or similarity between their pixel values or features. For instance, Mean Square Error, Structural Similarity Index, Cosine Similarity, Euclidian distance (L2 distance), Manhattan distance (LI distance) may be used.

[0127] The second component (consistency component) is derived from the calculation of n H&E generator loss functions computed for the H&E generative model 35. In other words, a H&E generator loss function is computed for each encoder-decoder pair formed by one of the IHC encoders 33 and the H&E decoder 36.

[0128] Each H&E generator loss function may be configured to measure how well the virtual IHC stain images 71 generated by the IHC generative model 30 match with true IHC images from the training dataset (e.g. give a feedback on the performances of the IHC encoders 33 and the H&E decoder 36). For instance, a distance between the true IHC images and the virtual IHC stain images 71 may be computed. Computing the distance between the true IHC image and the virtual IHC stain image involves quantifying the dissimilarity or similarity between their pixel values or features. For instance, Mean Square Error, Structural Similarity Index, Cosine Similarity, Euclidian distance (L2 distance), Manhattan distance (LI distance) may be used. Advantageously, an H&E generator loss function relying on IHC labeling allows to impose stricter penalties on the IHC generative model 30, which enables prioritizing IHC markers over structural markers in loss function minimization.

[0129] The global generator loss function may further comprise an adversarial component configured to measure how well the IHC generative model 30 can fool the n IHC discriminative models 27. To that end, the prediction of the n IHC discriminative models 27 (e.g. values comprised between 0 and 1) are compared to the value (e.g. 0 for a virtual IHC stain image and 1 for a real IHC image coming from the IHC stain image ensembles 22, 23, 24) associated with the images fed to the n IHC discriminative models 27.

[0130] The global generator loss function may further comprise an adversarial component configured to measure how well the H&E generative model 35 can fool the H&E discriminative model 37. To that end, the prediction of the H&E discriminative model 37 (e.g. values comprised between 0 and 1) are compared to the value (e.g. 0 for avirtual H&E stain image and 1 for a real H&E image coming from the H&E stain image ensembles 21) associated with the images fed to the H&E discriminative model 37.

[0131] The global generator loss function may further comprise a regularization component. For instance, the regularization component may include computation of the dissimilarity or similarity between two points or vectors in the latent space of the H&E generative model 35 and computation of the dissimilarity or similarity between two points or vectors in the latent space of the IHC generative model 30. The dissimilarity or similarity may be computed by calculating a distance such as Mean Square Error, Structural Similarity Index, Cosine Similarity, Euclidian distance (L2 distance), Manhattan distance (LI distance).

[0132] Additionally, the distance may be computed only on IHC-activated regions of the true IHC image and the virtual IHC stain image for the regularization component of the IHC generative model 30.

[0133] The global generator loss function may further comprise a regularization component for the verification generative model 38 formed by the H&E encoder 31 and the H&E decoder 36. For instance, a distance between a real H&E image fed to the verification generative model 38 and a virtual H&E image outputted by said verification generative model 38 may be computed. The distance may be calculated using Mean Square Error, Structural Similarity Index, Cosine Similarity, Euclidian distance (L2 distance), Manhattan distance (LI distance).

[0134] Moreover, the global generator loss function may further comprise a regularization component for the generative model formed by the n IHC encoders 33 and the n IHC decoders 32. For instance, a distance between the real IHC images fed to the generative model and the virtual IHC images outputted by said generative model may be computed. The distance may be calculated using Mean Square Error, Structural Similarity Index, Cosine Similarity, Euclidian distance (L2 distance), Manhattan distance (LI distance). The distance may be computed only on IHC-activated regions.

[0135] Regularization components are configured to prevent overfitting, which occurs when a model becomes too complex and starts fitting the training data too closely, including its noise and outliers. Regularization is typically added to the loss functionduring the training process to discourage overly complex models and promote simpler ones that generalize well to new, unseen data.

[0136] The parameters of the trainable machine learning system 40 may be iteratively updated either sequentially or parallelly.

[0137] When the parameters are updated sequentially, the global generator loss function is backpropagated sequentially through its corresponding each encoder-decoder pair.

[0138] For the H&E encoder 31 and the H&E decoder 36, the average of all the components is calculated and backpropagated to the H&E encoder-decoder pair.

[0139] When the parameters 30 are updated parallelly, the global generator loss function is backpropagated simultaneously to the corresponding encoder-decoder pair and to the H&E encoder-decoder pair.

[0140] Additionally, during training, n IHC discriminator loss functions are computed. In other words, a IHC discriminator loss function is computed for each ithIHC discriminative model 27.

[0141] During training, discriminator loss functions may be computed to measure how well the discriminative models 27, 37 discriminates between real images and virtual stain images generated by the generative models 30, 35.

[0142] To that end, a total of n IHC discriminator loss functions may be calculated, one for each ithIHC discriminative model 28. To obtain each ithdiscriminator loss function the average between a first component and a second component is calculated.

[0143] The first component is obtained by calculating a difference between the prediction of said ithIHC discriminative model 28 and the real images (e.g. coming from the IHC stain image ensembles 22, 23, 24) fed to said ithIHC discriminative model 28. For instance, a distance may be calculated using Mean Square Error, Structural Similarity Index, Cosine Similarity, Euclidian distance (L2 distance), or Manhattan distance (LI distance). Additionally, the distance may be computed only for IHC-activated regions of the true IHC image and the virtual IHC stain image.

[0144] The second component is obtained by calculating a difference between the prediction of said ithIHC discriminative model 28 and the ithvirtual IHC image 72 fed to said ithIHC discriminative model 28. For instance, a distance may be calculated usingMean Square Error, Structural Similarity Index, Cosine Similarity, Euclidian distance (L2 distance), or Manhattan distance (LI distance). Additionally, the distance may be computed only for IHC-activated regions of the true IHC image and the virtual IHC stain image.

[0145] Additionally, a H&E discriminator loss function may be calculated by averaging a third component and a fourth component.

[0146] The third component is obtained by calculating a difference between the prediction of said H&E discriminative model 37 and the real image (e.g. coming from the H&E stain image ensemble 21) fed to said H&E discriminative model 37. For instance, a distance may be calculated using Mean Square Error, Structural Similarity Index, Cosine Similarity, Euclidian distance (L2 distance), or Manhattan distance (LI distance).

[0147] The fourth component is obtained by calculating a difference between the prediction of said H&E discriminative model and the virtual H&E image 74 fed to said H&E discriminative model 37. For instance, a distance may be calculated using Mean Square Error, Structural Similarity Index, Cosine Similarity, Euclidian distance (L2 distance), or Manhattan distance (LI distance).

[0148] The parameters of the trainable machine learning system 40 may be updated by backpropagating each ithIHC discriminator loss function and the H&E discriminator loss function exclusively to its corresponding discriminative model (e.g. the ithIHC discriminative model 28 for the ithIHC discriminator loss function and the H&E discriminative model 37 for the H&E discriminator loss function).

[0149] The learning process typically involves iteratively training the generative model and the discriminative models until a certain convergence criterion is met. The satisfaction or stopping point depends on the specific goals of the training and the desired quality of generated virtual IHC stain images.

[0150] Once the training completed, module 14 is configured to output the obtained training parameters 42 and the trained IHC machine learning system 41.

[0151] Thus, as illustrated on Figure 7, the IHC machine learning system 41 may comprise the n IHC discriminative models 27 and the IHC generative model 30 with theH&E encoder 31 and the n IHC decoders 32 and optionally the H&E decoder 36 and the H&E discriminative model 37. Alternatively, the IHC machine learning system 41 may comprise the IHC generative model 30 with the H&E encoder 31 and the n IHC decoders 32 and optionally the H&E decoder 36 and the H&E discriminative model 37.

[0152] In the obtained IHC machine learning system 41, the n IHC decoders 32 from the IHC generative model 30 are all equivalent and so are the n IHC discriminative models 27.

[0153] The training parameters 42 and the IHC machine learning system may be stored in one or more local or remote database(s) 10. The latter can take the form of storage resources available from any kind of appropriate storage means, which can be notably a RAM or an EEPROM (Electrically-Erasable Programmable Read-Only Memory) such as a Flash memory, possibly within an SSD (Solid-State Disk).

[0154] In its automatic actions, the device 1 may for example execute the following process (Figure 2):- receiving the H&E stain images ensemble 21, the n IHC stain images ensembles 22, 23, 24 and the trainable machine learning system 40 (step 61),- optionally preprocessing the images from the H&E stain images ensemble 21 and the n IHC stain images ensembles 22, 23, 24 (step 62),- constructing the training dataset using the images of the H&E stain images ensemble 21 and the n IHC stain images ensembles 22, 23, 24 (step 63),- training the trainable machine learning system 40 using the constructed training dataset so as to obtain training parameters 42 and the IHC machine learning system 41 (step 64).

[0155] The present invention also relates to a device 2 for generating n virtual IHC stain images 71 from a single H&E stain image 25, using the IHC machine learning system 41 with the training parameters 42 obtained from the device 1, as described above. The device 2 will be described in reference to a particular function embodiment as illustrated in Figure 3.

[0156] The device 2 may comprise a module 15 for receiving the IHC machine learning system 41 and the H&E stain image 25, stored in one or more local or remotedatabase(s) 10. The latter can take the form of storage resources available from any kind of appropriate storage means, which can be notably a RAM or an EEPROM (Electrically- Erasable Programmable Read-Only Memory) such as a Flash memory, possibly within an SSD (Solid-State Disk). In advantageous embodiments, the IHC machine learning system 38 and all its parameters may have been previously generated by a system including the device 2 for training. Alternatively, the trainable IHC machine learning system 40 and its training parameters 42 (or the IHC machine learning system 41) may be received by the device 2 from a communication network.

[0157] The device 2 further comprises optionally a module 16 for preprocessing the H&E stain image. For instance, module 16 may be adapted to preprocess the H&E stain image 25 by performing automatic segmentation of the tissue regions, followed by normalization to reduce batch effects and variability.

[0158] Additionally, the H&E encoder 31 may be fed with the H&E stain image 25 and used to obtain a latent representation of the H&E image 25. The latent representation may then be passed on to the H&E decoder 36 to generate a denoised H&E image (e.g. a virtual H&E stain image). The denoised H&E images may then be used for training.

[0159] This process can address challenges such as variations in H&E chemical concentration and can also reduce background noise to a certain extent. However, for extreme cases, this method might be ineffective. In such instances, a heatmap may be provided to the user, highlighting the specific spatial origin of the issue and the user may choose to exclude the problematic H&E stain image from the H&E stain images ensemble 21.

[0160] The device 2 may further comprise a module 17 configured to provide the H&E stain image to the IHC machine learning system 41 so as to generate the corresponding n virtual IHC stain images 71.

[0161] Advantageously, the IHC machine learning system 41 may be patch-based or tilebased. In other words, the device 2 may comprise an additional module configured to divide the H&E stain image 25 received as an input into a set of m tiles (e.g. subsections of an image). Each tile is configured to be fed to the IHC machine learning system 41 soas to output n corresponding virtual IHC stain tiles. Based on the outputted virtual IHC stain tiles, it is possible to reconstruct the n virtual IHC stain images. Indeed, the kthvirtual IHC stain image corresponds to the association of the k,houtputted virtual IHC tile for each tile (among the m tiles) fed as an input to the IHC machine learning system 41.

[0162] The device 2 may interact with a user interface 18, via which information can be entered and retrieved by a user. The user interface 18 includes any means appropriate for entering or retrieving data, information or instructions, notably visual, tactile and / or audio capacities that can encompass any or several of the following means as well known by a person skilled in the art: a screen, a keyboard, a trackball, a touchpad, a touchscreen, a loudspeaker, a voice recognition system.

[0163] In its automatic actions, the device 2 may for example execute the following process (Figure 4):- receiving the H&E stain image 25 and the IHC machine learning system 41 (step 51),- optionally preprocessing the H&E stain image 25 (step 52),- providing the H&E stain image 25 to said IHC machine learning system 41 so as to generate the n virtual IHC stain images 71 (step 53).

[0164] In a further embodiment, the device 2 may be configured to provide virtual multiplex IHC stain images (e.g. images comprising multiple colors, wherein the colors correspond to different IHC stains overlaid on the same image and corresponding to different types of proteins). The multiplex IHC stain images may be obtained by segmenting the IHC-activated regions (e.g. regions of the image highlighted by a specific staining signal, such as a specific color) in the virtual IHC image that need to appear in the multiplex IHC stain image. The segmented regions obtained are then overlaid so as to obtain the virtual multiplex HC stain image.

[0165] The device 2 may be further configured to evaluate the quality of the data provided as an input to the device 2 and the data generated as an output by the device 2.

[0166] In one example, the device 2 may be configured to verify the quality of the H&E stain image 25. To that end, the verification generative model 38 may receive said H&E stain image 25 and output a virtual H&E stain image 26, said virtual H&E stain image 26being fed to the H&E discriminative model 37 so as to obtain a H&E reliability score based on the prediction of said H&E discriminative model 37. In other words, the device 2 is further configured to provide the H&E stain image 25 to the verification generative model 38 so as to obtain as output a virtual H&E stain image 26, which may be also normalized. The device 2 may be further configured to provide said virtual H&E stain image 26 to the H&E discriminative model 37 so as to obtain a H&E reliability score. The H&E reliability score may be an inception score, a Frechet inception distance Perceptual Path Length or Structural Similarity Index Measure.

[0167] When the H&E reliability score is calculated pixel- wise, a heatmap may be generated such as illustrated on Figure 8, in other to provide a representation of when the H&E reliability score does not satisfy a predefined first reliability criterion. Said predefined first reliability criterion may be a predefined threshold value to be compared with the H&E reliability score calculated for a pixel or a group of pixels. The heatmap may then comprise pixels of two different colors: a first color being representative of trustworthiness in the H&E stain image 25 (i.e., predefined first reliability criterion is satisfied, e.g. when the H&E reliability score is inferior to the predefined threshold) and a second color being representative of non-compliance of the H&E stain image 25 (i.e., predefined first reliability criterion is not satisfied, e.g. when the H&E reliability score is superior to the predefined threshold). In Figure 8, the H&E stain image 25 comprises a water drop distortion anomaly 175. This anomaly 175 is identifiable by the user in the heatmap thanks to an aggregation 176 of pixels of the second color, representative of non- compliance.

[0168] In another example, the device 2 may be configured to verify the quality of the virtual IHC images 171 generated by the IHC machine learning system 41. To that end, the prediction made by the ithdiscriminative model 28 based on the ithvirtual IHC stain image 172 from the n virtual IHC stain images 171 provided as input to said ithdiscriminative model 28 may be used to obtain a pixel-wise reliability score. When the obtained pixel-wise reliability scores do not satisfy a predefined second reliability criterion, the at least one processor is further configured to generate a heatmap to provide visual Al explicability for a user, for him to understand and explain the decision-making processes and outcomes of the IHC machine learning system 41.

[0169] The pixel- wise reliability score and the H&E reliability score may be obtained using the following strategies taken alone or in combination.

[0170] A predefined threshold (for example 0.9, representing a high trust level) may be defined as predefined first reliability criterion for the pixel- wise reliability score or H&E reliability score. The prediction of the n discriminative models 27 or H&E discriminative model 37 may be compared to the predefined threshold at a pixel level. If a substantial fraction (e.g., over 5% of the image) scores below this predefined threshold, it indicates that the stain doesn't meet the reliability standard.

[0171] Alternatively, the overall average or median reliability score of all the image's pixels may be calculated and compared to the predefined threshold (e.g. for instance 0.85), and a heatmap might be generated if the average or median is below the predefined threshold.

[0172] Alternatively, local variance in the predictions of the n discriminative models 27 or H&E discriminative model 37 may be evaluated and compared to a predetermined threshold.

[0173] Alternatively, Region of Interest (ROI) may be determined and a more rigorous predetermined threshold (e.g. for instance 0.95) could be set.

[0174] Also, past predictions made by the n discriminative models 27 or H&E discriminative model 37 may be taken into account.

[0175] To generate the heatmaps, it is possible to perform a retro propagation of the gradient of the t n discriminative models 27 or H&E discriminative model 37 (GRADCAM family), a drop out sampling or a Montecarlo sampling.

[0176] A particular apparatus may embody the device 1 as well as the device 2 described above. It corresponds for example to a workstation, a laptop, a tablet, a smartphone, or a head- mounted display (HMD).

[0177] That apparatus is suited to generation of segmentation mask and to related Machine Learning training. It comprises the following elements, connected to each other by a bus of addresses and data that also transports a clock signal:- a microprocessor (or CPU);- a graphics card comprising several Graphical Processing Units (or GPUs) and a Graphical Random Access Memory (GRAM); the GPUs are quite suited to image processing, due to their highly parallel structure;- a non-volatile memory of ROM type;- a RAM;- one or several I / O (Input / Output) devices such as for example a keyboard, a mouse, a trackball, a webcam; other modes for introduction of commands such as for example vocal recognition are also possible;- a power source; and- a radiofrequency unit.

[0178] According to a variant, the power supply is external to the apparatus.

[0179] The apparatus also comprises a display device of display screen type directly connected to the graphics card to display synthesized images calculated and composed in the graphics card. According to a variant, a display device is external to the apparatus and is connected thereto by a cable or wirelessly for transmitting the display signals. The apparatus, for example through the graphics card, comprises an interface for transmission or connection adapted to transmit a display signal to an external display means such as for example an LCD or plasma screen or a video -projector. In this respect, the RF unit can be used for wireless transmissions.

[0180] It is noted that the word "register" used hereinafter in the description of memories can designate in each of the memories mentioned, a memory zone of low capacity (some binary data) as well as a memory zone of large capacity (enabling a whole program to be stored or all or part of the data representative of data calculated or to be displayed). Also, the registers represented for the RAM and the GRAM can be arranged and constituted in any manner, and each of them does not necessarily correspond to adjacent memory locations and can be distributed otherwise (which covers notably the situation in which one register includes several smaller registers).

[0181] When switched-on, the microprocessor loads and executes the instructions of the program contained in the RAM.

[0182] As will be understood by a skilled person, the presence of the graphics card is not mandatory, and can be replaced with entire CPU processing and / or simpler visualization implementations.

[0183] In variant modes, the apparatus may include only the functionalities of the device 1, and not the learning capacities of the device 2. In addition, the device 1 and / or the device 2 may be implemented differently than a standalone software, and an apparatus or set of apparatus comprising only parts of the apparatus may be exploited through an API call or via a cloud interface.EXAMPLE

[0184] The present invention is further illustrated by the following example.Materials and Methods

[0185] The present example relies on a curated dataset comprising eight different paired Hematoxylin and Eosin (H&E) and Immunohistochemistry (IHC) stain images that is used to test a machine learning system trained using an unpaired training dataset. The H&E stain images and IHC stain images from the curated dataset were specific to pediatric Crohn’s disease at diagnosis (pre-treatment), with 30 slides available for each pair of H&E / IHC stain images. The digitized slides were scanned using a high-resolution scanner at a consistent magnification scale to maintain homogeneity across the dataset.

[0186] For data normalization, a standardized approach to reduce batch effects and any existing variability was applied. The preprocessing step involved automatic segmentation of the tissue regions from the digitized slides, followed by normalization of the H&E stain images to reduce batch effects and variability. Subsequently, non-overlapping patches from the normalized images were selected, with an appropriate balance between the different tissue types, to form a comprehensive representation of each stain type.

[0187] The deep learning model employed is designed around an encoder-decoder framework. A single encoder serving multiple decoders was used. The encoder is trained to identify and emphasize crucial regions in the H&E stain image tissues, while the decoders work on generating more precise synthetic stains based on the identified regions. For each decoder, a U-Net architecture was utilized given its effective performance inbiomedical image segmentation tasks. A self-inspection mechanism was also used in the system that uses the learned H&E distribution for real-time validation of the synthetic stains generated by the model. The dataset was divided into training dataset, validation dataset, and testing datasets, following a 70:15:15 percentage split. The training dataset was used to learn the model parameters, while the validation dataset was used for hyperparameter tuning and models election. The testing datasets were reserved to evaluate the final model’s performance.

[0188] In this example, a ComboGAN was employed as trainable machine learning system, originally designed for art style transfer, owing to its prominent scalability features. It was validated on Whole Slide Images (WSIs). In order to instill trust in the generated synthetic stains, XAI methods are integrated. These strategies are instrumental in uncovering the model’s decision-making process, thereby fostering confidence in the model’s predictions and enhancing the trustworthiness of the synthetic stains.

[0189] Training the trainable machine learning system involved a comprehensive strategy that employs annotation-free knowledge, loss functions, and regularization. Different loss functions and regularization strategies were incorporated to optimize the learning process. This example focuses on adaptation of the ComboGAN and CycleGAN architectures to transform images of H&E- stained histological slides into multiple images of other stain types (e.g. IHC stain images), thereby tailoring it specifically for multiplex immunohistochemical (IHC) staining images synthesis.

[0190] In computational model training for virtual IHC stain images synthesis, traditional loss functions and mean square error (MSE), have notable limitations. Chief among these is their equal treatment of all slide regions, whether it’s tissue, or IHC- activated areas. This approach is far from ideal, particularly when handling the intrinsic staining imbalance typical of IHC slides. Within the framework, the pivotal element is the innovative loss function, LIHC. Crafted to address the distinct challenges of IHC- stained slides, it mirrors the function of the cycle consistency loss. However, what sets LIHC apart is its keen consideration of the nuances of IHC staining.

[0191] Building on the aforementioned limitations of traditional loss functions in the synthesis of virtual IHC stain images, the LIHC loss is a game-changer. Its primaryadvantage lies in its nuanced differentiation between IHC-activated regions (foreground) and IHC-non-activated regions (background). By introducing dynamic weighting factors, LIHC ensures a more balanced representation of these regions in the synthesized virtual IHC stain images. This balance is pivotal for accurate representation and interpretation, as real IHC slides usually exhibit staining imbalances. Utilizing the foreground and background masks, LIHC offers a more refined loss computation, optimizing the model towards faithfully recreating the unique features of real IHC- stained slides. In the present approach, the total loss Ltotai is not only determined by the IHC loss LIHC, but also includes adversarial Ladv and regularization Lregcomponents. Specifically, the adversarial component aims to ensure that generated images cannot be distinguished from real ones, based on the principles of a min-max adversarial game. In terms of implementation, the discriminative models are trained in a manner similar to CycleGAN. The discriminative models are optimized using both actual stained and synthetic tiles, with training procedures carried out independently for each discriminative model. Experimental validation confirms its effectiveness, revealing a significant improvement in stain transformation accuracy. This validates LIHC as a robust and reliable loss function for IHC stain images analysis (virtual or real), thus enhancing both the model’s reliability and performance. As described previously, the present methodology is based on an encoderdecoder architecture, within which a dedicated generative model is designed to handle H&E stain images was used, alongside other stain types. This architectural enhancement not only yields performance improvements by enforcing regularization constraints on the encoder to preserve essential H&E information — critical for accurate tissue analysis — but also provides us a PatchGAN discriminative model explicitly designed to handle H&E stain images. To evaluate the effectiveness of the present model for impurity detection, a novel experimental setup was incorporated. Specifically, localized degradation to the H&E stain images was introduced, which simulates variations in radial distortion and chromatic aberrations, thereby creating artificial impurities. These artificially introduced irregularities are designed to challenge the model’s discriminatory capabilities.

[0192] To further extend the practical utility of the present approach, a cloud-based virtual staining system has been implemented. This system allows users to upload Whole Slide Images (WSI), generate virtual IHC stain images, and provide feedback. This user-centric design helps refining the present model based on real-world usage and feedback, ultimately leading to a more robust and reliable tool for computational pathology.

[0193] The experimental setup encompasses a series of processes, including data collection and preprocessing, model configuration, and procedures for training and testing the computational pathology model from this example.Results

[0194] The present example resulted in significant contributions to computational pathology, yielding promising results for stain transformations. Through rigorous testing and validation, it has been demonstrated the effectiveness of a unified H&E encoder in identifying critical regions within H&E tissue samples. This is a crucial advancement for the generation of highly accurate synthetic stains, made possible by the specialized decoding algorithms herein disclosed.

[0195] In order to ensure a fair comparison, the same encoder, decoder, and discriminative models for both ComboGAN and CycleGAN were used, maintaining identical configurations such as random seed and training time. This standardized approach enabled to focus solely on the benefits of the unified H&E encoder. Importantly, methods mentioned in related works were not included. Introducing these methods would have added multiple variables, such as changes in architecture and different loss functions, thereby complicating the analysis and making it difficult to derive conclusive results.

[0196] A quantitative evaluation reveals a notable advantage of the present system. Specifically, the synthetic stains generated using the present approach significantly outperformed those created via the CycleGAN method. To provide a concrete metric for this performance advantage, a MSE for comparing synthetic stains against their actual counterparts was used. Remarkably, the ComboGAN method, when tailored for virtual staining applications, reduces the MSE by an impressive 31.5% in a paired setting. This substantial improvement underscores ComboGAN’ s efficacy in generating synthetic stains that closely emulate real ones.

[0197] Additionally, the present approach yields significant gains in computational efficiency. By using a single encoder, decoder, and discriminative model for the entireH&E staining process, the present method requires 43.77% fewer trainable parameters compared to alternative techniques. This streamlined architecture not only enhances computational efficiency but also paves the way for scalable deployment. The present system can readily support a wide array of output stains and significantly accelerates the training process.

[0198] In summary, the ComboGAN methodology excels in two key aspects: it produces synthetic stains with higher accuracy and accomplishes this with enhanced computational efficiency. These attributes make it an ideal, scalable, and easy-to-train solution for generating synthetic stains in histopathological studies.

[0199] Moreover, the self-inspection feature played a crucial role in establishing trust in synthetic stains. The system from this example successfully validated the alignment of new H&E slides with the trained H&E distribution, demonstrating that the quality of synthetic stained slides either matched or surpassed the test distribution. This was validated by performing the Kolmogorov-Smirnov test between the distributions of the synthetic and real stains, with p-values exceeding 0.05, indicating a lack of significant difference between the two.

[0200] Additionally, the present method was designed to leverage the information present in the stained slides to improve the trust worthiness and robustness of the virtual staining training. This via the loss functions during the training phase. Where the activated regions are identified, then used to spatially weight the loss function. A consistent reduction in the loss function values across multiple iterations was observed, indicating that the system was able to learn and improve effectively. Moreover, the current approach resulted in an improvement in the overall accuracy of stain transformations across both paired and

[0201] The context-driven approach from the present invention was highly effective in capturing the complexity and variability of staining patterns, which was validated by pathologists’ assessments and quantitative metrics. For instance, the application of context-oriented virtual staining resulted in an improved understanding of patterns within the stains, as reflected in the lower error rates in the context-oriented settings compared to random magnification scales.

[0202] The empirical investigations validate the discriminative prowess of the H&E- adapted PatchGAN discriminative models ineffectively identifying these artificial impurities within H&E samples and precisely delineating their locations. In a specific experimental configuration, the PatchGAN discriminative model defines a distinct "fake" region within the degraded segment while consistently categorizing the remaining portions of the image as genuine.

[0203] This strategic adaptation of the PatchGAN discriminative model for H&E impurity detection holds significant promise for assessing the quality of input H&E samples and offers a valuable tool for medical diagnostics and research.

[0204] The cloud-based virtual staining system was successfully used by numerous pathologists, indicating its high accessibility and usability. Feedback from users was consistently positive, suggesting that the system was efficient and user-friendly.

[0205] Moreover, the feedback mechanism helped continually refine the tool, further enhancing its performance.

[0206] The new dataset of paired H&E / IHC stains specific to pediatric Crohn’s disease has been well received by the scientific community, stimulating further research in this area. This data will be instrumental in future advancements in computational pathology.

[0207] The innovative advancements introduced by the present work contribute significantly to the realm of computational pathology.

[0208] The focus on scalability, accuracy, trust, and practicality in the context of stain transformations has given rise to a robust and adaptable methodology that demonstrates excellent potential for real- world application.

[0209] A primary achievement is the introduction of a unified H&E encoder serving multiple stain decoders. The results obtained herein indicate this strategy amplifies the system’s ability to identify critical tissue regions, thereby enhancing synthetic stain precision. This finding aligns with and extends upon previous literature which suggests a mutual enhancement in paired staining and encoding techniques.

[0210] Moreover, the development of a self-inspection feature for real-time validation represents a significant stride towards this direction, ensuring the system generatessynthetic stains of optimal quality. The role of trustworthiness in medical imaging has been well documented in existing literature, highlighting the importance of the findings in this area. The utilization of loss functions and regularization, integrated within the training phase, exhibits significant performance improvements, in addition to enhancing the system’s trustworthiness. The present findings resonate with previous studies that have identified the value of annotation-free knowledge in model training Further, the scope was broadened by employing loss regularization, which ensures accuracy regardless of the chosen training setting.

[0211] The present approach, emphasizing context in virtual staining, offers a promising contribution to the field. By emulating the complex practices of pathologists, a deeper comprehension of IHC staining patterns was facilitated. This method substantiates recent research stressing the importance of context-oriented learning Finally, the cloud-based virtual staining computing and curation of a new dataset of pediatric Crohn’s disease are major milestones in the present study. By mitigating technical barriers and providing an expansive, disease-specific dataset, it paved the way for broader access to and deeper research within computational pathology.

[0212] The synthetic stains’ biological fidelity could be further verified by more extensive comparative studies with manual staining. Second, while the system currently supports eight types of IHC stains, the scalability and adaptability of the framework to accommodate more diverse stain types require additional investigation.

[0213] In terms of future research, the expansion of the dataset to include additional pathological conditions beyond pediatric Crohn’s disease could enhance the model’s generalizability. Further exploration of deep learning architectures and training methodologies could also bolster the system’s overall performance. Lastly, comprehensive user studies

[0214] involving pathologists would be valuable to gain insights into the system’s usability and potential areas of improvement.

Claims

CLAIMS1. A device (1) for generating a IHC machine learning system (41) for generating n virtual Immunohistochemical (IHC) stain images (61) from one Hematoxylin and Eosin (H&E) stain image (32), said device (1) comprising: at least one input configured to receive a training dataset of unpaired images comprising: o a H&E stain images ensemble (21) comprising a plurality of H&E stain images, and o n IHC stain images ensembles (22, 23, 24), wherein each ithIHC stain images ensemble comprises a plurality of IHC stain images obtained using an ithIHC stain, and wherein the ithIHC stain used to obtain the images of the ithIHC stain images ensemble is different from the IHC stain used to obtain the images of the others IHC stain images ensembles; at least one processor configured to train a trainable machine learning system (40) to obtain training parameters (42), wherein: o said trainable machine learning system (40) comprises:■ n IHC discriminative models (27) and a IHC generative model (30) comprising one H&E encoder (31) and n IHC decoders (32), and■ one H&E discriminative model (37) and a H&E generative model (35) comprising n IHC encoders (33) and one H&E decoder (36), and o during the training:■ the H&E encoder (31) from said IHC generative model (30) is fed with each image of the H&E stain images ensemble (21) from said training dataset and said IHC generative model (30) outputs, for each of said image of the H&E stain images ensemble (21), a first set of n virtual IHC stain images (71), wherein each ithIHC decoder (39) of the IHC generativemodel (30) provides as an output a ithvirtual IHC stain image (72) of the first set of n virtual IHC stain images (71), each of the ithIHC discriminative model (28) receiving as an input said ithvirtual IHC stain image (72) and the IHC stain images from the ithIHC stain images ensemble (23) from the training dataset; each ithIHC encoder (34) from the H&E generative model (35) receiving as input the ithvirtual IHC stain image (72) outputted by the ithIHC decoder; the H&E decoder (36) from said H&E generative model (35) providing as an output a first set of n virtual H&E stain images (73), the H&E discriminative model (37) receiving as an input said first set of n virtual H&E stain images (73) outputted by the H&E generative model (35) and the H&E stain images ensemble (21) from said training dataset; wherein each ithvirtual H&E stain image is generated using a latent representation generated by the ithIHC encoder (34) of the H&E generative model (35);■ each ithIHC encoder (34) from the H&E generative model (35) receives as an input the IHC stain images from the ithIHC stain images ensemble (23) from the training dataset and the ithvirtual IHC stain image (72) outputted by the ithIHC decoder (39), and the H&E decoder (36) from said H&E generative model (35) provides as an output a second set of n virtual H&E stain images (74), the H&E discriminative model (37) receiving as an input each image of the second set of n virtual H&E stain images (76) outputted by the H&E generative model (35), one by one, and the H&E stain images ensemble (21) from said training dataset; the H&E encoder (36) from said IHC generative model (30) further receiving as an input the second set of n virtual H&E stain images (76), said IHC generative model (30) providing as an output a second set of n virtual IHC stain images (77), wherein each ithIHC decoder (39) of the n IHC decoders (32) provides as an output a ithvirtual IHC stain image (78) of thesecond set of n virtual IHC stain images (77), the ithIHC discriminative model (28) receiving as an input the ithvirtual IHC stain image (72) outputted by the ithIHC decoder and the IHC stain images from the ithIHC stain images ensemble (23) from the training dataset; at least one output configured to provide the obtained training parameters (42) for said IHC machine learning system (41), wherein said IHC machine learning system (41) comprises the H&E encoder (31), the n IHC decoders (32) and the n IHC discriminative models (27).

2. The device (1) according to claim 1, wherein said at least one input is configured to receive a training dataset of paired images, said at least one processor being further configured, during training, to implement an adaptation phase wherein said trainable machine learning system (40) is re-trained using said training dataset of paired images.

3. The device (1) according to claim 1 or 2, wherein during training, the at least one processor is further configured to minimize a global generator loss functions comprising at least one consistency component, at least one adversarial component, at least one regularization component and at least one dynamic weighting factor.

4. The device (1) according to claim 3, wherein said dynamic weighting factor is calculated based on the total number of pixels in IHC-activated regions and in IHC- non-activated regions of the IHC stain images from the n IHC stain images ensembles (22, 23, 24).

5. The device (1) according to any one of claims 1 to 4, wherein during training, each IHC discriminative model of the n IHC discriminative models (27) is configured to output a prediction, the at least one processor being further configured to minimize n IHC discriminator loss functions associated with the n IHC discriminative models (27), wherein each ithIHC discriminator loss function comprises calculating, for each image fed to the corresponding ithIHC discriminativemodel (28), a difference between the prediction of said ithIHC discriminative model (28) and the image fed to said ithIHC discriminative model (28).

6. The device (1) according to claim 5, wherein said calculated difference is a pixelwise difference.

7. The device (1) according to any one of claims 1 to 6, wherein the H&E stain images from the H&E stain images ensemble (21) and / or the IHC stain images from the n IHC stain images ensembles (22, 23, 24) are obtained by capturing at least one microscope slide with a magnification of xlO.

8. A device (2) for generating n virtual Immunohistochemical (IHC) stain images (71) from one Hematoxylin and Eosin (H&E) stain image (25), using an IHC machine learning system (41) trained using the training parameters (42) obtained thanks to the device (1) of any of claims 1 to 7, wherein said device (2) comprises: at least one input configured to receive said H&E stain image (25); at least one processor configured to feed said H&E stain image (25) to said trained IHC machine learning system (41) and output n virtual IHC stain images (71); and at least one output configured to provide the n virtual IHC stain images (71).

9. The device (2) according to claim 8, wherein the IHC machine learning system (41) further comprises the H&E decoder (36) and the H&E discriminative model (37) and wherein the H&E encoder (31) and the H&E decoder (36) form a verification generative model (38), said verification generative model (38) being configured to receive said H&E stain image (25) and to output a virtual H&E stain image (26), said virtual H&E stain image (26) being fed to the H&E discriminative model (37) so as to obtain a H&E reliability score.

10. The device (2) according to claim 9, wherein when the obtained H&E reliability score does not satisfy a predefined reliability criterion, the at least one processor is further configured to generate a heatmap to provide visual Al explicability for a user.

11. The device (2) according to any one of claims 8 to 10, wherein the at least one processor is further configured to segment IHC-activated regions of at least two virtual IHC images from the n virtual IHC stain images (71) and overlay said segmented IHC-activated regions so as to obtain at least one virtual multiplex Immunohistochemical (IHC) stain image.

12. The device (2) according to any one of claim 8 to 11, wherein the at least one processor is further configured to: divide the H&E stain image (25) into a set of tiles, each tile being fed to said trained IHC machine learning system (41) so as to output n corresponding virtual IHC stain tiles; for each tile of the set, select the kth outputted virtual IHC stain tile so as to reconstruct the k,hvirtual IHC stain image; and output the reconstructed n virtual IHC stain images (71).

13. A computer-implemented method for generating a plurality of n virtual Immunohistochemical (IHC) stain images (71) from one Hematoxylin and Eosin (H&E) stain image (32), using a IHC machine learning system (41) trained using the training parameters (42) obtained according to any of claims 1 to 7, wherein said method comprises: receiving said H&E stain image (25); feeding said H&E stain image (25) to said IHC machine learning system (41) so as to generate said n virtual IHC stain images (71), and outputting said n virtual IHC stain images (71).

14. A computer program product comprising instructions which, when the program is executed by a computer, cause the computer to carry out the method of claim 13.

15. A computer-readable storage medium comprising instructions which, when executed by a computer, cause the computer to carry out the method of claim 13.

Citation Information

Patent Citations

  • Synthetic generation of immunohistochemical special stains

    US20230368504A1