Method for generating digital staining microscopic images
By digitally staining the fluorescent images of unlabeled tissues using deep neural networks, the time and cost of the histochemical staining process in the prior art was solved, rapid and economical tissue imaging was achieved, and the original state of the tissue sample was retained for subsequent analysis.
Patent Information
- Application Number
- CN202510094766.9
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Priority Date
- 2018-03-30
- Filing Date
- 2019-03-29
- Publication Date
- 2025-05-16
AI Technical Summary
The prior art is difficult to effectively bypass the histochemical staining process in tissue imaging, resulting in increased time and cost, while also limiting the possibility of preservation and subsequent analysis of tissue samples.
Using trained deep neural networks, especially convolutional neural networks (CNN) and generative adversarial networks (GAN) models, digital or virtual staining is used to use fluorescent images of unlabeled tissues to replace the traditional histochemical staining steps.
The rapid and economical generation of digital stained images matching the labeled tissue samples is achieved, saving time and cost, and retaining the original state of unlabeled tissue samples for subsequent analysis.
Smart Images

Figure CN120014640A_ABST
Abstract
Description
[0001] This application is a divisional application of the Chinese national phase application 201980029172.2 of the PCT application with an international filing date of March 29, 2019, an international application number of PCT / US2019 / 025020, and an invention name of “Method and system for digitally staining unlabeled fluorescent images using deep learning”, the entire contents of which are incorporated herein by reference.
[0002] CROSS-REFERENCE TO RELATED APPLICATIONS
[0003] This application claims priority to U.S. Provisional Patent Application No. 62 / 651,005, filed on March 30, 2018, which is incorporated herein by reference. Priority is claimed under 35 U.S.C. §119 and any other applicable statutes. Technical Field
[0004] The technical field generally relates to methods and systems for imaging unstained (i.e., unlabeled) tissue. In particular, the technical field relates to microscopy methods and systems for digitally or virtually staining images of unstained or unlabeled tissue using deep neural network learning. Deep learning (a class of machine learning algorithms) in neural networks is used to digitally stain images of unlabeled tissue sections into images that are equivalent to microscopy images of the same sample that have been stained or labeled. Background Art
[0005] Microscopic imaging of tissue samples is a fundamental tool for diagnosing a variety of diseases and forms a major tool in pathology and biology. The gold standard image of a tissue section determined clinically is the result of a laborious process that involves fixing the tissue specimen in formalin-fixed paraffin (FFPE), cutting it into thin slices (usually about 2-10 μm), labeling / staining and mounting it on a slide, and then imaging it microscopically using, for example, bright-field microscopy. All of these steps use a variety of reagents and produce irreversible effects on the tissue. Recently, there have been efforts to change this workflow using different imaging modalities. Attempts have been made to image fresh, non-paraffin-embedded tissue samples using nonlinear microscopy methods based on, for example, two-photon fluorescence, second harmonic generation, third harmonic generation, and Raman scattering. Other attempts have used controllable supercontinuum sources to acquire multimodal images for chemical analysis of fresh tissue samples. These methods require the use of ultrafast lasers or supercontinuum light sources, which may not be easy to use in most cases and require long scanning times due to the weak light signal. In addition to these, other microscopy methods have emerged for imaging unsectioned tissue samples by UV excitation of stained specimens or exploiting the fluorescence emission of biological tissues at short wavelengths.
[0006] Indeed, fluorescent signals have created some unique opportunities for imaging tissue samples by exploiting the fluorescence emitted from endogenous fluorophores. Such endogenous fluorescent markers have been shown to carry useful information that can be mapped to the functional and structural properties of biological samples and have therefore been widely used for diagnostic and research purposes. One of the major areas of focus of these efforts has been the spectroscopic study of the relationship between different biomolecules and their structural properties under different conditions. Some of these well-characterized biological components include vitamins (e.g., vitamin A, riboflavin, thiamine), collagen, coenzymes, fatty acids, etc.
[0007] Although some of the techniques discussed above have the unique ability to use various contrast mechanisms to distinguish, for example, cell types and subcellular components in a tissue sample, pathologists and tumor classification software are generally trained to examine "gold standard" stained tissue samples to make diagnostic decisions. Partially motivated by this, some of the techniques mentioned above have been enhanced to create pseudo-hematoxylin and eosin (H&E) images based on a linear approximation that relates the fluorescence intensity of the image to the dye concentration per tissue volume, using empirically determined constants to represent the average spectral response of the various dyes embedded in the tissue. These methods also use exogenous staining to enhance the contrast of the fluorescence signal to create a virtual H&E image of the tissue sample. Summary of the invention
[0008] In one embodiment, a system and method using a trained deep neural network is provided for digital or virtual staining of unlabeled thin tissue sections or other samples using fluorescence images obtained from chemically unstained tissue (or other samples). Chemically unstained tissue refers to the lack of standard stains or markers for tissue chemical staining. The fluorescence of chemically unstained tissue can include autofluorescence of tissue from naturally occurring or endogenous fluorescence or other endogenous light emitters, the frequency of which is different from the illumination frequency (i.e., frequency-shifted light). The fluorescence of chemically unstained tissue can further include fluorescence from tissues of exogenously added fluorescent markers or other exogenous light emitters. The sample is imaged using a fluorescence microscope such as a widefield fluorescence microscope (or a standard fluorescence microscope). The microscope can use a standard near-ultraviolet excitation / emission filter set or other excitation / emission light sources / filter sets known to those skilled in the art. In some embodiments, in a preferred embodiment, a single fluorescence image obtained from a sample is digitally or virtually stained using a trained deep neural network.
[0009] In one embodiment, the trained deep neural network is a convolutional neural network (CNN) trained using a generative adversarial network (GAN) model to match the bright field microscopic image of the corresponding tissue sample after marking the tissue sample with a certain tissue stain. In this embodiment, a fluorescent image of an unstained sample (e.g., tissue) is input into a trained deep neural network to produce a digital staining image. Therefore, in this embodiment, the tissue chemical staining and bright field imaging steps are completely replaced by the use of a trained deep neural network to generate a digital staining image. As explained herein, in some embodiments, using a standard desktop computer, using an imaging field of view of approximately 0.33mm×0.33mm using, for example, a 40x objective lens, the network inference speed performed by the trained neural network is less than one second. Using a 20x objective lens to scan the tissue, the network inference time reaches 1.9 seconds / mm 2 .
[0010] A deep learning-based digital / virtual histological staining method using autofluorescence has been demonstrated by imaging unlabeled human tissue samples including salivary glands, thyroid, kidney, liver, lung, and skin, where the trained deep neural network output can create equivalent images that substantially match images of the same sample labeled with three different stains, namely H&E (salivary glands and thyroid), Jones stain (kidney), and Masson's trichrome (liver and lung). Since the input images to the trained deep neural network were captured by a conventional fluorescence microscope with a standard filter set, this method has the transformative potential to use unstained tissue samples for pathology and histology applications, completely bypassing the histochemical staining process, saving time and the costs that go with it. This includes the labor costs involved in the staining process, reagents, additional time, etc. For example, for histological staining approximated using the digital or virtual staining processes described herein, each staining process for tissue sections takes an average of approximately 45 minutes (H&E) and 2-3 hours (Masson's trichrome and Jones staining), with an estimated cost (including labor) of $2-5 for H&E and over $16-35 for Masson's trichrome and Jones staining. In addition, some of these histochemical staining processes require time-sensitive steps that require an expert to monitor the process under a microscope, which not only makes the entire process lengthy and relatively expensive, but also laborious. The systems and methods disclosed herein bypass all of these staining steps and also allow for the preservation of unlabeled tissue sections for later analysis, such as micro-labeling of sub-regions of interest on unstained tissue specimens, which can be used for more advanced immunochemical and molecular analysis to facilitate, for example, customized treatment. In addition, a group of pathologists blindly evaluated the staining effects of the method on whole slide images (WSIs) corresponding to some of these samples, and they were able to identify tissue pathological features using digital / virtual staining techniques that were highly consistent with histological staining images of the same sample.
[0011] Furthermore, this deep learning-based digital / virtual histology staining framework can be broadly applied to other excitation wavelengths or fluorescence filter sets, as well as to other microscopy modalities that exploit other endogenous contrast mechanisms (e.g., nonlinear microscopy). In the experiments, sectioned and fixed tissue samples were used for meaningful comparisons with the results of standard histochemical staining procedures. However, the proposed method is also applicable to unfixed, unsectioned tissue samples, making it suitable for use in the operating room or biopsy site for rapid diagnosis or remote pathology applications. In addition to its clinical applications, this method can also broadly benefit the field of histology and its applications in life science research and education.
[0012] In one embodiment, a method for generating a digitally stained microscopic image of an unlabeled sample comprises: utilizing a trained deep neural network run by image processing software using one or more processors of a computing device, wherein the trained deep neural network is trained with multiple matching chemical staining images or image patches of the same sample and their corresponding fluorescent images or image patches. The unlabeled sample may include tissues, cells, pathogens, biological fluid smears, or other micro-objects of interest. In some embodiments, a combination of one or more tissue types / chemical staining types may be used to train the deep neural network. For example, this may include tissue type A with stain #1, stain #2, stain #3, etc. In some embodiments, a deep neural network may be trained using tissues that have been stained with multiple stains.
[0013] The fluorescence image of the sample is input to a trained deep neural network. Then, the trained deep neural network outputs a digitally stained microscopic image of the sample based on the input sample fluorescence image. In one embodiment, the trained deep neural network is a convolutional neural network (CNN). This can include a CNN using a generative adversarial network (GAN) model. The fluorescence input image of the sample is obtained using a fluorescence microscope and an excitation light source (e.g., an ultraviolet or near-ultraviolet emitting light source). In some alternative embodiments, multiple fluorescence images are input to a trained deep neural network. For example, one fluorescence image can be obtained at a wavelength or wavelength range of a first filter, and another fluorescence image can be obtained at a wavelength or wavelength range of a second filter. The two fluorescence images are then input to a trained deep neural network to output a single digital / virtual stained image. In another embodiment, the obtained fluorescence image can be subjected to one or more linear or nonlinear preprocessing operations selected from contrast enhancement, contrast inversion, and image filtering, and these preprocessing operations can be input to a trained deep neural network alone or together with the obtained fluorescence image.
[0014] For example, in another embodiment, a method for generating a digital stain microscopic image of an unlabeled sample includes: providing a trained deep neural network executed by image processing software using one or more processors of a computing device, wherein the trained deep neural network is trained with multiple matching chemical stain images or image patches of the same sample and their corresponding fluorescent images or image patches. A first fluorescent image of the sample is obtained using a fluorescent microscope, and wherein the endogenous fluorescence of the frequency-shifted light or other endogenous emitter in the sample emits fluorescence of a first emission wavelength or wavelength range. A second fluorescent image of the sample is obtained using a fluorescent microscope, and wherein the endogenous fluorescence of the frequency-shifted light or other endogenous emitter in the sample emits fluorescence of a second emission wavelength or wavelength range. The first fluorescent image and the second fluorescent image can be obtained by using different excitation / emission wavelength combinations. The first fluorescent image and the second fluorescent image of the sample are then input into the trained deep neural network, which outputs a digital stain microscopic image of the sample, which is substantially equivalent to the corresponding bright field image of the same sample that has been chemically stained.
[0015] In another embodiment, a system for generating a digital stained microscopic image of a chemically unstained sample includes a computing device having image processing software executed thereon or by it, the image processing software including a trained deep neural network executed using one or more processors of the computing device. The trained deep neural network is trained with a plurality of matched chemically stained images or image patches and their corresponding fluorescent images or image patches of the same sample. The image processing software is configured to receive one or more fluorescent images of the sample and output a digital stained microscopic image of the sample that is substantially equivalent to a corresponding bright field image of the same sample that has been chemically stained. BRIEF DESCRIPTION OF THE DRAWINGS
[0016] Figure 1 A system for generating a digitally / virtually stained output image of a sample from an unstained microscope image of the sample is schematically shown according to one embodiment.
[0017] Figure 2 Schematic diagram showing deep learning-based digital / virtual histology staining operations using fluorescent images of unstained tissue.
[0018] FIG. 3A to FIG. 3H The digital / virtual staining results that match the chemically stained H&E samples are shown. The first two (2) columns ( Figure 3A and Figure 3E ) shows autofluorescence images of unstained salivary gland tissue sections (used as input to the deep neural network), and the third column ( Figure 3C and Figure 3G ) shows the digital / virtual staining results. The last column ( Figure 3D and Figure 3H ) shows a bright field image of the same tissue section after the histochemical staining process. Figure 3C and Figure 3D Evaluation of the fibrous adipose tissue in the subcutaneous tissue shows small islands of infiltrating tumor cells. Note that the details of the nuclei, including the nucleoli ( Figure 3C and Figure 3D Arrows) and chromatin structure. Figure 3G and Figure 3H In the , H&E staining shows invasive squamous cell carcinoma. In both stains / panels, there is a deproliferative reaction with edematous glial changes in the adjacent stroma ( Figure 3G and Figure 3H ) are clearly identifiable.
[0019] 4A to 4H Digital / virtual staining results are shown to match the chemically stained Jones sample. The first two (2) columns ( Figure 4A , Figure 4E ) shows autofluorescence images of unstained kidney tissue sections (used as input to the deep neural network), and the third column ( Figure 4C and Figure 4G ) shows the digital / virtual staining results. The last column ( Figure 4D , Figure 4H ) shows a bright field image of the same tissue section after the histochemical staining process.
[0020] Figures 5A to 5P Results of digital / virtual staining of liver and lung tissue sections are shown to match the trichrome staining of Masson's trichrome. The first two (2) columns show unstained liver tissue sections (rows 1 and 2 - Figure 5A , Figure 5B , Figure 5E , Fig. 5F ) and unstained lung tissue sections (rows 3 and 4 - Fig.5I , Figure 5J , Figure 5M , Figure 5N ) is used as the input of the deep neural network. The third column ( Figure 5C , Figure 5G , Figure 5K , Fig.5O ) show the digital / virtual staining results of these tissue samples. The last column ( Figure 5D , Figure 5H , Figure 5L , Figure 5P ) shows a bright field image of the same tissue section after the histochemical staining process.
[0021] Fig. 6AA plot showing the combined loss function versus the number of iterations for random initialization and transfer learning initialization. Fig. 6A Shows how to achieve superior convergence using transfer learning. A new deep neural network is initialized using weights and biases learned from salivary gland tissue sections to achieve H&E virtual staining of thyroid tissue. Transfer learning converges faster than random initialization and also achieves lower local minima.
[0022] Figure 6B Images of the network output at different stages of the learning process are shown, for both random initialization and transfer learning, to better illustrate the impact of transfer learning on translating the proposed method to new tissue / staining combinations.
[0023] Figure 6C Corresponding bright field images of H&E chemical staining are shown.
[0024] Fig. 7A Virtual staining of skin tissue using only the DAPI channel (H&E staining) is shown.
[0025] Figure 7B Virtual staining of skin tissue using DAPI and Cy5 channels (H&E staining) is shown. Cy5 refers to a far-red fluorescent marker cyanine dye used to label biomolecules.
[0026] Figure 7C Corresponding histologically stained (ie, chemically stained with H&E) tissues are shown.
[0027] Figure 8 Shown is the process of field matching and registration of an autofluorescence image of an unstained tissue sample relative to a bright field image of the same sample after a chemical staining process.
[0028] Fig. 9 Schematic illustration of the training process of the virtual coloring network using GAN.
[0029] Fig.10 A generative adversarial network (GAN) architecture for the generator and discriminator is shown according to one embodiment. DETAILED DESCRIPTION
[0030] Figure 1One embodiment of a system 2 for outputting a digital staining image 40 from an input microscope image 20 of a sample 22 is schematically shown. As described herein, the input image 20 is a fluorescent image 20 of a sample 22 (such as a tissue in one embodiment) that is not stained or labeled with a fluorescent stain or marker. That is, the input image 20 is an autofluorescence image 20 of the sample 22, wherein the fluorescence emitted by the sample 22 is the result of one or more endogenous fluorescence or other endogenous emitters of frequency-shifted light contained therein. Frequency-shifted light is light emitted at a frequency (or wavelength) different from the incident frequency (or wavelength). Endogenous fluorescence or endogenous emitters of frequency-shifted light may include molecules, compounds, complexes, molecular species, biomolecules, pigments, tissues, etc. In some embodiments, the input image 20 (e.g., a raw fluorescent image) is subjected to one or more linear or nonlinear preprocessing operations selected from contrast enhancement, contrast inversion, image filtering. The system includes a computing device 100 including one or more processors 102 and image processing software 104 incorporating a trained deep neural network 10 (e.g., a convolutional neural network as explained herein in one or more embodiments). As explained herein, computing device 100 may include a personal computer, a laptop computer, a mobile computing device, a remote server, etc., although other computing devices (e.g., devices containing one or more graphics processing units (GPUs)) or other application specific integrated circuits (ASICs) may be used. GPUs or ASICs may be used to accelerate training and final image output. Computing device 100 may be associated with or connected to a monitor or display 106 for displaying digital staining images 40. Display 106 may be used to display a graphical user interface (GUI), which a user may use to display and view digital staining images 40. In one embodiment, a user may be able to manually trigger or switch between multiple different digital / virtual stainings of a particular sample 22 using, for example, a GUI. Alternatively, triggering or switching between different stainings may be automatically accomplished by computing device 100. In a preferred embodiment, trained deep neural network 10 is a convolutional neural network (CNN).
[0031] For example, in a preferred embodiment as described herein, a GAN model is used to train a trained deep neural network 10. In the deep neural network 10 trained by GAN, two models are used for training. The generative model is used to capture the data distribution, while the second model estimates the probability that the sample comes from the training data rather than from the generative model. Detailed information about GAN can be found in Goodfellow et al., Generative Adversarial Nets, Advances in Neural Information Processing Systems, Vol. 27, pp. 2672-2680 (2014), cited here as a reference. Network training of a deep neural network 10 (e.g., GAN) can be performed on the same or different computing devices 100. For example, in one embodiment, a personal computer can be used to train a GAN, although such training may take a lot of time. In order to speed up this training process, one or more dedicated GPUs can be used for training. As explained herein, such training and testing are performed on a GPU obtained from a commercially available graphics card. Once the deep neural network 10 is trained, it can be used or executed on a different computing device 110, which may include a computing device with fewer computing resources for the training process (although a GPU can also be integrated into the execution of the trained deep neural network 10).
[0032] The image processing software 104 can be implemented using Python and TensorFlow, although other software packages and platforms can be used. The trained deep neural network 10 is not limited to a specific software platform or programming language, and the trained deep neural network 10 can be executed using any number of commercially available software languages or platforms. The image processing software 104 that is combined with or runs in conjunction with the trained deep neural network 10 can run in a local environment or a cloud-free environment. In some embodiments, certain functions of the image processing software 104 can run in a specific language or platform (e.g., image normalization), while the trained deep neural network 10 can run in another specific language or platform. Nevertheless, both operations are performed by the image processing software 104.
[0033] like Figure 1As shown, in one embodiment, the trained deep neural network 10 receives a single fluorescence image 20 of an unlabeled sample 22. In other embodiments, for example, using multiple excitation channels (see the discussion of melanin herein), there may be multiple fluorescence images 20 of the unlabeled sample 22 (which are input into the trained deep neural network 10 (e.g., one image per channel)). The fluorescence image 20 may include a wide field of view fluorescence image 20 of the unlabeled tissue sample 22. Wide field of view means that a wide field of view (FOV) is obtained by scanning a smaller FOV, and the wide field of view is between 10-2,000 mm 2 For example, a smaller FOV can be obtained by using the image processing software 104 to digitally stitch together smaller FOVs to create a scanning fluorescence microscope 110 with a wider FOV. For example, a wide FOV can be used to obtain a whole slide image (WSI) of the sample 22. A fluorescence image is obtained using the imaging device 110. For the fluorescence embodiments described herein, it may include a fluorescence microscope 110. The fluorescence microscope 110 includes an excitation light source that illuminates the sample 22, and one or more image sensors (e.g., CMOS image sensors) for capturing fluorescence emitted by fluorophores or other endogenous emitters of frequency-shifted light contained in the sample 22. In some embodiments, the fluorescence microscope 110 may include the ability to illuminate the sample 22 with excitation light of multiple different wavelengths or wavelength ranges / bands. This can be achieved using multiple different light sources and / or different filter sets (e.g., standard ultraviolet or near-ultraviolet excitation / emission filter sets). In addition, in some embodiments, the fluorescence microscope 110 may include multiple filter sets that can filter different emission bands. For example, in some embodiments, multiple fluorescent images 20 may be captured, each fluorescent image 20 being captured at a different emission frequency band using a different filter set.
[0034] In some embodiments, sample 22 may include a portion of tissue disposed on or in substrate 23. In some embodiments, substrate 23 may include an optically transparent substrate (e.g., a glass or plastic slide, etc.). Sample 22 may include a tissue slice that is cut into thin slices using a microtome device, etc. Thin slices of tissue 22 may be considered as weakly scattering phase objects with limited amplitude contrast modulation under bright field illumination. Sample 22 may be imaged with or without a coverslip / cover glass. The sample may involve a frozen section or a paraffin (wax) section. Tissue sample 22 may be fixed (e.g., using formalin) or unfixed. Tissue sample 22 may include mammalian (e.g., human or animal) tissue or plant tissue. Sample 22 may also include other biological samples, environmental samples, etc. Examples include particles, cells, organelles, pathogens, or other microscale objects of interest (objects with a size of microns or less). Sample 22 may include smears of biological fluids or tissues. These include, for example, blood smears, Pap smears. As described herein, for fluorescence-based embodiments, sample 22 includes one or more naturally occurring or endogenous fluorophores that fluoresce and are captured by fluorescence microscopy device 110. Most plant and animal tissues exhibit autofluorescence when excited with ultraviolet or near-ultraviolet light. For example, endogenous fluorophores may include proteins such as collagen, elastin, fatty acids, vitamins, flavins, porphyrins, lipofuscin, coenzymes (e.g., NAD(P)H). In some optional embodiments, exogenously added fluorescent markers or other exogenous light emitters may also be added. As explained herein, sample 22 may also include other endogenous emitters of frequency-shifted light.
[0035] The trained deep neural network 10 outputs or generates a digitally stained or labeled output image 40 in response to the input image 20. The digitally stained output image 40 has a "stain" that has been digitally integrated into the stained output image 40 using the trained deep neural network 10. In some embodiments, such as those involving tissue sections, the trained deep neural network 10 appears to a skilled observer (e.g., a trained histopathologist) to be substantially equivalent to a corresponding bright field image of the same tissue section sample 22 that has been chemically stained. Indeed, as described herein, experimental results obtained using the trained deep neural network 10 indicate that a trained pathologist is able to identify histopathological features using both staining techniques (chemical staining and digital / virtual staining) and a high degree of agreement between the two techniques, without a clear preferred staining technique (virtual vs. histology). This digital or virtual staining of the tissue section sample 22 looks just like the tissue section sample 22 has been histochemically stained (even though no such staining operation has been performed).
[0036] Figure 2The operations involved in a typical fluorescence-based embodiment are schematically shown. Figure 2 As shown, a sample 22, such as an unstained tissue section, is obtained. This can be obtained from living tissue, such as by biopsy B. The unstained tissue section sample 22 is then fluorescently imaged using a fluorescence microscope 110, and a fluorescence image 20 is generated. This fluorescence image 20 is then input into a trained deep neural network 10, which then immediately outputs a digital staining image 40 of the tissue section sample 22. This digital staining image 40 is very similar to the appearance of a bright field image of the same tissue section sample 22, while the actual tissue section sample 22 is subjected to histochemical staining. Figure 2 A conventional process is shown (using dashed arrows) in which a tissue section sample 22 is subjected to histochemical staining 44 and then to conventional bright field microscopy 46 to produce a conventional bright field image 48 of the stained tissue section sample 22. Figure 2 As shown, the digital staining image 40 is very similar to the actual chemical staining image 48. Similar resolution and color profiles can be obtained using the digital staining platform described herein. Figure 1 As shown, the digital stain image 40 can be shown or displayed on a computer monitor 106, but it should be understood that the digital stain image 40 can be displayed on any suitable display (e.g., a computer monitor, a tablet computer, a mobile computing device, a cell phone, etc.). A GUI can be displayed on the computer monitor 106 so that a user can view and optionally interact with the digital stain image 40 (e.g., zoom, crop, highlight, mark, adjust exposure, etc.).
[0037] Experimental – Digital staining of unlabeled tissue using autofluorescence
[0038] Virtual staining of tissue samples
[0039] The system 2 and method described herein were tested and demonstrated using different combinations of tissue section samples 22 and stains. After training the CNN-based deep neural network 10, its inference was blindly tested by feeding it autofluorescence images 20 of unlabeled tissue sections 22 that did not overlap with the images used in the training or validation sets. 4A to 4H Results are shown for a salivary gland tissue section that was digitally / virtually stained to match an H&E-stained brightfield image 48 (i.e., ground truth image) of the same sample 22. These results demonstrate the ability of the system 2 to convert a fluorescent image 20 of an unlabeled tissue section 22 into a brightfield equivalent image 40 that displays the correct color scheme expected for H&E-stained tissue, including various components such as epithelial cells, nuclei, nucleoli, matrix, and collagen. Figure 3C and Figure 3DEvaluation of the H&E stain demonstrates small islands of infiltrating tumor cells within the subcutaneous fibroadipose tissue. Note that nuclear detail, including nucleoli (arrows) and distinction in chromatin texture, is clearly appreciated in both panels. Similarly, in Figure 3G and Figure 3H In the figure, H&E staining demonstrates invasive squamous cell carcinoma. A plasticizing reaction with edematous glial changes (asterisks) in the adjacent stroma can be clearly identified in both stains.
[0040] Next, the deep network 10 was trained to digitally / virtually stain other tissue types with two different stains, namely Jones dimethylamine methyl silver stain (kidney) and Masson's trichrome stain (liver and lung). 4A to 4H and Figures 5A to 5P The results of deep learning-based digital / virtual staining of these tissue sections 22 are summarized, which closely match the bright field images 48 of the same sample 22 captured after the histochemical staining process. These results demonstrate that the trained deep neural network 10 is able to infer staining patterns of different types of histological stains for different tissue types from a single fluorescent image 20 of an unlabeled specimen (i.e., without any histochemical staining). FIG. 3A to FIG. 3H With the same general conclusion, the pathologists also confirmed that Figure 4C and Figure 5G The neural network output in the image correctly shows the histological features corresponding to hepatocytes, sinusoidal spaces, collagen, and fat droplets ( Figure 5G ), consistent with the way they appear in a bright field image 48 of the same tissue sample 22 captured after chemical staining ( Figure 5D and Figure 5H ). Similarly, the same expert also confirmed that Figure 5K and Fig.5O The deep neural network output image 40 reported in (lung) reveals consistently stained histological features corresponding to blood vessels, collagen, and alveolar spaces as they appear in a bright field image 48 of the same tissue sample 22 imaged after chemical staining ( Figure 5L and Figure 5P ).
[0041] The digital / virtual stain output images 40 from the trained deep neural network 10 were compared to standard histochemical stain images 48 to diagnose various types of conditions on various types of tissues, which could be formalin-fixed paraffin-embedded (FFPE) or frozen sections. The results are summarized in Table 1 below. Fifteen (15) tissue sections were analyzed by four board-certified pathologists (who were unaware of the virtual staining technique) and the results showed 100% non-major disagreement, i.e., no significant differences in diagnosis between the professional observers. The "diagnosis time" varied greatly between observers, ranging from an average of 10 seconds per image for Observer 2 to 276 seconds per image for Observer 3. However, the intra-observer variability was small, and for all observers except Observer 2, there was a trend towards shorter diagnosis times using the virtually stained slide images 40, which was equal to, i.e., approximately 10 seconds per image for both the virtual slide images 40 and the histologically stained slide images 48. These indicate that the diagnostic utility between the two image modalities is very similar.
[0042] Table 1
[0043]
[0044]
[0045] Blind Assessment of Staining Efficacy on Whole Slide Images (WSI)
[0046] After evaluating the differences in tissue sections and staining, the capabilities of the Virtual Staining System 2 were tested in a dedicated staining histology workflow. In particular, the autofluorescence distribution of 15 unlabeled liver tissue section samples and 13 unlabeled kidney tissue sections were imaged with a 20x / 0.75NA objective. All liver and kidney tissue sections were obtained from different patients and included both small biopsies and larger resections. All tissue sections were obtained from FFPE but were not covered with coverslips. After autofluorescence scanning, tissue sections were histologically stained with Masson's trichrome (4μm liver tissue sections) and Jones stain (2μm kidney tissue sections). The WSIs were then divided into training and testing sets. For the liver section cohort, 7 WSIs were used to train the virtual staining algorithm and 8 WSIs were used for blind testing; for the kidney section cohort, 6 WSIs were used to train the algorithm and 7 WSIs were used for testing. The study pathologists were blinded to the staining technique used for each WSI and were asked to apply a 1-4 numerical rating scale for the quality of the different stains: 4 = perfect, 3 = very good, 2 = acceptable, and 1 = unacceptable. Second, the study pathologists applied the same scoring scale (1-4) for specific features: for liver only, nuclear detail (ND), cytoplasmic detail (CD), and extracellular fibrosis (EF). These results are summarized in Tables 2 (liver) and 3 (kidney) below (winners are in bold). The data demonstrate that pathologists were able to identify histopathological features using both staining techniques and a high degree of agreement between the two techniques without the need for a clear preferred staining technique (virtual vs. histological).
[0047] Table 2
[0048]
[0049] Table 3
[0050]
[0051] Quantification of network output image quality
[0052] Next, in addition to FIG. 3A to FIG. 3H , 4A to 4H , Figures 5A to 5PIn addition to the visual comparison provided in , the results of the trained deep neural network 10 are quantified by first calculating the pixel-level differences between the brightfield image 48 of the chemically stained sample 22 and the digital / virtual stained image 40 synthesized using the deep neural network 10 (without any labeling / staining). Table 4 below summarizes the comparison of different combinations of tissue types and staining using the YCbCr color space, where the chromaticity components Cb and Cr fully define the color, while Y defines the brightness component of the image. The comparison results show that the average differences between the two sets of images are < about 5% and < about 16% for the chromaticity (Cb, Cr) and brightness (Y) channels, respectively. Next, the comparison is further quantified using a second metric, namely the structural similarity index (SSIM), which is commonly used to predict the score given to an image by a human observer compared to a reference image (Equation 8). The range of SSIM is 0 to 1, where 1 defines the score of the same image. The results of this SSIM quantification are also summarized in Table 4, which well illustrates the strong structural similarity between the network output image 40 and the brightfield image 48 of the chemically stained sample.
[0053] Table 4
[0054]
[0055] It should be noted that the brightfield images 48 of the chemically stained tissue samples 22 do not actually provide a true gold standard for this particular SSIM and YCbCr analysis of the network output images 40, as the tissue undergoes uncontrolled changes and structural changes during the histochemical staining process and the associated dehydration and clearing steps. Another variation noted for some images was that the automated microscope scanning software selected different autofocus planes for the two imaging modalities. All of these variations present some challenges for absolute quantitative comparison of the two sets of images (i.e., network output 40 of unlabeled tissue vs. brightfield images 48 of the same tissue after the histological staining process).
[0056] Staining Standardization
[0057] An interesting byproduct of the digital / virtual staining system 2 may be staining standardization. In other words, the trained deep neural network 10 converges to a “common staining” coloring scheme, whereby the variation in the histologically stained tissue image 48 is higher than the variation in the virtually stained tissue image 40. The coloring of the virtual staining is simply a result of its training (i.e., the gold standard histological staining used during the training phase) and can be further adjusted according to the pathologist's preferences by retraining the network on new staining colorings. This “improved” training can be created from scratch or accelerated through transfer learning. Potential staining standardization using deep learning can correct the negative effects of inter-individual variation at different stages of sample preparation, establish common ground between different clinical laboratories, enhance clinicians' diagnostic workflow, and assist in the development of new algorithms, such as automated tissue metastasis detection or grading of different types of cancer.
[0058] Transfer of learning to other tissue-staining combinations
[0059] Using the concept of transfer learning, the training process for new tissues and / or staining types can converge faster while also reaching improved performance, i.e., better local minima of the training cost / loss function. This means that the deep neural network 10 can be initialized using pre-learned CNN models from different tissue-staining combinations to statistically learn virtual staining for new combinations. FIG. 6A to FIG. 6C The advantageous properties of this approach are illustrated by training a new deep neural network 10 to virtually stain an autofluorescence image 20 of an unstained thyroid tissue section and initializing it using the weights and biases of another deep neural network 10 that was previously trained for H&E virtual staining of salivary glands. Fig. 6A As shown, the evolution of the loss metric as a function of the number of iterations used in the training phase clearly shows that the new thyroid deep network10 converges quickly to a lower minimum compared to the same network architecture trained from scratch using random initialization. Figure 6B The output images 40 of the thyroid network 10 at different stages of its learning process are compared, which further illustrates the impact of transfer learning to quickly adapt the proposed method to new tissue / staining combinations. After the training phase (e.g., ≥6,000 iterations), the network output image 40 reveals that the cell nuclei show irregular contours, nuclear grooves, and pale chromatin, suggesting papillary thyroid carcinoma; the cells also show mild to moderate eosinophilic granular cytoplasm, and the fibrovascular core at the network output image shows an increase in inflammatory cells, including lymphocytes and plasma cells. Figure 6C A corresponding bright field image 48 of H&E chemical staining is shown.
[0060] Use multiple fluorescence channels at different resolutions
[0061] The method using the trained deep neural network 10 can be combined with other excitation wavelengths and / or imaging modalities to enhance its inference performance for different tissue components. For example, an attempt was made to detect melanin on a skin tissue section sample using virtual H&E staining. However, due to the weak autofluorescence signal of melanin at the DAPI excitation / emission wavelengths measured in the experimental system described in this article, melanin was not clearly identified in the output of the network. A potential way to increase the autofluorescence of melanin is to image the sample while it is in an oxidizing solution. However, a more practical alternative approach was used, where an additional autofluorescence channel (e.g., from a Cy5 filter (excitation 628nm / emission 692nm)) was used, which enabled the melanin signal to be enhanced and accurately inferred in the trained deep neural network 10. As 7A to 7C As shown in FIG. 1 , by using both the DAPI and Cy5 autofluorescence channels to train the network 10, the trained deep neural network 10 is able to successfully determine the location where melanin appears in the sample. In contrast, when only the DAPI channel is used ( Fig. 7A ), the network 10 cannot determine the areas containing melanin (these areas appear white). In other words, the network 10 uses the additional autofluorescence information from the Cy5 channel to distinguish melanin from background tissue. 7A to 7C The results shown use a low resolution objective (10x / 0.45NA) to acquire an image 20 for the Cy5 channel to complement a high resolution DAPI scan (20x / 0.75NA), assuming that the most necessary information is found in the high resolution DAPI scan, while other information (such as the presence of melanin) can be encoded with the lower resolution scan. In this way, two different channels are used, one of which is used to identify melanin at a lower resolution. This may require scanning the sample 22 multiple times with a fluorescence microscope 110. In yet another multi-channel embodiment, multiple images 20 can be fed to a trained deep neural network 10. For example, this may include a combination of an original fluorescence image with one or more images that have undergone linear or nonlinear image preprocessing (such as contrast enhancement, contrast inversion, and image filtering).
[0062] The system 2 and method described herein demonstrate the ability to digitally / virtually stain unlabeled tissue sections 22 using supervised deep learning techniques using as input a single fluorescence image 20 of the sample captured by a standard fluorescence microscope 110 and filter set (in other embodiments, multiple fluorescence images 20 are input when multiple fluorescence channels are used). This statistical learning-based approach has the potential to restructure the clinical workflow of tissue pathology and can benefit from a variety of imaging modalities, such as fluorescence microscopy, nonlinear microscopy, holographic microscopy, stimulated Raman scattering microscopy, and optical coherence tomography to provide a digital alternative to the standard practice of histochemical staining of tissue samples 22. Here, the method is demonstrated using fixed unstained tissue samples 22 to provide a meaningful comparison with chemically stained tissue samples, which is critical for training the deep neural network 10 and blindly testing the performance of the network output against clinically recognized methods. However, the proposed deep learning-based method is widely applicable to different types and states of samples 22, including unsectioned fresh tissue samples (e.g., according to a biopsy procedure), without the use of any labels or stains. After training, the deep neural network 10 can be used to digitally / virtually stain images of unlabeled fresh tissue samples 22 acquired using, for example, UV or deep UV excitation or even non-linear microscopy modalities. For example, Raman microscopy can provide very rich unlabeled biochemical features, which can further enhance the effectiveness of the virtual staining learned by the neural network.
[0063] An important part of the training process involves matching the fluorescent images 20 of the unlabeled tissue samples 22 and their corresponding bright field images 48 after the histochemical staining process (i.e., the chemically stained images). It should be noted that during the staining process and related steps, some tissue components may be lost or deformed, which can mislead the loss / cost function during the training phase. However, this is only a challenge related to training and validation, and there is no limitation on the practice of the trained deep neural network 10 for virtually staining the unlabeled tissue samples 22. In order to ensure the quality of the training and validation phases and minimize the impact of this challenge on the network performance, a threshold is established for the acceptable correlation value between the two sets of images (i.e., before and after the histochemical staining process), and unmatched image pairs are eliminated from the training / validation sets to ensure that the deep neural network 10 can learn the real signal, rather than the interference caused by the chemical staining process on the tissue morphology. In practice, the process of cleaning the training / validation image data can be done iteratively: one can start with roughly eliminating samples that are significantly changed and thus converge on the trained neural network 10. After this initial training phase, the output images 40 of each sample in the available image set can be screened relative to their corresponding bright field images 48 to set a more precise threshold to reject some other images and further clean up the training / validation image set. Through several iterations of this process, not only can the image set be further optimized, but the performance of the final trained deep neural network 10 can also be improved.
[0064] The above approach will alleviate some of the training challenges due to the random loss of some tissue features after the histological staining process. In fact, this highlights another motivation to skip the laborious and expensive steps involved in histochemical staining, as label-free methods can more easily preserve the histology of local tissues without the need for experts to handle some of the delicate procedures of the staining process and sometimes the need to observe the tissue under a microscope.
[0065] Using a desktop PC, the training phase of a deep neural network10 takes a considerable amount of time (e.g., about 13 hours for the salivary gland network). However, the entire process can be greatly accelerated by using specialized GPU-based computer hardware. FIG. 6A to FIG. 6CAs already emphasized in , transfer learning provides a “hot start” for the training phase of new tissue / staining combinations, making the entire process significantly faster. Once the deep neural network 10 is trained, the digital / virtual staining of the sample image 40 is performed in a single non-iterative manner, which does not require trial and error or any parameter adjustments to obtain the best results. Based on its feed-forward and non-iterative architecture, the deep neural network 10 quickly outputs a virtually stained image in less than one second (e.g., 0.59 seconds, corresponding to a sample field of view of approximately 0.33 mm×0.33 mm). Through further GPU-based acceleration, it is possible to achieve real-time or near real-time performance when outputting the digital / virtual staining image 40, which is particularly useful in operating room or in vivo imaging applications.
[0066] The implemented digital / virtual staining process is based on training a separate CNN deep neural network 10 for each tissue / staining combination. If autofluorescence images 20 with different tissue / staining combinations are fed into the CNN-based deep neural network 10, it will not function as expected. However, this is not a limitation because for histological applications, the tissue type and staining type are predetermined for each sample of interest 22, so the specific CNN selection used to create the digital / virtual staining image 40 from the autofluorescence image 20 of the unlabeled sample 22 does not require additional information or resources. Of course, more general CNN models can be learned for multiple tissue / staining combinations by, for example, increasing the number of training parameters in the model, but at the cost of potentially increasing training and inference time. Another avenue is the potential of system 2 and methods for performing multiple virtual stainings on the same unlabeled tissue type.
[0067] A significant advantage of System 2 is that it is extremely flexible. If diagnostic failures are detected through clinical comparison, feedback can be accommodated to statistically improve its performance by penalizing the failures found accordingly. This iterative training and transfer learning cycle will help optimize the robustness and clinical impact of the proposed method based on clinical evaluation of the network output performance. Finally, the method and System 2 can be used to perform micro-guided molecular analysis at the unstained tissue level by locally identifying regions of interest based on virtual staining and using this information to guide subsequent analysis of the tissue, such as micro-immunohistochemistry analysis or sequencing. Such virtual micro-guidance performed on unlabeled tissue samples can facilitate high-throughput disease subtype identification and also help develop customized therapies for patients.
[0068] Sample preparation
[0069] Formalin-fixed paraffin-embedded 2 μm thick tissue sections were deparaffinized using xylene and stained using Cytoseal TM(Thermo Fisher Scientific, Waltham, MA, USA) were mounted on standard slides and then covered with coverslips (Fisherfinest TM , 24x50-1, Fisher Scientific, Pittsburgh, PA, USA). After the initial autofluorescence imaging process of unlabeled tissue samples (using DAPI excitation and emission filter sets), the slides were placed in xylene for approximately 48 hours, or until the coverslip could be removed without damaging the tissue. Once the coverslip was removed, the slides were immersed (approximately 30 dips) in absolute alcohol, 95% alcohol, and then washed in DI water for approximately 1 minute. This step was followed by the corresponding staining procedures for H&E, Masson's trichrome, or Jones stain. This tissue processing path was only used for training and validation of the method and was no longer required after training the network. To test the system and method, different tissue and stain combinations were used: salivary gland and thyroid tissue sections were stained with H&E, kidney tissue sections were stained with Jones stain, and liver and lung tissue sections were stained with Masson's trichrome.
[0070] In the WSI study, 2–4 μm thick FFPE tissue sections were not covered with a coverslip during the autofluorescence imaging phase. After autofluorescence imaging, tissue samples were histologically stained as described above (Masson’s trichrome for liver and Jones’s for kidney tissue sections). Unstained frozen samples were prepared by embedding tissue sections in OCT (Tissue Technology, Finetek, Sakura, USA) and then immersing them in 2-methylbutane with dry ice. The frozen sections were then cut into 4 μm sections and placed in a freezer until imaging. After the imaging process, tissue sections were washed with 70% alcohol, H&E stained, and coverslipped. Samples were obtained from the Translational Pathology Core Laboratory (TPCL) and prepared by the Histology Laboratory at UCLA. Kidney tissue sections from diabetic and nondiabetic patients were obtained under IRB18-001029 (UCLA). All samples were obtained after de-identifying patient-related information and were prepared from existing samples. Therefore, this work did not interfere with standard operations of care or sample collection procedures.
[0071] Data collection
[0072] Unlabeled tissue autofluorescence images were captured using a conventional fluorescence microscope 110 (IX83, Olympus Corporation, Tokyo, Japan) equipped with a motorized stage, where the image acquisition process was controlled by microscope automation software (Molecular Devices, Inc.). Unstained tissue samples were excited with near-UV light and imaged using a DAPI filter (OSFI3-DAPI-5060C, excitation wavelength 377 nm / 50 nm bandwidth, emission wavelength 447 nm / 60 nm bandwidth) with a 40× / 0.95 NA objective (Olympus UPLSAPO 40X2 / 0.95 NA, WD0.18) or a 20× / 0.75 NA objective (Olympus UPLSAPO 20X / 0.75 NA, WD0.65). For melanin inference, additional autofluorescence images of the samples were acquired using a Cy5 filter (CY5-4040C-OFX, excitation wavelength 628 nm / 40 nm bandwidth, emission wavelength 692 nm / 40 nm bandwidth) with a 10× / 0.4NA objective (Olympus UPLSAPO10X2). Each autofluorescence image was captured with a scientific CMOS sensor (ORCA-flash4.0 v2, Hamamatsu Optoelectronics KK, Shizuoka, Japan) with an exposure time of approximately 500 ms. Brightfield images 48 (for training and validation) were acquired using a slide scanner microscope (Leica Biosystems) using a 20× / 0.75NA objective (Plan Apo) equipped with a 2× magnification adapter.
[0073] Image preprocessing and alignment
[0074] Since the deep neural network 10 aims to learn the statistical transformation between the autofluorescence image 20 of a chemically unstained tissue sample 22 and the bright field image 48 of the same tissue sample 22 after histochemical staining, it is important to accurately match the FOV of the input image and the target image (i.e., the unstained autofluorescence image 20 and the stained bright field image 48). Figure 8 The overall scheme describing the global and local image registration process is described in and has been implemented in MATLAB (The MathWorks Inc., Natick, MA, USA). The first step in this process is to find candidate features to match the unstained autofluorescence image and the chemically stained brightfield image. To this end, each autofluorescence image 20 (2048 × 2048 pixels) was downsampled to match the effective pixel size of the brightfield microscopy image. This resulted in a 1351 × 1351 pixel image of the unstained autofluorescence tissue, which was contrast enhanced by saturating the bottom 1% and top 1% of all pixel values and inverting the contrast ( Figure 820a) to better represent the color map of the grayscale converted whole slide image. Then, a correlation patch process 60 is performed, in which a normalized correlation score matrix is calculated by correlating each of the 1351×1351 pixel patches with a corresponding patch of the same size extracted from the whole slide grayscale image 48a. The highest scoring entry in this matrix represents the FOV that is most likely to match between the two imaging modalities. Using this information (defining a pair of coordinates), the matching FOV of the original whole slide brightfield image 48 is cropped 48c to create a target image 48d. Following this FOV matching procedure 60, the autofluorescence 20 and brightfield microscopy image 48 are roughly matched. However, due to slight mismatches in the sample placement of the two different microscopy imaging experiments (autofluorescence, followed by brightfield), they are still not accurately registered at the individual pixel level, which randomly results in a slight rotation angle (e.g., about 1-2 degrees) between the input image and the target image of the same sample.
[0075] The second part of the input object matching process involves a global registration step 64, which corrects for this small rotation angle between the autofluorescence and bright field images. This is done by extracting feature vectors (descriptors) and their corresponding positions from the image pairs and matching the features using the extracted descriptors. The M-estimated sample consensus (MSAC) algorithm, which is a variant of the random sample consensus (RANSAC) algorithm, is then used to find the transformation matrix corresponding to the matching pair. Finally, the angle-corrected image 48e is obtained by applying this transformation matrix to the original bright field microscope image patch 48d. After applying this rotation, the images 20b, 48e are further cropped by 100 pixels (50 pixels on each side) to accommodate the undefined pixel values at the image boundaries due to the rotation angle correction.
[0076] Finally, for the local feature registration operation 68, elastic image registration, the local features of the two sets of images (autofluorescence 20b and bright field 48e) are matched by hierarchical matching of corresponding blocks from large to small. A neural network 71 is used to learn the transformation between roughly matching images. This network 71 uses the same Fig.10 The same structure of the network 10 in . A small number of iterations is used so that the network 71 learns only the accurate color mapping without learning any spatial transformation between the input image and the labeled image. Finally, the transformation map calculated from this step is applied to each bright field image patch 48e. At the end of these registration steps 60, 64, 68, the autofluorescence image patches 20b and their corresponding bright field tissue image patches 48f are accurately matched to each other and can be used as input and label pairs for training the deep neural network 10, allowing the network to focus on and learn only the problem of virtual histology staining.
[0077] A similar process was used for the 20x objective images (used to generate the data for Tables 2 and 3). Instead of downsampling the autofluorescence images 20, the bright field microscopy images 48 were downsampled to 75.85% of their original size so that they matched the lower magnification images. In addition, additional shading correction and normalization techniques were applied to create whole slide images using these 20x images. Before being fed to the network 71, each field of view was normalized by subtracting the mean of the entire slide and dividing it by the standard deviation between pixel values. This normalizes the network input within each slide and between slides. Finally, shading correction was applied to each image to account for the lower relative intensities measured at the edges of each field of view.
[0078] Deep Neural Network Architecture and Training
[0079] In this work, a GAN architecture was used to learn the translation from an unlabeled unstained autofluorescence input image 20 to the corresponding bright field image 48 of a chemically stained sample. The training of a standard convolutional neural network aims to learn to minimize the loss / cost function between the network output and the target label. Therefore, the loss function 69 ( Fig. 9 and Fig.10 ) is a key component of deep network design. For example, simply choosing an l2-norm penalty as the cost function will tend to produce blurry results because the network averages the weighted probabilities of all plausible outcomes; therefore, additional regularization terms are often required to guide the network to retain the desired sharp sample features at the network output. GANs avoid this problem by learning a criterion that aims to accurately classify whether the deep network output image is real or fake (i.e., its virtual coloring is correct or incorrect). This makes it so that output images that are inconsistent with the desired labeling are not tolerated, which adapts the loss function to the data and the desired task at hand. To achieve this goal, the GAN training process involves training two different networks, such as Fig. 9 and Fig.10 As shown: (i) a generator network 70, whose purpose in this case is to learn the statistical transformation between an unstained autofluorescence input image 20 and a corresponding brightfield image 48 of the same sample 12 after a histological staining process; and (ii) a discriminator network 74, which learns how to discriminate between a true brightfield image of a stained tissue section and the output image of the generator network. Ultimately, the desired outcome of this training process is a trained deep neural network 10 that transforms an unstained autofluorescence input image 20 into a digital stained image 40 that will be difficult to distinguish from a stained brightfield image 48 of the same sample 22. To accomplish this task, the loss functions 69 for the generator 70 and the discriminator 74 are defined as follows:
[0080] l generator =MSE{zlabel ,z output}+λ×TV{z output}+α×(1-D(z output )) 2
[0081] l discrimnator =D(z output ) 2 +(1-D(z label )) 2 (0)
[0082] Where D is the output of the discriminator network, z label is a bright field image of chemically stained tissue, z output is the output of the generator network. The generator loss function is empirically set to different values (they adapt to pixel-level MSE loss of about 2% and about 20% and comprehensive generator loss of about 20%, respectively (l generator )) to balance the pixel-wise mean squared error (MSE) of the generator network output image relative to its label, the total variation (TV) operator of the output image, and the discriminator network prediction of the output image. The TV operator for image z is defined as:
[0083]
[0084] Where p, q are pixel indices. Based on equation (1), the discriminator attempts to minimize the output loss while maximizing the probability of correctly classifying the true marker (i.e., bright field image of chemically stained tissue). Ideally, the goal of the discriminator network is to achieve D(z label )=1 and D(z output )=0, but if GAN successfully trains the generator, then ideally D(z output ) will converge to 0.5.
[0085] exist Fig.10 The generator deep neural network architecture 70 is described in detail in . The network 70 processes the input image 20 in a multi-scale manner using downsampling and upsampling paths, thereby helping the network learn virtual coloring tasks at various scales. The downsampling path contains four separate steps (four blocks #1, #2, #3, #4), each of which contains a residual block, and each residual block converts the feature map x k Mapped to feature map x k+1 :
[0086] x k+1 =x k +LReLU[CONV k3 {LReLU[CONV k2 {LReLU[CONVk1 {x k}]}]}] (0)
[0087] where CONV{.} is the convolution operator (which includes the bias term), k1, k2, and k3 represent the sequence number of the convolutional layer, and LReLU[.] is the nonlinear activation function (i.e., leaky rectified linear unit) used throughout the network, defined as:
[0088]
[0089] The number of input channels at each level in the downsampling path is set to: 1, 64, 128, 256, and the number of output channels in the downsampling path is set to: 64, 128, 256, 512. To avoid the size mismatch of each block, the feature map x k Zero-padded to match x k+1 The connection between each downsampling level is a 2×2 average pooling layer with a stride of 2 pixels, which can downsample the feature map by a factor of 4 (2 times in each direction). After the output of the fourth downsampling block, another convolutional layer (CL) keeps the number of feature maps to 512 before connecting it to the upsampling path. The upsampling path consists of four symmetrical upsampling steps (#1, #2, #3, #4), each of which contains a convolutional block. The feature map y k Mapped to feature map y k+1 The convolution block operation is given by:
[0090] y k+1 =LReLU[CONV k6 {LReLU[CONV k5 {LReLU[CONV k4 {CONCAT(x k+1 ,US{y k})}]}]}](0)
[0091] Among them, CONCAT(.) is the concatenation between two feature maps that merge the number of channels, US{.} is the upsampling operator, and k4, k5, and k6 represent the sequence number of the convolutional layers. The number of input channels for each level in the upsampling path is set to 1024, 512, 256, 128, and the number of output channels for each level in the upsampling path is set to 256, 128, 64, 32, respectively. The last layer is a convolutional layer (CL) that maps the 32 channels represented by the YCbCr color map to 3 channels. Both the generator network and the discriminator network have been trained with a patch size of 256×256 pixels.
[0092] exist Fig.10As summarized in Figure 1, the discriminator network receives three (3) input channels which correspond to the YCbCr color space of the input image 40YCbCr, 48YCbCr. This input is then converted to a 64-channel representation using a convolutional layer, followed by 5 blocks of the following operators:
[0093] z k+1 =LReLU[CONV k2 {LReLU[CONV k1 {z k}]}] (0)
[0094] Among them, k1, k2, represent the serial number of the convolutional layer. The number of channels in each layer is 3, 64, 64, 128, 128, 256, 256, 512, 512, 1024, 1024, 2048. The next layer is an average pooling layer with a filter size equal to the patch size (256×256), which will produce a vector with 2048 entries. The output of this average pooling layer is then fed to two fully connected layers (FC) with the following structure:
[0095] z k+1 =FC[LReLU[FC{z k}]] (0)
[0096] Where FC represents a fully connected layer with learnable weights and biases. The first fully connected layer outputs a vector with 2048 entries, while the second outputs a scalar value. This scalar value is used as input to the sigmoid activation function D(z) = 1 / (1+exp(-z)), which calculates the probability (between 0 and 1) that the discriminator network input is real / true or fake, i.e., ideally Fig.10 The output 67 is shown in D(z label )=1.
[0097] The convolution kernels of the entire GAN are set to 3×3. These kernels are randomly initialized using a truncated normal distribution with a standard deviation of 0.05 and a mean of 0; all network biases are initialized to 0. The network is trained using the Adaptive Moment Estimation (Adam) optimizer (70 for the generator network) with a learning rate of 1×10 -4 , for the discriminator network 74 the learning rate is 1×10 -5 ) Back propagation (such as Fig.10 ), the learnable parameters are updated during the training phase of the deep neural network 10. Moreover, for each iteration of the discriminator 74, there are 4 iterations of the generator network 70 to avoid training stagnation after the discriminator network may potentially overfit to the labeling. The batch size used in training is 10.
[0098] Once all fields of view have passed through the network 10, the entire slide image is stitched together using the Fiji Grid / Collection stitching plug-in (e.g., see Schindelin, J. et al., Fiji: A source platform for biological image analysis, Nature Methods, Vol. 9, pp. 676-682, 2012, cited herein as a reference). The plug-in calculates the exact overlap between each tile and then linearly blends them into a large image. Overall, inference and stitching took approximately 5 minutes and 30 seconds per square centimeter, respectively, and can be greatly improved with hardware and software improvements. Before displaying to the pathologist, portions of the autofluorescence or brightfield images that were out of focus or had large aberrations (e.g., due to dust particles) were cropped out. Finally, the images were exported to Zoomify format (designed for viewing large images using a standard web browser; http: / / zoomify.com / ) and uploaded to the GIGAmacro website (https: / / viewer.gigamacro.com / ) for easy access and viewing by the pathologist.
[0099] Implementation details
[0100] Other implementation details, including the number of trained patches, the number of epochs, and the training time, are shown in Table 5 below. The digital / virtual colored deep neural network 10 was implemented using Python version 3.5.0. The GAN was implemented using the TensorFlow framework version 1.4.0. Other python libraries used included os, time, tqdm, Python Imaging Library (PIL), SciPy, glob, ops, sys, and numpy. The software was implemented on a desktop computer with a Core i7-7700K CPU @ 4.2 GHz (Intel) and 64 GB RAM, running the Windows 10 operating system (Microsoft). Dual A GTX 1080Ti GPU (NVIDIA) was used for network training and testing.
[0101] Table 5
[0102]
[0103] Although the embodiments of the present invention have been shown and described, various modifications may be made without departing from the scope of the present invention. Accordingly, the present invention should not be restricted except in accordance with the following claims and their equivalents.
Claims
1. A method for generating a digitally stained microscopic image, comprising: providing a neural network using one or more processors of a computing device, wherein the neural network is trained using a plurality of chemically stained images or image patches matched to corresponding unlabeled images or image patches of a training sample; obtaining one or more unlabeled images of an unlabeled test sample using a fluorescence microscope and one or more excitation light sources, wherein the detected light is emitted from at least one endogenous fluorophore or at least one endogenous emitter of the unlabeled test sample; inputting the one or more unlabeled images of the unlabeled test samples into the neural network; and A digitally stained microscopic image of the unlabeled test sample is output via the neural network.
2. The method according to claim 1, wherein: The neural network includes a plurality of neural networks.
3. The method according to claim 1, wherein: The neural network is trained using a generative adversarial network (GAN) model.
4. The method according to claim 1, wherein: The neural network is trained using a generator network configured to learn statistical transformations between matched unlabeled images or image patches and chemically stained images or image patches of the same training sample and a network configured to discriminate between ground truth chemically stained images of the training sample and output digitally stained microscopy images of the training sample.
5. The method according to claim 1, wherein: The unlabeled test sample includes animal tissue, plant tissue, cell, pathogen or biological fluid smear.
6. The method according to claim 1, wherein: The neural network outputs a digitally stained microscopic image in less than one second of inputting the one or more unlabeled images.
7. The method according to claim 1, wherein: The unlabeled test sample includes an unfixed tissue sample.
8. The method according to claim 1, wherein: The unlabeled test sample includes a fixed tissue sample.
9. The method according to claim 8, wherein: The fixed tissue samples were embedded in paraffin.
10. The method according to claim 1, wherein: The unlabeled test sample comprises a frozen tissue sample.
11. The method according to claim 1, wherein: The unlabeled test sample includes a fresh tissue sample.
12. The method according to claim 1, wherein: The excitation light source emits ultraviolet light or near-ultraviolet light.
13. The method according to claim 1, wherein: The one or more unlabeled images are obtained using one or more filters of a filter set.
14. The method according to claim 13, wherein: A plurality of unlabeled images are captured using the plurality of filters.
15. The method according to claim 14, wherein: Multiple images are obtained by using multiple excitation light sources to emit light at different wavelengths or wavelength bands.
16. The method according to claim 1, wherein: The one or more unlabeled images are subjected to one or more image pre-processing operations before being input into the neural network.
17. The method according to claim 16, wherein: The one or more image pre-processing operations include contrast enhancement, contrast inversion, image filtering, or a combination thereof.
18. The method according to claim 1, wherein: The neural network is trained using one or more GPUs or ASICs.
19. The method according to claim 1, wherein: The neural network is executed using one or more GPUs or ASICs.
20. The method according to claim 1, wherein: Obtaining one or more unlabeled images of the unlabeled test sample includes obtaining at least two fluorescent images of the unlabeled test sample.
21. The method according to claim 20, wherein: The at least two fluorescent images are acquired using different wavelengths.
22. The method according to claim 20, wherein: The at least two fluorescent images are acquired using different resolutions.
23. The method according to claim 1, wherein: The neural network includes a convolutional neural network.
24. The method according to claim 1, wherein: The detected light is emitted from one or both of the at least one intrinsic fluorescent light and the at least one intrinsic emitter of frequency-shifted light of the unlabeled test sample.
Citation Information
Cited By
Method of training a machine learning model in order to create at least one virtual histological stained image
US12682620B1