Method and system for digital staining of microscopic images using deep learning

Deep neural networks are used to virtually stain and autofocus unstained tissue samples, addressing the inefficiencies of traditional methods by producing high-quality images comparable to chemically stained samples, thus enhancing diagnostic speed and resource efficiency.

JP2026041721APending Publication Date: 2026-03-10RGT UNIV OF CALIFORNIA
View PDF 0 Cites 0 Cited by

Patent Information

Authority / Receiving Office
JP · JP
Patent Type
Applications
Current Assignee / Owner
Filing Date
2025-11-05
Publication Date
2026-03-10

AI Technical Summary

Technical Problem

Existing microscopic imaging techniques for tissue samples require laborious and resource-intensive processes, such as formalin-fixation, paraffin-embedding, and chemical staining, which are not readily available in all settings and are time-consuming, and there is a lack of efficient methods for imaging unsectioned or unstained tissues.

Method used

Utilizing deep neural networks trained with matched immunohistochemistry-stained and fluorescence lifetime imaging images to virtually stain unstained tissues, and employing computational autofocus to improve focus and convert images between different microscopy modalities.

Benefits of technology

Enables rapid, resource-efficient generation of high-quality, virtually stained images that resemble chemically stained samples, reducing processing time and preserving tissue integrity, while allowing for real-time diagnosis and improved imaging speed and accuracy.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 2026041721000001_ABST
    Figure 2026041721000001_ABST
Patent Text Reader

Abstract

It allows the creation of digitally / virtually stained microscopic images from unlabeled or unstained samples. [Solution] Fluorescence lifetime imaging (FLIM) images of a sample are used with a fluorescence microscope to generate digitally / virtually stained microscopic images of an unlabeled or unstained sample. In another embodiment, a method for digitally / virtually autofocusing is provided that uses machine learning to generate microscopic images with improved focus using a trained deep neural network. In another embodiment, the trained deep neural network generates digitally / virtually stained microscopic images of an unlabeled or unstained sample obtained with a microscope, the digitally / virtually stained microscopic images having multiple different stains. The multiple stains of the output image or subregions thereof are substantially equivalent to corresponding microscopic images or image subregions of the same sample that have been histochemically stained.
Need to check novelty before this filing date? Find Prior Art

Description

[Technical Field]

[0001] [Related Applications] This application claims priority to U.S. Provisional Patent Application No. 63 / 058,329, filed July 29, 2020, and U.S. Provisional Patent Application No. 62 / 952,964, filed December 23, 2019, which applications are incorporated herein by reference in their entireties. Priority is claimed pursuant to 35 U.S.C. § 119 and any other applicable statutes.

[0002]

[0002] The technical field generally relates to methods and systems used to image unstained (i.e., unlabeled) tissue. In particular, the technical field relates to microscopy methods and systems that utilize deep neural network learning to digitally or virtually stain images of unstained or unlabeled tissue. Deep learning with neural networks, a class of machine learning algorithms, is used to digitally stain images of unlabeled tissue sections into images equivalent to microscopic images of the same sample that have been stained or labeled. [Background technology]

[0003] Microscopic imaging of tissue samples is a fundamental tool used in the diagnosis of various diseases and forms a mainstay of pathology and biological science. Clinically established gold-standard images of tissue sections are the result of a laborious process that involves formalin-fixed, paraffin-embedded (FFPE) tissue specimens, sectioning them into thin slices (typically approximately 2-10 μm), labeling / staining, and mounting them on glass slides, followed by microscopic imaging using, for example, bright-field microscopy. All of these steps involve the use of multiple reagents and irreversible effects on the tissue. Recent efforts have attempted to modify this workflow using various imaging modalities. For example, attempts have been made to image fresh, non-paraffin-embedded tissue samples using nonlinear microscopy based on two-photon fluorescence, second-harmonic generation, third-harmonic generation, and Raman scattering. Other attempts have used controllable supercontinuum sources to acquire multimodal images for chemical analysis of fresh tissue samples. These methods require the use of ultrafast lasers or supercontinuum sources, which may not be readily available in most settings, and require relatively long scan times due to the weak optical signal.In addition to these, other microscopy methods have emerged for imaging unsectioned tissue samples by using UV excitation on stained samples or by exploiting the fluorescence emission of living tissue at shorter wavelengths.

[0004] Indeed, fluorescent signals create several unique opportunities for imaging tissue samples by utilizing fluorescence emitted from endogenous fluorophores. Such intrinsic fluorescence signatures have demonstrated useful information that can be mapped to functional and structural properties of biological specimens and have therefore been widely used for diagnostic and research purposes. One of the main focus areas of these efforts has been the spectroscopic investigation of the relationship between various biomolecules and their structural properties under various conditions. Some of these well-characterized biological components include vitamins (e.g., vitamin A, riboflavin, thiamine), collagen, coenzymes, and fatty acids, among others.

[0005] While some of the techniques described above have the unique ability to identify, for example, cell types and subcellular components of tissue samples using various contrast mechanisms, pathologists and tumor classification software are generally trained to examine "gold standard" stained tissue samples to make diagnostic decisions. Partly motivated by this, some of the techniques described above have been extended to create pseudo-hematoxylin-eosin (H&E) images, which are based on a linear approximation that relates the fluorescence intensity of the image to the dye concentration per tissue volume using empirically determined constants that represent the average spectral response of various dyes embedded in the tissue. These methods used exogenous stains to enhance the contrast of the fluorescent signal to create a virtual H&E image of the tissue sample. Summary of the Invention

[0006] In one embodiment, a method for generating a virtually stained microscopic image of a sample includes using one or more processors of a computing device to provide a trained deep neural network executed by image processing software, where the trained deep neural network is trained using a plurality of matched immunohistochemistry (IHC)-stained microscopic images or image patches and their corresponding fluorescence lifetime imaging (FLIM) microscopic images or image patches of the same sample obtained before immunohistochemistry (IHC) staining. The fluorescence lifetime (FLIM) images of the sample are obtained using a fluorescence microscope and at least one excitation light source, and the fluorescence lifetime (FLIM) images of the sample are input to the trained deep neural network. The trained deep neural network outputs a virtually stained microscopic image of the sample that is substantially equivalent to the corresponding image of the same immunohistochemistry (IHC)-stained sample.

[0007]

[0007] In another embodiment, a method for virtually autofocusing a microscopic image of a sample obtained using an incoherent microscope includes providing a trained deep neural network executed by image processing software using one or more processors of a computing device, the trained deep neural network being trained using multiple pairs of out-of-focus and / or in-focus microscopic images or image patches, which are used as input images to the deep neural network, and corresponding or matching in-focus microscopic images or image patches of the same sample obtained using an incoherent microscope, which are used as ground truth images for training the deep neural network. An out-of-focus or in-focus image of the sample is obtained using the incoherent microscope. The out-of-focus or in-focus image of the sample obtained from the incoherent microscope is then input to the trained deep neural network. The trained deep neural network outputs an output image with improved focus, which substantially matches the in-focus image (ground truth) of the same sample acquired by the incoherent microscope.

[0008] In another embodiment, a method for generating a virtually stained microscopic image of a sample using an incoherent microscope includes providing a trained deep neural network executed by image processing software using one or more processors of a computing device, the trained deep neural network being trained with multiple pairs of out-of-focus and / or in-focus microscopic images or image patches that all match corresponding in-focus microscopic images or image patches of the same sample obtained using the incoherent microscope after a chemical staining process, which are used as input images to the deep neural network and generate ground truth images for training the deep neural network. Out-of-focus and in-focus images of the sample are acquired using the incoherent microscope, and the out-of-focus and in-focus images of the sample obtained from the incoherent microscope are input to the trained deep neural network. The trained deep neural network outputs an output image of the virtually stained sample with improved focus and that is substantially similar to and matches the in-focus chemically stained image of the same sample obtained by the incoherent microscope after the chemical staining process.

[0009] In another embodiment, a method for generating a virtually stained microscopic image of a sample includes using one or more processors of a computing device to provide a trained deep neural network executed by image processing software, the trained deep neural network being trained with a plurality of pairs of stained microscopic images or image patches, each pair being virtually stained by at least one algorithm or chemically stained to have a first stain type, and all matched with corresponding stained microscopic images or image patches of the same sample that are virtually stained by at least one algorithm or chemically stained to have another, different stain type, the corresponding stained microscopic images or image patches of the same sample constituting ground truth images for training the deep neural network to convert input images histochemically or virtually stained with the first stain type to output images virtually stained with a second stain type. A histochemical or virtually stained input image of the sample stained with the first stain type is obtained. The histochemical or virtually stained input image of the sample is input to the trained deep neural network, which converts the input image stained with the first stain type to an output image virtually stained with the second stain type. The trained deep neural network outputs an output image of the sample having a virtual stain that substantially resembles and matches a chemically stained image of the same sample stained with a second stain type obtained by incoherent microscopy after the chemical staining process.

[0010]

[0010] In another embodiment, a method for generating virtually stained microscopic images of a sample with multiple different stains using a single trained deep neural network includes providing a trained deep neural network executed by image processing software using one or more processors of a computing device, wherein the trained deep neural network is trained using multiple matched chemically stained microscopic images or image patches using multiple chemical stains, which are used as ground truth images for training the deep neural network, and their corresponding matched fluorescent microscopic images or image patches of the same sample obtained before chemical staining, which are used as input images for training the deep neural network. Fluorescent images of the sample are obtained using a fluorescent microscope and at least one excitation light source. One or more class-conditional matrices are applied to condition the trained deep neural network. The fluorescent images of the sample, along with the one or more class-conditional matrices, are input to the trained deep neural network. The trained conditional deep neural network outputs a virtually stained microscopic image of the sample having one or more different stains, wherein the output image or subregion thereof is substantially equivalent to a corresponding microscopic image or image subregion of the same sample that is histochemically stained with the corresponding one or more different stains.

[0011]

[0011] In another embodiment, a method for generating virtually stained microscopic images of a sample with multiple different stains using a single trained deep neural network includes providing a trained deep neural network executed by image processing software using one or more processors of a computing device, wherein the trained deep neural network is trained with multiple matching chemically stained microscopic images or image patches using multiple chemical stains and their corresponding microscopic images or image patches of the same sample obtained before the chemical staining. An input image of the sample is obtained using a microscope. One or more class-conditional matrices are applied to condition the trained deep neural network. The input image of the sample along with the one or more class-conditional matrices are input to the trained deep neural network. The trained conditioned deep neural network outputs a virtually stained microscopic image of the sample with one or more different stains, and the output image or a subregion thereof is substantially equivalent to a corresponding microscopic image or image subregion of the same sample histochemically stained with the corresponding one or more different stains.

[0012] In another embodiment, a method for generating a virtually destained microscopic image of a sample includes using one or more processors of a computing device to provide a trained first deep neural network executed by image processing software, the trained first neural network being trained with a plurality of matching chemically stained microscopic images or image patches used as training inputs to the deep neural network and their corresponding unstained microscopic images or image patches of the same sample or plurality of samples obtained before chemical staining, which constitute ground truth during training of the deep neural network. The microscopic images of the chemically stained samples are obtained using a microscope. The images of the chemically stained samples are input to the trained first deep neural network. The trained first deep neural network outputs a virtually destained microscopic image of the sample that is substantially equivalent to a corresponding image of the same sample obtained before or without chemical staining. [Brief explanation of the drawings]

[0013] [Figure 1] FIG. 1 shows a schematic diagram of a system used to generate a digitally / virtually stained output image of a sample from an unstained microscope image of the sample, according to one embodiment. [Figure 2]

[0014] Figure 2 shows a schematic diagram of the deep learning-based digital / virtual histology staining process using fluorescent images of unstained tissue. [Figure 3]

[0015] Figures 3A-3H show the results of digital / virtual staining matched to chemically stained H&E samples. The first two columns (Figure 3A and Figure 3E) show autofluorescence images of unstained salivary gland tissue sections (used as input to the deep neural network), while the third column (Figure 3C and Figure 3G) shows the results of digital / virtual staining. The final columns (Figure 3D and Figure 3H) show brightfield images of the same tissue sections after the histochemical staining process. Evaluation in both Figures 3C and 3D reveals islands of invasive tumor cells within the subcutaneous fibroadipose tissue. Note that nuclear details, including nucleoli (arrows in Figure 3C and Figure 3D) and distinct chromatin textures, are clearly visible in both panels. Similarly, in Figures 3G and 3H, H&E staining reveals invasive squamous cell carcinoma. An interstitial fibrotic reaction with edematous myxoid changes in the adjacent stroma (asterisks in Figures 3G and 3H) is clearly discernible in both stains / panels. [Figure 4]

[0016] Figures 4A-H show the results of digital / virtual staining matched to the chemically stained Jones samples. The first two columns (Figure 4A, Figure 4E) show autofluorescence images of unstained kidney tissue sections (used as input to the deep neural network), while the third column (Figure 4C and Figure 4G) shows the results of digital / virtual staining. The final columns (Figure 4D, Figure 4H) show brightfield images of the same tissue sections after the histochemical staining process. [Figure 5]

[0017] Figures 5A-5P show digital / virtual staining results consistent with Masson's trichrome staining of liver and lung tissue sections. The first two columns show autofluorescence images of unstained liver tissue sections (rows 1 and 2—Figures 5A, 5B, 5E, and 5F) and unstained lung tissue sections (rows 3 and 4—Figures 5I, 5J, 5M, and 5N), which are used as inputs to the deep neural network. The third column (Figures 5C, 5G, 5K, and 5O) shows the digital / virtual staining results of these tissue samples. The final columns (Figures 5D, 5H, 5L, and 5P) show brightfield images of the same tissue sections after the histochemical staining process. [Figure 6]

[0018] Figure 6A shows a graph of the loss function versus iteration number combination for random initialization and transfer learning initialization. Figure 6A illustrates how superior convergence is achieved using transfer learning. A new deep neural network is initialized using weights and biases learned from salivary gland tissue sections and virtual staining of thyroid tissue using H&E. Compared to random initialization, transfer learning allows for much faster convergence and achieves fewer local minima.

[0019] Figure 6B shows network output images at various stages of the learning process for both random initialization and transfer learning, better illustrating the impact of transfer learning on transforming the presented approach to new tissue / stain combinations.

[0020] Figure 6C shows the corresponding H&E chemically stained bright-field image. [Figure 7]

[0021] FIG. 7A shows virtual staining (H&E staining) of skin tissue using only the DAPI channel.

[0022] Figure 7B shows virtual staining (H&E staining) of skin tissue using DAPI and Cy5 channels, where Cy5 refers to a far-red fluorescent cyanine dye used to label biomolecules.

[0023] FIG. 7C shows the corresponding histologically stained (ie, chemically stained with H&E) tissue. [Figure 8]

[0024] FIG. 8 shows the field matching and registration process of an autofluorescence image of an unstained tissue sample to a bright field image of the same sample after the chemical staining process. [Figure 9]

[0025] Figure 9 shows a schematic diagram of the training process of the virtual staining network using GAN. [Figure 10]

[0026] FIG. 10 illustrates a generative adversarial network (GAN) architecture for the generator and discriminator, according to one embodiment. [Figure 11]

[0027] Figure 11A shows machine learning-based virtual IHC staining using autofluorescence and fluorescence lifetime images of unstained tissue. Following training with a deep neural network model, the trained deep neural network rapidly outputs a virtually stained tissue image in response to an autofluorescence lifetime image of an unstained tissue section, bypassing the standard IHC staining procedures used in histology.

[0028] FIG. 11B illustrates a generative adversarial network (GAN) architecture for a generator (G) and a discriminator (D) used in fluorescence lifetime imaging (FLIM) according to one embodiment. [Figure 12]

[0029] Figure 12 shows the results of digital / virtual staining consistent with the HER2 staining diagram. The two leftmost columns show the autofluorescence intensity (first column) and lifetime images (second column) of unstained human breast tissue sections (used as input to the deep neural network 10) for two excitation wavelengths, while the center column shows the results of virtual staining (Virtual Staining). The final column shows a brightfield image of the same tissue section after the IHC staining process (IHC Staining). [Figure 13]

[0030] Figure 13 shows a comparison between standard (prior art) autofocusing methods and the virtual focusing method disclosed herein. Top right: Standard autofocusing methods require acquiring multiple images with the autofocus algorithm selecting the most focused image according to predefined criteria. In contrast, the disclosed method requires only a single aberrated image to virtually refocus on it using a trained deep neural network. [Figure 14]

[0031] Figure 14 shows a demonstration of the refocusing capability of the post-imaging computational autofocusing method for various imaging planes of a sample, where the value of z denotes the focal distance from the focusing plane, which resides at z = 0 (as the reference plane). The refocusing capability of the network can be evaluated both qualitatively and quantitatively using the structural similarity index displayed in the top left corner of all panels and comparing its image with the reference focused image (z = 0). [Figure 15]

[0032] Figure 15 shows a schematic of a machine learning-based approach that uses class conditioning to apply multiple stains to microscopic images to create virtually stained images. [Figure 16]

[0033] 16 shows a display illustrating a graphical user interface (GUI) used to display an output image from a trained deep neural network, according to one embodiment. Highlighted are various areas or subregions of the image that can be virtually stained with various stains. This can be done manually, automatically using image processing software, or some combination of the two (e.g., a hybrid approach). [Figure 17]

[0034] Figure 17 shows an example of stain microstructuring. A diagnostician can manually label sections of unstained tissue. These labels are used by the network to stain various areas of the tissue with the desired stain. For comparison, a co-registered image of histochemically stained H&E tissue is shown. [Figure 18]

[0035] Figures 18A-18G show examples of stain blending. Figure 18A shows an autofluorescence image used as input to the machine learning algorithm. Figure 18B shows a co-registered image of histochemically stained H&E tissue for comparison. Figure 18C shows kidney tissue virtually stained with H&E (no Jones). Figure 18D shows kidney tissue virtually stained with Jones stain (no H&E). Figure 18E shows kidney tissue virtually stained with a 3:1 input class conditioning ratio of H&E:Jones stain. Figure 18F shows kidney tissue virtually stained with a 1:1 input class conditioning ratio of H&E:Jones stain. Figure 18G shows kidney tissue virtually stained with a 1:3 input class conditioning ratio of H&E:Jones stain. [Figure 19]

[0036] 19A and 19B show another embodiment for virtually destaining an image obtained from a microscope (e.g., from a stained image to a destained image). Also shown is an optional machine learning-based process for restaining the image with a different chemical stain than that obtained from the microscope. This illustrates virtual destaining and restaining. [Figure 20]

[0037] FIG. 20A shows a virtual staining network capable of generating both H&E and special stain images.

[0038] FIG. 20B shows one embodiment of a style transfer network (e.g., a CycleGan network) used to transfer style to virtual and / or histochemically stained images.

[0039] Figure 20C shows the scheme used to train the stain transformation network (10stainTN). The stain transformation is randomly fed either a virtually stained H&E tissue or an image of the same field of view after passing through one of the K=8 stain transformation networks (10styleTN). A perfectly matched virtually stained tissue using the desired special stain (PAS in this case) is used as ground truth for training this neural network. [Figure 21]

[0040] Figure 21 shows a schematic of the transformations performed by various networks during the training phase of style transfer (e.g., CycleGAN). [Figure 22]

[0041] FIG. 22A shows the structure of the first generator network G(x) used for style transfer.

[0042] FIG. 22B shows the structure of the second generator network F(y) used for style transfer. DETAILED DESCRIPTION OF THE INVENTION

[0014]

[0043] FIG. 1 schematically illustrates one embodiment of a system 2 for outputting a digitally stained image 40 from an input microscopic image 20 of a sample 22. As described herein, the input image 20 is a fluorescence image 20 of a sample 22 (e.g., a tissue, in one embodiment) that has not been stained or labeled with a fluorescent stain or label. That is, the input image 20 is an autofluorescence image 20 of the sample 22, where the fluorescence emitted by the sample 22 is the result of one or more endogenous fluorophores or other endogenous emitters of frequency-shifted light contained therein. Frequency-shifted light is light emitted at a frequency (or wavelength) different from the incident frequency (or wavelength). Endogenous fluorophores or endogenous emitters of frequency-shifted light may include molecules, compounds, complexes, molecular species, biomolecules, dyes, tissues, etc. In certain embodiments, the input image 20 (e.g., a raw fluorescence image) is subjected to one or more linear or nonlinear pre-processing steps selected from contrast enhancement, contrast inversion, and image filtering. The system includes a computing device 100 including one or more processors 102 and image processing software 104 incorporating a trained deep neural network 10 (e.g., a convolutional neural network as described herein in one or more embodiments). The computing device 100 may include a personal computer, laptop, mobile computing device, remote server, etc., as described herein, although other computing devices (e.g., devices incorporating one or more graphics processing units (GPUs) or other application-specific integrated circuits (ASICs)) may be used. The GPU or ASIC may be used to speed up training and final image output. The computing device 100 may be associated with or connected to a monitor or display 106 used to display the digitally stained image 40. The display 106 may be used to display a graphical user interface (GUI) used by a user to display and view the digitally stained image 40.In one embodiment, a user may be able to manually trigger or switch between multiple different digital / virtual stains of a particular sample 22, for example, using a GUI. Alternatively, triggering or switching between different stains may be done automatically by the computing device 100. In a preferred embodiment, the trained deep neural network 10 is a convolutional neural network (CNN).

[0015]

[0044] For example, in a preferred embodiment described herein, the trained deep neural network 10 is trained using a GAN model. In a GAN-trained deep neural network 10, two models are used for training. A generative model is used to capture the data distribution, while a second model estimates the probability that a sample was obtained from the training data rather than from the generative model. More information about GANs is provided in Goodfellow et al., "Generative Adversarial Nets," Advances in Neural Information Processing Systems, 27, pp. 2672-2680 (2014), which is incorporated herein by reference. Network training of the deep neural network 10 (e.g., a GAN) may be performed on the same or a different computing device 100. For example, in one embodiment, a personal computer may be used to train a GAN, although such training may take a significant amount of time. To expedite the training process, one or more dedicated GPUs may be used for training. As described herein, such training and testing was performed on a GPU obtained from a commercially available graphics card. Once the deep neural network 10 is trained, it may be used or executed on a different computing device 110, which may include one that uses fewer computational resources for the training process (although a GPU may also be integrated into the execution of the trained deep neural network 10).

[0016]

[0045] The image processing software 104 can be implemented using Python and TensorFlow, although other software packages and platforms can be used. The trained deep neural network 10 is not limited to a particular software platform or programming language; the trained deep neural network 10 can be executed using any number of commercially available software languages ​​or platforms. The image processing software 104, which incorporates or runs in concert with the trained deep neural network 10, can be executed in a local environment or a remove cloud-type environment. In some embodiments, some functions of the image processing software 104 can be executed in one particular language or platform (e.g., image normalization), while the trained deep neural network 10 can be executed in another particular language or platform. Nevertheless, both processes are performed by the image processing software 104.

[0017]

[0046] As seen in FIG. 1 , in one embodiment, the trained deep neural network 10 receives a single fluorescence image 20 of an unlabeled sample 22. In other embodiments, for example, if multiple excitation channels are used (see the discussion of melanin herein), there may be multiple fluorescence images 20 (e.g., one image per channel) of the unlabeled sample 22 input to the trained deep neural network 10. The fluorescence image 20 may include a wide-field fluorescence image 20 of the unlabeled tissue sample 22. Wide-field is meant to indicate that a wide field of view (FOV) is obtained by scanning a smaller FOV, with a wide field of view typically ranging from 10 to 2000 mm. 2For example, smaller FOVs can be obtained by a scanning fluorescence microscope 110, which digitally stitches smaller FOVs together to create a wider FOV using image processing software 104. For example, a wider FOV can be used to obtain a whole slide image (WSI) of the sample 22. Fluorescence images can be obtained using an imaging device 110, which, in the case of the fluorescence embodiments described herein, can include a fluorescence microscope 110. The fluorescence microscope 110 includes at least one excitation light source that illuminates the sample 22 and one or more image sensors (e.g., CMOS image sensors) for capturing fluorescence emitted by frequency-shifted fluorophores or other endogenous emitters contained in the sample 22. The fluorescence microscope 110, in some embodiments, can include the capability to illuminate the sample 22 with excitation light of multiple different wavelengths or wavelength ranges / bands. This can be achieved using multiple different light sources and / or different filter sets (e.g., standard UV or near-UV excitation / emission filter sets). Additionally, the fluorescence microscope 110, in some embodiments, can include multiple filter sets capable of filtering different emission bands. For example, in some embodiments, multiple fluorescent images 20 may be captured, each captured in a different emission band using a different filter set.

[0018]

[0047] In some embodiments, the sample 22 may include a portion of tissue disposed on or within the substrate 23. In some embodiments, the substrate 23 may include an optically transparent substrate (e.g., a glass or plastic slide, etc.). The sample 22 may include a tissue section cut into thin sections using a microtome or the like. The tissue thin section 22 can be considered a weakly scattering phase object with limited amplitude contrast modulation under bright-field illumination. The sample 22 may be imaged with or without a coverslip / coverslip. The sample may include a frozen section or a paraffin (wax) section. The tissue sample 22 may be fixed (e.g., using formalin) or unfixed. The tissue sample 22 may include mammalian (e.g., human or animal) tissue or plant tissue. The sample 22 may also include other biological samples, environmental samples, etc. Examples include particles, cells, organelles, pathogens, parasites, fungi, or other microscale objects of interest (objects of micrometer size or smaller). The sample 22 may include a smear of bodily fluid or tissue. These include, for example, a blood smear, a Papanicolaou smear, or a Pap smear. As described herein, in fluorescence-based embodiments, the sample 22 contains one or more naturally occurring or endogenous fluorescent light that fluoresces and is captured by the fluorescence microscope device 110. Most plant and animal tissues exhibit some degree of autofluorescence when excited with ultraviolet or near-ultraviolet light. Endogenous fluorophores may include, by way of example, proteins such as collagen, elastin, butyric acid, vitamins, flavins, porphyrins, lipofuscin, and coenzymes (e.g., NAD(P)H). In an optional embodiment, exogenously added fluorescent labels or other exogenous illuminants may be added (for training the deep neural network 10, for testing new samples 12, or both). As described herein, the sample 22 may also contain other endogenous emitters of frequency-shifted light.

[0019]

[0048] In response to the input image 20, the trained deep neural network 10 outputs or generates a digitally stained or labeled output image 40. The digitally stained output image 40 has "staining" digitally integrated into it using the trained deep neural network 10. In certain embodiments, such as involving tissue sections, the trained deep neural network 10 produces images that appear to a skilled observer (e.g., a skilled histopathologist) to be substantially equivalent to a corresponding brightfield image of the same chemically stained tissue section sample 22. Indeed, as described herein, experimental results obtained using the trained deep neural network 10 indicate that a skilled pathologist was able to recognize histopathological features using both staining techniques (chemical staining vs. digital / virtual staining) and with a high degree of agreement between techniques (virtual vs. histological) without a clear preferred staining technique. This digital or virtual staining of the tissue section sample 22 appears the same as if the tissue section sample 22 had undergone histochemical staining, even if such a staining step had not been performed.

[0020]

[0049] FIG. 2 illustrates a schematic diagram of the steps involved in a typical fluorescence-based embodiment. As seen in FIG. 2, a sample 22, such as an unstained tissue section, is obtained. This may be obtained, for example, from biological tissue via biopsy B. The unstained tissue section sample 22 then undergoes fluorescent imaging using a fluorescent microscope 110 to generate a fluorescent image 20. This fluorescent image 20 is then input into a trained deep neural network 10, which then immediately outputs a digitally stained image 40 of the tissue section sample 22. This digitally stained image 40 closely resembles the appearance of a brightfield image of the same tissue section sample 22 after the actual tissue section sample 22 has undergone histochemical staining. FIG. 2 illustrates (using dashed arrows) a conventional process in which the tissue section sample 22 undergoes histochemical or immunohistochemical (IHC) staining 44 and subsequently undergoes conventional brightfield microscopy imaging 46 to generate a conventional brightfield image 48 of the stained tissue section sample 22. As seen in FIG. 2, the digitally stained image 40 closely resembles the actual chemically stained image 48. Similar resolutions and color profiles are obtainable using the digital staining platform described herein. This digitally stained image 40 may be shown or displayed on a computer monitor 106, as shown in FIG. 1, although it should be understood that the digitally stained image 40 may be displayed on any suitable display (e.g., a computer monitor, a tablet computer, a mobile computing device, a mobile phone, etc.). A GUI may be displayed on the computer monitor 106 to allow a user to view and optionally interact with (e.g., zoom, cut, highlight, mark, adjust exposure, etc.) the digitally stained image 40.

[0021]

[0050] In one embodiment, the fluorescence microscope 110 acquires a fluorescence lifetime image of an unstained tissue sample 22 and outputs an image 40 that closely matches a bright-field image 48 of the same field after IHC staining. Fluorescence lifetime imaging (FLIM) generates images based on differences in the decay rates of excited states from fluorescent samples. FLIM is thus a fluorescence imaging technique in which contrast is based on the lifetime or decay of individual fluorophores. Fluorescence lifetime is generally defined as the average time a molecule or fluorophore remains in an excited state before returning to its ground state by emitting a photon. Of all the intrinsic properties of unlabeled tissue samples, the fluorescence lifetime of endogenous fluorophores is one of the most informative channels for measuring the time a fluorophore remains in an excited state before returning to its ground state.

[0022]

[0051] It is well known that the lifetimes of endogenous fluorescent emitters, such as flavin adenine dinucleotide (FAD) and nicotinamide adenine dinucleotide (NAD or NADH), depend on the immersing chemical environment, such as oxygen availability, and therefore indicate physiological and biological changes within tissues that are not evident by bright-field or fluorescence microscopy. While existing literature has confirmed a close correlation between lifetime changes between benign and cancerous tissues, there is a lack of a cross-diagnostic image transformation method that allows pathologists or computer software to perform disease diagnosis on unlabeled tissues based on trained color contrast. In this embodiment of the present invention, a machine learning algorithm (i.e., a trained deep neural network 10) enables virtual IHC staining of unstained tissue samples 22 based on fluorescence lifetime imaging. Using this method, laborious and time-consuming IHC staining procedures can be replaced with virtual staining, thereby significantly speeding up the process and enabling tissue preservation for further analysis.

[0023]

[0052] In one embodiment, the trained neural network 10 is trained using a lifetime (e.g., decay time) fluorescence image 20 of an unstained sample 22, and the paired ground truth image 48 is a brightfield image of the same field after IHC staining. The trained neural network 10 may also be trained, in another embodiment, using a combination of a lifetime fluorescence image 20 and a fluorescence intensity image 20. Once the neural network 10 has converged (i.e., trained), it can be used to blindly infer a new lifetime image 20 from the unstained tissue sample 22 and convert or output the new lifetime image 20 to the equivalent of a post-stain brightfield image 40 without parameter adjustment, as shown in Figures 11A and 11B.

[0024]

[0053] To train the artificial neural network 10, a generative adversarial network (GAN) framework was used to perform virtual staining. The training dataset consisted of autofluorescence (endogenous fluorophore) lifetime images 20 of multiple tissue sections 22 for single or multiple excitation and emission wavelengths. The samples 22 were scanned using a standard fluorescence microscope 110 with photon counting capabilities, which output a fluorescence intensity image 20I and a lifetime image 20L for each field of view. Tissue samples 22 were also sent to a pathology lab for IHC staining and scanned using a brightfield microscope, which was used to generate ground truth training images 48. The fluorescence lifetime image 20L and the brightfield image 48 of the same field of view were paired. The training dataset consisted of thousands of such pairs 20L, 48, which served as input and output, respectively, for training the network 10. Typically, the artificial neural network model 10 converged after approximately 30 hours on two Nvidia 1080Ti GPUs. Once the neural network 10 has converged, the method allows for virtual IHC staining of unlabeled tissue sections 22 in real-time performance, as shown in Figure 12. Note that the deep neural network 10 can be trained using fluorescence lifetime images 20L and fluorescence intensity images 20I, as shown in Figure 11B.

[0025]

[0054] In another embodiment, a trained deep neural network 10a is provided that takes an aberrant and / or out-of-focus input image 20 and then outputs a corrected image 20a that substantially matches a focused image of the same field of view. For example, a key step for high-quality, rapid microscopic imaging of a tissue sample 22 is autofocus. Traditionally, autofocus is performed using a combination of optical and algorithmic methods. These methods are time-consuming because they image the specimen 22 at multiple focal depths. The ever-increasing demand for higher-throughput microscopy requires more assumptions to be made about the specimen's profile. In other words, the accuracy typically achieved by acquiring multiple focal depths is sacrificed under the assumption that the specimen's profile is uniform across adjacent fields of view. This type of assumption often results in image focusing errors. These errors may require reimaging the specimen, which is not always possible, for example, in life science experiments. In digital pathology, for example, such focusing errors can prolong the diagnosis of a patient's disease.

[0026]

[0055] In this particular embodiment, post-imaging computational autofocus is performed using a deep neural network 10a trained for incoherent imaging modalities. Thus, post-imaging computational autofocus can be used in conjunction with images obtained by fluorescence microscopy (e.g., fluorescent microscopy) as well as other imaging modalities. Examples include fluorescence microscopy, widefield microscopy, super-resolution microscopy, confocal microscopy with single-photon or multiphoton excitation fluorescence, second-harmonic or high-harmonic generation fluorescence microscopy, light-sheet microscopy, FLIM microscopy, bright-field microscopy, dark-field microscopy, structured illumination microscopy, total internal reflection microscopy, computational microscopy, ptychographic microscopy, synthetic aperture-based microscopy, or phase-contrast microscopy. In one embodiment, the output of the trained deep neural network 10a generates a modified input image 20a that is in focus or that is more focused than the raw input image 20. This modified input image 20a, with improved focus, is then input to a separate trained deep neural network 10 described herein that converts from a first imaging modality to a second imaging modality (e.g., from fluorescence microscopy to brightfield microscopy). In this regard, the trained deep neural networks 10a, 10b are coupled together in a "daisy chain" configuration, with the output of the trained autofocus neural network 10a being input to the trained deep neural network 10 for digital / virtual staining. In another embodiment, the machine learning algorithm used in the trained deep neural network 10 combines the autofocus function with the function described herein of converting images from one microscopy modality to another. In this latter embodiment, two separate trained deep neural networks are not required. Instead, a single trained deep neural network 10a is provided that performs virtual autofocus and digital / virtual staining. The functions of both networks 10a, 10a, are combined into a single network 10a. This deep neural network 10a follows the architecture described herein.

[0027]

[0056] Whether the method is implemented with a single trained deep neural network 10a or multiple trained deep neural networks 10, 10a, it can generate a virtually stained image 40 even from an out-of-focus input image 20 and increase the scanning speed of the imaging sample 22. The deep neural network 10a is trained using images acquired at different focal depths, while the output (either the input image 20 or the virtually stained image 40, depending on the implementation) is an in-focus image of the same field of view. The images used for training are acquired using a standard optical microscope. In training the deep neural network 10a, a "gold standard" or "ground truth" image is paired with various out-of-focus or abnormal images. The gold standard / ground truth images used for training can include, for example, in-focus images of the sample 22 that can be identified by any number of focusing criteria (e.g., sharp edges or other features). "Gold standard" images may also include extended depth-of-field images (EDOF), which are synthetically focused images based on multiple images that provide an in-focus view over a wider depth of field. For training of deep neural network 10a, some of the training images may themselves be in-focus images. A combination of out-of-focus and in-focus images may be used to train deep neural network 10a.

[0028]

[0057] After the training phase is completed, in contrast to standard autofocus techniques that require acquiring multiple images through multiple depth planes, a deep neural network 10a can be used to refocus the abnormal image from a single defocused image, as shown in FIG. 13. As seen in FIG. 13, a single defocused image 20d obtained from the microscope 110 is input into the trained deep neural network 10a, generating a focused image 20f. This focused image 20f can then be input into the trained deep neural network 10 in one embodiment described herein. Alternatively, the functionality of the deep neural network 10 (virtual staining) can be combined with the autofocus functionality into a single deep neural network 10a.

[0029]

[0058] To train the deep neural network 10a, a generative adversarial network (GAN) may be used to perform virtual focusing. The training dataset consists of autofluorescence (endogenous fluorophore) images of multiple tissue sections for multiple excitation and emission wavelengths. In alternative embodiments, the training images may be from other microscopy modalities (e.g., brightfield, super-resolution, confocal, light-sheet, FLIM, widefield, dark-field, structured illumination, computational, ptychographic, synthetic aperture-based or total internal reflection, and phase-contrast microscopy).

[0030]

[0059] The sample is scanned with an Olympus microscope, and a 21-layer image stack with 0.5 μm axial spacing is acquired for each field of view (although in other embodiments, different numbers of images can be acquired at different axial spacings). Defocused and in-focus images of the same field of view are paired. A training dataset consists of thousands of such pairs, which are used as input and output, respectively, for network training. On an Nvidia 2080Ti GPU, training 30,000 image pairs takes approximately 30 hours. Following training of the deep neural network 10a, the method allows for refocusing of an image 20d of the specimen 22 for multiple defocus distances into an in-focus image 20f, as shown in FIG. 14.

[0031]

[0060] The autofocus method is also applicable to thick specimens or samples 22, where the network 10a can be trained to refocus on features at specific depths of the specimen (e.g., the surface of a thick tissue section) and eliminate out-of-focus scattering, which significantly degrades image quality. Various user-defined depths or planes can be defined by the user, including the top surface of the sample 22, the middle surface of the sample 22, or the bottom surface of the sample 22. The output of this trained deep neural network 10a can then be used as input to a second, independently trained virtual staining neural network 10, as described herein, to virtually stain the label-free tissue sample 22. The output image 20f of the trained first deep neural network 10a is then input to the virtual staining trained neural network 10. In an alternative embodiment, a process similar to that outlined above can be used to train a single neural network 10 that can directly take an out-of-focus image 20d from an incoherent microscope 110, such as a fluorescent, bright-field, dark-field, or phase microscope, and directly output a virtually stained image 40 of the label-free sample 22, where the raw image 20d was out of focus (at the input of the same neural network). The virtually stained image 40 resembles an imaging modality other than the incoherent microscope 110 that obtained the out-of-focus image 20d. For example, the out-of-focus image 20d may be obtained using a fluorescent microscope, but the in-focus and digitally stained output image 40 substantially resembles a bright-field microscope image.

[0032]

[0061] In another embodiment, a machine learning-based framework is utilized in which a trained deep neural network 10 enables digital / virtual staining of a sample 22 with multiple stains. Multiple histological virtual stains can be applied to an image using a single trained deep neural network 10. Furthermore, this method allows for user-defined, region-of-interest specific virtual staining, as well as blending of multiple virtual stains (e.g., to generate other unique stains or stain combinations). For example, a graphical user interface (GUI) may be provided to allow a user to paint or highlight specific regions of an unlabeled histological tissue image with one or more virtual stains. This method uses a class-conditional convolutional neural network 10 to transform an input image composed of one or more input images 20, including, in one particular embodiment, an autofluorescence image 20 of an unlabeled tissue sample 22.

[0033]

[0062] As an example, to demonstrate its utility, a single trained deep neural network 10 was used to analyze images of unlabeled sections of tissue samples 22 stained with hematoxylin and eosin (H&E), hematoxylin, eosin, Jones silver, Masson's trichrome, Periodic Acid Schiff (PAS), Congo red, Alcian blue, iron blue, silver nitrate, trichrome, Ziehl-Nielsen, Grocott methenamine silver (GMS), and Gram stain. Colors were stained virtually using acid stains, basic stains, silver stains, Nissl, Weigert stains, Golgi stains, Luxol Fast Blue stains, Toluidine Blue, Genta, Mallory Trichrome stains, Gomori Trichrome, Van Gieson, Giemsa, Sudan Black, Perls Prussian, Best Carmine, Acridine Orange, immunofluorescence stains, immunohistochemical stains, Kinyon-Cold stains, Albert stains, flagella stains, endospore stains, Nigrosine or India Ink.

[0034]

[0063] The method can also be used to generate novel stains that are compositions of multiple virtual stains, as well as staining of specific tissue microstructures with these trained stains. In yet another alternative embodiment, image processing software can be used to automatically identify or segment regions of interest within an image of unlabeled tissue sample 22. These identified or segmented regions of interest can be presented to a user for virtual staining, or can be already stained by the image processing software. As an example, nuclei can be automatically segmented and "digitally" stained with specific virtual stains without having to be identified by a pathologist or other human operator.

[0035]

[0064] In this embodiment, one or more autofluorescence images 20 of unlabeled tissue 22 are used as input to a trained deep neural network 10. This input is transformed into an equivalent image 40 of a stained tissue section of the same field of view using a class-conditional generative adversarial network (c-GAN) (see FIG. 15). During network training, the class input to the deep neural network 10 is identified as the class of the ground truth image corresponding to that image. In one embodiment, class conditioning can be implemented as a set of "one-hot" encoding matrices ( FIG. 15 ) with the same vertical and horizontal dimensions as the network input images. During training, the classes can be varied to any number. Alternatively, by modifying the class encoding matrix M to use a mixture of multiple classes rather than a simple one-hot encoding matrix, multiple stains can be mixed to create an output image 40 with unique stains that have features emanating from the various stains learned by the deep neural network 10 ( FIGS. 18A-18G ).

[0036]

[0065] Because the deep neural network 10 aims to learn the transformation from an autofluorescence image 20 of an unlabeled tissue specimen 22 to an image of a stained specimen (i.e., the gold standard), accurately aligning the FOVs is important. Furthermore, when multiple autofluorescence channels are used as inputs to the network 10, the various filter channels need to be aligned. Because four different stains (H&E, Masson's Trichrome, PAS, and Jones) were used, image preprocessing and alignment was implemented for each input and target image pair (training pair) from these four different stain datasets. The image preprocessing and alignment followed the global and local registration processes described herein and shown in FIG. 8. However, one major difference is that when using multiple autofluorescence channels as network inputs (i.e., DAPI and TxRed as shown herein), they need to be aligned. It was determined that even though images from the two channels were captured using the same microscope, corresponding FOVs from the two channels were not accurately aligned, especially at the edges of the FOV. Therefore, the elastic registration algorithm described herein was used to accurately align the multiple autofluorescence channels. The elastic registration algorithm matches local features in both channels (e.g., DAPI and TxRed) of an image by hierarchically dividing the image into smaller and smaller blocks while matching corresponding blocks. The calculated transformation map is then applied to the TxRed image to ensure it is aligned to the corresponding image from the DAPI channel. Finally, the aligned images from the two channels are aligned to a whole-slide image containing both the DAPI and TxRed channels.

[0037]

[0066] At the end of the co-registration process, images 20 from one or more autofluorescence channels of the unlabeled tissue section are fully aligned to the corresponding brightfield image 48 of the histologically stained tissue section 22. Before feeding these aligned pairs to the deep neural network 10 for training, normalization is implemented on the DAPI and TxRed whole-slide images, respectively. This whole-slide normalization is performed by subtracting the mean value across the tissue sample and dividing it by the standard deviation between pixel values. Following the training procedure, multiple virtual stains can be applied to the image 20 using a single algorithm on the same input image 20 using class conditioning. In other words, additional networks are not required for each individual stain. A single trained neural network can be used to apply one or more digital / virtual stains to the input image 20.

[0038]

[0067] FIG. 16 shows a display illustrating a graphical user interface (GUI) used to display an output image 40 from a trained deep neural network 10, according to one embodiment. In this embodiment, a user is provided with a list of tools (e.g., pointers, markers, erasers, loops, highlighters, etc.) that can be used to identify and select specific regions of the output image 40 for virtual staining. For example, a user may use one or more of the tools to select specific areas or regions of tissue within the output image 40 for virtual staining. In this particular example, three areas are identified by hash lines (areas A, B, and C) manually selected by the user. The user may be provided with a palette of stains to select from to stain the regions. For example, the user may be provided with staining options for staining the tissue (e.g., Masson's Trichrome, Jones, H&E). The user can then select various areas for staining with one or more of these stains. This results in a microstructural output such as that shown in FIG. 17. In a separate embodiment, image processing software 104 may be used to automatically identify or segment specific regions of the output image 40. For example, image segmentation and computer-generated mapping may be used to identify specific histological features within the imaged sample 22. For example, cell nuclei, specific cells, or tissue types may be automatically identified by the image processing software 104. These automatically identified regions of interest in the output image 40 may be manually and / or automatically stained with one or more stains / stain combinations.

[0039]

[0068] In yet another embodiment, a blend of multiple stains may be generated in the output image 40. For example, multiple stains may be blended in various ratios or percentages to create unique stains or stain combinations. Examples are disclosed herein (FIGS. 18E-18G) using blends of stains with different input class conditions of stain ratios (e.g., a virtual 3:1 H&E:Jones stain in FIG. 18E) to generate a virtually stained network output image 40. FIG. 18F shows a virtually blended H&E:Jones stain at a 1:1 ratio. FIG. 18F shows a virtually blended H&E:Jones stain at a 1:3 ratio.

[0040]

[0069] While the digital / virtual staining method may be used for fluorescence images obtained from an unlabeled sample 22, it should be understood that the multi-stain digital / virtual staining method may also be used for other microscopic imaging modalities. These include, for example, brightfield microscopy images of stained or unstained sample 22. In other examples, the microscope may include a single-photon fluorescence microscope, a multi-photon microscope, a second-harmonic generation microscope, a high-order harmonic generation microscope, an optical coherence tomography (OCT) microscope, a confocal reflectance microscope, a fluorescence lifetime microscope, a Raman spectroscopy microscope, a brightfield microscope, a darkfield microscope, a phase contrast microscope, a quantitative phase microscope, a structured illumination microscope, a super-resolution microscope, a light sheet microscope, a computational microscope, a ptychographic microscope, a synthetic aperture-based microscope, and a total internal reflection microscope.

[0041]

[0070] Digital / virtual staining methods can be used with any number of stains, including, for example, hematoxylin and eosin (H&E) stain, hematoxylin, eosin, Jones silver stain, Masson's trichrome stain, Periodic Acid Schiff (PAS) stain, Congo red stain, Alcian blue stain, blue iron, silver nitrate, trichrome stain, Ziehl-Nielsen, Grocott methenamine silver (GMS) stain, Gram stain, acid stain, basic stain, silver stain, Nissl, Weigert stain, Golgi stain, Luxol fast blue stain, toluidine blue, Genta, Mallory trichrome stain, Gomori trichrome, Van Gieson, Giemsa, Sudan black, Perls-Prussian, Best carmine, acridine orange, immunofluorescence stain, immunohistochemical stain, Kinyon-Cold stain, Albert stain, flagella stain, endospore stain, Nigrosin, and India Ink stain. The sample 22 to be imaged may comprise a tissue section or cells / cellular structures.

[0042]

[0071] In another embodiment, the trained deep neural network 10', 10" may operate to virtually destain (and optionally virtually re-stain the sample with a different stain). In this embodiment, a first trained deep neural network 10' is provided that is executed by image processing software 104 using one or more processors 102 of a computing device 100 (see FIG. 1), and the first trained deep neural network 10' is configured to process a plurality of matching chemically stained microscopic images or image patches 80. train , and their corresponding unstained ground truth microscopic images or image patches 82 of the same sample obtained before chemical staining. GT 19A. Thus, in this embodiment, the ground truth images used to train the deep neural network 10' are unstained (or unlabeled) images, while the training images are chemically stained (e.g., IHC stained) images. In this embodiment, the microscope image 84 of the chemically stained sample 12 is test(i.e., the sample 12 to be tested or imaged) is obtained using a microscope 110 of any of the types described herein. An image 84 of the chemically stained sample is then obtained. test is input to the trained deep neural network 10′, which then outputs a virtually destained microscopic image 86 of the sample 12 that is substantially equivalent to a corresponding image of the same sample 12 obtained without chemical staining (i.e., unstained or unlabeled).

[0043]

[0072] Optionally, as seen in FIG. 19B , a trained second deep neural network 10″ is provided, executed by image processing software 104 using one or more processors 102 of a computing device 100 (see FIG. 1 ), wherein the trained second deep neural network 10″ is configured to generate a plurality of matched unstained or unlabeled microscopic images or image patches 88 of the same sample obtained before chemical staining. train , and their corresponding stained ground truth microscopic images or image patches 90 GT (Figure 19A). In this embodiment, the stained ground truth is trained on train The virtual destained microscopic image 86 of the sample 12 (i.e., the sample 12 being tested or imaged) is then input into the trained second deep neural network 10'', which then outputs a stained or labeled microscopic image 92 of the sample 12 that is substantially equivalent to a corresponding image of the same sample 12 obtained with a different chemical stain.

[0044]

[0073] For example, stains may be converted from / to one of the following: hematoxylin and eosin (H&E) stain, hematoxylin, eosin, Jones silver stain, Masson's trichrome stain, Periodic Acid Schiff (PAS) stain, Congo red stain, Alcian blue stain, iron blue, silver nitrate, trichrome stain, Ziehl-Nielsen, Grocott methenamine silver (GMS) stain, Gram stain, acid stain, basic stain, silver stain, Nissl, Weigert stain, Golgi stain, Luxol fast blue stain, toluidine blue, Genta, Mallory's trichrome stain, Gomori trichrome, Van Gieson, Giemsa, Sudan black, Perls-Prussian, Best carmine, acridine orange, immunofluorescence stain, immunohistochemical stain, Kinyon cold stain, Albert stain, flagella stain, endospore stain, Nigrosin, or India ink stain.

[0045]

[0074] It should be understood that this embodiment can be combined with machine learning-based training of out-of-focus and in-focus images. Thus, a network (e.g., deep neural network 10′) can be trained to focus or eliminate optical aberrations in addition to destaining / restaining. Furthermore, for all embodiments described herein, the input image 20 can, in some cases, have the same or substantially similar numerical aperture and resolution as the ground truth (GT) image. Alternatively, the input image 20 can have a lower numerical aperture and lower resolution compared to the ground truth (GT) image.

[0046]

[0075] Experimental - Digital staining of unlabeled tissue using autofluorescence

[0076] Virtual staining of tissue samples

[0077] The system 2 and methods described herein were tested and validated using various combinations of tissue section samples 22 and stains. Following training of the CNN-based deep neural network 10, its inferences were blindly tested by feeding autofluorescence images 20 of unlabeled tissue sections 22 that did not overlap with the images used in the training or validation sets. Figures 4A-4H show the results of a salivary gland tissue section digitally / virtually stained to match an H&E-stained brightfield image 48 (i.e., ground truth image) of the same sample 22. These results demonstrate the ability of system 2 to convert the fluorescent image 20 of the unlabeled tissue section 22 into a brightfield-equivalent image 40, displaying the correct color scheme expected from H&E-stained tissue, including various components such as epithelial-like cells, cell nuclei, nucleoli, stroma, and collagen. Evaluation of both Figures 3C and 3D shows that the H&E staining reveals islets of infiltrating tumor cells within the subcutaneous fibroadipose tissue. Note that nuclear details, including nucleoli (arrows) and distinct chromatin texture, are clearly visible in both panels. Similarly, in Figures 3G and 3H, H&E staining reveals invasive squamous cell carcinoma. Desmoplastic reaction with edematous myxoid change (asterisks) in the adjacent stroma is clearly discernible in both stains.

[0047]

[0078] Next, the deep network 10 was trained to digitally / virtually stain other tissue types with two different stains: Jones methenamine silver stain (kidney) and Masson trichrome stain (liver and lung). Figures 4A-4H and 5A-5P summarize the results of the deep learning-based digital / virtual staining of these tissue sections 22, which are in excellent agreement with brightfield images 48 of the same samples 22 captured after the histochemical staining process. These results demonstrate that the trained deep neural network 10 can infer the staining patterns of different types of histological stains used for different tissue types from a single fluorescent image 20 of an unlabeled specimen (i.e., without histochemical staining). With the same overall conclusion as in Figures 3A-3H, a pathologist confirmed that the neural network output images in Figures 4C and 5G correctly reveal histological features corresponding to hepatocytes, sinusoidal spaces, collagen, and lipid droplets, consistent with those appearing in brightfield images 48 (Figures 5D and 5H) of the same tissue sample 22 captured after chemical staining. Similarly, the same expert also confirmed that the deep neural network output images 40 reported in Figures 5K and 5O (lung) reveal concordantly stained histological features corresponding to blood vessels, collagen, and alveolar spaces, as seen in brightfield images 48 (Figures 6L and 6P) of the same tissue sample 22 imaged after chemical staining.

[0048]

[0079] Digitally / virtually stained output images 40 from the trained deep neural network 10 were compared with standard histochemically stained images 48 to diagnose multiple conditions in multiple types of tissues that were either formalin-fixed, paraffin-embedded (FFPE) or frozen sectioned. The results are summarized in Table 1 below. Analysis of 15 tissue sections by four board-certified pathologists (who were blinded to the virtual staining technique) demonstrated 100% non-major disagreement, defined as no clinically significant difference in diagnosis between expert observers. The "time to diagnosis" varied significantly between observers, ranging from an average of 10 seconds per image for observer 2 to 276 seconds per image for observer 3. However, intraobserver variability was very small, and there was a trend toward shorter times to diagnosis for virtually stained slide images 40 for all observers except observer 2, i.e., ~10 seconds per image for both virtual slide images 40 and histologically stained slide images 48. These results indicate very similar diagnostic utility between the two imaging modalities. TIFF2026041721000002.tif234170TIFF2026041721000003.tif55170

[0049]

[0080] Blind assessment of staining efficacy for whole slide images (WSI)

[0081] After evaluating the differences between tissue sections and staining, the capabilities of the Virtual Staining System 2 were tested in a specialized staining histology workflow. Specifically, the autofluorescence distribution of 15 unlabeled liver tissue sections and 13 unlabeled kidney tissue sections was imaged using a 20x / 0.75 NA objective. All liver and kidney tissue sections were obtained from various patients and included both small biopsies and larger resections. All tissue sections were obtained from FFPE tissue but were not coverslipped. After autofluorescence scanning, the tissue sections were histologically stained with Masson's trichrome (4 μm liver tissue sections) and Jones stain (2 μm kidney tissue sections). The WSIs were then divided into a training set and a test set. For the liver slide cohort, seven WSIs were used to train the virtual staining algorithm and eight WSIs were used for blind testing; for the kidney slide cohort, six WSIs were used to train the algorithm and seven WSIs were used for testing. Study pathologists were blinded to the staining technique for each WSI and were asked to apply a numerical grade of 1 to 4 to the quality of various stains: 4 = perfect, 3 = very good, 2 = acceptable, and 1 = unacceptable. Second, study pathologists applied the same score scale (1 to 4) to specific features: for liver only, nuclear detail (ND), cytoplasmic detail (CD), and extracellular fibrosis (EF). These results are summarized below in Table 2 (liver) and Table 3 (kidney) (winners are shown in bold). The data demonstrate that, without a clear preferred staining technique (virtual vs. histological), pathologists were able to recognize histopathological features by both staining technique and with a high degree of agreement between techniques. TIFF2026041721000004.tif145169TIFF2026041721000005.tif123169

[0050]

[0082] Quantifying network output image quality

[0083] Next, beyond the visual comparisons provided in Figures 3A-3H, 4A-4H, and 5A-5P, the results of the trained deep neural network 10 were quantified by first calculating pixel-level differences between brightfield images 48 of chemically stained samples 22 and digitally / virtually stained images 40 synthesized using the deep neural network 10 without the use of labels / stains. Table 4 below summarizes this comparison for various combinations of tissue type and stain using the YCbCr color space, where the chroma components Cb and Cr define the overall color and Y defines the luminance component of the image. The results of this comparison reveal that the average difference between these two sets of images is <5% and <16% in the chroma (Cb, Cr) and luminance (Y) channels, respectively. A second metric was then used to further quantify the comparison: the structural similarity index (SSIM), which is commonly used to predict the score a human observer would give an image when compared to a reference image (Equation 8 herein). The SSIM ranges from 0 to 1, with 1 defining a score for identical images. The results of this SSIM quantification are also summarized in Table 4 and very clearly demonstrate the strong structural similarity between the network output images40 and the bright-field images48 of the chemically stained samples. TIFF2026041721000006.tif66170

[0051]

[0084] Because of the uncontrolled variations and structural changes that tissue undergoes during the histochemical staining process and associated dehydration and clearing steps, brightfield images 48 of chemically stained tissue samples 22 do not, in fact, provide a true gold standard for specific SSIM and YCbCr analysis of network output images 40. Another variation noted in some of the images was that the automated microscope scanning software selected different autofocus planes for the two imaging modalities. All these variations create some challenges in absolute quantitative comparison of the two sets of images (i.e., network output 40 of unlabeled tissue vs. brightfield images 48 of the same tissue after the histological staining process).

[0052]

[0085] Staining standardization

[0086] An interesting by-product of the digital / virtual staining system 2 could be stain standardization. In other words, the trained deep neural network 10 converges to a “generic stain” coloring scheme, whereby the variability of histologically stained tissue images 48 is greater than the variability of virtually stained tissue images 40. The virtual stain coloring is solely a result of its training (i.e., the gold-standard histological stain used in the training phase) and can be further adjusted based on the pathologist's preferences by retraining the network with new stain colorings. Such “improved” training can be created from scratch or accelerated through transfer learning. This potential stain standardization using deep learning could ameliorate the adverse effects of interpersonal variability at various stages of sample preparation, create a common foundation across various clinical laboratories, enhance clinicians' diagnostic workflows, and aid in the development of new algorithms, such as automated tissue metastasis detection or grading of various types of cancer, among others.

[0053]

[0087] Transfer learning to other tissue-stain combinations

[0088] Using the concept of transfer learning, the training procedure for new tissues and / or stain types can converge much faster while improving performance, i.e., reaching a better minimum in the training cost / loss function. This means that pre-trained CNN models from different tissue-stain combinations can be used to initialize the deep neural network 10 to statistically learn the virtual stains of new combinations. Figures 6A-6C illustrate the favorable attributes of such an approach: a new deep neural network 10 was trained to virtually stain autofluorescence images 20 of unstained thyroid tissue sections and initialized using the weights and biases of another deep neural network 10 previously trained for H&E virtual staining of salivary glands. The evolution of the loss metric as a function of the number of iterations used in the training phase clearly demonstrates that the new thyroid deep network 10 rapidly converges to a lower minimum compared to the same network architecture trained from scratch using random initialization, as seen in Figure 6A. Figure 6B compares output images 40 of this thyroid network 10 at different stages of its training process, further demonstrating the impact of transfer learning to rapidly adapt the presented approach to new tissue / stain combinations. The network output image 40, after a training phase involving, for example, ≥6,000 iterations, shows that the cell nuclei exhibit irregular contours, nuclear grooves, and chromatin pallor suggestive of papillary thyroid carcinoma; the cells also exhibit mild to moderate eosinophilic granular cytoplasm, and the fibrovascular core of the network output image shows an increase in inflammatory cells, including lymphocytes and plasma cells. Figure 6C shows the corresponding H&E-stained brightfield image 48.

[0054]

[0089] Use of multiple fluorescence channels at various resolutions

[0090] The method using the trained deep neural network 10 can be combined with other excitation wavelengths and / or imaging modalities to enhance its inference performance for various tissue components. For example, melanin detection was attempted in skin tissue section samples using virtual H&E staining. However, melanin exhibits a weak autofluorescence signal at the DAPI excitation / emission wavelengths measured in the experimental system described herein, and therefore was not clearly identified in the network output. One potential method to increase melanin autofluorescence is to image the sample while it is in an oxidizing solution. However, a more practical alternative was used where an additional autofluorescence channel, such as from a Cy5 filter (excitation 628 nm / emission 692 nm), was employed so that the melanin signal could be enhanced and accurately inferred by the trained deep neural network 10. By training the network 10 using both the DAPI and Cy5 autofluorescence channels, the trained deep neural network 10 was able to successfully determine where melanin occurred in the sample, as shown in Figures 7A-7C. In contrast, when only the DAPI channel was used (Figure 7A), the network 10 was unable to distinguish areas containing melanin (areas that appear white). In other words, the additional autofluorescence information from the Cy5 channel was used by the network 10 to distinguish melanin from background tissue. For the results shown in Figures 7A-7C, it was assumed that the most necessary information would be detected in the high-resolution DAPI scan, and additional information (e.g., the presence of melanin) could be encoded in the low-resolution scan. Therefore, the image 20 was acquired using a low-resolution objective lens (10x / 0.45NA) for the Cy5 channel to complement the high-resolution DAPI scan (20x / 0.75NA). In this way, two different channels were used, with one of the channels used at low resolution to identify melanin. This may require multiple scans of the sample 22 using the fluorescence microscope 110. In yet another multi-channel embodiment, multiple images 20 may be fed into the trained deep neural network 10.This may include, for example, raw fluorescence images combined with one or more images that have undergone linear or non-linear image pre-processing such as contrast enhancement, contrast inversion and image filtering.

[0055]

[0091] The system 2 and method described herein demonstrate the ability to digitally / virtually stain unlabeled tissue sections 22 using supervised deep learning techniques that use as input a single fluorescent image 20 of a sample captured by a standard fluorescent microscope 110 and filter set (or, in other embodiments, multiple fluorescent images 20 are input when multiple fluorescent channels are used). This statistical learning-based method has the potential to reshape the clinical workflow of histopathology and potentially provide a digital alternative to standard techniques for histochemical staining of tissue samples 22, thereby benefiting from a variety of imaging modalities, including fluorescent microscopy, nonlinear microscopy, holographic microscopy, stimulated Raman scattering microscopy, and optical coherence tomography, among others. Here, the method demonstrates how fixed, unstained tissue samples 22 can be used to provide meaningful comparisons with chemically stained tissue samples, which is essential for training deep neural networks 10 and blindly testing the performance of the network output against clinically approved methods. However, the presented deep learning-based approach is broadly applicable to various types and conditions of samples 22, including unsectioned fresh tissue samples (e.g., following a biopsy procedure) without the use of labels or stains. Following its training, the deep neural network 10 can be used to digitally / virtually stain images of unlabeled fresh tissue samples 22, acquired, for example, using UV or deep-UV excitation or nonlinear microscopy. For example, Raman microscopy can provide a very rich set of unlabeled biochemical signatures that can further enhance the effectiveness of the virtual stains that the neural network learns.

[0056]

[0092] A key part of the training process is matching fluorescent images 20 of unlabeled tissue samples 22 with their corresponding brightfield images 48 after the histochemical staining process (i.e., chemically stained images). It should be noted that during the staining process and related steps, some tissue components may be lost or distorted in a way that misleads the loss / cost function during the training phase. However, this is only a challenge associated with training and validation and does not impose limitations on the practical application of a fully trained deep neural network 10 for virtual staining of unlabeled tissue samples 22. To ensure the quality of the training and validation phases and minimize the impact of this challenge on the network's performance, a threshold for acceptable correlation values ​​between the two sets of images (i.e., before and after the histochemical staining process) was set, and mismatched image pairs were eliminated from the training / validation sets to ensure that the deep neural network 10 was learning actual signals rather than perturbations to tissue morphology due to the chemical staining process. In fact, this process of cleaning the training / validation image data can be iterative: starting by roughly eliminating obviously altered samples and thus converging on the trained neural network 10. After this initial training stage, the output images 40 of each sample in the available image set can be screened against their corresponding brightfield images 48 to reject some additional images and further set a finer threshold to further clean the training / validation image set. Repeating this process several times can not only further refine the image set but also improve the performance of the final trained deep neural network 10.

[0057]

[0093] The methodology described above alleviates some of the training challenges caused by the random loss of some tissue features after the histological staining process. Indeed, this highlights another motivation for skipping the laborious and costly steps involved in histochemical staining, as it is easy to preserve local tissue histology in a non-labeling manner without the need for an expert to handle some of the delicate steps of the staining process, which sometimes require viewing the tissue under a microscope.

[0058]

[0094] Using a PC desktop, the training phase of the deep neural network 10 takes a significant amount of time (e.g., ~13 hours for the salivary gland network). However, this entire process can be significantly accelerated by using dedicated computer hardware based on a GPU. Furthermore, as already highlighted in Figures 6A-6C, transfer learning provides a "warm start" to the training phase of new tissue / stain combinations, significantly speeding up the entire process. Once the deep neural network 10 is trained, digital / virtual staining of the sample image 40 is performed in a single, non-iterative manner, without requiring a trial-and-error approach or the tuning of any parameters to achieve optimal results. Based on its feedforward and non-iterative architecture, the deep neural network 10 rapidly outputs a virtually stained image in less than 1 second (e.g., 0.59 seconds, corresponding to a sample field of view of ~0.33 mm × 0.33 mm). Furthermore, GPU-based acceleration has the potential to achieve real-time or near-real-time performance in outputting the digital / virtually stained image 40, which may be particularly relevant in operating room or in vivo imaging applications. It should be understood that this method may also be used in in vitro imaging applications such as those described herein.

[0059]

[0095] The implemented digital / virtual staining procedure is based on training a separate CNN deep neural network 10 for each tissue / stain combination. If autofluorescence images 20 with different tissue / stain combinations were fed into the CNN-based deep neural network 10, it would not function as desired. However, for histological applications, this is not a limitation because the tissue type and stain type are predetermined for each sample 22 of interest, and therefore, the specific CNN selection for creating a digitally / virtually stained image 40 from an autofluorescence image 20 of an unlabeled sample 22 does not require additional information or resources. Of course, more general CNN models can be trained for multiple tissue / stain combinations by, for example, increasing the number of trained parameters in the model, at the expense of longer training and inference times. Another approach is the possibility of the system 2 and method performing multiple virtual stainings on the same unlabeled tissue type.

[0060]

[0096] A key advantage of System 2 is its high flexibility. System 2 can respond to feedback to statistically repair its performance by penalizing accordingly when diagnostic failures are detected by clinical comparison. This iterative training and transfer learning cycle based on clinical evaluation of the performance of the network output helps optimize the robustness and clinical impact of the presented approach. Finally, this method and System 2 can be used for microguided molecular analysis at the unstained tissue level by locally identifying regions of interest based on virtual staining and using this information to guide subsequent analysis of the tissue, for example, for microimmunohistochemistry or sequencing. This type of virtual microguidance to unlabeled tissue samples could facilitate high-throughput identification of disease subtypes and also aid in the development of customized treatments for patients.

[0061]

[0097] Sample preparation

[0098] Formalin-fixed, paraffin-embedded tissue sections, 2 μm thick, were deparaffinized using xylene and mounted onto standard glass slides using Cytoseal™ (Thermo-Fisher Scientific, Waltham, MA, USA), followed by a coverslip (Fisherfinest™, 24x50-1, Fisher Scientific, Pittsburgh, PA, USA). Following an initial autofluorescence imaging process of the unlabeled tissue samples (using DAPI excitation and emission filter sets), the slides were placed in xylene for approximately 48 hours or until the coverslip could be removed without damaging the tissue. Once the coverslip was removed, the slides were immersed in absolute alcohol, 95% alcohol (approximately 30 times), and then washed in deionized water for approximately 1 minute. This step was followed by the corresponding staining procedure used for H&E, Masson's Trichrome, or Jones staining. This tissue processing pass is only used for training and validation of the approach and is not required after the network is trained. To test the system and method, various tissue and staining combinations were used: salivary gland and thyroid tissue sections were stained with H&E, kidney tissue sections were stained with Jones stain, while liver and lung tissue sections were stained with Masson's trichrome.

[0062]

[0099] In the WSI study, 2- to 4-μm-thick FFPE tissue sections were not coverslipped during the autofluorescence imaging stage. Following autofluorescence imaging, tissue samples were histologically stained as described above (Masson's Trichrome for liver tissue sections and Jones for kidney tissue sections). Unstained frozen samples were prepared by embedding the tissue sections in OCT (Tissue Tek, SAKURA FINETEK USA INC.) and immersing them in 2-methylbutane with dry ice. The frozen sections were then cut into 4-μm sections and placed in a freezer until imaging. Following the imaging process, the tissue sections were washed with 70% alcohol, stained with H&E, and coverslipped. Samples were obtained from the Translational Pathology Core Laboratory (TPCL) and prepared by the UCLA Histology Laboratory. Renal tissue sections from diabetic and non-diabetic patients were obtained under IRB18-001029 (UCLA). All samples were obtained after de-identification of patient-related information and prepared from existing specimens, therefore, this work did not interfere with standard practices of care or sample collection procedures.

[0063]

[0100] Data Acquisition

[0101] Autofluorescence images 20 of unlabeled tissue were captured using a conventional fluorescence microscope 110 (IX83, Olympus Corporation, Tokyo, Japan) equipped with a motorized stage, with the image acquisition process controlled by MetaMorph® microscope automation software (Molecular Devices, LLC). Unstained tissue samples were excited with near-UV light and imaged using a DAPI filter cube (OSFI3-DAPI-5060C, excitation wavelength 377 nm / 50 nm bandwidth, emission wavelength 447 nm / 60 nm bandwidth) with a 40× / 0.95 NA objective (Olympus UPLSAPO 40X2 / 0.95 NA, WD 0.18) or a 20× / 0.75 NA objective (Olympus UPLSAPO 20X / 0.75 NA, WD 0.65). For melanin inference, autofluorescence images of the samples were additionally acquired using a Cy5 filter cube (CY5-4040C-OFX, excitation wavelength 628 nm / 40 nm bandwidth, emission wavelength 692 nm / 40 nm bandwidth) with a 10x / 0.4 NA objective lens (Olympus UPLSAPO10X2). Each autofluorescence image was captured with an exposure time of ~500 ms using a Scientific CMOS sensor (ORCA-flash 4.0 v2, Hamamatsu Photonics K.K., Shizuoka, Japan). Bright-field images (used for training and validation) were acquired using a slide scanner microscope (Aperio AT, Leica Biosystems) with a 20x / 0.75 NA objective lens (Plan Apo) equipped with a 2x magnification adapter.

[0064]

[0102] Image preprocessing and alignment

[0103] Because the deep neural network 10 aims to learn a statistical transformation between an autofluorescence image 20 of a chemically unstained tissue sample 22 and a bright-field image 48 of the same tissue sample 22 after histochemical staining, it is important to accurately match the FOVs of the input and target images (i.e., the unstained autofluorescence image 20 and the stained bright-field image 48). An overall scheme illustrating the global and local image registration process implemented in MATLAB (The MathWorks Inc., Natick, Massachusetts, USA) is shown in Figure 8. The first step in this process is to find candidate features for matching the unstained autofluorescence image and the chemically stained bright-field image. To this end, each autofluorescence image 20 (2048 × 2048 pixels) is downsampled to match the effective pixel size of the bright-field microscopy image. This produces a 1351 x 1351 pixel unstained autofluorescent tissue image, which is contrast-enhanced by saturating the bottom and top 1% of all pixel values ​​and inverted to better represent the color map of the grayscale-converted whole slide image (image 20a in Figure 8). A correlation patch process 60 is then performed to calculate a normalized correlation score matrix by correlating each of the 1351 x 1351 pixel patches with a corresponding patch of the same size extracted from the whole slide grayscale image 48a. The entry in this matrix with the highest score represents the FOV most likely to match between the two imaging modalities. Using this information (which defines the coordinate pairs), the matched FOV of the original whole slide brightfield image 48 is cropped 48c to create a target image 48d. Following this FOV-matching procedure 60, the autofluorescent image 20 and the brightfield microscope image 48 are roughly matched. However, they are not precisely aligned at the individual pixel level due to slight discrepancies in sample placement in the two different microscopic imaging experiments (autofluorescence followed by bright field), which randomly results in slight rotation angles (e.g., ∼1–2 degrees) between the input and target images of the same sample.

[0065]

[0104] The second part of the input target matching process involves a global registration step 64, which corrects for this slight rotation angle between the autofluorescence and bright-field images. This is done by extracting feature vectors (descriptors) and their corresponding locations from the image pairs and matching the features using the extracted descriptors. The M-estimation sample consensus (MSAC) algorithm, a variant of the random sample consensus (RANSAC) algorithm, is then used to find the transformation matrix corresponding to the matched pair. Finally, the angle-corrected image 48e is obtained by applying this transformation matrix to the original bright-field microscope image patch 48d. Following the application of this rotation, the images 20b, 48e are cropped by an additional 100 pixels (50 pixels on each side) to accommodate undefined pixel values ​​at the image borders due to the rotation angle correction.

[0066]

[0105] Finally, for local feature registration step 68, elastic image registration matches local features in both sets of images (autofluorescence 20b vs. brightfield 48e) by hierarchically matching corresponding blocks from largest to smallest. A neural network 71 is used to learn the transformation between the roughly matched images. This network 71 uses the same structure as network 10 in FIG. 10. A small number of iterations is used so that network 71 learns only the exact color mapping and not the spatial transformation between the input and labeled images. The transformation map computed from this step is finally applied to each brightfield image patch 48e. At the end of these registration steps 60, 64, and 68, the autofluorescence image patches 20b and their corresponding brightfield tissue image patches 48f are precisely matched to each other and can be used as input and labeled pairs for training deep neural network 10, allowing the network to focus its learning solely on the virtual histological staining problem.

[0067]

[0106] A similar process was used for the 20x objective images (used to generate the data in Tables 2 and 3). Instead of downsampling the autofluorescence images 20, the brightfield microscope images 48 were downsampled to 75.85% of their original size so that they matched the lower magnification images. Furthermore, to create whole-slide images using these 20x images, additional shading correction and normalization techniques were applied. Before being fed to network 71, each field was normalized by subtracting the mean value across the slide and dividing it by the standard deviation between pixel values. This normalizes the network input both within each slide and across slides. Finally, shading correction was applied to each image to account for the lower relative intensity measured at the edges of each field.

[0068]

[0107] Deep Neural Network Architecture and Training

[0108] Here, a GAN architecture was used to learn the transformation of unlabeled, unstained autofluorescent input images 20 to corresponding bright-field images 48 of chemically stained samples. Standard convolutional neural network-based training learns to minimize a loss / cost function between the network's output and the target label. Therefore, the selection of this loss function 69 (Figures 9 and 10) is a critical component of deep network design. For example, simply selecting an l2-norm penalty as the cost function tends to produce blurry results, since the network averages the weighted probabilities of all plausible outcomes; therefore, an additional regularization term is typically required to guide the network to maintain the desired sharp sample features in the network's output. GANs circumvent this challenge by learning a criterion aimed at accurately classifying the deep network's output images as real or fake (i.e., whether the staining is hypothetically correct or incorrect). This disallows output images that are inconsistent with the desired label, allowing the loss function to adapt to the data and the desired task at hand. To achieve this goal, the GAN training procedure involves training two different networks, as shown in Figures 9 and 10: (i) a generator network 70 whose goal, in this case, is to learn a statistical transformation between an unstained autofluorescent input image 20 and a corresponding brightfield image 48 of the same sample 12 after a histological staining process; and (ii) a discriminator network 74 that learns how to distinguish between true brightfield images of stained tissue sections and the generator network's output image. Ultimately, the desired outcome of this training process is a trained deep neural network 10 that transforms an unstained autofluorescent input image 20 into a digitally stained image 40 that is indistinguishable from a stained brightfield image 48 of the same sample 22. For this task, the loss functions 69 for the generator 70 and discriminator 74 were defined as follows:

[0109] TIFF2026041721000007.tif15170

[0110] where D denotes the discriminator network output and z labelshows a bright-field image of chemically stained tissue, z output denotes the generator network. The generator loss function is the generator loss (l generator We balance the pixel-wise MSE of the generator network output image with respect to its label, the total variation (TV) operator of the output image, and the discriminator network prediction of the output image using regularization parameters (λ, α) empirically set to different values ​​corresponding to ~2% and ~20%, respectively, of the output image. The TV operator for image z is defined as:

[0111] TIFF2026041721000008.tif12170

[0112] where p and q are pixel indices. Based on equation (1), the discriminator seeks to minimize the output loss while maximizing the probability of correctly classifying the actual label (i.e., the bright-field image of the chemically stained tissue). Ideally, the discriminator network should be able to minimize the output loss D(z label )=1 and D(z output ) = 0, but if the generator is successfully trained by the GAN, D(z output ) ideally converges to 0.5.

[0069]

[0113] The generator deep neural network architecture 70 is detailed in Figure 10. The input image 20 is processed by the network 70 in a multi-scale manner using downsampling and upsampling passes to help the network learn the virtual staining task at various scales. The downsampling pass consists of four separate steps (four blocks #1, #2, #3, #4), each containing one residual block, and each of the steps produces a feature map x k+1 Feature map x k To map:

[0114] TIFF2026041721000009.tif10170

[0115] where CONV{.} is the convolution operator (including the bias term), k1, k2, and k3 denote the serial numbers of the convolutional layers, and LReLU[.] is the nonlinear activation function (i.e., leaky rectified linear unit) used throughout the network, defined as:

[0116] TIFF2026041721000010.tif13170

[0070]

[0117] The number of input channels at each level of the downsampling path was set to: 1, 64, 128, 256, while the number of output channels of the downsampling path was set to: 64, 128, 256, 512. To avoid the mismatch in the dimensions of each block, the feature map x k is the number of channels x k+1 The connection between each downsampling level is a 2x2 average pooling layer with a stride of 2 pixels, which downsamples the feature maps by a factor of 4 (by a factor of 2 in each direction). Following the output of the fourth downsampling block, another convolutional layer (CL) maintains the number of feature maps as 512 before connecting it to the upsampling path. The upsampling path consists of four symmetric upsampling steps (#1, #2, #3, #4), each containing one convolutional block. The feature map y k The feature map k+1 The convolution block operation that maps to is given by the following formula: TIFF2026041721000011.tif9170

[0118] Here, CONCAT(.) is the concatenation between two feature maps that merge the number of channels, US{.} is the upsampling operator, and k4, k5, and k6 indicate the serial numbers of the convolutional layers. The number of input channels at each level of the upsampling path was set to 1024, 512, 256, and 128, and the number of output channels at each level of the upsampling path was set to 256, 128, 64, and 32, respectively. The last layer is a convolutional layer (CL) that maps the 32 channels represented by the YcbCr color map to 3 channels. Both the generator network and the discriminator network were trained with a patch size of 256 × 256 pixels.

[0071]

[0119] The discriminator network, summarized in Figure 10, receives three input channels, which correspond to the YcbCr color space of the input image 40YcbCr, 48YcbCr. This input is then converted to a 64-channel representation using a convolutional layer, followed by five blocks of the following operators:

[0120] TIFF2026041721000012.tif8170

[0121] where k1 and k2 denote the serial numbers of the convolutional layers. The number of channels in each layer was 3, 64, 64, 128, 128, 256, 256, 512, 512, 1024, 1024, and 2048. The next layer was an average pooling layer with a filter size equal to the patch size (256 × 256) that yielded a vector of 2048 entries. The output of this average pooling layer was then fed into two fully connected layers (FC) with the following structure:

[0122] TIFF2026041721000013.tif7170

[0123] Here, FC stands for fully connected layer with learnable weights and biases. The first fully connected layer outputs a vector with 2048 entries, while the second layer outputs a scalar value. This scalar value is used as input to a sigmoid activation function D(z)=1 / (1+exp(-z)), which calculates the probability (between 0 and 1) of the discriminator network input being true / genuine or fake, i.e., ideally, D(z) = 1 / (1+exp(-z)), as shown by output 67 in Figure 10. label )=1.

[0072]

[0124] The convolutional kernels for the entire GAN were set to 3 × 3. These kernels were randomly initialized by using a truncated normal distribution with a standard deviation of 0.05 and a mean of 0; all network biases were initialized as 0. The learnable parameters were 1 × 10 for the generator network 70. -4 and 1 × 10 for discriminator network 74. -5 The deep neural network 10 was updated through the training phase by backpropagation (illustrated by the dotted arrows in Figure 10) using an adaptive moment estimation (Adam) optimizer with a learning rate of . Also, for each iteration of the discriminator 74, there were four iterations of the generator network 70 to avoid training plateaus following potential overfitting of the discriminator network to the labels. A batch size of 10 was used in training.

[0073]

[0125] Once all fields of view have been passed through the network 10, the whole slide images are stitched together using the Fiji Grid / Collection Stitching Plugin (see, e.g., Schindelin, J. et al., Fiji: an open-source platform for biological-image analysis. Nat. Methods 9, 676-682 (2012), which is incorporated herein by reference). This plugin calculates the exact overlap between each tile and linearly blends them into one large image. Overall, the inference and stitching are performed in a compact manner. 2 The process takes approximately 5 minutes and 30 seconds per section, respectively, and can be significantly improved using advances in hardware and software. Sections that are out of focus or have significant abnormalities (e.g., due to dust particles) in either the autofluorescence or brightfield images are cropped before being displayed to the pathologist. Finally, the images are exported to Zoomify format (designed to display large images using a standard web browser; http: / / zoomify.com / ) and uploaded to the GIGAmacro website (https: / / viewer.gigamacro.com / ) for easy access and viewing by the pathologist.

[0074]

[0126] Implementation details

[0127] Other implementation details, including the number of trained patches, number of epochs, and training time, are shown in Table 5 below. The digital / virtual dye deep neural network was implemented using Python version 3.5.0. The GAN was implemented using the TensorFlow framework version 1.4.0. Other Python libraries used were os, time, tqdm, Python Imaging Library (PIL), SciPy, glob, ops, sys, and numpy. The software was implemented on a desktop computer with a Core i7-7700K CPU @ 4.2 GHz (Intel) and 64 GB of RAM, running the Windows 10 operating system (Microsoft). Network training and testing were performed using dual GeForce® GTX 1080Ti GPUs (Nvidia). TIFF2026041721000014.tif93170

[0075]

[0128] Experiment - Virtual staining of samples using fluorescence lifetime imaging (FLIM)

[0129] In this embodiment, a trained neural network 10 is used that enables virtual IHC staining of an unstained tissue sample 22 based on fluorescence lifetime imaging. The algorithm acquires a fluorescence lifetime image 20L of the unstained tissue sample 22 and outputs an image 40 that is sufficiently consistent with a brightfield image 48 of the same field after IHC staining. Using this method, the tedious and time-consuming IHC staining procedure can be replaced with virtual staining, which is significantly faster and allows tissue preservation for further analysis.

[0076]

[0130] Data Acquisition

[0131] Referring to Figure 11A, unstained formalin-fixed, paraffin-embedded (FFPE) breast tissue (e.g., obtained by biopsy B) was cut into thin 4 μm slices and fixed on standard microscope slides. These tissue sections were obtained under IRB18-001029. Prior to imaging, the tissue was deparaffinized with xylene and mounted on standard slides using Cytoseal (Thermo-Fisher Scientific). A standard fluorescence lifetime microscope (SP8-DIVE, Leica Microsystems) was used, equipped with a 20× / 0.75 NA objective (Leica HC PL APO20× / 0.75 IMM) and two separate hybrid photodetectors receiving fluorescent signals in the wavelength ranges of 435–485 nm and 535–595 nm, respectively. The microscope 110 used a 700 nm wavelength laser with excitation at ~0.3 W to image the autofluorescence lifetime of these unlabeled tissue sections 22 to generate images 20L. The scanning speed was 200 Hz at 1024 × 1024 pixels with a pixel size of 300 nm. Once the autofluorescence lifetime images 20L were obtained, the slides were IHC stained using standard HER2, ER, or PR staining procedures. Staining was performed by the UCLA Translational Pathology Core Laboratory (TPCL). These IHC-stained slides were then imaged using a commercially available slide-scanning microscope with a 20x / 0.75 NA objective (Aperio AT, Leica Biosystems) to generate target images 48 used to train, validate, and test the neural network 10.

[0077]

[0132] Image preprocessing and alignment

[0133] Because the deep neural network 10 aims to learn a transformation from the autofluorescence lifetime images 20L of the unlabeled specimen 22, it is important to accurately align the FOV between them and the corresponding brightfield image 48 of the target. Image preprocessing and alignment follow a global and local registration process, as described herein and shown in FIG. 8. At the end of the registration process, images from the lifetime channels from single / multiple autofluorescent and / or unlabeled tissue sections 22 are fully aligned to the corresponding brightfield image 48 of the IHC-stained tissue section. Before feeding these aligned image pairs 20L, 48 to the neural network 10, slide normalization is implemented on the fluorescence intensity images by subtracting the mean value across the slide and dividing it by the standard deviation between pixel values.

[0078]

[0134] Deep Neural Network Architecture, Training and Validation

[0135] For the trained deep neural network 10, a conditional GAN ​​architecture was used to learn the transformation from unlabeled, unstained autofluorescence lifetime input images 20L with three different stains (HER2, PR, and ER) to corresponding brightfield images 48. After aligning the autofluorescence lifetime images 20L to the brightfield images 48, these precisely aligned FOVs were randomly divided into overlapping patches of 256 × 256 pixels, which were then used to train the GAN-based deep neural network 10.

[0079]

[0136] The GAN-based neural network 10 consists of two deep neural networks: a generator network (G) and a discriminator network (D). For this task, the loss functions of the generator and discriminator are defined as follows:

[0137]

[0138] TIFF2026041721000015.tif28170

[0139] where the anisotropic total variation (TV) operator and the L1 norm are defined as:

[0140]

[0141] TIFF2026041721000016.tif14170

[0142] where D(·) and G(·) denote the outputs of the discriminator and generator networks, and z label shows bright-field images of histologically stained tissues, and z output denotes the output of the generator network.

[0080]

[0143] The structural similarity index (SSIM) is defined as follows:

[0144] TIFF2026041721000017.tif10170

[0145] where μ x , μ y is the mean of image x, y, TIFF2026041721000018.tif8162 is a variable of x, y, and σ x、y is the covariance of x and y, and c1, c2 are variables used to stabilize the division with a small denominator. An SSIM value of 1.0 refers to the same image. The generator loss function balances the pixel-wise SSIM and L1 norm of the generator network output image with respect to its label, the total variation (TV) operator of the output image, and the discriminator network prediction of the output image. The regularization parameters (μ, α, ν, λ) were set to (0.3, 0.7, 0.05, 0.002).

[0081]

[0146] The deep neural network architecture of generator G follows the structure of the deep neural network 10 shown in FIG. 11B (and described herein). However, in this implementation, network 10 begins with a convolutional layer that maps 200 the input lifetime and / or autofluorescence image data into 16 channels, followed by a downsampling pass consisting of five individual downsampling steps 202. The number of input channels at each level of the downsampling pass was set to 16, 32, 64, 128, and 256, while the number of output channels at each level of the downsampling pass was set to 32, 64, 128, 256, and 512. Following the center block convolutional layer 203, an upsampling pass consists of five symmetric upsampling steps 204. The number of input channels at each level of the upsampling pass was set to 1024, 512, 256, 128, and 64, while the number of output channels at each level of the upsampling pass was set to 256, 128, 64, 32, and 16, respectively. The final layer 206 is a convolutional layer that maps 16 channels to 3 channels using a tanh() activation function 208, represented by an RGB color map. Both the generator (G) and discriminator (D) networks are trained using an image patch size of 256x256 pixels.

[0082]

[0147] The discriminator network (D) receives three (i.e., red, green, and blue) input channels corresponding to the RGB color space of the input image. This three channel input is then converted to a 16 channel representation using a convolutional layer 210, followed by five blocks 212 of the following operators:

[0148] TIFF2026041721000019.tif5170

[0149] where CONV{.} is the convolution operator (including the bias term), k1 and k2 denote the serial numbers of the convolutional layers, LReLU[.] is the nonlinear activation function (i.e., leaky rectified linear unit) used throughout the network, and POOL(.) is a 2 × 2 average pooling step defined as follows:

[0150] TIFF2026041721000020.tif9170

[0083]

[0151] The number of input and output channels at each level follows the downsampling path of the generator followed by a center block convolutional layer 213. The final level 214 is expressed as:

[0152] TIFF2026041721000021.tif18170Here, FC[.] represents a fully connected layer with learnable weights and biases, Sigmoid(.) represents a sigmoid activation function, and Dropout[.] randomly removes 50% of the connections from the fully connected layer.

[0084]

[0153] The convolutional filter size of the entire GAN-based deep neural network was set to 3 × 3. The learnable parameters were 1 × 10 for the generator network (G). -4 and 1×10 for the discriminator network (N). -5 The training batch size was set to 48.

[0085]

[0154] Implementation details

[0155] The virtual dye network 10 was implemented using Python version 3.7.1 with the Pytorch framework version 1.3. The software was implemented on a desktop computer equipped with a 3.30 GHz Intel i9-7900X CPU and 64 GB of RAM running the Microsoft Windows 10 operating system. Network training and testing were performed using two NVIDIA GeForce GTX 1080Ti GPUs.

[0086]

[0156] Experiment - Post-imaging computational autofocus

[0157] This embodiment involves post-imaging computational autofocus for incoherent imaging modalities such as brightfield and fluorescence microscopy. This method uses a trained deep neural network to virtually refocus only one aberrant image. This data-driven machine learning algorithm takes an aberrant and / or out-of-focus image and outputs an image that is well-matched to a focused image of the same field of view. This method can be used to increase the scanning speed of microscopes imaging samples such as tissue.

[0087]

[0158] Fluorescence image acquisition

[0159] Referring to Figure 13, tissue autofluorescence images 20 were acquired using an inverted Olympus microscope (IX83, Olympus) 110 controlled by MicroManager microscope automation software. Unstained tissue 22 was excited in the near-ultraviolet range and imaged using a DAPI filter cube (OSF13-DAPI-5060C, excitation wavelength 377 nm / 50 nm bandwidth, emission wavelength 447 nm / 60 nm bandwidth). Images 20 were acquired using a 20x / 0.75 NA objective lens (Olympus UPLSAPO 20x / 0.75 NA, WD 0.65). At each stage position, the automation software performed autofocus based on image contrast, and z-stacks were acquired from -10 μm to 10 μm at 0.5 μm axial intervals. Each image was captured using a Scientific CMOS sensor (ORCA-flash 4.0 v.2, Hamamatsu Photonics) with an exposure time of ∼100 ms.

[0088]

[0160] Image preprocessing

[0161] To correct for the exact shift and rotation from the microscope stage, the autofluorescence image stack (2048 × 2048 × 41) was first aligned using the ImageJ plugin "StackReg." Extended depth-of-field images were then generated for each stack using the ImageJ plugin "Extended Depth of Field." The stack and the corresponding extended depth-of-field (EDOF) image were cropped into small 512 × 512 patches with no lateral overlap, and the most in-focus plane (target image) was set as the plane with the highest structural similarity index (SSIM) with the EDOF image. Ten planes above and below the focal plane (corresponding to a defocus of + / - 5 μm) were then selected to lie within the stack, and each of the 21 planes generated an input image for training Network 10a.

[0089]

[0162] To generate training and validation datasets, defocused and in-focus images of the same field of view were paired and used as the input and output, respectively, for training network 10a. The original dataset consisted of ∼30,000 such pairs and was randomly split into training and validation datasets taking 85% and 15% of the data. The training dataset was augmented eight times by random flipping and rotation during training, while the validation dataset was not augmented. The test dataset was cropped from a separate FOV that was not represented in the training and validation datasets. Images were normalized by their mean and standard deviation across the FOV before being fed to network 10a.

[0090]

[0163] Deep Neural Network Architecture, Training, and Validation

[0164] To perform snapshot autofocus, a generative adversarial network (GAN) 10a is used here. The GAN network 10a consists of a generator network (G) and a discriminator network (D). The generator network (G) is a U-net with residual connections, and the discriminator network (D) is a convolutional neural network. During training, the network 10a iteratively minimizes the loss functions of the generator and discriminator, which are defined as follows:

[0165] TIFF2026041721000022.tif18170

[0166] TIFF2026041721000023.tif7170

[0167] where z label indicates the in-focus fluorescence image, and z outputdenotes the generator output, and D is the discriminator output. The generator loss function is a combination of the mean absolute error (MAE), the multi-scale structural similarity (MS-SSIM) index and the mean squared error (MSE), balanced by the regularization parameters λ, β, and α. In training, the parameters are empirically set as λ=50, β=1, and α=1. The multi-scale structural similarity index (MS-SSIM) is defined as

[0168] TIFF2026041721000024.tif13170

[0169] where x j and y j is 2 j-1 is the distorted reference image downsampled times; μ x , μ y is the mean of x, y; TIFF2026041721000025.tif8162 is the variance of x; σ xy is the covariance of x, y; and C1, C2, C3 are small constants to stabilize fractions with small denominators.

[0091]

[0170] An adaptive moment estimation (Adam) optimizer is used, with a 1×10 -4 , 1 × 10 for the discriminator (D) -6 The learnable parameters are updated with a learning rate of . Additionally, at each iteration, 6 updates to the generator loss and 3 updates to the discriminator loss are performed. A batch size of 5 was used during training. The validation set is tested every 50 iterations, and the best model is selected as the model with the lowest validation set loss.

[0092]

[0171] Implementation details

[0172] The network is implemented using TensorFlow on a PC with a 2.3GHz Intel Xeon Core W-2195 CPU and 256GB RAM, using an Nvidia GeForce 2080Ti GPU. Training on ~30,000 image pairs of size 512x512 takes approximately ~30 hours. Test time on a 512x512 pixel image patch is ~0.2 seconds.

[0093]

[0173] Experiment - Virtual staining with multiple stains using a single network

[0174] In this embodiment, a class-conditional convolutional neural network 10 is used to transform input images consisting of one or more autofluorescence images 20 of unlabeled tissue samples 22. As an example, to demonstrate its utility, a single network 10 was used to virtually stain images of unlabeled sections with hematoxylin and eosin (H&E), Jones silver stain, Masson's trichrome, and periodic acid-Schiff (PAS) stain. The trained neural network 10 is able to stain specific tissue microstructures with these trained stains as well as generate novel stains.

[0094]

[0175] Data Acquisition

[0176] Unstained formalin-fixed, paraffin-embedded (FFPE) kidney tissue was cut into thin 2 μm slices and mounted on standard microscope slides. These tissue sections were obtained under IRB 18-001029. A conventional widefield fluorescence microscope (IX83, Olympus) with a 20× / 0.75 NA objective (Olympus UPLSAPO 20X / 0.75NA, WD 0.65) and two separate filter cubes, DAPI (OSFI3-DAPI-5060C, EX 377 / 50 nm EM 447 / 60 nm, Semrock) and TxRed (OSFI3-TXRED-4040C, EX 562 / 40 nm EM 624 / 40 nm, Semrock), was used to image the autofluorescence of unlabeled tissue sections. The exposure time for the DAPI channel was ~50 ms, and the exposure time for the TxRed channel was ~300 ms. Once the autofluorescence of the tissue sections was imaged, the slides were histologically stained using standard H&E, Jones, Masson's Trichrome, or PAS staining procedures. Staining was performed by the UCLA Translational Pathology Core Laboratory (TPCL). These histologically stained slides were then imaged using an FDA-approved slide-scanning microscope (Aperio AT, Leica Biosystems, scanned using a 20x / 0.75 NA objective) to create target images used for training, validation, and testing the neural network.

[0095]

[0177] Deep Neural Network Architecture, Training, and Validation

[0178] A conditional GAN ​​architecture was used to train deep neural network 10 to learn the transformation of unlabeled, unstained autofluorescent input images 20 to corresponding brightfield images 48 using four different stains (H&E, Masson's Trichrome, PAS, and Jones). Of course, other or additional stains can be trained on deep neural network 10. After co-registering autofluorescent images 20 to brightfield images 48, these precisely aligned FOVs were randomly divided into overlapping patches of 256x256 pixels, which were then used to train GAN network 10. In our implementation of conditional GAN ​​network 10, the training process uses a one-hot encoding matrix M (Figure 15) concatenated to the network's 256x256 input image / image stack patches, with each matrix M corresponding to a different stain. One way to represent this conditioning is given by:

[0179] TIFF2026041721000026.tif5170

[0180] where [·] represents concatenation and c i represents a 256 × 256 matrix for the labels of the ith stain type (in this example, H&E, Masson's Trichrome, PAS, and Jones). For an input image and target image pair from the ith stain dataset, c i is set to a matrix of all ones, while all remaining other matrices are assigned zero values ​​accordingly (see FIG. 15). The conditional GAN ​​network 10, as described herein, consists of two deep neural networks: a generator (G) and a discriminator (D). For this task, the loss functions of the generator and discriminator are defined as follows:

[0181]

[0182] TIFF2026041721000027.tif24170

[0183] Here, the anisotropic TV operator and the L1 norm are defined as follows:

[0184]

[0185] TIFF2026041721000028.tif16170

[0186] where D(·) and G(·) represent the discriminator network output and the generator network output, respectively, and z label shows bright-field images of histologically stained tissues, z output denotes the output of the generator network. P and Q represent the vertical and horizontal pixel counts of the image patch, and p and q represent the sum exponents. The regularization parameters (λ, α) are set to 0.02 and 2000, which corresponds to a total variation loss term of approximately 2% of the L1 loss and a discriminator loss term of 98% of the total generator loss.

[0096]

[0187] The deep neural network architecture of the generator (G) follows the structure of the deep neural network 10 shown in FIG. 10 (and described herein). However, in one implementation, the number of input channels at each level of the downsampling path was set to: 1, 96, 192, 384. The discriminator network (D) receives seven input channels: three (YCbCr color maps) originate from either the generator output or the target, and four originate from the one-hot coded class conditioning matrix M. A convolutional layer is used to convert this input into a 64-channel feature map. The convolutional filter size for the entire GAN was set to 3x3. The learnable parameters were adjusted through training using an adaptive moment estimation (Adam) optimizer, with a maximum of 1x10 for the generator network (G). -4 , and 2 × 10 for the discriminator network (D). -6 The learning rate was updated with . For each discriminator step, the generator network was iterated 10 times. The training batch size was set to 9.

[0097]

[0188] Single virtual tissue staining

[0189] Once the deep neural network is trained, it generates one-hot coded labels. TIFF2026041721000029.tif7170 is used to condition the network 10 to generate the desired stained image 40. In other words, c i The matrix M is set to an all-ones matrix, and the remaining matrices are set to all zeros for the ith stain (in the case of a single stain embodiment). Thus, one or more conditional matrices can be applied to the deep neural network 10 to generate respective stains for all or subregions of the imaged sample. The conditional matrix / matrices M define the subregions or boundaries of each stain channel.

[0098]

[0190] Dye mixing and microstructuring

[0191] Following the training process, the conditional matrix can be used in the way the network 10 was not trained to create new or novel types of staining. The encoding rules that must be met can be summarized in the following formula:

[0192] TIFF2026041721000030.tif7170

[0099]

[0193] In other words, for a given set of exponents (in our example N stains = 4), the total number of stains on which the network 10 is trained must equal 1. One possible implementation is by modifying the class encoding matrix to use a mixture of multiple classes, as described in the following equation:

[0194] TIFF2026041721000031.tif6170

[0195] Various stains can be mixed to create a unique stain with features emanating from the various stains learned by an artificial neural network, as shown in Figure 17.

[0100]

[0196] Another option is to divide the tissue field into different regions of interest (ROI), where every region of interest can be virtually stained with a specific stain or a mixture of these stains:

[0197] TIFF2026041721000032.tif7170

[0101]

[0198] Here, an ROI is a region of interest defined within the field of view. Multiple non-overlapping ROIs can be defined throughout the field of view. In one implementation, different stains can be used for different regions of interest or microstructures. These can be user-defined, manually marked as described herein, or algorithmically generated. As an example, a user can manually define various tissue areas via a GUI and stain them with different stains (FIGS. 16 and 17). This results in different tissue components being stained differently, as shown in FIGS. 16 and 17. Selective staining of ROIs (microstructuring) was implemented using the Python segmentation package Labelme. Using this package, an ROI logical mask was generated and the labels of the microstructures were added. TIFF2026041721000033.tif7170. In another implementation, tissue structures can be stained based on computer-generated mapping, e.g., obtained by segmentation software. For example, this can stain virtually all cell nuclei with one (or more) stains, while the rest of the tissue 22 is stained with another stain or combination of stains. Other manual, software, or hybrid approaches can be used to implement the selection of various tissue structures.

[0102]

[0199] FIG. 16 illustrates an example GUI for the display 106 that includes a toolbar 120 that can be used to select specific regions of interest within the network output image 40 for staining with a particular stain. The GUI may also include various stain options or a stain palette 122 that can be used to select a desired stain for the network output image 40 or a selected region of interest within the network output image 40. In this particular example, three areas are identified by hash lines (areas A, B, and C) that are manually selected by the user. These areas may also be automatically identified by the image processing software 104 in other embodiments. FIG. 16 illustrates a system that includes a computing device 100 that includes one or more processors 102 therein and image processing software 104 that incorporates a trained deep neural network 10.

[0103]

[0200] Implementation details

[0201] The virtual dye network was implemented using Python version 3.6.0 and TensorFlow framework version 1.11.0. The software can be implemented on any computing device 100. In the experiments described herein, the computing device 100 was a desktop computer equipped with a 2.30 GHz Intel Xeon W-2195 CPU, 256 GB of RAM, and running the Microsoft Windows 10 operating system. Network training and testing were performed using four NVIDIA GeForce RTX 2080 Ti GPUs.

[0104]

[0202] Extensive training in multiple styles

[0203] 20A to 20C show the dye conversion network 10 stainTN (FIG. 20C) shows another embodiment in which the training of the deep neural network that functions as a dye transfer network 10 is extended in an additional style. staniTNImages intended to be used as input for training of the stain transfer network 10 are normalized (or standardized). However, whole slide images generated using standard histological stains and scans will exhibit inevitable variations as a result of changing staining procedures and reagents between different laboratories and specific digital whole slide scanner characteristics, so extension with additional styles is required in certain situations. Therefore, this embodiment incorporates multiple styles during the training process to develop a stain transfer network 10 that can accommodate a wide range of input images. stainTN Generate.

[0105]

[0204] Therefore, network 10 stainTN To generalize the performance of the proposed method to this staining variability, network training was extended with additional staining styles. The styles represent the various image variability that can appear in chemically stained tissue samples. In this implementation, although other style transfer networks could be used, it was decided to facilitate this extension with K=8 unique style transfer (stain normalization) networks trained using the CycleGAN approach. CycleGAN is a GAN approach that uses two generators (G) and two discriminators (D). One pair is used to translate images from a first domain to a second domain. The other pair is used to translate images from the second domain to the first domain. This can be seen, for example, in Figure 21. More details regarding the CycleGAN approach are described, for example, in Zhu et al., "Unpaired Image-to-Image Translation using Cycle-Consistent Adversarial Networks," arXiv:1703.10593v1-v7 (2017-2020), which is incorporated herein by reference. These style extensions allow for a wide sample space for subsequent dye transformation networks. stainTN and therefore ensures that stain transfer is effective when applied to chemically H&E stained tissue samples, regardless of inter-technician, inter-laboratory or inter-instrument variability.

[0106]

[0205] Style Transfer Network 10 styleTN is used to generate a virtual dye generator / transformation network 10 for virtual-to-virtual dye transformation or chemical-to-virtual dye transformation. stainTN The latter is likely to be used in stain conversion networks given the variability found in chemical stains (e.g., H&E stains). For example, in industry, there is a need to convert one type of chemical stain to another. Examples include the need to create chemical H&E stains and specialized stains such as PAS, MT, or JMS. For example, non-neoplastic renal diseases rely on these "special stains" to provide pathological evaluation of standard treatment. In many clinical practices, H&E stains are available prior to the specialized stains, and the pathologist can provide a "preliminary diagnosis" to enable the patient's nephrologist to initiate treatment. This is particularly useful in some disease settings, such as crescentic glomerulonephritis or transplant rejection, where rapid diagnosis followed by rapid initiation or treatment can significantly improve clinical outcomes. In settings where only H&E slides are initially available, the preliminary diagnosis is followed by a final diagnosis, typically provided the next business day. As described herein, improved deep neural networks 10 stainTN was developed to improve preliminary diagnosis by using H&E stained slides to generate three additional special stains: PAS, MT, and Jones methenamine silver (JMS), which can be simultaneously reviewed by a pathologist with the histochemically stained H&E stain.

[0107]

[0206] A set of supervised deep learning-based workflows is presented, which allows users to develop dye transformation networks. stainTNThis allows us to perform the transformation between the two stains using a stain transformation network. This is achieved by first generating aligned pairs of a virtually stained H&E image and a specific stain of the same autofluorescent image of an unlabeled tissue section (Figure 20A). These generated images can then be used to transform the chemically stained image to the specific stain. This facilitates the creation of fully spatially aligned (paired) datasets and allows us to perform the transformation between the two stains using a stain transformation network without relying solely on distributional matching loss and unpaired data. stainTN Furthermore, no abnormalities due to misalignment are created, thus improving the accuracy of the transformation. This is verified by evaluating kidney tissues with various non-neoplastic diseases.

[0108]

[0207] Deep Neural Networks 10 stainTN was used to perform the transformation between H&E stained tissue and special stains. To train this neural network, a set of additional deep neural networks 10 were used in combination with each other. This workflow relies on the virtual staining's ability to generate images of multiple different stains using a single unlabeled tissue section (FIG. 20A). By using a single neural network 10 to generate both H&E images along with one of the special stains (PAS, MT, or JMS), a perfectly matched (i.e., paired) data set can be created. Normalization of the images generated using the virtual staining network 10 allows the stain transformation network 10 to perform the transformation. stainTN The virtually stained image 40 intended to be used as input when training the style transfer network 10 is augmented with additional staining styles to ensure generalization (FIG. 20B). This is because the virtually stained image 40 in FIG. 20B is styleTN to produce an augmented or style-transferred output image 40'. In other words, the color transfer network 10 stainTNis designed to be able to handle the inevitable H&E staining variations that are a result of staining procedures and reagents between different laboratories, as well as the characteristics of specific digital whole-slide scanners. This extension utilizes K=8 unique style transfer (stain normalization) networks10 trained using CycleGAN (Figure 20B). styleTN These extensions allow the wide H&E sample space to be covered by the following stain transformation network: stainTN This ensures that the results are covered by the NIRS and are therefore useful when applied to H&E stained tissue samples, regardless of inter-technician, inter-laboratory, or inter-instrument variability. stainTN is used to convert the chemical stains into one or more virtual stains, while the neural network 10 stainTN Note that the extended training of is performed using a variety of virtual dye inputs for dye style conversion.

[0109]

[0208] Using this dataset, we can generate a dye transformation network 10 using the scheme shown in Figure 20C. stainTN can be trained. Network 10 stainTN is an image patch from a virtually stained tissue 40 (top path) or a K=8 style transfer network 10 styleTN The network 10 is then randomly fed a virtually stained image passing through one of the CycleGAN-style paths (left path) and generates an output image 40'' with the transferred histochemical stain. The corresponding special stain (a virtual special stain from the same unlabeled field) is used as the ground truth 48, regardless of the CycleGAN-style transfer. stainTN was tested blindly (on patients / cases on which the network was not trained) on a variety of digitized H&E slides taken from the UCLA repository, representing a cohort of diseases and staining alterations.

[0110]

[0209] A method for dyeing style transfer for data augmentation

[0210] Neural Networks 10 stainTNTo ensure that

[10] is applicable to a wide variety of H&E stained tissue sections, a CycleGan model is used to generate a style transfer network

[10] . styleTN The training dataset was expanded by performing style transfer using this neural network 10 styleTN learns to map between two domains X and Y given training samples x and y, where X is the domain for the original virtually stained H&E (image 40 in Figure 21). x ) and Y is the number of images produced by H&E from different laboratories or hospitals. y The domain for this model is G:X → Y and F:Y → X. Furthermore, two adversarial discriminators D x and D y A diagram showing the relationship between these various networks is shown in Figure 21. As can be seen in Figure 21, the virtual H&E image 40 x (Style X) is input to generator G(x), and a generated H&E image 40 with style Y is generated. y Then, create image 40 x , which is then mapped back to the X domain using a second generator F(G(x)) to generate . Referring to FIG. 21, the histochemically stained H&E image 40 y (Style Y) is input to generator F(y), and the generated H&E image 48 with style X is generated. x Then, create image 48 y The second generator G(F(y)) is used to generate the second generator G(x) and the second generator F(y) are used to generate the second generator G(x) and the second generator F(y) are used to generate the second generator G(x) and the second generator F(y) are used to generate the second generator F(y). Figures 22A and 22B show the configuration of two generator networks G(x) (Figure 22A) and F(y) (Figure 22B).

[0111]

[0211] The generator loss function l generator It has two terms: an adversarial loss l to match the staining style of the generated image to the style of the histochemically stained image in the target domain; advand a cycle consistency loss l to prevent the learned mappings G and F from contradicting each other. cycle Therefore, the overall loss can be described by the following formula:

[0212] TIFF2026041721000034.tif5170

[0213] where λ and φ are constants used to weight the loss function. For all networks, λ was set to 10 and φ was set equal to 1. Each generator in Figure 21 uses a discriminator (D x or D y ) The losses of each of the generator networks can be written as:

[0214] TIFF2026041721000035.tif9170

[0215] TIFF2026041721000036.tif9170

[0216] Then, the cycle consistency loss can be written as:

[0217] TIFF2026041721000037.tif6170

[0218] where the L1 loss, i.e., mean absolute error, is given by:

[0219] TIFF2026041721000038.tif8170

[0220] In this formula, p and q are pixel indices, and P and Q are the total number of pixels in the horizontal direction.

[0112]

[0221] D x and D y The adversarial loss term used in training is defined as:

[0222] TIFF2026041721000039.tif7170

[0223] TIFF2026041721000040.tif7170

[0113]

[0224] In these CycleGAN models, G and F use a U-net architecture. This architecture consists of three "down blocks" 220 followed by three "up blocks" 222. Each of the down blocks 220 and up blocks 222 contains three convolutional layers with a 3x3 kernel size activated by a leaky ReLU activation function. Each down block 220 doubles the number of channels and ends with an average pooling layer with a stride and kernel size of 2. The up blocks 222 begin with bicubic upsampling before applying the convolutional layers. Between each of the blocks in a particular layer, skip connections are used to pass data through the network without passing through every block.

[0114]

[0225] Discriminator D X and D Y consists of four blocks. These blocks contain two convolutional layers paired with leaky ReLUs, which together double the number of channels. These are followed by an average pooling layer with stride 2. After five blocks, two fully connected layers reduce the output dimensionality to a single value.

[0115]

[0226] During training, an adaptive moment estimation (Adam) optimizer was used to optimize both the generator (G) and discriminator (D) networks using 2 × 10 -5 The learnable parameters were updated with a learning rate of . At each step of discriminator training, one training iteration was performed on the generator network, and the batch size for training was set to 6.

[0116]

[0227] Style Transfer Network 10 styleTN In some embodiments, this may include normalization of the staining vectors. stainTNcan be trained in a supervised manner between two different variations of the same stained slide, where the variations result from imaging the sample with a different microscope or from histochemical staining followed by a second re-staining of the same slide. staniTN can be trained with a set of pairs of virtually stained images, each stained with a different virtual stain, generated by a single virtual stain neural network 10.

[0117]

[0228] While embodiments of the present invention have been shown and described, various modifications may be made without departing from the scope of the present invention. For example, while various embodiments have been described as generating digitally / virtually stained microscopic images of unlabeled or unstained samples, the methods may also be used when the samples are labeled with one or more exogenous fluorescent labels or other exogenous illuminants. Thus, these samples are labeled without traditional immunohistochemistry (IHC) staining. Accordingly, the present invention should not be limited, except as set forth in the following claims and their equivalents.

Claims

1. 1. A method for generating a virtually stained microscopic image of a sample, comprising: providing, using one or more processors of a computing device, a trained deep neural network executed by image processing software, the trained deep neural network being trained using a plurality of matched immunohistochemistry (IHC) stained microscopic images or image patches and their corresponding fluorescence lifetime (FLIM) microscopic images or image patches of the same samples obtained before immunohistochemistry (IHC) staining; obtaining a fluorescence lifetime (FLIM) image of the sample using a fluorescence microscope and at least one excitation light source; inputting the fluorescence lifetime (FLIM) images of the sample into the trained deep neural network; and the trained deep neural network outputs the virtually stained microscopic image of the sample that is substantially equivalent to a corresponding image of the same sample that has been immunohistochemistry (IHC) stained.

2. 10. The method of claim 1, wherein the sample is unlabeled or unstained, and the fluorescence emitted from the sample is emitted from endogenous fluorophores or endogenous emitters of light within the sample.

3. 10. The method of claim 1, wherein the sample is first labeled with one or more exogenous fluorescent labels or other exogenous emitters of light.

4. The method of any one of claims 1 to 3, wherein the trained deep neural network comprises a convolutional neural network.

5. The method of any one of claims 1 to 3, wherein the deep neural network is trained using a generative adversarial network (GAN) model.

6. 4. The method of claim 1, wherein the deep neural network is trained using a generator network and a discriminator network, the generator network being configured to learn a statistical transformation between the fluorescence lifetime (FLIM) images or image patches of the sample acquired during the IHC staining and the matched immunohistochemistry (IHC) stained images of the same sample, and the discriminator network being configured to digitally distinguish between ground truth chemically stained IHC images of the sample and the output virtually stained microscopy images of the same sample.

7. 4. The method of any one of claims 1 to 3, wherein the sample comprises mammalian tissue, plant tissue, cells, pathogens, bacteria, parasites, fungi, body fluid smear, liquid biopsy or other matter of interest.

8. 4. The method of claim 1, wherein the deep neural network is trained using a sample of the same type as the type of the sample to be virtually stained by the trained deep neural network.

9. 4. The method of claim 1, wherein the deep neural network is trained using a different type of sample compared to the type of sample to be virtually stained by the trained deep neural network.

10. The method of any one of claims 1 to 3, wherein the sample comprises an unfixed tissue sample.

11. The method of any one of claims 1 to 3, wherein the sample comprises a fixed tissue sample.

12. The method of claim 11 , wherein the fixed tissue sample is embedded in paraffin.

13. The method of any one of claims 1 to 3, wherein the sample comprises a fresh tissue sample or a frozen tissue sample.

14. The method of any one of claims 1 to 3, wherein the sample comprises tissue imaged in vivo or in vitro.

15. The method according to any one of claims 1 to 3, wherein at least one excitation light source emits ultraviolet or near ultraviolet radiation.

16. 4. The method of claim 1, wherein the fluorescence lifetime (FLIM) images are obtained in a filtered emission band or in a set of emission bands using at least one filter, and the obtained images are input into the trained deep neural network for virtual staining of the sample.

17. 17. The method of claim 16, wherein one or more filters are used to capture one or more fluorescence lifetime (FLIM) images, and the resulting images are input into the trained deep neural network for virtual staining of the sample.

18. 17. The method of claim 16, wherein the one or more fluorescence lifetime (FLIM) images are obtained with one or more excitation light sources emitting light at different wavelengths or wavelength bands, and the one or more obtained fluorescence lifetime (FLIM) images are input into the trained deep neural network for virtual staining of the sample.

19. 4. The method of claim 1, wherein the fluorescence lifetime (FLIM) images are subjected to one or more linear or non-linear image pre-processing steps including one or more of contrast enhancement, contrast inversion, image filtering before being input to the trained deep neural network.

20. 4. The method of claim 1, wherein the fluorescence lifetime (FLIM) image and one or more pre-processed images are input together into the trained deep neural network for virtual staining of the sample.

21. 4. The method of claim 1, wherein the plurality of matched immunohistochemistry (IHC) stained sample images and fluorescence lifetime (FLIM) images or image patches of the same samples obtained before immunohistochemistry (IHC) staining are subjected to registration during training, the registration comprising at least one global registration process that corrects for rotation and a local registration process that matches and aligns to local features found in the matched chemically stained sample images and the fluorescence lifetime (FLIM) images.

22. The method of any one of claims 1 to 3, wherein the trained deep neural network is trained using one or more GPUs or ASICs.

23. The method of any one of claims 1 to 3, wherein the trained deep neural network is executed using one or more GPUs or ASICs.

24. 4. The method of any one of claims 1 to 3, wherein the virtually stained microscopic image of the sample is output in real time or near real time after obtaining a single or set of fluorescence lifetime (FLIM) images of the sample.

25. 4. The method of any one of claims 1 to 3, wherein the trained deep neural network is trained for a second tissue / stain combination using weights and biases of an initial neural network from a first tissue / stain combination that are further optimized for the second tissue / stain combination using transfer learning or additional algorithms.

26. 4. The method of claim 1, wherein the trained deep neural network is trained for virtual staining of multiple tissue / stain combinations.

27. 4. The method of claim 1, wherein the trained deep neural network is trained for multiple immunohistochemistry (IHC) stain types for a given tissue type.

28. 1. A method for virtually autofocusing a microscopic image of a sample obtained using an incoherent microscope, comprising: providing a trained deep neural network executed by image processing software using one or more processors of a computing device, the trained deep neural network being trained using a plurality of pairs of out-of-focus and / or in-focus microscopic images or image patches used as input images to the deep neural network and corresponding or matched in-focus microscopic images or image patches of the same sample obtained using the incoherent microscope used as ground truth images for training the deep neural network; obtaining an out-of-focus or in-focus image of the sample using the incoherent microscope; inputting into the trained deep neural network out-of-focus or in-focus images of the sample obtained from the incoherent microscope; and the trained deep neural network outputs an output image with improved focus that substantially matches the in-focus image (ground truth) of the same sample acquired by the incoherent microscope.

29. 29. The method of claim 28, wherein the sample comprises a three-dimensional sample, the out-of-focus images of the sample comprise images including out-of-focus features and / or clutter, and the trained deep neural network outputs output images at different planes or depths within the three-dimensional sample that reject out-of-focus features and / or clutter.

30. 30. The method of claim 28 or 29, wherein the network output image comprises an extended depth of field (EDOF) image.

31. 30. The method of claim 28 or 29, wherein the incoherent microscope comprises one of a fluorescence microscope, a super-resolution microscope, a confocal microscope, a light sheet microscope, a FLIM microscope, a bright field microscope, a dark field microscope, a structured illumination microscope, a total internal reflection microscope, a computational microscope, a ptychographic microscope, a synthetic aperture-based microscope, and a phase contrast microscope.

32. 30. The method of claim 28 or 29, further comprising inputting the output image of the trained deep neural network with improved focus to a separate trained deep neural network executed by image processing software using one or more processors of a computing device to virtually stain the image with improved focus, wherein the separate trained deep neural network is trained using a plurality of matched chemically stained images or image patches used as ground truth images or image patches and their corresponding images or image patches of the same sample obtained using the incoherent microscope before histochemical staining of the sample, and wherein the trained deep neural network outputs a virtually stained microscopic image of the sample that is substantially equivalent to a corresponding microscopic image of the same sample that has been histochemically stained.

33. 1. A method for generating a virtually stained microscopic image of a sample using an incoherent microscope, comprising: providing, using one or more processors of a computing device, a trained deep neural network executed by image processing software, wherein the trained deep neural network is trained with a plurality of pairs of out-of-focus and / or in-focus microscopic images or image patches that all match corresponding in-focus microscopic images or image patches of the same sample obtained using the incoherent microscope after a chemical staining process, which are used as input images to the deep neural network and generate ground truth images for training the deep neural network; obtaining an out-of-focus or in-focus image of the sample using the incoherent microscope; inputting the out-of-focus or in-focus images of the sample obtained from the incoherent microscope into the trained deep neural network; and the trained deep neural network outputs an output image of the virtually stained sample having improved focus and substantially resembling and matching a chemically stained and in-focus image of the same sample obtained by the incoherent microscope after the chemical staining process.

34. 34. The method of claim 33, wherein the sample is unlabeled or unstained.

35. 34. The method of claim 33, wherein the sample is first labeled with one or more exogenous labels and / or light emitters.

36. 1. A method for generating a virtually stained microscopic image of a sample, comprising: providing a trained deep neural network executed by image processing software using one or more processors of a computing device, the trained deep neural network being trained with a plurality of pairs of stained microscopic images or image patches that are virtually stained or histochemically stained to have a first stain type by at least one algorithm, the pairs of microscopic images or image patches all being matched to corresponding stained microscopic images or image patches of the same sample that are virtually stained or chemically stained to have another, different stain type by at least one algorithm, constituting ground truth images for training of the deep neural network to convert input images that are histochemically or virtually stained using the first stain type to output images that are virtually stained using the second stain type; obtaining a histochemically or virtually stained input image of the sample stained with the first stain type; inputting histochemically or virtually stained images of the sample into the trained deep neural network, which converts input images stained with the first stain type into output images virtually stained with the second stain type; and the trained deep neural network outputs an output image of the sample having virtual staining that substantially resembles and matches a chemically stained image of the same sample stained with the second stain type obtained by incoherent microscopy after the chemical staining process.

37. 37. The method of claim 36, wherein the input stained image comprises a virtually stained image generated by the output of a separate trained neural network or at least one digital or machine learning algorithm.

38. 37. The method of claim 36, wherein the stained input image comprises a stained image of a histochemically stained sample acquired by an incoherent microscope.

39. 37. The method of claim 36, wherein the trained deep neural network is trained using augmented training images that have been augmented with a style transfer network or algorithm.

40. 40. The method of claim 39, wherein the style transfer network comprises a CycleGAN-based style transfer network.

41. 40. The method of claim 39, wherein the style transfer network includes stain vector normalization.

42. 40. The method of claim 39, wherein the trained deep neural network is trained in a supervised manner between two different variations of the same stained slide, the variations being the result of imaging the sample with a different microscopic or histochemical stain followed by a second re-staining of the same slide.

43. 40. The method of claim 39, wherein the trained deep neural network is trained with a set of paired virtually stained images, each stained with a different virtual stain generated by a single virtual stain neural network or by two different virtual stain neural networks.

44. 41. The method of any one of claims 28, 29, and 36 to 40, wherein the chemical stain is one of hematoxylin and eosin (H&E) stain, hematoxylin, eosin, Jones silver stain, Masson's trichrome stain, Periodic Acid Schiff (PAS) stain, Congo red stain, Alcian blue stain, blue iron, silver nitrate, trichrome stain, Ziehl-Nielsen, Grocott methenamine silver (GMS) stain, Gram stain, acid stain, basic stain, silver stain, Nissl, Weigert stain, Golgi stain, Luxol fast blue stain, toluidine blue, Genta, Mallory's trichrome stain, Gomori's trichrome, Van Gieson, Giemsa, Sudan black, Perls-Prussian, Best carmine, acridine orange, immunofluorescence stain, immunohistochemical stain, Kinyon-Cold stain, Albert's stain, flagella stain, endospore stain, Nigrosine, or India ink stain.

45. 34. The method of claim 33, wherein the incoherent microscope used to input the input image into the trained deep neural network comprises one of a fluorescence microscope, a widefield microscope, a super-resolution microscope, a confocal microscope, a confocal microscope with single-photon or multi-photon excited fluorescence, a second-harmonic or high-harmonic generation fluorescence microscope, a light-sheet microscope, a FLIM microscope, a bright-field microscope, a dark-field microscope, a structured illumination microscope, a total internal reflection microscope, a computational microscope, a ptychographic microscope, a synthetic aperture-based microscope, or a phase-contrast microscope.

46. 34. The method of claim 33, wherein the out-of-focus image of the sample obtained using the incoherent microscope is out of focus within an axial range about the in-focus image plane.

47. 1. A method for generating virtually stained microscopy images of a sample with a plurality of different stains using a single trained deep neural network, comprising: providing, using one or more processors of a computing device, a trained deep neural network executed by image processing software, the trained deep neural network being trained with a plurality of matched chemically stained microscopic images or image patches using a plurality of chemical stains, which are used as ground truth images for training the deep neural network, and their corresponding matched fluorescent microscopic images or image patches of the same sample obtained before chemical staining, which are used as input images for training the deep neural network; obtaining a fluorescence image of the sample using a fluorescence microscope and at least one excitation light source; applying one or more class-conditional matrices to condition the trained deep neural network; inputting the fluorescence image of the sample along with the one or more class conditional matrices into the trained deep neural network; and a step in which the trained deep neural network outputs the virtually stained microscopic image of the sample having one or more different stains, wherein the output image or subregion thereof is substantially equivalent to a corresponding microscopic image or image subregion of the same sample histochemically stained with the corresponding one or more different stains.

48. 48. The method of claim 47, wherein the sample is unlabeled or unstained, and the fluorescence emitted from the sample is initially labeled using endogenous fluorophores or other endogenous emitters of light within the sample.

49. 48. The method of claim 47, wherein the sample is first labeled with one or more exogenous fluorescent labels or other exogenous emitters of light.

50. 50. The method of any one of claims 47 to 49, wherein the fluorescence microscope comprises one of a super-resolution microscope, a confocal microscope, a light sheet microscope, a FLIM microscope, a wide-field microscope, a structured illumination microscope, a computed microscope, a ptychographic microscope, a synthetic aperture-based microscope, or a total internal reflection microscope.

51. 50. The method of any one of claims 28, 33, 36 and 47-49, wherein the fluorescent images or image patches of the sample obtained before staining and used to train a deep neural network and / or the fluorescent images of the sample obtained using the fluorescent microscope used as input to the trained neural network are out of focus, and the trained deep neural network learns and performs both autofocus and virtual staining.

52. 50. The method of any one of claims 28, 33, 36 and 47 to 49, wherein the fluorescence images or image patches of the sample obtained before staining and used to train a deep neural network and / or the fluorescence images of the sample obtained using the fluorescence microscope used as input to the trained neural network are in focus or computationally refocused or autofocused.

53. 50. The method of any one of claims 28, 33, 36 and 47-49, wherein the input image has the same or substantially similar numerical aperture and resolution as the ground truth image.

54. 50. The method of any one of claims 28, 33, 36 and 47-49, wherein the input image has a lower numerical aperture and a lower resolution compared to the ground truth image.

55. 50. The method of any one of claims 47 to 49, wherein the input fluorescence image of a new test sample is obtained by initially focusing the microscope using at least one fluorescence image obtained using a different emission filter.

56. 50. The method of any one of claims 47 to 49, wherein a plurality of conditional matrices are applied to condition the trained deep neural network.

57. 50. The method of any one of claims 47 to 49, wherein the plurality of conditional matrices correspond to different stainings representing spatially non-overlapping subregions.

58. 50. The method of any one of claims 47 to 49, wherein the plurality of conditional matrices correspond to different stainings and have at least some spatial overlap.

59. 50. The method of any one of claims 47 to 49, wherein the spatial boundaries of the one or more class-conditional matrices are manually defined.

60. 50. The method of any one of claims 47 to 49, wherein the spatial boundaries of the one or more class-conditional matrices are defined automatically using image processing software or another automated algorithm.

61. 50. A method according to any one of claims 47 to 49, wherein the one or more class-conditional matrices are applied to a field of view of a particular sample acquired by the microscope.

62. 50. The method of any one of claims 47 to 49, wherein the plurality of stains comprises at least two of hematoxylin and eosin (H&E) stain, hematoxylin, eosin, Jones silver stain, Masson's trichrome stain, Periodic Acid Schiff (PAS) stain, Congo red stain, Alcian blue stain, blue iron, silver nitrate, trichrome stain, Ziehl-Nielsen, Grocott methenamine silver (GMS) stain, Gram stain, silver stain, acid stain, basic stain, Nissl, Weigert stain, Golgi stain, Luxol fast blue stain, toluidine blue, Genta, Mallory's trichrome stain, Gomori's trichrome, Van Gieson, Giemsa, Sudan black, Perls-Prussian, Best carmine, acridine orange, immunofluorescence stain, immunohistochemical stain, Kinyon-Cold stain, Albert's stain, flagella stain, endospore stain, nigrosine, and India ink stain.

63. 50. The method of any one of claims 47 to 49, wherein the sample comprises mammalian tissue, plant tissue, cells, pathogens, bacteria, parasites, fungi, body fluid smear, liquid biopsy or other matter of interest.

64. 1. A method for generating virtually stained microscopy images of a sample with a plurality of different stains using a single trained deep neural network, comprising: providing, using one or more processors of a computing device, a trained deep neural network executed by image processing software, the trained deep neural network being trained using a plurality of matched chemically stained microscopic images or image patches using a plurality of chemical stains and their corresponding microscopic images or image patches of the same sample obtained before chemical staining; obtaining an input image of the sample using a microscope; applying one or more class-conditional matrices to condition the trained deep neural network; inputting the sample input images into the trained deep neural network along with the one or more class conditional matrices; and a step in which the trained and conditioned deep neural network outputs the virtually stained microscopic image of the sample having one or more different stains, wherein the output image or subregion thereof is substantially equivalent to a corresponding microscopic image or image subregion of the same sample histochemically stained with the corresponding one or more different stains.

65. 65. The method of claim 64, wherein the sample is unlabeled or unstained.

66. 65. The method of claim 64, wherein the sample is first labeled with one or more exogenous labels and / or stains.

67. 67. The method of any one of claims 64 to 66, wherein the images or image patches of the sample obtained before staining and used for training a deep neural network and / or the images of the sample obtained using the microscope and used as input to the trained neural network are out of focus, and the trained deep neural network learns and performs both autofocus and virtual staining.

68. 67. The method of any one of claims 64 to 66, wherein the images or image patches of the sample obtained before staining and used for training a deep neural network and / or the images of the sample obtained using the microscope and used as input to the trained neural network are in focus or are computationally autofocused or refocused.

69. 67. The method of any one of claims 64 to 66, wherein a plurality of conditioning matrices are applied to condition the trained deep neural network.

70. A method according to any one of claims 64 to 66, wherein the input image has the same or substantially similar numerical aperture and resolution as the ground truth image.

71. A method according to any one of claims 64 to 66, wherein the input image has a lower numerical aperture and a lower resolution compared to the ground truth image.

72. 70. The method of claim 69, wherein the plurality of conditional matrices correspond to different stains representing spatially non-overlapping subregions.

73. 70. The method of claim 69, wherein the plurality of conditional matrices correspond to different stainings and have at least some spatial overlap.

74. 67. The method of any one of claims 64 to 66, wherein the spatial boundaries of the one or more class-conditional matrices are manually defined.

75. 67. A method according to any one of claims 64 to 66, wherein the spatial boundaries of the one or more class-conditional matrices are defined automatically using image processing software or another automated algorithm.

76. 67. The method of any one of claims 64 to 66, wherein the one or more class-conditional matrices are applied to a field of view of a particular sample captured by the microscope.

77. 67. The method of any one of claims 64 to 66, wherein the plurality of stains comprises at least two of hematoxylin and eosin (H&E) stain, hematoxylin, eosin, Jones silver stain, Masson's trichrome stain, Periodic Acid Schiff (PAS) stain, Congo red stain, Alcian blue stain, blue iron, silver nitrate, trichrome stain, Ziehl-Nielsen, Grocott methenamine silver (GMS) stain, Gram stain, acid stain, basic stain, silver stain, Nissl, Weigert stain, Golgi stain, Luxol fast blue stain, toluidine blue, Genta, Mallory's trichrome stain, Gomori's trichrome, Van Gieson, Giemsa, Sudan black, Perls-Prussian, Best carmine, acridine orange, immunofluorescence stain, immunohistochemical stain, Kinyon-Cold stain, Albert's stain, flagella stain, endospore stain, nigrosine, and India ink stain.

78. 67. The method of any one of claims 64 to 66, wherein the sample comprises mammalian tissue, plant tissue, cells, pathogens, bacteria, parasites, fungi, body fluid smear, liquid biopsy or other matter of interest.

79. 67. The method of any one of claims 64 to 66, wherein the input image input to the trained neural network is obtained using a microscope selected from the group comprising: a single-photon fluorescence microscope, a multi-photon microscope, a second-harmonic generation microscope, a high-order harmonic generation microscope, an optical coherence tomography (OCT) microscope, a confocal reflectance microscope, a fluorescence lifetime microscope, a Raman spectroscopy microscope, a bright-field microscope, a dark-field microscope, a phase contrast microscope, a quantitative phase microscope, a structured illumination microscope, a super-resolution microscope, a light-sheet microscope, a computational microscope, a ptychographic microscope, a synthetic aperture-based microscope, and a total internal reflection microscope.

80. 65. The method of claim 64, wherein the input image comprises an image of a sample containing stained tissue and / or cells / cellular tissue.

81. 65. The method of claim 64, wherein the input image comprises an image of a sample containing unstained tissue or cells / cellular tissue.

82. 1. A method for generating a virtually destained microscopic image of a sample, comprising: providing, using one or more processors of a computing device, a trained first deep neural network executed by image processing software, said first trained deep neural network being trained using a plurality of matched chemically stained microscopic images or image patches used as training inputs to said deep neural network and their corresponding unstained microscopic images or image patches of the same sample or samples obtained before chemical staining, which constitute ground truth during training of said deep neural network; obtaining a microscopic image of the chemically stained sample using a microscope; inputting an image of the chemically stained sample into the trained first deep neural network; and the trained first deep neural network outputs the virtually destained microscopic image of the sample that is substantially equivalent to a corresponding image of the same sample obtained before or without chemical staining.

83. 83. The method of claim 82, further comprising inputting the virtually destained microscopic image output from the trained first deep neural network into a second trained deep neural network trained to transform the virtually destained microscopic image into a second output image that substantially matches the desired microscopic image of the same sample after being chemically stained with a different type of chemical stain.

84. The microscopic images of the same sample chemically stained with different chemical stains were analyzed using the following histological stains: hematoxylin and eosin (H&E), hematoxylin, eosin, Jones silver, Masson's trichrome, Periodic Acid Schiff (PAS), Congo red, Alcian blue, iron blue, silver nitrate, trichrome, Ziehl-Nielsen, Grocott methenamine silver (GMS), Gram, acid stain, basic stain, silver stain, Nissl. , Weigert stain, Golgi stain, Luxol Fast Blue stain, Toluidine Blue, Genta, Mallory Trichrome stain, Gomori Trichrome, Van Gieson stain, Giemsa stain, Sudan Black, Perls Prussian, Best Carmine, Acridine Orange, immunofluorescence stain, immunohistochemical stain, Kinyon-Cold stain, Albert stain, flagella stain, endospore stain, Nigrosine or India Ink stain.

85. 85. A method according to any one of claims 82 to 84, wherein the images or image patches of the sample obtained before virtual destaining and used to train the first deep neural network and / or the input images of the sample obtained using a microscope and input to the trained first neural network are out of focus, and the trained first deep neural network learns and performs both autofocus and virtual destaining.

86. 85. A method according to any one of claims 82 to 84, wherein the images or image patches of the sample obtained before virtual destaining and used to train the first deep neural network and / or the input images of the sample obtained using the microscope and input to the trained first deep neural network are in focus or computationally autofocused or refocused.

87. 85. The method of any one of claims 82 to 84, wherein the trained first deep neural network is further configured to output a virtually destained microscopic image of the sample with improved focus and extended depth of field.